Human-computer interaction user action intention recognition method and device, equipment and medium

Through the combination of K-Means clustering, TPT model, CNN-LSTM-Attention and Monte Carlo simulation algorithm, gesture action intentions are identified, and the cumbersome and complexity of traditional remote control technology in multi-device collaborative control is solved, and an intelligent and personalized user experience is achieved.

CN120295475APending Publication Date: 2025-07-11GUANGDONG UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510438061.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional remote control technology is complicated and complex when it is coordinated to control multiple devices, lacks immersion and fun, making it difficult to achieve an intelligent and personalized user experience.

Method used

The K-Means clustering algorithm and the TPT model are combined, combined with the CNN-LSTM-Attention model and the Monte Carlo simulation algorithm, and the action segments are segmented, the user's action intention is identified and the control instructions are generated.

Benefits of technology

It realizes efficient learning without manual labeling of data, reduces the cost of preliminary preparation, improves the accuracy and robustness of data processing, and provides a smooth, personalized and immersive control experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295475A_ABST
    Figure CN120295475A_ABST
Patent Text Reader

Abstract

The invention relates to a man-machine interaction user action intention recognition method, device and equipment and a medium, and the method comprises the steps: for continuous wave trough values of a wave trough value cluster, according to a first time index corresponding to a wave peak value in a wave peak value cluster and a second time index corresponding to the wave trough value in the wave trough value cluster, calculating the continuous wave trough values of the wave trough value cluster; finding out a middle crest value between the continuous trough values to segment a plurality of action segments in the gesture action; inputting the to-be-recognized gesture action into the gesture action recognition model to determine probability distribution corresponding to the gesture action category of the to-be-recognized gesture action; and adopting a preset Monte Carlo simulation algorithm to perform multiple random sampling on the probability distribution corresponding to the gesture action category, and taking the gesture action category with the maximum probability as the user action intention of the gesture action to be recognized so as to complete the recognition of the user action intention of human-computer interaction. According to the method and the device, smoother, more personalized and more immersive control experience can be provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent control, and in particular to a method for recognizing user action intentions in human-computer interaction, a corresponding device, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the development of the Internet of Things and human-computer interaction technology, users have put forward higher requirements for the naturalness, flexibility and immersion of the interaction method. Traditional control methods (such as remote control, mobile phone APP, voice assistant, etc.) have certain limitations. For example, the control object of traditional remote control technology is often limited to a single device. When it is necessary to achieve coordinated control of multiple devices, its limitations will be highlighted. Since each device may need to be equipped with a dedicated remote control, this not only leads to increased cumbersomeness in operation, but also significantly increases the complexity and management difficulty of the entire control system. In addition, the existence of multiple remote controls can easily cause confusion and inconvenience to users during use, thereby reducing the overall control efficiency and user experience. It is particularly important to note that with the people's growing needs for a better life, this traditional control method often lacks immersion and fun, and cannot meet the needs of modern users for intelligent and entertaining control experience.

[0003] Therefore, in the existing technical system, we are facing an obvious lack, that is, the lack of an efficient and intelligent wand control solution in multiple scenarios (smart home, theme park, etc.). In particular, there are still significant technical challenges and difficulties in how to use machine learning models to deeply optimize the control experience. At present, how to design and implement a machine learning model that can intelligently identify user intentions, autonomously learn, and efficiently detect and adjust control strategies, so as to provide users with a smoother and more personalized control experience, has become a key issue that needs to be solved urgently.

[0004] In summary, when traditional remote control technology needs to realize the coordinated control of multiple devices, its limitations will be highlighted. Since each device may need to be equipped with a dedicated remote control, this not only increases the complexity of operation, but also significantly increases the complexity and management difficulty of the entire control system. This application makes corresponding explorations to solve this problem. Summary of the invention

[0005] The purpose of the present application is to solve the above-mentioned problems and to provide a method for identifying user action intentions in human-computer interaction, a corresponding device, an electronic device and a computer-readable storage medium.

[0006] In order to meet the various objectives of this application, this application adopts the following technical solutions:

[0007] A method for recognizing user action intentions in human-computer interaction proposed to meet one of the purposes of the present application, including:

[0008] Responding to an instruction for recognizing action intentions of human-computer interaction actions, and obtaining first sensor timing signals corresponding to multiple gesture actions;

[0009] Using a preset K-Means clustering algorithm to cluster the first sensor timing signals to determine multiple clustering clusters, and determining a peak value cluster and a valley value cluster in each clustering cluster through a time sliding window based on a preset TPT model, where the clustering cluster represents a local action pattern in the gesture action;

[0010] For consecutive valley values in the valley value cluster, according to the first time index corresponding to the peak value in the peak value cluster and the second time index corresponding to the valley value in the valley value cluster, find the intermediate peak value between the consecutive valley values to segment multiple action segments in the gesture action;

[0011] Inputting the gesture action to be recognized into a gesture action recognition model trained by using the first sensor timing signal and a second sensor timing signal corresponding to the multiple action segments to determine a probability distribution corresponding to the gesture action category of the gesture action to be recognized;

[0012] Using a preset Monte Carlo simulation algorithm to perform multiple random samplings on the probability distribution corresponding to the gesture action category, and taking the gesture action category with the maximum probability as the user action intention of the gesture action to be recognized, so as to complete the recognition of the user action intention in human-computer interaction.

[0013] Optionally, the step of finding the intermediate peak value between the consecutive valley values according to the first time index corresponding to the peak value in the peak value cluster and the second time index corresponding to the valley value in the valley value cluster to segment multiple action segments in the gesture action includes:

[0014] Initializing an empty list to store the extracted action segments;

[0015] Traversing each pair of consecutive valley values in the second time index, where each pair of consecutive valley values represents the start point and the end point of a local action pattern;

[0016] Between each pair of consecutive valley values, check whether there is a peak value in the first time index within this pair of consecutive valley values. If so, it means that a valid action segment has occurred in this interval, and then add the action segment to the empty list;

[0017] The above steps are executed repeatedly until all the action segments in the gesture action are segmented.

[0018] Optionally, clustering the first sensor time series signal using a preset K-Means clustering algorithm to determine a plurality of clustering clusters includes:

[0019] Acquire multiple data points in the first sensor signal, and minimize the sum of squares of distances from each data point in the first sensor signal to the center of the cluster to which it belongs according to the objective function of the K-Means clustering algorithm;

[0020] The objective function of the K-Means clustering algorithm is iteratively optimized, and the cluster centers and the allocation relationships are updated. During the iterative optimization process, data points are continuously allocated to the nearest cluster centers, and the cluster centers are updated to determine multiple clusters.

[0021] Optionally, the step of using a preset Monte Carlo simulation algorithm to perform multiple random samplings on the probability distribution corresponding to the gesture action category, and taking the gesture action category with the maximum probability as the user action intention of the gesture action includes:

[0022] Based on the output data of the gesture recognition model, the gesture action category and its corresponding probability distribution of each gesture action are obtained, and the probability distribution corresponding to the gesture action category is randomly sampled multiple times using a preset Monte Carlo simulation algorithm, and the gesture action category with the maximum probability is used as the user action intention of the gesture action, wherein the expression of the Monte Carlo simulation algorithm predicting the probability distribution is:

[0023]

[0024] Among them, x is the input data, f represents the gesture action recognition model, ε i is the random noise of the i-th sampling, N is the number of sampling times of Monte Carlo simulation, and P(y|x) represents the maximum probability distribution of the gesture action category.

[0025] Optionally, after the step of using a preset Monte Carlo simulation algorithm to perform multiple random samplings on the probability distribution corresponding to the gesture action category and taking the gesture action category with the maximum probability as the user action intention of the gesture action to be identified, the method includes:

[0026] The intelligent terminal device obtains the user's gesture action category as the user action intention of the gesture action to be identified;

[0027] A terminal operation control instruction of the smart terminal device is generated according to the gesture action category, and terminal control of the smart terminal device is performed according to the terminal operation control instruction to complete the recognition of the user action intention of human-computer interaction.

[0028] Optionally, the intelligent terminal device includes an intelligent air conditioner, intelligent doors and windows, or intelligent lamps. Among them, the gesture action categories include drawing the letter O shape, drawing the letter C shape, drawing the letter U shape, or drawing the letter D shape. The terminal operation control instructions include turning on the light, opening the door, turning on the air conditioner, turning off the light, closing the door, turning off the air conditioner, increasing the brightness of the bulb, raising the temperature of the air conditioner, decreasing the brightness of the bulb, and decreasing the temperature of the air conditioner.

[0029] Optionally, the first sensor timing signal includes the gesture linear acceleration, gesture angular velocity, and gesture direction corresponding to each gesture action; the second sensor timing signal includes the gesture linear acceleration, gesture angular velocity, and gesture direction corresponding to each action segment; the gesture action is initiated by an intelligent control rod or the user's hand; the local action mode includes waving upward, waving left, waving downward, or waving right; the basic network architecture of the gesture action recognition model is a CNN-LSTM model based on the attention mechanism; the first time index represents the time point corresponding to each wave peak value, and the second time index represents the time point corresponding to each wave trough value.

[0030] A user action intention recognition device for human-computer interaction provided to meet another object of the present application includes:

[0031] A training set acquisition module, configured to respond to an instruction for recognizing the action intention of a human-computer interaction action, and acquire the first sensor timing signals corresponding to multiple gesture actions;

[0032] A clustering cluster determination module, configured to cluster the first sensor timing signals by using a preset K-Means clustering algorithm to determine multiple clustering clusters, and determine the wave peak value cluster and wave trough value cluster in each clustering cluster through a time sliding window based on a preset TPT model, where the clustering cluster represents the local action mode in the gesture action;

[0033] An action segment determination module, configured to, for the continuous wave trough values of the wave trough value cluster, find the intermediate wave peak value between the continuous wave trough values according to the first time index corresponding to the wave peak value in the wave peak value cluster and the second time index corresponding to the wave trough value in the wave trough value cluster, so as to segment multiple action segments in the gesture action;

[0034] A probability distribution determination module, configured to input the gesture action to be recognized into a gesture action recognition model trained by using the first sensor timing signals and the second sensor timing signals corresponding to the multiple action segments, so as to determine the probability distribution corresponding to the gesture action category of the gesture action to be recognized;

[0035] The action intention determination module is configured to perform multiple random samplings on the probability distribution corresponding to the gesture action category by using a preset Monte Carlo simulation algorithm, and take the gesture action category with the highest probability as the user action intention of the gesture action to be recognized, so as to complete the recognition of the user action intention for human-computer interaction.

[0036] An electronic device provided to meet another object of the present application includes a central processing unit and a memory. The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the method for recognizing the user action intention of the human-computer interaction described in the present application.

[0037] A computer-readable storage medium provided to meet another object of the present application stores a computer program implemented according to the method for recognizing the user action intention of the human-computer interaction in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the corresponding method.

[0038] Compared with the prior art, when the traditional remote control technology needs to achieve the coordinated control of multiple devices, its limitations will be highlighted. Since each device may need to be equipped with a dedicated remote control, this not only increases the complexity of operation, but also significantly improves the complexity and management difficulty of the entire control system. The present application includes, but is not limited to, the following beneficial effects:

[0039] First, the combination of K-Means clustering and the TPT model enables the solution to automatically learn using raw sensor data with less manual effort, significantly reducing the cost of pre-preparation. This means that developers can improve efficiency with less manual input. Through the K-Means clustering algorithm, even if the signal data is incomplete or contains noise, the algorithm can separate the noise into independent clusters, avoiding false detections and mislabeling, and improving the accuracy of data processing.

[0040] Second, during the data collection process, the signals of the sensors may be affected by the external environment, the hardware itself, or the operator, generating noise. The combination of the K-Means-TPT algorithm can effectively distinguish noise points from valid signals, thus avoiding such noise interference in subsequent analysis. For action signals with obvious periodicity or discreteness, this method can perform accurate action segmentation and feature extraction, thereby improving the accuracy and robustness of the subsequent classification model.

[0041] Thirdly, the TPT model directly relies on the physical characteristics of signals (such as wave peaks and wave valleys), enabling the K-Means clustering algorithm to effectively identify the local patterns of gesture actions. Even in a dynamically changing environment, this method still has strong robustness. Through precise action segmentation and feature extraction, this solution significantly improves the performance of subsequent gesture action classification models, especially in the recognition of complex actions, and can accurately distinguish different gesture action categories.

[0042] Fourthly, the combination of the CNN-LSTM-Attention model has powerful spatio-temporal feature capture capabilities, and can consider both the spatial features and time series features of actions simultaneously. In the task of the wand waving trajectory, the CNN can extract high-level features from the time series data of acceleration and angular velocity, the LSTM can model the long-term dependencies of actions, and the Attention mechanism helps the model focus on the most important time steps, improving the classification performance. This deep learning architecture can better adapt to complex action patterns and improve the accuracy and speed of recognition.

[0043] Fifthly, Monte Carlo simulation has unique advantages in dealing with complex non-linear relationships and uncertainties. By performing multiple random samplings of probability distributions, it can simulate the ambiguity and randomness in user behavior, thus more accurately predicting the intentions of users. The flexibility of Monte Carlo simulation enables this method to adapt to different user behavior patterns and provide personalized predictions and feedback for user interactions in different scenarios.

[0044] In summary, in multiple scenarios such as smart homes and theme parks, adopting this solution can provide users with a smoother, more personalized, and immersive control experience. By intelligently identifying the user's action intentions and providing real-time feedback, users can interact with devices in a more natural and flexible way, enhancing the fun and convenience of operation. Description of the Drawings

[0045] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0046] Figure 1 is a schematic flowchart of the method for identifying the user's action intention in the human-computer interaction of the embodiment of the present application;

[0047] Figure 2 is an exemplary structure of the wand in the embodiment of the present application;

[0048] Figure 3 is a schematic block diagram of the device for identifying the user's action intention in the human-computer interaction of the embodiment of the present application;

[0049] Figure 4It is a schematic structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners

[0050] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.

[0051] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when the present application states that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0052] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the field to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as herein.

[0053] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with a wireless signal receiver that only has the ability to receive and no ability to transmit, and devices with receiving and transmitting hardware that can perform two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers, tablet computers, etc., which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; conventional laptop and / or palmtop computers or other devices, which are conventional laptop and / or palmtop computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed manner at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, such as a PDA, MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or can also be a smart TV, a set-top box, etc.

[0054] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, which is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0055] It should be noted that the concept of "server" in this application can similarly be extended to apply to server clusters. According to the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can either be independent of each other but can be invoked through interfaces, or integrated into a single physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by this in the implementation of the network deployment method of this application.

[0056] One or several technical features of this application, unless expressly specified, can either be deployed on a server and accessed by a client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for access.

[0057] The neural network models cited or possibly cited in this application, unless expressly specified, can either be deployed on a remote server and remotely invoked on the client, or deployed on a client with sufficient device capabilities for direct invocation. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid over-occupying the client's hardware operating resources.

[0058] All kinds of data involved in this application, unless expressly specified, can either be remotely stored on a server or stored on a local terminal device, as long as it is suitable for being invoked by the technical solution of this application.

[0059] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can all be executed independently. Similarly, for each embodiment disclosed in this application, they are all proposed based on the same inventive concept. Therefore, concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, should be equivalently understood.

[0060] For each embodiment to be disclosed in this application, unless expressly pointed out that there is a mutually exclusive relationship between them, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0061] Please refer to Figure 1 , in one embodiment of the user action intention recognition method for human-computer interaction in this application, it includes:

[0062] Step S10: In response to an instruction for performing action intention recognition on a human-computer interaction action, obtain first sensor time-series signals corresponding to multiple gesture actions;

[0063] The user action intention recognition system in the intelligent terminal device can, in response to an instruction for performing action intention recognition on a human-computer interaction action, obtain first sensor time-series signals corresponding to multiple gesture actions; wherein, the gesture actions are initiated by an intelligent control stick or the user's hand; the intelligent control stick is named "magic wand" in this application, the gesture actions represent the waving trajectories of the magic wand or the user's hand, and they are used to generate terminal control instructions for the intelligent terminal device. The first sensor time-series signals include gesture linear acceleration, gesture angular velocity, and gesture direction corresponding to each gesture action.

[0064] In some embodiments, refer to Figure 2 , the magic wand includes a wand body 1, a switch button 2, a wand handle 3, a charging port 4, and an integrated PCB board 5 (the rest of the internal components are not shown). By analyzing the waving trajectory of the magic wand, this application believes that the waving trajectory of the magic wand has obvious similarity characteristics. Therefore, in the actual research process, for a certain type of action with certain rules, how to extract and recognize it from these data with action characteristics will be a difficulty and challenge. Through a large number of experiments, this application discovers that the best strategy to cope with such challenges lies in accurately recognizing and defining a complete action interval, which needs to accurately cover the entire trajectory of one waving of the magic wand and can clearly identify the key points of the start, transition, and end of the action and other landmark stages. On this basis, a model needs to be used to cognitively learn these landmark stages and strengthen the cognition in continuous learning to improve the cognitive accuracy. In order to better label the action data, it is necessary to standardize the action template setting. For example, define the action for the waving trajectory of the magic wand as shown in Table 1, where Table 1 is the definition of the waving trajectory of the magic wand.

[0065] Table 1 Definition of the waving trajectory of the magic wand

[0066]

[0067]

[0068] In order to accurately identify the defined waving trajectory, this application uses the IMU sensor inside the magic wand to collect data on the waving trajectory of the tester's magic wand, so as to determine the first sensor timing signals corresponding to multiple gesture actions. IMU is an inertial sensor, usually including an accelerometer, a gyroscope, a magnetometer, etc., which can measure the linear acceleration of the gesture, the angular velocity of the gesture, and the direction of the gesture at the same time. The linear acceleration of the gesture, the angular velocity of the gesture, and the direction of the gesture are used as the first sensor timing signals. Of course, in the design and testing stage, 3D modeling tool Blender can be considered to generate hand motion data to simulate the waving trajectory of the magic wand, so as to reduce cost expenditure.

[0069] Considering that in actual applications, the IMU is likely to be affected by external interference, such as large internal friction noise caused by poor manufacturing technology, signal interference such as high-frequency current generated by the human body, etc. Therefore, after obtaining the corresponding data, it is necessary to preprocess the data, and its main steps include filtering, smoothing and denoising, and normalization processing. In terms of filtering and denoising, this application uses a Butterworth low-pass filter. The Butterworth low-pass filter is a commonly used digital filter that can retain the low-frequency components in the signal while removing high-frequency noise. The system function of this filter is:

[0070]

[0071] where s is the complex frequency variable, Ω c is the passband cut-off frequency, and N is the order of the filter.

[0072] The Butterworth low-pass filter can be designed and applied using the scipy.signal library in Python. After removing high-frequency noise, the moving average smoothing operation is implemented using convolution to reduce noise and make the waveform smoother for subsequent processing. The formula is:

[0073]

[0074] where N is the window size, x(n) is the input signal, and y(n) is the smoothed signal. After filtering, smoothing and denoising, the sklearn.preprocessing library can be used to normalize the data, that is, scale the data to a unified range so that the data of different sensors have the same dimension.

[0075] In some embodiments, the first sensor timing signal includes gesture linear acceleration, gesture angular velocity and gesture direction corresponding to each gesture action; the second sensor timing signal includes gesture linear acceleration, gesture angular velocity and gesture direction corresponding to each action segment; the gesture action is initiated by a smart control stick or a user's hand; the local action mode includes waving up, waving to the left, waving down or waving to the right; the basic network architecture of the gesture action recognition model is a CNN-LSTM model based on an attention mechanism; the first time index represents the time point corresponding to each peak value, and the second time index represents the time point corresponding to each trough value.

[0076] Step S20: clustering the first sensor time series signal using a preset K-Means clustering algorithm to determine a plurality of clusters, and determining a peak value cluster and a trough value cluster in each cluster through a time sliding window based on a preset TPT model, wherein the clusters represent local action patterns in the gesture action;

[0077] After obtaining the first sensor timing signals corresponding to the plurality of gesture actions, clustering the first sensor timing signals using a preset K-Means clustering algorithm to determine a plurality of clustering clusters, and determining a peak value cluster and a trough value cluster in each clustering cluster through a time sliding window based on a preset TPT model, wherein the clustering clusters represent local action patterns in the gesture actions; the local action patterns include waving upward, waving to the left, waving downward, or waving to the right, etc.;

[0078] In some embodiments, clustering the first sensor time series signal using a preset K-Means clustering algorithm to determine a plurality of clustering clusters includes:

[0079] Step S201, obtaining a plurality of data points in the first sensor signal, and minimizing the sum of squares of distances from each data point in the first sensor signal to the center of the cluster to which it belongs according to the objective function of the K-Means clustering algorithm;

[0080] Step S202, iteratively optimize the objective function of the K-Means clustering algorithm, update the cluster centers and the allocation relationship, and continuously allocate data points to the nearest cluster centers during the iterative optimization process, and update the cluster centers to determine multiple clusters.

[0081] Specifically, the K-Means clustering algorithm is an unsupervised learning algorithm that can help identify feature points in the signal and cluster the obtained data. The present application divides the obtained data into K clusters, each of which corresponds to a local action mode, such as waving upwards, waving to the left, and other local action modes, thereby providing a basis for action segmentation.

[0082] Furthermore, the TPT model can extract action segments based on the "valley-peak-valley" pattern and segment independent action units from continuous data. After combining with the K-Means clustering algorithm, the TPT model can more accurately locate the start and end points of the action. Therefore, in the algorithm combining K-Means clustering with the TPT model, K-Means is mainly used as an auxiliary tool to optimize action segmentation and feature extraction, thereby improving the performance of the final classification model.

[0083] After the data is preprocessed, the data is used as the input of the K-Means clustering algorithm. The goal of the K-Means clustering algorithm is to divide the data points in the first sensor time series signal into K clusters to minimize the sum of the squares of the distances from each data point to the center of the cluster to which it belongs. The objective function of the K-Means clustering algorithm is:

[0084]

[0085] Where N is the total number of data points, K is the number of clusters, and x i is the i-th data point, μ j is the center of the jth cluster, r ij is an indicator variable, indicating that if the data point x i Belong to μ j When ij is 1 if the value is true, otherwise it is 0.

[0086] K-Means can also iteratively optimize the objective function, update the cluster center and the distribution relationship, and continuously iteratively identify the significant feature points in the signal, which can be more conducive to improving the subsequent segmentation of hand waving movements. In the iterative optimization process, data points will be continuously assigned to the nearest cluster center, and the cluster center will be updated. The expression is:

[0087]

[0088] By iteratively optimizing the objective function of the K-Means clustering algorithm, updating the cluster centers and the allocation relationship, data points are continuously allocated to the nearest cluster centers during the iterative optimization process, and the cluster centers are updated to determine multiple clusters.

[0089] In the TPT model, it is necessary to define the local maximum value (peak value) x t >x t-1 And x t >x t+1 and the local minimum (trough value) x t <x t-1 And x t <x t+1, the model can detect peaks and valleys in each cluster by comparing the current data point with the surrounding data within a sliding window. With the assistance of K-Means, the peak value cluster and the valley value cluster can be defined here, and the data points belonging to the peak value cluster and the valley value cluster can be filtered out. The expression is as follows:

[0090]

[0091] Among them, C p is the peak cluster, that is, the cluster indicated by the maximum cluster center maxμ j , and C t is the valley cluster, that is, the cluster indicated by the minimum cluster center minμ j .

[0092] Step S30: For the continuous valley values of the valley value cluster, according to the first time index corresponding to the peak value in the peak value cluster and the second time index corresponding to the valley value in the valley value cluster, find the intermediate peak value between the continuous valley values to segment multiple action segments in the gesture action;

[0093] After clustering the first sensor time series signal by using the preset K-Means clustering algorithm to determine multiple clusters, and determining the peak value cluster and the valley value cluster in each cluster through a time sliding window based on the preset TPT model, for the continuous valley values of the valley value cluster, according to the first time index corresponding to the peak value in the peak value cluster and the second time index corresponding to the valley value in the valley value cluster, find the intermediate peak value between the continuous valley values to segment multiple action segments in the gesture action;

[0094] In some embodiments, in the TPT model, {p1, p2,..., p m} can be set as the time index of the peak value in a certain cluster, {t1, t2,..., t m} can be set as the time index of the valley value in a certain cluster. For the continuous valley values {t i , t i+1} in each cluster, find the intermediate peak value p j to segment its action segments. The expression of the action segment in this application is defined as:

[0095] S i = [t i , p j , t i+1 (6)

[0096] Among them, S i represents the time range of the i-th action segment.

[0097] In some embodiments, for consecutive valley values of the valley value cluster, according to the first time index corresponding to the peak value in the peak value cluster and the second time index corresponding to the valley value in the valley value cluster, the step of finding an intermediate peak value between the consecutive valley values to segment multiple action segments in the gesture action includes:

[0098] Step S301: Initialize an empty list to store the extracted action segments;

[0099] Step S302: Traverse each pair of consecutive valley values in the second time index, where each pair of consecutive valley values represents the start point and the end point of a local action pattern;

[0100] Step S303: Between each pair of consecutive valley values, check whether there is a peak value in the first time index that is within this pair of consecutive valley values. If so, it means that a valid action segment has occurred in this interval, and then add the action segment to the empty list;

[0101] Step S304: Loop and execute the above steps until all action segments in the gesture action are segmented.

[0102] Step S40: Input the gesture action to be recognized into a gesture action recognition model trained with the first sensor timing signal and the second sensor timing signals corresponding to the multiple action segments to determine the probability distribution corresponding to the gesture action category of the gesture action to be recognized;

[0103] For consecutive valley values of the valley value cluster, after finding an intermediate peak value between the consecutive valley values according to the first time index corresponding to the peak value in the peak value cluster and the second time index corresponding to the valley value in the valley value cluster to segment multiple action segments in the gesture action, input the gesture action to be recognized into a gesture action recognition model trained with the first sensor timing signal and the second sensor timing signals corresponding to the multiple action segments to determine the probability distribution corresponding to the gesture action category of the gesture action to be recognized, where the second sensor timing signal includes the gesture linear acceleration, gesture angular velocity, and gesture direction corresponding to each action segment;

[0104] After completing the extraction of the action segments, then extract feature values for each action segment. The extraction of the feature values depends on the prominent features in different scenarios, such as time-domain features, frequency-domain features, statistical features, etc., so as to train a gesture action recognition model as the basis for gesture action classification.

[0105] For the action segments that have completed feature extraction, this application uses a CNN-LSTM model based on the attention mechanism as the basic network architecture of the gesture action recognition model to classify and train the actions. Among them, Convolutional Neural Networks (CNN) are good at capturing local spatial features, such as waveform features near peaks and valleys, while the Long Short-Term Memory (LSTM) model is good at modeling long-term dependencies in time series and can capture dynamic changes between different time steps. For the wand waving trajectory, the Convolutional Neural Network (CNN) can extract high-level feature representations from the time series of acceleration and angular velocity, and the Long Short-Term Memory (LSTM) can learn the overall dynamic pattern of the action. Specifically, the attention mechanism can weight important time step features, enabling the model to focus on the most important parts of the input sequence, thereby improving the classification performance. Of course, in deep learning, reinforcement learning can also be incorporated to form a deep reinforcement learning method, which continuously learns and optimizes its strategy through interaction with the environment to facilitate continuous accumulation of rewards.

[0106] More specifically, the architecture of the gesture action recognition model is constructed by a Convolutional Neural Network (CNN), a Long Short-Term Memory (LSTM), an Attention Mechanism, as well as a fully connected layer and an output layer. Among them, in the Convolutional Neural Network (CNN), local spatial features of the gesture action, such as waveform features near wave peaks and wave valleys, are extracted through convolutional layers. By means of multiple convolutional layers, features at different levels can be captured. Each convolutional layer applies multiple convolutional kernels to extract local temporal features from the temporal data of acceleration and angular velocity. For example, the size of the convolutional kernel can be adjusted according to the frequency characteristics of the gesture action to capture the instantaneous changes of the gesture action. In the Pooling Layers, pooling operations are used to reduce the dimension of the convolutional results, retaining important spatial information while reducing computational complexity.

[0107] In the long short-term memory network (LSTM), the LSTM layers are used to process the extracted temporal features and model the long-term dependencies in gesture actions. The LSTM can remember and utilize the information from previous time steps, thereby capturing the dynamic changes of the actions. The input of the LSTM layers is the temporal features extracted by the CNN, and the output is the hidden state at each time step. Through multiple layers of LSTM, the model can better learn the overall dynamic patterns of the actions. Based on the LSTM network, an attention mechanism is introduced to assign different weights to the features at different time steps, emphasizing the important time-step features. The attention mechanism helps the model focus on the key parts in the sequence, increase the weights of important features, and enable the model to pay more attention to the parts most relevant to the action category during classification. The time-step features weighted by the attention mechanism will be passed as input to the subsequent classification layer to improve the accuracy of action classification.

[0108] The features extracted from the CNN and LSTM are aggregated into a vector and input into the fully connected layer for further processing. The output of the Output Layer is a vector containing the probability distributions of each category, representing the probabilities that the gesture action to be recognized belongs to each category. The Softmax activation function is used for the output to ensure that the sum of the output probabilities is 1.

[0109] The second sensor temporal signals corresponding to the multiple action segments and the first sensor temporal signal are input into the constructed gesture action recognition model for training. After the gesture action recognition model is trained to convergence, the gesture action to be recognized is input into the gesture action recognition model that has been trained to the convergence state to determine the gesture action category of the gesture action to be recognized and its corresponding probability distribution.

[0110] Step S50: Use a preset Monte Carlo simulation algorithm to perform multiple random samplings on the probability distribution corresponding to the gesture action category, and take the gesture action category with the highest probability as the user action intention of the gesture action to be recognized, so as to complete the recognition of the user action intention in human-computer interaction.

[0111] After inputting the gesture action to be recognized into the gesture action recognition model trained with the first sensor temporal signal and the second sensor temporal signals corresponding to the multiple action segments to determine the probability distribution corresponding to the gesture action category of the gesture action to be recognized, use a preset Monte Carlo simulation algorithm to perform multiple random samplings on the probability distribution corresponding to the gesture action category, and take the gesture action category with the highest probability as the user action intention of the gesture action to be recognized, so as to complete the recognition of the user action intention in human-computer interaction.

[0112] In some embodiments, the step of performing multiple random samplings on the probability distribution corresponding to the gesture action category by using a preset Monte Carlo simulation algorithm and taking the gesture action category with the highest probability as the user action intention of the gesture action includes:

[0113] After the model has completed action recognition and classification, a user intention prediction method based on Monte Carlo simulation can be adopted. This method can construct the behavior probability distribution of the user according to historical action data, further improving the intelligence level of the system. The principle of Monte Carlo Simulation is mainly to generate a large number of random samples and perform statistical analysis on these samples based on the law of large numbers to estimate the solution of the problem or the behavior of the system, which is especially suitable for dealing with problems with uncertainty or complex probability distributions.

[0114] By performing random sampling on the probability distribution corresponding to the gesture action category and calculating the probability distribution, the system can accurately predict the user's action intention, thereby improving the intelligence level of human-computer interaction. This method has high accuracy, good adaptability, and can be widely applied in fields such as smart home, virtual reality, and robot control.

[0115] Based on the output data of the gesture action recognition model, obtain the gesture action category of each gesture action and its corresponding probability distribution. Perform multiple random samplings on the probability distribution corresponding to the gesture action category by using a preset Monte Carlo simulation algorithm, and take the gesture action category with the highest probability as the user action intention of the gesture action. Among them, the expression for the Monte Carlo simulation algorithm to predict the probability distribution is:

[0116]

[0117] where x is the input data, f represents the gesture action recognition model, ε i is the random noise of the i-th sampling, N is the number of samplings in the Monte Carlo simulation, and P(y|x) represents the maximum probability distribution of the gesture action category.

[0118] In a further embodiment, after the step of performing multiple random samplings on the probability distribution corresponding to the gesture action category by using a preset Monte Carlo simulation algorithm and taking the gesture action category with the highest probability as the user action intention of the gesture action to be recognized, it includes:

[0119] Step S501: The intelligent terminal device obtains the gesture action category of the user as the user action intention of the gesture action to be recognized;

[0120] Step S502: Generate a terminal operation control instruction for the intelligent terminal device according to the gesture action category, and perform terminal control of the intelligent terminal device according to the terminal operation control instruction to complete the recognition of the user action intention of the human-computer interaction.

[0121] In some embodiments, the intelligent terminal device includes an intelligent air conditioner, intelligent doors and windows, or intelligent lights. Among them, the gesture action categories include drawing the letter O shape, drawing the letter C shape, drawing the letter U shape, or drawing the letter D shape, and the terminal operation control instructions include turning on the light, opening the door, turning on the air conditioner, turning off the light, closing the door, turning off the air conditioner, increasing the brightness of the light bulb, raising the temperature of the air conditioner, decreasing the brightness of the light bulb, and decreasing the temperature of the air conditioner.

[0122] Specifically, the creation of the control rule table is the last step of controlling the intelligent terminal device with the magic wand after recognizing the action. This application needs to map each gesture category to a specific device, and generate corresponding control instructions through communication means after recognizing the specified action. Taking the smart home as an example, please refer to Table 2, which is an example of the control rule table.

[0123] Table 2 Example of Control Rule Table

[0124]

[0125]

[0126] After the magic wand recognizes the action, it generates a control signal and sends it to the intelligent terminal device through the corresponding communication protocol and communication module to achieve device control. The establishment of these rule tables can be stored in a database for easy subsequent expansion and modification. The above is only an example for reference.

[0127] In some embodiments, the user action intention recognition method of the human-computer interaction of this application can be applied to the following scenarios:

[0128] 1. VR and AR in theme parks: Users can interact with the virtual environment through gestures or using carrier tools to achieve gesture interaction and motion capture. Use the action recognition model to process the sensor data, and map the recognized actions to interaction instructions in the virtual environment, facilitating users to control virtual objects through gestures (such as grasping, moving, rotating, etc.).

[0129] 2. Motion monitoring in the gym: Use wearable devices to collect sensor data, collect sensor data, and monitor the user's motion actions in real time. Analyze the user's motion patterns through the action recognition model, provide motion data analysis, such as action standardization and calorie consumption, and provide real-time feedback and corrective suggestions according to the suggestions.

[0130] 3. Medical treatment and nursing: The doctor performs surgery with a scalpel carrying an identification model, maps the identified actions to medical or nursing equipment, monitors the doctor's surgical actions, and provides real-time feedback. For example, in minimally invasive surgery, the system can detect whether the doctor's operation is accurate.

[0131] In summary, compared with the prior art, when the traditional remote control technology needs to achieve the coordinated control of multiple devices, its limitations will be prominent. Since each device may need to be equipped with a dedicated remote control, this not only increases the complexity of operation, but also significantly improves the complexity and management difficulty of the entire control system. The present application includes, but is not limited to, the following beneficial effects:

[0132] First, the combination of K-Means clustering and TPT model enables the solution to automatically learn using raw sensor data with little manual input, greatly reducing the cost of preliminary preparation. This means that developers can improve efficiency. Through the K-Means clustering algorithm, even if the signal data is incomplete or contains noise, the algorithm can separate the noise into independent clusters to avoid misdetection and mislabeling, improving the accuracy of data processing.

[0133] Second, during the data collection process, the signals of the sensors may be affected by the external environment, the hardware itself, or the operator, generating noise. The combination of the K-Means-TPT algorithm can effectively distinguish noise points from valid signals, thus avoiding the interference of this noise on subsequent analysis. For action signals with obvious periodicity or discreteness, this method can perform accurate action segmentation and feature extraction, thereby improving the accuracy and robustness of the subsequent classification model.

[0134] Third, the TPT model directly depends on the physical characteristics of the signal (such as wave peaks and valleys), enabling the K-Means clustering algorithm to effectively identify the local patterns of gesture actions. Even in a dynamically changing environment, this method still has strong robustness. Through accurate action segmentation and feature extraction, this solution significantly improves the performance of the subsequent gesture action classification model, especially in the recognition of complex actions, and can accurately distinguish different gesture action categories.

[0135] Fourthly, the combination of the CNN-LSTM-Attention model has a powerful spatio-temporal feature capture ability, which can consider both the spatial features and time series features of actions simultaneously. In the task of the wand waving trajectory, the CNN can extract high-level features from the time series data of acceleration and angular velocity, the LSTM can model the long-term dependencies of actions, and the Attention mechanism helps the model focus on the most important time steps, improving the classification performance. This deep learning architecture can better adapt to complex action patterns, enhancing the accuracy and speed of recognition.

[0136] Fifthly, Monte Carlo simulation has unique advantages in dealing with complex non-linear relationships and uncertainties. By performing multiple random samplings of probability distributions, it can simulate the ambiguity and randomness in user behavior, thus predicting the user's intention more accurately. The flexibility of Monte Carlo simulation enables this method to adapt to different user behavior patterns, providing personalized predictions and feedback for user interactions in different scenarios.

[0137] In summary, in multiple scenarios such as smart homes and theme parks, adopting this solution can provide users with a smoother, more personalized, and immersive control experience. By intelligently recognizing the user's action intention and providing real-time feedback, users can interact with devices in a more natural and flexible way, enhancing the fun and convenience of operation.

[0138] Please refer to the figure. A user action intention recognition device for human-computer interaction provided to meet one of the purposes of this application includes a training set acquisition module 1100, a clustering cluster determination module 1200, an action segment determination module 1300, a probability distribution determination module 1400, and an action intention determination module 1500. Among them, the training set acquisition module 1100 is configured to respond to an instruction for recognizing the action intention of a human-computer interaction action and acquire first sensor time series signals corresponding to multiple gesture actions; the clustering cluster determination module 1200 is configured to perform clustering on the first sensor time series signals by using a preset K-Means clustering algorithm to determine multiple clustering clusters, and determine the peak value cluster and valley value cluster in each clustering cluster through a time sliding window based on a preset TPT model, where the clustering cluster represents the local action pattern in the gesture action; the action segment determination module 1300 is configured to, for the continuous valley values of the valley value cluster, find the intermediate peak values between the continuous valley values according to the first time index corresponding to the peak values in the peak value cluster and the second time index corresponding to the valley values in the valley value cluster to segment multiple action segments in the gesture action; the probability distribution determination module 1400 is configured to input the gesture action to be recognized into a gesture action recognition model trained by using the first sensor time series signals and second sensor time series signals corresponding to the multiple action segments to determine the probability distribution corresponding to the gesture action category of the gesture action to be recognized; the action intention determination module 1500 is configured to perform multiple random samplings on the probability distribution corresponding to the gesture action category by using a preset Monte Carlo simulation algorithm, and use the gesture action category with the highest probability as the user action intention of the gesture action to be recognized to complete the recognition of the user action intention of the human-computer interaction.

[0139] Based on any embodiment of this application, please refer to Figure 4 , another embodiment of this application further provides an electronic device, which can be implemented by a computer device. As shown in Figure 4 , is a schematic internal structure diagram of the computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a method for recognizing the user action intention of human-computer interaction. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the method for recognizing the user action intention of human-computer interaction of this application. The network interface of the computer device is used to connect and communicate with the terminal. Those skilled in the art can understand,Figure 4 The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0140] In this embodiment, the processor is used to execute Figure 3 the specific functions of each module in. The memory stores the program codes and various types of data required to execute the above-mentioned modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program codes and data required to execute all modules in the user action intention recognition device for human-computer interaction of this application, and the server can call the program codes and data of the server to execute the functions of all modules.

[0141] This application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the method for recognizing a user action intention in human-computer interaction according to any embodiment of this application.

[0142] This application also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by one or more processors, the steps of the method for recognizing a user action intention in human-computer interaction according to any embodiment of this application are implemented.

[0143] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of this application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium may be a computer-readable storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0144] The above are only some embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A method for recognizing user action intentions in human-computer interaction, characterized in that, include: In response to an instruction to recognize the action intention of a human-computer interaction action, a first sensor timing signal corresponding to a plurality of gesture actions is obtained; Clustering the first sensor time series signal using a preset K-Means clustering algorithm to determine a plurality of clusters, and determining a peak value cluster and a trough value cluster in each cluster through a time sliding window based on a preset TPT model, wherein the clusters represent local action patterns in the gesture action; For the continuous trough values ​​of the trough value cluster, according to the first time index corresponding to the peak value in the peak value cluster and the second time index corresponding to the trough value in the trough value cluster, finding the intermediate peak value between the continuous trough values ​​to segment the multiple action segments in the gesture action; Inputting the gesture action to be recognized into a gesture action recognition model trained by using the first sensor timing signal and the second sensor timing signal corresponding to the plurality of action segments to determine a probability distribution corresponding to the gesture action category of the gesture action to be recognized; A preset Monte Carlo simulation algorithm is used to perform multiple random samplings on the probability distribution corresponding to the gesture action category, and the gesture action category with the maximum probability is used as the user action intention of the gesture action to be identified, so as to complete the identification of the user action intention of human-computer interaction.

2. The user action intention recognition method for human-computer interaction according to claim 1, characterized in that, For the continuous trough values ​​of the trough value cluster, according to the first time index corresponding to the peak value in the peak value cluster and the second time index corresponding to the trough value in the trough value cluster, the step of finding the intermediate peak value between the continuous trough values ​​to segment the multiple action segments in the gesture action includes: Initialize an empty list to store the extracted action fragments; Traversing each pair of consecutive trough values ​​in the second time index, wherein each pair of consecutive trough values ​​represents a starting point and an end point of a local motion pattern; Between each pair of consecutive trough values, check whether there is a peak value in the first time index located within the consecutive trough values. If so, it indicates that a valid action segment has occurred within this interval, and the action segment is added to the empty list; The above steps are executed repeatedly until all the action segments in the gesture action are segmented.

3. The method for identifying user action intention of human-computer interaction according to claim 1, characterized in that, The step of clustering the first sensor time series signal using a preset K-Means clustering algorithm to determine a plurality of clusters includes: Acquire multiple data points in the first sensor signal, and minimize the sum of squares of distances from each data point in the first sensor signal to the center of the cluster to which it belongs according to the objective function of the K-Means clustering algorithm; The objective function of the K-Means clustering algorithm is iteratively optimized, and the cluster centers and the allocation relationships are updated. During the iterative optimization process, data points are continuously allocated to the nearest cluster centers, and the cluster centers are updated to determine multiple clusters.

4. The method for identifying a user action intention of human-computer interaction according to claim 1, characterized in that, The step of using a preset Monte Carlo simulation algorithm to perform multiple random samplings on the probability distribution corresponding to the gesture action category, and taking the gesture action category with the maximum probability as the user action intention of the gesture action, includes: Based on the output data of the gesture action recognition model, obtain the gesture action category of each gesture action and its corresponding probability distribution. Use a preset Monte Carlo simulation algorithm to perform multiple random samplings on the probability distribution corresponding to the gesture action category, and take the gesture action category with the highest probability as the user action intention of the gesture action. Wherein, the expression for predicting the probability distribution by the Monte Carlo simulation algorithm is: where x is the input data, f represents the gesture action recognition model, ε i is the random noise of the i-th sampling, N is the number of samplings for Monte Carlo simulation, and P(y|x) represents the maximum probability distribution of the gesture action category.

5. The method for identifying user action intention in human-computer interaction according to claim 1, characterized in that, After the step of using a preset Monte Carlo simulation algorithm to perform multiple random samplings on the probability distribution corresponding to the gesture action category and taking the gesture action category with the highest probability as the user action intention of the gesture to be recognized, it includes: The intelligent terminal device obtains the gesture action category of the user as the user action intention of the gesture to be recognized; Generate a terminal operation control instruction for the intelligent terminal device according to the gesture action category, and perform terminal control of the intelligent terminal device according to the terminal operation control instruction to complete the recognition of the user action intention of the human-computer interaction.

6. The method for identifying user action intentions in human-computer interaction according to any one of claims 1 to 5, characterized in that, The intelligent terminal device includes an intelligent air conditioner, intelligent doors and windows or intelligent lamps. Wherein, the gesture action category includes drawing the letter O shape, drawing the letter C shape, drawing the letter U shape or drawing the letter D shape, and the terminal operation control instructions include turning on the light, opening the door, turning on the air conditioner, turning off the light, closing the door, turning off the air conditioner, increasing the brightness of the bulb, increasing the temperature of the air conditioner, decreasing the brightness of the bulb, and decreasing the temperature of the air conditioner.

7. The user action intention recognition method for human-computer interaction according to any one of claims 1 to 5, characterized in that, The first sensor timing signal includes the gesture linear acceleration, gesture angular velocity and gesture direction corresponding to each gesture action; the second sensor timing signal includes the gesture linear acceleration, gesture angular velocity and gesture direction corresponding to each action segment; the gesture action is initiated by an intelligent control rod or the user's hand; the local action mode includes waving upward, waving leftward, waving downward or waving rightward; the basic network architecture of the gesture action recognition model is a CNN-LSTM model based on the attention mechanism; the first time index represents the time point corresponding to each wave peak value, and the second time index represents the time point corresponding to each wave trough value.

8. A user action intention recognition device for human-computer interaction, characterized in that It includes: A training set acquisition module, configured to respond to an instruction for recognizing an action intention of a human-computer interaction action and acquire first sensor timing signals corresponding to multiple gesture actions; A clustering cluster determination module, configured to use a preset K-Means clustering algorithm to cluster the first sensor timing signals to determine multiple clustering clusters, and determine a wave peak cluster and a wave trough cluster in each clustering cluster through a time sliding window based on a preset TPT model, wherein the clustering cluster represents a local action mode in the gesture action; An action segment determination module, configured to, for the continuous wave trough values of the wave trough cluster, find the intermediate wave peak between the continuous wave trough values according to the first time index corresponding to the wave peak in the wave peak cluster and the second time index corresponding to the wave trough in the wave trough cluster, so as to segment multiple action segments in the gesture action; A probability distribution determination module, configured to input a gesture action to be recognized into a gesture action recognition model trained using the first sensor timing signal and the second sensor timing signals corresponding to the multiple action segments, so as to determine the probability distribution corresponding to the gesture action category of the gesture action to be recognized; An action intention determination module, configured to perform multiple random samplings on the probability distribution corresponding to the gesture action category by using a preset Monte Carlo simulation algorithm, and use the gesture action category with the highest probability as the user action intention of the gesture action to be recognized, so as to complete the recognition of the user action intention for human-computer interaction.

9. An electronic device, comprising a central processing unit and a memory, characterized in that, The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to the method according to any one of claims 1 to 7. When the computer program is called and run by the computer, it executes the steps included in the corresponding method.

Citation Information

Cited By

  • Intelligent magic wand interaction method and system based on gyroscope gesture recognition

    CN121364786A

  • An intelligent magic wand interaction method and system based on gyroscope gesture recognition

    CN121364786B