A human-computer interaction method, device, equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2026-03-12
- Publication Date
- 2026-08-04
AI Technical Summary
压力信号主要反映用户发力的程度,但由于不同用户的肌肉分布和发力习惯存在差异,导致压力信号的幅值和频率特征的适配性低,相同肢体动作在不同用户间采集的信号存在差异,增加特征提取的难度,导致动作识别的准确率低
[0019] The technical solution of this invention broadens the signal dimensions of action recognition by acquiring multiple types of biomechanical signals transmitted by wearable devices. It acquires at least pressure signals and inertial signals for action recognition, solving the problem of low accuracy when using a single pressure signal for action recognition. Based on multiple types of biomechanical signals, more accurate action recognition results can be obtained. Using accurate action recognition results as the basis for business response, the business processing can be accurately matched with the user's real needs, reducing business execution errors and thus improving the reliability of human-computer interaction.
Smart Images

Figure CN122507264A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a human-computer interaction method, apparatus, device, and medium. Background Technology
[0002] To meet hygiene and safety requirements in public settings, contactless interaction has gradually become a major area of optimization for human-computer interaction methods. One way to achieve contactless interaction is for users to control terminal devices through body movements.
[0003] Current limb movement recognition interaction technologies mainly rely on collecting pressure signals. Pressure signals primarily reflect the degree of force exerted by the user. However, due to differences in muscle distribution and force exertion habits among different users, the adaptability of the amplitude and frequency characteristics of pressure signals is low. The signals collected for the same limb movement differ among different users, increasing the difficulty of feature extraction and resulting in low accuracy of movement recognition. Summary of the Invention
[0004] This invention provides a human-computer interaction method, device, equipment, and medium that can improve the accuracy of action recognition and enhance the reliability of human-computer interaction.
[0005] According to one aspect of the present invention, an embodiment of the present invention provides a human-computer interaction method, the method comprising:
[0006] During the human-computer interaction process at the business terminal, multiple types of biomechanical signals transmitted through the wearable device worn by the user are acquired; the multiple types of biomechanical signals include at least: pressure signals and inertial signals;
[0007] Action recognition is performed on each of the biomechanical signals to obtain action recognition results;
[0008] The system responds to the current business processing request based on the action recognition result.
[0009] According to another aspect of the present invention, embodiments of the present invention also provide a human-computer interaction device, the device comprising:
[0010] The signal acquisition module is used to acquire multiple types of biomechanical signals transmitted through the wearable device worn by the user during the human-computer interaction process of the business terminal; the multiple types of biomechanical signals include at least: pressure signals and inertial signals;
[0011] The action recognition module is used to perform action recognition on each of the biomechanical signals to obtain action recognition results;
[0012] The request-response module is used to respond to the current business processing request based on the action recognition result.
[0013] According to another aspect of the present invention, embodiments of the present invention also provide a human-computer interaction device, the human-computer interaction device comprising:
[0014] At least one processor; and
[0015] A memory that is communicatively connected to at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the human-computer interaction method of any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the human-computer interaction method of any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the human-computer interaction method described in any embodiment of the present invention.
[0019] The technical solution of this invention broadens the signal dimensions of action recognition by acquiring multiple types of biomechanical signals transmitted by wearable devices. It acquires at least pressure signals and inertial signals for action recognition, solving the problem of low accuracy when using a single pressure signal for action recognition. Based on multiple types of biomechanical signals, more accurate action recognition results can be obtained. Using accurate action recognition results as the basis for business response, the business processing can be accurately matched with the user's real needs, reducing business execution errors and thus improving the reliability of human-computer interaction.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a human-computer interaction method provided according to an embodiment of the present invention;
[0023] Figure 2This is a flowchart of a human-computer interaction method provided according to an embodiment of the present invention;
[0024] Figure 3 This is a structural diagram of a human-computer interaction device according to an embodiment of the present invention;
[0025] Figure 4 This is a structural schematic diagram of a human-computer interaction device provided in an embodiment of the present invention. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] The acquisition, storage, and application of biomechanical signals involved in the technical solutions of this invention comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0029] This invention, in its embodiments, acquires, stores, and applies biomechanical signals after obtaining explicit authorization from the user. The authorization method employs prominent means such as pop-up confirmation or checkboxes, clearly specifying the scope of information collection, usage methods, and retention period. Biomechanical signals are acquired only after user confirmation of authorization. Authorized biomechanical signals are used only within the authorized scope. The acquired biomechanical signals are protected throughout their entire lifecycle using high-level encryption technology.
[0030] Figure 1This is a flowchart illustrating a human-computer interaction method provided in an embodiment of the present invention. This embodiment is applicable to human-computer interaction based on action recognition. The method can be executed by a human-computer interaction device, which can be implemented in hardware and / or software. The human-computer interaction device can be configured in a server.
[0031] See Figure 1 The human-computer interaction methods shown include:
[0032] S101. During the human-computer interaction process of the business terminal, multiple types of biomechanical signals transmitted through the wearable device worn by the user are acquired, wherein the multiple types of biomechanical signals include at least: pressure signals and inertial signals.
[0033] The service terminal can be a smart terminal for handling business. Service terminals include ATMs, self-service card issuing machines, or smart counters. Wearable devices can be data collection devices worn on the user for signal acquisition. Wearable devices include wearable bracelets, wearable armbands, or fingertip patches.
[0034] Biomechanical signals can be signals generated by user limb movements. Multiple types of biomechanical signals include pressure signals and inertial signals. At least two types of biomechanical signals should be acquired for movement recognition. Optionally, both pressure and inertial signals can be acquired for movement recognition. Pressure signals can be signals generated by changes in the contact force between the user's limb and the wearable device. During user interaction, the contact force between the user's limb and the wearable device changes as the interaction progresses. This change can be reflected by pressure signals acquired on the wearable device. Pressure signals reflect the force exertion state and are suitable for recognizing fine motor skills. Inertial signals can be signals generated by limb movement. Inertial signals include acceleration and angular velocity signals. Inertial signals are used to capture the trajectory and velocity of limb movements. Optionally, pressure and inertial signals can be acquired through a wearable device that integrates pressure and inertial signal sensors.
[0035] Users can interact with the service terminal through body movements. During the interaction, the service terminal prompts the user to wear a wearable device. After wearing the device, the user performs an interactive action. The wearable device can acquire the biomechanical signals generated by the user's actions and transmit them to the service terminal. For example, when a user uses the service terminal to conduct business, they wear a wearable armband on their forearm and perform an interactive action; the armband acquires the biomechanical signals from the user's forearm movements.
[0036] In an optional embodiment, during the human-computer interaction process of the business terminal, acquiring multiple types of biomechanical signals transmitted through the wearable device worn by the user includes: during the human-computer interaction process of the business terminal, demonstrating multiple action videos and action instructions corresponding to each action video through the business terminal; during the process of the user following the action video to perform corresponding actions, acquiring multiple types of biomechanical signals transmitted through the wearable device worn by the user, wherein the user performs the corresponding actions to input the action instructions corresponding to the corresponding action video.
[0037] Action commands can be instructions that control the business terminal to perform specific operations. For example, action commands include instructions such as "Previous", "Next", "Cancel", or "Confirm".
[0038] The service terminal demonstrates multiple interactive actions to the user via video on its display interface, simultaneously showing the action command represented by each action. The user performs the same interactive actions as demonstrated in the video according to their needs, and the wearable device collects multiple types of biomechanical signals during the user's actions. The user then inputs action commands to the service terminal by performing the corresponding actions.
[0039] For example, if a video demonstration shows an arm moving to the left, the corresponding action command is "previous step". The user follows the video and moves their arm to the left. The system performs action recognition by acquiring multiple types of biomechanical signals related to this movement. If the action recognition result indicates the action type is "arm to the left", and "arm to the left" corresponds to the "previous step" action command, then the system sends the command "previous step", and the system executes the "return to previous step" operation.
[0040] Optionally, before demonstrating multiple action videos and their corresponding action commands on the business terminal, a video showing how to wear the wearable device can be shown. Optionally, when the wearable device is designed for wearing on the upper limb, the user can choose to wear it on the left or right upper limb and confirm the wearing location on the system. For example, when the wearable device is a wearable armband, the system demonstrates how to wear it and prompts the user to choose between wearing it on the left or right arm. After the user has worn the armband, a voice prompt indicates whether the armband is worn on the left or right arm.
[0041] It is evident that by demonstrating action videos and having users perform the corresponding actions according to the videos, the actions performed by users become more standardized and the characteristics of the actions become more obvious, thus enabling biomechanical signals to carry more accurate information.
[0042] In an optional embodiment, the wearable device is communicatively connected to the service terminal, and the wearable device includes at least one electrode acquisition position.
[0043] The electrode acquisition location can be a point on the wearable device used to collect biomechanical signals. One electrode acquisition location corresponds to a physiological location of the user, and different electrode acquisition locations can collect biomechanical signals from different locations on the user's limbs. For example, a wearable armband has two electrode acquisition locations, one on the inner side of the forearm and one on the outer side. The electrode acquisition location on the inner side of the forearm can collect biomechanical signals from the inner side of the user's forearm, and the electrode acquisition location on the outer side of the forearm can collect biomechanical signals from the outer side of the user's forearm. The wearable device includes at least one electrode acquisition location. In a specific embodiment, the wearable device is a wearable armband that integrates a 16-channel pressure signal sensor and a 48-channel inertial signal sensor. Here, "channel" can refer to an electrode acquisition location.
[0044] The wearable device communicates with the business terminal, transmitting the collected biomechanical signals to the terminal. This connection can be either wired or wireless.
[0045] It is evident that by defining the communication connection between the wearable device and the business terminal, and by including at least one electrode acquisition position on the wearable device, the transmission direction and spatial location of the signal can be clearly defined.
[0046] S102. Perform motion recognition on each of the biomechanical signals to obtain motion recognition results.
[0047] This involves extracting signal features from multiple types of biomechanical signals, classifying the extracted signal features, and obtaining action recognition results.
[0048] Optionally, a time window can be set to retain the historical action recognition results within the previous time window. For example, if 0-5 seconds is time window 1, 5-10 seconds is time window 2, and the current time is 12 seconds, then the action recognition results of time window 2 will be retained. This is to facilitate subsequent correction of action recognition results for non-standard interactive actions.
[0049] S103. Respond to the current business processing request based on the action recognition result.
[0050] In this context, a business processing request can be an operation request initiated by the user to handle the current business. Based on the action recognition result, the corresponding operation is executed on the business processing request initiated by the user.
[0051] In an optional embodiment, responding to the current business processing request based on the action recognition result includes: determining the action instruction input by the user based on the action type in the action recognition result; and responding to the current business processing request based on the action instruction.
[0052] The action type can be a type of limb movement. For example, limb movements such as moving the arm to the left, moving the arm to the right, waving the arm left and right, moving the arm upward, or moving the hand downward can be used as action types. A mapping relationship is established between action types and action commands. Based on the action type in the action recognition results, the corresponding action command is determined, and the system responds to the current business process according to the action command. For example, when a user withdraws cash from an ATM, the current ATM interface is the withdrawal amount input interface. If the system recognizes the user's action as moving the arm to the right, and the corresponding action command for moving the arm to the right is "next step," then the system proceeds to the next step and jumps to the password input interface.
[0053] A specific mapping relationship between action types and action commands is as follows: Moving the arm to the left corresponds to the previous step, moving the arm to the right corresponds to the next step, swinging the arm left and right corresponds to cancel, moving the arm upward corresponds to the previous item, moving the hand downward corresponds to the next item, curling the arm corresponds to re-entering, and extending the arm outward corresponds to ending the input. Optionally, in the system's numeric input interface, moving the arm upward can correspond to the previous digit, and moving the hand downward can correspond to the next digit.
[0054] As can be seen, by determining the action command based on the action type in the action recognition result, and responding to the current business processing request according to the action command, the user's operation intention is clarified through the accurate mapping between action type and action command, ensuring that the business processing follows the user's actual operation needs.
[0055] The technical solution of this invention broadens the signal dimensions of action recognition by acquiring multiple types of biomechanical signals transmitted by wearable devices. It acquires at least pressure signals and inertial signals for action recognition, solving the problem of low accuracy when using a single pressure signal for action recognition. Based on multiple types of biomechanical signals, more accurate action recognition results can be obtained. Using accurate action recognition results as the basis for business response, the business processing can be accurately matched with the user's real needs, reducing business execution errors and thus improving the reliability of human-computer interaction.
[0056] Figure 2 This is a flowchart of a human-computer interaction method provided by an embodiment of the present invention. Based on the above embodiments, this embodiment of the present invention further defines the action recognition of each biomechanical signal to obtain an action recognition result as follows: fusing the biomechanical signals of each type to obtain a target signal; extracting channel spatial features from the target signal to obtain spatial features of the target signal; extracting amplitude-time features from the target signal to obtain temporal features of the target signal; fusing the spatial and temporal features of the target signal to obtain signal features of the target signal; and classifying the signal features of the target signal to obtain the action recognition result.
[0057] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments.
[0058] See Figure 2 The human-computer interaction methods shown include:
[0059] S201. During the human-computer interaction process of the business terminal, multiple types of biomechanical signals transmitted through the wearable device worn by the user are acquired, wherein the multiple types of biomechanical signals include at least: pressure signals and inertial signals.
[0060] S202. The biomechanical signals of each type are fused to obtain the target signal.
[0061] This involves fusing multiple types of biomechanical signals. A fusion algorithm integrates these signals into a target signal, which, as the fused signal, overcomes the limitations of a single signal.
[0062] In an optional embodiment, fusing the biomechanical signals of each type to obtain the target signal includes splicing the pressure signal and inertial signal included in each type of biomechanical signal to obtain the target signal.
[0063] Biomechanical signals include pressure signals and inertial signals. Pressure signals excel at reflecting a user's exertion state and are well-suited for recognizing small, fine movements, while inertial signals can capture changes in the speed and displacement of limb movements and are well-suited for perceiving large-amplitude movement trajectories. By concatenating pressure and inertial signals, the target signal can accurately recognize both small, fine movements and large-amplitude movements such as arm translation and swinging. Optionally, concatenating pressure and inertial signals can be done by keeping the time dimension constant and directly superimposing the channel dimensions of the pressure and inertial signals.
[0064] It is evident that by splicing pressure signals and inertial signals, the target signal can both perceive motion trajectories and capture fine movements. The two signals complement each other, ensuring stable recognition results.
[0065] Optionally, the biomechanical signals of each type may be preprocessed before fusion.
[0066] When biomechanical signals include pressure signals and inertial signals, a specific preprocessing method is as follows: use a sliding window with a duration of 200ms and a time interval of 100ms along the time axis to segment the pressure signal, acceleration signal and angular velocity signal.
[0067] A 5Hz high-pass filter is used to remove the DC component from the pressure signal, a Gaussian filter is used to filter out high-frequency noise from the acceleration signal, and a zero-phase low-pass filter is used to smooth the angular velocity signal to eliminate abnormal jitter, thereby achieving signal standardization and noise reduction.
[0068] S203. Extract the channel spatial features of the target signal to obtain the spatial features of the target signal.
[0069] The channel spatial features can be signal features from multiple electrode acquisition locations. These multiple electrode acquisition locations reflect the spatial attributes of the target signal. The spatial features of the target signal are then extracted.
[0070] In an optional embodiment, the step of extracting channel spatial features from the target signal to obtain the spatial features of the target signal includes: extracting local channel spatial features from the target signal to obtain a first channel feature; extracting global channel spatial features from the target signal to obtain a second channel feature; and fusing the first channel feature and the second channel feature to obtain the spatial features of the target signal.
[0071] Local channel spatial features can be used to capture the channel spatial features of a local part of the target signal. Global channel spatial features can be used to capture the channel spatial features of the target signal as a whole.
[0072] The target signal is input into a spatial feature extraction module with a small convolutional kernel size to extract local channel spatial features, resulting in the first channel feature. The target signal is then input into a spatial feature extraction module with a large convolutional kernel size to extract global channel spatial features, resulting in the second channel feature.
[0073] The spatial features of the target signal are obtained by fusing the features of the first channel and the features of the second channel.
[0074] Optionally, the target signal can be input into convolutional neural networks with kernels of different sizes to obtain first-channel features and second-channel features. In a specific embodiment, the target signal is input into two neural networks with three convolutional layers, and the parameter settings of the two convolutional neural networks are shown in the table below:
[0075]
[0076]
[0077] Wherein: the kernel size of convolutional neural network 1 is smaller than that of convolutional neural network 2; convolutional neural network 1 is used for local channel spatial feature extraction, and convolutional neural network 2 is used for global channel spatial feature extraction.
[0078] It is evident that by fusing the features of the first channel and the features of the second channel, the spatial features of the target signal are obtained, which include both local and global features, thereby improving the accuracy of the spatial features in representing the target signal features.
[0079] S204. Extract the amplitude-time features of the target signal to obtain the time features of the target signal.
[0080] The amplitude-time feature can be the amplitude characteristics at different sampling times. The temporal features of the target signal are extracted.
[0081] Optionally, amplitude temporal features of the target signal can be extracted based on the visual conversion module. Specifically, the target signal is divided into multiple target signal blocks according to the time dimension, analogous to pixel blocks in an image, and each signal block is mapped to an embedding vector of a fixed dimension; the correlation weights between all signal blocks are calculated to capture the temporal features of the target signal in different dimensions.
[0082] In an optional embodiment, the step of extracting amplitude-time features from the target signal to obtain the time features of the target signal includes: extracting local amplitude-time features from the target signal to obtain a first amplitude feature; extracting global amplitude-time features from the target signal to obtain a second amplitude feature; and fusing the first amplitude feature and the second amplitude feature to obtain the time features of the target signal.
[0083] Local amplitude-time characteristics can capture the amplitude-time characteristics of a local part of the target signal. Global amplitude-time characteristics can focus on the amplitude-time characteristics of the target signal as a whole.
[0084] The target signal is divided into multiple target signal blocks, each with a small time span. Local amplitude time features are extracted from the target signal to obtain the first amplitude feature. The target signal is then divided into multiple target signal blocks, each with a larger time span, and global amplitude time features are extracted to obtain the second amplitude feature. The first and second amplitude features are then concatenated to obtain the time feature of the target signal.
[0085] In one specific embodiment, the target signal is input into two visual conversion modules with different target signal block division sizes. The two visual conversion modules are configured as follows:
[0086] Visual Transformation 1: Block segmentation dimension: [4, n / 2]; Number of attention heads: 4; Number of transformer layers: 2; Hidden layer dimension of multilayer perceptron: 4
[0087] Visual Transformation 2: Block segmentation dimension: [8, n]; Number of attention heads: 6; Number of transformer layers: 6; Hidden layer dimension of multilayer perceptron: 4
[0088] In this context, the block segmentation dimension is the same as the size of the target signal block. A larger block segmentation dimension results in a larger target signal block size and a longer time span. Here, n represents the time dimension of the input data. The number of attention heads is the number of parallel computation branches in the self-attention mechanism. The number of transformer layers is the number of core feature extraction layers. More transformer layers result in stronger feature abstraction capabilities and the ability to uncover more complex global dependencies. The hidden layer dimension of the multilayer perceptron is the number of hidden layer neurons in each transformer layer.
[0089] Visual transformation 1 is used for local amplitude time feature extraction, and visual transformation 2 is used for global amplitude time feature extraction.
[0090] It is evident that by fusing the first amplitude feature and the second amplitude feature, the temporal feature of the target signal is obtained, which includes both local and global features, thereby improving the accuracy of the temporal feature in representing the target signal features.
[0091] S205. The spatial and temporal features of the target signal are fused to obtain the signal features of the target signal.
[0092] Among them, spatial and temporal features are fused to obtain signal features that can reflect both the temporal variation pattern and the spatial distribution pattern of the target signal.
[0093] Optionally, the weights of spatial and temporal features are calculated separately, and the values of each element in the spatial and temporal features are summed in a weighted manner to obtain the signal features of the target signal.
[0094] In one specific embodiment, the target signal is input into a convolutional neural network with two convolutional kernels of different sizes to obtain a first channel feature C1 and a second channel feature C2; the target signal is input into a visual conversion module with two target signal blocks of different partition sizes to obtain a first amplitude feature V1 and a second amplitude feature V2.
[0095] The first channel feature C1, the second channel feature C2, the first amplitude feature V1, and the second amplitude feature V2 are concatenated along the feature dimension to form a global joint feature vector F. Then, a fully connected layer is used to learn an intermediate representation H, and another fully connected layer is used to map the intermediate representation H to four scalar scores E. The scalar scores E are normalized, transforming them into a probability distribution with values ranging from 0 to 1, thus obtaining the attention weights A for each branch. The attention weights A include the attention weight a1 for the first channel feature C1, the attention weight a2 for the second channel feature C2, the attention weight a3 for the first amplitude feature V1, and the attention weight a4 for the second amplitude feature V2.
[0096] The signal features F_fused of the target signal are obtained by weighting C1, C2, V1, and V2 using attention weight A.
[0097] F_fused=a1×C1+a2×C2+a3×V1+a4×V2
[0098] S206. Classify the signal features of the target signal to obtain the action recognition result.
[0099] The target signal is input into a classifier for classification, and the classification result is used as the action recognition result.
[0100] S207. Respond to the current business processing request based on the action recognition result.
[0101] The technical solution of this invention fuses multiple types of biomechanical signals to obtain a target signal, extracts the temporal and spatial features of the target signal, and can capture the temporal variation pattern and spatial correlation information of the action respectively. The temporal and spatial features are further fused to obtain signal features for classification. The signal features contain richer action feature expressions, effectively improving the accuracy and robustness of the classification task.
[0102] Figure 3 This is a schematic diagram of a human-computer interaction structure provided by an embodiment of the present invention. This embodiment of the present invention can be applied to the human-computer interaction of a certain application to be tested. The device can execute a human-computer interaction method, and the device can be implemented in hardware and / or software.
[0103] See Figure 3 The human-computer interaction device shown includes:
[0104] The signal acquisition module 301 is used to acquire multiple types of biomechanical signals transmitted through a wearable device worn by the user during the human-computer interaction process of the business terminal; the multiple types of biomechanical signals include at least: pressure signals and inertial signals;
[0105] Action recognition module 302 is used to perform action recognition on each of the biomechanical signals to obtain action recognition results;
[0106] The request response module 303 is used to respond to the current business processing request based on the action recognition result.
[0107] The technical solution of this invention broadens the signal dimensions of action recognition by acquiring multiple types of biomechanical signals transmitted by wearable devices. It acquires at least pressure signals and inertial signals for action recognition, solving the problem of low accuracy when using a single pressure signal for action recognition. Based on multiple types of biomechanical signals, more accurate action recognition results can be obtained. Using accurate action recognition results as the basis for business response, the business processing can be accurately matched with the user's real needs, reducing business execution errors and thus improving the reliability of human-computer interaction.
[0108] In an optional embodiment, the wearable device is communicatively connected to the service terminal, and the wearable device includes at least one electrode acquisition position.
[0109] In an optional embodiment, the signal acquisition module 301 includes:
[0110] An action demonstration unit is used to demonstrate multiple action videos and corresponding action instructions for each action video through the business terminal during the human-computer interaction process of the business terminal.
[0111] The signal acquisition unit is used to acquire multiple types of biomechanical signals transmitted through the wearable device worn by the user during the process of the user performing corresponding actions while following the action video; the user performs the corresponding actions to input the action instructions corresponding to the action video.
[0112] In an optional embodiment, the request-response module 303 includes:
[0113] The instruction determination unit is used to determine the user-inputted action instruction based on the action type in the action recognition result;
[0114] The request response unit is used to respond to the current business processing request according to the action instruction.
[0115] In an optional embodiment, the action recognition module 302 includes:
[0116] The target signal acquisition unit is used to fuse the biomechanical signals of each type to obtain the target signal;
[0117] A spatial feature extraction unit is used to extract channel spatial features from the target signal to obtain the spatial features of the target signal;
[0118] A time feature extraction unit is used to extract the amplitude time features of the target signal to obtain the time features of the target signal;
[0119] The feature fusion unit is used to fuse the spatial and temporal features of the target signal to obtain the signal features of the target signal;
[0120] An action recognition unit is used to classify the signal features of the target signal to obtain action recognition results.
[0121] In an optional embodiment, the target signal acquisition unit includes:
[0122] The target signal acquisition subunit is used to splice together the pressure signals and inertial signals of each type of biomechanical signal to obtain the target signal.
[0123] In an optional embodiment, the spatial feature extraction unit includes:
[0124] A spatial local feature extraction subunit is used to extract local channel spatial features from the target signal to obtain the first channel features;
[0125] A spatial full-local feature extraction subunit is used to extract global channel spatial features from the target signal to obtain second channel features;
[0126] The spatial feature fusion subunit is used to fuse the first channel features and the second channel features to obtain the spatial features of the target signal.
[0127] In an optional embodiment, the time feature extraction unit includes:
[0128] A local temporal feature extraction subunit is used to extract local amplitude temporal features from the target signal to obtain a first amplitude feature;
[0129] A global time feature extraction subunit is used to extract the global amplitude time features of the target signal to obtain the second amplitude feature;
[0130] The time feature fusion subunit is used to fuse the first amplitude feature and the second amplitude feature to obtain the time feature of the target signal.
[0131] The human-computer interaction device provided in the embodiments of the present invention can execute the human-computer interaction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the human-computer interaction method.
[0132] Figure 4 A schematic diagram of the structure of a human-computer interaction device 400 that can be used to implement an embodiment of the present invention is shown.
[0133] like Figure 4As shown, the human-computer interaction device 400 includes at least one processor 401 and a memory, such as a read-only memory 402 or a random access memory 403, communicatively connected to the at least one processor 401. The memory stores computer programs executable by the at least one processor. The processor 401 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 402 or loaded from the storage unit 408 into the random access memory 403. The random access memory 403 can also store various programs and data required for the operation of the human-computer interaction device 400. The processor 401, read-only memory 402, and random access memory 403 are interconnected via a bus 404. An input / output interface 405 is also connected to the bus 404.
[0134] Multiple components in the human-computer interaction device 400 are connected to the input / output interface 405, including: an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, optical disk, etc.; and a communication unit 409, such as a network card, modem, wireless transceiver, etc. The communication unit 409 allows the human-computer interaction device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0135] Processor 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 401 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 401 performs the various methods and processes described above, such as human-computer interaction methods.
[0136] In some embodiments, the human-computer interaction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed onto the human-computer interaction device 400 via read-only memory 402 and / or communication unit 409. When the computer program is loaded into random access memory 403 and executed by processor 401, one or more steps of the human-computer interaction method described above may be performed. Alternatively, in other embodiments, processor 401 may be configured to perform the human-computer interaction method by any other suitable means (e.g., by means of firmware).
[0137] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), complex programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0138] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0139] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on an operational detection device, which includes: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the human-computer interaction device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0141] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0142] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product within the cloud computing service system. This addresses the shortcomings of traditional physical hosts and virtual private servers, such as high management difficulty and weak business scalability.
[0143] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0144] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A human-computer interaction method, characterized in that, The method includes: During the human-computer interaction process at the business terminal, multiple types of biomechanical signals transmitted through the wearable device worn by the user are acquired; the multiple types of biomechanical signals include at least: pressure signals and inertial signals; Action recognition is performed on each of the biomechanical signals to obtain action recognition results; The system responds to the current business processing request based on the action recognition result.
2. The method according to claim 1, characterized in that, The process of performing action recognition on each of the biomechanical signals to obtain action recognition results includes: The biomechanical signals of each type are fused to obtain the target signal; Channel spatial feature extraction is performed on the target signal to obtain the spatial features of the target signal; The amplitude-time features of the target signal are extracted to obtain the time features of the target signal; The spatial and temporal features of the target signal are fused to obtain the signal features of the target signal; The signal features of the target signal are classified to obtain the action recognition result.
3. The method according to claim 2, characterized in that, The step of extracting channel spatial features from the target signal to obtain the spatial features of the target signal includes: Local channel spatial feature extraction is performed on the target signal to obtain the first channel feature; Global channel spatial feature extraction is performed on the target signal to obtain the second channel feature; The spatial features of the target signal are obtained by fusing the features of the first channel and the features of the second channel.
4. The method according to claim 2, characterized in that, The step of extracting the amplitude-time features of the target signal to obtain the time features of the target signal includes: The target signal is subjected to local amplitude time feature extraction to obtain the first amplitude feature; Global amplitude time feature extraction is performed on the target signal to obtain the second amplitude feature; The first amplitude feature and the second amplitude feature are fused to obtain the time feature of the target signal.
5. The method according to claim 2, characterized in that, The process of fusing biomechanical signals of each type to obtain the target signal includes: The target signal is obtained by splicing together the pressure signals and inertial signals of each type of biomechanical signal.
6. The method according to claim 1, characterized in that, The wearable device is communicatively connected to the service terminal, and the wearable device includes at least one electrode acquisition position.
7. The method according to claim 1, characterized in that, During the human-computer interaction process at the business terminal, multiple types of biomechanical signals transmitted through the wearable device worn by the user are acquired, including: During the human-computer interaction process of the business terminal, multiple action videos and corresponding action instructions for each action video are demonstrated through the business terminal. During the process of the user following the action video and performing corresponding actions, multiple types of biomechanical signals transmitted through the wearable device worn by the user are acquired, and the user performs the corresponding actions to input the action instructions corresponding to the action video; The step of responding to the current business processing request based on the action recognition result includes: The user-inputted action command is determined based on the action type in the action recognition result; Respond to the current business processing request according to the action instruction.
8. A human-computer interaction device, characterized in that, The device includes: The signal acquisition module is used to acquire multiple types of biomechanical signals transmitted through the wearable device worn by the user during the human-computer interaction process of the business terminal; the multiple types of biomechanical signals include at least: pressure signals and inertial signals; The action recognition module is used to perform action recognition on each of the biomechanical signals to obtain action recognition results; The request-response module is used to respond to the current business processing request based on the action recognition result.
9. A human-computer interaction device, characterized in that, The human-computer interaction device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the human-computer interaction method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the human-computer interaction method according to any one of claims 1-7.