Voice recognition remote control handle and voice and joystick collaboration method

By acquiring and fusing multi-scale features of voice and joystick operation, and combining them with historical interaction information for dynamic matching, the problem of recognizing collaborative control intentions of voice recognition and joystick operation is solved, thereby improving the accuracy and intelligence of control.

CN121963728APending Publication Date: 2026-05-01FOCUS CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FOCUS CLOUD COMPUTING CO LTD
Filing Date
2026-01-07
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, voice recognition and joystick operation are mostly processed independently, without fully considering the multi-scale temporal correlation between voice and joystick operation. This leads to deviations in the matching of semantic intent and operation logic, high error rates, and insufficient control stability in complex interaction scenarios.

Method used

By acquiring multi-scale semantic and temporal features of voice and joystick operation, hierarchical information complementarity and fusion are performed. Combined with historical interaction information, temporal constraints are determined, and control commands consistent with intent and operation are dynamically matched and generated.

Benefits of technology

It improves the accuracy and smoothness of human-computer interaction with the remote control, ensures the recognition of coordinated control intentions through voice and joystick operation, and enhances the level of intelligence in operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963728A_ABST
    Figure CN121963728A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition remote control handle and a voice and joystick collaboration method, and the method comprises the steps: extracting multi-scale semantic voice features from a voice input signal of the remote control handle in a human-computer interaction process; extracting operating rod operating characteristics of a multi-scale time sequence intention from a voice input signal of the remote control handle in a man-machine interaction process; determining collaborative fusion features of the remote control handle according to the semantic voice features and the operating rod operating features; determining constraint conditions of time sequence interaction between voice and operation of the remote control handle; and dynamically matching the semantic intention and the operation logic of the remote control handle in the man-machine interaction process through the constraint condition and the collaborative fusion feature, generating a control instruction with consistent intention and operation, and transmitting the control instruction to the remote control handle. By adopting the scheme of the invention, the cooperative control intention of the voice instruction and the joystick operation can be identified based on the multi-scale time sequence relevance between the voice and the joystick operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, and more specifically, to a voice recognition remote control and a voice and joystick coordination method. Background Technology

[0002] Human-computer interaction is an interdisciplinary field that studies the information transmission, instruction exchange, and feedback response between humans and computer systems through specific media and methods. It integrates knowledge from multiple disciplines such as computer science, psychology, design, and cognitive science, and uses diverse interactive interfaces such as keyboards, mice, touch screens, voice recognition, gesture control, and eye tracking to build efficient, natural, and user-friendly interactive modes.

[0003] With the development of intelligent control technology, voice recognition remote control handles are widely used in drones, smart cars, smart homes and other fields due to their ease of operation. The core requirement is to achieve coordinated linkage between voice commands and joystick operations to improve control accuracy. In existing technologies, voice recognition and joystick operations are mostly processed independently. Even if there are simple fusion solutions, they generally suffer from problems such as single feature extraction dimensions, insufficient consideration of the multi-scale temporal correlation between voice and joystick operations, and lack of constraint mechanisms based on historical interaction habits. This leads to phenomena such as semantic intent and operation logic mismatch, high error rate, and insufficient control stability in complex interaction scenarios, making it difficult to meet users' needs for precise and intelligent coordinated control. Therefore, how to recognize the coordinated control intent of voice commands and joystick operations based on the multi-scale temporal correlation between voice and joystick operations has become a problem faced by the industry. Summary of the Invention

[0004] This invention provides a voice recognition remote control handle and a voice and joystick coordination method, which can recognize the coordinated control intention of voice commands and joystick operations based on the multi-scale temporal correlation between voice and joystick operations.

[0005] In a first aspect, the present invention provides a voice recognition remote control handle voice and joystick coordination method, comprising the following steps: Acquire the voice input signals and operation input signals of the remote control during human-computer interaction; Multi-scale semantic speech features are extracted from the speech input signal, and multi-scale temporal intention joystick operation features are extracted from the operation input signal. The semantic speech features and the joystick operation features are fused together by hierarchical information to obtain the collaborative fusion features of local details and global context of the remote control during human-computer interaction. The constraints of the temporal interaction between the voice and operation of the remote control are determined based on the historical interaction information of the remote control during the human-computer interaction process. The semantic intent and operation logic of the remote control handle during human-computer interaction are dynamically matched using the constraints and the collaborative fusion features to generate control commands that are consistent with the intent and operation, and then the control commands are transmitted to the remote control handle.

[0006] In some embodiments, the voice input signal includes the original voice waveform signal and voice frame data in different time dimensions.

[0007] In some embodiments, extracting multi-scale semantic speech features from the speech input signal specifically includes: The voice input signal is preprocessed to obtain a preprocessed voice input signal; Multi-scale semantic analysis is performed on the preprocessed speech input signal to obtain multi-scale semantic speech features.

[0008] In some embodiments, extracting joystick operation features with multi-scale temporal intent from the operation input signal specifically includes: The operation input signal is preprocessed to obtain the preprocessed operation input signal; Multi-scale temporal intent analysis is performed on the preprocessed operation input signal to obtain the joystick operation characteristics of multi-scale temporal intent.

[0009] In some embodiments, the complementary fusion of hierarchical information between the semantic speech features and the joystick operation features to obtain the collaborative fusion features of local details and global context of the remote control during human-computer interaction specifically includes: The semantic speech features and the joystick operation features are hierarchically aligned to obtain hierarchically aligned features. The hierarchical alignment features are complementaryly fused to obtain complementary fused features for each level; Based on all complementary fusion features, the collaborative fusion features of local details and global context of the remote control during human-computer interaction are determined.

[0010] In some embodiments, determining the constraints for the temporal interaction between the voice and operation of the remote control handle based on the historical interaction information of the remote control handle during human-computer interaction specifically includes: Obtain the historical interaction information of the remote control during the human-computer interaction process; The historical interaction information is preprocessed to obtain preprocessed historical interaction information; Based on the preprocessed historical interaction information, a temporal analysis is performed on the correlation between the voice and operation of the remote control handle to obtain the constraints of the temporal interaction between the voice and operation of the remote control handle.

[0011] In some embodiments, dynamically matching the semantic intent and operational logic of the remote control handle during human-computer interaction through the constraints and the collaborative fusion features to generate control commands that are consistent with the intent and operation specifically includes: The global semantic intent of the remote control in the voice dimension and the long-sequence operation logic in the joystick dimension during the human-computer interaction process are extracted from the collaborative fusion features. The global semantic intent and the long-time operation logic are constrained and adjusted according to the constraints to obtain the adjusted global semantic intent and long-time operation logic. The adjusted global semantic intent and long-term operation logic are verified for consistency, and control instructions consistent with the intent and operation are generated.

[0012] In a second aspect, the present invention provides a voice recognition remote control handle, which includes a voice and joystick coordination unit, wherein the voice and joystick coordination unit includes: The acquisition module is used to acquire the voice input signals and operation input signals of the remote control during the human-computer interaction process; The processing module is used to extract multi-scale semantic speech features from the speech input signal and to extract multi-scale temporal intention joystick operation features from the operation input signal. The processing module is also used to perform hierarchical information complementary fusion of the semantic speech features and the joystick operation features to obtain the collaborative fusion features of local details and global context of the remote control handle during human-computer interaction. The processing module is further configured to determine the constraints of the temporal interaction between the voice and operation of the remote control handle based on the historical interaction information of the remote control handle during the human-computer interaction process. The execution module is used to dynamically match the semantic intent and operation logic of the remote control handle during human-computer interaction through the constraints and the collaborative fusion features, generate control commands that are consistent with the intent and operation, and transmit the control commands to the remote control handle.

[0013] Thirdly, the present invention provides a computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described voice recognition remote control handheld controller voice and joystick coordination method.

[0014] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described voice recognition remote control handheld device voice and joystick coordination method.

[0015] The technical solutions provided by the embodiments disclosed in this invention have the following beneficial effects: The voice recognition remote control and voice-and-joystick coordination method provided by this invention includes the following steps: acquiring voice and operation input signals; extracting multi-scale semantic voice features and multi-scale temporal intention joystick features; focusing on the core requirements of "multi-scale" and "temporal," this method captures semantic information at different levels in the voice signal and uncovers the temporal variation patterns of joystick operation, providing a precise feature carrier for the association and fusion of the two, and solving the problem that a single feature dimension is insufficient to reflect the coordination relationship; and the hierarchical information complementary fusion step, by integrating the local details and global context of the two types of features, effectively compensates for the limitations of a single feature and strengthens the voice-and-joystick coordination. The temporal correlation between joystick operations makes the fused features more closely aligned with the core meaning of collaborative control intentions. The step of determining temporal constraints based on historical interaction information constructs a temporal logic framework using past interaction patterns, providing a reasonable constraint basis for intent matching and avoiding misjudgments caused by deviating from historical operation logic, thus improving the temporal rationality of the matching. The step of dynamically matching and generating control commands, combined with constraints and fused features, achieves precise alignment between semantic intent and operational logic, ensuring a high degree of consistency between output commands and collaborative control intentions. Ultimately, this effectively solves the problem of collaborative control intent recognition, improving the accuracy, fluency, and intelligence level of the remote control's human-computer interaction. Using the above scheme, the collaborative control intent of voice commands and joystick operations can be recognized based on the multi-scale temporal correlation between voice and joystick operations. Attached Figure Description

[0016] Figure 1 This is an exemplary flowchart of a voice recognition remote control handheld device with voice and joystick coordination method according to some embodiments of the present invention; Figure 2 This is an exemplary flowchart illustrating the determination of speech features according to some embodiments of the present invention; Figure 3 This is an exemplary flowchart illustrating the determination of constraints according to some embodiments of the present invention; Figure 4 This is a schematic diagram of the structure of a voice and joystick coordination unit according to some embodiments of the present invention; Figure 5 This is a schematic diagram of the structure of a computer device that implements a voice and joystick coordination method for a voice recognition remote control, according to some embodiments of the present invention. Detailed Implementation

[0017] To better understand the technical solution of the present invention, the technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] refer to Figure 1 The figure is an exemplary flowchart of a voice recognition remote control handheld device's voice and joystick coordination method according to some embodiments of the present invention. This voice recognition remote control handheld device's voice and joystick coordination method mainly includes the following steps: In step 101, the voice input signal and operation input signal of the remote control handle during the human-computer interaction process are acquired.

[0019] It should be noted that, in this invention, the voice input signal refers to the audio electrical signal that carries the user's semantic command intent, collected by the sound pickup device built into the remote control during human-computer interaction, reflecting the user's control needs expressed by voice. The voice input signal includes the original voice waveform signal and voice frame data in different time dimensions. The operation input signal refers to the sensor electrical signal of the user's physical control action, reflecting the user's control tendency implemented by the joystick. The operation input signal includes the X / Y axial displacement of the joystick, the button pressing pressure value, the duration of a single operation, the displacement change rate during continuous operation, and the time sequence data of the operation action.

[0020] In step 102, multi-scale semantic speech features are extracted from the speech input signal, and multi-scale temporal intention joystick operation features are extracted from the operation input signal.

[0021] In some embodiments, reference Figure 2 As shown, this figure is an exemplary flowchart for determining speech features in some embodiments of the present invention. In this embodiment, the extraction of multi-scale semantic speech features from the speech input signal can be achieved by the following steps: In step 1021, the voice input signal is preprocessed to obtain a preprocessed voice input signal; In step 1023, multi-scale semantic analysis is performed on the preprocessed speech input signal to obtain multi-scale semantic speech features.

[0022] In specific implementation, the voice input signal is preprocessed to obtain a preprocessed voice input signal. Specifically, a first-order high-pass filter is first used for pre-emphasis processing to enhance the high-frequency components in the voice signal and compensate for the high-frequency attenuation of the signal during acquisition and transmission. Then, the pre-emphasized voice signal is divided into frames according to a frame length of 20ms and a frame shift of 10ms. A Hamming window is applied to each voice frame to reduce the spectral leakage effect at the frame edges. Finally, an adaptive filtering algorithm is used to eliminate background noise and electromagnetic interference in the environment to obtain the preprocessed voice input signal.

[0023] In addition, in specific implementation, multi-scale semantic analysis is performed on the preprocessed speech input signal to obtain multi-scale semantic speech features. Specifically, the multi-scale includes three scales: short-time, medium-time, and long-time. The short-time scale corresponds to the instantaneous acoustic characteristics of the speech signal, with a time range set to a standard frame length of 20ms. The medium-time scale corresponds to the continuous semantic unit characteristics at the phoneme level, with a time range set to cover the duration of a single phoneme, typically 50ms to 200ms. The long-time scale corresponds to the global semantic characteristics of a complete control command statement, with a time range set to cover the duration of the entire remote control voice command, typically 1s to 10s. In the short-time scale, for each frame of preprocessed speech signal, Mel-frequency cepstral coefficients, linear prediction coefficients, zero-crossing rate, and short-time energy are calculated. Mel-frequency cepstral coefficients are obtained by performing a Fast Fourier Transform on the speech frame to obtain the spectrum, then mapping it to a Mel-scale filter bank and calculating the logarithmic energy, finally obtaining the result through a Discrete Cosine Transform (DCT). These coefficients characterize the short-time spectral properties of the speech. Linear prediction coefficients are obtained by fitting the vocal tract model of the speech signal using linear prediction analysis to reflect the vocal tract resonance characteristics. The zero-crossing rate distinguishes between voiced and unvoiced sounds in the speech signal, and the short-time energy determines the phonic and nonvoiced segments of the speech signal. This process yields the frame... At the intermediate temporal scale, phoneme segmentation is performed on continuous frame-level acoustic features based on Hidden Markov Models (HMMs). The states of the HMMs are set as different phoneme categories of the speech signal. Frame-level acoustic features are used as observation sequences, and the model is trained using a forward-backward algorithm. Then, the trained model is decoded using the Viterbi algorithm to achieve phoneme boundary division of continuous speech frames and extract phoneme-level semantic unit features, reflecting the syllable composition and pronunciation rules of the speech signal. At the long temporal scale, semantic word segmentation based on maximum word length matching is first used to segment the text corresponding to the entire speech signal. The system matches the speech text with a pre-defined dictionary of control terms, traversing the dictionary in descending order of word length to complete word segmentation. The segmented results are then matched with a pre-defined control intent tag library, which covers common control commands for remote controllers such as forward, backward, left turn, and right turn. Valid semantic words are filtered out by setting a 90% matching threshold, and global semantic features at the sentence level are extracted to represent the user's overall control needs. Finally, feature data from short, medium, and long time scales are integrated to obtain multi-scale semantic speech features. Other methods can be used in other embodiments, which are not limited here.

[0024] It should be noted that the speech feature representation in this invention is a quantitative representation of the characteristics of speech input signals in different time dimensions. The speech features are hierarchically coordinated to achieve a full-dimensional representation of speech input signals from basic acoustic properties and local semantic composition to global control intentions.

[0025] In some embodiments, extracting the joystick operation features of multi-scale timing intent from the operation input signal can be achieved by the following steps: The operation input signal is preprocessed to obtain the preprocessed operation input signal; Multi-scale temporal intent analysis is performed on the preprocessed operation input signal to obtain the joystick operation characteristics of multi-scale temporal intent.

[0026] In specific implementation, the operation input signal is preprocessed to obtain a preprocessed operation input signal. Specifically, the operation input signal is processed using a moving average filtering algorithm with a length of 5 sampling points. The filter value of the current sampling point is calculated to remove glitch abnormal values ​​caused by sensor jitter. Then, the minimum-maximum normalization method is used to map sensor data of different dimensions such as displacement and pressure values ​​to the [0,1] interval, thereby eliminating the dimensional differences of output signals from different sensors and obtaining a preprocessed operation input signal with noise removal and dimensional uniformity.

[0027] In addition, in specific implementation, multi-scale temporal intent analysis is performed on the preprocessed operation input signal to obtain the joystick operation features of multi-scale temporal intent. Specifically: at the short time scale, for the preprocessed operation input signal at a single sampling moment, four sampling point-level physical parameter features are extracted: normalized value of joystick X-axis displacement, normalized value of joystick Y-axis displacement, normalized value of button pressing pressure, and operation trigger state, representing the basic attributes of instantaneous operation actions; at the medium time scale, with a sliding window duration of 500ms and a window step size of 200ms, five trend features are calculated on the continuous sampling data within the window: displacement change rate, pressure change gradient, duration of a single operation, consistency of operation direction, and number of consecutive operation triggers, representing the dynamic change law of staged operations; at the long time scale, the trend features of multiple consecutive sliding windows are used as an observation sequence and input to... The trained Hidden Markov Model (HMM) has a preset state set of five typical remote control operation intentions: forward, backward, left turn, right turn, and stop. During the model training phase, 1000 sets of manually labeled typical operation sequences are used as training data. The initial state probability, state transition probability, and observation probability of the model are optimized through 100 iterations of the forward-backward algorithm. Then, the Viterbi algorithm is used to search for the optimal path of the input observation sequence and output the probability value of each typical operation intention corresponding to the sequence. The operation intention with the highest probability value is extracted as the operation sequence-level intention tendency feature, which represents the user's long-term control intention. Finally, the short-term sampling point-level features, medium-term trend features, and long-term sequence-level intention tendency features are concatenated according to dimensions to form a 14-dimensional multi-scale temporal intention joystick operation feature. Other methods can be used in other embodiments, which are not limited here.

[0028] It should be noted that the joystick operation features in this invention represent the quantitative characterization of the control attributes of the joystick operation input signal in different time dimensions. The joystick operation features are hierarchically coordinated to achieve a full-dimensional characterization of the joystick operation input signal from instantaneous micro-actions, phased meso-level trends to long-term macro-level intentions.

[0029] In step 103, the semantic speech features and the joystick operation features are fused together by hierarchical information to obtain the collaborative fusion features of local details and global context of the remote control during human-computer interaction.

[0030] In some embodiments, the complementary fusion of hierarchical information between the semantic speech features and the joystick operation features to obtain the collaborative fusion features of local details and global context of the remote control during human-computer interaction can be achieved through the following steps: The semantic speech features and the joystick operation features are hierarchically aligned to obtain hierarchically aligned features. The hierarchical alignment features are complementaryly fused to obtain complementary fused features for each level; Based on all complementary fusion features, the collaborative fusion features of local details and global context of the remote control during human-computer interaction are determined.

[0031] In specific implementation, the semantic speech features and the joystick operation features are hierarchically aligned to obtain hierarchically aligned features. Specifically, the short-time frame-level acoustic features, mid-time phoneme-level semantic unit features, and long-time sentence-level global semantic features of the speech features are first matched one-to-one with the short-time sampling point-level physical parameter features, mid-time sliding window trend features, and long-time sequence-level intent tendency features of the joystick operation features. To address the issue of inconsistent dimensions between the two types of features at different scales, linear interpolation is used to expand the dimensions of the low-dimensional features, thereby increasing the dimensionality of the speech features at the short-time scale. The 12-dimensional Mel-Cepstral coefficients of the features and the 4-dimensional physical parameter features of the joystick features are uniformly interpolated to 16 dimensions. The 8-dimensional phoneme features of speech at the mid-time scale and the 5-dimensional trend features of the joystick are uniformly interpolated to 13 dimensions. The 6-dimensional global semantic features of speech at the long-time scale and the 5-dimensional intention tendency features of the joystick are uniformly interpolated to 11 dimensions. At the same time, according to the timestamp synchronization principle, the speech features and joystick operation features within the same time interval are bound together to obtain hierarchical alignment features with unified dimensions and temporal synchronization. Other methods can be used in other embodiments, which are not limited here.

[0032] Furthermore, in specific implementation, the hierarchical alignment features are complementaryly fused to obtain complementary fused features at each level. Specifically, during the complementary fusion of hierarchical alignment features, at the bottom-level fusion stage, a combination of feature splicing and min-max normalization is used for the short-term hierarchical alignment features to map the spliced ​​16-dimensional features to the [0,1] interval, eliminating the dimensional differences between acoustic features and physical parameter features, thus obtaining the bottom-level complementary fused features. At the mid-level fusion stage, for the mid-term hierarchical alignment features, attention weights are determined by calculating the Pearson correlation coefficients between the speech features and the joystick operation features across each dimension; that is, all correlation coefficients are normalized. The normalized values ​​are used as attention weights. The attention weights are multiplied by the corresponding feature dimensions and then the features are fused to highlight the feature components that are strongly related to the manipulation intent, thus obtaining the mid-level complementary fusion features. In the high-level fusion stage, for the hierarchical alignment features with long time scales, a 3×3 convolution kernel is used to perform convolution operations on the features to extract the spatial correlation information between the features. Then, the convolutional features are temporally smoothed with a sliding window of 1 second to mine the temporal dependency between global semantic features and manipulation intent features, thus obtaining the high-level complementary fusion features. Finally, the complementary fusion features of each level are obtained. Other methods can be used in other embodiments, which are not limited here.

[0033] Furthermore, in specific implementation, the collaborative fusion features of local details and global context of the remote control handle during human-computer interaction are determined based on all complementary fusion features. Specifically, the complementary fusion features of the bottom 16-dimensional, middle 13-dimensional, and top 11-dimensional layers are concatenated in dimensional order to obtain a 50-dimensional fusion feature matrix. Principal component analysis is then used to reduce the dimensionality of this 50-dimensional fusion feature matrix. First, a decentralization operation is performed on the feature matrix, then the covariance matrix of the decentralized feature matrix is ​​calculated. Finally, the eigenvalues ​​of the covariance matrix and their corresponding positive eigenvalues ​​are solved using an eigenvalue decomposition algorithm. Intersecting feature vectors, the first k feature vectors are selected in descending order of feature values ​​to construct a projection matrix. At the same time, the cumulative contribution rate of the first k feature values ​​is calculated, and the cumulative contribution rate threshold is set to 95%. Through iterative calculation, it is determined that the cumulative contribution rate satisfies ≥95% when k=25. The decentralized feature matrix is ​​multiplied by the projection matrix to obtain an m×25 dimensionality-reduced feature matrix. Each row of feature vectors in this matrix is ​​the collaborative fusion feature that takes into account both the local details of voice and joystick operation and the global context. Other methods can be used in other embodiments, which are not limited here.

[0034] It should be noted that the hierarchical alignment features in this invention represent the alignment features of the two original features, speech and joystick operation, at the same time scale, reflecting the consistency between the two types of features in terms of time sequence and dimension; the complementary fusion features represent the three types of features obtained by fusing the hierarchical alignment features at the bottom, middle, and top levels, which can be used to analyze the complementary correlation information of speech and joystick operation features at various scales; the collaborative fusion features represent the comprehensive feature information that takes into account both the local details and global context of speech and joystick operation, reflecting the complete intention tendency of collaborative voice and joystick operation, and providing core feature support for the subsequent dynamic matching of semantic intent and operation logic.

[0035] In step 104, the constraints of the timing interaction between the voice and operation of the remote control are determined based on the historical interaction information of the remote control during the human-computer interaction process.

[0036] In some embodiments, reference Figure 3 As shown, this figure is an exemplary flowchart of determining constraints in some embodiments of the present invention. In this embodiment, the constraints of the temporal interaction between the voice and operation of the remote control handle, based on the historical interaction information of the remote control handle during the human-computer interaction process, can be achieved by the following steps: In step 1041, the historical interaction information of the remote control handle during the human-computer interaction process is obtained; In step 1042, the historical interaction information is preprocessed to obtain preprocessed historical interaction information; In step 1043, a temporal analysis is performed on the correlation between the voice and operation of the remote control based on the preprocessed historical interaction information to obtain the constraints of the temporal interaction between the voice and operation of the remote control.

[0037] It should be noted that the historical interaction information in this invention represents the user's long-term voice control habits, the physical operation rules of the joystick, and the historical full-process status information of command execution when using the remote control. It reflects the inherent correlation pattern between voice command type and joystick operation type, the timing matching rules of voice and operation, and the differences in the execution effectiveness of different collaborative commands. The historical interaction information includes the original audio signal of the voice command issued by the user, the text content of the command after speech-to-text processing, the start and end timestamps of the voice command, the displacement of the joystick in the X / Y axis, the duration of continuous operation, the button pressing pressure value, the operation trigger and end timestamps, and the execution feedback results of the remote control device for the voice-operation collaborative command, and the execution feedback timestamp.

[0038] In specific implementation, the historical interaction information is preprocessed to obtain preprocessed historical interaction information. Specifically, invalid records are first removed using preset invalid data filtering rules, including removing voice data with empty or semantic noise text, unintentional shaking operation data where the joystick displacement and pressure values ​​are both zero and last for more than 5 seconds, and unresponsive feedback data caused by hardware failure. Then, the remaining valid data is timestamped and corrected using linear interpolation to correct the time deviation between voice commands and joystick operations caused by differences in sensor sampling frequencies, ensuring that the timestamps of voice commands and corresponding joystick operations are completely synchronized. Then, numerical data such as displacement and pressure values ​​are mapped to the [0,1] interval to complete normalization processing, resulting in preprocessed historical interaction information with uniform format, synchronized timing, and valid data. Other methods can also be used in other embodiments, which are not limited here.

[0039] In addition, in specific implementation, based on the preprocessed historical interaction information, a temporal analysis is performed on the association between the voice and operation of the remote control handle to obtain the constraints of the temporal interaction between the voice and operation of the remote control handle. That is: firstly, a prior algorithm in association rule mining is used to perform association analysis on the preprocessed data, setting the voice command type as the antecedent of the frequent itemset and the joystick operation type as the consequent, and calculating the support and confidence of each voice-operation correspondence. The support calculation formula is Support(A→B)=count(A∩B) / count(all), where count(A∩B) is the number of samples where voice command A and joystick operation B occur simultaneously, and count(all) is the total number of samples. The confidence calculation formula is Confidence(A→B)=count(A∩B) / count(A). The support threshold is set to 10%, and the confidence threshold is set to 85%, and the conditions that meet the requirements are selected. The system establishes a high-correlation voice-operation correspondence with a threshold requirement. Then, based on timestamp information, it calculates the time difference between the start time of the voice command and the trigger time of the joystick operation in each high-correlation voice-operation correspondence. A histogram statistical method is used to analyze the distribution of the time difference data, selecting the interval with the most concentrated time difference distribution as the initial effective timing interaction window. This initial effective timing interaction window is then optimized based on the command execution feedback results. Voice-operation correspondences with an execution success rate higher than 90% retain their effective timing interaction windows; correspondences with an execution success rate between 60% and 90% have their effective timing interaction windows narrowed by 20%; and correspondences with an execution success rate lower than 60% are directly eliminated. Finally, the high-correlation voice-operation correspondence, the optimized effective timing interaction window, and the execution feedback judgment rules are integrated to form the constraints for the timing interaction between the remote control's voice and operation. Other methods can be used in other embodiments, which are not limited here.

[0040] It should be noted that the constraints in this invention represent the matching priority of different voice command types and joystick operation types, the effective time difference range between voice commands and joystick operations, and the execution success rate threshold of different matching combinations. These constraints reflect the user's collaborative control habits of voice and joystick operations when using a remote control, the temporal correlation between voice commands and joystick operations, and the criteria for filtering effective collaborative commands and excluding unintentional misoperations. They can be used to provide a clear temporal judgment basis for the dynamic matching of voice and joystick operations.

[0041] In step 105, the semantic intent and operation logic of the remote control handle during human-computer interaction are dynamically matched using the constraints and the collaborative fusion features to generate control commands that are consistent with the intent and operation, and the control commands are transmitted to the remote control handle.

[0042] In some embodiments, the following steps can be used to dynamically match the semantic intent and operational logic of the remote control handle during human-computer interaction through the constraints and the collaborative fusion features to generate control commands that are consistent with the intent and operation: The global semantic intent of the remote control in the voice dimension and the long-sequence operation logic in the joystick dimension during the human-computer interaction process are extracted from the collaborative fusion features. The global semantic intent and the long-time operation logic are constrained and adjusted according to the constraints to obtain the adjusted global semantic intent and long-time operation logic. The adjusted global semantic intent and long-term operation logic are verified for consistency, and control instructions consistent with the intent and operation are generated.

[0043] In specific implementation, the global semantic intent of the remote control handle in the human-computer interaction process and the long-term operation logic of the joystick are extracted from the collaborative fusion features. That is, firstly, the feature components corresponding to the long-term scale speech semantics in the collaborative fusion features are matched with a preset control intent label library. The preset control intent label library includes five types of labels: forward, backward, left turn, right turn, and stop. By calculating the Euclidean distance between the feature components and the feature templates of each label, the label with the smallest distance is selected as the global semantic intent in the speech dimension. At the same time, the feature components corresponding to the long-term scale joystick operation in the collaborative fusion features are extracted, and the parameters of the average direction of the joystick's X / Y axial displacement, the steady-state trend of pressure change, and the proportion of continuous operation time are parsed out. Then, the parameters are retrieved from the previous extraction based on historical interaction information. A rule base for judgment is established, which contains parameter threshold ranges corresponding to typical control intentions such as forward, backward, left turn, right turn, and stop. Then, the extracted parameters are matched one by one with the thresholds of each intention in the rule base to calculate the parameter matching compliance. That is, first, the total number of core joystick parameters participating in the matching is counted, and each extracted parameter is compared with the corresponding threshold range of typical control intentions in the rule base. Parameters that match the threshold range are scored 1 point, and those that do not are scored 0 points. The sum of the scores of all parameters is divided by the total number of parameters, and the ratio obtained is the parameter matching compliance. The compliance threshold is set to 0.8 and the confidence threshold is set to 85%. Typical control intentions that simultaneously meet both thresholds and have the best matching degree are selected and finally determined as the long-term operation logic of the joystick dimension. Other methods can also be used in other embodiments, which are not limited here.

[0044] It should be noted that the global semantic intent of the voice dimension in this invention represents the complete control direction conveyed by the user through voice commands, reflecting the user's core control purpose in the human-computer interaction process; the long-term operation logic of the joystick dimension represents the operation tendency constituted by the displacement direction, pressure change trend, and operation duration of the joystick over a long period of time, reflecting the long-term and stable control behavior pattern implemented by the user through the joystick. Together, the two provide the core judgment basis for the dynamic matching of semantic intent and operation logic.

[0045] In addition, in specific implementation, the global semantic intent and the long-term operation logic are constrained and adjusted according to the constraints to obtain the adjusted global semantic intent and long-term operation logic. That is, based on the constraints, it is verified whether the time difference between the current voice command and the joystick operation is within the 0.5-2.0s effective time window defined by the constraints, and then it is verified whether the correlation between the two meets the threshold requirements of support of not less than 10% and confidence of not less than 85%. Matching combinations with time differences exceeding the effective window or correlation of not reaching the threshold are directly judged as invalid operations and eliminated. For combinations with time differences within the effective window but with slight deviations between operation logic and semantic intent, linear interpolation combined with historical matching data is used to fine-tune the displacement direction and pressure value of the joystick to make it closer to the direction of the global semantic intent of the voice, thus obtaining the adjusted global semantic intent and long-term operation logic. Other methods can also be used in other embodiments, which are not limited here.

[0046] Furthermore, in specific implementation, the adjusted global semantic intent and long-term operation logic are subjected to consistency verification to generate control instructions that are consistent with the intent and operation. That is, the cosine similarity algorithm is used to calculate the similarity between the adjusted feature vector and the effective instruction feature vector in the historical interaction information database, and the formula is as follows: Where A is the adjusted feature vector and B is the historical valid feature vector, the similarity threshold is set to 85%, and historical valid instructions with similarity higher than the threshold are selected as the matching benchmark. If the adjusted global semantic intent and the long-term operation logic are completely consistent, the verification is deemed to be successful. Then, according to the preset control instruction encoding rules, such as the serial communication instruction format: start bit + intent encoding + operation parameter + check bit + stop bit, the verified semantic intent and operation logic are converted into standardized control instructions that can be recognized by the remote control device. Other methods can also be used in other embodiments, which are not limited here.

[0047] It should be noted that after the control command is transmitted to the remote control, the main control chip built into the remote control will quickly parse the semantic intent code and joystick operation parameters in the command. On the one hand, it drives the status indicator component of the remote control to output the corresponding prompt signal, so that the user can intuitively know that the command has been successfully received and taken effect. On the other hand, it forwards the parsed control command to the controlled device, drives the controlled device to perform the action matching the user's intention, and collects the real-time operating status of the controlled device and feeds it back to the remote control, forming a closed-loop control process of "command transmission-parsing-execution-feedback", ensuring that the user can accurately and stably remotely control the controlled device through voice and joystick coordination.

[0048] Furthermore, in another aspect of the present invention, in some embodiments, the present invention provides a voice and joystick coordination unit, with reference to... Figure 4The figure is a schematic diagram of the structure of a voice and joystick coordination unit 400 according to some embodiments of the present invention. The voice and joystick coordination unit 400 includes: an acquisition module 401, a processing module 402, and an execution module 403, which are described below: The acquisition module 401 in this invention is mainly used to acquire the voice input signal and operation input signal of the remote control handle during the human-computer interaction process; Processing module 402, in this invention, is used to extract multi-scale semantic speech features from the speech input signal and to extract multi-scale temporal intention joystick operation features from the operation input signal. It should be noted that the processing module 402 in this invention is also used to perform complementary fusion of hierarchical information on the semantic speech features and the joystick operation features to obtain the collaborative fusion features of local details and global context of the remote control handle during human-computer interaction. Additionally, it should be noted that the processing module 402 in this invention is also used to determine the constraints of the temporal interaction between the voice and operation of the remote control handle based on the historical interaction information of the remote control handle during the human-computer interaction process. The execution module 403 in this invention is mainly used to dynamically match the semantic intent and operation logic of the remote control handle in the human-computer interaction process through the constraints and the collaborative fusion features, generate control instructions that are consistent with the intent and operation, and transmit the control instructions to the remote control handle.

[0049] In addition, the present invention also provides a computer device, the computer device including a memory and a processor, the memory storing code, the processor being configured to acquire the code and execute the above-described voice recognition remote control handheld voice and joystick coordination method.

[0050] In some embodiments, reference Figure 5 This figure is a schematic diagram of the structure of a computer device implementing a voice recognition remote control handheld device and a joystick coordination method according to some embodiments of the present invention. The voice recognition remote control handheld device and joystick coordination method in the above embodiments can be achieved through… Figure 5 The computer device shown is used to implement this, and the computer device 500 includes at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504.

[0051] Processor 501 can be a general-purpose central processing unit (CPU) or an application-specific integrated circuit (ASIC).

[0052] The communication bus 502 can be used to transmit information between the aforementioned components.

[0053] Memory 503 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory 503 may exist independently and be connected to processor 501 via communication bus 502. Memory 503 may also be integrated with processor 501.

[0054] The memory 503 stores program code for executing the present invention, and its execution is controlled by the processor 501. The processor 501 executes the program code stored in the memory 503. The program code may include one or more software modules. The method used in the above embodiments can be implemented by the processor 501 and one or more software modules in the program code in the memory 503.

[0055] Communication interface 504 uses any transceiver-like device to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0056] In a specific implementation, as one example, a computer device may include multiple processors, each of which may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, a processor may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).

[0057] The aforementioned computer device can be a general-purpose computer device or a special-purpose computer device. In specific implementations, the computer device can be a desktop computer, a portable computer, a network server, a handheld digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. This embodiment of the invention does not limit the type of computer device.

[0058] In addition, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described voice recognition remote control handheld device voice and joystick coordination method.

[0059] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0060] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A voice recognition remote control method for voice and joystick coordination, characterized in that, Includes the following steps: Acquire the voice input signals and operation input signals of the remote control during human-computer interaction; Multi-scale semantic speech features are extracted from the speech input signal, and multi-scale temporal intention joystick operation features are extracted from the operation input signal. The semantic speech features and the joystick operation features are fused together by hierarchical information to obtain the collaborative fusion features of local details and global context of the remote control during human-computer interaction. The constraints of the temporal interaction between the voice and operation of the remote control are determined based on the historical interaction information of the remote control during the human-computer interaction process. The semantic intent and operation logic of the remote control handle during human-computer interaction are dynamically matched using the constraints and the collaborative fusion features to generate control commands that are consistent with the intent and operation, and then the control commands are transmitted to the remote control handle.

2. The method as described in claim 1, characterized in that, The voice input signal includes the original voice waveform signal and voice frame data in different time dimensions.

3. The method as described in claim 1, characterized in that, Extracting multi-scale semantic speech features from the speech input signal specifically includes: The voice input signal is preprocessed to obtain a preprocessed voice input signal; Multi-scale semantic analysis is performed on the preprocessed speech input signal to obtain multi-scale semantic speech features.

4. The method as described in claim 1, characterized in that, Extracting the joystick operation features with multi-scale temporal intent from the operation input signal specifically includes: The operation input signal is preprocessed to obtain the preprocessed operation input signal; Multi-scale temporal intent analysis is performed on the preprocessed operation input signal to obtain the joystick operation characteristics of multi-scale temporal intent.

5. The method as described in claim 1, characterized in that, The semantic speech features and the joystick operation features are fused hierarchically to obtain the collaborative fusion features of local details and global context of the remote control during human-computer interaction. Specifically, this includes: The semantic speech features and the joystick operation features are hierarchically aligned to obtain hierarchically aligned features. The hierarchical alignment features are complementaryly fused to obtain complementary fused features for each level; Based on all complementary fusion features, the collaborative fusion features of local details and global context of the remote control during human-computer interaction are determined.

6. The method as described in claim 1, characterized in that, The constraints for determining the temporal interaction between the voice and operation of the remote control handle based on the historical interaction information of the remote control handle during human-computer interaction specifically include: Obtain the historical interaction information of the remote control during the human-computer interaction process; The historical interaction information is preprocessed to obtain preprocessed historical interaction information; Based on the preprocessed historical interaction information, a temporal analysis is performed on the correlation between the voice and operation of the remote control handle to obtain the constraints of the temporal interaction between the voice and operation of the remote control handle.

7. The method as described in claim 1, characterized in that, The dynamic matching of semantic intent and operational logic of the remote control handle during human-computer interaction, based on the constraints and collaborative fusion features, to generate control commands that align intent and operation specifically includes: The global semantic intent of the remote control in the voice dimension and the long-sequence operation logic in the joystick dimension during the human-computer interaction process are extracted from the collaborative fusion features. The global semantic intent and the long-time operation logic are constrained and adjusted according to the constraints to obtain the adjusted global semantic intent and long-time operation logic. The adjusted global semantic intent and long-term operation logic are verified for consistency, and control instructions consistent with the intent and operation are generated.

8. A voice recognition remote control, comprising a voice and joystick coordination unit, characterized in that, The voice and joystick coordination unit includes: The acquisition module is used to acquire the voice input signals and operation input signals of the remote control during the human-computer interaction process; The processing module is used to extract multi-scale semantic speech features from the speech input signal and to extract multi-scale temporal intention joystick operation features from the operation input signal. The processing module is also used to perform hierarchical information complementary fusion of the semantic speech features and the joystick operation features to obtain the collaborative fusion features of local details and global context of the remote control handle during human-computer interaction. The processing module is further configured to determine the constraints of the temporal interaction between the voice and operation of the remote control handle based on the historical interaction information of the remote control handle during the human-computer interaction process. The execution module is used to dynamically match the semantic intent and operation logic of the remote control handle during human-computer interaction through the constraints and the collaborative fusion features, generate control commands that are consistent with the intent and operation, and transmit the control commands to the remote control handle.

9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing code, and the processor being configured to retrieve the code and execute the voice and joystick coordination method for a voice recognition remote control as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the voice and joystick coordination method for a voice recognition remote control as described in any one of claims 1 to 7.