An electromyography bracelet-based interaction system and method

CN122837633APending Publication Date: 2026-09-29HISENSE ELECTRONICS TECH SHENZHEN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611007014.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]本申请提供了一种基于肌电手环的交互系统及方法,以解决肌电交互系统无法在全场景下实现高精度连续双手手语识别的问题

Benefits of technology

[0007]上述技术方案具有以下有益效果或优点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837633A_ABST
    Figure CN122837633A_ABST
Patent Text Reader

Abstract

The application discloses an electromyography bracelet-based interaction system and method. The interaction system comprises a first electromyography bracelet, a second electromyography bracelet and a processing terminal. The first and second electromyography bracelets respectively collect surface electromyography signals generated by forearm muscle groups on both sides of a target user, and the processing terminal establishes wireless communication connections with the two electromyography bracelets. The processing terminal synchronously acquires two channels of surface electromyography signals, and performs feature fusion on the two channels of surface electromyography signals to generate fused feature data. Then, a preset spatio-temporal fusion recognition model is called to perform sign language recognition on the fused feature data, and corresponding sign language text information is obtained; and the sign language text information is converted into a voice signal and played. The system can improve the continuous sign language recognition accuracy by cooperatively collecting complete double-hand sign language action features through bilateral electromyography signals, realize end-to-end conversion from sign language to voice, and provide a convenient and natural voice interaction experience for sign language users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wearable devices and human-computer interaction technology, and in particular to an interaction system and method based on an electromyography (EMG) wristband. Background Technology

[0002] Electromyography (EMG) interaction systems are intelligent systems that analyze and convert motor intentions by collecting electromyographic (EMG) signals from the human body. Wearable interaction solutions centered around EMG wristbands can acquire muscle bioelectrical signals through devices worn on the limbs, analyze the user's motor intentions, and can be applied to various scenarios such as barrier-free communication, intelligent device control, and sports rehabilitation monitoring to meet diverse and personalized user interaction needs.

[0003] Electromyography (EMG) interaction systems support sign language recognition and speech conversion. They employ different signal acquisition architectures, acquiring sign language movement data through corresponding sensing units, and then processing this data using algorithms to convert sign language into speech. In related technologies, EMG wristbands for EMG interaction systems are divided into visual acquisition devices and single-arm EMG acquisition devices. Visual acquisition devices use a camera as the core sensing component, acquiring optical images of hand movements and then recognizing sign language words through image feature extraction and dictionary matching algorithms. Single-arm EMG acquisition devices use an EMG wristband worn on one forearm, acquiring local muscle electrical signals through a surface electrode array. After classification algorithms recognize gesture commands, these signals are transmitted to the terminal to drive the speech synthesis module to output the corresponding speech.

[0004] However, visual acquisition devices rely on optical imaging, which is easily affected by lighting conditions and hand obstruction, and continuous image acquisition poses a risk of privacy leakage. Single-arm electromyography (EMG) acquisition devices can only acquire muscle signals from one side of the limb and cannot capture the characteristics of coordinated movement of both arms. They are unable to cover complex sign language expressions completed by both hands, resulting in EMG interaction systems being unable to achieve high-precision continuous biphasic sign language recognition in all scenarios. This makes it difficult to meet the daily natural communication needs of hearing-impaired people and affects the user experience. Summary of the Invention

[0005] This application provides an interactive system and method based on an electromyography (EMG) wristband to solve the problem that EMG interactive systems cannot achieve high-precision continuous biphasic sign language recognition in all scenarios.

[0006] In a first aspect, this application provides an interactive system based on an electromyography (EMG) wristband, comprising: The first electromyography (EMG) wristband is configured to collect a first surface electromyography (EMG) signal, which is generated by the first forearm muscle group of the target user. The second electromyography (EMG) bracelet is configured to acquire a second surface electromyography (EMG) signal, which is generated by the second forearm muscle group of the target user. The processing terminal establishes wireless communication connections with the first and second electromyography (EMG) wristbands, respectively, and the processing terminal is configured as follows: Simultaneously acquire the first surface electromyography signal and the second surface electromyography signal; The first surface electromyography (EMG) signal and the second surface EMG signal are fused to generate fused feature data; A preset spatiotemporal fusion recognition model is invoked to perform sign language recognition on the fused feature data in order to obtain the corresponding sign language text information; The sign language text information is converted into a speech signal, and the speech signal is played.

[0007] The above technical solution has the following beneficial effects or advantages: By simultaneously acquiring and fusing electromyographic (EMG) signals from both forearms, the system fully captures the EMG features of coordinated hand gestures in sign language, solving the problem that single-arm acquisition cannot cover coordinated hand gestures and that recognition accuracy is insufficient. The spatiotemporal fusion recognition model takes into account the spatial correlation of muscles in multiple channels and the temporal dependence of movements, which can improve the recognition accuracy of continuous sign language. The system realizes end-to-end conversion of EMG signals into sign language text and then into speech output, providing sign language users with low-latency and highly natural voice interaction capabilities, thereby improving the convenience of sign language communication.

[0008] In some embodiments of this application, before the processing terminal performs feature fusion on the first surface electromyography signal and the second surface electromyography signal, it is further configured to: The first surface electromyography (EMG) signal and the second surface EMG signal are filtered to obtain two pre-processed EMG signals with interference removed. The two preprocessed electromyographic signals are framed using a sliding window with a preset window length and sliding step size to obtain two framed electromyographic signals. Extract the time-domain and frequency-domain features corresponding to each frame of the two-channel segmented electromyography signals; The time-domain features and the frequency-domain features are concatenated to obtain two initial feature vectors. Dimensionality reduction is performed on the two initial feature vectors to obtain single-path feature sequences with unified dimensions. The feature fusion is performed based on the single-path feature sequence.

[0009] The above technical solution has the following beneficial effects or advantages: Through a multi-level preprocessing chain of filtering, framing, feature extraction, and dimensionality reduction, environmental noise and power frequency interference are filtered out, while retaining the effective frequency band of the electromyography (EMG) signal. The extraction method that combines time-domain and frequency-domain features can comprehensively characterize the EMG signal from two dimensions: amplitude variation and frequency distribution, providing richer features for subsequent recognition. Dimensionality reduction compresses redundant feature dimensions, reducing the computational load of subsequent models while retaining core information, thereby improving the recognition response speed. Parallel preprocessing of the two signals before fusion ensures the integrity and temporal consistency of the bilateral motion features.

[0010] In some embodiments of this application, before the processing terminal performs the feature fusion based on the single-path feature sequence, it is further configured to: Read the single-frame features from the two single-channel feature sequences as the current frame features; Based on the amplitude changes of the current frame features, the limb movement state of the target user is detected; If the limb movement state is detected to be in a static state, and the duration of the static state reaches a first preset duration, then the static feature data for a second preset duration is resampled, and the baseline statistical parameters are reset; wherein, the second preset duration is less than the first preset duration, and the baseline statistical parameters are used to characterize the statistical quantity of the target user's electromyographic feature baseline level in a static state; If the target user is detected to be in a normal motion state, the baseline statistical parameters are updated by sliding through a feature buffer of a preset length; Based on the baseline statistical parameters, dynamic calibration is performed on the single-frame features in the two single-channel feature sequences to obtain the two calibrated feature sequences. The feature fusion is performed based on the two calibrated feature sequences; When the processing terminal updates the baseline statistical parameters by sliding through a feature buffer of a preset length, it is also configured to: Monitor the average confidence accuracy of sign language text information; If the average confidence accuracy is lower than a preset threshold, the update rate of the baseline statistical parameters is increased.

[0011] The above technical solution has the following beneficial effects or advantages: By detecting motion, baseline parameters are adaptively updated. The baseline is reset in a static state to ensure calibration accuracy, and the slow drift of the adaptive signal is updated during normal motion to compensate for the signal baseline offset caused by muscle fatigue and changes in skin impedance due to sweating. The baseline update rate is dynamically adjusted in combination with the sign language recognition accuracy. The update speed is accelerated when the recognition accuracy decreases, taking into account both the sensitivity and stability of the calibration. This can improve the recognition accuracy during long-term wear and use, and overcome the defect of accuracy decay of static baseline solutions after long-term use.

[0012] In some embodiments of this application, the processing terminal invokes a preset spatiotemporal fusion recognition model to perform sign language recognition on the fused feature data, specifically configured as follows: The spatial dimension features of the fused feature data are extracted by the one-dimensional convolutional neural network module in the spatiotemporal fusion recognition model to obtain a spatial feature sequence. The spatial feature sequence is modeled using the bidirectional long short-term memory network module in the spatiotemporal fusion recognition model to obtain the temporal feature sequence. The attention weights of the temporal feature sequence are calculated using the attention mechanism module in the spatiotemporal fusion recognition model. The temporal feature sequence is weighted and summed based on the attention weights to obtain the context feature vector; The context feature vector is input into the fully connected classification layer of the spatiotemporal fusion recognition model to obtain the probability distribution of sign language words; The corresponding sign language text information is determined based on the probability distribution.

[0013] The above technical solution has the following beneficial effects or advantages: One-dimensional convolutional neural networks are used to extract spatial dimensional features, which can uncover the synergistic relationships between different acquisition channels and different muscle groups at the same time. Bidirectional long short-term memory networks are used to model temporal dependencies, which can capture the logical and transitional relationships between consecutive sign language movements. Attention mechanisms can improve the recognition weight of core movements by weighting and strengthening key action frames and weakening redundant frames. Four-layer spatiotemporal fusion can improve the recognition accuracy and robustness of complex two-handed coordinated sign language.

[0014] In some embodiments of this application, the spatiotemporal fusion recognition model further includes a connectionist time classification module, which is configured with a preset sign language vocabulary list, which includes sign language vocabulary entries and blank symbols used to identify vocabulary boundaries. The processing terminal determines the corresponding sign language text information based on the probability distribution, specifically configured as follows: For continuously input fused feature data, the connectionist temporal classification module inserts blank symbols into a preset sign language vocabulary list and determines the alignment path between each frame feature in the fused feature data and the sign language vocabulary entry and the blank symbol; Calculate the sum of probabilities of the alignment paths to obtain the corresponding word probability sequence; Decoding is performed on the word probability sequence to obtain an initial decoding result containing repeating characters and whitespace symbols; The initial decoding result is subjected to deduplication and whitespace removal processing to obtain discrete word fragments; Based on the location of the blank symbol, word boundaries are detected, and the discrete word fragments are spliced ​​into a continuous word sequence. The sign language text information is generated based on the continuous word sequence.

[0015] The above technical solution has the following beneficial effects or advantages: Automatic word segmentation for continuous sign language is achieved based on a connectionist temporal classification mechanism. This eliminates the need to pre-label the start and end positions of each word, reducing the labeling cost of training data and adapting to natural and coherent sign language expression scenarios. By marking word boundaries with blank symbols and combining deduplication and splicing processes, continuous feature sequences can be accurately converted into discrete sign language word sequences, achieving precise decomposition from continuous action flow to independent words, enabling the system to support complete continuous sign language recognition.

[0016] In some embodiments of this application, the processing terminal is configured with a lexical mapping library, a syntax rule engine, and a language model; The processing terminal generates the sign language text information based on the continuous word sequence, specifically configured as follows: The continuous vocabulary sequence is transformed word by word using the lexical mapping library to obtain a natural language vocabulary string; The grammar rule engine is used to detect the word order structure of the natural language vocabulary string to obtain the word order structure detection result. Based on the word order structure detection results, word order is reorganized to obtain the initial sentence; The language model is used to perform probability optimization on the initial statement, and the context disambiguation of the initial statement is performed in combination with the dialogue history of a preset window length to obtain the sign language text information.

[0017] The above technical solution has the following beneficial effects or advantages: Lexical mapping facilitates the conversion of sign language vocabulary to natural language vocabulary, bridging the gap between sign language expression and everyday language. The grammar rule engine corrects word order differences between sign language and natural language, adjusting the unique expression order of sign language to a subject-verb-object word order that conforms to everyday communication habits. By combining language models and dialogue history for context disambiguation, the reasonable semantics of polysemous words can be distinguished, outputting fluent and natural text content, thereby ensuring the semantic accuracy and intelligibility of subsequent speech synthesis.

[0018] In some embodiments of this application, the processing terminal converts the sign language text information into a speech signal, specifically configured as follows: Detect the network connection status of the processing terminal; If the network connection is active, the cloud-based speech synthesis service is invoked to process the sign language text information into speech and generate the speech signal. If the network connection is offline, the locally integrated lightweight speech synthesis engine is invoked to perform speech synthesis processing on the sign language text information to generate the speech signal; In response to a user-inputted speech rate adjustment command, the playback speech rate of the voice signal is set within a preset adjustment range, and the voice signal is played according to the playback speech rate.

[0019] The above technical solution has the following beneficial effects or advantages: The adaptive dual-mode speech synthesis architecture can dynamically switch synthesis paths according to network status. When connected to the network, it calls cloud services to output highly natural speech, while when offline, it enables a local lightweight engine to ensure the availability of basic functions, taking into account both speech quality and usability in all scenarios. It supports users to customize the playback speed, which can adapt to users of different ages and hearing habits, improving the adaptability and comfort of the interactive experience.

[0020] In some embodiments of this application, the first electromyography (EMG) wristband and the second EMG wristband are further configured as follows: The surface electromyography (EMG) signal is converted into a digital format signal, wherein the surface EMG signal is either the first surface EMG signal or the second surface EMG signal. The digital format signal is synchronously transmitted to the processing terminal according to the preset sampling frequency.

[0021] The above technical solution has the following beneficial effects or advantages: The electromyography (EMG) wristband completes the conversion of the original analog signal to a digital signal, improving the anti-interference capability during signal transmission and ensuring the stability of data transmission; it synchronously transmits two signals at a unified sampling frequency, providing a consistent time reference for feature fusion and timing modeling on the terminal side.

[0022] In some embodiments of this application, the processing terminal further establishes a communication connection with the server; the processing terminal is further configured to: The first surface electromyography (EMG) signal, the second surface EMG signal, and the sign language text information are uploaded to the server; Receive updated model parameters from the server based on the first surface electromyography signal, the second surface electromyography signal, and the sign language text information; The spatiotemporal fusion recognition model is incrementally updated based on the updated model parameters.

[0023] The above technical solution has the following beneficial effects or advantages: End-to-cloud collaboration can leverage cloud computing power to complete model training and optimization, breaking through the computing power limitations of local devices and continuously iterating to improve model recognition accuracy. Incremental updates do not require a full replacement of model files, resulting in fast updates, low terminal resource consumption, and seamless model upgrades for users. Uploading local data to the cloud for training allows the model to continuously adapt to the electromyographic features of more users and a richer vocabulary of sign language, thereby expanding the system's applicability and recognition capabilities.

[0024] Secondly, this application also provides an interaction method based on an electromyography (EMG) wristband, comprising: Simultaneously acquire a first surface electromyography (EMG) signal and a second surface EMG signal; the first EMG signal is generated by the first forearm muscle group of the target user, and the second EMG signal is generated by the second forearm muscle group of the target user; The first surface electromyography (EMG) signal and the second surface EMG signal are fused to generate fused feature data; A preset spatiotemporal fusion recognition model is invoked to perform sign language recognition on the fused feature data in order to obtain the corresponding sign language text information; The sign language text information is converted into a speech signal, and the speech signal is played.

[0025] The above technical solution has the following beneficial effects or advantages: It can cover the feature information of coordinated sign language movements of both hands, with high recognition accuracy; the spatiotemporal fusion recognition model takes into account the feature extraction of spatial and temporal dimensions, and has strong robustness in recognizing continuous sign language movements; the end-to-end processing of the entire link can quickly convert the user's sign language movements into speech output, providing a convenient and natural daily communication experience for the hearing-impaired community. Attached Figure Description

[0026] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 A schematic diagram illustrating application scenarios of an interactive system based on electromyography (EMG) wristband provided in some embodiments of this application; Figure 2 A schematic diagram of the hardware configuration of an electromyography (EMG) wristband provided in some embodiments of this application; Figure 3 This application provides a schematic diagram of the software configuration of an interactive system for some embodiments. Figure 4 A flowchart illustrating an interaction method based on an electromyography (EMG) wristband, provided for some embodiments of this application; Figure 5This is a schematic diagram illustrating the signal preprocessing and feature extraction process provided in some embodiments of this application; Figure 6 A schematic diagram illustrating the dynamic baseline adaptive calibration process provided for some embodiments of this application; Figure 7 A schematic diagram of the hierarchical architecture of a sign language recognition model provided in some embodiments of this application; Figure 8 A schematic diagram illustrating the sign language to natural language mapping process provided in some embodiments of this application; Figure 9 A schematic diagram of the adaptive speech synthesis process provided in some embodiments of this application; Figure 10 A schematic diagram illustrating the process of updating the end-to-cloud collaborative model provided in some embodiments of this application; Figure 11 Interaction timing diagrams of an electromyography (EMG) wristband provided for some embodiments of this application. Detailed Implementation

[0028] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0029] In this application embodiment, the interactive system based on electromyography (EMG) wristbands generally refers to a wearable interactive system with the ability to acquire EMG signals, process data, and output results. For example, the interactive system includes, but is not limited to, sign language recognition and conversion systems, gesture control systems, and motion analysis systems, and can be adapted to various application scenarios such as barrier-free communication, smart device control, and rehabilitation monitoring.

[0030] Figure 1 These are schematic diagrams illustrating application scenarios of the interactive system provided in some embodiments of this application. For example... Figure 1 As shown, the user can wear the electromyography (EMG) bracelet 100 on the left and right limbs respectively. The EMG bracelet 100 establishes a connection with the processing terminal 200 via wireless communication, and transmits the collected EMG signals to the processing terminal 200 for processing and output. For example, the processing terminal 200 can be a cloud processing engine, smartphone, tablet computer, embedded processing device, smart display device (such as a TV), or other software or hardware device.

[0031] The processing terminal 200 can serve as a data processing and interaction terminal, used to receive electromyographic data transmitted by the electromyography (EMG) bracelet 100 and perform human-computer interaction functions such as signal analysis, intent recognition, and result conversion. In some embodiments, the processing terminal 200 can install supporting software applications with the EMG bracelet 100 to achieve bidirectional data transmission and control operations through network communication protocols. The processing results can also be output in the form of voice, text, control commands, etc., to achieve corresponding interactive functions.

[0032] like Figure 1 The diagram also shows that the processing terminal 200 communicates with the server 300 via various communication methods. The processing terminal 200 can connect via a local area network (LAN), wireless local area network (WLAN), and other networks to upload electromyography (EMG) data to the server 300 for cloud computing, model training, data storage, and other processing.

[0033] The interactive system based on the EMG bracelet 100 can provide basic EMG signal acquisition and motion recognition functions, and can also provide extended intelligent functions such as voice conversion, device control, and data statistics, including but not limited to sign language voice conversion, gesture control, and motion status monitoring.

[0034] Figure 2 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of the Zhongmyoelectric Bracelet 100.

[0035] In some embodiments, the electromyography (EMG) wristband 100 includes a left EMG wristband and a right EMG wristband, which are symmetrically configured to collect EMG signals from the corresponding side of the limb. Each EMG wristband 100 may include at least one of the following: an electrode acquisition module 110, a signal conditioning module 120, a wireless communication module 130, a controller 140, a power supply 150, a status feedback module 160, and a memory.

[0036] In some embodiments, the electrode acquisition module 110 is used to acquire electromyographic signals from the human body surface. For example, the electrode acquisition module 110 may include a multi-channel surface electrode array, which acquires bioelectrical signals generated by muscle activity through contact with the skin, covering the signal acquisition range of the corresponding muscle groups of the limb.

[0037] In some embodiments, the signal conditioning module 120 is used to amplify, filter, and perform analog-to-digital conversion on the acquired raw electromyographic signals. For example, the signal conditioning module 120 may include a front-end amplification circuit, a filtering circuit, and an analog-to-digital conversion unit to convert the analog electromyographic signals into digital signals and output them to the controller for further processing.

[0038] In some embodiments, the wireless communication module 130 is a component for communicating with a processing terminal or other devices according to various communication protocol types. The myoelectric bracelet 100 may be equipped with corresponding communication modules depending on the supported communication methods. For example, when the myoelectric bracelet 100 supports Bluetooth connection communication, it is equipped with a wireless communication module 130 containing Bluetooth functionality; when it supports wireless network communication, it is equipped with a wireless communication module 130 containing WiFi functionality.

[0039] The wireless communication module 130 enables the electromyography (EMG) bracelet 100 to communicate with the processing terminal 200 via wireless connection, realizing the uplink transmission of EMG data and the downlink reception of control commands. The EMG bracelet 100 can establish a direct connection with the processing terminal, or it can establish a connection indirectly through devices such as gateways and routers.

[0040] In some embodiments, the controller 140 may include at least one of a microprocessor, a digital signal processing unit, and a power management unit, as well as various interfaces for input / output (first interface to nth interface). The controller 140 controls the operation of the electromyography (EMG) bracelet and responds to interactive commands through various software control programs stored in the memory, thereby controlling the overall operation of the EMG bracelet.

[0041] In some embodiments, the power supply 150 provides power to the various components of the myoelectric bracelet. For example, the power supply 150 may include a built-in battery and charging management circuitry to support device battery life and charging management.

[0042] In some embodiments, the status feedback module 160 is used to output feedback information on the device's operating status to the user. For example, the status feedback module 160 may include a vibration unit and an indicator light unit, which can indicate the device's connection status, operating mode, battery status, and other information through vibration and light status.

[0043] To perform full interactive functions, in some embodiments, the processing terminal 200 is configured with an edge computing processing module, a speech synthesis and output module, and mobile APP / embedded control software. The processing terminal 200 may run an operating system, which is a computer program used to manage and control the hardware and software resources in the processing terminal, and can support the operation of functions such as electromyography signal processing and result output.

[0044] It should be noted that the operating system can be a native operating system based on a specific operating platform, or a third-party operating system that is heavily customized for that platform. When the processing terminal is a mobile device, a mobile operating system can be adapted; when the processing terminal is an embedded device, an embedded operating system can be adapted.

[0045] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 3 As shown, in some embodiments, the system is divided into four layers, which from top to bottom are an Applications layer (referred to as "application layer"), an Application Framework layer (referred to as "framework layer"), a system library layer and a kernel layer.

[0046] In some embodiments, the application layer is configured to provide services and interfaces for applications, so that the processing terminal 200 can run matching control software and functional applications, and interact with users based on the applications. At least one application can be run in the application layer, and these applications can be system setting programs that come with the operating system, or can be functional application programs matched with electromyographic interaction. In specific implementation, the application packages in the application layer are not limited to the above examples.

[0047] The framework layer provides applications with Application Programming Interfaces (APIs) and programming frameworks. The application framework layer includes some pre-defined functions. The application framework layer is equivalent to a processing center, which determines that the applications in the application layer perform actions. Through the API interfaces, applications can access resources in the system and obtain system services during execution.

[0048] As Figure 3 shown, in the embodiments of the present application, the application framework layer includes a signal processing component, a manager, a content provider, etc. Among them, the signal processing component can design and implement the processing flow and algorithm call of electromyographic signals, including functional components such as signal preprocessing, feature extraction, recognition and inference. The manager includes at least one of the following modules: a device manager configured to interact with and manage an electromyographic bracelet connected to the system; a data manager configured to provide storage and access interfaces of electromyographic data for system services or applications; a notification manager configured to control the display and output of notification messages; a power manager configured to manage the power consumption and power supply status of the processing terminal.

[0049] In some embodiments, the device manager is configured to manage the status of the connected electromyographic bracelet device and the function switching logic, for example, controlling the start, stop and parameter adjustment of the acquisition function. A speech synthesis manager is configured to manage the speech output function and control the trigger, parameter adjustment and playback logic of speech synthesis.

[0050] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction libraries included in the system runtime library layer, for example, a C / C++ instruction library, to implement the functions to be achieved by the framework layer.

[0051] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the processing terminal 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 3 As shown, hardware drivers can be configured in the kernel layer. The drivers included in the kernel layer can be at least one of the following: communication driver, audio driver, display driver, sensor driver, power driver, etc.

[0052] In some embodiments, the myoelectric bracelet-based interactive system supports sign language recognition and speech conversion. Sign language recognition relies on visual acquisition devices (such as RGB cameras or depth cameras) to capture hand movements and facial expressions, and then identifies sign language words through image feature extraction and dictionary matching algorithms. However, visual acquisition devices are highly sensitive to lighting conditions, obstructions, and shooting angles, and their accuracy decreases in low-light, backlit, or obstructed scenarios. Furthermore, visual acquisition involves capturing images of the user's face and surrounding environment, posing a privacy risk, and it is difficult to capture subtle muscle activity information that occurs synchronously with hand movements, resulting in insufficient ability to distinguish sign language words with small movements and delicate finger postures.

[0053] In some embodiments, sign language recognition is also based on a single-arm electromyography (EMG) acquisition device. This device is worn on one forearm and collects local muscle electrical signals through a surface electrode array. After the signal is identified by a classification algorithm, the signal is transmitted to the terminal to drive the speech synthesis module to output the corresponding speech. However, the muscle activity information that a single-arm EMG acquisition device can acquire is limited, making it difficult to fully cover the symmetrical and asymmetrical muscle activity patterns involved in coordinated hand gestures. Furthermore, during prolonged use, the recognition performance will significantly decrease over time due to factors such as muscle fatigue, skin sweating, and changes in electrode contact impedance.

[0054] To address the aforementioned issues, some embodiments of this application provide a sign language recognition and speech conversion system and method based on a dual electromyography (EMG) wristband. By collaboratively acquiring bilateral EMG signals, the system can fully capture the characteristics of hand gestures, improve the accuracy of continuous sign language recognition, achieve end-to-end conversion between sign language and speech, and provide sign language users with a convenient and natural voice interaction experience.

[0055] Figure 4 A flowchart illustrating an interaction method based on an electromyography (EMG) wristband, as provided in some embodiments of this application, is shown below. Figure 4 As shown, the processing terminal 200 is configured to perform the following steps: S401. Simultaneously acquire the first surface electromyography signal and the second surface electromyography signal.

[0056] The processing terminal 200 synchronously receives the first surface electromyography (EMG) signal collected by the first EMG wristband and the second surface EMG signal collected by the second EMG wristband via a wireless communication connection. The two signal links use a unified sampling frequency and a synchronized timestamp to ensure the timing alignment of the bilateral EMG signals. The first surface EMG signal is generated by the target user's first forearm muscle group, and the second surface EMG signal is generated by the target user's second forearm muscle group.

[0057] For example, the first side is the left forearm and the second side is the right forearm. By collecting electromyographic signals from both sides, it is possible to fully cover the activity information of all muscle groups involved in hand gestures, fully capture the symmetrical and asymmetrical hand gesture features, and avoid recognition errors caused by incomplete information collected from one side.

[0058] S402. Perform feature fusion on the first surface electromyography signal and the second surface electromyography signal to generate fused feature data.

[0059] The processing terminal 200 performs preprocessing operations such as filtering, framing, feature extraction, and dimensionality reduction on the two surface electromyography signals to obtain two single-channel feature sequences with unified dimensions. Then, the two single-channel feature sequences are spliced ​​or attention-weighted fused to generate fused feature data.

[0060] S403. Call the preset spatiotemporal fusion recognition model to perform sign language recognition on the fused feature data and obtain the corresponding sign language text information.

[0061] The processing terminal 200 inputs the fused feature data into the spatiotemporal fusion recognition model. The spatiotemporal fusion recognition model includes a one-dimensional convolutional neural network (1D-CNN) spatial feature extraction module, a bidirectional long short-term memory network (Bi-LSTM) temporal modeling module, an attention mechanism weighting module, and a connectionist temporal classification (CTC) sequence decoding module. It can parse the sign language word sequence corresponding to continuous sign language actions from the fused feature data, and then generate sign language text information in natural language form through lexical mapping, grammatical recombination, and language model optimization.

[0062] S404. Convert the sign language text information into a speech signal and play the speech signal.

[0063] Based on the current network connection status, the processing terminal 200 adaptively selects either a cloud-based speech synthesis service or a local lightweight speech synthesis engine to convert sign language text information into speech signals, which are then played through a speaker or external audio device.

[0064] In this embodiment, by simultaneously acquiring electromyographic signals from the surface of both hands using bilateral electromyographic wristbands, the activity information of muscle groups corresponding to coordinated sign language movements of both hands can be completely captured. Compared with visual acquisition, it is not affected by lighting or occlusion, does not require capturing environmental images, and has higher privacy and security. Compared with single-arm acquisition, it can obtain more comprehensive movement features, reduce the impact of muscle fatigue and changes in electrode impedance on recognition performance, and improve the recognition accuracy of continuous sign language movements in different scenarios, providing a stable and smooth sign language-to-speech conversion interaction experience for hearing-impaired and speech-impaired users.

[0065] In practical applications, raw surface electromyography (EMG) signals typically contain environmental noise, power frequency interference, and motion artifacts, and the EMG signals from different channels differ in amplitude range and frequency distribution. Therefore, this application introduces a multi-level preprocessing chain before feature fusion to improve signal quality and provide highly representative feature vectors for subsequent identification. Figure 5 This is a schematic flowchart illustrating signal preprocessing and feature extraction provided in some embodiments of this application. In some embodiments, the processing terminal 200 is configured to perform the following steps: S501. Filter the first surface electromyography (EMG) signal and the second surface EMG signal to obtain two preprocessed EMG signals with interference removed.

[0066] The processing terminal 200 performs cascaded filtering on the two raw surface electromyography signals in sequence.

[0067] In some implementations, the filtering process includes: a Butterworth 4th-order bandpass filter to preserve the effective frequency band of the electromyographic signal and filter out low-frequency motion artifacts and high-frequency environmental noise; an adaptive notch filter to remove power line frequency interference; and full-wave rectification combined with low-pass filtering to extract the envelope of the electromyographic signal, reflecting the changing trend of muscle activation. After the above filtering process, two pre-processed electromyographic signals with interference removed are obtained.

[0068] S502. The two preprocessed electromyographic signals are framed using a sliding window with a preset window length and sliding step size to obtain two framed electromyographic signals.

[0069] The processing terminal 200 performs sliding window frame division processing on the filtered signals from each channel of the two wristbands.

[0070] In some implementations, the window length is set to 200ms and the sliding step size is set to 50ms, meaning there is a 75% overlap between adjacent windows to fully capture the transition information between adjacent gestures. After frame-by-frame processing, each signal outputs a segment of electromyographic signal corresponding to the number of frames.

[0071] S503. Extract the time-domain and frequency-domain features corresponding to each frame of the two-channel segmented electromyography signals, and concatenate the time-domain and frequency-domain features to obtain two initial feature vectors.

[0072] For each frame of electromyography signal, the processing terminal 200 simultaneously extracts time-domain features and frequency-domain features.

[0073] In some implementations, the time-domain features include the following six dimensions: MAV (Mean Absolute Value), WL (Waveform Length), ZC (Zero Crossing), SSC (Slope Sign Change), VAR (Variance), and RMS (Root Mean Square); the frequency-domain features include the following two dimensions: MF (Mean Frequency) and MDF (Median Frequency). Eight dimensions are extracted per channel. For a single-arm bracelet with a 16-channel electrode array, 128 dimensions are generated per frame. If each bracelet is configured with 16 channels, and two EMG bracelets have a total of 32 channels, then a 256-dimensional initial feature vector is generated per frame.

[0074] The time-domain features and frequency-domain features are concatenated sequentially along the channel dimension to obtain two initial feature vectors.

[0075] S504. Perform dimensionality reduction on the two initial feature vectors to obtain a single-path feature sequence with unified dimensions.

[0076] The processing terminal 200 performs dimensionality reduction on the initial feature vector of each path to compress the feature dimension, remove redundant information, and reduce the computational load of subsequent models.

[0077] In some implementations, Principal Component Analysis (PCA) is used to reduce the dimensionality of the initial 256-dimensional eigenvectors to 128 dimensions.

[0078] S505, Perform feature fusion based on single-path feature sequence.

[0079] The processing terminal 200 concatenates the two single-channel feature sequences (e.g., 128-dimensional) with unified dimensions to obtain fused feature data (e.g., 256-dimensional), which is then input into the subsequent spatiotemporal fusion recognition model for sign language recognition.

[0080] In some implementations, weighted splicing or attention mechanism fusion can also be used to dynamically adjust the fusion weights based on the contribution of each channel to the current gesture recognition.

[0081] In this embodiment, a cascaded filtering chain of bandpass filtering, notch filtering, and envelope extraction is used to filter out environmental noise and power frequency interference, while retaining the effective frequency band and muscle activation envelope information of the electromyographic signal. The extraction strategy combining time-domain and frequency-domain features can comprehensively characterize the electromyographic signal from two dimensions: signal amplitude variation and spectral distribution, providing rich and effective feature representations for subsequent recognition. PCA dimensionality reduction retains the main information while compressing the feature dimensions, which can reduce the computational complexity of the spatiotemporal fusion recognition model and improve the real-time response speed of sign language recognition. The two signals are fused in parallel after a completely symmetrical preprocessing process to ensure the temporal consistency of the bilateral sign language action features.

[0082] Because electromyographic signals are affected by factors such as muscle fatigue, skin sweating, and changes in electrode-skin contact impedance, a slow baseline drift can occur during prolonged use. Using fixed baseline parameters will cause the accuracy of sign language recognition to gradually decrease with increasing usage time.

[0083] Therefore, such as Figure 6 As shown, Figure 6 This is a schematic diagram of a dynamic baseline adaptive calibration process provided for some embodiments of this application. In some embodiments, the processing terminal 200 is further configured to perform the following steps before performing feature fusion: S601. Read the single-frame features from the two single-channel feature sequences as the current frame features, and detect the target user's limb movement state based on the amplitude change of the current frame features.

[0084] The processing terminal 200 continuously reads the MAV feature values ​​of each frame in the single-channel feature sequence and calculates the signal abrupt change amplitude between adjacent frames. If the signal abrupt change amplitude is greater than a preset threshold (e.g., greater than 30), it is determined that the target user is currently in motion, and the update of baseline parameters is paused to avoid contamination of baseline statistics by electromyography data in motion. If the signal abrupt change amplitude does not exceed the preset threshold and remains in the low amplitude range, it is determined that the target user is in a stationary state.

[0085] S602. Based on the detection results of limb movement status, execute the corresponding baseline update strategy.

[0086] If the target user's limb movement is detected to be in a static state, and the duration of this static state reaches a first preset duration (e.g., 2 seconds), the processing terminal 200 resamples the static feature data for a second preset duration (e.g., 1 second), and calculates and resets the baseline statistical parameters based on the resampled data. The second preset duration is shorter than the first preset duration. The baseline statistical parameters are statistical measures characterizing the baseline level of the target user's electromyographic features in a static state, including the mean μ. baseline and standard deviation σ baselineResetting the baseline parameters can eliminate the contamination of the baseline estimation by the motion state in the previous stage, ensuring the accuracy of the calibration benchmark.

[0087] If the target user is detected to be in a normal motion state, the processing terminal 200 updates the baseline statistical parameters by sliding through a feature buffer of a preset length (e.g., 100 frames).

[0088] In some implementations, the sliding update method involves adding each new frame of feature data to the tail of the feature buffer and removing the old data from the head of the buffer upon receiving the new frame of feature data, then recalculating μ based on the updated buffer data. baseline and σ baseline This enables adaptive tracking of slow drift in the signal baseline.

[0089] S603. Based on baseline statistical parameters, perform dynamic calibration on single-frame features in the two single-channel feature sequences to obtain the two calibrated feature sequences.

[0090] During the calibration phase, a dynamic calibration formula is used to adjust the current frame features F. t Perform calibration; the calibration formula is F. t '=(F t -μ baseline ) / σ baseline This means subtracting the baseline mean from each feature dimension of the current frame and then dividing by the baseline standard deviation to eliminate baseline offset and scale changes caused by muscle fatigue, electrode impedance changes, etc.

[0091] S604. Perform feature fusion based on the feature sequences after two-way calibration.

[0092] After completing the dynamic baseline calibration, two calibrated feature sequences are obtained. The calibrated features eliminate the adverse effects of signal baseline drift and the feature distribution is more stable. Then, the two calibrated feature sequences are fused according to the aforementioned fusion method to generate fused feature data with higher stability and stronger representation ability. This data is then input into the subsequent spatiotemporal fusion recognition model for sign language recognition.

[0093] In some embodiments, when the processing terminal 200 updates the baseline statistical parameters by sliding through a feature buffer of a preset length, it also monitors the average confidence accuracy corresponding to the sign language text information. When classifying features in each frame, the sign language recognition model outputs probability values ​​for each category, and takes the highest probability value as the confidence level for that frame. The processing terminal 200 calculates the average confidence accuracy over a recent period using a sliding window approach. If the average confidence accuracy is detected to be lower than a preset threshold (e.g., 85%), it determines that there is a deviation between the current baseline parameters and the user's actual electromyographic features. In this case, the update rate of the baseline statistical parameters is increased, for example, by doubling the sliding update frequency of the buffer or reducing the buffer length, to accelerate the adaptation speed of the baseline parameters to changes in the user's state. If the average confidence accuracy is greater than or equal to the preset threshold, the current baseline update strategy remains unchanged.

[0094] This embodiment of the disclosure achieves adaptive switching of baseline parameters through motion state detection. In a static state, the baseline is reset to eliminate motion contamination; in a moving state, the baseline is updated via sliding to track signal drift, compensating for baseline offset issues caused by muscle fatigue, skin sweating, and changes in electrode impedance. Simultaneously, the baseline update rate is dynamically adjusted based on the sign language recognition confidence accuracy, and baseline calibration is actively accelerated when recognition accuracy decreases, balancing calibration stability and sensitivity, thus improving the robustness of sign language recognition in long-term continuous use scenarios.

[0095] Since the sign language movements described in step S403 involve both the coordinated activation of different hand muscle groups at the same time (i.e., spatial dimension) and the temporal dependence between consecutive movements (i.e., time dimension).

[0096] Therefore, embodiments of this application also provide a spatiotemporal fusion recognition model that takes into account both spatial feature extraction and temporal dependency modeling. For example... Figure 7 As shown, Figure 7 This is a schematic diagram of the hierarchical architecture of a sign language recognition model provided in some embodiments of this application. In some embodiments, the processing terminal 200 invokes a spatiotemporal fusion recognition model to perform sign language recognition on the fused feature data, specifically configured as follows: S701. Using the one-dimensional convolutional neural network module in the spatiotemporal fusion recognition model, spatial dimension features are extracted from the fused feature data to obtain a spatial feature sequence.

[0097] The fused feature data is first fed into a one-dimensional convolutional neural network (1D-CNN) module. In some implementations, the 1D-CNN module contains two convolutional layers.

[0098] In some implementations, the first convolutional layer has 64 kernels of size 3 and uses ReLU activation; the second layer has 128 kernels of size 3 and uses ReLU activation. A batch normalization layer is placed between the two convolutional layers to accelerate model convergence. The convolutional operation is performed along the channel dimension (i.e., the feature dimension of different muscle groups), which can uncover the spatial synergistic relationships between different acquisition channels and different forearm muscle groups at the same time. After processing by the 1D-CNN module, the spatial feature sequence is output.

[0099] S702. By using the bidirectional long short-term memory network module in the spatiotemporal fusion identification model, the spatial feature sequence is modeled for temporal dependence to obtain the temporal feature sequence.

[0100] Spatial feature sequences are input into the Bidirectional Long Short-Term Memory (Bi-LSTM) network module.

[0101] In some implementations, the Bi-LSTM module comprises two layers: a first layer with 256 hidden units and a second layer with 128 hidden units. Both layers have `return_sequences=True` set to output the complete time series. Each Bi-LSTM layer is followed by a Dropout operation to prevent overfitting. Bi-LSTM scans the feature sequence in both forward and backward directions, enabling it to simultaneously capture both forward dependencies (the influence of one action on the next) and backward dependencies (the supplementary information from the next action to the previous action) of consecutive sign language movements.

[0102] S703. Through the attention mechanism module in the spatiotemporal fusion recognition model, calculate the attention weight of the temporal feature sequence, and perform weighted summation on the temporal feature sequence based on the attention weight to obtain the context feature vector.

[0103] The temporal feature sequence input attention mechanism module uses a learnable weight matrix to weight each time step of the temporal feature sequence.

[0104] In some implementations, attention weights are calculated using the formula α = softmax(W·h + b), where h is the temporal feature vector, and W and b are learnable parameters. The attention mechanism adaptively strengthens the attention weights for key action frames in sign language recognition (such as gesture start and end frames) while suppressing the influence of redundant frames (such as gesture hold frames and non-critical transition frames). After weighted summation, a fixed-dimensional contextual feature vector is obtained, which aggregates the global semantic information of the entire sign language action sequence.

[0105] S704. Input the context feature vector into the fully connected classification layer of the spatiotemporal fusion recognition model to obtain the probability distribution of sign language words.

[0106] The context feature vector is then fed into a fully connected classification layer. In some implementations, the fully connected classification layer comprises: a first fully connected layer with 512 units, using ReLU activation function and followed by Dropout (dropout rate of 0.5); a second fully connected layer with num... classes Units, num classes The number of words in the sign language vocabulary is equal to the number of words in the vocabulary list, and the activation function is Softmax. Softmax outputs the probability value of each sign language word, forming a probability distribution of the sign language words.

[0107] Based on the above four-layer cascaded architecture, the spatiotemporal fusion recognition model can achieve end-to-end mapping from the original fusion features to the probabilities of sign language words.

[0108] In some embodiments, the spatiotemporal fusion recognition model further includes a connectionist temporal classification (CTC) module for enabling automatic segmentation of continuous sign language without the need for pre-annotated lexical boundaries. The CTC module is configured with a preset sign language vocabulary, which includes sign language word entries and blank symbols used to identify word boundaries.

[0109] S705. Determine the corresponding sign language text information based on probability distribution.

[0110] After the processing terminal 200 obtains the probability distribution, it determines the corresponding sign language text information based on the probability distribution.

[0111] In some embodiments, when the processing terminal 200 determines the corresponding sign language text information based on the probability distribution, for continuously input fused feature data, the CTC module inserts blank symbols into a preset sign language vocabulary list and determines the alignment path between each frame feature in the fused feature data and the sign language vocabulary entry and blank symbol. CTC calculates the sum of probabilities of all possible alignment paths using a forward-backward algorithm, with a time complexity of O(T×U), where T is the number of frames and U is the length of the vocabulary sequence. The corresponding vocabulary probability sequence is then obtained.

[0112] Decoding is performed on the word probability sequence. In some implementations, the decoding process may employ a greedy decoding strategy. In each frame, the tag with the highest probability (including sign language words or blank symbols) is selected, and after post-processing to remove consecutive duplicate tags and blank symbols, discrete word fragments are obtained.

[0113] In other implementations, the decoding process may also employ a beam search decoding strategy, retaining the Top-K optimal paths to improve decoding accuracy.

[0114] The processing terminal 200 detects word boundaries based on the location of blank symbols in the decoding results, splices discrete word fragments into a continuous word sequence in chronological order, and finally generates sign language text information based on the continuous word sequence.

[0115] This embodiment of the disclosure achieves automatic word segmentation for continuous sign language based on the CTC mechanism. It eliminates the need to pre-annotate the start and end timestamps of each sign language word in the training data, reducing the annotation cost of the training data. CTC automatically learns word boundaries using whitespace symbols, and combined with post-decoding processing, it can accurately decompose continuous action streams into independent sign language words, enabling the system to support complete continuous sign language recognition and adapt to natural and fluent sign language expression scenarios.

[0116] When generating sign language text information, there are differences between the grammatical rules of sign language and natural language. For example, the order of expression in sign language is often subject-object-verb, while that in Chinese natural language is subject-verb-object. In addition, sign language vocabulary may also have one-to-many semantic ambiguity.

[0117] Therefore, such as Figure 8 As shown, Figure 8 This is a schematic diagram illustrating the sign language to natural language mapping process provided in some embodiments of this application. In some embodiments, the processing terminal 200 is configured with a lexical mapping library, a grammar rule engine, and a language model. When the processing terminal 200 generates sign language text information based on a continuous word sequence, it is specifically configured as follows: S801. A word-by-word mapping transformation is performed on the continuous word sequence using a lexical mapping library to obtain a natural language word string.

[0118] The lexical mapping library stores the mapping relationship between sign language vocabularies and natural language vocabularies, supporting one-to-one or one-to-many mappings. The processing terminal 200 performs word-by-word lookup on the sign language vocabulary sequence output by CTC decoding, mapping each sign language vocabulary to one or more corresponding natural language candidate words, resulting in a natural language vocabulary string. For one-to-many mappings, all candidate words are retained and submitted to the subsequent statistical optimization module for disambiguation selection.

[0119] S802. The word order structure of the natural language vocabulary string is detected by the grammar rule engine to obtain the word order structure detection result. Based on the word order structure detection result, the word order is reorganized to obtain the initial sentence.

[0120] The grammar rule engine has a built-in set of predefined grammar rules used to adjust the word order of sign language expressions to conform to the word order structure of the target natural language (such as Chinese). For example, for the sign language word order "I-food-eat", the grammar rule engine detects the structure of subject (I), object (food), and verb (eat), and then reorganizes it into "I-eat-food" according to the Chinese SV word order rules. The grammar rule engine is also responsible for handling the natural language expression of grammatical elements such as tense marking and negation marking.

[0121] S803, performing probability optimization on the initial statement through a language model, and performing contextual disambiguation on the initial statement in combination with dialogue history of a preset window length, so as to obtain sign language text information.

[0122] The language model adopts an n-gram statistical language model or a neural language model, calculates language probabilities for a plurality of candidate statements generated after grammatical reorganization, and selects the statement with the highest probability as the optimal initial statement. Meanwhile, the processing terminal 200 maintains a dialogue history window with a preset window length (for example, the last 5 statements), and selects the most reasonable interpretation for polysemous words or ambiguous words based on context information in the dialogue history. For example, when the sign language word "xing" can be mapped to "walk" or "okay", if semantics related to travel or transportation appears in the dialogue history, the interpretation of "walk" is preferentially selected. Finally, after language model optimization and contextual disambiguation, smooth and natural sign language text information is output.

[0123] In the embodiment of the present disclosure, a conversion bridge between sign language expression and natural language is built through a three-level processing link of lexical mapping, grammatical rule reorganization and language model optimization. Lexical mapping completes the corresponding conversion at the lexical level, the grammatical rule engine corrects the unique word order structure of sign language, and the language model combines dialogue history to eliminate word meaning ambiguity and optimize expression fluency, ensuring that the output sign language text information has accurate semantics and smooth word order, thereby improving the intelligibility and naturalness of speech synthesis.

[0124] In order to meet users' voice interaction requirements in different network environments, such as Figure 9 shown, Figure 9 it is a schematic flow chart of adaptive speech speech speech speech speech speech speech synthesis provided by some embodiments of the present application. In some embodiments, when the processing terminal 200 converts sign language text information into a speech signal, it is specifically configured to: S901, detecting a network connection state of the processing terminal.

[0125] After receiving the sign language text information, the processing terminal 200 firstly detects the current network connection state, including a WiFi connection state and a mobile data network connection state, and determines whether it has the capability of accessing a cloud service.

[0126] S902, selecting a corresponding speech synthesis mode according to the network connection state.

[0127] If the network connection state is an online state, the processing terminal 200 invokes a cloud speech synthesis service, and sends the sign language text information to the cloud for high-quality speech synthesis processing. The cloud speech synthesis service adopts advanced neural network speech synthesis technology, can output human voice with high naturalness, and supports adjustment of various timbres, languages and intonations. After the cloud synthesis is completed, the processing terminal 200 receives the audio data stream returned by the cloud and generates a corresponding speech signal.

[0128] If the network connection is offline, such as without Wi-Fi, mobile data, or with a poor network signal, the processing terminal 200 calls the locally integrated lightweight speech synthesis engine to perform speech synthesis processing on the sign language text information. The local lightweight engine is pre-installed and runs locally when the device leaves the factory, without relying on a network connection, ensuring the availability of basic voice output functions, and is suitable for dialogue scenarios with extremely high real-time requirements.

[0129] S903, Play voice signal.

[0130] The processing terminal 200 plays the generated voice signal through a built-in speaker or an external audio output device.

[0131] In some implementations, the processing terminal 200 responds to a user-inputted speech rate adjustment command, dynamically setting the playback speed of the voice signal within a preset adjustment range (e.g., 0.8x-1.2x), and plays the voice signal at the adjusted playback speed. Users can input speech rate adjustment requests through the interface controls or voice commands of the processing terminal 200; for example, the default speech rate is 1.0x. This speech rate adjustment function can adapt to users of different ages and with different hearing habits, enhancing the personalized adaptability of the interactive experience.

[0132] This embodiment of the disclosure can dynamically switch the synthesis path according to the network status. When connected to the internet, it calls cloud services to output highly natural-sounding speech, while when offline, it enables a local lightweight engine to ensure the availability of basic functions in environments without a network, balancing speech synthesis quality and usability across all scenarios. Simultaneously, it supports user-defined playback speed, enhancing the adaptability and comfort of the interactive experience.

[0133] In the signal acquisition and transmission stage, to ensure the acquisition quality and transmission stability of bilateral electromyographic signals, the first electromyographic wristband and the second electromyographic wristband in this application embodiment complete signal digitization and synchronous transmission at the hardware level.

[0134] In some embodiments, the front-end amplifier circuit of each electromyography (EMG) bracelet receives the analog EMG signals collected by the surface electrode array and then converts the surface EMG signals into digital signals via an analog-to-digital converter (ADC). The digitized EMG signals can improve the anti-interference capability during transmission and avoid the problem of analog signals being susceptible to electromagnetic interference during long-distance transmission.

[0135] Two electromyography (EMG) wristbands transmit digital signals synchronously to the processing terminal 200 via a wireless communication module at a preset sampling frequency (e.g., 1000Hz). To ensure the time synchronization of the two signals, the processing terminal 200 uses a unified timestamp reference for frame alignment when receiving signals, so that the EMG signal frames collected by the left and right wristbands at the same time correspond in time, thus ensuring the integrity and temporal consistency of bilateral sign language movement features.

[0136] To continuously improve the recognition accuracy and generalization ability of the spatiotemporal fusion recognition model, and to support adaptive sign language features for more users, such as... Figure 10 As shown, Figure 10 This is a schematic diagram illustrating the process of updating the end-to-cloud collaborative model provided in some embodiments of this application. In some embodiments, the processing terminal 200 establishes a communication connection with the server 300, and the processing terminal 200 is further configured to: S1001. Upload the first surface electromyography signal, the second surface electromyography signal, and the sign language text information to the server.

[0137] With user authorization, the processing terminal 200 uploads the collected first surface electromyography (EMG) signals, second surface EMG signals, and the recognized and corrected sign language text information (as labeled data) to the server 300 via a secure communication connection. The uploaded data will be used for model retraining and parameter optimization on the server 300.

[0138] In some implementations, the uploaded electromyographic signal data may be anonymized first, such as by removing user identification information, to protect user privacy.

[0139] S1002, Receive the updated model parameters sent by the server based on electromyographic signals and sign language text information.

[0140] Server 300 aggregates electromyographic signal data and corresponding sign language text annotations uploaded from multiple processing terminals. It then uses cloud computing power to centrally train and fine-tune the spatiotemporal fusion recognition model, generating updated model parameters. Server 300 pushes the updated model parameters down to each processing terminal 200 via OTA (Over-The-Air). The distribution process uses incremental updates, transmitting only the changed parameters in the model, rather than completely replacing the model file.

[0141] S1003. Incrementally update the spatiotemporal fusion recognition model based on the updated model parameters.

[0142] After receiving the updated model parameters from server 300, terminal 200 performs an incremental update on the spatiotemporal fusion recognition model locally, and then hot-loads the updated parameters into the model's runtime environment. This incremental update method eliminates the need to restart the application or reload the entire model, resulting in low terminal resource consumption and seamless model upgrades for users. With the continuous aggregation of more user data and the ongoing iteration of the model, the recognition accuracy and generalization ability of the spatiotemporal fusion recognition model for new users' sign language features will continuously improve.

[0143] This embodiment utilizes the powerful computing capabilities of a cloud server to perform centralized training and parameter optimization of the model, overcoming the limitations of local processing terminals' computing resources and enabling continuous iterative improvement of model recognition accuracy. The incremental parameter update method balances the speed of model updates with the continuity of user experience. Uploading local electromyography (EMG) data to the cloud allows the model to continuously adapt to the EMG characteristics of more users and a richer vocabulary of sign language, thereby expanding the applicability and recognition capabilities of the interactive system.

[0144] In other embodiments, the number of electrode channels of the electromyography (EMG) bracelet is not limited to 16 channels, but can also be configured with 8 channels, 24 channels or 32 channels to meet the needs of different cost budgets and application scenarios.

[0145] In other embodiments, the wireless communication protocol is not limited to Bluetooth 5.0, but can also employ communication protocols such as WiFi 6, ZigBee, and UWB (Ultra Wide Band). WiFi mode can support higher data transmission rates, making it suitable for scenarios requiring the transmission of high-density electromyography (EMG) array data.

[0146] In other embodiments, the temporal features in the above feature extraction can be supplemented with additional feature dimensions such as IEMG (Integrated Electromyography) and MNF (Mean Frequency). The dimensionality reduction method is not limited to PCA, but can also use dimensionality reduction techniques such as LDA (Linear Discriminant Analysis), t-SNE (t-Distributed Stochastic Neighbor Embedding), or autoencoders.

[0147] In other embodiments, the temporal modeling module in the spatiotemporal fusion recognition model can also be replaced by a gated loop unit or a Transformer-based temporal encoder.

[0148] In other embodiments, a pre-recorded audio splicing scheme can be added to the speech synthesis mode: pre-record the corresponding audio segments of commonly used sign language words, and play the recognition results directly by splicing the pre-recorded audio segments, which can achieve extremely low speech output latency and is suitable for real-time dialogue scenarios with extremely high latency requirements.

[0149] In multi-user recognition scenarios, in some embodiments, the processing terminal 200 can establish an independent electromyographic feature profile for each user based on the differences in electromyographic signal characteristics among different users through a user identification module. The processing terminal 200 performs identification and signal separation on the electromyographic signals of multiple users collected at the same time, and maintains an independent sign language recognition process for each user, thereby realizing personalized sign language recognition and speech conversion in scenarios where multiple users use the technology simultaneously.

[0150] Regarding privacy protection, in some embodiments, the processing terminal 200 encrypts the locally stored electromyography signal data and user sign language text information using the Advanced Encryption Standard (AES-256) algorithm, with the key managed in the secure hardware module of the processing terminal. Users can trigger data deletion operations through the settings interface of the processing terminal 200. After receiving the deletion command, the processing terminal 200 securely erases all user data using a multi-overwrite method.

[0151] Figure 11 Interaction timing diagrams based on electromyography (EMG) wristbands provided for some embodiments of this application. For example... Figure 11 As shown, in some embodiments, after the first and second electromyography (EMG) wristbands are initialized, they respectively collect the raw surface EMG signals (first surface EMG signal and second surface EMG signal) of the target user's bilateral forearms. After being converted from analog to digital to generate digital format signals, they are synchronously transmitted to the processing terminal 200 according to a unified timestamp.

[0152] The processing terminal 200 is equipped with a processor capable of running a spatiotemporal fusion recognition model. The processor receives two channels of surface electromyography (EMG) signals, sequentially performs filtering and denoising, framing, time-frequency feature extraction and dimensionality reduction, and then obtains two calibrated feature sequences through dynamic baseline calibration. Feature fusion is performed on the two calibrated feature sequences to generate fused feature data. The fused feature data is input into the spatiotemporal fusion recognition model, and after spatial feature extraction, temporal modeling, attention weighting and classification decoding, a continuous sign language vocabulary sequence is obtained. Lexical mapping, word order reordering and context disambiguation are sequentially performed on the sign language vocabulary sequence to generate sign language text information conforming to natural language standards, and the sign language text information is converted into a speech signal and played.

[0153] Based on the above-described interactive system, some embodiments of this application also provide an interactive method based on an electromyography (EMG) wristband, which can be applied to the processing terminal 200 described in any of the above embodiments, such as... Figure 4 As shown, the method includes: S401. Simultaneously acquire the first surface electromyography signal and the second surface electromyography signal.

[0154] The first surface electromyography (EMG) signal is generated by the first forearm muscle group of the target user, and the second surface EMG signal is generated by the second forearm muscle group of the target user.

[0155] S402. Perform feature fusion on the first surface electromyography signal and the second surface electromyography signal to generate fused feature data.

[0156] S403. Call the preset spatiotemporal fusion recognition model to perform sign language recognition on the fused feature data to obtain the corresponding sign language text information.

[0157] S404. Convert the sign language text information into a speech signal and play the speech signal.

[0158] Based on the above interaction method, by simultaneously acquiring and fusing electromyographic signals from both forearms, the electromyographic features of coordinated sign language movements of both hands are fully captured. Combined with a spatiotemporal fusion recognition model that takes into account both spatial and temporal dimensions of feature extraction, and then through mapping from sign language to natural language and adaptive speech synthesis output, end-to-end processing from electromyographic signals to sign language text and then to speech output is achieved, providing sign language users with a low-latency, high-precision natural speech interaction experience.

[0159] It should be noted that the interaction method described in the embodiments of this application can adopt the same principle and implementation as the embodiments of the above-mentioned interaction system, and will not be described again in this application.

[0160] Similar parts between the embodiments provided in this application can be referred to mutually. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods extended from the solution of this application without creative effort shall fall within the scope of protection of this application.

Claims

1. An interactive system based on an electromyography (EMG) wristband, characterized in that, include: The first electromyography (EMG) wristband is configured to collect a first surface electromyography (EMG) signal, which is generated by the first forearm muscle group of the target user. The second electromyography (EMG) bracelet is configured to acquire a second surface electromyography (EMG) signal, which is generated by the second forearm muscle group of the target user. The processing terminal establishes wireless communication connections with the first and second electromyography (EMG) wristbands, respectively, and is configured as follows: Simultaneously acquire the first surface electromyography signal and the second surface electromyography signal; The first surface electromyography (EMG) signal and the second surface EMG signal are fused to generate fused feature data; A preset spatiotemporal fusion recognition model is invoked to perform sign language recognition on the fused feature data in order to obtain the corresponding sign language text information; The sign language text information is converted into a speech signal, and the speech signal is played.

2. The interactive system based on electromyography (EMG) wristband according to claim 1, characterized in that, Before performing feature fusion on the first surface electromyography (SEMG) signal and the second surface electromyography (SEMG) signal, the processing terminal is further configured to: The first surface electromyography (EMG) signal and the second surface EMG signal are filtered to obtain two pre-processed EMG signals with interference removed. The two preprocessed electromyographic signals are framed using a sliding window with a preset window length and sliding step size to obtain two framed electromyographic signals. Extract the time-domain and frequency-domain features corresponding to each frame of the two-channel segmented electromyography signals; The time-domain features and the frequency-domain features are concatenated to obtain two initial feature vectors. Dimensionality reduction is performed on the two initial feature vectors to obtain single-path feature sequences with unified dimensions. The feature fusion is performed based on the single-path feature sequence.

3. The interactive system based on electromyography (EMG) wristband according to claim 2, characterized in that, Before performing the feature fusion based on the single-channel feature sequence, the processing terminal is further configured as follows: Read the single-frame features from the two single-channel feature sequences as the current frame features; Based on the amplitude changes of the current frame features, the limb movement state of the target user is detected; If the limb movement state is detected to be in a static state, and the duration of the static state reaches a first preset duration, then the static feature data for a second preset duration is resampled, and the baseline statistical parameters are reset; wherein, the second preset duration is less than the first preset duration, and the baseline statistical parameters are used to characterize the statistical quantity of the target user's electromyographic feature baseline level in a static state; If the target user is detected to be in a normal motion state, the baseline statistical parameters are updated by sliding through a feature buffer of a preset length; Based on the baseline statistical parameters, dynamic calibration is performed on the single-frame features in the two single-channel feature sequences to obtain the two calibrated feature sequences. The feature fusion is performed based on the two calibrated feature sequences; When the processing terminal updates the baseline statistical parameters by sliding through a feature buffer of a preset length, it is also configured to: Monitor the average confidence accuracy of sign language text information; If the average confidence accuracy is lower than a preset threshold, the update rate of the baseline statistical parameters is increased.

4. The interactive system based on electromyography (EMG) wristband according to claim 1, characterized in that, The processing terminal invokes a preset spatiotemporal fusion recognition model to perform sign language recognition on the fused feature data, specifically configured as follows: The spatial dimension features of the fused feature data are extracted by the one-dimensional convolutional neural network module in the spatiotemporal fusion recognition model to obtain a spatial feature sequence. The spatial feature sequence is modeled using the bidirectional long short-term memory network module in the spatiotemporal fusion recognition model to obtain the temporal feature sequence. The attention weights of the temporal feature sequence are calculated using the attention mechanism module in the spatiotemporal fusion recognition model. The temporal feature sequence is weighted and summed based on the attention weights to obtain the context feature vector; The context feature vector is input into the fully connected classification layer of the spatiotemporal fusion recognition model to obtain the probability distribution of sign language words; The corresponding sign language text information is determined based on the probability distribution.

5. The interactive system based on electromyography (EMG) wristband according to claim 4, characterized in that, The spatiotemporal fusion recognition model also includes a connectionist time classification module, which is configured with a preset sign language vocabulary list, which includes sign language vocabulary entries and blank symbols used to mark the boundaries of the vocabulary. The processing terminal determines the corresponding sign language text information based on the probability distribution, specifically configured as follows: For continuously input fused feature data, the connectionist temporal classification module inserts blank symbols into a preset sign language vocabulary list and determines the alignment path between each frame feature in the fused feature data and the sign language vocabulary entry and the blank symbol; Calculate the sum of probabilities of the alignment paths to obtain the corresponding word probability sequence; Decoding is performed on the word probability sequence to obtain an initial decoding result containing repeating characters and whitespace symbols; The initial decoding result is subjected to deduplication and whitespace removal processing to obtain discrete word fragments; Based on the location of the blank symbol, word boundaries are detected, and the discrete word fragments are spliced ​​into a continuous word sequence. The sign language text information is generated based on the continuous word sequence.

6. The interactive system based on electromyography (EMG) wristband according to claim 5, characterized in that, The processing terminal is equipped with a lexical mapping library, a syntax rule engine, and a language model; The processing terminal generates the sign language text information based on the continuous word sequence, specifically configured as follows: The continuous vocabulary sequence is transformed word by word using the lexical mapping library to obtain a natural language vocabulary string; The grammar rule engine is used to detect the word order structure of the natural language vocabulary string to obtain the word order structure detection result. Based on the word order structure detection results, word order is reorganized to obtain the initial sentence; The language model is used to perform probability optimization on the initial statement, and the context disambiguation of the initial statement is performed in combination with the dialogue history of a preset window length to obtain the sign language text information.

7. The interactive system based on electromyography (EMG) wristband according to claim 1, characterized in that, The processing terminal converts the sign language text information into a speech signal, specifically configured as follows: Detect the network connection status of the processing terminal; If the network connection is active, the cloud-based speech synthesis service is invoked to process the sign language text information into speech and generate the speech signal. If the network connection is offline, the locally integrated lightweight speech synthesis engine is invoked to process the sign language text information into speech and generate the speech signal. In response to a user-inputted speech rate adjustment command, the playback speech rate of the voice signal is set within a preset adjustment range, and the voice signal is played according to the playback speech rate.

8. The interactive system based on electromyography (EMG) wristband according to claim 1, characterized in that, The first and second electromyographic wristbands are further configured as follows: The surface electromyography (EMG) signal is converted into a digital format signal, wherein the surface EMG signal is either the first surface EMG signal or the second surface EMG signal. The digital format signal is synchronously transmitted to the processing terminal according to the preset sampling frequency.

9. The interactive system based on electromyography (EMG) wristband according to any one of claims 1-8, characterized in that, The processing terminal also establishes a communication connection with the server; the processing terminal is further configured to: The first surface electromyography (EMG) signal, the second surface EMG signal, and the sign language text information are uploaded to the server. Receive updated model parameters from the server based on the first surface electromyography signal, the second surface electromyography signal, and the sign language text information; The spatiotemporal fusion recognition model is incrementally updated based on the updated model parameters.

10. An interaction method based on an electromyography (EMG) wristband, characterized in that, include: Simultaneously acquire the first surface electromyographic signal and the second surface electromyographic signal; The first surface electromyography signal is generated by the first forearm muscle group of the target user, and the second surface electromyography signal is generated by the second forearm muscle group of the target user. The first surface electromyography (EMG) signal and the second surface EMG signal are fused to generate fused feature data; A preset spatiotemporal fusion recognition model is invoked to perform sign language recognition on the fused feature data in order to obtain the corresponding sign language text information; The sign language text information is converted into a speech signal, and the speech signal is played.