Neuromuscular regulation and control treatment system with multi-modal feedback and self-adaptive adjustment
Through the neuromuscular regulation system of multimodal feedback data fusion and stratified reinforcement learning, the problems of insufficient personalized parameters, inaccurate positioning and low comfort in the existing technology are solved, and precise electrical stimulation and personalized rehabilitation path planning are realized, which improves treatment effect and patient compliance.
Patent Information
- Application Number
- CN202510498007.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-22
AI Technical Summary
The existing neuromuscular electrical stimulation technology lacks personalized parameter adjustment, inaccurate positioning, poor stimulation timing, low patient comfort, lack of dynamic synergistic stimulation, and lack of online learning dynamic optimization, resulting in limited treatment effect.
Multimodal feedback data fusion technology is adopted, combined with a layered reinforcement learning architecture, adaptive personalized optimization of millisecond-level dynamic electrical stimulation parameters are realized, precise regulation is carried out through phase synchronization pulse technology, personalized rehabilitation path planning is driven with the help of digital twin technology, and a coordinated working mode of safety feedback module and multi-device are set up to carry out real-time data acquisition and parameter adjustment.
Accurate electrical stimulation based on individual differences and real-time changes in patients is achieved, which improves the treatment effect and patient comfort, enhances the compliance and persistence of treatment, and improves the rehabilitation efficiency and safety.
Smart Images

Figure CN120346447A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neuromuscular electrical stimulation, and particularly to a neuromuscular regulation and treatment system with multi-modal feedback adaptive regulation. Background Art
[0002] As a common physical therapy method, neuromuscular electrical stimulation plays an important role in the field of rehabilitation medicine. Its main functions include muscle rehabilitation and atrophy prevention, muscle strength and endurance enhancement, nerve function recovery improvement, muscle spasm relief, blood circulation and lymphatic return promotion, movement pattern regulation, auxiliary movement training, pain management, etc. However, there are many defects that cannot be ignored in the current neuromuscular electrical stimulation technology, which seriously limits its treatment effect and patient experience.
[0003] (1) Fixed parameters, lack of personalization Traditional neuromuscular electrical stimulation devices are designed with a fixed electrical stimulation parameter mode, which means that they are completely unable to be flexibly adjusted according to the unique individual differences of patients and their real-time changing physiological states. It should be noted that the skin characteristics, nerve sensitivities, and muscle basal states of different patients are very different, and the required treatment parameters will also vary greatly. Any factor such as the point of action of electrical stimulation to regulate neuromuscular, output waveform, intensity, frequency, time, etc. will have a significant impact on the final treatment effect and the patient's experience. However, the existing product devices ignore this difference and are all put into use with fixed parameters and points, resulting in a significant reduction in the treatment effect, and even being completely ineffective in some cases.
[0004] For example, when facing patients with high muscle sensitivity, the fixed and possibly excessive stimulation intensity of traditional devices often directly causes strong pain and discomfort to the patients. In the long run, this over-stimulation is very likely to cause substantial damage to the patient's muscle tissue, not only failing to achieve the treatment purpose, but also causing secondary harm to the patient's body. On the contrary, for patients with naturally weak muscle responses, the fixed and too small stimulation intensity simply cannot effectively stimulate muscle contraction, and the treatment process is difficult to achieve the expected effect. (2) Imprecise positioning, poor stimulation timing Traditional electrical stimulation devices have significant deficiencies in the localization of the diseased site and are difficult to achieve precise positioning. This results in a scattered state of the stimulation energy during electrical stimulation treatment, a large amount of energy is wasted ineffectively, and at the same time, it is very likely to have an adverse impact on the functions of surrounding normal tissues. More importantly, traditional equipment lacks the necessary flexibility in choosing the timing of stimulation. It cannot keenly capture the patient's active movement intention and stimulate it in coordination with it. In rehabilitation therapy, especially when adjusting movement patterns, the location, timing, and intensity of stimulation play a decisive role. For example, when reconstructing complex movement patterns, such as gait, grasping and other movements, the lack of precise control of the timing of stimulation will lead to difficulties in reconstructing movement patterns. Taking the gait cycle as an example, the error may even exceed 50ms, which seriously hinders the further improvement of the treatment effect. (3) Low patient comfort In the actual use of existing neuromuscular electrical stimulation equipment, the patient's comfort level is generally at a low level. Long-term exposure to such inappropriate stimulation will cause patients to have strong psychological resistance. Once this resistance arises, it will inevitably have a negative impact on treatment compliance. Patients may not cooperate with treatment, or reduce the duration and frequency of treatment. At the same time, it will also greatly affect the sustainability of the treatment effect. Even if certain treatment results are achieved in the early stage, it may be difficult to maintain due to the patient's non-cooperation in the later stage. (4) Lack of dynamic synergistic stimulation Most of the similar devices currently on the market use static stimulation methods and fail to effectively incorporate autonomous movement into the stimulation system. For patients with nerve damage, this static stimulation mode makes it difficult for them to effectively establish neural circuits. In addition, on the technical level, the existing single-modal feedback (such as relying solely on sEMG) has obvious limitations, and it is difficult to accurately locate the source of neuromuscular abnormalities. This directly leads to a high error rate when identifying neuromuscular states and movement intentions. Moreover, due to the lack of effective timing coordination control, there are many difficulties in reconstructing complex movement patterns (such as gait and grasping). Not only that, there is a delay of more than 5 seconds between the adjustment of stimulation parameters and changes in the patient's physiological state, which makes it impossible for the device to adjust the stimulation parameters in time according to the patient's physical state, further reducing the treatment effect.
[0005] (5) No dynamic optimization based on online learning according to treatment progress Existing reinforcement learning algorithms and adaptive adjustment mechanisms often require a large amount of labeled data as a training basis. In the absence of data, it is difficult for the algorithm to converge to an effective strategy and to achieve dynamic optimization of treatment parameters. Even if some systems have certain adaptive functions, they are mostly based on limited offline training models, making it difficult to respond to dynamic changes in the patient's condition in real time during treatment, resulting in the inability to adjust the treatment plan synchronously with the rehabilitation process, further weakening the sustainability and effectiveness of the rehabilitation effect. This over-reliance on historical data and the lack of online adaptive capabilities make it difficult for traditional rehabilitation systems to meet the personalized and dynamic treatment needs of patients, and has become a key technical obstacle to improving the clinical application effect of neuromuscular electrical stimulation technology.
[0006] In summary, traditional electrical stimulators in the prior art adopt a fixed parameter mode and cannot adapt to personalized pathological and sensory characteristics. Single-modal feedback (such as only sEMG) is difficult to accurately locate the root cause of abnormalities, resulting in a high error rate in the recognition of neuromuscular states and movement intentions. The lack of sequential collaborative control makes it difficult to reconstruct complex movement patterns (such as gait, grasping, and gait cycle error > 50 ms). There is a delay of > 5 seconds between the adjustment of stimulation parameters and changes in physiological states. And there is no online learning and dynamic optimization according to the treatment progress. Summary of the Invention
[0007] In view of the above problems, the purpose of the present invention is to provide a neuromuscular regulation and treatment system with multi-modal feedback adaptive adjustment. Through multi-modal feedback data fusion technology (integrating sEMG surface electromyogram signals, IMU inertial measurement unit data, and biomechanical parameters), combined with a hierarchical reinforcement learning architecture, it realizes the adaptive personalization optimization of millisecond-level dynamic electrical stimulation parameters; adopts phase synchronization pulse technology to effectively suppress pathological abnormal rhythms while accurately triggering the regulation of normal movement patterns; finally, through digital twin technology, it drives personalized rehabilitation path planning to form a closed-loop intelligent rehabilitation system of data collection-intelligent decision-making-precise stimulation-path optimization.
[0008] The above invention purpose of the present invention is achieved through the following technical solutions: A neuromuscular regulation and treatment system with multi-modal feedback adaptive adjustment, comprising: A wearable working device, configured to output pulsed electrical signals to adjust nerves and muscles to form movements or adjust sensations, and collect electromyogram signals, motion data, and user sensory inputs to perform feedback adjustment on the output parameters of the pulsed electrical signals output; A handheld control device, configured to connect to the wearable working device and control the output working mode and working state of the wearable working device; A server-side device, configured to manage user information and preferences of the handheld control device, set function permissions of the handheld control device, and at the same time obtain data collected by the wearable working device and user usage feedback, and realize the adaptive output of the output parameters of the pulsed electrical signals of the wearable working device based on an adaptive reinforcement learning algorithm.
[0009] Further, any one of the wearable working devices includes stimulation electrodes and acquisition electrodes, and the electrode combination is carried out in any one of the following ways: Method 1: Each of the positive and negative stimulation electrodes has a path of acquisition electrodes including complete positive and negative poles, and the acquisition electrodes and the stimulation electrodes are insulated from each other; Method 2: The positive and negative electrodes of the stimulation electrode and the acquisition electrode are respectively on the same electrode, and the acquisition electrode and the stimulation electrode are insulated from each other; Method 3: The stimulation electrode and the acquisition electrode are the same electrode, the acquisition electrode and the stimulation electrode are not insulated, and time-division multiplexing is performed through a digital switch in the circuit.
[0010] Furthermore, any one of the wearable working devices includes: A signal acquisition and processing module that uses a motion sensor to feedback and collect motion data, provides motion state information, and acquires electromyogram signals through electromyogram signal acquisition. Then, the electromyogram signals are analyzed and processed, combined with hardware integration and threshold comparison to judge the muscle state, providing a basis for the output of the pulsed electrical signal serving as the stimulation signal; Among them, the motion includes two types of motions: muscle contraction motion where only muscle contraction occurs but no limb movement is driven, and limb movement driven after muscle contraction. If it is muscle contraction motion, the output of the pulsed electrical signal is provided based on the current motion state information and electromyogram signals. If it is limb movement, not only the current data but also the data for a period of time before the current moment, that is, the historical time series, is required. Based on the historical motion and electromyogram time series, a long short-term memory time network is used to predict the motion state and electromyogram information at the next moment, and the output parameters of the pulsed electrical signal are adjusted according to the difference between the predicted value and the reference value; A stimulation signal control and output module that uses a control unit to receive the processed electromyogram signals and motion data, coordinates and sends instructions to control the generation of the stimulation signal. At the same time, the stimulation signal control adjusts the output parameters of the stimulation signal according to the instructions of the control unit, transmits the electrical stimulation to the target muscle through the stimulation signal output, and at the same time uses a digital switching switch to assist in realizing time-division multiplexing of the same electrode and the on / off of the signal path; A safety feedback module that combines current and impedance feedback to monitor the changes in current and impedance during the stimulation process, ensures stimulation safety, avoids damage to the human body caused by abnormal intensity, detects the range of motion speed and acceleration, avoids damage caused by excessive movement, and detects changes in electromyogram frequency to prevent continuous muscle fatigue; A power supply module that is powered by a lithium battery and supports wireless charging function. At the same time, it uses 3V to boost to 0 - 200V controllable boost to raise the 3V voltage of the lithium battery to 0 - 200V to provide the required voltage support for the stimulation signal output; A communication module that uses broadcastable and networkable wireless communication to support wireless communication between devices, realizes data broadcast or networking functions, and facilitates multi-device collaborative work or remote control.
[0011] Furthermore, the working modes of multiple wearable working devices adopt any one of the following: Mode 1: Based on the synchronous interconnection of the controller, the handheld control device is connected to each of the wearable working devices respectively, the wearable working device feeds back detection parameters including electromyographic signals and motion data to the handheld control device, and the handheld control device sets the training output parameters and the timing of the wearable working device; Mode 2: Based on controller-free synchronous interconnection, one of the wearable work devices is connected to the handheld control device as the main device, and the other wearable work devices are connected to the main device as sub-devices. The sub-devices feed back detection parameters including electromyographic signals and motion data, and the main device sets the training output parameters and the timing of the wearable work devices.
[0012] Furthermore, the workflow of the neuromuscular regulation treatment system is as follows: Selecting a working mode, and setting the initialization parameters including current intensity and pulse width output by the system as the starting point for parameter adjustment, adjusting the scanning current intensity and pulse width in turn with the motion sensor and sensory input as feedback, drawing a pulse width-current intensity curve including dimensions of sensation, movement, and pain, determining a reference intensity and a reference pulse width through the pulse width-current intensity curve, and determining the initial current intensity, pulse width, and frequency range under the current working mode according to the working mode; Combined with the personal feelings input by the user, the joint angle, speed, acceleration, torque, muscle activation and fatigue degree are detected through motion sensors and surface electromyographic signals. The parameter regulator based on rules and thresholds or / and the parameter regulator based on adaptive reinforcement learning, as well as the safety and parameter change rate constraints, output new current intensity, pulse width and frequency range parameters. If it is a muscle contraction movement, only the data detected by the current motion sensor and surface electromyographic signal are used to output parameters. If it is a limb movement, in addition to obtaining the current data, it is also necessary to use the historical movement and electromyographic time series data obtained for a period of time before the current moment, and use the long short-term memory time network to predict the movement state and electromyographic information at the next moment, and output the parameters based on the difference between the predicted value and the reference value.
[0013] Further, determining the initial current intensity, pulse width and frequency range in the current working mode according to the working mode also includes: Call the data of the reference action library, extract the data in the library, and set the initial trigger start output time point and end output time based on rules and thresholds.
[0014] Further, a parameter regulator based on rules or / and adaptive reinforcement learning, and combining safety and parameter change rate constraints, outputs new current intensity, pulse width and frequency range parameters, and also includes: Detect the time characteristics of parameters including joint angle, velocity, acceleration, and torque through a motion sensor and compare them with the reference actions in the reference action library; Based on the trigger time point regulator of adaptive reinforcement learning, and combined with safety and parameter change rate constraint conditions, update the trigger time point of the output parameter adjustment.
[0015] Furthermore, the parameter regulator based on rules and thresholds is specifically: For neuromuscular regulation and treatment of the motion state, different timing output parameters, trigger conditions, and dynamic adjustment strategies are set according to the motion phases at different stages, the target action, and the stimulated muscles; For neuromuscular regulation and treatment of the non - motion state, different output parameters are set according to different treatment sites and the target stimulated muscles.
[0016] Furthermore, the trigger time point regulator based on adaptive reinforcement learning is specifically: Adopt a hierarchical reinforcement learning framework for learning, including the upper - layer strategy for mode switching, the middle - layer strategy for timing control, and the lower - layer strategy for parameter control; The upper - layer strategy for mode switching selects the mode according to the recognized current task and outputs the task probability distribution, and uses the LSTM network to process the timing data of the motion + surface electromyogram sensor as rules and algorithms; The middle - layer strategy for timing control inputs the IMU motion trajectory and velocity according to the recognized current motion phase and muscle activation timing, outputs discrete motion phase labels and muscle activation timing, and uses the LSTM + attention mechanism, the proximal policy optimization (PPO) algorithm, and designs a timing reward function for reinforcement learning as rules and algorithms; The lower - layer strategy for parameter control generates pulse parameters according to the phase label and timing, inputs the phase label timing history + sEMG / kinematic feedback, outputs parameter values + output timing, adopts the proximal policy optimization (PPO) algorithm, designs a total timing reward function for reinforcement learning, and sets the adjustment change range as rules and algorithms.
[0017] Furthermore, adopt the proximal policy optimization (PPO) algorithm and design a timing reward function for reinforcement learning, specifically: Establish the state space and action space of the PPO algorithm. The state space is the input condition of the action, and the action space is the operation means to change the state. The state space and the action space interact with each other, and under the guidance of the reward function, continuously optimize the electrical stimulation parameters; The state space: st = [sEMG t ,IMU t ,θ joint ,stim_historyt Among them, sEMGt is the surface electromyogram signal at time t, which reflects the muscle electrical activity state and is used to evaluate the muscle activation degree. IMUt is the inertial measurement unit data at time t, which records information such as motion posture and acceleration and helps to judge the motion state. θjoint is the joint angle, which describes the spatial position state of the joint and is used to track the accuracy of joint motion. stim_historyt is the electrical stimulation history record before time t, which contains the used stimulation parameters and provides a reference for subsequent parameter adjustment; The action space: at = [ΔI, ΔPW, ΔF, Δdelay] Among them, ΔI is the change in current intensity and is used to dynamically adjust the intensity parameter of electrical stimulation. ΔIPW is the change in pulse width, that is, the adjustment amount of the electrical pulse width, which affects the stimulation duration. ΔF is the change in frequency and is used to adjust the frequency parameter of electrical stimulation. Δdelay is the change in delay time and controls the time relationship between the triggering time of electrical stimulation and the motion; The temporal reward function is: Among them, are weight coefficients, which respectively adjust the influence degree of each reward item; is the joint motion error tracking, which measures the deviation between the actual joint motion and the target motion; is the standard deviation of the joint motion error, which reflects the fluctuation degree of the error; is the ratio of the actual electromyogram signal to the target value and is used to evaluate the matching degree between the muscle electrical activity and the expectation; is the total stimulation energy, is the current intensity, is the pulse width, and the less the total stimulation energy, the better; is the pain or discomfort degree; is the temporal accuracy reward, is the actual delay from channel i to j, is the reference delay from the rule base or natural motion data.
[0018] Compared with the prior art, the present invention includes at least one of the following beneficial effects: (1) Multimodal feedback data fusion and adaptive personalized optimization: The system uses multimodal feedback data fusion technology to integrate sEMG surface electromyography signals, IMU inertial measurement unit data and biomechanical parameters (such as joint angle, speed, acceleration, torque, etc.) collected by wearable work equipment. These multi-source data are collected and analyzed through the signal acquisition and processing module to provide comprehensive and accurate information for subsequent parameter adjustment. At the same time, combined with the server-side device based on adaptive reinforcement learning algorithm (such as hierarchical reinforcement learning architecture), it can achieve adaptive personalized optimization of millisecond-level dynamic electrical stimulation parameters. In the workflow, whether it is a parameter regulator based on rules and thresholds or a parameter regulator based on adaptive reinforcement learning, it can accurately adjust the output parameters of the pulse electrical signal (such as current intensity, pulse width, frequency, etc.) according to different motion states, treatment sites and personal feelings of users to meet the unique needs of each user.
[0019] (2) Precise control using phase-synchronized pulse technology: The system uses phase-synchronized pulse technology. In the stimulation signal control and output module, the control unit coordinates and sends instructions to precisely control the generation and output of stimulation signals. This technology can effectively suppress pathological abnormal rhythms while accurately triggering normal movement pattern control. For example, in the parameter regulator based on rules and thresholds, different timing output parameters, trigger conditions, and dynamic adjustment strategies are set for neuromuscular regulation therapy in motion and non-motion states to ensure that the stimulation signal matches the physiological rhythm of the human body and promotes the recovery and improvement of neuromuscular function.
[0020] (3) Digital twin technology drives personalized rehabilitation path planning: The system uses server-side devices to manage user information and preferences, set functional permissions, and obtain data collected by wearable work devices. On this basis, digital twin technology is used to drive personalized rehabilitation path planning. In the workflow, the data of the reference action library is called, and the parameters detected by the motion sensor are compared with the reference action. The trigger time point of the output parameter adjustment is updated by the trigger time point regulator based on adaptive reinforcement learning. This method can tailor a personalized rehabilitation plan for the user based on the real-time status and rehabilitation process, realize a closed-loop intelligent rehabilitation system of "data collection-intelligent decision-making-precise stimulation-path optimization", and improve the effect and efficiency of rehabilitation treatment.
[0021] (4) Multi-device collaboration and flexible working mode: Multiple wearable working devices can adopt two working modes: controller-based synchronous interconnection or controller-free synchronous interconnection. The handheld control device can flexibly control the output working mode and working state of the wearable working device. This multi-device collaborative working mode is not only convenient for combination and adjustment according to different treatment needs, but also enables data sharing and interaction, further improving the adaptability and treatment effect of the system.
[0022] (5)Safety guarantee and improvement of user experience: The system is equipped with a safety feedback module, which combines current and impedance feedback to monitor the changes in current and impedance during the stimulation process in real time, ensuring the safety of the stimulation and avoiding damage to the human body caused by abnormal intensity. At the same time, the power supply module uses a lithium battery for power supply and supports wireless charging function, and the communication module supports wireless communication that can be broadcast and networked, improving the convenience of device use and user experience. In addition, parameter adjustment is combined with the personal feelings input by the user in the work process, fully considering the subjective experience of the user and enhancing the user's compliance and satisfaction with the treatment.
[0023] (6)Online learning and dynamic optimization according to the treatment progress: The online learning mechanism implemented through the model-based deep reinforcement learning framework based on safety constraints effectively breaks through the bottleneck of the traditional rehabilitation system's dependence on historical data. In the scenario where the initial training data is extremely scarce or even missing, the system can quickly initialize the base model with the help of prior biomedical knowledge, achieve the preliminary adaptation of treatment parameters, and dynamically update the model parameters in an online learning manner by real-time collecting the patient's treatment data during use, accurately capturing the subtle changes in the patient's individual neuromuscular state. Description of the Drawings
[0024] Figure 1 It is the overall structure diagram of the neuromuscular regulation and treatment system with multi-modal feedback adaptive regulation of the present invention; Figure 2 It is the schematic diagram of the system composition of the neuromuscular regulation and treatment system with multi-modal feedback adaptive regulation of the present invention; Figure 3 It is the schematic diagram of the system working principle of the neuromuscular regulation and treatment system with multi-modal feedback adaptive regulation of the present invention; Figure 4 It is the schematic diagram of the combined electrode of the present invention; Figure 5 It is the principle framework diagram of any one of the wearable working devices of the present invention; Figure 6 It is the schematic diagram of the working modes of multiple wearable working devices of the present invention; Figure 7 It is the working flow chart of the neuromuscular regulation and treatment system of the present invention; Figure 8 It is the schematic diagram of the application of the neuromuscular regulation and treatment system with multi-modal feedback adaptive regulation of the present invention in cranial nerve regulation; Figure 9 It is the schematic diagram of the array arrangement of the cranial nerves of the present invention. Detailed Description of the Invention
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0026] Those skilled in the art of this technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural form. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0027] In view of the many shortcomings of existing neuromuscular electrical stimulation technologies, the present invention has made a series of important improvements. First, by adopting multi-modal biofeedback technology, it is possible to monitor various aspects of information such as the patient's movement state, nerve signals, and physiological indicators in real time. The multi-modal biofeedback adaptive reinforcement learning technology covers rich content, such as kinematic parameters, including movement trajectories, accelerations, joint angles, and biomechanical parameters such as forces and torques calculated based on these parameters; electrophysiological parameter sEMG, including electromyography data and fatigue levels, muscle activation levels, etc. calculated based on these parameters. The feedback methods include rule-based conditional feedback and feedback based on hierarchical reinforcement learning control. In terms of data processing algorithms, we have adopted adaptive filtering algorithms, Kalman filtering algorithms, etc., which can more accurately extract valuable information from complex biological signals, thereby providing more accurate data support for subsequent adjustment of electrical stimulation parameters. These information are used to adaptively adjust the electrical stimulation parameters, thus realizing true on-demand stimulation and personalized stimulation. Specifically, according to the patient's muscle sensitivity and response ability, combined with the accurate data obtained from the feedback algorithm, the stimulation intensity is accurately adjusted to ensure that while effectively stimulating muscle contraction, pain and discomfort are minimized to the greatest extent. For the localization of the lesion site, advanced sensors and imaging technologies are used to more accurately identify and focus, so that the stimulation energy is concentrated on the target area, improving the treatment efficiency and reducing the adverse effects on surrounding normal tissues. In addition, through reinforcement learning, the personalization and adaptability of treatment parameters are continuously improved, enabling the device to have intelligent learning capabilities. In terms of stimulation timing, it is combined with the patient's active movement intention. When the patient is ready to perform an active movement, the electrical stimulation device can respond in a timely manner to provide coordinated stimulation, thereby enhancing muscle strength and coordination, promoting the completion of the movement, and further improving the rehabilitation effect. For muscle activation that interferes with normal movement, it can be inhibited through stimulation signals to complete normal behavioral functions. In addition, in order to improve the patient's comfort, the device is designed with a more ergonomic structure and materials to reduce pressure and friction on the skin. And, by optimizing the waveform and frequency of the electrical stimulation, it is made closer to the natural nerve electrical signals of the human body, thereby reducing discomfort.
[0028] The following is illustrated by specific embodiments: The first embodiment As Figure 1 shown, this embodiment provides a neuromuscular regulation and treatment system with multi-modal feedback adaptive adjustment, including: A wearable working device for outputting pulsed electrical signals to regulate nerves and muscles to form movements or regulate sensations, and collecting electromyography signals, motion data, and user sensation inputs to perform feedback adjustment on the output parameters of the pulsed electrical signals. Among them, the motion data includes angles, speeds, accelerations, etc., and the user sensation inputs include pain, soreness, proprioception, position sense, etc.
[0029] A handheld control device for connecting to the wearable working device and controlling the output working mode and working state of the wearable working device. The handheld control device can be connected to multiple wearable working devices simultaneously, and the handheld control device includes bio-signals composed of flexible electronic skin, an IMU cluster, and an sEMG array.
[0030] A server-side device for managing user information and preferences of the handheld control device, setting function permissions of the handheld control device, adding, editing, and deleting function modules. At the same time, in order to be able to adaptively learn output parameters and achieve precise personalized treatment, the server-side device obtains the data collected by the wearable working device and the user's usage feedback, and realizes the adaptive output of the output parameters of the pulsed electrical signal of the wearable working device based on an adaptive reinforcement learning algorithm.
[0031] For example, as Figure 2 shown, it is a schematic diagram of the system composition of a specific example of the multi-modal feedback adaptive regulation neuromuscular regulation treatment system in this embodiment. Among them, the wearable working device adopts a wearable working unit, the handheld control device adopts a mobile phone terminal, and the server-side device adopts a server. In addition, as Figure 3 shown is a schematic diagram of the working principle of the system. The wearable working device is connected to the human body, the handheld control device is connected to the wearable working device through wireless communication, and the server-side control and self-learning center is connected to the mobile phone terminal.
[0032] Furthermore, in this embodiment, as Figure 4 shown, for any one of the wearable working devices, it includes stimulation electrodes and acquisition electrodes, and any one of the following methods is used for electrode combination: Method 1: Each of the positive and negative stimulation electrodes has a path of acquisition electrodes including complete positive and negative electrodes, and the acquisition electrodes and the stimulation electrodes are insulated from each other; Method 2: The positive and negative electrodes of the stimulation electrodes and the acquisition electrodes are respectively on the same electrode, and the acquisition electrodes and the stimulation electrodes are insulated from each other; Method 3: The stimulation electrodes and the acquisition electrodes are the same electrode, the acquisition electrodes and the stimulation electrodes are not insulated, and time-division multiplexing is performed through a digital switch in the circuit.
[0033] Furthermore, in this embodiment, as Figure 5 shown, any one of the wearable working devices includes: The signal acquisition and processing module collects motion data through the feedback of motion sensors to provide motion state information, and acquires electromyographic signals through electromyographic signal acquisition. Then, the electromyographic signals are analyzed and processed, and combined with hardware integration and threshold comparison to judge the muscle state, providing a basis for the output of the pulsed electrical signal serving as the stimulation signal; Among them, the motion includes two types of motion: muscle contraction motion where only muscle contraction occurs but no limb motion is driven, and limb motion driven by muscle contraction. If it is muscle contraction motion, the output of the pulsed electrical signal is based on the current motion state information and electromyographic signals. If it is limb motion, not only the current data but also the data for a period of time before the current moment, that is, the historical time series, are required. Based on the historical motion and electromyographic time series, a long short-term memory time network is used to predict the motion state and electromyographic information at the next moment, and the output parameters of the pulsed electrical signal are adjusted according to the difference between the predicted value and the reference value; The stimulation signal control and output module uses a control unit to receive the processed electromyographic signals and motion data, coordinate and send instructions to control the generation of the stimulation signal. At the same time, the stimulation signal control adjusts the output parameters of the stimulation signal according to the instructions of the control unit, and transmits the electrical stimulation to the target muscle through the stimulation signal output. At the same time, the digital switch is used to assist in realizing time-division multiplexing of the same electrode and the on / off of the signal path; The safety feedback module combines current and impedance feedback to monitor the changes in current and impedance during the stimulation process to ensure stimulation safety, avoid damage to the human body caused by abnormal intensity, detect the range of motion speed and acceleration, avoid damage caused by excessive movement, and detect the change in electromyographic frequency to prevent continuous muscle fatigue; The power supply module is powered by a lithium battery and supports wireless charging function. At the same time, it uses a 3V to 0 - 200V controllable boost to boost the 3V voltage of the lithium battery to 0 - 200V to provide the required voltage support for the stimulation signal output; The communication module uses broadcastable and networkable wireless communication to support wireless communication between devices, realizing data broadcast or networking functions, facilitating multi-device collaborative work or remote control.
[0034] Further, in this embodiment, as Figure 6 shown, the working modes of multiple said wearable working devices adopt any one of the following: Mode 1: Synchronous interconnection based on a controller. The handheld control device is respectively connected to each of the wearable working devices. The wearable working devices feedback detection parameters including electromyographic signals and motion data to the handheld control device, and the handheld control device sets the training output parameters and the timing of the wearable working devices; Mode 2: Based on controllerless synchronous interconnection, one of the wearable working devices is connected to the handheld control device as the master device, and the other wearable working devices are connected to the master device as slave devices. The slave devices feedback detection parameters including electromyogram signals and motion data, and the master device sets the training output parameters and the timing of the wearable working devices.
[0035] In addition, as Figure 7 shown, the working process of the neuromuscular regulation and treatment system in this embodiment is summarized as follows: Select the working mode, and set the initialization parameters including current intensity and pulse width output by the system as the starting point for parameter adjustment (such as setting the output current intensity to 0 mA and the pulse width to 1 ms), alternately adjust the scanning current intensity and pulse width with the motion sensor and sensory input as feedback, draw the pulse width-current intensity curve in dimensions including sensation, motion, and pain, determine the reference intensity and reference pulse width through the pulse width-current intensity curve, and determine the initial current intensity, pulse width, and frequency range in the current working mode according to the working mode; Combined with the personal feelings input by the user, detect the joint angle, speed, acceleration, torque, muscle activation, and fatigue degree through the motion sensor and surface electromyogram signal. Based on the rule- and threshold-based parameter regulator or / and the adaptive reinforcement learning-based parameter regulator, and combined with the safety and parameter change rate constraint conditions, output new current intensity, pulse width, and frequency range parameters. Among them, if it is a muscle contraction movement, only output parameters based on the data detected by the current motion sensor and surface electromyogram signal. If it is a limb movement, in addition to obtaining the current data, it is also necessary to predict the motion state and electromyogram information at the next moment using the long short-term memory time network based on the historical motion and electromyogram time series data obtained for a period of time before the current moment, and output parameters based on the difference between the predicted value and the reference value.
[0036] Meanwhile, in Figure 7 it, determining the initial current intensity, pulse width, and frequency range in the current working mode according to the working mode further includes: Call the data in the reference action library, and extract the data in the library to set the initial trigger start output time point and end output time based on rules and thresholds.
[0037] Furthermore, based on the rule- or / and adaptive reinforcement learning-based parameter regulator, and combined with the safety and parameter change rate constraint conditions, outputting new current intensity, pulse width, and frequency range parameters further includes: Detect the time characteristics of parameters including joint angle, velocity, acceleration, and torque through a motion sensor and compare them with the reference actions in the reference action library; based on the trigger time point regulator of adaptive reinforcement learning, and combined with safety and parameter change rate constraint conditions, update the trigger time point for adjusting the output parameters.
[0038] In this embodiment, the parameter regulator based on rules and thresholds is specifically: For neuromuscular regulation therapy of the motion state (regulation therapy with timing control), different timing output parameters, trigger conditions, and dynamic adjustment strategies are set according to the motion phases at different stages, the target action, and the stimulated muscles. Timing control can enable muscle electrical stimulation to be combined with motion to achieve the reconstruction and optimization of motor functions; for neuromuscular regulation therapy of the non-motion state (regulation therapy without timing control), different output parameters are set according to different treatment sites and the target stimulated muscles.
[0039] As shown in Table 1, taking gait as a specific example of neuromuscular regulation therapy of the motion state, the parameter settings of the parameter regulator based on rules and thresholds are shown in Table 1.
[0040] Table 1 Again, taking spinal scoliosis treatment as a specific example of neuromuscular regulation therapy of the non-motion state, the parameter settings of the parameter regulator based on rules and thresholds are as follows: (1) Concave side strengthening stimulation: Target muscles: Erector spinae, multifidus, and transverse abdominal muscles on the concave side.
[0041] Parameters: Low-frequency tetanic stimulation (20 - 30 Hz, 0.4 - 0.6 ms pulse width, intensity is 70 - 90% of the maximum tolerance).
[0042] (2) Convex side inhibition strategy: Target muscles: Quadratus lumborum and latissimus dorsi on the convex side.
[0043] Parameters: High-frequency intermittent inhibition (50 Hz, 0.1 - 0.2 ms pulse width, intensity is 30% of the motion threshold).
[0044] (3) Respiratory co-stimulation: Target muscles: Diaphragm and external intercostal muscles.
[0045] Parameters: Synchronized with the respiratory rhythm (5 Hz during inspiration, 1 Hz during expiration).
[0046] Again, for example, Figure 8As shown, taking cranial nerve regulation as a specific example of neuromuscular regulation treatment in a non - motor state, after initially determining the target through various examinations, a three - dimensional model of the head and the lesion is established based on a head CT or MRI. The arrangement of the electrical stimulation electrode array is initially designed, and then the initial stimulation parameters are output based on rules and thresholds, and the array output arrangement and stimulation parameter adjustment are performed based on adaptive reinforcement learning. As Figure 9 shown, the adaptive positioning of the target is achieved by feedback - adjusting which points in the array arrangement output. The detection target parameters are determined according to the treatment purpose. For example, the target parameters for treating sleep can be sleep time, deep sleep time, and activity frequency. The target parameters for treating depression can be the score of a depression scale.
[0047] Further, in this embodiment, as shown in Table 2, the trigger time - point regulator based on adaptive reinforcement learning is specifically: Hierarchical reinforcement learning framework is used for learning, including the upper - layer strategy for mode switching, the middle - layer strategy for timing control, and the lower - layer strategy for parameter control; The upper - layer strategy for mode switching selects the mode according to the recognized current task and outputs the task probability distribution, and uses an LSTM network to process the timing data of the motion + surface electromyography sensor as rules and algorithms; The middle - layer strategy for timing control inputs the IMU motion trajectory and speed according to the recognized current motion phase and muscle activation timing, and outputs discrete motion phase labels and muscle activation timing. It uses an LSTM + attention mechanism, the proximal policy optimization (PPO) algorithm, and designs a timing reward function for reinforcement learning as rules and algorithms; The lower - layer strategy for parameter control generates pulse parameters according to the phase label and timing, inputs the phase - label timing history + sEMG / kinematic feedback, and outputs parameter values + output timing. It uses the proximal policy optimization (PPO) algorithm, designs a total timing reward function for reinforcement learning, and sets the adjustment range as rules and algorithms.
[0048] Table 2 Hierarchy Function Input Output Rules and algorithms Upper-level decision-making (mode switching) Identify the current task (muscle building, pain, prevention of blood clots, suppression of tremors, restoration of movement, muscle coordination rebalancing) Mode selection Output task probability distribution (recognize movement intention involving movement) LSTM network processes motion + surface electromyography sensor time series data Middle-level strategy (temporal control) Identify the current motion phase and muscle activation time series (such as touchdown, mid-stance, swing in the gait cycle) IMU motion trajectory and speed Motion phase label (discrete), muscle activation time series LSTM + attention mechanism, proximal policy optimization of PPO algorithm, design of temporal reward function for reinforcement learning Lower-level strategy (parameter control) Generate pulse parameters (intensity, pulse width, frequency) and multi-channel temporal arrangement according to phase label and time series Phase label time series history + sEMG / kinematic feedback Parameter value + output time series Proximal policy optimization of PPO algorithm, design of total temporal reward function for reinforcement learning. Set the adjustment range of the adjustment amount: ΔI ∈ [-1, +1] mA for intensity adjustment, ΔPW ∈ [-0.05, +0.05] ms for pulse width adjustment, ΔFreq ∈ [-5, +5] Hz for frequency adjustment Further, in this embodiment, the proximal policy optimization (PPO) algorithm is used, and a timing reward function is designed for reinforcement learning, specifically: The state space and action space of the PPO algorithm are established. The state space is the input condition of the action, and the action space is the operation means to change the state. The state space and the action space interact with each other, and under the guidance of the reward function, the electrical stimulation parameters are continuously optimized; The state space: st = [sEMG t ,IMU t ,θ joint ,stim_historyt Among them, sEMGt is the surface electromyogram signal at time t, reflecting the muscle electrical activity state and used to evaluate the muscle activation degree; IMUt is the inertial measurement unit data at time t, recording information such as motion posture and acceleration to assist in judging the motion state; θjoint is the joint angle, describing the spatial position state of the joint and used to track the accuracy of joint motion; stim_historyt is the electrical stimulation history record before time t, including the used stimulation parameters, providing a reference for subsequent parameter adjustment; The action space: at = [ΔI, ΔPW, ΔF, Δdelay] Among them, ΔI is the change in current intensity, used to dynamically adjust the intensity parameter of electrical stimulation; ΔIPW is the change in pulse width, that is, the adjustment amount of the electrical pulse width, affecting the stimulation duration; ΔF is the change in frequency, used to adjust the frequency parameter of electrical stimulation; Δdelay is the change in delay time, controlling the time relationship between the triggering time of electrical stimulation and the motion; The time-sequence reward function is: Among them, , , , are weight coefficients, respectively adjusting the influence degree of each reward item; is the joint motion error tracking, measuring the deviation between the actual joint motion and the target motion; is the standard deviation of the joint motion error, reflecting the fluctuation degree of the error; is the ratio of the actual electromyogram signal to the target value, used to evaluate the matching degree between the muscle electrical activity and the expectation; is the total stimulation energy, , is the current intensity, is the pulse width, and the less the total stimulation energy, the better; is the pain or discomfort degree; is the time-sequence accuracy reward, , is the actual delay from channel i to j, is the reference delay from the rule base or natural motion data.
[0049] Furthermore, this embodiment further includes: realizing online learning based on a safety-constrained model-based deep reinforcement learning framework to solve the problem of adaptive optimization during the use process of the system under the condition of less initial training data or no training data, so as to meet more personalized treatment parameters. Specifically: Initialize the base model with prior biomedical knowledge. The initialization parameters include the state: the working mode adopted by the patient, and the corresponding motion state and surface electromyogram state, the corresponding stimulation intensity, frequency, duration, and stimulation trigger condition; During the use of the device, record the treatment parameters and the corresponding motion state, electromyogram state, and user feeling input, update the model parameters through online learning by regularly collecting data, and based on the updated model in the subsequent treatment process, realize the real-time optimization of the output through a safety-constrained reinforcement learning algorithm.
[0050] A computer-readable storage medium stores computer code, and when the computer code is executed, the above method is executed. Those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0051] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements can be made without departing from the principle of the present invention, and these improvements and refinements should also be regarded as the protection scope of the present invention.
[0052] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0053] It should be noted that the above embodiments can be freely combined according to needs. The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art in this technical field, several improvements and refinements can be made without departing from the principle of the present invention, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A neuromuscular regulation therapy system with multimodal feedback adaptive regulation, characterized in that Including: A wearable working device for outputting pulsed electrical signals to regulate nerves and muscles to form movements or regulate sensations, and collecting electromyography signals, motion data, and user sensory inputs to perform feedback regulation on the output parameters of the output pulsed electrical signals; A handheld control device for connecting to the wearable working device and controlling the output working mode and working state of the wearable working device; A server-side device for managing user information and preferences of the handheld control device, setting function permissions of the handheld control device, and at the same time obtaining data collected by the wearable working device and user usage feedback to adaptively output the output parameters of the pulsed electrical signals of the wearable working device based on an adaptive reinforcement learning algorithm.
2. The neuromuscular regulation treatment system with multi-modal feedback adaptive regulation according to claim 1, wherein Any one of the wearable working devices includes stimulating electrodes and collecting electrodes, and the electrode combination is carried out in any one of the following ways: Method 1: There is a path of collecting electrodes including complete positive and negative poles on each of the positive and negative stimulating electrodes, and the collecting electrodes and the stimulating electrodes are insulated from each other; Method 2: The positive and negative poles of the stimulating electrodes and the collecting electrodes are respectively on the same electrode, and the collecting electrodes and the stimulating electrodes are insulated from each other; Method 3: The stimulating electrodes and the collecting electrodes are the same electrode, the collecting electrodes and the stimulating electrodes are not insulated, and time-division multiplexing is carried out through a digital switch in the circuit.
3. The neuromuscular regulation therapy system with multi-modal feedback adaptive regulation according to claim 1, wherein Any one of the wearable working devices includes: A signal acquisition and processing module that uses a motion sensor to feedback and collect motion data, provides motion state information, and acquires electromyography signals through electromyography signal acquisition. Then, the electromyography signals are analyzed and processed, combined with hardware integration and threshold comparison to judge the muscle state, providing a basis for the output of the pulsed electrical signals serving as stimulation signals; Among them, the motion includes two types of motions: muscle contraction motion where only muscle contraction occurs but no limb movement is driven, and limb movement driven after muscle contraction. If it is muscle contraction motion, the output of the pulsed electrical signals is provided based on the current motion state information and electromyography signals. If it is limb movement, not only the current data but also the data for a period of time before the current moment, that is, the historical time series, is required. Based on the historical motion and electromyography time series, a long short-term memory time network is used to predict the motion state and electromyography information at the next moment, and the output parameters of the pulsed electrical signals are adjusted according to the difference between the predicted value and the reference value; A stimulation signal control and output module that uses a control unit to receive the processed electromyography signals and motion data, coordinates and sends instructions to control the generation of stimulation signals. At the same time, the stimulation signal control adjusts the output parameters of the stimulation signals according to the instructions of the control unit, transmits electrical stimulation to the target muscle through the stimulation signal output, and at the same time assists in realizing time-division multiplexing of the same electrode and the on-off of the signal path through a digital switching switch. The safety feedback module combines current and impedance feedback to monitor the current and impedance changes during the stimulation process to ensure the safety of stimulation and avoid damage to the human body caused by abnormal intensity. It also detects the range of movement speed and acceleration to avoid damage caused by excessive movement, detects changes in myoelectric frequency, and prevents continuous muscle fatigue. The power supply module is powered by a lithium battery and supports wireless charging. It also uses a 3V to 0-200V controllable boost to increase the 3V voltage of the lithium battery to 0-200V, providing the required voltage support for the stimulation signal output; The communication module adopts broadcastable and networkable wireless communication, supports wireless communication between devices, realizes data broadcasting or networking functions, and facilitates multi-device collaborative work or remote control.
4. The neuromuscular regulation therapy system with multi-modal feedback adaptive regulation according to claim 1, wherein, The working modes of the plurality of wearable working devices are any of the following: Mode 1: Based on the synchronous interconnection of the controller, the handheld control device is connected to each of the wearable working devices respectively, the wearable working device feeds back detection parameters including electromyographic signals and motion data to the handheld control device, and the handheld control device sets the training output parameters and the timing of the wearable working device; Mode 2: Based on controller-free synchronous interconnection, one of the wearable work devices is connected to the handheld control device as the main device, and the other wearable work devices are connected to the main device as sub-devices. The sub-devices feed back detection parameters including electromyographic signals and motion data, and the main device sets the training output parameters and the timing of the wearable work devices.
5. The neuromuscular regulation therapy system with multimodal feedback adaptive regulation according to claim 3, wherein The workflow of the neuromuscular regulation therapy system is as follows: Selecting a working mode, and setting the initialization parameters including current intensity and pulse width output by the system as the starting point for parameter adjustment, adjusting the scanning current intensity and pulse width in turn with the motion sensor and sensory input as feedback, drawing a pulse width-current intensity curve including dimensions of sensation, movement, and pain, determining a reference intensity and a reference pulse width through the pulse width-current intensity curve, and determining the initial current intensity, pulse width, and frequency range under the current working mode according to the working mode; Combined with the personal feelings input by the user, the joint angle, speed, acceleration, torque, muscle activation and fatigue degree are detected through motion sensors and surface electromyographic signals. The parameter regulator based on rules and thresholds or / and the parameter regulator based on adaptive reinforcement learning, as well as the safety and parameter change rate constraints, output new current intensity, pulse width and frequency range parameters. If it is a muscle contraction movement, only the data detected by the current motion sensor and surface electromyographic signal are used to output parameters. If it is a limb movement, in addition to obtaining the current data, it is also necessary to use the historical movement and electromyographic time series data obtained for a period of time before the current moment, and use the long short-term memory time network to predict the movement state and electromyographic information at the next moment, and output the parameters based on the difference between the predicted value and the reference value.
6. The neuromuscular regulation therapy system with multi-modal feedback adaptive regulation according to claim 5, characterized in that, Determining the initial current intensity, pulse width and frequency range in the current working mode according to the working mode also includes: Call the data of the reference action library, extract the data in the library, and set the initial trigger start output time point and end output time based on rules and thresholds.
7. The neuromuscular regulation therapy system with multimodal feedback adaptive regulation according to claim 6, characterized in that, A parameter regulator based on rules or / and adaptive reinforcement learning, which combines safety and parameter change rate constraint conditions to output new current intensity, pulse width, and frequency range parameters, further includes: Detecting the time characteristics of parameters including joint angle, velocity, acceleration, and torque through a motion sensor and comparing them with the reference actions in the reference action library; An adaptive reinforcement learning-based trigger time point regulator, which combines safety and parameter change rate constraint conditions to update the trigger time point for adjusting the output parameters.
8. The neuromuscular regulation treatment system with multi-modal feedback adaptive regulation according to claim 5, wherein A parameter regulator based on rules and thresholds, specifically: For neuromuscular regulation therapy of the motion state, different timing output parameters, trigger conditions, and dynamic adjustment strategies are set according to the motion phases at different stages, the target action, and the stimulated muscles; For neuromuscular regulation therapy of the non-motion state, different output parameters are set according to different treatment sites and the target stimulated muscles.
9. The neuromuscular regulation therapy system with multi-modal feedback adaptive regulation according to claim 7, the adaptive reinforcement learning-based trigger time point regulator, specifically: Using a hierarchical reinforcement learning framework for learning, including an upper layer policy for mode switching, a middle layer policy for timing control, and a lower layer policy for parameter control; The upper layer policy for mode switching selects a mode according to the recognized current task and outputs a task probability distribution, and uses an LSTM network to process the timing data of the motion + surface electromyogram sensor as rules and algorithms; The middle layer policy for timing control inputs the IMU motion trajectory and velocity according to the recognized current motion phase and muscle activation timing, outputs discrete motion phase labels and muscle activation timing, and uses an LSTM + attention mechanism, the proximal policy optimization (PPO) algorithm, and designs a timing reward function for reinforcement learning as rules and algorithms; The lower layer policy for parameter control generates pulse parameters according to the phase labels and timing, inputs the phase label timing history + sEMG / kinematic feedback, outputs parameter values + output timing, uses the PPO algorithm for proximal policy optimization, designs a total timing reward function for reinforcement learning, and sets the adjustment change range as rules and algorithms.
10. The neuromuscular regulation therapy system with multimodal feedback adaptive regulation according to claim 9, characterized in that, Using the PPO algorithm for proximal policy optimization and designing a timing reward function for reinforcement learning, specifically: Establishing the state space and action space of the PPO algorithm. The state space is the input condition of the action, and the action space is the operation means for changing the state. The state space and the action space interact with each other and continuously optimize the electrical stimulation parameters under the guidance of the reward function; The state space: st = [sEMG t , IMU t , θ joint , stim_history t Among them, sEMGt is the surface electromyogram signal at time t, reflecting the muscle electrical activity state and used to evaluate the muscle activation degree. IMUt is the inertial measurement unit data at time t, recording information such as motion posture and acceleration, assisting in judging the motion state. θjoint is the joint angle, describing the spatial position state of the joint and used to track the accuracy of joint motion. stim_historyt is the electrical stimulation history record before time t, containing the previously used stimulation parameters, providing a reference for subsequent parameter adjustment; The action space: at = [ΔI, ΔPW, ΔF, Δdelay] Wherein, ΔI is the change in current intensity, which is used to dynamically adjust the intensity parameter of the electrical stimulation; ΔIPW is the change in pulse width, that is, the adjustment amount of the electrical pulse width, which affects the stimulation duration; ΔF is the change in frequency, which is used to adjust the frequency parameter of the electrical stimulation; Δdelay is the change in delay time, which controls the time relationship between the trigger timing of the electrical stimulation and the movement; The timing reward function is: Among them, , , , are weighting coefficients, which respectively adjust the influence degrees of each reward item; For tracking joint motion errors, which measure the deviation between the actual joint motion and the target motion; is the standard deviation of the joint motion error, reflecting the degree of error fluctuation; It is the ratio of the actual electromyogram signal to the target value and is used to evaluate the matching degree between muscle electrical activity and the expectation; is the total stimulation energy, , is the current intensity, is the pulse width, and the less total stimulation energy, the better; is the degree of pain or discomfort; For the timing accuracy reward, , is the actual delay from channel i to j, is the reference delay from the rule base or natural motion data.
11. The neuromuscular regulation therapy system with multimodal feedback adaptive regulation according to claim 1, wherein, It further includes: The model-based deep reinforcement learning framework based on safety constraints realizes online learning to solve the problem of adaptive optimization during the use of the system under the condition of few initial training data or no training data, so as to meet more personalized treatment parameters. Specifically: The basic model is initialized with prior biomedical knowledge. The initialized parameters include the state State: the working mode adopted by the patient, and the corresponding movement state and surface electromyogram state, the corresponding stimulation intensity, frequency, duration, and stimulation trigger conditions; During the use of the device, the treatment parameters and the corresponding movement state, electromyogram state, and user experience input are recorded. The model parameters are updated through online learning by regularly collecting data. In subsequent treatment processes, based on the updated model, real-time optimization of the output is realized through the reinforcement learning algorithm with safety constraints.
Citation Information
Cited By
Dynamic monitoring system for spinal cord injury reconstruction process
CN120732444A
Wireless micro-current tactile feedback game control equipment interaction method and system
CN120733343A
Wireless micro-current haptic feedback game control device interaction method and system
CN120733343B
Physical state detection and muscle electrical stimulation method based on intelligent wearable device
CN120939452A
Nerve regulation and control system and method for spinal cord ventral epidural area
CN121314072A