Intelligent artificial limb control method and system with feedback mechanism
By combining sensor clusters with biomechanical cognitive models and reinforcement learning strategies, the feedback signal of the intelligent prosthesis is dynamically adjusted, solving the problem of the feedback signal's inability to adapt in existing technologies and achieving precise control and improved safety in complex environments.
Patent Information
- Application Number
- CN202510924208.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-04
AI Technical Summary
The existing intelligent prosthetic feedback mechanism is unable to achieve adaptive real-time adjustment of feedback signal type, intensity and timing in complex dynamic environments, resulting in distorted, delayed or information overloaded user feedback information, and unable to accurately reflect the actual state and deviation of the prosthesis in the process of executing the intention, increasing the user's cognitive burden and weakening walking confidence and control effectiveness in complex environments.
By deploying a sensor cluster to collect terrain features, joint motion parameters and residual limb electromyographic signals, a feedback parameter set is generated by combining the biomechanical cognitive model with the reinforcement learning strategy model. Physiological tolerance screening and control stability verification are performed through the mutual inspection arbitration unit, and the feedback type, intensity and timing are dynamically adjusted to ensure accurate matching of feedback signals.
It achieves adaptive and precise matching of feedback signals in complex dynamic environments, improves the user's control accuracy and safety in complex terrain, and reduces the user's cognitive burden and the risk of control misalignment.
Smart Images

Figure CN120753844A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent prosthesis control, in particular to an intelligent prosthesis control method and system with a feedback mechanism. BACKGROUND
[0002] Modern intelligent prosthesis control technology has made significant progress, especially in integrating feedback mechanisms. Existing solutions generally can obtain prosthesis state or environmental information through integrated sensors and deliver basic feedback signals (such as pressure, position or simple vibration) to the user, aiming to fill the user's perception gap and assist in control decisions. However, these technologies have significant limitations in dealing with complex dynamic environments in the real world. Users' daily activities involve variable and unstructured terrain, such as uneven road surfaces, slopes or stairs, while gait requirements also quickly switch between actions such as walking, running, turning, etc. This results in the interaction state of the prosthesis and the environment, including force patterns, kinematics and dynamics parameters, presenting high-speed, nonlinear and complex changes. The core deficiency of current feedback mechanisms is the lack of intelligence and adaptability. Feedback strategies usually rely on pre-set fixed rules or simple linear mapping relationships, and cannot intelligently analyze and decide based on real-time perception of multi-dimensional, dynamic interaction states (such as simultaneous impact, side slip and joint angle changes). The modal selection, intensity adjustment, action location and output timing of the feedback signal often remain static or respond slowly, making it difficult to accurately match the rapidly changing actual situation. As a result, the feedback information received by the user is distorted, delayed or overloaded, and cannot accurately reflect the real state and deviation during the execution of the user's intention.
[0003] This mismatch between feedback and state directly leads to inaccurate decisions by the user or control system, increases the user's cognitive burden, and severely weakens the walking confidence and control effectiveness in complex environments. Therefore, it is urgent to break through the static feedback mode of existing technologies and improve the real-time intelligence and adaptability of the feedback mechanism. The technical problem to be solved is: how to realize the adaptive real-time adjustment of the type, intensity and timing of the intelligent prosthesis feedback signal in a complex dynamic environment, so as to accurately match the instantaneous state of the prosthesis-environment interaction and the user's control requirements. SUMMARY
[0004] To achieve the above purpose, the present application is implemented by the following technical scheme: an intelligent prosthesis control method with a feedback mechanism, comprising the following steps: S1: acquisition step: through the sensor cluster deployed on the prosthesis foot bottom, ankle joint and knee joint, real-time acquisition of terrain features, joint motion parameters and residual limb electromyographic signals; S2: decision step: input multi-source data into a dynamic feedback decision module, simultaneously execute a biomechanics cognitive model and a reinforcement learning strategy model, respectively generate a theoretical optimal feedback parameter set and a real-time adaptive feedback parameter set; S3: mutual inspection arbitration step: physiological tolerance screening and control stability simulation verification are performed on the two parameter sets, a fusion instruction is output when the parameter difference is less than a safety threshold, otherwise a weighted arbitration based on user operation habits is triggered; S4: execution step: the multi-modal actuator array of the residual limb interface is driven according to the final instruction, and the feedback type, intensity, position and output timing are adjusted.
[0005] Preferably, the dynamic feedback decision module comprises: a biomechanics cognition unit that stores a mapping relationship library of joint kinematic chain and neural response; a reinforcement learning strategy unit that adopts a deep deterministic policy gradient algorithm with a hierarchical reward mechanism; a mutual inspection arbitration unit connected to the above two units and comprising a physiological tolerance screener and a control stability evaluator.
[0006] Preferably, the operation of the biomechanics cognition unit comprises: matching a pre-defined feedback mode combination rule according to the terrain characteristics; outputting a parameter set containing feedback type priority, intensity safety interval and timing reference window.
[0007] Preferably, the operation of the reinforcement learning strategy unit comprises: real-time receiving a six-dimensional tensor representing the environment-prosthesis-user interaction state; outputting feedback type switching threshold, intensity gradient coefficient and timing offset.
[0008] Preferably, the mutual inspection arbitration unit step comprises the following steps: Step a: physiological tolerance screening: comparing the feedback intensity parameter with the user's pain threshold database; Step b: control stability simulation: predicting the plantar pressure trajectory after executing the parameter; Step c: generating a weighted fusion instruction when both model outputs pass the verification and the Euclidean distance of the parameters is less than the tolerance; Step d: if the verification fails or the distance exceeds the limit, then the weight arbitration is allocated according to the user's high-frequency operation scene preference.
[0009] Preferably, the execution step comprises: modulating the tactile vibration frequency to a preset vibration frequency range within a gait cycle, and / or controlling the electric stimulation current within a preset current intensity range; the phase deviation of the feedback pulse sequence is less than a preset phase threshold proportion of the gait cycle.
[0010] Preferably, it further comprises a closed-loop verification step: Real-time monitoring of plantar pressure trajectory after feedback intervention, if the trajectory deviates more than three times in a row, the set threshold is rolled back to the previous valid parameters; Periodically optimize decision model parameters based on user walking confidence score.
[0011] An intelligent prosthesis control system implementing an intelligent prosthesis control method with a feedback mechanism, comprising: A sensor cluster module containing a plantar three-dimensional force sensor, a joint inertial measurement unit, and a myoelectricity acquisition electrode; A dynamic feedback decision module, which is communicatively connected to the sensor cluster module and contains a biomechanics cognition unit, a reinforcement learning strategy unit, and a mutual inspection arbitration unit; A multi-modal actuator array module, which embeds a piezoelectric vibration plate and a neuroelectricity stimulation electrode in the residual limb receiving cavity.
[0012] Preferably, the mutual inspection arbitration unit comprises: A physiological tolerance screener, which accesses a user historical pain threshold database; A control stability evaluator, which is built-in with a gait trajectory prediction algorithm; A weight distributor, which stores a user scene operation preference weight table.
[0013] Preferably, it further comprises a safety switching module: When the mutual inspection arbitration unit outputs conflicting instructions for three times in a row, switch to a conservative mode that only enables the biomechanics cognition unit, and activate a remote expert communication interface.
[0014] The present application provides an intelligent prosthesis control method and system with a feedback mechanism. It has the following beneficial effects: The intelligent prosthesis control method and system with a feedback mechanism, through the dual-drive architecture of the biomechanics cognition model and the reinforcement learning strategy model, combined with the dual verification mechanism of the mutual inspection arbitration unit, realizes the adaptive and accurate matching of feedback signal types, intensity and timing in complex dynamic environments. The biomechanics model ensures that the feedback parameters meet the clinical safety boundaries, the reinforcement learning model optimizes the environmental adaptability in real time, and the mutual inspection arbitration resolves model conflicts through physiological tolerance screening and control stability prediction, solving the problems of control inaccuracy, high user cognitive load and walking confidence decline caused by traditional static rule feedback. Users can obtain real-time synchronized tactile / electricity stimulation feedback with the prosthesis-environment interaction state, effectively improving the control accuracy and safety in complex terrain. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 It is an intelligent prosthesis control module interaction diagram with a feedback mechanism of the present application; Figure 2 It is a flowchart of the present application. DETAILED DESCRIPTION
[0016] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0017] Please refer to Figure 1 and Figure 2 The present application provides a technical solution: an intelligent prosthesis control method with a feedback mechanism, comprising the following steps: S1: acquisition step: through the sensor cluster deployed on the prosthesis foot bottom, ankle joint and knee joint, real-time acquisition of terrain features, joint motion parameters and residual limb electromyographic signals; S2: decision step: input the multi-source data into the dynamic feedback decision module, simultaneously execute the biomechanics cognitive model and the reinforcement learning strategy model, respectively generate the theoretical optimal feedback parameter set and the real-time adaptive feedback parameter set; S3: mutual inspection arbitration step: physiological tolerance screening and control stability simulation verification are performed on the two parameter sets, when the parameter difference is less than the safety threshold, the fusion instruction is output, otherwise the weighted arbitration based on user operation habits is triggered; S4: execution step: according to the final instruction, drive the multi-modal actuator array of the residual limb interface, adjust the feedback type, intensity, position and output timing.
[0018] It needs to be further explained that in the specific implementation process, through the three-dimensional force sensor of the prosthesis foot bottom, the inertial measurement unit at the joint and the electromyographic electrode on the surface of the residual limb receiving cavity, the terrain ruggedness, joint angle and angular velocity change data, and user muscle activation signal are collected in real time; the above multi-source heterogeneous data is input into the dynamic feedback decision module, and the biomechanics cognitive model and the reinforcement learning strategy model are started simultaneously: the biomechanics cognitive model matches the current terrain features and motion state according to the pre-constructed joint motion chain-neural response mapping library, outputs the theoretical optimal feedback parameter set, including the recommended feedback type combination, intensity safety interval, timing reference window, wherein the feedback type combination includes vibration + electric stimulation, the intensity safety interval includes electric stimulation 0.1-2mA, and the timing reference window includes triggering 50ms after heel strike; the reinforcement learning strategy model dynamically calculates the feedback type switching threshold, intensity gradient coefficient and timing offset based on the real-time generated six-dimensional space-time state tensor through the hierarchical reward mechanism, and generates the real-time adaptive parameter set.
[0019] The two sets of parameters are subjected to mutual inspection arbitration: first, physiological tolerance screening is performed, the feedback intensity value is compared with the user's historical pain threshold database, and if it exceeds the upper limit of tolerance, the parameter instruction is automatically intercepted; second, control stability simulation is performed, the trajectory of the center of pressure on the sole after executing the parameter is predicted based on the current prosthesis state, and if the trajectory deviation exceeds the gait stability tolerance, the scheme is rejected; when both model outputs pass the verification and the parameter difference value is less than the preset safety tolerance, the two parameter sets are weighted to generate a fused instruction; if the verification fails or the difference exceeds the limit, the weight is allocated according to the user's high-frequency operation scene preference, for example, the reinforcement learning strategy weight is significantly higher than the biomechanics model in the stair climbing scene, and a weighted arbitration instruction is generated.
[0020] The final decision instruction drives the multi-modal actuator array module embedded in the residual limb receiving cavity: if the instruction requires tactile feedback, the frequency of the piezoelectric vibration piece is modulated to the range of 100-300 Hz within 20 ms; if the instruction contains electrical stimulation, the current intensity is accurately controlled in the safe range of 0.1-5 mA; the deviation of the trigger time point of the feedback pulse sequence from the reference time of the gait cycle is strictly less than one-twentieth of the gait cycle.
[0021] Through the progressive logic of double-model generation-double-verification-dynamic arbitration, real-time and accurate matching of feedback parameters in complex environments is ensured, solving the problem of control inaccuracy caused by feedback static lag.
[0022] The dynamic feedback decision module includes: a biomechanics cognition unit that stores a mapping relationship library of joint kinematic chains and neural responses; a reinforcement learning strategy unit that uses a deep deterministic policy gradient algorithm with a hierarchical reward mechanism; a mutual inspection arbitration unit that connects the two units and includes a physiological tolerance screener and a control stability evaluator.
[0023] It should be further noted that in the specific implementation process, the intelligent prosthesis control system includes a multi-physical field sensor cluster module, which is deployed on the three-dimensional force sensor of the prosthesis sole to capture the ground reaction force distribution and terrain inclination in real time, the inertial measurement unit of the ankle joint and the knee joint to collect joint angular acceleration and motion trajectory, and the electromyography electrode on the surface of the residual limb receiving cavity to monitor the user's muscle activation pattern; all sensor data is transmitted to the dynamic feedback decision module through a high-speed bus. The biomechanics cognition unit of the module has a joint kinematic chain-neural response mapping library, which calls a predefined feedback rule set according to the terrain classification result; the reinforcement learning strategy unit runs a deep deterministic policy gradient algorithm with a hierarchical reward mechanism, inputs the fused spatiotemporal state tensor into the policy network, and dynamically outputs the feedback adjustment parameters, wherein the spatiotemporal state tensor includes the terrain risk level, the joint load coefficient, and the user's intention label; The mutual inspection arbitration unit as the core hub receives the output parameters of the two units in real time, the physiological tolerance screening device accesses the user's historical pain threshold database such as the individualized upper limit of electrical stimulation of amputees, the stability evaluator predicts the foot pressure trajectory deviation based on the gait dynamics model, when the difference between the two model parameters exceeds the safety tolerance, the user's high-frequency operation scene preference is generated, such as prioritizing reinforcement learning strategy when going up and down stairs, and the weighted arbitration instruction is generated. Through the multi-source perception-double model parallel decision-making-mutual inspection arbitration closed-loop system architecture, it is ensured that the feedback control not only meets the biomechanical safety constraints, but also can adapt to complex environmental changes in real time.
[0024] The operation of the biomechanical cognitive unit includes: According to the terrain feature, the pre-defined feedback mode combination rule is matched; Output parameter set containing feedback type priority, intensity safety interval and timing reference window.
[0025] It needs to be further explained that in the specific implementation process, the operation process of the biomechanical cognitive unit is: when the sensor cluster module detects that the terrain feature is uneven road surface, the preset "rough terrain mode" is automatically activated, the corresponding feedback rule set is retrieved from the joint motion chain-nerve response mapping library, if the forefoot area load increases suddenly and the ankle inversion angle exceeds the limit, it is determined that there is a side slip risk, and the feedback type priority is "high-frequency vibration + medium-intensity electrical stimulation"; For the stair climbing scene, when the knee flexion angle is greater than 50 degrees and the angular velocity is positive, the "climbing stairs mode" rule is triggered, and the timing reference window is set to start vibration feedback 30 milliseconds after the toe touches the ground, and superimpose electrical stimulation confirmation signal when the knee is stretched to the maximum angle; All output parameters contain intensity safety interval, for example, the upper limit of electrical stimulation intensity is dynamically set according to user's historical pain data, if the user's maximum tolerance current last week is 2.1 mA, the output interval is automatically limited to 0.1-2.0 mA.
[0026] Through the four-layer progressive logic of terrain mode, risk diagnosis, timing anchoring and individualized constraints, clinical medical knowledge is converted into executable control rules, the timing reference window is bound with the action phase, and the problem of feedback disconnection with action is solved.
[0027] The operation of the reinforcement learning strategy unit includes: Real-time reception of six-dimensional tensor representing the state of environment-prosthesis-user interaction; Output feedback type switching threshold, intensity gradient coefficient and timing offset.
[0028] It needs to be further explained that in the specific implementation process, the reinforcement learning strategy unit receives the fusion generated six-dimensional space-time state tensor in real time, which contains interactive features such as terrain risk level, joint load coefficient, and user electromyographic intention label; when detecting that the user electromyographic signal shows an acceleration intention and the terrain radar identifies a rain-wet subway staircase, the deep deterministic policy gradient algorithm of the layered reward mechanism starts multi-objective optimization: the first layer reward function calculates the deviation of the foot pressure center trajectory from the ideal gait, and the smaller the deviation, the higher the reward value; The second layer reward function evaluates the user's real-time comfort score, and dynamically adjusts the weight according to the pain feedback record in the historical data under similar scenes; the third layer reward function is associated with the energy efficiency of the prosthetic motor, and suppresses the high-power parameter combination.
[0029] Based on the weighted sum of the triple rewards, the policy network outputs the feedback type switching threshold, for example, adjusting the vibration warning threshold from the regular 15 Newton to 8 Newton in the rain-slip staircase scene, realizing earlier risk warning; At the same time, the intensity gradient coefficient is generated, and if the current ankle joint load reaches 80% of the safety limit, the electric stimulation intensity is automatically increased to 90% of the upper limit of the tolerance interval; For the time offset, when the user's step frequency suddenly increases by 20%, the algorithm compresses the feedback window by 10 milliseconds to trigger the pulse sequence in advance to ensure synchronization.
[0030] Through the reinforcement learning decision chain of environment perception-layered optimization-dynamic output, real-time accurate adaptation of feedback parameters within the biomechanical safety boundary is realized. The environmental response logic of the three-layer reward function is highlighted, especially the threshold adjustment and step frequency-driven time compression in the rain-slip scene, solving the problem that fixed rules cannot cope with sudden environmental changes. The load-related design of the intensity gradient coefficient reflects the safety cooperation with the biomechanical model, laying the foundation for subsequent mutual inspection arbitration.
[0031] The mutual inspection arbitration unit step includes the following steps: Step a: physiological tolerance screening: compare the feedback intensity parameters with the user's pain threshold database; Step b: control stability simulation: predict the foot pressure trajectory after executing the parameters; Step c: when both model outputs pass the verification and the parameter Euclidean distance is less than the tolerance, generate a weighted fusion instruction; Step d: if the verification fails or the distance is out of limit, assign weights according to the user's high-frequency operation scene preference to arbitrate.
[0032] It needs to be further explained that in the specific implementation process, when the biomechanical cognitive unit outputs the theoretical optimal parameter set and the adaptive parameter set generated by the reinforcement learning strategy unit conflict, the mutual inspection arbitration unit starts the double verification process, wherein the theoretical optimal parameter set contains the recommended foot vibration intensity 3 levels + electric stimulation 1.5 mA in the stair mode, and the adaptive parameter set generated by the reinforcement learning strategy unit contains the recommended vibration intensity 4 levels + electric stimulation 2.0 mA in the rain slip environment.
[0033] The mutual inspection arbitration unit starts the double verification process, including: first performing physiological tolerance screening, calling the user's last week electric stimulation pain threshold record, if the historical data shows that 2.0 mA stimulation has caused the pain score to exceed 7 points, wherein the full score is 10, the parameter instruction is intercepted; After screening, enter the control stability simulation, predict the foot pressure trajectory after executing vibration intensity 4 levels based on the current ankle angle and ground friction coefficient, if the simulation shows that the pressure center offset exceeds the gait stability tolerance, such as the forefoot area load drops by 30%, the reinforcement learning scheme is denied.
[0034] When both schemes pass the verification and the parameter difference value is less than the safety tolerance, such as vibration intensity difference 1 level, electric current difference 0.3 mA, the user's set stair scene preference is followed, that is, the reinforcement learning weight dominates, and the weighted instruction is generated, such as the final vibration intensity 3.8 levels + electric stimulation 1.7 mA; If the simulation shows that the biomechanical scheme will cause trajectory deviation, such as insufficient vibration intensity to warn slip, the reinforcement learning scheme is completely adopted and online fine tuning is triggered to reduce the vibration intensity switching threshold in similar subsequent scenes.
[0035] Through the four-order decision chain of "physiological screening-stability verification-scene arbitration-model evolution", the risk response is realized in the emergency case. The "stability prediction based on physical simulation" and "scene preference driven dynamic arbitration" are highlighted, which solves the pain point that fixed rules cannot resolve model conflicts.
[0036] The execution steps include: Modulate the tactile vibration frequency to a preset vibration frequency range of 100 to 300 Hz within a gait cycle, and / or control the electric stimulation current in a preset current intensity range of 0.1 to 5 mA; The phase deviation of the feedback pulse sequence is less than one-twentieth of the preset phase threshold proportion of the gait cycle.
[0037] It needs to be further explained that in the specific implementation process, the multi-modal actuator array module is driven according to the final arbitration instruction: when the instruction requires tactile feedback, the piezoelectric vibration piece is started at the heel landing phase of the gait cycle, and the vibration frequency is accurately modulated to the range of 100 to 300 Hz within 20 milliseconds through a phase-locked loop, and if the arbitration instruction specifies the "rain slip warning mode", a high-frequency vibration of 280 Hz is used to enhance the risk perception; for the electric stimulation channel, the current intensity is strictly limited in the safety interval of 0.1 to 5 mA, when the biomechanical model output intensity upper limit is 2.0 mA and the arbitration instruction requires 1.7 mA, the electric stimulation controller automatically loads the gradient rise algorithm to avoid sudden stimulation.
[0038] The timing control core is to synchronize the pulse sequence with the gait phase: in the stair climbing scene, the vibration feedback is strictly constrained to be triggered within 30±5 milliseconds after the toe touches the ground, and the timing reference is calibrated in real time through the insole switch sensor; when the user's step frequency accelerates and the gait cycle is shortened to 80% of the original cycle, the feedback window length is automatically compressed to one-twentieth of the gait cycle, ensuring that the vibration pulse is completed before the knee joint flexion peak. Through the timing control chain of "event anchoring-environment adaptation-safety boundary", the abstract time accuracy is converted into detectable biomechanical events and physical tolerance. It solves the defect that fixed timing cannot adapt to gait mutations.
[0039] It also includes a closed-loop verification step: Real-time monitoring of the insole pressure trajectory after feedback intervention, if the trajectory deviation exceeds the set threshold for three consecutive times, roll back to the previous valid parameters; Periodically optimize decision model parameters based on user walking confidence score.
[0040] It needs to be further explained that in the specific implementation process, the insole pressure center trajectory change is monitored in real time after the feedback instruction is executed, if it is detected that the trajectory deviation of three consecutive steps exceeds the dynamically set threshold, such as the lateral drift of the pressure center being greater than 15 mm in the stair climbing scene, then automatically roll back to the previous valid decision parameters and mark this exception; simultaneously start decision traceability analysis, when the rollback event is caused by the biomechanical model parameters, freeze the scene rule base and strengthen the learning strategy weight to 90%; if the deviation is caused by the reinforcement learning parameters, trigger the high penalty term in the reward function and reduce the exploration rate. Aggregate the user walking confidence score and the prosthetic energy consumption benchmark value under the same terrain every week: when the average score of a specific scene is less than 6 for two consecutive weeks and the energy consumption increases by more than 20%, optimize the rule base of the biomechanical cognitive unit offline, expand the criterion of gravel road side slip to include the ground humidity factor, and recalibrate the hierarchical reward weight of the reinforcement learning strategy, so that the comfort factor proportion is increased to the safety limit extreme value, wherein the user walking confidence score is based on a 10-point scale.
[0041] Through the three-order closed loop of "real-time monitoring-failure tracing-long-term iteration", the dynamic optimization of the core decision mechanism is visualized. By using the dynamic tolerance of terrain self-adaptation and the differentiated compensation of model failure, the problem that the static system cannot continuously adapt to user changes is solved.
[0042] An intelligent prosthesis control system implementing an intelligent prosthesis control method with a feedback mechanism, comprising: A sensor cluster module including a plantar three-dimensional force sensor, a joint inertial measurement unit, and a myoelectricity acquisition electrode; A dynamic feedback decision module in communication connection with the sensor cluster module and including a biomechanics cognition unit, a reinforcement learning strategy unit, and a mutual inspection arbitration unit; A multi-modal actuator array module embedded with a piezoelectric vibration sheet and a nerve electrical stimulation electrode in a residual limb receiving cavity.
[0043] It needs to be further explained that in the specific implementation process, the system captures the ground reaction force distribution of the forefoot, heel, and lateral edge area in real time through the embedded plantar three-dimensional force sensor matrix, synchronously collects three-axis acceleration, three-axis angular velocity, and three-axis magnetic force data from the nine-axis inertial measurement unit of the ankle joint and knee joint, and extracts surface electromyography signal features at a sampling rate of 2000 Hz from the myoelectricity electrode array covering the proximal nerve-rich area of the residual limb receiving cavity. The above multi-source data is transmitted to the core processor of the dynamic feedback decision module through the anti-interference shielded cable. The processor has a dual-core architecture, the biomechanics cognition unit runs on the real-time operating system core, and the pre-compiled joint kinematic chain mapping library is called; the reinforcement learning strategy unit is deployed on the Linux core and executes the deep deterministic policy gradient algorithm; the space-time state tensor and parameter proposal are exchanged between the two cores through shared memory.
[0044] The mutual inspection arbitration unit, as an independent coprocessor, receives the output of the dual core, starts the physiological tolerance screener and control stability evaluator, and sends the arbitration result to the piezoelectric vibration sheet array and programmable electrical stimulation electrode embedded in the receiving cavity via the CAN bus. The electrical stimulation channel is equipped with an overcurrent protection circuit to ensure that the output is strictly limited to the 0.1-5 mA safe range. The physiological tolerance screener is used to access the encrypted user pain database, and the stability evaluator is used to call the gait dynamics simulation engine.
[0045] The mutual inspection arbitration unit includes: A physiological tolerance screener accessing a user historical pain threshold database; A control stability evaluator with a built-in gait trajectory prediction algorithm; A weight distributor storing a user scenario operation preference weight table.
[0046] It should be further explained that, in the specific implementation process, the mutual inspection arbitration unit accesses the encrypted user pain threshold curve database in real time through the physiological tolerance screener. When the reinforcement learning strategy unit outputs the electrical stimulation intensity parameter, it automatically matches the historical pain record of the same anatomical point of the current user: if the parameter value exceeds the highest value of the last three painless stimulations, such as 1.8mA at the user's gastrocnemius point is the safety upper limit, the instruction is immediately intercepted and marked as a high-risk operation; the control stability evaluator is input through the screened parameter set, and the module calls the gait dynamics lightweight engine to simulate the plantar pressure trajectory within the next 200 milliseconds based on the current ankle joint angle, ground friction coefficient and feedback intensity parameters: if the prediction shows that executing vibration intensity level 4 in an ice environment will cause the lateral displacement of the pressure center to exceed the ice surface tolerance limit, the plan is automatically rejected and the high penalty item of the reinforcement learning reward function is triggered. The weight allocator dynamically adjusts the arbitration weight based on the pre-stored risk-scenario matrix: when it detects that the user is in the peak scenario of subway transfer stairs, the "efficiency-safety" preference label is activated, and the weight of the reinforcement learning strategy is increased to 3 times that of the biomechanical model; if the environmental radar identifies a quiet home environment, it switches to "comfort priority" mode, and the weight of the biomechanical model increases to 2.5 times that of reinforcement learning.
[0047] Through a triple-checking mechanism of "physiological real-time interception, physical simulation prediction, and scenario-based weight allocation," abstract decisions are transformed into quantifiable safety operations. Highlighting "painless, historically high-value dynamic defense" and "multi-parameter gait simulation veto," this solves the problem of a single safety mechanism being unable to address complex risks.
[0048] Also includes the Safe Switch module: When the mutual inspection arbitration unit outputs a conflict instruction for three consecutive times, switch to a conservative mode that only enables the biomechanics cognitive unit, and activate the remote expert communication interface. It needs to be further explained that, in the specific implementation process, when the mutual inspection arbitration unit outputs a conflict instruction for three consecutive times, such as the biomechanics model continuously suggesting 1.6 mA electric stimulation and the reinforcement learning strategy insisting on a 2.0 mA scheme and both double verification failures, the safety switching module immediately starts a three-level emergency response: first, forcibly disconnect the decision-making authority of the reinforcement learning strategy unit, and downgrade the system to a conservative mode that only runs the biomechanics cognitive unit. In this mode, the feedback type is only retained for the basic vibration channel, and the upper limit of the electric stimulation intensity is locked at 80% of the user's historical painless average, such as 1.44 mA if the user's average tolerance is 1.8 mA, and the timing window is expanded to one-tenth of the gait cycle to reduce the synchronization accuracy requirement. At the same time, activate the remote expert communication interface, transmit the current environment multi-modal data stream through an encrypted channel, including terrain radar point cloud, joint motion trajectory and failed decision log, receive temporary rule patches issued remotely, such as special vibration frequency schemes for subway gate metal grid terrain, and if no manual instructions are received within 10 minutes, automatically enable a minimalist feedback strategy based on terrain classification, which fixes the 200 Hz standard vibration at the 50% phase point of each step cycle when smooth hard ground is detected, and gradually restores the dual-model decision-making after the system self-check is passed.
[0049] Through the four-stage emergency system of "conflict diagnosis - safety degradation - human-machine cooperation - autonomous recovery", the core decision failure risk is controlled within the clinically acceptable range, solving the defect problem of user loss of control caused by system deadlock.
[0050] It needs to be further explained that, in the specific implementation process, the intelligent prosthetic control system captures environmental terrain features, prosthetic kinematics parameters and user electromyographic intention signals in real time through sensor cluster modules deployed on the sole, joints and residual limb receiving cavity. After fusion, multi-source heterogeneous data generates a spatio-temporal tensor input dynamic feedback decision module representing the environment-prosthesis-user interaction state. The module runs the biomechanics cognitive model and the reinforcement learning strategy model in parallel: the biomechanics model matches the current scene based on the pre-constructed joint motion chain and neural response mapping library to output the theoretically optimal feedback parameters including type combination, intensity safety interval and timing reference window; the reinforcement learning model dynamically calculates the feedback type switching threshold, intensity gradient coefficient and timing offset based on the deep deterministic policy gradient algorithm combined with the hierarchical reward function.
[0051] The parameter set output by the two models is input into the mutual inspection arbitration unit to perform double verification: the physiological tolerance screening device calls the user's historical pain threshold database to compare the intensity parameter. If it exceeds the individual's upper limit of painlessness, the instruction is immediately intercepted; the control stability evaluator predicts the foot pressure trajectory after executing the parameter based on the gait dynamics model. When the deviation exceeds the terrain self-adaptive tolerance, the scheme is rejected. The parameter set that passes the verification is fused and output according to the user's scene preference weight if the difference value is less than the safety tolerance; if the difference is over the limit or the verification fails, the weight distributor is started to dynamically tilt the arbitration weight according to the high-frequency operation scene type, for example, the reinforcement learning weight is significantly higher than the biomechanical model in the subway transfer stair scene.
[0052] The final arbitration instruction drives the multi-modal actuator array module: the piezoelectric vibration piece is modulated to the target frequency range within twenty milliseconds, and the electrode outputs a safe current intensity under the constraint of the gradient loading algorithm. The feedback timing is strictly anchored to the gait key phase event, such as triggering within a fixed delay window after the toe touches the ground, and the window duration is dynamically compressed according to the user's step frequency to ensure synchronization with the action.
[0053] The system monitors the foot pressure trajectory after feedback intervention in real time. If the deviation exceeds the limit for multiple steps in a row, it will roll back to the previous valid parameter and trace back to mark the failed model. Long-term aggregation of user walking confidence scores and prosthetic energy consumption data: when the score in a specific scene continuously falls below the set threshold and the energy consumption abnormally rises, offline optimize the biomechanical rule library and recalibrate the reinforcement learning reward weight.
[0054] When the mutual inspection arbitration unit outputs conflicting instructions for multiple times in a row, the safety switching module forces a downgrade to a conservative mode that only enables the biomechanical cognitive unit, locks the basic vibration channel, limits the intensity upper limit to the historical painless average proportion value, relaxes the timing accuracy requirement, and activates the remote expert interface to transmit environmental data streams to receive artificial rule patches. If it does not respond within the timeout, it enables the terrain classification minimal strategy to maintain basic tactile feedback until the system self-check passes, and then gradually restores the dual-model decision weight.
[0055] Further explanation is needed. In the implementation process, the biomechanical cognitive model calls the predefined rule set according to the terrain classification result, for example: when uneven road is detected and the foot pressure distribution is abnormal, activate the rugged terrain mode to output high-frequency vibration and electric stimulation combination; the stair climbing scene is bound to the toe touch event to set the vibration trigger window. The reinforcement learning strategy model dynamically adjusts the parameters according to the real-time space-time tensor, such as automatically lowering the vibration warning threshold in a slippery rain environment and increasing the electric stimulation intensity gradient coefficient when walking on ice. When the outputs of the two models conflict, the mutual inspection arbitration unit prioritizes physiological tolerance screening: if the electric stimulation intensity proposed by the reinforcement learning exceeds the highest value of the user's painless record at the same anatomical point in the previous week, it is intercepted and the output upper limit of that point is lowered.
[0056] The control stability evaluator constructs a virtual gait environment, inputting the current ankle angle, the ground friction coefficient, and the feedback strength parameters. Using a lightweight dynamics engine, it simulates the plantar pressure trajectory over the next 200 milliseconds. If the prediction indicates that executing these parameters will cause the center of pressure to shift beyond the ground tolerance limit—for example, a shift greater than 15 millimeters on icy surfaces—the proposed solution is automatically rejected, and the shift is converted into a penalty term in the reinforcement learning reward function. The simulation results are also used for timing optimization. As the user's cadence increases, the feedback window duration is proportionally compressed to ensure that the vibration pulse is output before the peak of knee flexion.
[0057] Three consecutive mutual inspection arbitration failures trigger a Level 3 response: Reinforcement learning decision-making authority is immediately revoked, and the system switches to a biomechanically conservative mode. In this mode, only a single vibration channel is retained, electrical stimulation is disabled, and the intensity limit is locked to a fixed percentage of the individual's historical pain-free average. A remote expert collaboration interface is simultaneously activated, encrypting the transmission of terrain point cloud data, joint motion trajectories, and decision failure logs. After remote manual analysis of special scenarios such as metal grating terrain, a vibration frequency adjustment patch is issued. If no response is received after a communication timeout, the system automatically matches the terrain type, employing a minimalist feedback strategy for smooth and hard surfaces, triggering standard frequency vibration at a fixed phase point within each step cycle, until the system self-checks and confirms that the environmental risk has been eliminated. The dual-model weights are then restored using a three-stage gradient.
[0058] A real-time rollback mechanism is associated with model responsibility determination: If trajectory deviation is caused by biomechanical parameters, the scenario rule base is frozen and the reinforcement learning weight is increased to a dominant position. If the reinforcement learning parameters cause failure, the exploration penalty coefficient is increased. Long-term optimization is based on a multi-dimensional evaluation matrix. When the user's walking confidence score continues to fall below the set threshold and the energy consumption on the same terrain is significantly higher than the baseline value, offline rule base iteration is triggered. For example, the gravel road judgment criteria can be expanded to include surface moisture factors, and the reinforcement learning comfort level is increased, with the reward weight increased to the maximum safe limit.
[0059] Through the architecture of "dual-model parallel decision-making, dual safety verification, scenario-based dynamic arbitration, and human-machine collaboration in fault conditions," the system covers the entire technology chain from environmental perception to safety control. It emphasizes individualized defense mechanisms through physiological tolerance screening and stability prediction technology based on physical simulation, transforming clinical medical knowledge into computable safety rules. Furthermore, it uses reinforcement learning to achieve real-time optimization of feedback parameters in complex environments.
[0060] A method for controlling an intelligent prosthesis with a feedback mechanism comprises the following steps: Step S1: Using the plantar 3D force sensor, joint inertial measurement unit, and residual limb electromyographic electrodes, terrain features, joint motion parameters, and user electromyographic intention signals are collected in real time and fused to generate a spatiotemporal state tensor. Step S2: Synchronously input the spatiotemporal state tensor into the biomechanical cognitive model and the reinforcement learning strategy model: The biomechanics cognitive model calls the joint kinematic chain-neural response mapping library, matches the terrain scene output theoretical optimal feedback parameter set, including type combination, intensity safety interval and timing reference window; the reinforcement learning strategy model dynamically calculates the feedback type switching threshold, intensity gradient coefficient and timing offset based on the hierarchical reward mechanism, and generates a real-time adaptive parameter set; Step S3: mutual inspection arbitration is performed on the two model output parameter sets: Physiological tolerance screening: compare the feedback intensity parameters with the user's historical pain threshold database to intercept out-of-limit instructions; Control stability simulation: input joint angle and ground friction coefficient to predict foot pressure trajectory offset, and reject out-of-tolerance schemes; When the parameter difference is less than the safety tolerance, the output is fused according to the user's scene preference weight; if the difference is out of limit or the verification fails, the arbitration weight is dynamically allocated according to the high-frequency operation type; Step S4: drive the piezoelectric vibration piece and the electric stimulation electrode according to the final instruction: The vibration frequency is modulated to the target range within twenty milliseconds, and the timing anchor is set to the gait key phase event, such as the delay window after the toe touches the ground; The electric stimulation intensity is safely output by the gradient loading algorithm, and the window length is compressed in proportion with the step frequency acceleration; Step S5: real-time monitoring of foot pressure trajectory after feedback intervention: If the continuous multiple-step offset exceeds the dynamic threshold, roll back to the previous valid parameters; Traceability determines the failure responsibility model, triggers corresponding weight compensation or rule freezing; Step S6: periodically aggregate user walking confidence scores and prosthetic energy consumption data: When the specific scene score continuously falls below the set threshold and the energy consumption abnormally rises, offline optimize the biomechanics rule library and calibrate the reinforcement learning reward weight; Step S7: if the mutual inspection arbitration fails continuously for multiple times, start the safety bottom: Force switch to only biomechanics conservative mode, lock the basic vibration channel and individualized intensity upper limit; Activate the remote expert interface to transmit environmental data streams and load artificial rule patches; Step S8: when the remote collaboration times out and does not respond, enable the terrain classification minimalist strategy: trigger fixed phase standard vibration on smooth hard ground; maintain basic feedback until the system self-check passes; Step S9: gradually restore the dual-model decision weight, and increase the reinforcement learning authority in three stages.
[0061] Through the dual driving architecture of the biomechanical cognitive model and the reinforcement learning strategy model, combined with the double verification mechanism of the mutual inspection arbitration unit, the adaptive and accurate matching of the feedback signal type, intensity and timing in the complex dynamic environment is realized. The biomechanical model ensures that the feedback parameters meet the clinical safety boundary, the reinforcement learning model optimizes the environmental adaptability in real time, and the mutual inspection arbitration resolves model conflicts through physiological tolerance screening and control stability prediction, solving the problems of control inaccuracy, high user cognitive load and decreased walking confidence caused by traditional static rule feedback. Users can obtain real-time synchronous tactile / electric stimulation feedback with the prosthetic-environment interaction state, effectively improving the control accuracy and safety in complex terrain.
[0062] It should be noted that the relational terms herein such as first and second and the like are used solely to distinguish one from another entity or action, without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.
[0063] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for controlling an intelligent prosthesis with a feedback mechanism, characterized in that: The steps include: S1: Acquisition step: Through the sensor cluster deployed on the sole, ankle and knee joints of the prosthetic limb, terrain features, joint motion parameters and residual limb electromyographic signals are acquired in real time; S2: Decision-making step: Multi-source data is input into the dynamic feedback decision module, and the biomechanical cognitive model and the reinforcement learning strategy model are synchronously executed to generate the theoretical optimal feedback parameter set and the real-time adaptive feedback parameter set respectively; S3: Mutual inspection and arbitration step: Physiological tolerance screening and control stability simulation verification are performed on the two parameter sets. When the parameter difference is less than the safety threshold, a fusion instruction is output; otherwise, a weighted arbitration based on user operation habits is triggered; S4: Execution step: driving the multimodal actuator array of the residual limb interface according to the final instruction, and adjusting the feedback type, intensity, position and output timing.
2. The intelligent prosthesis control method with a feedback mechanism according to claim 1, characterized in that: The dynamic feedback decision module includes: Biomechanical cognitive unit, which stores the mapping relationship between joint motion chain and neural response; Reinforcement learning policy unit, using a deep deterministic policy gradient algorithm with a hierarchical reward mechanism; The mutual inspection arbitration unit connects the two aforementioned units and includes a physiological tolerance screener and a control stability evaluator.
3. The intelligent prosthesis control method with a feedback mechanism according to claim 2, characterized in that: The operations of the biomechanical cognition unit include: Matching predefined feedback modal combination rules according to terrain features; The output includes a parameter set of feedback type priority, strength safety interval, and timing reference window.
4. The intelligent prosthesis control method with a feedback mechanism according to claim 2, characterized in that: The operations of the reinforcement learning strategy unit include: Receive in real time a six-dimensional tensor representing the state of the environment-prosthesis-user interaction; Output feedback type switching threshold, intensity gradient coefficient and timing offset.
5. The intelligent prosthesis control method with a feedback mechanism according to claim 2, characterized in that: The mutual inspection arbitration unit step includes the following steps when used: Step a: Physiological tolerance screening: Comparing feedback intensity parameters with the user pain threshold database; Step b: Control stability simulation: predict the plantar pressure trajectory after executing the parameters; Step c: When the outputs of both models pass the verification and the parameter Euclidean distance is less than the tolerance, a weighted fusion instruction is generated; Step d: If verification fails or the distance exceeds the limit, weight arbitration is performed based on the user's preference for high-frequency operation scenarios.
6. The intelligent prosthesis control method with a feedback mechanism according to claim 1, characterized in that: The execution steps include: Modulating the tactile vibration frequency to a preset vibration frequency range and / or controlling the electrical stimulation current to a preset current intensity range during the gait cycle; The phase deviation of the feedback pulse train is less than a preset phase threshold ratio of the gait cycle.
7. The intelligent prosthesis control method with feedback mechanism according to claim 1, characterized in that: Also includes closed-loop verification steps: Real-time monitoring and feedback of the plantar pressure trajectory after intervention. If the trajectory deviation exceeds the set threshold for three consecutive times, it will roll back to the previous valid parameters; Regularly optimize the decision model parameters based on the user's walking confidence score.
8. An intelligent prosthetic control system implementing any one of the methods of claims 1-7, characterized in that: include: Sensor cluster module, including plantar 3D force sensor, joint inertial measurement unit and myoelectric acquisition electrodes; Dynamic feedback decision module, which communicates with the sensor cluster module and includes a biomechanical cognition unit, a reinforcement learning strategy unit, and a mutual inspection and arbitration unit; Multimodal actuator array module, piezoelectric vibrating piece and nerve electrical stimulation electrode embedded in the residual limb socket.
9. The system according to claim 8, characterized in that: The mutual inspection arbitration unit includes: a physiological tolerance screener that accesses a database of users' historical pain thresholds; Control stability evaluator with built-in gait trajectory prediction algorithm; Weight allocator, which stores the user scenario operation preference weight table.
10. The system according to claim 9, characterized in that Also includes the Safe Switch module: When the mutual inspection arbitration unit outputs conflicting instructions three times in a row, it switches to a conservative mode in which only the biomechanical cognitive unit is enabled, and the remote expert communication interface is activated.
Citation Information
Patent Citations
Control system based on slippage sensing and spontaneous grasping of artificial limb
CN115300194A
Intelligent artificial limb control system and method with feedback mechanism
CN115645123A
Muscle stimulation adjusting method and system based on gait detection
CN120132218A
Simulation-real world feedback loop for learning robotic control policies
US10800040B1
Systems and methods for reinforcement learning control of a powered prosthesis
US20230066952A1