Spine posture correction control method based on reinforcement learning

CN122822218APending Publication Date: 2026-09-25NANJING SPORT INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610992010.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明的一个目的在于提出基于强化学习的脊柱矫姿控制方法,针对现有技术难以根据实时脊柱和躯干姿态变化对气囊贴合压力、热敷时长、热敷温度、电刺激输出和训练提醒进行动态协同调节的问题,提出了采集多源穿戴数据并构建多模态状态向量,利用时序特征提取网络获得姿态变化趋势和干预响应趋势,再通过约束强化学习控制器、预测安全屏蔽层和个体化耐受边界更新形成闭环控制的技术方案,本发明具备兼顾矫姿效果、佩戴舒适性和长期依从性的技术效果

Benefits of technology

1、通过将惯性测量、气囊接触压力、贴肤温度、电刺激回路、佩戴时长和训练提醒响应数据统一构建为多模态状态向量,并由时序特征提取网络输出姿态变化趋势和干预响应趋势,能够使后续控制动作基于连续姿态反馈和干预响应进行生成,减少仅依赖单一姿态提醒造成的控制滞后。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822218A_ABST
    Figure CN122822218A_ABST
Patent Text Reader

Abstract

The application discloses a spine posture correction control method based on reinforcement learning and belongs to the field of intelligent wearable posture correction control. In order to solve the problem that existing posture products are difficult to adjust air bags, hot compresses, electric stimulation and training reminders according to real-time spine and torso posture changes, the application forms a closed-loop posture correction control through multi-modal state construction, time sequence trend extraction, constraint reinforcement learning decision-making, prediction safety shielding and individualized tolerance boundary updating, and realizes the technical effects of giving consideration to posture correction effect, wearing comfort and long-term compliance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent wearable posture control, and more particularly to a spinal posture control method based on reinforcement learning. Background Technology

[0002] In the context of spinal and trunk posture management for children and adolescents, existing posture correction products typically rely on mechanical support structures, fixation straps, or simple posture reminders for intervention. While these solutions can constrain the wearer's posture to some extent, the intervention actions are mostly executed according to preset rules, making it difficult to coordinate adjustments based on real-time changes in the wearer's posture, local pressure, skin temperature changes, electrical stimulation circuit status, and training cooperation.

[0003] In actual use, airbag pressure, heat application temperature, heat application duration, electrical stimulation output, and training reminders interact with each other. If reminders are triggered based solely on a single postural deviation or if pressure, heating, and electrical stimulation actions are performed in a fixed manner, problems such as excessive pressure, local heat load accumulation, decreased tolerance to electrical stimulation, or frequent reminders causing interruptions can easily occur. This results in a lack of real-time feedback loop in the posture correction process and makes it difficult to balance comfort and long-term user compliance.

[0004] Therefore, there is a need for a spinal posture control method that can overcome the shortcomings of the existing technologies. Summary of the Invention

[0005] One objective of this invention is to propose a spinal posture correction control method based on reinforcement learning. Addressing the problem that existing technologies struggle to dynamically and collaboratively adjust airbag pressure, heat application duration, heat application temperature, electrical stimulation output, and training reminders based on real-time changes in spinal and trunk posture, this invention proposes a technical solution that collects multi-source wearable data and constructs a multimodal state vector. It then utilizes a temporal feature extraction network to obtain posture change trends and intervention response trends, and finally forms a closed-loop control system through a constrained reinforcement learning controller, a predicted safety shielding layer, and individualized tolerance boundary updates. This invention achieves a balance between posture correction effectiveness, wearing comfort, and long-term compliance.

[0006] This invention provides a reinforcement learning-based method for spinal posture control, comprising: S1. Collect inertial measurement data, airbag contact pressure data, skin temperature data, electrical stimulation circuit data, wearing time data, training reminder response data, and comfort feedback data from multiple torso positions of the wearer, and align them into periodic observation data according to the control cycle. S2. Construct a multimodal state vector based on the periodic observation data and obtain the individualized tolerance boundary corresponding to the wearer. The multimodal state vector includes spinal posture offset, trunk rotation angle, duration of abnormal posture, local contact pressure, heat load, electrical stimulation tolerance, compliance status and comfort feedback status. The individualized tolerance boundary includes target posture range, pressure tolerance boundary, heat compress tolerance boundary and electrical stimulation tolerance boundary. The electrical stimulation tolerance boundary is the upper limit of safe stimulation output. S3. Input the multimodal state vectors of multiple consecutive control cycles into the trained temporal feature extraction network to obtain the posture change trend and intervention response trend; S4. The posture change trend, the intervention response trend, and the individualized tolerance boundary input constraint reinforcement learning controller are used to generate candidate collaborative intervention actions based on comfort-safety dual constraint rewards. The candidate collaborative intervention actions include airbag pressure adjustment amount, target heat application temperature, heat application duration, stimulation output level, and training reminder timing. S5. Input the candidate collaborative intervention action into the prediction safety shielding layer, predict the risks of overpressure, overheating and overstimulation based on the current pressure, temperature, electrical stimulation circuit status, comfort feedback and change trend, and correct the candidate collaborative intervention action based on the individualized tolerance boundary to obtain the safe collaborative intervention action. S6. The safety collaborative intervention action is sent to the airbag pump valve, heating pad, electrical stimulation module and reminder module, and the multimodal state vector, the comfort-safety dual constraint reward and the individualized tolerance boundary are updated according to the data of the next control cycle to form a closed-loop posture correction control.

[0007] Optionally, S1 includes: The triaxial angular velocity and triaxial acceleration in the inertial measurement data are converted into a torso attitude angle sequence; The airbag contact pressure data is mapped into a local pressure sequence according to airbag partitions; The skin-contact temperature data is mapped into a temperature sequence according to the heating element partition; The electrical stimulation circuit data is mapped into a sequence of stimulation current, impedance, and intermittent states. The wearing duration data, the training reminder response data, and the comfort feedback data are mapped into a compliance observation sequence; When any sequence is missing in the current control period, the validity flag of that sequence is set to invalid and a similar sequence from the previous valid control period is used as a placeholder input. The validity flag is then entered into the periodic observation data along with the placeholder input.

[0008] Optionally, S2 includes: The spinal posture offset is determined based on the deviation of the trunk posture angle sequence from the target posture range; The torso rotation angle is determined based on the difference between the posture angles of adjacent torso positions; The duration of the posture abnormality is determined based on the number of control cycles in which the spinal posture offset is continuously outside the target posture range. The local bonding pressure is determined based on the local pressure sequence; The heat load is determined based on the temperature sequence, the heating duration determined by the heating element execution record of the previous control cycle, and the historical target heat application temperature. The tolerance to electrical stimulation is determined based on the stimulation current, impedance, intermittent state, and user pause records determined by user interaction records from the previous control cycle. The compliance status is determined based on the cumulative wearing time, training reminder response delay, number of intervention cancellations, and training completion rate. The comfort feedback state is determined based on the comfort feedback data; The corresponding state component is set with an invalidation flag or a weighting factor based on the validity flag.

[0009] Optionally, S3 includes: The multimodal state vectors of the multiple consecutive control cycles are input into the temporal feature extraction network, which includes temporal convolutional units and gated recurrent units; The temporal convolutional unit extracts the attitude shift, pressure, thermal load, and electrical stimulation tolerance changes between adjacent control cycles. The gated loop unit performs time-series fusion of the posture deviation change, the pressure change, the heat load change, the electrical stimulation tolerance change, and the compliance state change, and outputs the posture change trend and the intervention response trend, wherein the intervention response trend includes the pressure change trend, the heat load change trend, and the electrical stimulation tolerance change trend.

[0010] Optionally, S4 includes: The comfort-safety dual-constraint reward includes postural improvement benefits, training completion benefits, local pressure penalties, thermal load penalties, electrical stimulation abnormality penalties, user pause penalties, and wearing interruption penalties. The posture improvement benefit is determined based on the reduction in the spinal posture offset relative to the previous control cycle; The training completion benefit is determined based on the record of training reminders being responded to and training actions being completed; The local pressure penalty is determined based on the extent to which the local bonding pressure exceeds the pressure tolerance boundary and the number of durations. The heat load penalty is determined based on the extent to which the heat load exceeds the heat therapy tolerance boundary and the number of durations. The penalty for abnormal electrical stimulation is determined based on the magnitude by which the stimulation output level exceeds the upper limit of the stimulation output safety limit, the impedance abnormality marker, and the intermittent state. The user suspension penalty is determined based on records of user-triggered suspensions during intervention execution; The penalty for interruption of wear is determined based on the number of durations of the interruption. Furthermore, the constraint reinforcement learning controller includes a policy network, a value network, and a constraint cost network; The policy network outputs the probability distribution of actions based on the posture change trend, the intervention response trend, the individualized tolerance boundary, and the comfort-safety dual-constraint reward distribution. The constraint cost network generates constraint cost values ​​based on the local pressure penalty, the thermal load penalty, and the electrical stimulation anomaly penalty; When the constraint cost exceeds the preset constraint cost threshold, the selection probability of the corresponding action of pressurization, heating, prolonged hot compress or increased stimulation output level in the action probability distribution is reduced according to the preset probability decay coefficient, and the candidate collaborative intervention action is output.

[0011] Optionally, S5 includes: The predictive security shielding layer includes a risk prediction model and an action correction table; The risk prediction model outputs the overpressure risk, overheat risk, overstimulation risk, and reminder load risk within the future control cycle based on the current local application pressure, skin temperature, electrical stimulation circuit status, comfort feedback status, pressure change trend, thermal load change trend, electrical stimulation tolerance change trend, compliance status, and the predicted local application pressure, predicted thermal load, predicted stimulation output level, candidate training reminder timing, and predicted reminder load after the candidate synergistic intervention action. When the predicted value corresponding to any risk exceeds its preset risk threshold, the corresponding action in the candidate collaborative intervention action is corrected to a limiting action, a delayed action, or a replacement action according to the action correction table, so as to obtain the safety collaborative intervention action. Furthermore, the limiting action includes limiting the airbag pressure adjustment amount so that the adjusted local application pressure does not exceed the pressure tolerance boundary, limiting the target heat application temperature or heat application duration so that the predicted heat load does not exceed the heat application tolerance boundary, and limiting the stimulation output level so that the predicted stimulation output level does not exceed the stimulation output safety upper limit. The delayed action includes maintaining the target heat application temperature, heat application duration, or stimulation output level of the previous control cycle and re-evaluating it in the next control cycle. The replacement actions include replacing the pressurization action with the pressure maintenance action, replacing the heating action with the heating stop action, replacing the action of increasing the stimulation output level with the action of decreasing the stimulation output level, and replacing the continuous training reminder with the interval training reminder when the reminder load risk exceeds the preset reminder risk threshold.

[0012] Optionally, S6 includes: At the end of the continuous wearing cycle, the baseline update is calculated based on the posture improvement rate determined by the multimodal state vector, the comfort feedback determined by the comfort feedback state, the number of intervention cancellations determined by the compliance state or the training reminder response data, and the training completion rate. The target posture range, the pressure tolerance boundary, the heat therapy tolerance boundary, and the electrical stimulation tolerance boundary are updated based on the baseline update amount. The updated individualized tolerance boundary is fed back to the constrained reinforcement learning controller and the predicted safety shielding layer, and the updated comfort-safety dual-constraint reward is used for the action generation or policy update of the constrained reinforcement learning controller in the next control cycle. Furthermore, when the comfort feedback indicates discomfort, the number of times the intervention is canceled exceeds a preset cancellation threshold within the continuous wearing period, or the training completion rate does not reach a preset completion rate threshold, the baseline update amount is used to narrow the single adjustment range of the target posture range, reduce the pressure tolerance boundary, reduce the heat compress tolerance boundary, or reduce the upper limit of the stimulation output safety. When the posture improvement rate reaches a preset improvement rate threshold and the training completion rate reaches the preset completion rate threshold, the baseline update amount is used to maintain or gradually tighten the target posture range, and to keep the pressure tolerance boundary, the heat therapy tolerance boundary and the stimulation output safety upper limit from exceeding the preset safety upper limit.

[0013] The beneficial effects of this invention are: 1. By unifying inertial measurement, airbag contact pressure, skin temperature, electrical stimulation circuit, wearing time, and training reminder response data into a multimodal state vector, and extracting the posture change trend and intervention response trend from the temporal feature extraction network, subsequent control actions can be generated based on continuous posture feedback and intervention response, reducing control lag caused by relying solely on a single posture reminder.

[0014] 2. By introducing comfort-safety dual-constraint rewards into the constrained reinforcement learning controller, and by using posture improvement benefits, training completion benefits, local pressure penalties, thermal load penalties, electrical stimulation abnormality penalties, user pause penalties, and wearing interruption penalties together for motion decision-making, airbag pressure, heat application, electrical stimulation, and training reminders can form a coordinated control relationship.

[0015] 3. By predicting the risks of overpressure, overheating and overstimulation through the safety shielding layer, and by limiting, delaying or replacing candidate synergistic intervention actions in combination with individualized tolerance boundaries, the risk of wearing interruption due to local pressure, heat load accumulation or intolerance to electrical stimulation during intervention execution can be reduced. Attached Figure Description

[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a reinforcement learning-based spinal posture control method.

[0017] Figure 2 This is a flowchart of step S5 of the present invention, which predicts the risk of the security shielding layer, determines the threshold, and performs limiting / delaying / replacement correction. Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0019] refer to Figures 1-2 Reinforcement learning-based methods for spinal posture control include: S1. Collect inertial measurement data, airbag contact pressure data, skin temperature data, electrical stimulation circuit data, wearing time data, training reminder response data, and comfort feedback data from multiple torso positions of the wearer, and align them into periodic observation data according to the control cycle. S2. Construct a multimodal state vector based on the periodic observation data and obtain the individualized tolerance boundary corresponding to the wearer. The multimodal state vector includes spinal posture offset, trunk rotation angle, duration of abnormal posture, local contact pressure, heat load, electrical stimulation tolerance, compliance status and comfort feedback status. The individualized tolerance boundary includes target posture range, pressure tolerance boundary, heat compress tolerance boundary and electrical stimulation tolerance boundary. The electrical stimulation tolerance boundary is the upper limit of safe stimulation output. S3. Input the multimodal state vectors of multiple consecutive control cycles into the trained temporal feature extraction network to obtain the posture change trend and intervention response trend; S4. The posture change trend, the intervention response trend, and the individualized tolerance boundary input constraint reinforcement learning controller are used to generate candidate collaborative intervention actions based on comfort-safety dual constraint rewards. The candidate collaborative intervention actions include airbag pressure adjustment amount, target heat application temperature, heat application duration, stimulation output level, and training reminder timing. S5. Input the candidate collaborative intervention action into the prediction safety shielding layer, predict the risks of overpressure, overheating and overstimulation based on the current pressure, temperature, electrical stimulation circuit status, comfort feedback and change trend, and correct the candidate collaborative intervention action based on the individualized tolerance boundary to obtain the safe collaborative intervention action. S6. The safety collaborative intervention action is sent to the airbag pump valve, heating pad, electrical stimulation module and reminder module, and the multimodal state vector, the comfort-safety dual constraint reward and the individualized tolerance boundary are updated according to the data of the next control cycle to form a closed-loop posture correction control.

[0020] In this specific embodiment, S1 includes: The wearable system has an inertial measurement unit fixed in the upper chest, lower chest, and waist segments, and timestamps all sensor data using a unified system clock. The system operates on a fixed control cycle. Trigger an observation alignment and record the start time of each control cycle as . And the sequence number of this control cycle is recorded as The inertial measurement data is acquired at 100Hz, including triaxial angular velocity and triaxial acceleration. A second-order Butterworth low-pass filter is applied in each control cycle, and the cutoff frequency is determined. After suppressing jitter, complementary filtering is used to solve the attitude and output the torso attitude angle sequence. The complementary filtering uses the attitude increment obtained by angular velocity integration as the short-term attitude update and uses the gravity direction calculated by triaxial acceleration to correct the drift of pitch and roll angles. The initial heading angle in the stationary state at the start of wearing is set as zero reference, thereby obtaining the torso attitude angle sequence composed of pitch, roll and heading angles of each torso position in each control cycle. Airbag contact pressure data are acquired at 50 Hz by an array of pressure sensors arranged in each airbag zone. A zero-point calibration is performed before wearing to record the baseline pressure when there is no contact. During wearing, the baseline pressure of all pressure sensors in each airbag zone is arithmetically averaged to obtain the local pressure sequence of the airbag zone tissue. Skin temperature data is collected at 10Hz by a skin-touch digital temperature sensor in each heating element zone and directly outputs the Celsius temperature value to obtain the temperature sequence organized by heating element zones; The electrical stimulation circuit data is acquired at 100Hz by sampling the voltage across the sampling resistor at the output of the stimulation module and the output voltage, combined with the known resistance value of the sampling resistor. The stimulation current is calculated and the loop impedance is calculated from the ratio of the output voltage to the stimulation current. At the same time, the discontinuous state is generated based on the deviation between the effective pulse count and the command pulse count of the stimulation module in this control cycle, thus obtaining the stimulation current, impedance and discontinuous state sequence. Wearing time data is accumulated by the device's running timer and incremented in each control cycle. The current cumulative wearing time is obtained. The training reminder response data is recorded by the reminder module at the reminder trigger time and the user confirmation time, and the response delay is calculated and the training action is recorded. The comfort feedback data is generated by the device according to the single-choice feedback popped up according to the control cycle and uses five levels of discrete values ​​to represent "comfortable, slight discomfort, moderate discomfort, obvious discomfort, and unable to continue", thus jointly mapping it into the compliance observation sequence. To achieve alignment according to the control cycle, the system uniformly resamples inertial measurement data, local pressure, temperature, electrical stimulation circuit, and compliance observations into a format of length [length missing] within each control cycle. The sequence is processed using linear interpolation on equally spaced time grids, and the resampling rule is as follows: ; in Indicates the first Within the first control cycle, the first Aligned observations of each resampled point This represents the original data stream with timestamps or the data stream mapped from the original data stream. Indicates the first The start time of each control cycle Indicates the duration of the control cycle. This indicates the length of the aligned sequence within each control cycle. Indicates the resampling point number and its value range is Indicates the control cycle number; When any data stream does not receive new data or cannot form the above-mentioned length within the current control cycle When a sequence is obtained, the system sets the validity flag corresponding to the data stream to invalid and uses the same type of sequence from the previous valid control period as a placeholder input. The validity flag and the placeholder sequence are then encapsulated together to enter the periodic observation data of the current control period and used in subsequent steps to invalidate or reduce the weight of the modal observation.

[0021] In this specific embodiment, S2 includes: In each control cycle, the controller The system reads periodic observation data and retrieves the wearer's individualized tolerance boundary from non-volatile memory. In this embodiment, the individualized tolerance boundary consists of four types of boundaries and is stored in a structure as follows: ; in The target posture range is represented by upper and lower limits of pitch and roll angles for three segments: upper chest, lower chest, and waist. The average pitch and roll angles for each segment, collected continuously for 30 seconds in a static, upright position during the initial calibration phase, are used as the center and taken as... Forming upper and lower limits, This indicates the pressure tolerance boundary and sets the upper limit of the local bonding pressure at 18 kPa. Indicates the heat therapy tolerance boundary and takes the upper limit of heat load. , The electrical stimulation tolerance boundary is indicated, and the upper limit of the safe stimulation output is set to stimulation output level 6. The controller averages the torso posture angle sequence segment by segment within the current control cycle to obtain the average pitch and roll angles for the "upper thoracic segment, lower thoracic segment, and lumbar segment," and calculates the spinal posture offset based on the target posture range. To characterize the degree to which the posture deviates from the target range, the spinal posture offset is calculated using a piecewise linear penalty for deviations outside the range, accumulated segment by segment. The formula is as follows: ; in Indicates the first The spinal posture offset per control cycle, expressed in degrees. Indicates the use of trunk segments They refer to the upper thoracic segment, lower thoracic segment, and lumbar segment, respectively. Indicate attitude angle type and use Indicates pitch angle and uses Indicates the roll angle. Indicates the first Within the first control cycle, the first Segment corresponding attitude angle type The mean, and These represent the target attitude ranges respectively. The given first The lower and upper limits of this attitude angle type. To take the larger value function; The controller calculates the torso rotation angle based on the difference between the posture angles of adjacent torso positions. The adjacent trunk positions are defined as "upper thoracic segment and lower thoracic segment" and "lower thoracic segment and waist segment", and the trunk rotation angle is represented by the larger of the absolute difference of the mean of the heading angle of the two pairs of adjacent segments, which is used to reflect the degree of axial torsion of the trunk. The controller is based on the spinal posture offset. Is the duration of the zero-update pose anomaly? , among which when When Increase by 1 based on the previous control cycle, and when When Set to 0, the The unit is the number of control cycles and it is used to characterize the duration of continuous anomalies; The controller calculates the local adhesion pressure based on the local pressure sequence mapped by airbag zones obtained from S1. The mean value for each airbag zone within the current control cycle is taken, and the maximum value among all zone means is used as the mean value. And the unit is kPa, and the above With pressure tolerance boundary Direct comparison to support subsequent penalties and safety assessments; The controller determines the heat load based on the temperature sequence mapped by heating element zone obtained in step S1 and the heating element execution record of the previous control cycle. The heating element execution record includes the heating duration of the previous control cycle. Historical target heat therapy temperature Heat load is defined as the product of temperature rise and heating duration. As a baseline for skin temperature, specifically, the average skin temperature of each heating element zone within the current control cycle is taken and compared with... Multiply by the difference The increase in heat load for the current cycle is obtained, and then exponentially accumulated with the heat load of the previous cycle using a fixed forgetting factor of 0.9. This allows the heat load to simultaneously reflect the current skin-contact temperature and the impact of previous heating on heat accumulation; The controller determines the electrical stimulation tolerance based on the stimulation current, impedance, and intermittent state sequence obtained in step S1, as well as the user pause record from the previous control cycle. The range of electrical stimulation tolerance values ​​is as follows: And the initial value is 1. If the number of valid pulses indicating the discontinuous state is lower than the number of command pulses during the current control cycle... Or the impedance mean is not If the user's pause record is set to trigger pause, then... Reduce the threshold by 0.2 from the previous control cycle and cut off the lower limit to 0. If none of the above conditions are met, [the following will occur]. Increase by 0.05 based on the previous control cycle and truncate the upper limit to 1, thereby... It can rapidly decrease with circuit malfunctions and user-initiated pauses, and slowly recover during stable stimulation phases; The controller is based on the cumulative wearing time. Training reminder response delay Cancel the number of interventions and training completion rate Calculate compliance status The range of values ​​for compliance status is as follows: And it is generated using fixed rule scoring: when Record it as 0.25 points, otherwise... Linear scoring, when Record it as 0.25 points, otherwise... Scoring, when Record it as 0.25 points, otherwise... Scoring, when Record it as 0.25 points, otherwise... Scoring is done by adding the scores from the four categories. This allows wearing time, response speed, cancellation behavior, and training completion to be mapped together into continuous quantities that can be used as input for reinforcement learning; The controller determines the comfort feedback state based on the comfort feedback data obtained in step S1. The comfort feedback uses five levels of discrete values ​​and maps them to integers. Furthermore, the terms "comfortable, mild discomfort, moderate discomfort, significant discomfort, and unable to continue" are mapped to... And take the most recent feedback as the control cycle. ; The controller reads the validity flag of the periodic observation data that enters with the S1 placeholder input and sets an invalidation flag and a weighting coefficient for the corresponding state components, wherein an invalidation flag is set for each mode. And when the validity of this modality is marked as invalid, let And set the weighting coefficients of the state components corresponding to this mode to . When the modality validity is marked as valid, let And place This enables the subsequent temporal feature extraction network and constraint reinforcement learning controller to explicitly identify occupied data and suppress its influence; The final controller in the Output multimodal state vector in each control cycle. And fix its constituent components as Simultaneously outputting the individualized tolerance boundary bound to the wearer. This is used for steps S3 to S6.

[0022] In this specific embodiment, S3 includes: The controller in The control cycle will be the most recent consecutive The multimodal state vectors of each control cycle are arranged in chronological order to form an input sequence, which is then fed into a trained and parameter-fixed temporal feature extraction network to obtain the posture change trend and intervention response trend. The input sequence consists of... arrive Constructed by sequential splicing and each Consistent with and containing the output of step S2 , And the invalidation flag for the corresponding mode. With weighting coefficient Before being fed into the network, the controller performs fixed-scale normalization on the continuous components to ensure dimensional consistency and stabilize the network input range. The normalization method is to... and Divide by respectively ,Will Divide by 20 kPa, Divide by ,Will Divide by 50 control cycles, and Keep Original value and Divide by 4 and map to At the same time, invalid flags corresponding to each mode are set. With weighting coefficient The weighting coefficient is retained as an explicit input component and is the corresponding factor when a mode is marked as invalid in step S2. The continuous component of this mode is suppressed on the input side and invalidated by an invalidation flag. The network is informed that this suppression originates from missing placeholders; The temporal feature extraction network consists of a cascaded temporal convolutional unit and a gated recurrent unit, employing a causal structure to meet online inference requirements. The temporal convolutional unit uses two layers of one-dimensional causal convolutions to extract features along the temporal dimension and maintains a stable gradient with residual connections. The first convolutional layer has 64 output channels, a kernel size of 3, a dilation coefficient of 1, and uses ReLU activation. The second convolutional layer also has 64 output channels, a kernel size of 3, a dilation coefficient of 2, and uses ReLU activation. Both convolutional layers use left-side zero padding aligned with the input length to ensure the output remains the same length. The temporal feature sequence is obtained and the short-term dynamics of attitude shift, pressure change, thermal load change and electrical stimulation tolerance change between adjacent control cycles are extracted by the temporal convolution unit in the local temporal neighborhood. The gated recurrent unit uses a single-layer GRU with the hidden state dimension set to 64. Its input is the length of the output of the temporal convolutional unit. The feature sequence is recursively fused with changes in posture offset, pressure, heat load, electrical stimulation tolerance, and compliance status in chronological order, and the hidden state of the last time step is taken as the global temporal representation. The network is configured with two linear output heads following this global temporal representation. The attitude trend head outputs the attitude change trend vector and passes it through... Limit the output to The intervention response trend head outputs the intervention response trend vector and passes it through... Limit the output to The posture change trend vector includes the spinal posture offset. with torso rotation angle The direction and relative magnitude of change are encoded and used to characterize the attitude in recent times. The trend of improvement or deterioration within each control period, the intervention response trend vector includes the trends of pressure change, heat load change, and electrical stimulation tolerance change, and is used to characterize the recent trend. The upward or downward trend of physiological and interactive responses after intervention of each actuator within a control cycle; The temporal feature extraction network is denoted as a function as follows: And its reasoning process is expressed by the following formula: ,in Indicates the first The attitude change trend vector output in each control cycle Indicates the first The intervention response trend vector output in each control cycle. This represents a temporal feature extraction network composed of temporal convolutional units and a GRU. This represents the complete set of weight parameters of the network, which was trained and solidified using recorded wearable history data before deployment. Indicates by arrive The length of the structure is A multimodal state vector sequence Indicates the control cycle number. Indicates the length of the timing input window.

[0023] In this specific embodiment, S4 includes: The controller in Each control cycle receives the attitude change trend vector output from step S3. With intervention response trend vector And read the individualized tolerance boundary output in step S2. and the current multimodal state vector Multimodal state vector from the previous control cycle Based on this, first calculate the comfort-safety dual-constraint reward. And this is used as a unified scalar feedback for decision-making and strategy updates in this cycle, whereby... It is obtained by a linear combination of posture improvement gains, training completion gains, and multiple penalty terms, and is determined by the following formula: ; in Indicates the first Comfort-safety dual-constraint reward for each control cycle. Indicates the weight of posture improvement gains and takes This represents the benefit of posture improvement, expressed as the reduction in spinal posture deviation. The interval was truncated and used to suppress abnormal spikes. Indicates the weight of the training completion result and takes The value is 1 if the training is completed and a training reminder is responded to in the current period, and the training action is completed; otherwise, it is 1. Indicates the local pressure penalty weight and takes This indicates localized pressure penalty and is caused by localized adhesion pressure. Exceeding the pressure tolerance boundary The magnitude of the over-limit is determined by the number of consecutive over-limit control cycles, and it increases monotonically as the over-limit magnitude increases or the duration of the over-limit increases. Indicates the heat load penalty weight and takes Indicates heat load penalty and is caused by heat load Exceeding the limits of heat tolerance The magnitude of the over-limit is determined by the number of consecutive over-limit control cycles, and it increases monotonically as the over-limit magnitude increases or the duration of the over-limit increases. Indicates the weight of the abnormal punishment for electrical stimulation and takes This indicates the level of electrical stimulation as an abnormal punishment, and the stimulation output level actually executed in the previous control cycle. Exceeding the tolerance limit of electrical stimulation The amplitude, along with the abnormal impedance markers and discontinuous state markers of the electrical stimulation circuit in this cycle, are used to determine the value, and a positive value is taken when any abnormality exists. This indicates that the user has paused the penalty weight and taken The value is 1 if the user suspends the penalty and triggers the suspension during the current intervention period, and 0 otherwise. Indicates the weight of the penalty for interruption of wearing and takes This indicates a penalty for interruption of wear, and is determined by a monotonically increasing number of continuous control cycles during wear interruption. In obtaining Subsequently, in this embodiment, the constraint reinforcement learning controller adopts a three-network structure including a policy network, a value network, and a constraint cost network. It is trained before deployment and its parameters are fixed on the device for online inference. The policy network is... The concatenated vector is used as input and a two-layer fully connected structure is adopted. The first hidden layer has 128 neurons and the activation function is ReLU. The second hidden layer has 64 neurons and the activation function is ReLU. The output end adopts five action heads corresponding to the airbag pressure adjustment amount, target heat application temperature, heat application duration, stimulation output level and training reminder timing, respectively. Softmax is used on each action head to obtain the action probability distribution. The output set of the airbag pressure regulation actuator is fixed as follows: This indicates that the pressure increment command issued to the airbag pump valve during this control cycle is fixed at the target heat therapy temperature. It indicates the target temperature setting value for the closed-loop temperature control of the heating element, and the output set of the heating duration actuator is fixed at [value missing]. It indicates the duration of the planned hot compress treatment starting from this cycle, with the stimulation output level and head output set fixed at a certain value. It also indicates the amplitude setting of the electrical stimulation module, and that the output set of the motion head during training reminders is fixed. It also indicates that the reminder will be triggered a corresponding number of seconds later starting from this cycle; The value network and policy network share the same input and use the same two-layer fully connected backbone, and output a single scalar to evaluate the state value to support the estimation of the advantage function during the training phase. The constraint cost network penalizes local pressure in this cycle. Heat load penalty Abnormal punishment for electrical stimulation As input, a two-layer fully connected structure is used. The first hidden layer has 64 neurons with the ReLU activation function, and the second hidden layer has 32 neurons with the ReLU activation function. Output constraint cost at the output end It is also used to characterize the intensity of the accumulation of safety risks caused by the action choices made in this cycle; when Exceeding the preset constraint cost threshold At that time, the controller operates according to a preset probability attenuation coefficient. The risk action probability distribution output by the policy network is reduced and renormalized. The reduced risk actions are limited to positive pressure increase actions in airbag pressure regulation, heating actions in target heat application temperature, extended heat application duration actions in heat application duration, and increased intensity actions in stimulus output intensity. This reduces the probability of the reinforcement learning policy selecting high-load intervention when the constraint cost is too high and outputs candidate collaborative intervention actions. In this embodiment, the controller uses a deterministic selection rule to generate candidate collaborative intervention actions. That is, for each action head, the action value corresponding to the highest probability after attenuation and normalization is selected and combined to form candidate collaborative intervention actions. ,in This indicates the airbag pressure adjustment amount. Indicates the target heat therapy temperature. Indicates the duration of the hot compress. Indicates the stimulation output level. This indicates the timing for training reminders, and the candidate collaborative intervention action is output to the prediction safety shield layer in step S5.

[0024] In this specific embodiment, S5 includes: Predicted security shielding layer in the first Each control cycle receives candidate collaborative intervention actions. And simultaneously read individualized tolerance boundaries and the current local bonding pressure output in steps S2 and S3 Heat load Comfort feedback status Compliance status With intervention response trend vector , among which Extract pressure change trend by fixed index Heat load variation trend Trends in tolerance to electrical stimulation And as a priori for changes in risk; The predicted safety shielding layer calculates the skin-contact temperature scalar by averaging the skin-contact temperature data across heating element zones within the current control cycle and taking the maximum average value across each zone. The circuit impedance scalar is obtained by averaging the electrical stimulation circuit data over the current control cycle. And generate discontinuity markers from the discontinuous state sequence. And when the effective pulse count is lower than the command pulse count season Otherwise ; To enable risk prediction to take into account the consequences of candidate actions, the predicted safety shield layer uses a deterministic forward predictor to generate the predicted local bonding pressure after the candidate actions are performed within a single prediction range. Predicting heat load Predictive stimulus output level and predictive reminder load ,in Generate and take the response coefficient based on the linear response of the airbag pressure regulation. Used to Converted into localized bonding pressure changes and after generation Perform non-negative truncation. The heat load is accumulated exponentially and a fixed forgetting coefficient of 0.9 is used, and the planned heat application duration for this cycle is considered as... and based on skin temperature baseline and target heat application temperature Calculate the heat dose increment for this cycle to obtain , Directly select candidate stimulus output level , Timing of reminders Compliance status jointly determined and when or The system will then treat the load as high and increase the alert level. ; The risk prediction model employs a two-layer perceptron and runs on the device with floating-point parameters. Its input vector is defined as follows: It also outputs the overpressure risk for the next control cycle. Overheating risk Risk of overstimulation With reminders of load risk Its reasoning calculation is ,in This represents the output vector of the risk prediction model, and the range of values ​​for each component is... This represents the element-wise Sigmoid function. This represents the element-wise linear rectified function. This represents the input vector of the risk prediction model. This represents the weight matrix of the first fully connected layer with dimension . , This represents the first-level bias vector with dimension . , This represents the weight matrix of the second fully connected layer with dimension . , This represents the second-layer bias vector with dimension . The above The parameters were fixed after offline training, and the training data was obtained by using "exceeding the tolerance boundary, triggering user pause, or causing loop abnormality or reminder rejection" from real wearing records as risk labels and training with binary cross-entropy loss. Predicting the risk threshold for the security shield layer , ,when or The limiting action is executed and the airbag pressure adjustment is corrected to... The rule was amended to be First restrict to not making Exceeding the pressure tolerance boundary The maximum allowed increment, and when the maximum allowed increment is negative, directly... Set to -1 kPa to enter the depressurization direction and simultaneously replace the candidate pressurization action with the pressure holding action; when or The limiting action is executed and the target heat therapy temperature and duration are corrected. and The revised rule is to apply the rule to a given set of actions. and The internal temperature should be gradually reduced in the order of "first lowering the temperature, then shortening the duration" until... Not exceeding the heat tolerance boundary If lowered to and Still satisfied Then, the delayed action is executed, and the target heat application temperature and duration of the previous control cycle are maintained, and the results are reassessed in the next control cycle. when or or or At that time, a combination of limiting and replacement actions is performed, and the stimulation output level is adjusted to... The rule was amended to be Limit to not exceeding the electrical stimulation tolerance boundary And when discontinuous markings or impedance out-of-bounds errors occur, directly... Set to 0 and replace the action of increasing the stimulation output level with the action of decreasing the stimulation output level; when Execute the replacement action and adjust the timing of the training reminder to The rule was amended to be The interval has been changed to 60 seconds, and interval training reminders have been enabled on the reminder module side with a fixed minimum reminder interval of 300 seconds, thus replacing continuous training reminders with interval training reminders. Predicting the final output of the security shielding layer for coordinated security intervention actions Then send it to step S6 for actual execution.

[0025] In this specific embodiment, S6 includes: The controller in Each control cycle will involve safety coordination intervention actions. The command is sent to the airbag pump valve, heating element, electrical stimulation module, and alert module, forming a closed-loop execution. The airbag pump valve side will... The airbag pressure setpoint is superimposed on the previous control cycle and a micro air pump and exhaust valve are driven by a PID pressure closed loop running at 20Hz to make the local fit pressure track the setpoint. When the pressure sensor reading is below 1kPa for 3 consecutive control cycles, it is determined that the wearing is interrupted and all pressurization, heating and electrical stimulation outputs are stopped. heating element side Write the closed-loop temperature control setting and adjust the PWM duty cycle using skin temperature as feedback in the 10Hz temperature control cycle, while recording the actual heating duration of this cycle. And when The heating output is forcibly shut off at any time. The electrical stimulation module side will Convert to constant current pulse amplitude setting and set a fixed pulse frequency of 50Hz and pulse width. The output is sampled in real time during each control cycle, and the stimulation output level is immediately set to zero when a discontinuity mark or impedance over-limit occurs. The reminder module side is based on Establish a delayed trigger timer and record the reminder time and user confirmation time when it is triggered to form training reminder response data, and fix the minimum interval of continuous reminders to 300s to constrain the reminder load; After the above execution, the controller in the... Each control cycle re-acquires and aligns the cycle observation data according to steps S1 to S3, and updates the multimodal state vector. With trend vector , And calculate the comfort-safety dual-constraint reward according to the unified rules in step S4. For use in action generation and strategy updates in the next control cycle; In this embodiment, the continuous wearing period is defined as a continuous 1800-second time period starting from the start of wearing without any "wearing interruption" event, and the number of control cycles contained in this time period is denoted as . When the continuous wearing cycle ends, the controller calculates the baseline update based on the posture improvement rate, comfort feedback, number of intervention cancellations, and training completion rate during that continuous wearing cycle. And update the individualized tolerance boundary accordingly. The attitude improvement rate is denoted as And from the period before this continuous wearing cycle Mean and posterior spinal posture deviation during the control period The relative decrease rate of the mean spinal posture deviation within the control period was determined and truncated to Comfort feedback is taken from the comfort feedback status within the continuous wearing period. The arithmetic mean is denoted as And the mapping range is The number of times intervention was cancelled is recorded as The training completion rate is determined by the cumulative number of times the user triggers a pause or cancellation intervention within the continuous wearing period, and is recorded as follows: Furthermore, the ratio of "number of reminders to complete training" to "total number of reminders triggered" is used to determine and truncate the threshold. The baseline update amount is calculated using the following formula: ; in Indicates the baseline update amount and its value range is , This indicates a function that truncates and restricts the result within the parentheses to a specified value. Indicates the weight of the attitude improvement rate and takes Represents the training completion rate weights and takes Indicates the comfort feedback weight and takes This indicates that the intervention weight has been cancelled and taken. Indicates the attitude improvement rate. Indicates training completion rate. This represents the average value of the comfort feedback state. Indicates the number of times intervention is cancelled. This indicates a preset threshold for the number of cancellations, set to 4. When the comfort feedback indicates "moderate discomfort or worse," it means that there is discomfort within that continuous wearing period. Control cycle, or ,or and At that time, the controller uses the baseline update amount for safe and conservative updates and performs convergence amplitude narrowing and tolerance boundary reduction, specifically by adjusting the target attitude range. The single adjustment range parameter is fixed as follows: Furthermore, it prohibits further tightening of the target attitude range during the next consecutive wearing cycle, while simultaneously setting the pressure tolerance boundary. Lower And not less than 10 kPa, which will be the boundary of heat tolerance. Lower and not less than To limit electrical stimulation tolerance That is, the upper limit of the safe output of the stimulus is lowered. The gear level must be at least 3 levels; When the attitude improvement rate and and At this time, the controller uses the baseline update amount to progressively tighten the target attitude range while keeping the tolerance boundary from exceeding the preset safety limit. Specifically, this involves adjusting the target attitude range... The upper and lower limits of each pitch and roll angle are tightened towards their calibration center. And ensure that the half-width of each interval at each angle is not less than and the pressure tolerance boundary Heat therapy tolerance boundary Boundary with electrical stimulation tolerance Keep it unchanged and force it not to exceed the preset safety limit. With 6 gears; After the update is complete, the controller will update the individualized tolerance boundary. Write back to non-volatile memory and synchronously feed back to the constraint reinforcement learning controller and predictive safety shield for action generation and risk correction invocation in the next control cycle.

[0026] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

[0027] This invention achieves the corresponding technical effect by continuously transforming between multimodal state vectors, posture change trends, intervention response trends, candidate collaborative intervention actions, and safety collaborative intervention actions, thus enabling real-time posture feedback, comfort feedback, and actuator control to form a closed loop. This addresses the problem of lacking multi-actuator collaborative feedback control in posture correction products.

[0028] This invention combines comfort-safety dual-constraint rewards, a predicted safety shielding layer, and individualized tolerance boundary updates, enabling the actions output by the constraint reinforcement learning controller to respond to posture improvement needs while also being constrained by pressure, temperature, electrical stimulation tolerance, and compliance states, thus making it more suitable for long-term posture correction scenarios.

Claims

1. A spinal posture correction control method based on reinforcement learning, characterized in that, include: S1. Collect inertial measurement data, airbag contact pressure data, skin temperature data, electrical stimulation circuit data, wearing time data, training reminder response data, and comfort feedback data from multiple torso positions of the wearer, and align them into periodic observation data according to the control cycle. S2. Construct a multimodal state vector based on periodic observation data and obtain the individualized tolerance boundary corresponding to the wearer. The multimodal state vector includes spinal posture offset, trunk rotation angle, duration of abnormal posture, local contact pressure, heat load, electrical stimulation tolerance, compliance status and comfort feedback status. The individualized tolerance boundary includes target posture range, pressure tolerance boundary, heat compress tolerance boundary and electrical stimulation tolerance boundary. The electrical stimulation tolerance boundary is the safe upper limit of stimulation output. S3. Input the multimodal state vectors of multiple consecutive control cycles into the trained temporal feature extraction network to obtain the posture change trend and intervention response trend; S4. The posture change trend, intervention response trend and individualized tolerance boundary input constraint reinforcement learning controller are used to generate candidate collaborative intervention actions based on comfort-safety dual constraint rewards. The candidate collaborative intervention actions include airbag pressure adjustment amount, target heat application temperature, heat application duration, stimulation output level and training reminder timing. S5. Input the candidate synergistic intervention action into the prediction safety shield layer, predict the risks of overpressure, overheating and overstimulation based on the current pressure, temperature, electrical stimulation circuit status, comfort feedback and changing trends, and modify the candidate synergistic intervention action based on the individualized tolerance boundary to obtain the safe synergistic intervention action. S6. Send the safety collaborative intervention action to the airbag pump valve, heating pad, electrical stimulation module and reminder module, and update the multimodal state vector, comfort-safety dual-constraint reward and individualized tolerance boundary according to the data of the next control cycle to form a closed-loop posture correction control.

2. The spinal posture control method based on reinforcement learning according to claim 1, characterized in that, S1 includes: The triaxial angular velocity and triaxial acceleration in the inertial measurement data are converted into a trunk attitude angle sequence; the airbag contact pressure data are mapped into a local pressure sequence according to airbag partitions; the skin temperature data are mapped into a temperature sequence according to heating pad partitions; the electrical stimulation circuit data are mapped into a stimulation current, impedance, and intermittent state sequence; the wearing duration data, the training reminder response data, and the comfort feedback data are mapped into a compliance observation sequence; when any sequence is missing in the current control cycle, the validity mark of the sequence is set to invalid and the same type of sequence from the previous valid control cycle is used as a placeholder input, and the validity mark is entered into the cycle observation data along with the placeholder input.

3. The spinal posture correction control method based on reinforcement learning according to claim 2, characterized in that, S2 includes: The spinal posture offset is determined based on the deviation of the trunk posture angle sequence from the target posture range; the trunk rotation angle is determined based on the difference between adjacent trunk posture angles. The duration of the posture abnormality is determined based on the number of control cycles in which the spinal posture offset is continuously outside the target posture range; The local pressure is determined based on the local pressure sequence; the heat load is determined based on the temperature sequence, the heating duration determined by the heating pad execution record of the previous control cycle, and the historical target heat application temperature; the electrical stimulation tolerance is determined based on the stimulation current, impedance, intermittent state, and user pause records determined by the user interaction records of the previous control cycle; the compliance status is determined based on the cumulative wearing time, training reminder response delay, number of intervention cancellations, and training completion rate; and the comfort feedback status is determined based on the comfort feedback data. The corresponding state component is set with an invalidation flag or a weighting factor based on the validity flag.

4. The spinal posture control method based on reinforcement learning according to claim 1, characterized in that, S3 includes: inputting the multimodal state vectors of the multiple consecutive control cycles into the temporal feature extraction network containing a temporal convolutional unit and a gated recurrent unit; extracting the posture shift change, pressure change, thermal load change, and electrical stimulation tolerance change between adjacent control cycles by the temporal convolutional unit; and performing temporal fusion of the posture shift change, pressure change, thermal load change, electrical stimulation tolerance change, and compliance state change by the gated recurrent unit to output the posture change trend and the intervention response trend, wherein the intervention response trend includes the pressure change trend, thermal load change trend, and electrical stimulation tolerance change trend.

5. The spinal posture control method based on reinforcement learning according to claim 1, characterized in that, In step S4, the comfort-safety dual-constraint reward includes posture improvement benefit, training completion benefit, local pressure penalty, thermal load penalty, electrical stimulation abnormality penalty, user pause penalty, and wearing interruption penalty; the posture improvement benefit is determined based on the reduction in spinal posture deviation relative to the previous control cycle; the training completion benefit is determined based on records of training reminders being responded to and training actions being completed; the local pressure penalty is determined based on the magnitude and duration of the local contact pressure exceeding the pressure tolerance boundary; the thermal load penalty is determined based on the magnitude and duration of the thermal load exceeding the heat application tolerance boundary; the electrical stimulation abnormality penalty is determined based on the magnitude of the stimulation output level exceeding the stimulation output safety upper limit, impedance abnormality markers, and intermittent status; the user pause penalty is determined based on records of user-triggered pauses during intervention execution; and the wearing interruption penalty is determined based on the duration of wearing interruptions.

6. The spinal posture control method based on reinforcement learning according to claim 5, characterized in that, In step S4, the constraint reinforcement learning controller includes a policy network, a value network, and a constraint cost network. The policy network outputs the probability distribution of the posture change trend, the intervention response trend, the individualized tolerance boundary, and the comfort-safety dual constraint reward output action. The constraint cost network generates a constraint cost value based on the local pressure penalty, the thermal load penalty, and the electrical stimulation abnormality penalty. When the constraint cost value exceeds a preset constraint cost threshold, the selection probability of the corresponding pressurization, heating, prolonged heat application, or increased stimulation output level action in the action probability distribution is reduced according to a preset probability decay coefficient, and the candidate collaborative intervention action is output.

7. The spinal posture control method based on reinforcement learning according to claim 4, characterized in that, In step S5, the predicted safety shielding layer includes a risk prediction model and an action correction table. The risk prediction model outputs the overpressure risk, overheat risk, overstimulation risk, and reminder load risk within the future control cycle based on the current local contact pressure, skin temperature, electrical stimulation circuit status, comfort feedback status, pressure change trend, heat load change trend, electrical stimulation tolerance change trend, compliance status, and the predicted local contact pressure, predicted heat load, predicted stimulation output level, candidate training reminder timing, and predicted reminder load after the candidate synergistic intervention action is performed. When the predicted value corresponding to any risk exceeds its preset risk threshold, the corresponding action in the candidate synergistic intervention action is corrected to a limited action, a delayed action, or a replacement action according to the action correction table to obtain the safe synergistic intervention action.

8. The spinal posture control method based on reinforcement learning according to claim 7, characterized in that, In step S5, the limiting action includes limiting the airbag pressure adjustment amount so that the adjusted local application pressure does not exceed the pressure tolerance boundary, limiting the target heat application temperature or heat application duration so that the predicted heat load does not exceed the heat application tolerance boundary, and limiting the stimulation output level so that the predicted stimulation output level does not exceed the stimulation output safety upper limit; the delay action includes maintaining the target heat application temperature, heat application duration, or stimulation output level of the previous control cycle and re-evaluating in the next control cycle; the replacement action includes replacing the pressurization action with the pressure maintenance action, replacing the heating action with the heating stop action, replacing the stimulation output level increase action with the stimulation output level decrease action, and replacing the continuous training reminder with the interval training reminder when the reminder load risk exceeds the preset reminder risk threshold.

9. The spinal posture control method based on reinforcement learning according to claim 8, characterized in that, In step S6, at the end of the continuous wearing cycle, the baseline update amount is calculated based on the posture improvement rate determined by the multimodal state vector, the comfort feedback determined by the comfort feedback state, the number of intervention cancellations determined by the compliance state or the training reminder response data, and the training completion rate. The target posture range, the pressure tolerance boundary, the heat therapy tolerance boundary, and the electrical stimulation tolerance boundary are updated based on the baseline update amount. The updated individualized tolerance boundary is fed back to the constrained reinforcement learning controller and the predicted safety shielding layer, and the updated comfort-safety dual-constraint reward is used for the action generation or policy update of the constrained reinforcement learning controller in the next control cycle.

10. The spinal posture control method based on reinforcement learning according to claim 9, characterized in that, In step S6, when the comfort feedback indicates discomfort, the number of times the intervention is canceled exceeds a preset cancellation threshold within the continuous wearing period, or the training completion rate does not reach a preset completion rate threshold, the baseline update amount is used to narrow the single adjustment range of the target posture range, reduce the pressure tolerance boundary, reduce the heat compress tolerance boundary, or reduce the upper limit of the stimulation output safety. When the posture improvement rate reaches a preset improvement rate threshold and the training completion rate reaches the preset completion rate threshold, the baseline update amount is used to maintain or gradually tighten the target posture range, and to keep the pressure tolerance boundary, the heat therapy tolerance boundary and the stimulation output safety upper limit from exceeding the preset safety upper limit.