Hypertension personalized intervention system combining diet and exercise analysis

By designing a personalized hypertension intervention system combining diet and exercise analysis, using multi-source data and deep learning technology, the problem of lack of personalization and real-time in the existing system is solved, and efficient hypertension management and control is achieved.

CN120164571APending Publication Date: 2025-06-17安徽省宿州市立医院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510233442.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing hypertension management system lacks personalization and real-timeness, and cannot dynamically adjust diet and exercise suggestions based on the user's actual living conditions, resulting in unsatisfactory treatment results.

Method used

Design a personalized intervention system for hypertension combining diet and exercise analysis, and generate dynamically adjusted intervention solutions through multi-source data collection, data processing and analysis, personalized intervention generation and real-time feedback optimization, and use advanced technologies such as deep learning to generate dynamically adjusted intervention solutions.

Benefits of technology

The system can fully integrate diet, exercise and physiological data, generate and adjust personalized intervention plans, significantly improve the effect and accuracy of hypertension control, and improve the overall effect of patient compliance and health management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164571A_ABST
    Figure CN120164571A_ABST
Patent Text Reader

Abstract

The invention discloses a hypertension personalized intervention system combining diet and exercise analysis, and relates to the technical field of health management. According to the system, diet, exercise and physiological index data of a user are acquired through a multi-source data acquisition module; the data processing and analyzing module performs preprocessing and feature extraction on the data; and the personalized intervention generation module dynamically generates personalized diet adjustment suggestions and exercise plans according to the blood pressure target of the user and feedback. The multi-terminal synchronization module pushes an intervention scheme to a mobile phone and a smart watch of a user in real time, supports voice, image-text and push reminding, and improves user compliance. By deeply fusing diet and exercise data and combining a dynamic adjustment mechanism, the problems of data splitting, scheme staticizing and low user compliance in an existing hypertension intervention system are solved, and scientificity and effectiveness of hypertension management are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of health management, and particularly to a personalized hypertension intervention system combining diet and exercise analysis. Background Art

[0002] With the change of global lifestyle, especially the unhealthy eating habits and lack of sufficient exercise, hypertension has become a major public health problem worldwide. Hypertension is not only the main risk factor for various cardiovascular diseases such as heart disease and stroke, but also the main cause of kidney damage and retinopathy. Therefore, the management and control of hypertension have become an urgent problem to be solved in the global medical system.

[0003] Currently, the traditional treatment methods for hypertension mainly rely on drug treatment and lifestyle intervention. Lifestyle intervention includes diet adjustment and exercise intervention. However, the existing intervention methods lack personalized regulation for individuals and are difficult to dynamically adjust according to the real-time health status of patients, resulting in unsatisfactory treatment effects, poor patient compliance, and inaccurate personalized intervention during the treatment process.

[0004] With the rapid development of Internet of Things, artificial intelligence and big data technologies, more and more intelligent devices and applications can monitor the health status of users in real time, such as smart bracelets, smart weighing scales, smart sphygmomanometers, etc. These devices can provide health data in multiple aspects including diet, exercise, blood pressure, etc. However, how to combine these multi-source data and provide a personalized and dynamically adjusted hypertension intervention plan through scientific analysis and optimization methods is still a technical problem to be solved urgently.

[0005] Existing hypertension intervention systems mostly rely on rules set manually or simple algorithms, such as providing diet and exercise suggestions based on the basic health information of users. However, these systems cannot be dynamically adjusted according to the actual living conditions of users and lack personalization and real-time nature. Therefore, it is particularly important to design a system that can obtain and analyze users' diet, exercise and physiological data in real time and generate a personalized intervention plan through artificial intelligence algorithms.

[0006] Some current hypertension management systems have tried to use machine learning algorithms for data analysis, but most of these methods have not fully utilized the synergistic effect of multi-source data. In particular, existing systems often only analyze a single type of health data (such as blood pressure data or exercise data), and fail to effectively integrate the interaction of multi-faceted data such as diet, exercise and physiological indicators, and cannot provide a comprehensive and personalized health management plan.

[0007] Therefore, there is an urgent need for a new type of personalized hypertension intervention system that can acquire multi-source health data, perform intelligent analysis through advanced technologies such as deep learning, and dynamically generate and adjust diet and exercise plans based on the user's health status and feedback, thereby improving the intervention effect, helping users better manage and control hypertension, and reducing the risk of complications.

[0008] The present invention aims to solve the above problems and proposes a personalized hypertension intervention system that combines diet and exercise analysis. Through multi-source data collection, data processing and analysis, personalized intervention generation, and real-time feedback optimization, this system can provide users with a dynamically adjusted hypertension intervention plan, thereby effectively helping users control blood pressure and improving the effect of health management. Summary of the Invention

[0009] The object of the present invention is to provide a personalized hypertension intervention system that combines diet and exercise analysis. Through real-time collection and comprehensive analysis of multi-source data, and combined with artificial intelligence algorithms to generate dynamic intervention plans, it overcomes the deficiencies of existing hypertension management systems in personalized intervention, real-time adjustment, and multi-dimensional data fusion. Existing systems often cannot be adjusted in real time according to the actual health status of users and lack precise personalized intervention plans. Through the present invention, it is possible to comprehensively integrate diet, exercise, and physiological data, and generate and adjust diet and exercise plans. This system can not only be optimized in real time according to the user's feedback, but also dynamically adjust the intervention content according to the user's health goals, thereby significantly improving the effect and accuracy of hypertension control. It is particularly suitable for scenarios such as long-term health management, personalized disease intervention, and intelligent health monitoring, ensuring that users obtain more accurate personalized treatment plans.

[0010] To achieve the above object, the present invention provides a personalized hypertension intervention system that combines diet and exercise analysis, including the following main modules: a multi-source data collection module, a data processing and analysis module, a personalized intervention generation module, and a multi-terminal synchronization module.

[0011] Multi-source data collection module: used to obtain the user's diet, exercise, and physiological index data in real time. The collection methods include manual upload by the user and collection and upload through smart devices. Smart devices include wearable devices, blood pressure monitors, and weighing scales, ensuring that the data comprehensively and accurately reflects the user's health status.

[0012] Data processing and analysis module: performs multi-modal analysis on multi-source data, including data preprocessing and feature extraction.

[0013] Personalized Intervention Generation Module: Based on the data processed by the Data Processing and Analysis Module, the user's health goals, and feedback, the system generates personalized diet and exercise intervention plans through a pre-trained reinforcement learning model. After the user executes the plan, the system collects feedback data (including diet data, exercise data, physiological indicators, and user ratings), and optimizes the reinforcement learning model based on the feedback to dynamically adjust subsequent intervention plans.

[0014] Multi-terminal Synchronization Module: Real-time synchronization among multiple devices, including smartphones and smartwatches, to ensure that users can obtain the latest health data and intervention suggestions anytime and anywhere.

[0015] Furthermore, the diet data includes calorie, sodium, and potassium intake, the exercise data includes the type and duration of exercise, and the physiological indicator data includes blood pressure value and weight.

[0016] Furthermore, the data collected by the Multi-source Data Acquisition Module includes aspects of diet, exercise, and physiological indicators. Diet data is manually input by the user. The methods for collecting exercise data include manual input by the user and using smart bracelets or watches. The methods for collecting physiological indicator data include manual input by the user and using sphygmomanometers and weighing scales.

[0017] Furthermore, the Data Processing and Analysis Module performs missing value filling, outlier removal, and normalization on the collected data, and finally extracts data features for generating intervention plans.

[0018] Furthermore, the Personalized Intervention Generation Module uses the reinforcement learning method to dynamically generate personalized diet and exercise intervention plans according to the user's current state and feedback. When the system detects that the user's blood pressure remains high, it recommends that the user increase the frequency of aerobic exercise or adjust sodium intake.

[0019] Furthermore, the reinforcement learning algorithm involves state space, action space, reward function, policy, and value function.

[0020] Furthermore, the state space can include the user's current exercise data, diet data, and physiological indicators, expressed as follows:

[0021] S = {s1, s2, …, s d’}

[0022] where s d’ represents the d'-th state, and each state contains the user's health data and feedback information.

[0023] Furthermore, the action space includes recommended diet and exercise adjustment, expressed as follows:

[0024]

[0025] Among them, represents the th action.

[0026] Furthermore, the reward function defines the goodness or badness of taking a certain action in a specific state, which is expressed as follows:

[0027]

[0028] Among them, represents the state at time step ; represents the action taken at time step ; represents the reward obtained after executing a certain action at time step ; BP real is the actual blood pressure value of the user after executing the action , which is related to the state ; BP target is the target blood pressure value of the user, is the execution situation of the user for the action, and the value range is [0,1]; α and β are weight coefficients used to balance the contributions of prediction accuracy and user compliance to the reward;

[0029] Furthermore, the update method of the weights of α and β is as follows:

[0030]

[0031] Among them, η is the learning rate, and the value is 0.01 to 0.1; is the partial derivative of the reward function with respect to α, is the partial derivative of the reward function with respect to β.

[0032] The policy is a mapping from state to action, which is expressed as follows:

[0033]

[0034] Among them, represents the probability of selecting the action in the state .

[0035] Furthermore, the value function is used to evaluate the goodness or badness of a certain state or state-action pair. The value function in the state s is expressed as:

[0036]

[0037] Among them, γ is the discount factor (between 0 and 1), which represents the attenuation of future rewards; is from the state Take action The reward obtained after represents the expected value for all possible future state and action sequences.

[0038] Furthermore, use Q-learning to update the policy and value function, and the update formula is:

[0039]

[0040] where μ is the learning rate and γ is the discount factor, is the maximum Q value in state

[0041] Furthermore, the multi-terminal synchronization module supports data synchronization between multiple devices, and the multiple devices include smartphones, tablets, and smartwatches.

[0042] Compared with the prior art, the advantages of the present invention are as follows:

[0043] (1) Precise personalized intervention plan: By collecting the user's diet, exercise, and physiological data in real time and combining advanced algorithms for precise analysis, the present invention can dynamically generate and adjust the diet and exercise plan according to the user's health status and feedback. Compared with traditional systems, the present invention can not only provide personalized treatment plans but also adjust according to real-time health changes to ensure more effective blood pressure control and optimized health status.

[0044] (2) Intelligent optimization and adaptive adjustment: The present invention adopts a reinforcement learning algorithm to automatically optimize the model based on user feedback and continuously adjust the content and intensity of the intervention plan. Different from traditional static intervention methods, the present invention can intelligently adapt to and adjust the intervention strategy according to the user's blood pressure changes, exercise performance, and health feedback, greatly improving the accuracy and long-term effectiveness of the intervention, and enhancing the patient's compliance and the overall effect of health management.

[0045] (3) Multi-terminal real-time synchronization: Through the multi-terminal synchronization module, it is ensured that the user's health data and personalized intervention plan can be synchronized in real time between multiple devices such as smartphones, tablets, and smartwatches, and a seamless user experience is provided. Users can view and adjust health data at any time, and can effectively manage their health status whether at home, at work, or during exercise, meeting the personalized, convenient, and cross-device health management needs. Brief Description of the Drawings

[0046] ​To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0047] Figure 1 It is the system module diagram of the present invention;

[0048] Figure 2 It is the structure diagram of the multi-source data acquisition module of the present invention;

[0049] Figure 3 It is the structure diagram of the data processing and analysis module of the present invention. Specific embodiments

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0051] Embodiment 1: Please refer to Figure 1 As shown, the hypertension personalized intervention system combining diet and exercise analysis in this embodiment includes a multi-source data acquisition module, a data processing and analysis module, a personalized intervention generation module, a user feedback and optimization module, and a multi-terminal synchronization module.

[0052] Multi-source data acquisition module: Refer to Figure 2 As shown, the multi-source data acquisition module obtains the diet, exercise, and physiological index data of the user in real time. The acquisition methods include manual upload by the user and acquisition and upload through smart devices. The smart devices include wearable devices, sphygmomanometers, and weighing scales to ensure that the data comprehensively and accurately reflects the health status of the user.

[0053] The diet data (calories, sodium, and potassium intake in this embodiment) is collected by the user manually inputting. The user manually inputs the daily diet situation through the system interface, including meal times, food types, food portions, and nutritional components. The manual input method depends on the user's active record to ensure the detail and accuracy of the diet data. In addition, to improve the convenience of input, the system uses technologies such as voice recognition to assist the user in quickly recording food information.

[0054] The collection methods of the motion data (the motion type and duration in this embodiment) include manual input by the user and automatic monitoring through a smart bracelet or smart watch. The user can manually input information such as the daily motion type and duration into the system to facilitate recording of daily exercise. The smart bracelet or watch monitors the motion data in real time through built-in sensors and automatically records physiological parameters such as steps, heart rate, exercise intensity, and calories burned, helping the system to understand the user's exercise situation in real time. By connecting to the cloud platform, the system can synchronously obtain this data.

[0055] The collection methods of the physiological index data (blood pressure value and weight in this embodiment) also include the combination of manual input by the user and automated devices. The user records physiological data such as daily blood pressure and weight through manual input, which is particularly applicable to indicators that are easy to track in daily health management. To ensure the accuracy of the data, the system also supports automatic measurement and uploading of data through devices such as smart blood pressure monitors and weighing scales. The smart blood pressure monitor can monitor blood pressure changes in real time, automatically record and upload them to the system, reducing the possibility of human error. The smart weighing scale helps the system analyze the user's health status and generate corresponding adjustment suggestions by measuring data such as weight and body fat percentage.

[0056] Data processing and analysis module: Refer to Figure 3 As shown, the data processing and analysis module uses machine learning algorithms to preprocess and extract features from the collected data; the data preprocessing uses missing value filling, outlier removal, and normalization operations.

[0057] In the missing value filling step, this embodiment uses the mean filling method to fill the missing values in the user dataset {x1, x2, …, x N}:

[0058]

[0059] where n is the number of valid samples, x i is the valid data point in the dataset, and the missing value is filled with .

[0060] Outlier detection uses the first Z-score algorithm:

[0061]

[0062] where x i is the i-th data point, μ is the mean of the data, and σ is the standard deviation of the data; if |z i | > 3, then x i is considered an outlier and is deleted.

[0063] Data standardization uses the second Z-score algorithm:

[0064]

[0065] where x' i is the standardized value, μ is the mean of the data, and σ is the standard deviation of the data.

[0066] After data preprocessing, different feature extractions are adopted for different data types. In this embodiment, for diet data, the daily total calorie intake is calculated, and the formula is as follows:

[0067]

[0068] where Claories ck is the calorie of the kth food on the cth day, K is the number of foods on that day, and the same applies to sodium intake, potassium intake, etc. Take the sodium or potassium content ingested per meal and sum to obtain the total daily intake.

[0069] In this embodiment, for exercise data, the METs (metabolic equivalent) method is used to convert the amount of exercise into calories burned, and the formula is as follows:

[0070] Calories Burned = MET · Body Weight(kg) · Duration(hours)

[0071] where MET is the metabolic equivalent of the exercise, determined according to different exercise types and intensities (for example, the MET value of jogging is 7.0); Body Weight(kg) is the user's weight, and Duration(hours) is the duration of the exercise, in hours. 1 MET represents the energy consumed by the body at rest, that is, the calories consumed per kilogram of body weight per hour.

[0072] The blood pressure data in this embodiment includes systolic blood pressure and diastolic blood pressure, which are expressed as follows:

[0073] BP c = [SBP c , DBP c

[0074] where SBP c is the systolic blood pressure on the cth day, and DBP c is the diastolic blood pressure on the cth day.

[0075] ​Personalized Intervention Generation Module: Based on the data processed by the Data Processing and Analysis Module, the user's health goals, and feedback, the system generates personalized diet and exercise intervention plans through a pre-trained reinforcement learning model. After the user executes the plan, the system collects feedback data (including diet data, exercise data, physiological indicators, and user ratings), and optimizes the reinforcement learning model based on the feedback to dynamically adjust subsequent intervention plans.

[0076] When the system detects that the user's blood pressure remains high (systolic blood pressure > 130 mmHg or diastolic blood pressure > 80 mmHg), in terms of exercise, it is recommended that the user increase the frequency and intensity of aerobic exercise, including brisk walking, running, swimming, or cycling; the exercise intensity is maintained at 50%-65% of the maximum heart rate.

[0077] When the blood pressure is high, reduce the sodium intake in the diet, control the daily sodium intake below 1500 mg, reduce the intake of high-sodium foods, including highly salted processed foods, canned foods, and snacks, and at the same time increase the consumption of low-sodium foods such as fresh vegetables, fruits, and whole grains; potassium helps lower blood pressure, and foods rich in potassium can also be increased, including bananas, spinach, and potatoes.

[0078] After each intervention, the system adjusts the plan according to the user's feedback and health data (including blood pressure and weight). When the user feedback indicates that the exercise intensity is too high, the system reduces the exercise intensity according to the user's feedback and recommends appropriate low-intensity exercises.

[0079] The reinforcement learning algorithm involves a state space, an action space, a reward function, a policy, and a value function; the state space defines all possible states of the environment; the state space can include the user's current blood pressure value, exercise data, diet data, physiological indicators, and user ratings, expressed as follows:

[0080] S = {s1, s2, …, s d’}

[0081] where s d’ represents the d'-th state, and each state contains the user's health data and feedback information. s includes systolic blood pressure, diastolic blood pressure, exercise calories, sodium intake, potassium intake, diet calories, weight, execution compliance rate, and user ratings.

[0082] The action space refers to all possible actions that can be taken in a certain state. In this implementation, the action space includes recommending diet or adjusting exercise, expressed as follows:

[0083]

[0084] where represents the -th action.

[0085] The reward function defines the goodness of taking a certain action in a specific state as follows:

[0086]

[0087] where, represents the state at time step ; represents the action taken at time step ; represents the reward obtained after executing a certain action at time step ; BP real is the actual blood pressure value of the user after executing the action , related to the state ; BP target is the target blood pressure value of the user, is the execution situation of the action by the user, and the value range is [0, 1]; α and β are weight coefficients used to balance the contributions of prediction accuracy and user compliance to the reward;

[0088] The update method of the weights α and β is as follows:

[0089]

[0090] where, η is the learning rate, and the value is 0.01 to 0.1; is the partial derivative of the reward function with respect to α, is the partial derivative of the reward function with respect to β.

[0091] The policy is a mapping from state to action, that is, in each state, which action should be selected, expressed as follows:

[0092]

[0093] where, represents the probability of selecting action in state .

[0094] The value function is used to evaluate the goodness of a certain state or state-action pair. The value function in state s is expressed as:

[0095]

[0096] where, γ is the discount factor (between 0 and 1), representing the attenuation of future rewards; is the reward obtained after taking action from state , represents the expected value of all possible future state and action sequences.

[0097] The training process of reinforcement learning is as follows:

[0098] (1) Initialize the state space S, action space A, reward function R, initialize the value function V and policy π.

[0099] (2) At each time step Obtain the current state Based on the current policy π, select an action Execute the action After that, and the reward Then update the policy according to the new state and reward.

[0100] (3) Use Q-learning to update the policy and value function. The update formula is:

[0101]

[0102] where μ is the learning rate, γ is the discount factor, is the maximum Q-value in state

[0103] (4) Repeat the above process until the policy converges, that is, it no longer changes its policy or value function. After the training is completed, the reinforcement learning model will select appropriate actions according to the current value function when the state changes.

[0104] Multi-terminal synchronization module: Real-time synchronization among multiple devices, including smartphones, tablets and smart watches, to ensure that users can obtain the latest health data and intervention suggestions anytime and anywhere.

[0105] The above formulas are all in dimensionless form and only use numerical values for calculation. These formulas are obtained based on a large amount of data and through software simulation, aiming to be as close to the actual situation as possible. The preset parameters in the formulas can be adjusted by those skilled in the art according to specific requirements.

[0106] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0107] ​The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation manners. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A personalized hypertension intervention system combining diet and exercise analysis, characterized in that: include: Multi-source data acquisition module: used to obtain the user's diet, exercise and physiological index data in real time. The acquisition methods include manual upload by the user and collection and uploading through smart devices; the smart devices include wearable devices, blood pressure monitors and weight scales; Data processing and analysis module: multi-modal analysis of multi-source data, including data preprocessing and feature extraction; Personalized intervention generation module: Based on the data processed by the data processing and analysis module, the user's health goals and feedback, the system generates a personalized diet and exercise intervention plan through a pre-trained reinforcement learning model; after the user executes the plan, the system collects feedback data, optimizes the reinforcement learning model based on the feedback, and dynamically adjusts the subsequent intervention plan; the feedback data includes diet data, exercise data, physiological indicators and user scores; Multi-terminal synchronization module: Real-time synchronization between multiple devices, including smartphones, tablets and smart watches, ensures that users can obtain the latest health data and intervention recommendations in real time.

2. The system according to claim 1, characterized in that The dietary data includes calorie, sodium and potassium intake, the exercise data includes the type and duration of exercise, and the physiological index data includes blood pressure value and weight.

3. The system according to claim 1, characterized in that The dietary data is manually input by the user, the exercise data is collected by manually inputting by the user and using a smart bracelet or watch; the physiological index data is collected by manually inputting by the user and using a blood pressure monitor or a weight scale.

4. The system according to claim 1, characterized in that The data processing and analysis module fills missing values, removes outliers, and normalizes the collected data, and finally extracts data features for intervention plan generation.

5. The system according to claim 1, characterized in that The personalized intervention generation module adopts a reinforcement learning method to dynamically generate a personalized diet and exercise intervention plan based on the user's current status and feedback; when the system detects that the user's blood pressure continues to be high, it recommends that the user increase the frequency of aerobic exercise or adjust the sodium intake.

6. The system according to claim 5, characterized in that The reinforcement learning algorithm involves state space, action space, reward function, strategy and value function; the action taken when the user state changes is determined according to the strategy and value function.

7. The system according to claim 6, characterized in that The state space includes the user's current exercise data, diet data, physiological indicators, and user scores, which are expressed as follows: S={s1,s2,…,s d’ } Among them, s d’ represents the d'th state, each state contains the user's health data and feedback information; the action space includes recommended diet and adjusted exercise, which is expressed as follows: in, Representative action.

8. The system according to claim 6, characterized in that The reward function defines how good it is to take an action in a particular state and is expressed as follows: R(s t ,a t )=-α|BP real (s t ,a t )-BP target |+β·(1-|E exec (a t )|) Among them, s t represents the state at time step t; a t represents the action taken at time step t; R(s t ,a t ) represents the reward obtained after performing an action at time step t; BP real For the user performing action a t The actual blood pressure value after the state s t Related, BP target is the user's target blood pressure value, E exec (a t ) is the user's execution of the action, and its value range is [0,1]; α and β are weight coefficients, which are used to balance the contribution of prediction accuracy and user compliance to the reward; The strategy is a mapping from states to actions, expressed as follows: π(a t |s t )=P(a t |s t ) Among them, P(a t |s t ) means in state s t Next select action a t probability.

9. The system according to claim 6, characterized in that The value function is used to evaluate the quality of a state or state-action pair; the value function in state s is expressed as: Among them, γ is the discount factor, which indicates the decay of future rewards and takes values ​​between 0 and 1; From the state Take Action After the reward, Represents the expected value of all possible states and action sequences in the future; Q-learning is used to update the strategy and value function, and the update formula is: Where μ is the learning rate, γ is the discount factor, is the state t+1 The maximum Q value.

10. The system according to claim 1, characterized in that The multi-terminal synchronization module supports data synchronization between multiple devices, including but not limited to smartphones, tablet computers and smart watches.