Health management method and device based on artificial intelligence, equipment and storage medium
By acquiring users' multimodal states and psychological profiles, and using reinforcement learning algorithms to generate personalized incentive actions, the problem of lacking personalized communication and incentive mechanisms in existing health management engines is solved, achieving more natural and intelligent human-computer interaction and improving the effectiveness of users' health management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-14
- Publication Date
- 2026-05-01
AI Technical Summary
Existing health management intelligent engines lack personalized communication and incentive mechanisms at the human-computer interaction level, resulting in low user engagement in the long term and affecting the effectiveness of health management.
By acquiring users' multimodal integrated state vectors, multidimensional psychological profiles, and contextual state vectors, reinforcement learning algorithms are used to match personalized incentive actions, generate target incentive actions, and proactively motivate users to engage in healthy behaviors.
It enables more natural and intelligent human-computer interaction, increases user participation in health behaviors, promotes the formation of good health habits, and significantly improves the long-term effects of health management.
Smart Images

Figure CN121964173A_ABST
Abstract
Description
AI-based health management methods, devices, equipment, and storage media Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a health management method, apparatus, device, and storage medium based on artificial intelligence. Background Technology
[0002] In the field of medical technology, intelligent health management engines, such as intelligent health management assistants, intelligent health management customer service, and health management robots, are widely used for the daily monitoring and management of individual health status. These intelligent health management engines, relying on wearable devices and mobile terminals, can continuously track physiological data such as heart rate, steps, and sleep, and provide health management support to users through mobile reminders and health reports. However, despite some progress in data collection and presentation, these intelligent health management engines still fall short in human-computer interaction, thus affecting the effectiveness of individual health management.
[0003] Specifically, firstly, the interaction between these health management smart engines and users is primarily limited to static feedback of physiological data, usually presented in numerical or graphical form, lacking personalized communication methods. This makes it difficult for users to feel a close connection to their own health. Secondly, they often use incentive mechanisms such as "points rewards" or "achievement badges" to motivate users to participate in health management, which lacks sufficient appeal. Over time, users tend to lose interest, thus affecting the long-term effectiveness of health management. Furthermore, they are prone to pushing information at inappropriate times or in inappropriate scenarios, which may cause user resentment and resistance, further reducing user engagement and satisfaction. Summary of the Invention
[0004] This invention provides a health management method, device, equipment, and storage medium based on artificial intelligence, aiming to solve the technical problems of existing health management intelligent engines, such as lack of personalized communication with users, insufficient incentive mechanisms to encourage users to participate in health management for a long time, and inappropriate timing of interaction, which affect the long-term effectiveness of individual health management.
[0005] Firstly, an artificial intelligence-based health management method is provided, comprising: acquiring a user's multimodal integrated state vector, multidimensional psychological profile, and contextual state vector during an interaction, wherein the multimodal integrated state vector, the multidimensional psychological profile, and the context of the current health task constitute the interaction state; matching an incentive action applicable to the user under the current health task based on the contextual state vector; calculating the reward generated by taking the incentive action for the user in the interaction state; learning and generating a target incentive action for the user based on the interaction state, the incentive action, and the reward using a reinforcement learning algorithm; and incentivizing the user to perform health behaviors based on the target incentive action.
[0006] Secondly, an artificial intelligence-based health management device is provided, comprising: an acquisition module for acquiring a user's multimodal integrated state vector, multidimensional psychological profile, and contextual state vector during an interaction, wherein the multimodal integrated state vector, the multidimensional psychological profile, and the context of the current health task constitute the interaction state; a matching module for matching an incentive action applicable to the user under the current health task based on the contextual state vector; a calculation module for calculating the reward generated by taking the incentive action for the user in the interaction state; a generation module for learning and generating a target incentive action for the user based on the interaction state, the incentive action, and the reward using a reinforcement learning algorithm; and an incentive module for incentivizing the user to perform health behaviors based on the target incentive action.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based health management method.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned artificial intelligence-based health management method.
[0009] In the aforementioned AI-based health management methods, devices, equipment, and storage media, user health behavior incentives are abstracted into reinforcement learning tasks. On one hand, by acquiring the user's multimodal integrated state vector, multidimensional psychological profile, and contextual state vector during the interaction process, the multimodal integrated state vector, multidimensional psychological profile, and current health task context are combined to form an interaction state. On the other hand, the contextual state vector is used to match suitable incentive actions for the user under the current health task, ensuring that the matched incentive actions are consistent with the user's own condition. Furthermore, the reward generated after taking the incentive action is calculated. The interaction state, incentive action, and reward are combined to form reinforcement learning experience. Reinforcement learning algorithms are then used to learn based on this experience. By automatically generating personalized goal-oriented incentive actions, this process comprehensively considers the user's overall state reflected by the multimodal integrated state vector, the psychological state presented by the multidimensional psychological profile, and the specific context provided by the current health task when generating personalized goal-oriented incentive actions, thereby improving the adaptability of goal-oriented incentive actions. Ultimately, by proactively motivating users to engage in health behaviors through personalized goal-oriented incentive actions, the health management approach is effectively promoted from "passive health reminders" to "empathic health companionship," achieving a more natural and intelligent human-computer interaction. It can increase user participation in health behaviors at appropriate times and in a way that is closer to the user's emotional state and needs, thereby prompting users to form good health habits and significantly improving the long-term effects of individual health management. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 is a flowchart of a health management method based on artificial intelligence according to an embodiment of the present invention; Figure 2 is a flowchart of a specific implementation of step S20 in Figure 1; Figure 3 is a flowchart of a specific implementation of step S40 in Figure 1; Figure 4 is a structural schematic diagram of a health management device based on artificial intelligence according to an embodiment of the present invention; Figure 5 is a structural schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] The AI-based health management method provided in this invention can be applied to a server. The server can acquire the user's multimodal comprehensive state vector, multidimensional psychological profile, and contextual state vector during the interaction process. The multimodal comprehensive state vector, multidimensional psychological profile, and context of the current health task constitute the interaction state. The server matches appropriate incentive actions for the user under the current health task based on the contextual state vector; calculates the reward generated by taking incentive actions for the user in the interaction state; and uses a reinforcement learning algorithm to learn and generate target incentive actions for the user based on the interaction state, incentive actions, and rewards. The server then motivates the user to perform health behaviors based on the target incentive actions. In this way, user health behavior incentives are abstracted into reinforcement learning tasks. On the one hand, by acquiring the user's multimodal integrated state vector, multidimensional mental profile, and contextual state vector during the interaction process, the multimodal integrated state vector, multidimensional mental profile, and current health task context are combined to form an interaction state. On the other hand, the contextual state vector is used to match appropriate incentive actions for the user under the current health task, ensuring that the matched incentive actions are consistent with the user's own condition. Furthermore, the reward generated after taking the incentive action is calculated. Combining the interaction state, incentive action, and reward forms reinforcement learning experience. Using reinforcement learning algorithms, the user learns from this experience to automatically generate personalized target incentive actions. This process enables the generation of personalized goal-oriented incentive actions to comprehensively consider the user's overall state reflected by the multimodal integrated state vector, the psychological state presented by the multi-dimensional psychological profile, and the specific context provided by the current health task, thereby improving the adaptability of the goal-oriented incentive actions. Ultimately, by proactively motivating users to engage in health behaviors through personalized goal-oriented incentive actions, it effectively promotes the transformation of health management from "passive health reminders" to "empathic health companionship," achieving more natural and intelligent human-computer interaction. It can increase user participation in health behaviors at appropriate times and in a way that is closer to the user's emotional state and needs, thereby encouraging users to form good health habits and significantly improving the long-term effects of individual health management. The server-side can be implemented using a separate server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] To facilitate understanding, the terms involved in this invention will first be explained: Artificial intelligence (AI): is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. AI also refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0015] Reinforcement Learning (RL) uses a trial-and-error approach to guide an agent to select actions in a specific state and adjust those actions based on reward or penalty feedback to maximize the reward. The core elements of reinforcement learning include the agent, environment, state, action, and reward. The agent is the learner performing the action; the environment is the external system the agent interacts with; the state is the specific circumstances of the environment, i.e., the situation the agent is in at a given moment; the action is the behavior the agent can take in a specific state; and the reward is the feedback the agent receives after performing an action, used to evaluate the quality of that action.
[0016] Based on this, the health management method provided by the present invention will be described in detail below.
[0017] Please refer to Figure 1. Figure 1 is a flowchart of a health management method based on artificial intelligence provided in an embodiment of the present invention, including the following steps: S10: Obtain the user's multimodal integrated state vector, multidimensional psychological profile and contextual state vector during the interaction process, wherein the multimodal integrated state vector, multidimensional psychological profile and the context of the current health task constitute the interaction state.
[0018] The AI-based health management method provided by this invention can be applied to intelligent health management engines such as intelligent health management assistants, intelligent health management customer service, and health management robots in individual health management scenarios. The intelligent health management engine is implemented through a server-side component, and users interact with this server through a client-side component.
[0019] For step S10, the server can obtain the user's multimodal integrated state vector, multidimensional psychological profile, and contextual state vector during the interaction process.
[0020] A multimodal integrated state vector is a unified representation of a user's multimodal state data. It integrates information from emotional, behavioral, and semantic modalities, reflecting the user's overall state, and is denoted as [vector name missing]. .
[0021] A multi-dimensional psychological profile mainly consists of three parts: mood spectrum, behavioral preferences, and cognitive style. It can reflect the user's psychological state and is recorded as follows: .
[0022] The context state vector is a representation of the user's state on the client side, reflecting the user's short-term state during the interaction, denoted as . .
[0023] Current health tasks refer to specific activities that users need to complete by interacting with the health management intelligent engine, such as "sleep monitoring tasks", "sleep suggestions", "medication reminder tasks", or "early morning check-in challenges", etc.
[0024] The context of the current health task refers to the situation or context of the current health task. Essentially, it's information about the current health task dimension, used to help the health management intelligent engine understand the background of the motivational action. It is denoted as... .
[0025] This invention abstracts the incentive of user health behaviors into a state-action task in reinforcement learning. Reinforcement learning algorithms are applied to learn how to select the most effective incentive action under different emotional states or needs, adapting to individual user differences and maximizing user health improvement, long-term engagement, and satisfaction. Specifically, the core element of reinforcement learning—the state—is abstracted into an interaction state. The interaction state consists of a multimodal comprehensive state vector, a multi-dimensional psychological profile, and the context of the current health task, denoted as... .
[0026] In step S10 of some embodiments, multimodal state data of the user during the interaction process can be obtained, wherein the multimodal state data includes emotional data, behavioral data, semantic data, cognitive data and environmental data; a multimodal comprehensive state vector is obtained based on the emotional data, behavioral data and semantic data; a multidimensional psychological profile is constructed based on the emotional data, behavioral data and cognitive data; and a contextual state vector is obtained based on the emotional data, behavioral data and environmental data.
[0027] Multimodal state data of users can be stably collected through various sensors (such as wearable devices, home sensors, and user-end apps).
[0028] Multimodal state data includes emotional data, behavioral data, semantic data, cognitive data, and environmental data.
[0029] Emotional data includes psychological parameters such as facial expressions (e.g., joy, anger, sadness, or anxiety), tone of voice, and emotional content in written text; behavioral data includes motor parameters such as gait and step count, as well as physiological parameters such as heart rate and sleep; semantic data includes the meaning of written and spoken expressions; cognitive data includes language expression patterns (e.g., vocabulary selection or sentence complexity) and interaction response (e.g., response speed); and environmental data includes environmental factors such as time, location, and weather.
[0030] Based on the user's multimodal state data, on the one hand, the user's multimodal comprehensive state vector is obtained according to emotion data, behavioral data and semantic data.
[0031] Specifically, a pre-defined multimodal fusion model can be used to fuse emotional data, behavioral data, and semantic data into a multimodal comprehensive state vector. The multimodal fusion model is shown below:
[0032] in, This represents data from the m-th modality, namely sentiment data, behavioral data, and semantic data; Modal encoders include ResNet for encoding emotion data, Convolutional Neural Network (CNN) for encoding behavioral data, and Bidirectional Encoder Representations from Transformers (BERT) for encoding semantic data. The weight of a modality is used to balance the contribution of different modalities. This weight can be dynamically adjusted based on the attention score. For example, if the emotion data reflects extreme emotional changes, the weight of this modality may be increased, thus having a greater impact on the user's overall multimodal state assessment.
[0033] In other words, through the multimodal fusion model, emotional data can be encoded to obtain emotional features; behavioral data can be encoded to obtain behavioral features; semantic data can be encoded to obtain semantic features; and emotional features, behavioral features, and semantic features can be fused to obtain a multimodal comprehensive state vector, thereby achieving effective perception and in-depth understanding of the user's emotional intensity, behavioral activity, and psychological stress.
[0034] On the other hand, multi-dimensional psychological profiles are constructed based on emotional data, behavioral data, and cognitive data.
[0035] Specifically, the Plutchik emotion model can be used to map emotion data to the user's emotion spectrum, denoted as... Emotional spectrum, for example, is an 8-dimensional emotion vector captured by the Plutchik emotion model, which measures the intensity of eight emotions such as joy, trust, fear, surprise, sadness, disgust, anger, and anticipation.
[0036] Behavioral data can be processed to extract features and obtain user behavioral preferences (such as "quieter" or "regular schedule"), which are then recorded as follows. This enables a deeper understanding of users' individual behavioral tendencies.
[0037] Cognitive data can be analyzed to obtain users' cognitive styles, which are denoted as... For example, cognitive data can be analyzed to determine a user's information processing methods, thinking patterns, and reaction speed, thus clarifying their cognitive style. A cognitive style could be described as "logical + quick-response + straightforward expression".
[0038] In this way, a multi-dimensional psychological profile of a user can be constructed from sentiment spectrum, behavioral preferences, and cognitive style, denoted as:
[0039] On the other hand, contextual state vectors are obtained based on emotion data, behavioral data, and environmental data.
[0040] Specifically, a pre-defined dual-stream neural network (Dual-Stream Context Encoder, DSCE) can be used to encode emotional data to obtain an emotional encoding vector; to encode behavioral data to obtain a behavioral encoding vector; and to encode environmental information to obtain an environmental encoding vector. The emotional encoding vector, behavioral encoding vector, and environmental encoding vector constitute the user's context state vector, such as the vector corresponding to a user who is "fatigued, depressed, and in their bedroom at 22:30".
[0041] Therefore, by using a dual-stream neural network to process complex emotional, behavioral, and environmental data, a deeper understanding of the user's contextual state can be achieved.
[0042] S20: Match the appropriate incentive action for the user under the current health task based on the context state vector.
[0043] For step S20, based on the context state vector, the appropriate incentive action for the user under the current health task is matched from the preset interaction intent library to ensure that the matched incentive action is both consistent with the specific context and fits the user's short-term state.
[0044] In some embodiments, please refer to Figure 2. Step S20 may include, but is not limited to, the following steps: S21: Filter out candidate incentive actions available for the current health task from a preset interaction intent library; S22: Calculate the matching degree between the context state vector and the vector of the candidate incentive actions; S23: Filter out the incentive actions applicable to the user from the candidate incentive actions based on the matching degree.
[0045] Specifically, the core element of reinforcement learning—action—is abstracted into selectable incentive actions. These selectable incentive actions can be understood as various dialogue strategies that motivate users to engage in healthy behaviors, such as the dialogue methods, content, and style used when interacting with users.
[0046] These available motivational actions include, but are not limited to, various types such as reminders, encouragement, empathy, challenges, and social incentives. For example, encouragement-based motivational actions include textual encouragement, verbal encouragement, and so on. Exciting actions such as self-challenge, friend competition, etc.
[0047] To facilitate differentiation and matching, a mapping relationship between health tasks and their available incentive actions can be established in advance. This mapping relationship could be, for example, a mapping between reminder-type health tasks and reminder-type incentive actions, a mapping between challenge-type health tasks and challenge-type incentive actions, and so on.
[0048] This mapping relationship is stored in a preset interaction intent library, which also stores the vector of each stimulus action.
[0049] In this way, the current health task can be compared with the interaction intent library. Based on the mapping relationship between the health task and the available incentive actions in the interaction intent library, all available incentive actions for the current health task can be selected as candidate incentive actions.
[0050] Next, the matching degree between the context state vector and the vector of each candidate stimulus action is calculated, specifically using the matching degree calculation formula shown below: in, Represents the context state vector The vector of the k-th candidate activation action The degree of matching between them; This represents the weight matrix on the user side. The weight matrices represent the vectors of candidate incentive actions. These weight matrices can be learned through pre-training to maximize the matching effect.
[0051] By calculating the degree of matching between the context state vector and the vector of each candidate stimulus, the vector of the candidate stimulus with the highest matching degree is determined. :
[0052] Will The corresponding candidate incentive action is the incentive action applicable to the user, denoted as . .
[0053] In this way, by gaining a deep understanding of the user's contextual state, appropriate incentive actions can be matched to the user, so as to provide experience for subsequent reinforcement learning.
[0054] S30: Calculate the reward generated when an incentive action is taken against the user in an interactive state.
[0055] For step S30, obtain the specific interaction state. The rewards generated from the user's incentive actions are used to provide experience for subsequent reinforcement learning.
[0056] In step S30 of some embodiments, a preset reward function can be obtained; the reward generated by taking the incentive action is calculated and processed through the reward function to obtain the reward.
[0057] The core of reinforcement learning lies in evaluating the effectiveness of incentive actions through a reward function. In other words, the reward function evaluates the effect of the selected incentive action in a specific interaction state. The reward function is defined as follows:
[0058] in, Indicates a reward; Indicates the intensity of emotional response, such as a positivity score calculated based on tone of voice, facial expressions, or written expression (e.g., [+1, -1]). This indicates the completion rate of health tasks, such as check-in, step count achievement rate, etc. Indicates the number of times the user actively interacted; They represent The weighting coefficients.
[0059] In a specific interaction state The following calculations are performed using a reward function to determine the appropriate incentive actions for the user. The resulting rewards are used to measure the motivational actions. The effect.
[0060] S40: Through reinforcement learning algorithms, generate target incentive actions for users based on interaction states, incentive actions, and rewards.
[0061] For step S40, by integrating reinforcement learning and psychological motivation theory, and using reinforcement learning algorithms, the system learns to automatically determine suitable target motivation actions for users based on interaction states, motivational actions, and rewards, forming appropriate and effective personalized health behavior guidance. This provides accurate and scientific guidance for users' health management, thereby significantly improving the long-term effectiveness of health management.
[0062] Please refer to Figure 3. In some embodiments, step S40 may include, but is not limited to, the following steps: S41: Obtain a new interaction state generated by taking an incentive action for the user in the interaction state; S42: Train a preset reinforcement learning model for reinforcement learning based on the interaction state, incentive action, reward and new interaction state to obtain a trained reinforcement learning model; S43: Generate a target incentive action for the user through the trained reinforcement learning model.
[0063] For steps S41-S43, obtain the specific interaction state. Next, take incentive actions for users. The resulting new interactive state is denoted as .
[0064] Based on the interaction state, incentive actions, rewards, and new interaction states, the pre-set reinforcement learning model is trained to learn how to select more effective incentive actions, resulting in a well-trained reinforcement learning model.
[0065] The trained reinforcement learning model automatically selects personalized target incentive actions for users to guide them in engaging in healthy behaviors.
[0066] In step S42 of some embodiments, the interaction state, incentive action, reward and new interaction state can be input into the reinforcement learning model for reinforcement learning training; during the training process, the policy function of the reinforcement learning model is iteratively updated until the number of iterations is greater than a preset threshold, and a trained reinforcement learning model is obtained.
[0067] Specifically, the preset reinforcement learning model can be a DNQ (Deep Q-Network) model. The DNQ model combines deep learning and the traditional Q-learning algorithm. Its core idea is to use deep neural networks to approximate the state-action value function (Q-value function, which is abstracted as a policy function in this embodiment of the invention) in the traditional Q-learning algorithm through reinforcement learning, thereby continuously optimizing the expected reward (i.e., Q-value) of the action to decide the most appropriate action (in this embodiment of the invention, this is reflected in selecting the most effective incentive action).
[0068] The interaction state, incentive action, reward, and new interaction state are input into the DNQ model for reinforcement learning training, and the policy function is iteratively updated during the training process.
[0069] in, Represents the policy function; Indicates the learning rate; Indicates the discount factor; This indicates the stimulus action that may be selected in the new interaction state.
[0070] By iteratively updating the policy function, the DNQ model gradually learns which excitation actions will bring higher Q values under different interaction states, until the number of iterations exceeds a preset threshold, forming a more accurate policy function and obtaining a well-trained DNQ model.
[0071] In this way, by inputting interaction states, incentive actions, rewards, and new interaction states into the DNQ model for training, the DNQ model can continuously adjust and optimize its policy function, enabling the trained DNQ model to more intelligently select target incentive actions that better match the user's emotional state or emotional needs.
[0072] A well-trained DNQ model can automatically determine which target incentive action to use in different situations and user emotions. For example, when a user's heart rate is high and they are emotionally stressed, it can choose to output an empathetic voice prompt and recommend relaxing light music.
[0073] If a user's sleep duration consistently meets the target for 3 consecutive days, a reward point and encouraging messages will be sent to them.
[0074] If a user is detected to have been inactive for an extended period, invite friends to participate in a step-counting challenge.
[0075] S50: Guide users to adopt healthy behaviors based on goal-oriented incentive actions.
[0076] Ultimately, based on goal-oriented incentive actions, users are guided to adopt healthy behaviors. For example, when a user's heart rate is high and they are emotionally stressed, an empathetic voice prompt is output and relaxing light music is recommended to help the user relieve stress.
[0077] If a user's sleep duration meets the target for 3 consecutive days, a reward point and encouraging message will be sent to enhance the user's sense of accomplishment and increase the user's motivation to continue to maintain a good sleep duration.
[0078] If a user is detected to have been inactive for an extended period, a social challenge can be launched to invite friends to participate in step counting together, thereby encouraging the user to exercise.
[0079] Therefore, targeted goal-oriented incentive actions can enhance users' intrinsic motivation and sense of participation, thereby prompting users to continuously maintain healthy behaviors.
[0080] In step S50 of some embodiments, a virtual health role can be assigned to the trained reinforcement learning model; the virtual health role can then be used to present incentive actions to the user in order to encourage the user to engage in health behaviors.
[0081] To achieve the goals of long-term companionship and self-motivation, a virtual health role is assigned to the trained reinforcement learning model, and a Bayesian update model is used to adjust the characteristics of the virtual health role, thereby better establishing an emotional connection with the user and enhancing the user's sense of continuous engagement.
[0082] Specifically, the virtual health persona is constructed based on the user's multi-dimensional psychological profile and historical interaction dynamics, and has evolvable personality characteristics, including: tone: it can be gentle, formal, friendly, encouraging, etc., and gradually adjusts as the user interacts, reflecting the user's preference for the interaction method.
[0083] Emotional intensity: This can adjust the intensity of the emotions expressed by the character, such as using a positive tone when encouraging the user, or using appropriate concern when reminding the user.
[0084] Interaction style: Based on user feedback and interaction history, adjust the frequency of interaction, the complexity of statements, and the level of detail in responses to maintain user engagement and meet their needs.
[0085] These characteristics enable the virtual health persona to go beyond passive responses, allowing it to continuously grow and evolve through interaction with users. This growth and evolution process is based on a Bayesian update model, the formula of which is shown below:
[0086] in, These represent the personality parameters of the virtual health avatar (such as tone of voice, emotional intensity, and interaction style). This represents user interaction data (such as user input, feedback, and historical conversation content). Indicates that in a given Under these conditions, observed The possibility; express The prior distribution reflects the initial assumptions about the character's personality traits.
[0087] Bayesian updates allow virtual characters to update their personality parameters with each user interaction, achieving both personalization and dynamic adaptation. This adjustment process is gradual, not abrupt, ensuring that users always feel a sense of "resonance" with the virtual character—the character's changes are natural and in line with user expectations.
[0088] In this way, by interacting with users through virtual health avatars, users can be motivated to take action based on their own goals, which can inspire a sense of companionship and growth, thereby encouraging them to engage in healthy behaviors.
[0089] The reason why virtual health avatars can evoke a sense of companionship for users is that as they evolve, they become more and more aligned with users' needs and personalities, thus making users feel understood and accompanied.
[0090] The reason for resonating with users' growth is that as the virtual health avatar evolves, users also feel themselves growing through interaction with it. The adaptive feedback from the virtual health avatar not only meets current needs but also encourages users to develop in a better direction, thus creating a shared experience of growth.
[0091] Therefore, the sense of companionship and growth resonance brought by this virtual health role based on Bayesian updates can significantly enhance users' stickiness to the virtual health role, making the relationship between users and the virtual health role closer and deeper, forming a long-term companionship effect. Users will gradually form a self-driven health behavior pattern under the guidance of the virtual role, thereby stimulating their intrinsic motivation and enhancing their willingness to participate in health management in the long term.
[0092] After step S50 in some embodiments, feedback information after the user takes a target incentive action can be obtained; the multi-dimensional mental profile is updated based on the feedback information to obtain an updated multi-dimensional mental profile; and the trained reinforcement learning model is optimized and trained based on the updated multi-dimensional mental profile.
[0093] To ensure the long-term effectiveness of personalized incentives and optimize target incentive actions, the multi-dimensional psychological profile can be dynamically updated based on feedback information (such as the degree of improvement in physiological parameters) after each user takes the target incentive action. This continuous updating can be achieved using the following profile update formula:
[0094] in, A multi-dimensional psychological profile representing time step t+1; This represents a multi-dimensional psychological profile at time step t. This represents the time decay coefficient, used to control the effect of time. This indicates changes in the multi-dimensional psychological profile learned from recent interactions, reflecting changes in the user's emotional spectrum, behavioral preferences, or cognitive style during recent interactions.
[0095] In this way, the user's multi-dimensional psychological profile can be accurately updated, resulting in an updated multi-dimensional psychological profile. Then, based on the updated multi-dimensional psychological profile, the trained reinforcement learning model can be optimized and trained to improve the policy function of the trained reinforcement learning model. This enables the trained reinforcement learning model to achieve self-evolutionary learning, making the optimized reinforcement learning model increasingly adaptable to the user and able to determine more personalized and appropriate incentive actions, thereby further improving the effectiveness of health guidance.
[0096] As can be seen, the above solution realizes an end-to-end adaptive closed-loop health management mechanism, including: multimodal state data collection - multimodal state feature fusion - multi-dimensional psychological profile construction - context and incentive action matching - personalized incentive action generation - user feedback - model optimization. This closed-loop mechanism means that the technical solution of this invention effectively promotes the transformation of health management from "passive health reminders" to "empathic health companionship," achieving more natural and intelligent human-computer interaction; it also significantly improves users' long-term compliance, allowing users to enjoy a unique "health companion" experience and enhance emotional connection; and it can intervene at appropriate times and in appropriate ways, ensuring the accuracy of contextual responses, reducing user aversion, and increasing acceptance rates. Furthermore, this mechanism can continuously guide users in self-health management in non-clinical scenarios, thereby effectively reducing the risk of chronic diseases and medical costs, and demonstrating significant long-term value. Therefore, the technical solution of this invention redefines health management through intelligent and personalized methods, providing users with a more user-friendly and humanized long-term health management experience.
[0097] The AI-based health management method provided in the above embodiments has the following effective effects: This embodiment abstracts user health behavior incentives into reinforcement learning tasks. By acquiring the user's multimodal integrated state vector, multidimensional psychological profile, and contextual state vector during the interaction process, the multimodal integrated state vector, multidimensional psychological profile, and current health task context are combined to form an interaction state. Simultaneously, the contextual state vector is used to match suitable incentive actions for the user under the current health task, ensuring that the matched incentive actions are consistent with the user's own condition. Next, the reward generated after taking the incentive action is calculated, and the interaction state, incentive action, and reward are combined to form reinforcement learning experience. Using a reinforcement learning algorithm, learning is performed based on this experience to automatically... The process of generating personalized goal-oriented incentive actions allows for a comprehensive consideration of the user's overall state reflected by the multimodal integrated state vector, the psychological state presented by the multidimensional psychological profile, and the specific context provided by the current health task. This enhances the adaptability of the goal-oriented incentive actions. Ultimately, by proactively motivating users to engage in healthy behaviors through personalized goal-oriented incentive actions, the health management approach is effectively transformed from "passive health reminders" to "empathic health companionship." This achieves a more natural and intelligent human-computer interaction, increasing user participation in health behaviors at appropriate times and in a way that is closer to the user's emotional state and needs. This, in turn, encourages users to form good health habits, thereby significantly improving the long-term effectiveness of individual health management.
[0098] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0099] It should be noted that the software tools or components not belonging to our company that appear in the embodiments of this invention are merely illustrative examples and do not represent actual use.
[0100] In one embodiment, an AI-based health management device is provided, which corresponds one-to-one with the AI-based health management method described in the above embodiments. As shown in Figure 4, the AI-based health management device includes an acquisition module 101, a matching module 102, a calculation module 103, a generation module 104, and an incentive module 105. The functional modules are described in detail below: The acquisition module 101 is used to acquire the user's multimodal integrated state vector, multidimensional psychological profile, and contextual state vector during the interaction process, wherein the multimodal integrated state vector, the multidimensional psychological profile, and the context of the current health task constitute the interaction state; the matching module 102 is used to match an incentive action applicable to the user under the current health task based on the contextual state vector; the calculation module 103 is used to calculate the reward generated by taking the incentive action for the user in the interaction state; the generation module 104 is used to learn and generate a target incentive action for the user based on the interaction state, the incentive action, and the reward using a reinforcement learning algorithm; the incentive module 105 is used to incentivize the user to perform health behaviors based on the target incentive action.
[0101] In one embodiment, the generation module 104 is specifically configured to: acquire a new interaction state generated by taking the incentive action for the user in the interaction state; perform reinforcement learning training on a preset reinforcement learning model based on the interaction state, the incentive action, the reward and the new interaction state to obtain a trained reinforcement learning model; and generate a target incentive action for the user through the trained reinforcement learning model.
[0102] In one embodiment, the generation module 104 is further configured to: input the interaction state, the incentive action, the reward and the new interaction state into the reinforcement learning model for reinforcement learning training; and iteratively update the policy function of the reinforcement learning model during the training process until the number of iterations is greater than a preset threshold, thereby obtaining the trained reinforcement learning model.
[0103] In one embodiment, the incentive module 105 is specifically used to: assign a virtual health role to the trained reinforcement learning model; and present the target incentive action to the user through the virtual health role to incentivize the user to perform health behaviors.
[0104] In one embodiment, the matching module 102 is specifically used to: filter candidate incentive actions available for the current health task from a preset interaction intent library; calculate the matching degree between the context state vector and the vector of the candidate incentive action; and filter the incentive action applicable to the user from the candidate incentive actions according to the matching degree.
[0105] In one embodiment, the calculation module 103 is specifically used to: obtain a preset reward function; and calculate the reward generated by taking the incentive action through the reward function to obtain the reward.
[0106] In one embodiment, the acquisition module 101 is specifically configured to: acquire the user's multimodal state data during the interaction process, wherein the multimodal state data includes emotional data, behavioral data, semantic data, cognitive data, and environmental data; acquire the multimodal comprehensive state vector based on the emotional data, the behavioral data, and the semantic data; construct the multidimensional psychological profile based on the emotional data, the behavioral data, and the cognitive data; and acquire the contextual state vector based on the emotional data, the behavioral data, and the environmental data.
[0107] This invention provides an artificial intelligence-based health management device that abstracts user health behavior incentives into reinforcement learning tasks. On one hand, it acquires the user's multimodal integrated state vector, multidimensional psychological profile, and contextual state vector during the interaction process, combining these elements with the context of the current health task to form an interaction state. On the other hand, it uses the contextual state vector to match suitable incentive actions for the user under the current health task, ensuring that the matched incentive actions are consistent with the user's own condition. Furthermore, it calculates the reward generated after taking the incentive action, combining the interaction state, incentive action, and reward to form reinforcement learning experience. Using a reinforcement learning algorithm, it learns based on this experience to automatically generate personalized health management tasks. Personalized goal-oriented incentive actions, in this process, comprehensively consider the user's overall state reflected by the multimodal integrated state vector, the psychological state presented by the multidimensional psychological profile, and the specific context provided by the current health task when generating personalized goal-oriented incentive actions, thereby improving the adaptability of goal-oriented incentive actions. Ultimately, by actively motivating users to engage in health behaviors through personalized goal-oriented incentive actions, the health management approach is effectively promoted from "passive health reminders" to "empathic health companionship," achieving a more natural and intelligent human-computer interaction. It can increase user participation in health behaviors at appropriate times and in a way that is closer to the user's emotional state and needs, thereby prompting users to form good health habits and significantly improving the long-term effects of individual health management.
[0108] For specific limitations regarding AI-based health management devices, please refer to the limitations of AI-based health management methods described above, which will not be repeated here. The modules in the aforementioned AI-based health management device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0109] In one embodiment, a computer device is provided, which can be a server, and its internal structure diagram is shown in Figure 5. The computer device includes a processor, memory, a network interface, and a database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side health management method based on artificial intelligence.
[0110] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it performs the following steps: acquiring a user's multimodal integrated state vector, multidimensional mental profile, and contextual state vector during an interaction, wherein the multimodal integrated state vector, the multidimensional mental profile, and the context of the current health task constitute the interaction state; matching an incentive action applicable to the user under the current health task based on the contextual state vector; calculating a reward generated by taking the incentive action for the user in the interaction state; learning and generating a target incentive action for the user based on the interaction state, the incentive action, and the reward using a reinforcement learning algorithm; and incentivizing the user to perform health behaviors based on the target incentive action.
[0111] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program performs the following steps: acquiring a user's multimodal integrated state vector, multidimensional mental profile, and contextual state vector during an interaction, wherein the multimodal integrated state vector, the multidimensional mental profile, and the context of the current health task constitute the interaction state; matching an incentive action applicable to the user under the current health task based on the contextual state vector; calculating the reward generated by taking the incentive action for the user in the interaction state; learning and generating a target incentive action for the user based on the interaction state, the incentive action, and the reward using a reinforcement learning algorithm; and incentivizing the user to perform health behaviors based on the target incentive action.
[0112] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0113] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0114] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0115] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A health management method based on artificial intelligence, characterized in that, include: The system acquires the user's multimodal integrated state vector, multidimensional psychological profile, and contextual state vector during the interaction process, wherein the multimodal integrated state vector, the multidimensional psychological profile, and the context of the current health task constitute the interaction state; it matches an incentive action applicable to the user under the current health task based on the contextual state vector; it calculates the reward generated by taking the incentive action for the user in the interaction state; it learns and generates a target incentive action for the user based on the interaction state, the incentive action, and the reward using a reinforcement learning algorithm; and it motivates the user to perform health behaviors based on the target incentive action.
2. The health management method as described in claim 1, characterized in that, The step of generating a target incentive action for the user based on the interaction state, the incentive action, and the reward using a reinforcement learning algorithm includes: acquiring a new interaction state generated by taking the incentive action for the user in the interaction state; training a preset reinforcement learning model using reinforcement learning based on the interaction state, the incentive action, the reward, and the new interaction state to obtain a trained reinforcement learning model; and generating a target incentive action for the user using the trained reinforcement learning model.
3. The health management method as described in claim 2, characterized in that, The step of training a preset reinforcement learning model based on the interaction state, the incentive action, the reward, and the new interaction state to obtain a trained reinforcement learning model includes: inputting the interaction state, the incentive action, the reward, and the new interaction state into the reinforcement learning model for reinforcement learning training; and iteratively updating the policy function of the reinforcement learning model during the training process until the number of iterations exceeds a preset threshold, thereby obtaining the trained reinforcement learning model.
4. The health management method as described in claim 2, characterized in that, The step of incentivizing the user to perform health behaviors based on the target incentive action includes: assigning a virtual health role to the trained reinforcement learning model; and presenting the target incentive action to the user through the virtual health role to incentivize the user to perform health behaviors.
5. The health management method as described in claim 1, characterized in that, The step of matching the incentive action applicable to the user under the current health task based on the context state vector includes: filtering candidate incentive actions available for the current health task from a preset interaction intent library; calculating the matching degree between the context state vector and the vector of the candidate incentive action; and filtering the incentive action applicable to the user from the candidate incentive actions based on the matching degree.
6. The health management method as described in claim 1, characterized in that, The calculation of the reward generated by the user taking the incentive action in the interactive state includes: obtaining a preset reward function; and calculating the reward generated by taking the incentive action through the reward function to obtain the reward.
7. The health management method according to any one of claims 1 to 6, characterized in that, The acquisition of the user's multimodal integrated state vector, multidimensional psychological profile, and contextual state vector during the interaction process includes: acquiring the user's multimodal state data during the interaction process, wherein the multimodal state data includes emotional data, behavioral data, semantic data, cognitive data, and environmental data; acquiring the multimodal integrated state vector based on the emotional data, behavioral data, and semantic data; constructing the multidimensional psychological profile based on the emotional data, behavioral data, and cognitive data; and acquiring the contextual state vector based on the emotional data, behavioral data, and environmental data.
8. A health management device based on artificial intelligence, characterized in that, include: The system includes an acquisition module for acquiring the user's multimodal integrated state vector, multidimensional psychological profile, and contextual state vector during the interaction process, wherein the multimodal integrated state vector, the multidimensional psychological profile, and the context of the current health task constitute the interaction state; a matching module for matching an incentive action applicable to the user under the current health task based on the contextual state vector; a calculation module for calculating the reward generated by taking the incentive action for the user in the interaction state; a generation module for learning and generating a target incentive action for the user based on the interaction state, the incentive action, and the reward using a reinforcement learning algorithm; and an incentive module for incentivizing the user to perform health behaviors based on the target incentive action.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the artificial intelligence-based health management method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the artificial intelligence-based health management method as described in any one of claims 1 to 7.