Recommendation method and training method of medical care and health care service, electronic equipment and medium

By training medical care and health care service recommendation agents, using user tags to build demand portraits and conduct reinforcement learning, the problems of supply and demand matching and experience perception of existing medical care and health care services are solved, and accurate, efficient and intelligent service recommendations are achieved.

CN120179892AActive Publication Date: 2025-06-20CENT SOUTH UNIV

Patent Information

Application Number
CN202510150164.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-20
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The existing intelligent medical and health care services have difficulties in supply and demand matching, aging-friendly, experience perception and subsequent service implementation, which has affected service quality and efficiency.

Method used

A training method for recommending medical and health care services is proposed. By obtaining training data of medical and health care services and user tags, a user needs portrait is constructed, user tag correlation is calculated, and medical and health care services are predicted and recommended based on reinforcement learning methods.

Benefits of technology

It has achieved accurate, efficient and intelligent recommendations of medical and health care services, and improved the quality of services and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179892A_ABST
    Figure CN120179892A_ABST
Patent Text Reader

Abstract

The invention provides a medical care health service recommendation method, a training method, an electronic device and a medium, and the method comprises the steps: taking a medical care health service and at least two types of user tags as training data, training an intelligent agent, enabling the intelligent agent to complete the construction of a user demand portrait and the calculation of the correlation degree of the user tags, and carrying out the prediction of the medical care health service, through the training implementation of the intelligent agent, the medical care and health care service can be more accurate, efficient and intelligent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a recommendation method, a training method, an electronic device, and a medium for medical, health, and elderly care services. Background Art

[0002] The demand of the elderly population for promoting health is becoming increasingly diversified. They not only need medical care but also health care and elderly care, and more importantly, they need diversified integrated services that combine medical care, health care, and elderly care. However, the existing intelligent medical, health, and elderly care services face the problem of effective and efficient matching of supply and demand, which in turn leads to problems such as insufficient aging adaptation, poor experience, lack of follow-up services, and difficulty in implementing the "last mile", affecting the quality and efficiency of the services. Summary of the Invention

[0003] This application proposes a recommendation method, a training method, an electronic device, and a medium for medical, health, and elderly care services, which can solve the problems of precision, efficiency, and intelligence of medical, health, and elderly care services.

[0004] To achieve the above object, this application adopts the following technical solutions:

[0005] In a first aspect, a training method for a medical, health, and elderly care service recommendation agent is provided. The training method includes:

[0006] Obtaining training data, where the training data includes: medical, health, and elderly care services, and at least two user tags related to the medical, health, and elderly care services; and

[0007] Using the training data to train an agent, where the agent is specifically configured to: construct a user demand profile using the at least two user tags; calculate the correlation degree between the user tags; and, based on the user demand profile and the correlation degree, perform prediction of medical, health, and elderly care services.

[0008] Based on the above technical solution, using medical, health, and elderly care services, and at least two user tags as training data to train an agent, the agent can complete the construction of a user demand profile, the calculation of the correlation degree of user tags, and perform prediction of medical, health, and elderly care services accordingly. The realization of this training of the agent can make medical, health, and elderly care services more precise, efficient, and intelligent.

[0009] In a possible design of the first aspect, constructing a user demand profile using the at least two user tags specifically includes:

[0010] Performing feature fusion on the at least two user tags to obtain a high-dimensional feature vector;

[0011] Based on the user tags, obtaining at least two demand predictions; and

[0012] Using a reinforcement learning method with the high-dimensional feature vector and the demand prediction as inputs, the user demand portrait is obtained.

[0013] In a possible design of the first aspect, the reinforcement learning method includes an attention mechanism and a task adaptive reward mechanism.

[0014] In a possible design of the first aspect, calculating the correlation between the user tags specifically includes:

[0015] Extracting the structured data in the user tags; and

[0016] Using a graph neural network to process the structured data to obtain the correlation.

[0017] In a possible design of the first aspect, an attention mechanism is added to the processing of the graph neural network and combined with a long short-term memory network, and the long short-term memory network is used to complete the dynamic expression of the correlation according to the unstructured data and time series data in the user tags.

[0018] In a possible design of the first aspect, based on the user demand portrait and the correlation, medical, elderly care, and rehabilitation service prediction is performed. Specifically: a reinforcement learning method including an attention mechanism and a task adaptive reward mechanism and combined with the A* algorithm is used to obtain the medical, elderly care, and rehabilitation prediction service from the user demand portrait and the correlation.

[0019] In a possible design of the first aspect, the medical, elderly care, and rehabilitation services are subject to QoS constraints; the user tags are health tags, behavior tags, environment tags, or service preference tags; the user tags are obtained through user data collection, data cleaning, noise reduction, normalization, and analysis.

[0020] In a second aspect, a method for recommending medical, elderly care, and rehabilitation services is provided. The recommendation method includes:

[0021] Obtaining the current user tags; and

[0022] Using the trained agent as above to obtain the current medical, elderly care, and rehabilitation prediction service recommended currently.

[0023] In a third aspect, an electronic device is provided. The electronic device includes: a processor and a memory coupled to the processor. The memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the training method in any possible implementation manner of the first aspect or executes the recommendation method in the second aspect.

[0024] Fourthly, a computer-readable storage medium is provided, including a computer program or instructions. When the computer program or instructions are run on a computer, the computer is caused to execute the training method according to any possible implementation manner in the first aspect, or execute the recommendation method according to the second aspect. Description of the Drawings

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the related art descriptions. Obviously, the drawings in the following descriptions are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 is the overall flowchart provided by the embodiments of the present application;

[0027] Figure 2 is the flowchart of medical data preprocessing provided by the embodiments of the present application;

[0028] Figure 3 is the flowchart of preprocessing of rehabilitation data provided by the embodiments of the present application;

[0029] Figure 4 is the flowchart of preprocessing of behavior data provided by the embodiments of the present application;

[0030] Figure 5 is the flowchart of preprocessing of environmental data provided by the embodiments of the present application;

[0031] Figure 6 is the decision-level fusion flowchart provided by the embodiments of the present application;

[0032] Figure 7 is the flowchart of processing structured data provided by the embodiments of the present application;

[0033] Figure 8 is the flowchart of processing unstructured data provided by the embodiments of the present application;

[0034] Figure 9 is the flowchart of processing time series data provided by the embodiments of the present application;

[0035] Figure 10 is the service recommendation flowchart provided by the embodiments of the present application;

[0036] Figure 11 is the feedback optimization mechanism flowchart provided by the embodiments of the present application;

[0037] Figure 12 is the management and execution flowchart of the service chain provided by the embodiments of the present application. Detailed implementation manners

[0038] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0039] It should be noted that although functional module division is performed in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different sequence in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0041] Before introducing the embodiments of the present application, a brief description of the current technical research of the present application is given first:

[0042] Demand portrait: A multi-dimensional attribute description constructed for users through multi-source data analysis and fusion, including health status, behavior preferences, environmental characteristics, etc.

[0043] Deep learning: A method of machine learning that uses a multi-layer neural network to learn high-order feature representations from data.

[0044] QoS (Quality of Service): A comprehensive index that measures service response time, reliability, availability, etc. in service recommendation.

[0045] A* algorithm: A heuristic search algorithm used to find the shortest path from the starting point to the ending point in a weighted graph.

[0046] Reinforcement learning: A method of machine learning that learns a behavioral strategy that can achieve the goal through interaction with the environment.

[0047] Multi-source heterogeneous data: Data from different data sources, formats, structures and time distributions, such as medical records, behavioral data and environmental data.

[0048] Next, refer to the attached Figures 1 to 12 , and an exemplary description will be given of the training method of the intelligent agent for medical, elderly care and health service recommendation and the medical, elderly care and health service recommendation method of the embodiments of the present application.

[0049] This technical solution includes multiple modules such as information collection, data preprocessing, image construction, service recommendation, feedback optimization, and service chain management. Through the collaboration of these modules, it is possible to efficiently construct a dynamic and accurate user demand profile, and through the improved A* algorithm and QoS optimization, achieve personalized and real-time service recommendations. At the same time, the system also continuously improves the recommendation accuracy through reinforcement learning and feedback optimization mechanisms to ensure the efficient execution of the service chain and the reasonable scheduling of resources. The specific process is as Figure 1 shown.

[0050] S1: The data sources are mainly the data collected based on the medical care, rehabilitation, and health service system, including medical data (such as electronic medical records, drug usage records, and inspection reports), rehabilitation data (such as rehabilitation assessment records, treatment plans, and training logs), behavioral data (such as activity frequency, exercise data, especially data from wearable devices, and living habits), and environmental data (such as residential location, air quality, noise level, and community service distribution). The comprehensive analysis of these multi-dimensional data can comprehensively evaluate the user's health status and provide personalized health management suggestions.

[0051] S11: The collection of medical data UM is mainly obtained through the electronic health record system (EHR) of hospitals and clinics, including the user's electronic medical records, diagnosis records, and treatment plans. In addition, drug usage records are extracted through the hospital information system (HIS), covering prescription information, drug usage frequency, and dosage, etc. Inspection reports come from the hospital's inspection and testing system, including medical test results such as blood tests and imaging tests, ensuring the comprehensive recording of the user's health information.

[0052] S12: The collection of rehabilitation data UR is mainly through the rehabilitation assessment records carried out by medical professionals. These assessments usually use standardized tools (such as the Barthel index, FIM scale, etc.) to evaluate functional recovery and track disease progression. Treatment plans are obtained through the electronic health record or rehabilitation management system and are formulated based on the doctor's diagnosis and personalized needs. In addition, training logs are recorded by the patient himself or the staff of the rehabilitation center, and the recorded content includes daily training items, time, and progress, which are recorded in the form of paper logs, mobile applications, or spreadsheets, etc.

[0053] S13: The collection of behavioral data UB is carried out through various methods, including smartphones, wearable devices (such as smart bracelets, smart watches, etc.), and motion sensors, to real-time monitor the user's activity frequency, steps, exercise time, and exercise intensity, etc. Exercise data is provided by wearable devices, covering physiological indicators such as heart rate, exercise type, and gait. Living habit data is collected through questionnaires, user self-reports, or smart home devices (such as smart refrigerators, smart lighting, etc.), recording aspects such as the user's sleep pattern, eating habits, and social activities.

[0054] S14: The collection of environmental data UE obtains the user's residential location information through the positioning function of GPS or mobile devices. The air quality data is collected by air quality monitoring instruments, recording the levels of pollutants such as PM2.5, PM10, and carbon dioxide concentration to ensure comprehensive monitoring of environmental health indicators. The noise level is monitored by noise sensors or smart devices (such as smart speakers, smartphones), recording the noise data in the living environment. The community service distribution data is obtained through the public service data platform or government information system, providing location information and service distances of facilities such as hospitals, pharmacies, gyms, parks, etc.

[0055] S2: Preprocess the data collected in the S1 stage to improve the data quality and consistency. This process includes steps such as data cleaning, noise reduction, normalization, and feature extraction.

[0056] S21: Perform data preprocessing on the medical data UM collected in S1. Medical data usually includes electronic medical records, inspection reports, and drug usage records, and these data may have duplicates, missing values, or invalid values. First, remove duplicate records and invalid information (such as expired drug information) through automated algorithms, and fill in the missing values. The missing values will be filled with industry standard values. In terms of format unification, since different hospitals or systems may use different units or formats, normalization is required during preprocessing to ensure data consistency. The preprocessed data is UMs2 = {UMs, UMns}, where the subscript s represents structured data and ns represents unstructured data. For example, for the disease diagnosis codes (such as ICD codes) used by different medical institutions, standardization is required to ensure unified data management and analysis. At the same time, the unification of drug names and the coordination of inspection result units also need to be normalized to ensure data quality. The specific process is as Figure 2 shown.

[0057] S22: Perform data preprocessing on the rehabilitation data UR collected in S1. Rehabilitation data includes rehabilitation assessment records, treatment plans, and training logs, and these data may have missing values or measurement errors. In the data cleaning stage, first remove invalid records (such as incorrect assessment results or records lacking treatment). The missing data is filled through the mean imputation method. In the noise reduction stage, the motion data in the rehabilitation training log may be affected by equipment measurement errors, and the moving average method is used to smooth the data and remove noise. The preprocessed data is URs2 = {URs, URns, URts}, where the subscript s represents structured data, ns represents unstructured data, and ts represents time series data. As Figure 3 shown.

[0058] S23: Preprocess the behavioral data UB collected in S1. Behavioral data, such as gait and movement frequency, usually comes from wearable devices or sensors, and these data are vulnerable to noise interference (such as device errors or external environmental factors). During the data cleaning process, first remove the missing data points or outliers (such as unreasonable step frequency or heart rate). When performing noise reduction, use a Kalman filter to smooth the heart rate and movement trajectory data to reduce the impact of noise on the data. In the normalization process, due to the scale differences of behavioral data (such as the number of steps, heart rate, exercise intensity, etc.), it is necessary to unify these data into the same range. By normalizing the data to the 0-1 interval and ensuring that the data has zero mean and unit variance, the differences in different scales are avoided from affecting the training effect of subsequent machine learning models. The preprocessed data is UBs2 = {UBs, UBns, UBts}, where the subscript s represents structured data, ns represents unstructured data, and ts represents time series data. As Figure 4 shown.

[0059] S24: Preprocess the environmental data UE collected in S1. Environmental data such as air quality, noise level, and community service distribution may contain missing values or abnormal fluctuations. During the data cleaning phase, remove the invalid records and fill in the missing values by mean imputation. For air quality and noise data, noise reduction is particularly important, and a smoothing algorithm needs to be used to remove short-term sudden fluctuations. In the normalization process, unify the dimensions of environmental data (such as PM2.5 concentration, noise decibels, etc.) into a standardized range for comprehensive analysis with other health data. The preprocessed data is UEs2 = {UEs, UEns}, where the subscript s represents structured data and ns represents unstructured data. As Figure 5 shown.

[0060] S3: Construct a multi-dimensional demand portrait label system for the clean data processed in the S2 stage. By analyzing the health, behavior, environment, and service preference data, a multi-dimensional user portrait label system is constructed.

[0061] S31: Input the health data UMs2, URs2 obtained from the S2 preprocessing stage, including the user's disease type, treatment records, and key health indicators (such as blood pressure, blood sugar, etc.). These data help to comprehensively depict the user's health status and disease management needs, providing basic information for personalized recommendations. By analyzing the health data, output the user's health status labels HL (such as hypertension, hyperglycemia, diabetes, healthy population, etc.), disease management requirements (such as whether long-term monitoring is needed, whether regular treatment is needed, etc.), and health risk assessment (such as potential risk assessment based on health indicators).

[0062] S32: Input the behavioral data UBs2 obtained from the S2 preprocessing stage, including the user's living habits, hobbies, and daily activity patterns. By monitoring and analyzing the user's behavior patterns, output the user lifestyle label WL (such as a healthy lifestyle person, a sports enthusiast, a socially active person, etc.), behavior pattern recognition (such as a high-risk behavior person, a low-activity person, etc.), and personalized health advice (such as increasing exercise, adjusting diet, etc.).

[0063] S33: Input the environmental data UEs2 obtained from the S2 preprocessing stage, including the user's geographical location information, service accessibility, and traffic conditions, etc. By analyzing these environmental data, output the user's environmental impact label EL (such as convenient transportation, rich service resources, etc.), recommended surrounding services (such as nearby hospitals, rehabilitation centers, etc.), and environmental optimization suggestions (such as recommending convenient transportation solutions, health facilities, etc.).

[0064] S34: (The input is the feedback data UF obtained in S7), including the user's historical service usage records and the priority ranking of different service types. By analyzing the patterns and frequencies of the user's past service usage, output the user's service preference label SL (such as preferring telemedicine, regular physical examinations, etc.), personalized service recommendations (such as preferentially recommending the service types preferred by the user), and the trend of user demand changes (predicting future demands based on historical records).

[0065] S4: Through data fusion technology, integrate the multi-dimensional demand portrait labels constructed in the S3 stage to finally form a comprehensive and accurate user demand portrait. This process includes multiple steps such as feature-level fusion, decision-level fusion, weighted sum, and label optimization.

[0066] S41: Perform feature-level fusion on the labels of each dimension (health HL, behavior WL, environment EL, and service preference SL) generated in the S3 stage. By splicing these labels into a high-dimensional feature vector, it can comprehensively express the various needs and health status of the user. Health labels (such as high blood pressure, high blood sugar, etc.) will be combined with behavior labels (such as a healthy lifestyle person, a sports enthusiast) and environmental labels (such as convenient transportation, rich service resources) to generate a comprehensive user portrait feature vector.

[0067] (Here in S41, no feature extraction is required, only the user portrait labels in S3 need to be spliced)

[0068] S42: On the basis of feature-level fusion, adopt a decision-level fusion method to further integrate the results of multiple prediction models. First, for each dimension (health, behavior, environment, service preference), train multiple sub-models for dimension-level prediction.

[0069] For all sub - models, the logistic regression method is used. The input is the user profile tags of S3 (Health HL, Behavior WL, Environment EL, and Service Preference SL), and the outputs include: Health Model Prediction HP: Predict the user's health status (such as diabetes, hypertension, etc.). Behavior Model Prediction BP: Predict the user's behavior pattern (such as whether they like sports, eating habits, etc.). Environment Prediction Model EP: According to the user's geographical information, predict their environmental needs (such as whether there is convenient medical service, transportation, etc.). Service Preference Prediction Model SP: Predict the type of medical service the user prefers (such as telemedicine, physical examination, etc.).

[0070] Then, based on the prediction results of different models, a comprehensive decision - making is carried out through an improved reinforcement learning fusion method. The result of the comprehensive decision - making is FD.

[0071] After combining the adaptive decision - making strategy with the attention mechanism and adding it to the reinforcement learning, the improvement steps are as follows: Train the agent under the meta - learning framework, and at the same time introduce the attention mechanism to weight the key features in each task stage. And dynamically adjust the update strategy of the Q - value function for different tasks and environmental states, and use the attention mechanism to weight the contributions of different features. Design a reward function that includes attention weights and task adaptability, so that the agent not only pays attention to the prediction error but also rewards those behaviors that can effectively identify important features. After combination, the Q - value update process of the agent is expressed as:

[0072]

[0073] Among them, is the current Q - value (i.e., the expected return) of executing action a in state s; R(s,a) is the reward function, representing the immediate reward obtained by the agent after executing action a in state s; θ is the parameter of the policy function in reinforcement learning. By optimizing these parameters, the agent can improve its decision - making ability. The reward function takes into account the importance of the current task and whether the key features are correctly identified; γ is the discount factor, representing the discount degree of future rewards, and its value is between [0,1]. It measures the degree of the agent's emphasis on future rewards;

[0074]

[0075] is the maximum Q - value of all possible actions a' in the next state s', representing the maximum expected return that the agent may obtain in the next state; α is the learning rate, controlling the step size of each update. It determines the degree of the agent's emphasis on the current experience.

[0076] Q - value update after combining the attention mechanism:

[0077]

[0078] Among them, Q final (s,a) is the final Q value, which is obtained by weighting the Q values ​​of multiple sub-models; Q i (s,a) The Q value of the i-th sub-model, which indicates the expected return of performing action a in state s. Each sub-model corresponds to a different dimension prediction (health, behavior, environment, service preference); α is the attention weight, which indicates the contribution of each sub-model Q value in the final decision. The attention mechanism dynamically adjusts these weights by learning the key features in different task stages; n: The number of sub-models, usually equal to the number of task dimensions.

[0079] The final decision strategy selects the optimal action by maximizing the weighted Q value:

[0080]

[0081] Among them, a* is the optimal action chosen by the agent;

[0082]

[0083] represents the sum of weighted Q values, which represents the expected return after weighting the Q values ​​of all sub-models in state s.

[0084] By combining adaptive decision-making strategies and attention mechanisms, the performance of reinforcement learning models is significantly improved. The adaptive decision-making strategy enables the agent to flexibly respond to different task requirements through meta-learning, while the attention mechanism helps the agent focus on key input features to avoid information overload and redundancy.

[0085] S5: Input the label data of the user demand image obtained in the S3 stage (health HL, behavior WL, environment EL and service preference SL), and use the deep learning model to support the construction and data processing of multi-dimensional demand portraits to more efficiently mine the complex features and correlations in user data.

[0086] S51: Input the labeled data obtained in the S3 stage, including the structured data in health HL, behavior WL, environment EL, and service preference SL. Through an improved graph neural network (GNN), this application models the complex non-Euclidean relationships among various user features, thereby enhancing the ability to understand and predict user needs. Then, features such as user health, behavior, and environment are modeled as a graph structure, where each feature (such as disease type, health indicator, behavior habit, etc.) is a node in the graph, and the relationships between them (such as the association between diseases and environmental factors, and the correlation between health data and behavior patterns) are the edges in the graph. Next, the improved graph neural network (GNN) is used to learn the complex non-Euclidean relationships among these nodes, especially the association between user health data and environmental factors. Finally, the graph structure relationship UserGraph_Structure between user health data and behavior data constructed based on the GNN is output, enhancing the accuracy and personalization of the demand portrait. The specific process is as Figure 7 shown.

[0087] 1) Introduce the graph attention mechanism to enhance the model's attention to important features. First, calculate the attention coefficients between nodes for weighted aggregation of the feature information of neighbor nodes. Then, calculate the attention weight e vu :

[0088] e vu = LeakyReLU(a T [Wh v ||Wh u )

[0089] where hv and hu are the feature vectors of nodes v and u, W is the weight matrix, a is the weight vector for calculating attention, || represents the feature concatenation operation, T represents the transpose, and LeakyReLU represents the leaky rectified linear function.

[0090] Normalize the attention coefficient α vu through the following formula:

[0091]

[0092] where, is the set of neighbor nodes of node v, representing all nodes directly connected to node v; exp(e vu' ) applies the exponential function to amplify the attention coefficient, increasing the influence of neighbor nodes with larger weights;

[0093]

[0094] represents normalizing the attention coefficients of all neighbor nodes to ensure the sum is 1.

[0095] Finally, perform feature update:

[0096]

[0097] Among them, h' v represents the feature representation of the updated node v; N(v) represents the set of neighbor nodes of node v; σ is the activation function, using ReLU or LeakyReLU; α vu represents the attention coefficient between node v and node u, which is used to weighted-aggregate the feature h of neighbor node u u .

[0098] 2) Combine with LSTM to facilitate the simultaneous processing of complex relationships in the graph structure and dynamic changes in the time series. First, use the Graph Attention Network (GAT) to extract the features of the graph structure data and update the representation of each node. Then, use the node representation extracted by GAT as the input of LSTM to capture the dynamic changes in the time series data. Finally, predict the user requirements through the output of LSTM.

[0099] Use the node representation hv′ obtained by GAT as the input sequence of LSTM:

[0100]

[0101] Among them, is the node representation updated by the graph neural network at time step t; {h'1, h'2, …, h' n} is the feature representation of all nodes 1, 2, …, n at time step t.

[0102] LSTM dynamically models these feature vectors:

[0103]

[0104] Among them, is the input representation of LSTM, which is usually obtained by further normalizing the node features updated by the graph neural network (GAT).

[0105] Predict the user requirements based on the output of LSTM.

[0106]

[0107] Among them, is the prediction result at time step t, W out and b out are the output weights and biases for prediction. This combination method can not only handle the non-Euclidean structure of the graph data but also capture the temporality in the data, thus improving the performance of the model in tasks such as user requirement prediction.

[0108] S52: Input the unstructured data {UMns, URns, UBns, UEns} preprocessed in the S2 stage, including medical images uploaded by users (such as skin disease photos, CT scan images) or voice records of users (such as voice call records for health consultations). First is the encoder part. For image data, a convolutional neural network (CNN) is used to extract the features of the image. The CNN automatically learns and extracts discriminative local features (such as lesion areas, textures, colors of skin diseases, etc.) from the image through multiple convolutional layers and pooling layers. For voice data, a recurrent neural network (RNN) is used to process the temporal information and extract the temporal features in the voice signal. Common RNN variants such as LSTM (Long Short-Term Memory network) can capture long-term dependencies in the voice (such as emotional fluctuations, health problem prompt signals, etc.). The processed data (image or voice) is mapped to a low-dimensional latent space through the encoder part, generating a compact latent vector representing the core features of the input data. Then is the decoder part. For image data, a deconvolution layer is used to gradually restore the size of the image. The deconvolution layer performs an upsampling operation to restore the low-dimensional feature vector to the original image size and tries to retain the structural information of the image. For voice data, an RNN (such as LSTM) is used to restore the latent vector to the temporal structure of the voice signal, thereby reconstructing the spectrum or waveform of the original voice. The goal of the decoder is to reconstruct an output as close as possible to the original input data, that is, to minimize the error between the input data and the reconstructed data. The autoencoder is trained using self-supervised learning methods, that is, in the absence of labeled data, the parameters of the model are optimized by minimizing the reconstruction error. During the training process, the system continuously adjusts the parameters of the encoder and decoder until the difference between the reconstructed data and the original input data is minimized, and the mean squared error (MSE) or other loss functions are used for optimization. Output UserGraph_NStructure, the process is as Figure 8 shown.

[0109] S53: Input the time series data {URts, UBts} preprocessed in the S2 stage, including the user's health monitoring data (such as blood pressure, blood sugar, body temperature, etc.) and behavior data (such as exercise, sleep, etc.). First, use historical data to train the LSTM model, and update the weights of the LSTM model by minimizing the loss function (such as mean square error). During the training process, the LSTM will learn the correlation and long-term trends between the health monitoring data and behavior data in the input sequence. After training, capture the long-term dependencies in these time series data through the LSTM, analyze the past health and behavior data, and predict the future health status and behavior patterns. For example, based on past blood sugar data, the LSTM can predict the change trend of future blood sugar levels to help identify potential health risks; at the same time, the LSTM can also identify the long-term trends in exercise and sleep data and provide personalized behavior suggestions. Through these predictions, the system can provide accurate health management suggestions for users, optimize personalized services, and improve the effect of health management. The process is as Figure 9 shown.

[0110] S6: Input the user demand portrait processed in the S4 stage and the complex correlation obtained in S5, and combine the intelligent matching algorithm and the real-time recommendation process to achieve dynamic service recommendation and optimization.

[0111] S61: When the user demand portrait is updated, the system automatically triggers the service recommendation process. If the user's health status changes (such as an increase in blood sugar, abnormal exercise, etc.), the system re-evaluates the user's needs based on the new health data and activates the recommendation process. Improve the A* algorithm,

[0112] 1) Introduce Qos constraints, that is, in the original cost function, consider the quality of the service (such as response time, reliability, bandwidth, etc.) to adjust the search strategy, so that the finally found path is not only optimal in terms of cost, but also meets the QoS constraint conditions. Improvement steps: Each node represents a service instance, and the state is represented by the QoS characteristics of the service (such as response time, cost, bandwidth, reliability, etc.). The edge represents the conversion or dependency relationship between services, and the cost is the switching cost or the relationship between service selections. The cost function is:

[0113]

[0114] where g(n) is the cost of the current path, representing the actual cost from the starting point to the current node; h(n) is the heuristic function, representing the expected cost from the current node to the target node (usually the estimated distance or time); QoS i (n) is the i-th QoS characteristic of node n (such as response time, bandwidth, availability, etc.), and each QoS characteristic will affect the total cost; w iis the weight corresponding to each QoS feature, indicating the contribution degree of the QoS feature to the cost function; m is the number of QoS features considered.

[0115] 2) The heuristic function h(n) is static in the traditional A* algorithm, but in the improved version of this application, h(n) should take into account the dynamic changes of service QoS features. Therefore, this application designs a dynamically adjusted heuristic function:

[0116]

[0117] Among them, estimated.cost(n) is the expected cost from the current node to the target node estimated based on the heuristic method; α and β are weights that control the relative importance of the heuristic method estimation and QoS constraints.

[0118] 3) Reinforcement learning introduces an adaptive decision-making strategy. When QoS parameters change, reinforcement learning can dynamically adjust the service selection strategy to achieve optimal service recommendation. A state-action pair and a reward mechanism are introduced, and the Q-learning algorithm is used to adjust the service selection strategy. At each time step t, the Q value is adjusted through environmental feedback (user feedback on the service, QoS satisfaction):

[0119]

[0120] Among them, St is the current state, At is the currently selected service, Rt is the immediate reward, and γ is the discount factor.

[0121] 4) Combining reinforcement learning with the A* algorithm can combine global optimization and local optimization. The A* algorithm is used for global planning and searching for the optimal path, while reinforcement learning performs dynamic adjustment at the local level, enabling the system to make more intelligent and faster decisions during real-time service recommendation.

[0122] The service recommendation process is as Figure 10 shown.

[0123] S62: During the service recommendation process, the system filters service resources that meet the QoS constraints from the database. The filtering criteria include not only service quality (such as response time, success rate, treatment effect, etc.), but also resource limitations (such as service capacity, device availability). For example, if a user needs to conduct remote medical consultation and the doctors on a certain platform have been fully booked, the system will dynamically adjust according to the available resources and recommend other platforms or doctors. In this way, users can obtain timely medical services, avoid delays caused by resource limitations, and ensure the continuity and availability of services.

[0124] S63: When planning the service chain, the improved A* algorithm dynamically adjusts the recommended path according to the user's geographical location, service timeliness, and demand priority. This algorithm comprehensively considers multiple factors, such as distance, service availability, traffic conditions, etc., to optimize the order and process of services. For example, if a user lives in an area with inconvenient transportation, the system will preferentially select a medical or rehabilitation center that is closer and has higher timeliness, reducing the user's waiting time and travel burden. If the user's health condition requires urgent intervention, the system will ensure that the nearest available medical facility is reached in the shortest time, thereby improving service efficiency and user experience.

[0125] S64: After optimizing the service path, the system selects suitable service types such as medical, rehabilitation, and elderly care according to the user profile and optimization results, and generates a personalized service chain. For example, for users in the recovery period, the system will recommend targeted rehabilitation training services; for users with chronic disease management needs, the system will recommend regular health monitoring, remote medical consultations, and other services. The service recommendation plan will generate a service chain based on the optimized path and output a personalized recommendation plan, covering detailed information, implementation time, priority order, and execution method of each service step. Through this process, users can not only obtain services that meet their health needs but also enjoy a personalized and precise service experience during actual execution.

[0126] S7: Feedback optimization mechanism. At this stage, the system continuously improves the accuracy and effectiveness of service recommendations and dynamically updates the user demand profile tags through the collection of feedback data and model optimization. The specific process is as Figure 11 shown.

[0127] S71: Collect the feedback data, including objective data and subjective data. Among them, objective data includes service completion status, changes in user health indicators, etc. These data can accurately reflect the effect of service execution. For example, health monitoring data (such as blood glucose level, weight change, exercise data, etc.) can help evaluate the effect of recommended services and determine whether the expected health improvement goals have been achieved. Subjective data includes user satisfaction scores, complaint information, etc. These data help understand the user's subjective feelings about the service and provide real-time feedback on service quality. For example, if a user is dissatisfied with a certain medical service (such as remote consultation, health check), the system will collect the user's specific complaints and feedback and further analyze the advantages and disadvantages of the service. Through the collection of these two types of data, the system comprehensively understands the actual effect of the service and provides basic data for subsequent model optimization.

[0128] S72: Through the collected feedback data, the system dynamically optimizes the recommendation model using a reinforcement learning framework (DQN). The update of the model follows a reward mechanism and a punishment mechanism. The reward mechanism is that when the user is satisfied with the recommended service or the service completion is good, the model will give a higher reward score to the recommendation strategy. This reward mechanism encourages the system to continuously recommend service combinations with better effects. For example, if a specific service path (such as the telemedicine service of a certain hospital) can improve the user's health indicators and is highly evaluated by the user, then this service path will be regarded as a preferred option, and similar paths will be preferentially recommended in the future. The punishment mechanism is that if the user's satisfaction is low or the service completion rate is low (for example, a certain health management service fails to be completed as planned, or the user is dissatisfied with the response of the telemedicine), then the system will punish the corresponding service combination or path, reducing its priority in future recommendations. For example, if a service provider has a slow service response and poor treatment effect, resulting in user dissatisfaction, then the priority of this service provider will be reduced, and the system will preferentially recommend other service providers with better performance. Through the reward and punishment mechanisms, the system continuously learns from user feedback and gradually adjusts the recommendation strategy to improve service quality and user satisfaction.

[0129] S73: The demand portrait is dynamically updated. By taking the feedback data (including objective health data and subjective satisfaction data) as incremental input, the system will regularly update and optimize the user's demand portrait tags. This update can help the system better understand the user's health needs, service preferences, and satisfaction with services. For example, if the user feedbacks that the problem of high blood sugar has been gradually controlled after using the health management service for a period of time, then the system will adjust its health portrait, update the disease management demand tags, and accordingly recommend more efficient diabetes management services. As the user's demand portrait is continuously updated, the recommendation system can become more accurate and personalized, thereby improving the overall recommendation quality. By dynamically updating the demand portrait, the system can better respond to the user's health changes and feedback, and provide continuously optimized service recommendations.

[0130] S8: In the service chain management and execution stage, ensure the standardized execution, efficient monitoring of the service chain, and cross-institutional resource coordination, and improve the accuracy and efficiency of service execution. Specifically as Figure 12 shown.

[0131] S81: First of all, the system needs to establish a unified service process framework to ensure that each recommended service chain can be executed seamlessly. To this end, the system will define each service node in the service chain (such as medical services, rehabilitation services, elderly care services, etc.), and clarify the input and output requirements and time requirements of each node. For example, in the medical node, the input data may include health examination data (such as blood sugar, blood pressure, etc.), and the output is the diagnosis and treatment plan and drug advice, and it is required to complete the preliminary diagnosis and treatment within the specified time; in the rehabilitation node, the input data is the user's health assessment result, and the output is a personalized rehabilitation training plan, which requires the training plan to be completed within the specified time. Through such a standardized process framework, it can be ensured that the service execution of each node is carried out in accordance with unified specifications, avoiding service execution deviations.

[0132] S82: The system needs to monitor the execution of the service chain in real time, focusing on monitoring key indicators such as task completion time, user waiting time, and service resource utilization efficiency. The task completion time is used to ensure that each service node completes the task according to the specified time, and the user waiting time is used to monitor the time from when the user initiates a request to the actual start of the service, ensuring that the waiting time is within an acceptable range. At the same time, the service resource utilization efficiency will evaluate the allocation and use of resources such as doctors, beds, and equipment to ensure that resources are optimally configured. If an abnormal execution of a certain node is found during the service execution (such as service timeout, task not completed, etc.), the system will start an emergency plan, automatically switch to backup service resources, adjust the service execution order, or recommend other available resources. For example, if a certain hospital cannot provide remote medical services due to equipment failure, the system can automatically switch to other hospitals to ensure the continuity and timeliness of the service.

[0133] S83: At this stage, the system needs to build a service collaboration platform based on the microservices architecture to support data sharing and service interaction among medical, rehabilitation, and elderly care institutions. The platform can ensure the data flow between different institutions, avoiding information silos, thereby improving the collaborative efficiency of resources. By dynamically allocating resources, the system will flexibly adjust the resource allocation of each node according to the actual situation of the service chain execution to ensure that the service chain can be executed efficiently. For example, if the bed resources of a certain rehabilitation center are in short supply, the system will automatically transfer the user to other available rehabilitation institutions nearby according to the user's health condition and geographical location, thus avoiding service delays caused by insufficient resources and ensuring the efficiency and timeliness of the service.

[0134] Other optional embodiments:

[0135] 1. During the process of constructing the demand portrait, the LLMs model can be used to generate user portraits. By generating user portraits through LLMs, a more comprehensive and detailed understanding of users can be provided for the service recommendation system, not only limited to basic health data and behavior habits, but also able to deeply understand users' emotional backgrounds, life challenges, and goals. This method enhances the personalization ability of the recommendation system, helps improve the accuracy of service recommendations and user satisfaction, and at the same time promotes the continuous dynamic optimization of user needs. [6]

[0136] 2. In cross-institutional resource collaboration and execution, the method of using a cloud service-based collaboration platform can be adopted, and cloud computing platforms (such as AWS, Azure, etc.) are used to provide a unified data storage and service scheduling platform. In cross-institutional resource collaboration, the cloud platform can provide a unified resource pool and scheduling mechanism to ensure that data of each service node can be shared, and service execution can be quickly deployed and coordinated.

[0137] 3. In service chain management and execution, each task node in the service chain can be managed and scheduled by using a mature workflow management system (such as Apache Airflow or Camunda). The workflow engine can assign tasks to each service node, monitor the service status, and make dynamic adjustments according to the execution situation. In this way, the system can flexibly manage cross-service workflows and data transfer to ensure the smooth completion of tasks.

[0138] The embodiment of this application also provides an electronic device, including: a processor, and a memory coupled to the processor, the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in any one of the above embodiments.

[0139] The electronic device can be a computing device such as a desktop computer, a notebook, a handheld computer, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory.

[0140] The so-called processor may be a Central Processing Unit (CPU), or it may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device and connects various parts of the entire device using various interfaces and circuits.

[0141] The memory can be used to store the computer program. The processor realizes various functions of the electronic device by running or executing the computer program stored in the memory and calling the data stored in the memory.

[0142] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory and may also include non-volatile memory, such as hard disks, memory, plug-in hard disks, Smart Media Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0143] The embodiments of the present application also provide a storage medium. The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, Read-Only Memory (ROM), Random Access Memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0144] The above are the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.

Claims

1. A training method for a medical and health care service recommendation agent, characterized in that: The training method comprises: Obtaining training data, the training data including: medical care and health care services, and at least two user tags related to the medical care and health care services; and The training data is used to train an intelligent agent, and the intelligent agent is specifically used to: construct a user demand portrait using the at least two user tags; calculate the correlation between the user tags; and, based on the user demand portrait and the correlation, make medical and health service predictions.

2. The training method according to claim 1, characterized in that: Using the at least two user tags to construct a user demand profile specifically includes: Performing feature fusion on the at least two user tags to obtain a high-dimensional feature vector; Based on the user tags, obtaining at least two demand forecasts; and The user demand portrait is obtained by using a reinforcement learning method with the high-dimensional feature vector and the demand forecast as input.

3. The training method according to claim 2, characterized in that: The reinforcement learning method includes an attention mechanism and a task-adaptive reward mechanism.

4. The training method according to claim 1, characterized in that: Calculating the correlation between the user tags specifically includes: extracting structured data from the user tags; and The structured data is processed using a graph neural network to obtain the degree of association.

5. The training method according to claim 4, characterized in that: An attention mechanism is added to the graph neural network processing and combined with a long short-term memory network. The long short-term memory network is used to complete the dynamic expression of the association degree based on the unstructured data and time series data in the user tags.

6. The training method according to claim 5, characterized in that: Based on the user demand portrait and the correlation, medical care and health care service prediction is performed, specifically: a reinforcement learning method including an attention mechanism and a task adaptive reward mechanism combined with an A* algorithm is adopted to obtain medical care and health care prediction services from the user demand portrait and the correlation.

7. The training method according to claim 1, characterized in that: The medical and health care services are subject to QoS constraints; the user tags are health tags, behavior tags, environment tags, service preference tags or satisfaction tags; the user tags are obtained through user data collection, data cleaning, noise reduction, normalization and analysis.

8. A method for recommending medical and health services, characterized in that: The recommended methods include: Get the current user tag; and Utilize the intelligent agent trained as claimed in any one of claims 1 to 7 to obtain the currently recommended current medical care and health prediction service.

9. An electronic device, characterized in that: The electronic device comprises: a processor, and a memory coupled to the processor, The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory, so that the electronic device performs the training method as described in any one of claims 1 to 7, or performs the recommendation method as described in claim 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a computer program or instructions, and when the computer program or instructions are executed on a computer, the computer executes the training method according to any one of claims 1 to 7, or executes the recommendation method according to claim 8.

Citation Information

Patent Citations

  • Recommendation system based on dynamic attention and hierarchical reinforcement learning

    CN112597392A

  • Multimedia resource recommendation method and device, electronic equipment and storage medium

    CN113190757A

  • Training method of graph neural network model for recommending Web service combination

    CN114065033A

  • User personalized recommendation method based on multi-modal data

    CN114647787A

  • User portrait-based product recommendation method, apparatus and device, and storage medium

    CN114663198A

Cited By

  • Health monitoring management method based on big data

    CN121075668A