Online medical recommendation system and implementation method thereof
By using an improved PSVD method and knowledge graph technology, a patient-medical item preference matrix is constructed, marginal data is removed, and semantic information from the knowledge graph is combined to solve the problems of single recommendation results and data sparsity in existing medical information systems, thus achieving efficient and accurate medical information recommendation.
Patent Information
- Application Number
- CN202210132352.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-14
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-02-14
AI Technical Summary
Existing medical information systems face challenges in data storage, information utilization, and security. They are unable to efficiently recommend appropriate medical plans and information to patients, and the recommendations are limited, data is sparse, and the cold start problem is serious.
By employing an improved PSVD method and knowledge graph technology, we construct a patient-medical item preference matrix, remove marginal data, optimize the recommendation model using semantic information from the knowledge graph, and improve recommendation accuracy by combining patient interest attributes and positive and negative correlation entity models.
This enables more accurate and efficient recommendations of appropriate medical plans and information to patients, improves system security and information utilization efficiency, and reduces recommendation errors.
Smart Images

Figure CN114547444B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to an online medical recommendation system and method, in particular to an online medical recommendation system based on an improved knowledge graph and PSVD method and an implementation method thereof. BACKGROUND
[0002] At present, medical personnel in the field of medical and health information still face many problems and challenges in the process of collecting and processing massive data. First, the traditional hospital information storage method cannot meet the huge medical information storage demand. Second, the lack of system information, excessive information, information conflict, information dispersion and information error lead to the failure to fully exploit the potential value of hospital information data. Third, once the medical data shows explosive growth, it means that the data may face the risk of leakage, and the security and privacy of medical data are difficult to guarantee. In today's rapid development of information technology, in the massive medical and health information data, how to quickly and efficiently recommend appropriate medical solutions, doctors, rehabilitation plans and other information to patients has become a problem that needs to be solved in the medical field. SUMMARY
[0003] The present application designs a patient-medical project preference matrix for the problem of excessive medical information, and defines a matrix singular value critical parameter, which is set according to a variance model and a matrix singular value set, and removes the edge data in the matrix. In order to quickly and efficiently recommend projects to patients, a set of knowledge graph KI is used to analyze the multiple entities and attributes in KI to establish a vector triple, establish a patient interest attribute preference and a positive and negative correlation entity method, and give a recommendation model scheme. In view of the single recommendation result, data sparsity and data cold start problems existing in the existing medical recommendation method, the present application designs a medical recommendation based on a knowledge graph and an improved PSVD method. In summary, the present application provides an online medical recommendation system based on an improved knowledge graph and a PSVD method and an implementation method thereof, which is intended to recommend appropriate doctors and effective rehabilitation treatment suggestions to patients who log in to the system, while providing information publishing and questioning in the online community, information recommendation on the home page of the system, and information recommendation in the personal homepage of the logged-in user according to their interest preferences and other functions.
[0004] The present application provides an online medical recommendation system, which comprises a home page module, a personal information page module, a medical exchange feedback module, a disease community module and a rehabilitation therapy suggestion module. The home page module is used to aggregate and display the entrances of various sub-modules of the system, the personal information page module displays the basic information of patients, the medical exchange feedback module is used for medical information exchange of system users, the disease community module is used to display medical information of diseases, and the rehabilitation therapy suggestion module is used to display recommended information of disease rehabilitation therapy to system users.
[0005] The application firstly registers and logs in the system, inputs personal medical information, including: basic information, disease name, disease duration, treatment duration, etc. After the patient logs in the system, the patient can select: homepage module, personal information page module, medical exchange feedback module, disease community module, rehabilitation physiotherapy suggestion module to enter. The construction of the recommendation model is based on the improved PSVD method and the knowledge graph method, which is specifically divided into two parts. First, the PSVD method part uses the established patient-medical item preference matrix, noise data critical value, singular value set, etc. to construct a prediction score model of the patient for the unscored medical item. Then, considering that the PSVD method only depends on the data and ignores the semantic information of the patient and the item itself, which may cause inaccurate recommendation results, therefore, the application introduces the knowledge graph method to optimize the PSVD method result, constructs a knowledge graph KI according to the data set, and sets the vector triplets of (entity, attribute, relationship) in KI. Based on the triplets, the KI patient interest preference attribute level model and the medical item positive and negative related entity model are constructed respectively. Finally, the recommendation model is constructed by fusing the PSVD method constructed above, and the final recommendation result list set SET is obtained. In summary, the application uses an improved PSVD method, and uses the knowledge graph technology to optimize the prediction score of the patient for the medical item obtained by the PSVD method. Finally, the improved knowledge graph "patient interest attribute model" and "positive and negative related entity model" and the PSVD method prediction score model are fused to give the final medical recommendation model.
[0006] The application also provides an implementation method of an online medical recommendation system, which comprises the following steps:
[0007] Step 1, based on the data set patient-Medical-Data, define the preference matrix R of the patient patient for the medical item Medical, and define as follows:
[0008]
[0009] In the formula, a plurality of patients and medical items in the system are represented by two multidimensional vectors respectively, P=[p1,p2,p3,...p u ,...p p ] and M=[M1,M2,M3,...M i ,...M m ], The meaning of the representative is the preference score of the patient u for the medical item i;
[0010] Step 2, according to the idea of the SVD method, considering that the data amount in the preference matrix R is large and is mixed with a lot of edge data, in order to reduce the time complexity of the recommendation method and reduce unnecessary method calculation, an improved method is considered to remove the redundant data in the preference matrix R to obtain a new preference matrix R′ ;
[0011] Step 3, after the above series of steps, a new preference matrix R is obtained ′ , the parameter TOP-K defines the dimension of the matrix R ′ , compared with the initial preference matrix in step 1, the data dimension in the matrix R ′ is reduced, and the edge data is reduced, reducing the existence of noise data, which is beneficial to improve the accuracy of the recommendation result. In this step, according to the new preference matrix R ′ , the predicted preference score of the patient for the un-scored medical project is calculated;
[0012] Step 4, the predicted score obtained by the PSVD method is optimized and improved, since the PSVD only considers the data level and ignores the semantic attributes of the patient and the medical project entity, the medical project entity is first converted into a triple vector based on the TransH knowledge representation method in the application, and the vector representation of the entity attribute triple is obtained, that is
[0013] T u ={(h,r,a)|h∈History u}
[0014] Wherein, History u represents a set of historical medical project entities associated with the patient, T u represents the triple information contained in the historical medical project entity set of the patient u, h represents an entity in the recommendation system, a represents an attribute of the entity, and r represents a relationship between the entity and the attribute.
[0015] Step 5, using the a attribute in the vector triple, the method is used to obtain the weight of a certain medical entity attribute in all medical project entity attributes related to the patient, and the weight is used to judge the interest preference level of the patient to a certain entity attribute;
[0016] Step 6, based on the established medical project attribute a weight model and the vector triple in the knowledge graph, a level judgment model of the patient's interest preference for the medical project is constructed;
[0017] Step 7, based on the established knowledge vector triple, a positive and negative correlation medical project entity model is constructed;
[0018] Step 8, the fusion recommendation algorithm is obtained by fusing the improved PSVD prediction score model, the patient interest attribute preference level model and the positive and negative correlation entity model;
[0019] Step 9, return the final recommendation result set SET.
[0020] The recommended method of the application improves the traditional SVD method, and proposes a new PSVD method. In the PSVD method, a new method for removing the edge data of the preference matrix is proposed. The method is based on the established patient-item preference matrix R, and divides the medical data in the data set into two categories: patient data and medical item data, which are represented in the form of multidimensional vectors P and M: P=[P1, P2, P3,...P p ], M=[M1, M2, M3,...M m ], and then sets the edge data threshold E, which is defined according to the singular values in R, and forms a vector Q=[Q1, Q2, Q3,...Q m ]. The specific representation of E is as follows. According to E, the edge data in the preference matrix R can be removed, so as to reduce the distribution of noise data in the preference matrix R. At the same time, the least square method is introduced to reduce the error of the preference score of the patient to the medical item. Specifically, the patient factor vector p u , the medical item factor vector I i , two score deviation values δ u , δ i , and a rating deviation parameter μ are defined. The patient factor and the item factor deviation values in the predicted score preference score are represented as δ u,i , respectively. SS ′ is a preliminary defined prediction preference score method. After removing the edge data in the initial preference matrix, the new obtained preference matrix R p*k is decomposed into a patient factor matrix and a medical item factor matrix, which are denoted as U u*kThis invention analyzes patients' preferences for k factors and the presence rate of these potential factors in a specific medical item, combining this with the correlation between patients and an error reduction model to arrive at a final prediction score. Furthermore, a knowledge graph approach is incorporated to optimize the prediction score model obtained from the PSVD method. Since the PSVD algorithm relies solely on the "patient-medical item" preference matrix for recommendation, it depends only on data and ignores the semantic information of patients and medical items themselves. Using only the improved PSVD algorithm for medical recommendations would result in low accuracy. Therefore, this invention, based on the aforementioned steps, introduces knowledge graph technology, adding semantic association information between medical items and patients, item similarity, and other attributes from the knowledge graph to the recommendation model. This highly utilizes the implicit information of medical items and patients in the recommendation system, thereby improving the accuracy of recommendations. Finally, combining the semantic attributes of patients and medical projects, the project entities are transformed into triple vectorization using knowledge representation methods. Based on the established vector triples, the 'a' attribute in the vector triples is used to define a method to obtain the weight of a certain medical entity attribute among all medical project entity attributes related to the patient. Based on this weight, the patient's interest preference level for a certain entity attribute is judged.
[0021] The further optimized technical solution of the present invention is shown below:
[0022] The specific steps for step 2 are as follows:
[0023] Step 2.1: Based on the established matrix R, set a critical value E. The critical value determination depends on the one-dimensional vector dataset formed by the data on the subdiagonal of matrix R, defined as Q = [Q1, Q2, Q3, ..., Q...]. i ,...Q m Based on the concept of SVD, the data on the diagonal of the preference matrix are singular values. Removing data based on these singular values effectively ensures the accuracy of removing marginal data. The definition of the critical value E is as follows:
[0024]
[0025] In the formula, Let m be the mean of the vector dataset Q, and m be the vector dimension of Q;
[0026] Step 2.2: Based on E in Step 2.1, determine the values in matrix R. Relationship with E:
[0027] like In matrix R, retain The data in the row and column;
[0028] like In matrix R, delete data of the rows and columns;
[0029] Finally, a new preference matrix R is obtained ′ .
[0030] The specific operation of step 3 is as follows:
[0031] Step 3.1, decompose the new preference matrix R into a patient factor matrix and a medical project factor matrix, and denote them as U and V respectively ′ p*k , I u*k , where k represents the latent factor attribute of the patient's preference for the medical project, and by analyzing the patient's preference degree for k factors and the existence rate of these latent factors in a medical project, PSVD can predict the user's preference score for the corresponding item. The predicted preference score of patient u for medical project i is defined as follows:
[0032] SS u,i ′ = p u * I i + μ * (δ u + δ i )
[0033] In the formula, p u , I i represent the single patient factor vector and the single medical project factor vector respectively, δ u , δ i represent the rating deviation value when predicting the preference score of patient u for medical project i, the former represents the patient factor deviation value, and the latter represents the medical project factor deviation value; μ represents the set rating deviation parameter;
[0034] Step 3.2, considering that there is an error in the predicted score of the patient for the medical project, therefore, the least squares method is introduced to reduce the deviation rate, which is specifically represented as follows:
[0035]
[0036] In the formula, u, i ∈ SS represents a certain user u and medical project i in the selected data set, SS u,i , SS u,i ′ respectively represent the predicted preference score of patient u for medical project i before and after the dimension reduction of the preference matrix R, and γ represents the set threshold parameter;
[0037] Step 3.3, after processing the deviation error of the predicted score, the error of the predicted score can be reduced, and the target user is scored based on the historical score of the similar user, and the final predicted score is defined as follows:
[0038]
[0039] where the first addend represents the correlation between patient u and patient v, here the pcc theorem is cited to measure the similarity between roles, SS u,i , SS v,i represent the score of patient i on medical project i, respectively the average score of patient u, v on medical project, N represents the total number of medical projects in the system. Thus the prediction preference score based on the SVD method of multiple patient correlation is obtained.
[0040] Since the PSVD algorithm only depends on the "patient-medical project" preference matrix to realize recommendation, only data is relied on and the semantic information of patients and medical projects themselves is ignored, if only the improved PSVD algorithm is used to realize medical recommendation, the accuracy of the recommended result is not high. Therefore, based on the foregoing steps, the knowledge graph technology is introduced, the semantic correlation information between the medical projects and patients in the knowledge graph, the medical project similarity and the like are added to the recommendation model, the implicit information of the medical projects and patients in the recommendation system is highly utilized, and thus the accuracy of the recommendation is improved.
[0041] The weight of the a attribute is obtained in the step 5, and the specific method is as follows:
[0042] Based on the triple (h, r, a) in the knowledge graph, the weight of the a attribute in the medical project attribute related to patient u is defined, and the specific method is as follows:
[0043]
[0044] wherein, represents the weight value of the a attribute in all medical project entity attributes of patient u, h a T represents the previous entity associated with the a attribute, r a represents the relationship vector associated with the a attribute, (h, r, a) belongs to T u represents that the triple containing the a attribute in the knowledge graph is selected, h and r respectively represent the previous entity and the relationship vector in all triples containing the a attribute.
[0045] In the step 6, the grade determination model is constructed and includes the following steps:
[0046] Step 6.1, the weight of the entity attribute is obtained Then, the interest preference grade of patient u on the attribute a is determined based on this, and the specific method is as follows:
[0047]
[0048] wherein, Interestu->a represents the interest preference attribute value of the patient u to the entity attribute a, and the weight of the attribute a is represented as
[0049] Step 6.2, three level parameters are set, and the set level parameters are compared with the interest preference attribute value of the patient to the entity attribute obtained in step 6.1 to obtain the interest preference level;
[0050] Step 6.3, a first level parameter γ1 is set, and according to the vector triplets in the knowledge graph KI, only the front entity vector h a and the relationship vector r a of the entity attribute a are taken out to perform vector dot product operation. Since it only depends on the entity attribute a, it can be explained that the current attribute a matches the lowest degree of patient interest preference, and therefore the parameter represents that the patient is least interested, and is specifically represented as follows:
[0051]
[0052] Step 6.4, a second level parameter γ2 is set, and according to the vector triplets in the knowledge graph KI, as long as the triplet contains the attribute a, the front entity h and the relationship vector r are taken out and vector dot product operation is performed. Setting this parameter indicates that the current attribute a matches the medium degree of patient interest preference, and therefore the parameter represents that the patient is interested. Specifically represented as follows:
[0053]
[0054] Step 6.5, a third level parameter γ3 is set, which represents that the current attribute a matches the highest degree of patient interest preference, and therefore the parameter represents that the patient is very interested. Specifically represented as follows:
[0055]
[0056] Step 6.6, the preference level of the patient to the medical project attribute a is determined:
[0057] If Interest u->a < γ1, it is considered that the patient u is not interested in the attribute a;
[0058] If γ1 < Interest u->a < γ2, it is considered that the patient u is interested in the attribute a;
[0059] If Interest u->a > γ3, it is considered that the patient u is very interested in the attribute a, and thus the interest preference level model of the patient to the attribute a is obtained based on the knowledge graph KI.
[0060] The subject involved in the formulation of the recommendation scheme is the project entity, and for the project entity in the knowledge graph, it is classified as a positive correlation entity and a negative correlation entity, and in the final formation of the recommendation, the recommendation scheme list with the negative correlation entity as the core and the sub-entity attributes related to the negative correlation entity are automatically removed, thereby reducing the complexity of the method and unnecessary calculation.
[0061] In step 7, the subject involved in the formulation of the recommendation scheme is the project entity, and for the project entity in the knowledge graph, it is classified as a positive correlation entity and a negative correlation entity, and in the final formation of the recommendation, the recommendation scheme list with the negative correlation entity as the core and the sub-entity attributes related to the negative correlation entity are automatically removed, thereby reducing the complexity of the method and unnecessary calculation. The positive and negative correlation medical project entity model comprises the following steps:
[0062] Step 7.1, defining a random entity in the knowledge graph KI: for the triple formed by a plurality of entities in the knowledge graph KI, a single triple (h, r, t), wherein h represents a front entity, t represents a rear entity, and the front entity h is replaced with any random entity h in the knowledge graph KI random The rear entity is also replaced as t random The newly generated triples (h random , r, t), (h, r, t random ) do not exist in the knowledge graph KI;
[0063] Step 7.2, based on the front and rear entities h and t defined in step 7.1, define the W h Vector representation maps entities from entity space to relationship space, and sets an edge function f, and the function f obtains the square value of the vector module, which is specifically represented as:
[0064] f(h, t) = ‖h-t-(W h T ·h·W h -W h T ·t·W h ||; 2 ;
[0065] Step 7.3, establishing a judgment model of positive and negative correlation entities, based on the edge function f defined in step 7.2, designing a method model for distinguishing positive and negative correlation entities;
[0066] The positive and negative correlation entity model is specifically represented as follows:
[0067]
[0068] In the formula, class h,tThe precondition for model execution is: (h, r, t) ∈ T u and and
[0069] Step 7.4, analyze the two result values of class h,t to determine the positive and negative correlation entities.
[0070] In step 8, the fusion recommendation algorithm includes the following steps:
[0071] Step 8.1, for the two tuples (u, i) in the knowledge graph KI, if that is, the preference level of patient u for medical item i is not interested, the algorithm terminates, otherwise step 8.3 is executed;
[0072] Step 8.2, for the two tuples (u, i) in the knowledge graph KI, if
[0073]
[0074] that is, the preference level of patient u for medical item i is interested, the algorithm terminates, otherwise step 8.3 is executed;
[0075] Step 8.3, for the two tuples (u, i) in the knowledge graph KI, if f(I h ,I t ) is not satisfied f(I hrandom ,I t ) && f(I h ,I t ) > f(I h ,I trandom ), that is, the entities h, t are negative correlation entities, the algorithm terminates, otherwise step 8.4 is executed;
[0076] Step 8.4, if the number of recommended schemes in SET does not reach K, it is added to the recommended result set SET, and step 8.1 is executed, otherwise step 8.5 is executed;
[0077] Step 8.5, if the data in the SET set has reached k, output SET to form a recommended result set, and the recommendation ends; otherwise, step 8.1 is executed. BRIEF DESCRIPTION OF DRAWINGS
[0078] Figure 1 is the medical knowledge graph KI based on "diabetes" in the present application.
[0079] Figure 2 is the improved PSVD method flowchart in the present application.
[0080] Figure 3 The recommendation model flowchart in which the PSVD, the interest preference attribute and the positive and negative correlation entity are fused in the present application.
[0081] Figure 4 The architecture diagram of the system in the present application. DETAILED DESCRIPTION
[0082] The technical solutions of the present application will be further described in detail below in combination with the drawings: the present embodiment is implemented on the premise of the technical solutions of the present application, and detailed implementation manners and specific operation processes are given, but the protection scope of the present application is not limited to the following embodiments.
[0083] Embodiment 1
[0084] The present embodiment takes “diabetes” in medical treatment as an example to establish a knowledge graph (see Figure 1 ), and takes this as an example to perform the steps of the PSVD prediction score of the recommendation model and the interest preference attribute and the positive and negative correlation entity, compared with the medical recommendation method based on the traditional SVD method, the present system adds the knowledge graph method, so that the recommendation model combines the medical project and the semantic of the patient itself, not only relies on the preference score, but also is more conducive to improving the accuracy of recommendation and reducing the recommendation error.
[0085] The online medical recommendation system of the present embodiment includes a homepage module, a personal information page module, a medical exchange feedback module, a disease community module and a rehabilitation therapy suggestion module; the homepage module is used to aggregate and display the entrances of various sub-modules of the system, the personal information page module displays the basic information of the patient, the medical exchange feedback module is used for the medical information exchange of the system users, the disease community module is used to display the medical information of the disease, and the rehabilitation therapy suggestion module is used to display the recommended information of the disease rehabilitation therapy to the system users.
[0086] When using the above medical recommendation system, the specific steps are as follows: the patient registers and logs in the system, fills in the personal basic information, including: the name of the disease, the duration, the treatment record, etc.; the recommendation system module is composed of: the homepage module, the personal information page module, the medical exchange feedback module, the disease community module, the rehabilitation therapy suggestion module; the patient can enter each sub-module for viewing, among which the medical exchange feedback module can be used for the patients to exchange the medical disease conditions in the community, the disease community module can recommend the related medical research information of the current disease to the patients, including: the treatment doctor, the treatment drug, the complication caused, etc., and the rehabilitation therapy suggestion module can display some treatment suggestions, prevention measure schemes, rehabilitation treatment suggestions, etc. (see Figure 4 ) of the current disease to the patients.
[0087] As Figure 2 and Figure 3As shown, an implementation method of an online medical recommendation system, comprising the following steps:
[0088] Step 1, obtaining a medical data set. Download the relevant data set from the "Jinshu Li" website, name the data set: proDatas, define the preference matrix R of the diabetic patients Diabetics for the diabetes-related medical projects MedicalPros (diabetes treatment drugs, treatment instruments, treatment plans, preventive measures, etc.), and define as follows:
[0089]
[0090] In the formula, two types of data in the system proDatas: diabetic patients and projects are represented by two multidimensional vectors, that is,
[0091]
[0092] The meaning of the representative is the preference score of the diabetic patient u for the diabetes project i.
[0093] Step 2, reduce the data volume in the diabetic patient preference matrix, and remove the edge and noise data of the matrix.
[0094] Step 2.1, according to the established matrix R, set a critical value E, define a one-dimensional vector data set Q = [Q1, Q2, Q3,..., Q i ,...Q m ] formed by the data on the sub-diagonal line of the matrix R. According to the idea of SVD, the data on the diagonal line of the preference matrix is the singular value, and the removal can effectively guarantee the accuracy of removing the edge data. The definition of the critical value E is as follows:
[0095]
[0096] In the formula, is the mean of the vector data set Q, and m is the vector dimension of Q;
[0097] Step 2.2, according to E in step 2.1, judge the size relationship between the value in the matrix R and E:
[0098] If , in the matrix R, the data in the row and column where is located is retained;
[0099] If , in the matrix R, the data in the row and column where is located is deleted;
[0100] Finally, a new preference matrix R ′ is obtained.
[0101] Step 3, define the new matrix obtained in step 2.2 as R ′ , define the parameter TOP-K to represent the dimension of the matrix R ′ , and then calculate the predicted preference score of the diabetic patient for the un-scored diabetes medical project (such as: insulin treatment satisfaction, dependence on gastric surgery precautions, bone marrow stem cell, peripheral blood stem cell transplantation, acceptance of strict lifestyle habits, etc.) according to the new preference matrix.
[0102] Step 3.1, decompose R ′ into a diabetes, patient factor matrix including patient personal appearance information, disease duration, diabetes type, historical treatment record, etc. and a diabetes project factor matrix including related possible complications, diabetes prevention measures, and related symptoms of diabetes, etc., respectively denoted as U p*k , I u*k , where k represents the patient's preference for the underlying factor attributes of the diabetes project, such as: low-sugar diet, exercise fitness, smoking prohibition, etc. By analyzing the patient's preference for the k factors and the presence rate of these underlying factors in the diabetes medical project, PSVD can predict the patient's preference score for the corresponding project.
[0103] Preference score definition: SS u,i ′ = p u *I i + μ*(δ u + δ i ), where p u , I i represent the single patient factor vector and the single diabetes project factor vector respectively, δ u , δ i represent the rating deviation value when predicting the patient's u preference score for the medical project i, the former represents the patient factor deviation value, and the latter represents the medical project factor deviation value. μ represents the set rating deviation parameter.
[0104] Step 3.2, introduce the least squares method to reduce the deviation rate, which is specifically represented as follows:
[0105]
[0106] where u, i ∈ SS represents a certain user u and medical project i in the selected data set, SS u,i , SS u,i ′ represent the predicted preference score of the patient u for the medical project i before and after the dimension reduction of the preference matrix R, and γ represents the set threshold parameter.
[0107] Step 3.3, score prediction for the target diabetic patient based on the historical scores of similar patients, and the final prediction score is defined as follows:
[0108]
[0109] where the first addend represents the correlation degree between the diabetic patients u and v, SS u,i , SS v,i represent the scores of the patient on the medical project i, is the average score of the patient u and v on the medical project, and N represents the total number of diabetic medical projects. Thus, the prediction preference score based on the PSVD method of multiple diabetic patient correlation degree is obtained.
[0110] Step 4, establish a knowledge graph KI based on diabetes, and obtain a vector representation of entity attribute triples, T u = {(h, r, a) | h ∈ History u}, where History u represents the set of historical medical project entities associated with the diabetic patient u, T u represents the triple information contained in the set of historical medical project entities of the patient u, h represents the entity in KI, a represents the attribute of the entity, and r represents the relationship between the entity and the attribute.
[0111] Step 5, use the a attribute in the vector triple to define the method to obtain the weight of the diabetic medical entity attribute in all medical project entity attributes related to the diabetic patient, and judge the interest preference level of the patient on the diabetic entity attribute based on the weight.
[0112] Based on the triple (h, r, a) in the knowledge graph, the weight of the a attribute in the medical project attribute related to the patient u is defined. The weight value of the a attribute in all medical project entity attributes of the patient u is defined as follows:
[0113]
[0114] where represents the weight value of the a attribute in all medical project entity attributes of the patient u. h a T represents the previous entity associated with the a attribute, r a represents the relationship vector associated with the a attribute, (h, r, a) ∈ T u represents the selection of triples containing the a attribute in the knowledge graph, and h and r represent the previous entity and relationship vector in all triples containing the a attribute.
[0115] Step 6, use the triple T uthe a attribute in the interest of the patient, by analyzing the entity attribute in the interest of the patient, a method for calculating the patient's interest attribute preference is established:
[0116] Interest u->a represents the interest preference value of the patient u to the diabetes project attribute, represents the single diabetes project entity attribute vector value, r a represents the weight of the entity attribute a as W a , which is (h,r,t)∈T a represents a set of triple vector representations containing diabetes project entity attribute a, while calculating three level parameters γ1, γ2, γ3, and the final interest preference level is obtained according to the parameters.
[0117] Step 6.1, obtain the weight of the entity attribute After that, the interest preference level of the patient u to the attribute a is determined based on this, and the specific representation method is as follows:
[0118]
[0119] In the formula, Interest u->a represents the interest preference attribute value of the patient u to the entity attribute a, and the weight of the attribute a is represented as
[0120] Step 6.2, set three level parameters, and compare the set level parameters with the interest preference attribute value of the patient to the entity attribute obtained in step 6.1 to obtain the interest preference level;
[0121] Step 6.3, set the first level parameter γ1, according to the vector triplets in the knowledge graph KI, only take out the front entity vector h a and the relationship vector r a of the entity attribute a, and perform vector dot product operation. Because it only depends on the entity attribute a, it can be explained that the current attribute a matches the lowest interest preference of the patient, so this parameter represents that the patient is least interested, and the specific representation is as follows:
[0122]
[0123] Step 6.4, set the second level parameter γ2, according to the vector triplets in the knowledge graph KI, as long as the triplets contain attribute a, take out the front entity h and relationship vector r, and perform vector dot product operation. Set this parameter to represent that the current attribute a matches the medium interest preference of the patient, so this parameter represents that the patient is interested. The specific representation is as follows:
[0124]
[0125] Step 6.5, set the third level parameter γ3, which represents that the current attribute a matches the patient's interest preference degree the highest, so this parameter represents that the patient is very interested. The specific representation is as follows:
[0126]
[0127] Step 6.6, determine the patient's preference level for the medical project attribute a:
[0128] If Interest u->a <γ1, it is considered that the patient u is not interested in the interest preference level of attribute a;
[0129] If γ1< Interest u->a <γ2, it is considered that the patient u is interested in the interest preference level of attribute a;
[0130] If Interest u->a >γ3, it is considered that the patient u is very interested in the interest preference level of attribute a, and thus the patient's interest preference level model for attribute a is obtained based on the knowledge graph KI.
[0131] Step 7, after the above step 5, the positive and negative related entities in the method judgment data set proDatas are established with "diabetes" as the core word. Remove the recommended scheme corresponding to the "negative related entity", and add the recommended scheme containing the "positive related entity" to the final output recommended list set SET.
[0132] Step 7.1, for the triplets formed by several entities in the knowledge graph KI, a single triplet (h, r, t), where h represents the front entity and t represents the back entity, the front entity h is replaced by any one random entity h random in the knowledge graph KI, and the back entity is also replaced by t random , the newly generated triplets (h random , r, t) and (h, r, t random ) do not exist in the knowledge graph KI.
[0133] Step 7.2, define the W h vector representation maps entities from entity space to relationship space, and sets the edge function f, which obtains the square value of the vector module, and the specific representation is as follows:
[0134] f(h, t) = ‖h - t - (W h ·h·W T -W h h ·t·W T h)| 2 .
[0135] Step 7.3, based on the edge function f defined in step 7.2, design a method model for distinguishing positive and negative correlation entities, which is specifically represented as follows:
[0136]
[0137] In the formula, class h,t The precondition for model execution is: (h, r, t) ∈ T u And And
[0138] In the above formula, for the front entity h, h random in the knowledge graph KI, the values of the corresponding edge functions f(h, t) and f(h random , t) are calculated, for the back entity t, t random in the knowledge graph KI, the values of the corresponding edge functions f(h, t) and f(h, t random ) are calculated, if f(h, t) is greater than f(h random , t), it is considered that the entity h in KI is a positive correlation entity, if f(h, t) is greater than f(h, t random ), it is considered that the entity t in KI is a positive correlation entity, and the rest is a negative correlation entity.
[0139] Step 8, execute the fusion recommendation method to obtain the final recommendation result set SET.
[0140] Step 8.1, for the two tuples (u, i) in the knowledge graph KI, if That is, the preference level of the patient u for the medical project i is not interested, then the algorithm terminates, otherwise step 8.3 is executed;
[0141] Step 8.2, for the two tuples (u, i) in the knowledge graph KI, if
[0142]
[0143] That is, the preference level of the patient u for the medical project i is interested, then the algorithm terminates, otherwise step 8.3 is executed;
[0144] Step 8.3, for the two tuples (u, i) in the knowledge graph KI, if That is, the entities h and t are negative correlation entities, then the algorithm terminates, otherwise step 8.4 is executed;
[0145] Step 8.4, if the number of recommended schemes in the SET does not reach K, it is added to the recommended result set SET, and step 8.1 is executed, otherwise step 8.5 is executed;
[0146] Step 8.5, if the data in the SET set reaches k, the SET is output to form a recommended result set, and the recommendation ends; otherwise, step 8.1 is executed.
[0147] Step 9, return the final recommended result set SET.
[0148] The specific operation steps are as follows:
[0149]
[0150]
[0151] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited to this, any person skilled in the art can understand and think of the transformation or replacement within the technical range disclosed by the present application, which should be covered in the inclusive scope of the present application, therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for implementing an online medical recommendation system, characterized in that, This online medical recommendation system includes a homepage module, a personal information page module, a medical communication and feedback module, a disease community module, and a rehabilitation and physiotherapy suggestion module. The homepage module summarizes and displays the entry points to each sub-module. The personal information page module displays basic patient information. The medical communication and feedback module allows users to exchange medical information. The disease community module displays disease-related medical information. The rehabilitation and physiotherapy suggestion module provides users with recommended rehabilitation and physiotherapy treatments. The method includes the following steps: Step 1: Based on the dataset, define the patient preference matrix R for medical services as follows: In the formula, several patients and medical items in the system are represented by two multidimensional vectors, P = [p1, p2, p3, ... p u ,...p p M = [M1, M2, M3, ... M] i ,...M m ], The meaning represented is the patient u's preference score for medical item i; Step 2: Based on the idea of SVD, remove redundant data from the preference matrix R to obtain a new preference matrix R. ′ ; Step 3: Define the parameter TOP-K to represent matrix R. ′ The dimension, based on the new preference matrix R ′ Calculate the patient's predicted preference score for unrated medical items; Step 4: Convert the medical project entity into a vectorized triple, obtaining the vector representation of the entity attribute triple, i.e. T u ={(h,r,a)|h∈History u } Among them, History u T represents the set of historical medical records associated with a patient. u Let represent the triplet information contained in the patient u's historical medical item entity set, h represent the entity in the recommendation system, a represent the entity's attribute, and r represent the relationship between the entity and the attribute. Step 5: Using the 'a' attribute in the vector triple, define a method to obtain the weight of a certain medical entity attribute among all medical item entity attributes related to the patient, and judge the patient's interest preference level for a certain entity attribute based on this weight. Step 6: Based on the established weight model of medical item attribute a and the vector triples in the knowledge graph, construct a level determination model for patients' interest and preference for medical items. Step 7: Based on the established knowledge vector triples, construct entity models of positively and negatively correlated medical projects; Step 8: Integrate the improved PSVD prediction score model, patient interest attribute preference level model, and positive and negative correlation entity model to derive the fusion recommendation algorithm.
2. The implementation method of the online medical recommendation system according to claim 1, characterized in that, The specific steps for step 2 are as follows: Step 2.1: Based on the established matrix R, set a critical value E. The critical value determination depends on the one-dimensional vector dataset formed by the data on the subdiagonal of matrix R, defined as Q = [Q1, Q2, Q3, ..., Q...]. i ,...Q m Based on the concept of SVD, the data on the diagonal of the preference matrix are singular values. Removal is achieved using these singular values, effectively ensuring the accuracy of removing marginal data. The definition of the critical value E is as follows: In the formula, Let m be the mean of the vector dataset Q, and m be the vector dimension of Q; Step 2.2: Based on E in Step 2.1, determine the values in matrix R. Relationship with E: like In matrix R, retain The data in the row and column; like In matrix R, delete The data in the row and column; Finally, the new preference matrix R is obtained. ′ .
3. The method for implementing an online medical recommendation system according to claim 2, characterized in that, The specific steps for step 3 are as follows: Step 3.1: Convert the new preference matrix R ′ It is decomposed into a patient factor matrix and a medical project factor matrix, and denoted as U respectively. p*k ,I u*k The prediction preference score of patient u for medical item i is defined as follows: SS u,i ′ =p u *I i +μ*(δ u +δ i ) In the formula, p u I i δ represents the factor vector of a single patient and the factor vector of a single medical item, respectively. u δ i This represents the rating deviation value when predicting the preference score of patient u for medical item i. The former represents the patient factor deviation value, and the latter represents the medical item factor deviation value; μ represents the set rating deviation parameter. Step 3.2: Considering the potential error in patients' predicted scores for medical procedures, the least squares method is introduced to reduce the deviation rate, as shown below: In the formula, u,i∈SS represents selecting a user u and medical item i from the dataset, SS u,i SS u,i ′ Let represent the predicted preference scores of patient u for medical item i before and after dimensionality reduction of the preference matrix R, respectively, and γ represent the set threshold parameter; Step 3.3: After processing the deviation error of the predicted score, the error of the predicted score is reduced. The target user's score is predicted based on the historical scores of similar users. The final predicted score is defined as follows: In the formula, the first addend represents the degree of association between patient u and patient v. Here, the PCC theorem is used to measure the similarity between roles. u,i SS v,i All of these represent the patient's rating of medical service item i. , where are the average ratings of patients u and v for medical items, respectively, and N represents the total number of medical items in the system.
4. The implementation method of the online medical recommendation system according to claim 3, characterized in that, The weight of the medical attribute 'a' is obtained in step 5 as follows: Based on the triple (h, r, a) in the knowledge graph, the weight of attribute 'a' in the medical item attributes related to patient u is defined as follows: In the formula, h a T Represents the preceding entity associated with attribute 'a', r a Represents the relation vector associated with attribute 'a', (h,r,a)∈T u This indicates the selection of triples in the knowledge graph that contain the attribute 'a', where h and r represent the preceding entity and relation vector in the triples containing the attribute 'a', respectively.
5. The method for implementing an online medical recommendation system according to claim 4, characterized in that, Step 6, constructing the level determination model, includes the following steps: Step 6.1: Obtain the weights of entity attributes. Then, based on this, the patient u's level of interest preference for attribute a is determined, and the specific representation method is as follows: In the formula, Interest u->a This represents the interest preference attribute value of patient u for entity attribute a, and the weight of attribute a is expressed as: Step 6.2: Set three level parameters, and compare the set level parameters with the patient's interest preference attribute values for entity attributes obtained in Step 6.1 to obtain the interest preference level; Step 6.3: Set the first level parameter γ1, specifically as follows: Step 6.4: Set the second level parameter γ2, as follows: Step 6.5: Set the third level parameter γ3, as follows: Step 6.6: Determine the patient's preference level for medical service attribute a: If Interest u->a If γ < 1, then patient u is considered to have no interest in attribute a. If γ1 <Interest u->a If γ < 2, then patient u is considered to have an interest preference level of 'interest' for attribute a; If Interest u->a If γ > 3, then the patient u is considered to have a very high level of interest in attribute a. Thus, the patient's level of interest in attribute a is obtained based on the knowledge graph KI.
6. The method for implementing an online medical recommendation system according to claim 5, characterized in that, Step 7, constructing the entity model of positive and negative correlations in medical projects, includes the following steps: Step 7.1: Define random entities in the knowledge graph KI: For a triple formed by several entities in the knowledge graph KI, a single triple (h, r, t) is defined, where h represents the preceding entity and t represents the following entity. The preceding entity h is then assigned to any random entity h in the knowledge graph KI. random The entity is replaced by t. random The newly generated triple (h) random ,r,t)(h,r,t) random It does not exist in the knowledge graph KI; Step 7.2: Based on the preceding and following entities h and t defined in Step 7.1, define W. h Vector representation maps entities from entity space to relation space. A boundary function f is defined, which yields the squared value of the vector's magnitude, specifically expressed as: f(h,t)=‖h-t-(W h T ·h·W h -W h T ·t·W h )|| 2 ; Step 7.3: Establish a judgment model for positively correlated and negatively correlated entities. Based on the marginal function f defined in Step 7.2, design a method model to distinguish between positively and negatively correlated entities. The positively and negatively correlated entity model is specifically represented as follows: In the formula, class h,t The prerequisite for model execution is: (h,r,t)∈T u and and Step 7.4, Analyze the class h,t The two result values are used to determine the positive and negative related entities.
7. The method for implementing an online medical recommendation system according to claim 6, characterized in that, Step 8 of the fusion recommendation algorithm includes the following steps: Step 8.1: For the pair (u,i) in the knowledge graph KI, if If patient u's preference level for medical item i is "uninterested", the process terminates; otherwise, proceed to step 8.
3. Step 8.2: For the two-tuple (u,i) in the knowledge graph KI, if If patient u's preference level for medical item i is "interested", the process terminates; otherwise, proceed to step 8.
3. Step 8.3: For the pair (u,i) in the knowledge graph KI, if the following conditions are not met... If entities h and t are negatively related entities, the process terminates; otherwise, proceed to step 8.
4. Step 8.4: If the number of recommended solutions in SET does not reach K, add them to the recommendation result set SET and execute step 8.1; otherwise, execute step 8.
5. Step 8.5: If the number of data in the SET set has reached k, then output SET to form the recommendation result set, and the recommendation process ends; Otherwise, proceed to step 8.1.
Citation Information
Patent Citations
Community health medical interaction system based on internet and implementation method thereof
CN106530168A