Patient relation association and information matching method and device based on knowledge graph
By constructing a patient relationship correlation and information matching method based on knowledge graph, the problem of patient identity identification difficulties caused by high name overlap in Tibetan people is solved, and information matching with high accuracy and reliability is achieved, reducing the cost and delay of medical services.
Patent Information
- Application Number
- CN202510136083.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-07
AI Technical Summary
Tibetan people have difficulty identifying patients due to their high overlap in names in medical services. Traditional name-based search methods fail to deal with high-profile populations, increasing the additional costs of patients and medical institutions.
Using a knowledge graph-based patient relationship association and information matching method, a triple between patients, patients and families is generated by constructing a knowledge graph of hospital history information, and matching and verification is performed using autoencoder and related vector machine algorithms to improve the accuracy and reliability of the matching.
It solves the problem of patient identification difficulties caused by high name overlap among Tibetan people, improves the accuracy and reliability of patient information matching, and reduces the cost and delay of medical services.
Smart Images

Figure CN120067157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent healthcare, and more specifically, to a method, device, medium, and program product for patient relationship association and information matching based on a knowledge graph. Background Art
[0002] The present invention relates to the field of data processing technologies, and particularly to a method for patient relationship association and information matching based on a knowledge graph.
[0003] The background of the present invention is based on specific problems in the medical services of the Tibetan population, namely, the relatively high name coincidence rate caused by naming traditions, religious cultures, and language characteristics. This phenomenon causes difficulties in patient identity recognition in medical scenarios. Especially when patients do not carry identity cards or other supporting documents, it is difficult to quickly and accurately retrieve their historical medical records based solely on their names, directly affecting the continuity and accuracy of medical services.
[0004] Existing methods for retrieving patient information based on names often fail when dealing with high name coincidence rates. Patients may need to go home to get their identity cards or other supporting materials, resulting in increased time costs and even potential delays in necessary medical services. In addition, in remote areas or emergency medical scenarios, the difficulty of obtaining these materials further exacerbates the problem. Simply relying on patients' self-reported information also brings unnecessary manual searches and operational burdens due to communication barriers or input errors, increasing the pressure on medical institutions.
[0005] The Chinese invention patent with the publication number CN113688255B proposes a method for constructing a knowledge graph based on Chinese electronic medical records. Most of the currently constructed knowledge graphs contain a small number of medical record corpora, have a small scale, and are often only applicable to a single department or disease, with poor generality. Moreover, some relatively complete medical record knowledge graphs require a large amount of manual participation, are time-consuming and laborious, and have poor scalability. Due to different descriptions of disease categories between different departments and diseases in electronic medical records, the corresponding language environments for a series of examinations and treatments are also different, and the habitual expressions of doctors corresponding to different disease categories are different. These characteristics make the effects of some deep learning methods decline, and it is not easy to expand the knowledge graph construction framework. A data analysis and processing method for knowledge graphs based on Chinese electronic medical records, a corpus annotation process specification, and an entity relationship extraction scheme are formulated for the above existing problems.
[0006] The Chinese invention patent with the publication number CN118629571B proposes a clinical test result review method and system based on artificial intelligence and big data, which relates to the technical field of medical information processing, including: constructing a first knowledge graph, obtaining physiological disease characteristics and quality control information, historical test data and medication conditions, generating a heterogeneous data set, performing intelligent fusion, extracting semantic features and expanding them to obtain high-quality sample data; obtaining real-time message data of test equipment, performing security authentication, comparing adjacent sample data, generating a sample association verification result, constructing an initial automatic review model and outputting an automatic review result, comparing the automatic review result with the manual approval result to obtain a high-precision automatic review model; obtaining sample data and inputting it into the high-precision automatic review model, performing semantic reasoning, judging the abnormal risk and outputting an intelligent review report, calculating the confidence level and comparing it with the confidence level threshold, and outputting the intelligent review report with a confidence level higher than the confidence level threshold as a credible review result.
[0007] The Chinese invention patent with the publication number CN119046421A proposes an intelligent question-answering system and method based on deep semantic understanding and knowledge graph fusion. The system includes a receiving end, an analysis end, and a reply end. The method includes the receiving end collecting questions about specific diseases raised by users, the analysis end performing deep semantic understanding and knowledge graph mining operations on the questions based on an intelligent question-answering model to obtain relevant medical knowledge and data structures of the questions, and the reply end performing knowledge fusion based on the relevant medical knowledge and data structures to obtain a reply result and output it to the user. The present invention can solve the problem of the lack of deep semantic understanding and knowledge graph fusion functions in existing intelligent question-answering systems, improve the accuracy and response speed of the system, provide users with an efficient and reliable medical consultation channel, can realize the full process automation from the processing of user queries to the output of replies, reduce the risk of manual intervention, and promote the popularization and development of intelligent technologies.
[0008] The existing technologies have the following deficiencies:
[0009] 1. In the patient relationship association and information matching task, the traditional patient information retrieval method based on name cannot cope with the high rate of name repetition, easily leads to identity recognition failure, and increases the additional costs of patients and medical institutions.
[0010] 2. In the patient relationship association and information matching task, the existing methods do not make full use of patient relationship information, only rely on simple text or basic attribute comparison, and it is difficult to fully capture the associations among patients, accompanying persons, and family members, resulting in low matching accuracy.
[0011] 3. In the patient relationship association and information matching task, there is a lack of an effective verification mechanism, and it is impossible to further judge the matching results, which is easily affected by noise data or false information in medical decision-making and reduces the credibility of the information. Summary of the Invention
[0012] In view of the above problems, the present invention provides a method for patient relationship association and information matching based on a knowledge graph, which constructs a knowledge graph using historical hospital information to achieve rapid matching of patient information.
[0013] This application (in the first aspect) discloses a method for patient relationship association and information matching based on a knowledge graph, including:
[0014] S1: Obtain patient visit information records;
[0015] S2: Obtain patient escort person triples based on the visit information records, where the three elements of the patient escort person triples are the patient, the escort relationship, and the escort person;
[0016] S3: Match the patient escort person triples with the visit information knowledge graph to obtain a matching result; wherein, the acquisition method of the visit information knowledge graph is:
[0017] Step 1, obtain a patient visit information record data set;
[0018] Step 2, construct a knowledge graph for the visit information record data set to obtain the visit information knowledge graph.
[0019] Further, the method further includes: S4: Verify the relationship of the matching result to obtain a verified matching result.
[0020] Further, the specific method of matching in S3 is: after vectorizing the patient escort person triples to obtain a triple vector, the triple vector is converted by an encoder to obtain a low-dimensional space expression of the triple vector; after vectorizing all the triple sets in the visit information knowledge graph and inputting them into the encoder for conversion to obtain a low-dimensional space expression of the knowledge graph triple vector, the low-dimensional space expression of the triple vector is matched one by one with the low-dimensional space expression of the knowledge graph triple vector set to obtain a matching degree, and the knowledge graph triple vector with the highest matching degree is used as the matching result;
[0021] Optionally, the encoder is an autoencoder with adaptive spatial transformation projection, and the construction method includes:
[0022] Step 1: Obtain patient relationship triple training set data, and label the true association relationship of the patient relationship triples; construct an initial autoencoder with a multi-layer structure and initialize the parameters of each layer;
[0023] Step 2: Vectorize the patient relationship triples to obtain patient relationship triple vectors;
[0024] Step 3: The patient relationship triple vector is transformed by a projection matrix to obtain a reduced-dimensional low-dimensional space;
[0025] Step 4: The patient relationship triple vector is input into an autoencoder adjusted by a projection matrix to obtain a compressed low-dimensional space;
[0026] Step 5: The parameters are updated iteratively by backpropagation through a loss function until a stopping condition is reached, and the autoencoder with adaptive space transformation projection is obtained. The loss function is calculated based on the patient relationship triple vector, the reduced-dimensional low-dimensional space, the parameters of the autoencoder, and the output;
[0027] Optionally, the patient relationship triple vector is normalized and then input into the space transformation module and the autoencoder respectively.
[0028] Furthermore, the normalization processing method is expressed as:
[0029]
[0030] where μ a and Σ a are the mean vector and standard deviation vector of X a respectively; X' a is the normalized patient triple data;
[0031] Optionally, the dimensionality reduction process of the space transformation module is expressed as:
[0032]
[0033] where is the dimensionality-reduced patient triple data, X' a is the normalized patient triple data, and P r is a projection matrix dynamically optimized according to the characteristics of the input normalized patient triple data;
[0034] Optionally, the projection matrix is calculated using the principal component analysis method.
[0035] Optionally, the covariance matrix of the normalized patient triple data X' a is decomposed, and the first n ty eigenvectors of the obtained covariance matrix are selected to form the projection matrix;
[0036] Optionally, if the dimension of the normalized patient triple data X' a is n a1 ×n a2 , then the dimension of the projection matrix is n a2 ×n ty ;
[0037] Optionally, iteratively optimize the projection matrix based on error backpropagation learning to obtain an optimized projection matrix, and use the optimized projection matrix for the spatial module dimensionality reduction;
[0038] Optionally, optimize the projection matrix by using a gradient-based adaptive adjustment term, and the optimization method is expressed as:
[0039]
[0040] In the formula, is the projection matrix of the t-th iteration; is the projection matrix of the (t + 1)-th iteration; η r is the projection matrix learning rate, is the gradient of the loss function of the autoencoder with respect to P r δ is r the adjustment term; is the local sensitivity gradient adjustment factor of the t-th iteration; α r is the adjustment factor;
[0041] Optionally, α r is set to 0.2, and δ r is the information entropy of the projection matrix;
[0042] Optionally, the update method of the local sensitivity gradient adjustment factor is expressed as:
[0043]
[0044] In the formula, is the local sensitivity gradient adjustment factor of the (t + 1)-th iteration; η tre is the learning rate of the adjustment factor, represents the gradient of the loss function with respect to the projection matrix of the t-th iteration, is the L2 norm of the gradient of the loss function with respect to the projection matrix, which is used to measure the change amplitude of the gradient.
[0045] Furthermore, the network structure of the autoencoder includes multiple layers, and the parameters of each layer include a weight matrix and a bias term. The initialization of the parameters of each layer is expressed as: Let the number of layers of the autoencoder be L red layers, and the initialization methods of its weight matrix and bias term are expressed as:
[0046]
[0047] In the formula, represents the weight matrix of the l-th layer of the autoencoder, is the bias term of the l-th layer of the autoencoder; ~ represents being subject to a specific distribution; Represents a normal distribution with a mean of 0 and a variance of 0.01;
[0048] Optionally, the output of the autoencoder is expressed as:
[0049]
[0050] In the formula, is the activation output of the l-th layer of the encoder; is the activation output of the (l - 1)-th layer of the encoder; Sig() is the Sigmoid activation function; is the underestimation correction factor for the t-th iteration; represents the weight matrix of the l-th layer of the autoencoder, is the bias term of the l-th layer of the autoencoder;
[0051] Optionally, the underestimation correction factor is continuously updated using a dynamic adjustment method based on error feedback, and the update method is expressed as:
[0052]
[0053] In the formula, is the underestimation correction factor for the (t + 1)-th iteration; η λ is the learning rate of the underestimation correction factor, and are the first and second derivatives of the autoencoder loss function with respect to the underestimation correction factor for the t-th iteration, respectively;
[0054] Furthermore, the loss function is expressed as:
[0055]
[0056] The parameter update rule of the autoencoder is:
[0057]
[0058] In the formula, the vectorized patient triplet data input to the autoencoder is X a , is the patient triplet data after dimensionality reduction, ∥∥ 1 represents the L1 norm; ∥∥ 2 represents the L2 norm; η der is the learning rate of the autoencoder; L r is the loss function of the autoencoder; is the activation output of the last layer of the encoder; is the weight matrix of the last layer of the autoencoder; and are the partial derivatives of the autoencoder loss function with respect to the weight and bias, respectively; Denotes the gradient of the loss function with respect to the projection matrix at the t-th iteration; is the L2 norm of the gradient of the loss function with respect to the projection matrix; pes is the power adjustment parameter used to control the amplification effect of the gradient; α red is the first weighting factor; β red is the second weighting factor.
[0059] Furthermore, the relationship verification is performed using a relevance vector machine algorithm based on adaptive feature learning; the acquisition method of the adaptive feature learning relevance vector machine algorithm includes:
[0060] Step 1: Obtain a training patient - escort relationship dataset and label high - confidence and low - confidence;
[0061] Step 2: Initialize the kernel function parameters, learning rate, and regularization parameters of the relevance vector machine;
[0062] Step 3: Vectorize the patient data features in the patient - escort relationship dataset to obtain data feature vectors, and perform feature mapping on the data feature vectors to obtain feature vectors;
[0063] Step 4: Perform training iterations of the relevance vector machine in the Bayesian framework based on the feature vector space until the stop condition is reached to obtain the adaptive feature learning relevance vector machine algorithm;
[0064] Optionally, the feature mapping is expressed as:
[0065] where q i is the i - th patient data feature vector input to the relevance vector machine; q′ i is the i - th patient data feature vector after mapping; M() is the mapping function, and Θ qM is the parameter set required for mapping;
[0066] Optionally, the mapping function is adjusted by optimizing the following learning objective function, expressed as:
[0067]
[0068] where w ij is the similarity weight based on q i and q j in the original feature space; q j is the j - th patient data feature vector input to the relevance
[0069] q′ i = M(q i , Θ qM )
[0070] machine.
[0071] Optionally, the calculation method of the mapping function is expressed as:
[0072]
[0073] In the formula, represents the set of samples that are the nearest neighbors to the i-th patient data feature vector input to the relevance vector machine, and w ij is the contribution weight of the i-th patient data feature vector input to the relevance vector machine to the j-th patient data feature vector input to the relevance vector machine.
[0074] Optionally, the calculation method of the contribution weight is expressed as:
[0075]
[0076] In the formula, ∥∥ is the L2 norm; q k is the k-th patient data feature vector input to the relevance vector machine;
[0077] Optionally, the relevance vector machine of the Bayesian framework selects the patient data point with the highest posterior probability as the support vector by calculating the posterior probability distribution, and the calculation method is expressed as:
[0078]
[0079] In the formula, y is the label vector of the input patient data. For example, the label vector includes 0 and 1, corresponding to "high confidence" and "low confidence" respectively; p(y|q′ i , θ qk ) represents the posterior probability distribution, which characterizes the probability predicted based on the parameter θ qk ; θ qk is the kernel function parameter of the relevance vector machine.
[0080] Optionally, in each iteration, the relevance vector machine of the Bayesian framework calculates the energy entropy of the model and optimizes the parameters of the model. The calculation method of the energy entropy is expressed as:
[0081]
[0082] In the formula, H q represents the energy entropy; Ncs is the number of samples input to the relevance vector machine in the current batch.
[0083] Optionally, the relevance vector machine of the Bayesian framework adaptively adjusts the model parameters, and the dynamic update method of the kernel function parameter is expressed as:
[0084]
[0085] In the formula, is the kernel function parameter of the updated relevance vector machine, γ q is the learning rate of the relevance vector machine, represents the gradient of the kernel function parameter of the relevance vector machine.
[0086] Optionally, the gradient is calculated as:
[0087]
[0088] where y i is the label of the i-th patient data.
[0089] The second aspect of this application discloses a patient relationship association and information matching system based on a knowledge graph, including:
[0090] An acquisition module 201: used to acquire patient visit information records;
[0091] An extraction module 202: used to obtain patient escort person triples based on the visit information records, and the three elements of the patient escort person triples are patients, escort relationships, and escorts;
[0092] A matching module 203: used to match the patient escort person triples with the visit information knowledge graph to obtain a matching result; wherein, the acquisition method of the visit information knowledge graph is:
[0093] Step 1, acquire a patient visit information record data set;
[0094] Step 2, construct a knowledge graph for the visit information record data set to obtain the visit information knowledge graph.
[0095] The third aspect of this application discloses a computer device, the device includes: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, it is used to execute the steps of the above method.
[0096] The fourth aspect of this application discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above method.
[0097] The fifth aspect of this application discloses a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the above method.
[0098] This application has the following beneficial effects:
[0099] 1. In the patient relationship association and information matching task, the knowledge graph construction technology is adopted to generate triples by extracting the relationship data among patients, companions, and family members, solving the problem that it is difficult to identify the identity of patients in the Tibetan population due to the high coincidence rate of names. Through the attribute storage and standardization processing of relationship nodes, the structuring and organization of personnel identity data are enhanced, providing a basis for subsequent rapid matching.
[0100] 2. In the patient relationship association and information matching task, the autoencoder algorithm based on adaptive spatial transformation projection is adopted to reduce the dimension of patient-companion triples to a low-dimensional space, capture the non-linear relationship of data, and eliminate redundant information. Through the optimization of the dynamic projection matrix of the spatial transformation module, the problem that it is difficult to unify the semantics of triple data is solved, and the accuracy of matching is improved.
[0101] 3. In the patient relationship association and information matching task, the relevance vector machine algorithm based on adaptive feature learning is adopted to verify the relationship of the matching results. Using multi-dimensional features such as co-occurrence frequency, regional similarity, and language rules, high-confidence relationships are screened, overcoming the problem of high misjudgment rate of simple name similarity judgment in actual scenarios and improving the reliability of matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0103] Figure 1 is a schematic flowchart of the method provided in the first aspect of the embodiment of the present invention;
[0104] Figure 2 is a schematic diagram of the program product provided in the second aspect of the embodiment of the present invention;
[0105] Figure 3 is a schematic diagram of the computer device provided in the embodiment of the present invention;
[0106] Figure 4 is a schematic diagram of the architecture of the exemplary computing device provided in the embodiment of the present invention;
[0107] Figure 5 is a schematic diagram of the storage medium provided in the embodiment of the present invention;
[0108] Figure 6 is a schematic diagram of the patient relationship association and information matching between patient triples and the knowledge graph provided in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0109] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0110] In some processes described in the specification and claims of the present invention and the above-mentioned accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0111] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0112] Figure 1 It is a schematic flow chart of a method for patient relationship association and information matching based on an evaluation of a knowledge graph provided by an embodiment of the present invention. Specifically, the method includes the following steps:
[0113] S101: Obtain the patient's medical record information;
[0114] S102: Obtain a patient escort triple based on the medical record information. The three elements of the patient escort triple are the patient, the escort relationship, and the escort;
[0115] S103: Match the patient escort triple with the medical record information knowledge graph to obtain a matching result. Among them, the acquisition method of the medical record information knowledge graph is as follows:
[0116] Step 1: Obtain a dataset of the patient's medical record information;
[0117] Step 2: Construct a knowledge graph for the medical record information dataset to obtain the medical record information knowledge graph.
[0118] The present invention aims to solve the problem of difficult patient identification caused by a high degree of name coincidence among the Tibetan population during medical treatment. In the Tibetan population, the phenomenon of high name coincidence is determined by specific naming traditions, religious and cultural influences, and language characteristics. For example, naming methods using Buddhist terms, auspicious words, or family traditions are likely to result in duplicate names. This phenomenon is particularly prominent in medical scenarios. When a patient does not carry an ID card or other documents to prove their identity, it is impossible to quickly and accurately retrieve their historical medical records based solely on their name, posing a significant challenge to the continuity and accuracy of medical services.
[0119] Traditional name-based retrieval methods often fail due to the phenomenon of duplicate names, which may require patients to go home to get ID cards and other proof materials. This not only increases the time cost for patients but may also delay medical services, especially in remote areas or emergency medical scenarios. In addition, simply relying on patients' self-report may bring unnecessary manual searching and communication costs, further increasing the operating burden on medical institutions.
[0120] Therefore, the present invention proposes a patient relationship association and information matching method based on a knowledge graph ( Figure 6 as shown), and the main steps are as follows:
[0121] 1. Construct a knowledge graph of personnel relationships for all patients
[0122] In the process of constructing the knowledge graph of personnel relationships, data collection and preprocessing are first carried out;
[0123] Data collection is from the hospital's electronic health record system to obtain information such as patient names, ID numbers, names of accompanying persons, names of family members, and contact information;
[0124] Furthermore, data cleaning is performed on the data, including: removing redundant records (such as duplicate information entered multiple times for the same patient), handling polyphonic characters, pinyin, or typos, and unifying all names into a standard format (such as full pinyin or regional language form) to ensure data consistency.
[0125] Furthermore, for each patient, extract the personnel relationships related to them to generate triples of the knowledge graph. The generation rules for the triples are as follows:
[0126] Patient and accompanying person: (patient name, relationship, accompanying person name), the relationship is "accompanying person"
[0127] Patient and family member: (patient name, relationship, family member name), the relationship is "family member";
[0128] Family member and accompanying person: (family member name, relationship, accompanying person name), the relationship is "contact person".
[0129] It should be noted that during the generation process of triples, each triple should be able to uniquely describe a node and its relationship to ensure the accuracy of subsequent graph matching.
[0130] Furthermore, use a graph database (such as Neo4j) to store the generated triples, and the database structure is as follows:
[0131] Nodes: Represent patients, companions, family members, etc. Attributes include name, identity (patient / companion / family member), contact information, etc.;
[0132] Edges: Represent the relationships between two nodes, and the types are "companion", "family member", etc.
[0133] For clear representation, in one embodiment, simplify Tibetan names, and the example triples constructed are as follows:
[0134] (Tashi, companion, Tsering Dargye)
[0135] (Tashi, family member, Drolma)
[0136] (Drolma, contact person, Tsering)
[0137] (Lhamo, family member, Tsering)
[0138] (Dorje, companion, Rinchen)
[0139] It should be noted that the complexity, length, and frequency of rare words in Tibetan names are higher than the example data in this embodiment;
[0140] In an ideal situation, the knowledge graph in the graph database accumulates and updates continuously, and finally forms a complete knowledge graph of the personnel relationships of all patients in the hospital.
[0141] 2. Construct triples for the current patient and their companion
[0142] In the case where the patient does not carry an ID card, the hospital registration system needs to collect the patient's basic information, provided by the patient's self-report or the companion, including:
[0143] Patient name: The name of the current patient;
[0144] Companion name: If there is a companion, collect their name;
[0145] Relationship description: Mark the relationship between the companion and the patient (such as companion, family member, etc.).
[0146] Furthermore, generate patient-related triples based on the collected information. In one embodiment, the generated triples are as follows:
[0147] (Tashi, companion, Tsering)
[0148] 3. Triplet Matching
[0149] The triplets composed of patients and companions are mapped into a low-dimensional space through an autoencoder for efficient matching with the triplets in the knowledge graph;
[0150] Each part of the triplet, including two nodes (patient name, companion name) and the relationship type, is vectorized to adapt to the input requirements of the autoencoder. The present invention uses the Word2Vec algorithm to vectorize the text. The Word2Vec algorithm is a commonly used vectorization algorithm in the art. According to a preset large-scale corpus, it scans the text to be vectorized and represents each word as a one-hot encoded vector. The dimension of the one-hot encoded vector is equal to the size of the vocabulary in the corpus.
[0151] Furthermore, the encoder of the autoencoder compresses the high-dimensional vector into a low-dimensional space to extract the core semantic features;
[0152] And the decoder of the autoencoder restores the original triplet from the low-dimensional space to ensure that the model understands the relationship semantics.
[0153] The autoencoder uses the existing patient relationship triplets in the hospital as training data for model training, and labels the true association relationships in the historical triplets.
[0154] The present invention adopts an autoencoder algorithm based on adaptive space transformation projection. The adaptive space transformation projection finds the optimal representation of the vectorized patient triplet data in the low-dimensional space, and uses the projection matrix generated by the space transformation module to adjust each layer in the encoding process, so that the processing process of the autoencoder for the vectorized patient triplet data is more adaptable to the internal structure of the data, effectively removing redundant information and better capturing the non-linear relationship of the vectorized patient triplet data.
[0155] Specifically, the training process of the autoencoder algorithm based on adaptive space transformation projection is as follows:
[0156] 1) Construct the multi-layer structure of the autoencoder and initialize the parameters of each layer, including the weight matrix and the bias term. Let the number of layers of the autoencoder be L red layers, and the initialization methods of its weight matrix and bias term are expressed as:
[0157]
[0158] In the formula, represents the weight matrix of the l-th layer of the autoencoder, is the bias term of the l-th layer of the autoencoder; ~ represents following a specific distribution; represents a normal distribution with a mean of 0 and a variance of 0.01.
[0159] 2) The vectorized patient triple data input into the autoencoder is first standardized to eliminate the influence of different dimensions on training. Let the vectorized patient triple data input into the autoencoder be X a , and its standardization method is expressed as:
[0160]
[0161] In the formula, μ a and Σ a are the mean vector and standard deviation vector of X a respectively; X′ a is the standardized patient triple data.
[0162] 3) The standardized patient triple data passes through the spatial transformation module, which learns an adaptive projection matrix P r , used to map the high-dimensional standardized patient triple data to a low-dimensional space. The dimensionality reduction process is expressed as:
[0163]
[0164] In the formula, is the patient triple data after dimensionality reduction, and P r is a projection matrix dynamically optimized according to the characteristics of the input standardized patient triple data.
[0165] The projection matrix is learned based on the error backpropagation method. The initialization generation method of the projection matrix is to calculate the projection matrix using the principal component analysis method. Specifically, first perform covariance matrix decomposition on the standardized patient triple data X′ a , and select the first n ty eigenvectors of the obtained covariance matrix to form the projection matrix. If the dimension of the standardized patient triple data X′ a is n a1 ×n a2 , then the dimension of the projection matrix is n a2 ×n ty .
[0166] To enhance the projection ability, an adaptive adjustment term based on the gradient is used to optimize the projection matrix. The optimization method is expressed as:
[0167]
[0168] In the formula, is the projection matrix at the t-th iteration; is the projection matrix at the (t + 1)-th iteration; η r is the projection matrix learning rate, is the gradient of the loss function of the autoencoder with respect to P r , where δ r is the adjustment term; is the local sensitivity gradient adjustment factor for the t-th iteration; α r is the adjustment factor. Preferably, α r is set to 0.2, and δ r is the information entropy of the projection matrix.
[0169] Furthermore, the role of the local sensitivity gradient adjustment factor is to dynamically adjust the learning rate of the projection matrix in each training, enabling the model to make appropriate local adjustments and amplifications when facing high-noise or highly sensitive features to avoid information loss. The update method is expressed as:
[0170]
[0171] In the formula, is the local sensitivity gradient adjustment factor for the (t + 1)-th iteration; η tre is the learning rate of the adjustment factor, represents the gradient of the loss function with respect to the projection matrix for the t-th iteration, is the L2 norm of the gradient of the loss function with respect to the projection matrix, used to measure the change amplitude of the gradient. Preferably, η tre is set to 0.05.
[0172] 4) The autoencoder encodes the normalized patient triple data through several layers of neurons, gradually compressing it into a low-dimensional space. During this process, it learns the latent representation of the patient triple data by minimizing the reconstruction error, that is, finds the optimal expression of the patient triple data in the low-dimensional space. The projection matrix generated by the space transformation module will adjust each layer during the encoding process, making the dimensionality reduction process more adaptable to the internal structure of the patient triple data. During the dimensionality reduction process, a low-estimation adjustment optimization strategy is adopted to perform weighted correction on the output features. The output of the encoder is expressed as:
[0173]
[0174] In the formula, is the activation output of the l-th layer of the encoder; is the activation output of the (l - 1)-th layer of the encoder; Sig() is the Sigmoid activation function; is the low-estimation correction factor for the t-th iteration.
[0175] 5) The role of the low-estimation correction factor is to prevent information loss during the dimensionality reduction process. The present invention adopts a dynamic adjustment method based on error feedback to continuously update the low-estimation correction factor. The update method is expressed as:
[0176]
[0177] Wherein, is the low-estimation correction factor for the (t + 1)-th iteration; η λ is the learning rate of the low-estimation correction factor, and are respectively the first and second derivatives of the autoencoder loss function with respect to the low-estimation correction factor for the t-th iteration. Preferably, η λ is set to 0.3.
[0178] 6) During the training process, the autoencoder adjusts its parameters through the backpropagation algorithm to optimize the dimensionality reduction effect, and adopts the norm to constrain the sparsity of features, which can not only optimize the reconstruction error, but also promote the sparsity of features, thereby improving the effectiveness of the patient triple data representation after dimensionality reduction. The calculation method of the loss function is expressed as:
[0179]
[0180] And, the parameter update rule of the autoencoder is:
[0181]
[0182] Wherein, ∥∥ 1 represents the L1 norm; ∥∥ 2 represents the L2 norm; η der is the learning rate of the autoencoder; L r is the loss function of the autoencoder; is the activation output of the last layer of the encoder; is the weight matrix of the last layer of the autoencoder; and are respectively the partial derivatives of the loss function of the autoencoder with respect to the weight and bias; represents the gradient of the loss function with respect to the projection matrix for the t-th iteration; is the L2 norm of the gradient of the loss function with respect to the projection matrix; pes is the power adjustment parameter used to control the amplification effect of the gradient; α red is the first weighting factor; β red is the second weighting factor. Preferably, η der is set to 0.01, α red is set to 0.3, β red is set to 0.7.
[0183] 7) Repeat the above steps iteratively until the preset iteration stop condition is met, which indicates that the model training is completed. In one embodiment, the preset iteration stop condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.
[0184] After the autoencoder model training is completed, the triples input by the current patient are reduced to a low-dimensional vector space by the trained autoencoder, and the obtained low-dimensional features are used as the input triple features. At the same time, all the triples in the knowledge graph are reduced to the low-dimensional vector space by the trained autoencoder, and the obtained low-dimensional features are used as the graph triple features.
[0185] Further, the cosine similarity or Euclidean distance is used to calculate the similarity between the input triples and the graph triples. A high similarity indicates a high matching possibility of the input triples in the knowledge graph.
[0186] Further, sort by similarity, select several of the most similar triples as candidates, and return the patient node information corresponding to the candidate triples. In one embodiment, for example:
[0187] Input vs (Tashi, accompanier, Tsering Dargye): Similarity = 0.98 (highest match);
[0188] Input vs (Tashi, family member, Droma): Similarity = 0.92 (higher match);
[0189] Input vs (Dorje, accompanier, Rinchen): Similarity = 0.45 (low match);
[0190] Input vs (Lhamo, family member, Tsering): Similarity = 0.78 (average match).
[0191] 4. Relationship verification
[0192] Verify the relationship between the patient and the accompanier or family member for the matching result with the highest similarity, so as to improve the accuracy of knowledge graph matching.
[0193] The present invention uses the relevance vector machine algorithm based on adaptive feature learning for verification. As a classifier model, the relevance vector machine algorithm mainly uses the relevant information of patients, accompaniers, family members, etc. in the hospital electronic health record system, the relationship data in the historical medical records, and the existing patient-family member-accompanier triple data in the knowledge graph as the training data source;
[0194] The data features cover multi-dimensional information, including historical medical features (such as the frequency of historical co-occurrence of the patient and the accompanier and the time of the most recent co-occurrence), geographical features (such as the similarity of their addresses), relationship type features (whether the provided relationship is logical, such as family members, accompaniers, etc.), demographic features (such as whether the age difference between the patient and the accompanier is reasonable), language features (the language consistency of the names, such as the naming rules of Tibetan pinyin), and data source features (such as whether the record is manually input or system-generated).
[0195] Furthermore, by manually annotating data, a dataset containing high-confidence and low-confidence relationships is constructed, where positive samples represent known high-confidence relationships, such as records of patients with their real family members or companions, and negative samples represent low-confidence relationships, such as mis-matched or forged data. After each sample is vectorized, it is represented in the form of a feature vector, for example, including features such as co-occurrence times, address similarity, relationship type, etc., as well as manually annotated confidence labels, such as "high confidence" or "low confidence".
[0196] In the adopted Relevance Vector Machine (RVM) algorithm based on adaptive feature learning as the classification algorithm, the RVM is a sparse learning method based on Bayesian inference. In the present invention, adaptive feature learning is adopted in the traditional RVM algorithm, and by combining the energy entropy model, the classification decision boundary is optimized, so that the classifier can not only reflect the local structural characteristics of patient data, but also effectively capture the global distribution of patient data.
[0197] Specifically, the training process of the RVM algorithm based on adaptive feature learning is as follows:
[0198] 1) Initialize the kernel function parameters of the RVM and the basic configuration of the model. The basic configuration includes the learning rate and the regularization parameter. The initialization method is expressed as:
[0199]
[0200] γ q = 0.01
[0201] λ q = 0.1
[0202] In the formula, θ qk represents the parameter of the kernel function, is the initialization variance, γ q is the learning rate, λ q is the regularization parameter; ~ means following a specific distribution; represents the normal distribution.
[0203] 2) Process the input patient data feature vectors so that each data feature vector is mapped to a new feature space to obtain new feature vectors. This space can better reflect the internal structure and correlation of the features. The mapping method is expressed as:
[0204] q′ i = M(q i , Θ qM )
[0205] In the formula, q i is the i-th patient data feature vector input to the RVM; q′ iis the i-th patient data feature vector after mapping; M() is the mapping function, and Θ qM is the set of parameters required for mapping.
[0206] Furthermore, the mapping function is adjusted by optimizing the following learning objective function, expressed as:
[0207]
[0208] In the formula, w ij is the similarity weight based on q i and q j in the original feature space; q j is the j-th patient data feature vector input to the relevance vector machine.
[0209] Furthermore, the calculation method of the mapping function is expressed as:
[0210]
[0211] In the formula, represents the sample set closest to the i-th patient data feature vector input to the relevance vector machine, and w ij is the contribution weight of the i-th patient data feature vector input to the relevance vector machine to the j-th patient data feature vector input to the relevance vector machine.
[0212] Furthermore, the calculation method of the contribution weight is expressed as:
[0213]
[0214] In the formula, ∥∥ is the L2 norm; q k is the k-th patient data feature vector input to the relevance vector machine.
[0215] 3) In the new feature space, train using the relevance vector machine of the Bayesian framework. By calculating the posterior probability distribution, select the patient data point with the highest posterior probability as the support vector. The calculation method is expressed as:
[0216]
[0217] In the formula, y is the label vector of the input patient data. For example, the label vector includes 0 and 1, corresponding to "high confidence" and "low confidence" respectively; p(y|q′ i ,θ qk ) represents the posterior probability distribution, characterizing the probability predicted based on the parameter θ qk ; θ qk is the kernel function parameter of the relevance vector machine.
[0218] 4) In each iteration, by calculating the energy entropy of the model and optimizing the model parameters, the calculation method of the energy entropy is expressed as:
[0219]
[0220] In the formula, H q represents the energy entropy; Ncs is the number of samples input to the relevance vector machine in the current batch.
[0221] 5) Adaptively adjust the model parameters. In one embodiment, the dynamic update method of the kernel function parameters is expressed as:
[0222]
[0223] In the formula, is the kernel function parameter of the updated relevance vector machine, γ q is the learning rate of the relevance vector machine, represents the gradient of the kernel function parameter of the relevance vector machine.
[0224] Furthermore, the gradient is calculated as:
[0225]
[0226] In the formula, y i is the label of the i-th patient data.
[0227] 6) Repeat the above steps until the preset iteration stop condition is met, which means the model training is completed. In one embodiment, the preset iteration stop condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 500 times.
[0228] After the training of the relevance vector machine based on adaptive feature learning is completed, for each new patient - escort record, extract the feature vector and input it into the trained relevance vector machine for classification, and output "high - confidence relationship" or "low - confidence relationship".
[0229] Furthermore, send the patient relationship data with high - confidence relationship to the doctor to assist in judging the true information of the patient.
[0230] Figure 3 is a schematic diagram of a computer device provided by an embodiment of the present invention. As Figure 3 shown, the device 2000 may include: one or more processors 2010 and one or more memories 2020; wherein, computer - readable code is stored in the memory, and when the computer - readable code is run by the one or more processors, the above - mentioned method can be executed.
[0231] The processor in this embodiment may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, operations, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.
[0232] Generally speaking, the various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When the aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.
[0233] For example, the method or device according to the embodiments of the present disclosure may also be implemented by means of Figure 4 the architecture of the computing device 3000 shown. As Figure 4 shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for the processing and / or communication of the methods provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 4 the architecture shown is only exemplary, and when implementing different devices, one or more components shown in the Figure 4 computing device may be omitted according to actual needs.
[0234] The embodiment of the present invention also provides a computer-readable storage medium, such as Figure 5As shown, it is a schematic diagram of a storage medium 4000 provided by an embodiment of the present invention. Computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the method according to the embodiments of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory for the methods described herein is intended to include but not be limited to these and any other suitable types of memory. It should be noted that the memory for the methods described herein is intended to include but not be limited to these and any other suitable types of memory.
[0235] Embodiments of the present disclosure also provide a computer program product or a computer program, which when executed by a processor implements the steps of the above method, as Figure 2 shown, the computer program product or the computer program includes:
[0236] An acquisition module 201: configured to acquire medical visit information records of patients;
[0237] An extraction module 202: configured to obtain a patient escort triple based on the medical visit information record, and the three elements of the patient escort triple are a patient, an escort relationship, and an escort;
[0238] A matching module 203: configured to match the patient escort triple with a medical visit information knowledge graph to obtain a matching result; wherein, the acquisition method of the medical visit information knowledge graph is:
[0239] Step 1, acquire a medical visit information record data set of patients;
[0240] Step 2, perform knowledge graph construction on the medical visit information record data set to obtain the medical visit information knowledge graph.
[0241] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0242] Generally speaking, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.
[0243] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0244] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other may be indirect couplings or communication connections through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0245] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0246] In addition, each functional unit in various embodiments of the present invention may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0247] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A patient relationship association and information matching method based on knowledge graph, characterized in that: The method comprises: S1: Obtain patient medical information records; S2: obtaining a patient-accompanying person triplet based on the medical consultation information record, wherein the three elements of the patient-accompanying person triplet are patient, accompanying relationship, and accompanying person; S3: Match the patient's accompanying person triple with the medical consultation information knowledge graph to obtain a matching result; wherein the medical consultation information knowledge graph is obtained in the following manner: Step 1: Obtain the patient medical information record data set; Step 2: construct a knowledge graph for the medical information record dataset to obtain the medical information knowledge graph.
2. The patient relationship association and information matching method based on knowledge graph according to claim 1 is characterized in that: The method further includes: S4: performing relationship verification on the matching result to obtain a verified matching result.
3. The patient relationship association and information matching method based on knowledge graph according to claim 1 is characterized in that: The specific matching method in S3 is: vectorizing the triples of the patient's companion to obtain a triple vector, converting the triple vector through an encoder to obtain a low-dimensional space expression of the triple vector; vectorizing all triple sets in the medical information knowledge graph and inputting them into the encoder to obtain a low-dimensional space expression of the knowledge graph triple vector set, matching the low-dimensional space expression of the triple vector with the low-dimensional space expression of the knowledge graph triple vector set one by one to obtain a matching degree, and taking the knowledge graph triple vector with the highest matching degree as the matching result; Optionally, the encoder is an autoencoder of adaptive spatial transformation projection, and the construction method includes: Step 1: Obtain patient relationship triple training set data and mark the real correlation relationship of patient relationship triples; construct an initial autoencoder with a multi-layer structure and initialize the parameters of each layer; Step 2: vectorize the patient relationship triplet to obtain the patient relationship triplet vector; Step 3: The patient relationship triplet vector is transformed by a projection matrix to obtain a reduced-dimensional low-dimensional space; Step 4: The patient relationship triplet vector is input into the autoencoder adjusted by the projection matrix to obtain a compressed low-dimensional space; Step 5: The self-encoder of the adaptive spatial transformation projection is obtained by back-propagating the loss function to update the parameters and iterating until the stopping condition, wherein the loss function is calculated based on the patient relationship triplet vector, the reduced low-dimensional space, the parameters and output of the self-encoder; Optionally, the patient relationship triplet vector is normalized to obtain standardized patient triplet data, which are then input into the spatial transformation module and the autoencoder respectively.
4. The patient relationship association and information matching method based on knowledge graph according to claim 3 is characterized in that: The standardized processing method is expressed as: Where, X a represents the patient relationship triple vector, μ a and Σ a They are X a The mean vector and standard deviation vector of X′ a is the standardized patient triple data; Optionally, the dimension reduction process of the space transformation module is expressed as: or in, is the patient triplet data after dimension reduction, X a Represents the patient relationship triple vector, X′ a is the standardized patient triple data, P r It is a projection matrix that is dynamically optimized based on the input standardized patient triplet data characteristics; Optionally, the projection matrix is calculated by using a principal component analysis method; Optionally, for the standardized patient triple data X′ a Perform covariance matrix decomposition and select the first n ty The eigenvectors form the projection matrix; Optionally, if the standardized patient triplet data X′ a The dimension is n a1 ×n a2 , then the dimension of the projection matrix is n a2 ×n ty ; Optionally, the projection matrix is iteratively optimized based on error back propagation learning to obtain an optimized projection matrix, and the reduced low-dimensional space is obtained using the optimized projection matrix; Optionally, a gradient-based adaptive adjustment term is used to optimize the projection matrix. The optimization method is expressed as: In the formula, is the projection matrix of the tth iteration; is the projection matrix of the t+1th iteration; η r is the projection matrix learning rate, is the loss function of the autoencoder for P r The gradient of r It is an adjustment item; is the local sensitivity gradient adjustment factor of the tth iteration; α r is the regulating factor; Optional, α r Set to 0.2, δ r is the information entropy of the projection matrix; Optionally, the update method of the local sensitivity gradient adjustment factor is expressed as: In the formula, is the local sensitivity gradient adjustment factor for the t+1th iteration; is the local sensitivity gradient adjustment factor of the tth iteration; η tre is the learning rate of the adjustment factor, represents the gradient of the loss function with respect to the projection matrix at the tth iteration, is the L2 norm of the gradient of the loss function with respect to the projection matrix.
5. The patient relationship association and information matching method based on knowledge graph according to claim 4 is characterized in that: The network structure of the autoencoder includes multiple layers, and the parameters of each layer include a weight matrix and a bias term. The initialization of the parameters of each layer is expressed as follows: Assume that the number of layers of the autoencoder is L red The weight matrix and bias of the layer are initialized as follows: In the formula, represents the weight matrix of the lth layer of the autoencoder, is the bias term of the lth layer of the autoencoder; ~ means obeying a specific distribution; represents a normal distribution with a mean of 0 and a variance of 0.01; Optionally, the output of the autoencoder is represented as: In the formula, is the activation output of the encoder layer l; is the activation output of the encoder layer l-1; Sig() is the Sigmoid activation function; is the underestimation correction factor of the tth iteration; represents the weight matrix of the lth layer of the autoencoder, is the bias term of the lth layer of the autoencoder; Optionally, a dynamic adjustment method based on error feedback is used to continuously update the underestimation correction factor. The update method is expressed as: In the formula, are the underestimation correction factors for the t+1th and tth iterations respectively; η λ is the learning rate of the underestimation correction factor, and are the first and second order derivatives of the autoencoder loss function with respect to the underestimation correction factor at the t-th iteration, respectively.
6. The patient relationship association and information matching method based on knowledge graph according to claim 4 is characterized in that: The loss function is expressed as: Where, X a represents the patient relationship triple vector, is the patient triplet data after dimensionality reduction, ||||1 represents the L1 norm; ||||2 represents the L2 norm; α red is the first weighting factor; β red is the second weighting factor; is the weight matrix of the last layer of the autoencoder; is the activation output of the last layer of the encoder; Optionally, the parameter update rule of the autoencoder is: in, is the weight matrix of the lth layer of the autoencoder; η der is the learning rate of the autoencoder; L r is the loss function of the autoencoder; and They are the partial derivatives of the loss function of the autoencoder with respect to weights and biases respectively; Represents the gradient of the loss function with respect to the projection matrix of the tth iteration; is the L2 norm of the gradient of the loss function with respect to the projection matrix; pes is the power adjustment parameter.
7. The patient relationship association and information matching method based on knowledge graph according to claim 2 is characterized in that: The relationship verification is performed using a relevance vector machine algorithm based on adaptive feature learning; a method for obtaining the relevance vector machine algorithm based on adaptive feature learning includes: Step 1: Obtain the training patient-companion relationship dataset and annotate high-confidence and low-confidence data. Step 2: Initialize the kernel function parameters, learning rate, and regularization parameters of the correlation vector machine; Step 3: Quantize the patient data feature vectors in the patient-companion relationship data set to obtain a data feature vector, and perform feature mapping on the data feature vector to obtain a feature vector; Step 4: Performing the relevance vector machine training iteration of the Bayesian framework based on the feature vector space until the stopping condition is reached to obtain the relevance vector machine algorithm for the adaptive feature learning; Optionally, the feature map is expressed as: q′ i =M(q i ,Θ qM ) In the formula, q i is the feature vector of the i-th patient data input to the relevance vector machine; q′ i is the feature vector of the i-th patient after mapping; M() is the mapping function, Θ qM is the set of parameters required for mapping; Optionally, the mapping function is adjusted by optimizing the following learning objective function, expressed as: In the formula, w ij is based on q in the original feature space i and q j Similarity weight of j is the jth patient data feature vector input to the relevance vector machine. Optionally, the mapping function is calculated as: In the formula, represents the sample set of the nearest neighbors of the feature vector of the i-th patient data input to the RVM, w ij is the contribution weight of the i-th patient data feature vector input to the relevance vector machine to the j-th patient data feature vector input to the relevance vector machine; Optionally, the contribution weight is calculated as: In the formula, |||| is the L2 norm; q i ,q j ,q k They represent the i-th, j-th, and k-th patient data feature vectors input to the relevance vector machine, respectively; Optionally, the relevance vector machine of the Bayesian framework selects the patient data point with the highest posterior probability as the support vector by calculating the posterior probability distribution, and the calculation method is expressed as: Where y is the label vector of the input patient data; q′ i is the feature vector of the i-th patient after mapping; p(y|q′ i ,θ qk ) represents the posterior probability distribution, which represents the probability distribution based on the parameter θ qk The predicted probability; θ qk is the kernel function parameter of the correlation vector machine; Optionally, the relevant vector machine of the Bayesian framework calculates the energy entropy of the model and optimizes the parameters of the model in each iteration. The calculation method of the energy entropy is expressed as: In the formula, H q represents energy entropy; Ncs is the number of samples input to the relevant vector machine in the current batch; p(y|q′ i ,θ qk ) represents the posterior probability distribution; Optionally, the correlation vector machine of the Bayesian framework adaptively adjusts model parameters, and the dynamic update method of the kernel function parameters is expressed as follows: In the formula, is the kernel function parameter of the updated relevance vector machine, γ q is the learning rate of the RVM, Represents the gradient of the kernel function parameters of the relevant vector machine; H q represents energy entropy; q represents the regularization parameter; Optional, gradient The calculation method is expressed as: In the formula, y i is the label of the i-th patient data; Ncs is the number of samples input to the relevant vector machine in the current batch; p(y|q′ i ,θ qk ) represents the posterior probability distribution; H q represents energy entropy; q represents the regularization parameter; θ qk is the kernel function parameter of the correlation vector machine.
8. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
A knowledge graph construction method based on Chinese electronic medical records
CN113688255B
Clinical test result review method and system based on artificial intelligence and big data
CN118629571B
Intelligent question answering system and method based on deep semantic understanding and knowledge graph fusion
CN119046421A
Patient doctor-seeing full-process accompanying diagnosis system and method and related equipment
CN116246760A
Small sample knowledge graph completion method based on graph structure information
CN119227790A
Cited By
Equipment fault diagnosis method based on deep learning, storage medium and program product
CN120611275A
Deep learning-based device fault diagnosis methods, storage media, and software products
CN120611275B
Diagnosis accompanying service reservation method and system
CN121054217A
Accompanying diagnosis service reservation method and system
CN121054217B