Cooperative sharing method for medical information of chronic disease management
By constructing a tree-like relationship map and retrieval index encryption technology based on core word vectors, the problem of excessive ciphertext index construction and low retrieval efficiency in the existing technology is solved, and efficient medical information sharing and query experience is achieved.
Patent Information
- Application Number
- CN202510463307.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the prior art uses the hash table index method to establish ciphertext index of health and medical big data, the index constructed is too large, resulting in low retrieval efficiency and inability to perform fuzzy queries, which affects the user experience.
By constructing a tree-like relationship map to reflect the similarity between different word vectors, the accurate classification of similar word vectors is achieved, and the search index is constructed for each cluster class based on the core word vectors, and the search index is encrypted, thereby achieving efficient query of ciphertext data.
It improves the search speed of data in medical information sharing, enhances the user experience, and avoids the problem of excessive hash-based indexes and inefficient retrieval.
Smart Images

Figure CN119993365A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical information sharing, and in particular to a medical information collaborative sharing method for chronic disease management. Background Art
[0002] Chronic disease management is a medical behavior and process for chronic non-communicable diseases, which includes regular detection, continuous monitoring, evaluation and comprehensive intervention management of diseases and their risk factors. The key content of chronic disease management covers early screening, risk prediction, early warning and comprehensive intervention of chronic diseases, as well as comprehensive management and effect evaluation of chronic disease populations. Under the background of "Internet +", the design of the cross-regional medical information sharing and service collaboration platform aims to promote the construction of a regional medical integrated information platform, realize the vertical connection of regional medical resources, information exchange and sharing, and efficient business collaboration, thereby improving the quality and accessibility of regional medical and health services.
[0003] Since the treatment of chronic diseases is a long process, the patient's medical records may exist in multiple hospitals, and most of the patient's medical information is sensitive information. In the process of sharing health and medical big data, it is usually necessary to encrypt and transmit medical data from different sources, and create a ciphertext index to associate the plaintext data with the ciphertext data, so as to realize ciphertext query in data sharing. In the prior art, the hash table index method is a common way to establish a ciphertext index. In the process of using the hash table index method to establish a ciphertext index for health and medical big data, if the hash value is simply used as the index of all word vectors in the data, the constructed index will be too large and the retrieval efficiency will be low. At the same time, fuzzy queries cannot be performed during retrieval, which ultimately affects the user experience. Summary of the invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a method for collaborative sharing of medical information for chronic disease management.
[0005] In a first aspect, the present invention provides a method for collaborative sharing of medical information for chronic disease management, comprising the following steps: Obtain the word vectors in each medical record in the database, and determine the similar word vectors of each word vector based on the similarity between different word vectors; Based on the distribution of similar word vectors of each word vector in all medical records, determine the usage screening index of each word vector; Determine a tree relationship map according to the usage screening index of each word vector and the similarity between different word vectors, wherein the tree relationship map includes nodes corresponding to different word vectors and connections between different nodes; According to the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes connected to it, all word vectors are divided into clusters centered on the core word vector; Constructing a retrieval index for each cluster based on the core word vector, and encrypting each medical record and the retrieval index in the database, respectively, to obtain ciphertext data and encrypted retrieval indexes; Based on the ciphertext data and the encrypted search index, a ciphertext search of the medical records is performed.
[0006] In combination with the first aspect, in some possible implementations, determining similar word vectors for each word vector includes: The similarity between each word vector and any other word vector is obtained, and other word vectors whose similarity is greater than a first set similarity threshold are determined as similar word vectors corresponding to each word vector.
[0007] In combination with the first aspect above, in some possible implementations, determining the usage screening index for each word vector includes: Determine the associated medical records of each word vector in the database, where similar word vectors of the corresponding word vector exist in the associated medical records; Determine the average number of all similar word vectors for each word vector in the associated medical records, normalize the average number, and obtain a usage frequency index for each word vector; Based on the usage frequency index of each word vector and the proportion of medical records associated with each word vector in all medical records in the database, a usage screening index for each word vector is determined, and the usage frequency index and the proportion are both positively correlated with the usage screening index.
[0008] In conjunction with the first aspect above, in some possible implementations, determining the tree relationship graph includes: Determine each node in the tree relationship graph, each of the nodes corresponds to a word vector; According to the usage screening indicators of various word vectors, candidate word vectors are screened out from all the word vectors; Determine the initial parent node word vectors in all the candidate word vectors according to the similarities between the different candidate word vectors, wherein the similarities between the different initial parent node word vectors are all less than a second set similarity threshold; Based on the initial parent node word vector, the similarity between different word vectors, and the usage screening indicators corresponding to different word vectors, the connection between different nodes in the tree relationship graph is determined.
[0009] In combination with the first aspect above, in some possible implementations, determining the connection between different nodes in the tree relationship graph includes: For any initial parent node word vector, connect the initial parent node word vector and the nodes corresponding to its similar word vectors in the tree relationship graph, and determine the next parent node word vector from all similar word vectors of the initial parent node word vector based on the use screening index of each similar word vector of the initial parent node word vector; Connecting the next parent node word vector with the node corresponding to the untraversed similar word vector in the tree relationship graph, and determining the next parent node word vector from all the untraversed similar word vectors of the next parent node word vector based on the use screening index of the untraversed similar word vector of the next parent node word vector; In the tree relationship graph, the next parent node word vector is connected to the node corresponding to its similar word vector that has not been traversed, and new parent node word vectors are continuously determined in the same manner as the above-mentioned determination of the next parent node word vector, and the new parent node word vector is connected to the node corresponding to its similar word vector that has not been traversed in the tree relationship graph until the convergence condition is met.
[0010] In combination with the first aspect above, in some possible implementations, the step of determining other parent node word vectors except the initial parent node vector includes: Determine all usage screening indicators whose corresponding normalized values among the usage screening indicators corresponding to various other word vectors are greater than the first set value and less than the second set value to obtain a candidate usage screening indicator, and determine the maximum value among all the candidate usage screening indicators to obtain a maximum usage screening indicator; The similar word vector corresponding to the maximum usage screening index is used as the parent node word vector.
[0011] In combination with the first aspect above, in some possible implementations, screening out candidate word vectors from all the word vectors includes: Normalize the usage screening indicators of various word vectors to obtain normalized usage screening indicators; The normalized use screening index is judged for size, and a word vector corresponding to a normalized use screening index greater than a first set value and less than a second set value is determined as a candidate word vector.
[0012] In combination with the first aspect above, in some possible implementations, all word vectors are divided into clusters centered around core word vectors, including: Determine the cumulative value of the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes to which it is connected, to obtain the cumulative similarity value; The nodes corresponding to the normalized values of the similarity accumulation values that are greater than the set similarity accumulation threshold are regarded as core nodes, and the word vectors corresponding to the core nodes are determined as core word vectors; Based on the core word vector and the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes connected thereto, all word vectors are divided into clusters centered around the core word vector.
[0013] In combination with the first aspect above, in some possible implementations, all word vectors are divided into clusters centered around core word vectors, including: Determine the distance between different nodes in the tree relationship map based on the similarity between the word vectors corresponding to each node in the tree relationship map and other nodes to which it is connected; Based on the distance value, determining the local reachable density of the core word vector corresponding to each core node in the tree relationship graph; Based on the local reachable density, all word vectors are divided into clusters centered around the core word vector.
[0014] In combination with the first aspect above, in some possible implementations, performing a ciphertext search of the medical records based on the ciphertext data and the search index includes: Obtain word vectors from medical record retrieval data and generate ciphertext retrieval data; Based on the ciphertext retrieval data, querying the encrypted retrieval index for a matching cluster class that matches the word vector in the medical record retrieval data; Based on the ciphertext retrieval data and the ciphertext data in the matching cluster, retrieval result data is obtained.
[0015] In order to solve the above technical problems, in a second aspect, the present invention also provides a medical information collaborative sharing device for chronic disease management, comprising: A similar word vector acquisition module is used to obtain the word vectors in each medical record in the database, and determine the similar word vectors of each word vector based on the similarity between different word vectors; A screening index acquisition module is used to determine the use screening index of each word vector based on the distribution of similar word vectors of each word vector in all medical records; A tree relationship map acquisition module is used to determine a tree relationship map according to the usage screening index of each word vector and the similarity between different word vectors, wherein the tree relationship map includes nodes corresponding to different word vectors and connections between different nodes; A cluster acquisition module, used to divide all word vectors into clusters centered on a core word vector according to the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes to which it is connected; An index building encryption module is used to build a retrieval index for each cluster based on the core word vector, and to encrypt each medical record and the retrieval index in the database, respectively, to obtain ciphertext data and encrypted retrieval indexes; A retrieval module is used to perform ciphertext retrieval of medical records based on the ciphertext data and the encrypted retrieval index.
[0016] To solve the above technical problems, in a third aspect, the present invention further provides a medical information collaborative sharing system for chronic disease management, comprising a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, so that the device executes the method in the above first aspect or any possible implementation of the first aspect.
[0017] To solve the above technical problems, in a fourth aspect, the present invention further provides a computer program product, which includes: a computer program code, when the computer program code runs on a computer, enables the computer to execute the method in the above first aspect or any possible implementation of the first aspect.
[0018] To solve the above technical problems, in a fifth aspect, the present invention further provides a computer-readable storage medium, which stores a computer program code. When the computer program code runs on a computer, the computer executes the method in the above first aspect or any possible implementation of the first aspect.
[0019] The present invention has the following beneficial effects: by constructing a tree-like relationship map to reflect the similarity between different word vectors, accurate classification of similar word vectors is achieved, avoiding the inaccuracy of classification based only on the similarity between different word vectors; in the process of constructing the tree-like relationship map, the usage and specificity of the word vectors are analyzed based on the distribution of similar word vectors of each word vector in all medical records, so as to determine the usage screening index of each word vector to achieve node screening in the tree-like relationship map. A retrieval index is constructed through the constructed tree-like relationship map, and the retrieval index is encrypted, so that the ciphertext data obtained after encrypting each medical record in the database can be queried through the encrypted retrieval index without decryption, effectively improving the retrieval speed of data in medical information sharing, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 A flowchart of the steps of a method for collaborative sharing of medical information for chronic disease management according to an embodiment of the present invention; Figure 2 A flowchart of the steps of determining the use screening index of each word vector according to an embodiment of the present invention; Figure 3 A flowchart of the steps of determining a tree relationship graph according to an embodiment of the present invention; Figure 4 A flowchart of the steps of dividing all word vectors into clusters centered on core word vectors according to an embodiment of the present invention; Figure 5 A flowchart of the steps of performing ciphertext retrieval of medical records according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a medical information collaborative sharing device for chronic disease management according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of a medical information collaborative sharing system for chronic disease management according to an embodiment of the present invention; Among them: 701 represents a memory; 702 represents a processor; 703 represents a computer program. DETAILED DESCRIPTION
[0022] In order to clearly illustrate the technical features of the present invention, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.
[0023] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.
[0024] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0025] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0026] It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0027] Although operations or steps are described in a specific order in the drawings in the embodiments of the present invention, it should not be understood that it is required to perform these operations or steps in the specific order shown or in a serial order, or that all the operations or steps shown must be performed to obtain the desired results. In the embodiments of the present invention, these operations or steps may be performed in series; these operations or steps may also be performed in parallel; or some of these operations or steps may be performed.
[0028] At the same time, it is understood that the data involved in the technical solution of the present invention (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of the relevant laws, regulations and relevant provisions. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meanings as those generally understood by technicians in the technical field of the present invention, and all parameters or indicators in the formulas involved in the present invention are normalized values that eliminate the influence of dimensions.
[0029] In order to solve the problem that in the process of establishing a ciphertext index of health and medical big data using a hash table index method, if the hash value is simply used as the index of all word vectors in the data, the constructed index will be too large and the retrieval efficiency will be low. At the same time, fuzzy query cannot be performed during retrieval, which ultimately affects the user experience, an embodiment of the present invention provides a medical information collaborative sharing method for chronic disease management. The method reflects the similarities between different word vectors by constructing a tree relationship graph, thereby achieving accurate classification of similar word vectors, obtaining clusters centered on a core word vector, and constructing a retrieval index for each cluster based on the core word vector, and encrypting the retrieval index, so that the ciphertext data obtained after encrypting each medical record in the database can be queried through the encrypted retrieval index without decryption, effectively improving the retrieval speed of data in medical information sharing, and thus improving the user experience.
[0030] A medical information collaborative sharing method for chronic disease management provided by an embodiment of the present invention will be described in detail below in conjunction with the accompanying drawings.
[0031] Figure 1 A basic flow chart of a method for collaborative sharing of medical information for chronic disease management provided by an embodiment of the present invention is shown. Figure 1 As shown, the method specifically comprises the following steps: Step S100: Obtain the word vectors in each medical record in the database, and determine the similar word vectors of each word vector based on the similarity between different word vectors.
[0032] Obtain each medical record in the hospital medical information database in the server, and use the word embedding model to extract the word vector of each medical record, so as to obtain all the word vectors in each medical record. In the embodiment of the present invention, the Jieba word segmentation method is used to extract the keywords in each medical record, and the obtained keywords are input into the word vector model, which maps the keywords into a fixed-dimensional vector in a vector space, so as to obtain a vector representation of each keyword, which is the word vector in each medical record.
[0033] In order to improve the efficiency of subsequent data retrieval, similar word vectors are divided into one category, and corresponding retrieval indexes are created based on representative word vectors in each category, so that fast retrieval can be effectively achieved. Therefore, based on the similarity between different word vectors, similar word vectors for each word vector are determined.
[0034] In an embodiment of the present invention, similar word vectors of each word vector are determined, including: obtaining the similarity between each word vector and any other word vector, and determining the similar word vectors corresponding to each word vector from other word vectors whose similarity is greater than a first set similarity threshold. Among them, the embodiment of the present invention adopts cosine similarity as a measure of the similarity between each word vector and any other word vector; the first set similarity threshold can be reasonably set as needed, and the embodiment of the present invention sets the value of the first set similarity threshold to 0.75. If the cosine similarity between any two different word vectors is greater than the first set similarity threshold, the corresponding two different word vectors are set to be similar word vectors.
[0035] Step S200: Determine the usage screening index for each word vector based on the distribution of similar word vectors of each word vector in all medical records.
[0036] Considering that when classifying similar word vectors, all the obtained word vectors cannot be classified simply by the similarity between different word vectors. For example, the similarity between word vector A and word vector B exceeds a certain threshold, and the similarity between word vector B and word vector C also exceeds a certain threshold, but the similarity between word vector A and word vector C does not exceed a certain threshold. At this time, it is impossible to determine whether word vector A and word vector C can be classified into the same category. Therefore, it is also necessary to determine the usage frequency index of each word vector based on the distribution of similar word vectors of each word vector in all medical records. The usage frequency index reflects the usage frequency of similar word vectors of each word vector in the medical records in which they exist. Since it is necessary to screen word vectors according to the usage of word vectors when classifying similar word vectors to construct a tree relationship map, it is not possible to simply analyze the frequency of use of word vectors during screening. It is also necessary to obtain the specific situation of word vectors. For example, if the obtained word vector exists in every case, the current word vector cannot achieve the purpose of screening. Therefore, based on the usage frequency index of each word vector, it is also necessary to consider the distribution of word vectors in the entire case during screening, so as to determine the usage screening index of each word vector. The usage screening index reflects the usage of word vectors. The wider the usage range and the higher the frequency of use, the worse the distinction between word vectors. At the same time, the smaller the usage range and the smaller the frequency of use, the worse the effect of distinguishing word vectors. Subsequently, based on the usage screening index of each word vector and the similarity between different word vectors, word vectors can be screened to determine the tree relationship map, and finally the classification of all word vectors can be achieved.
[0037] In the embodiment of the present invention, Figure 2 As shown in the figure, determine the use and screening indicators of each word vector. The implementation steps include: Step S201: determining the associated medical records of each word vector in the database, wherein a similar word vector of the corresponding word vector exists in the associated medical records; Step S202: determining the average number of all similar word vectors of each word vector in the associated medical records, normalizing the average number, and obtaining a usage frequency index of each word vector; Step S203: Determine a usage screening index for each word vector based on the usage frequency index of each word vector and the percentage of medical records associated with each word vector in all medical records in the database, wherein the usage frequency index and the percentage are both positively correlated with the usage screening index.
[0038] For the above steps, as an example, for any word vector, Take the word vector as an example. The distribution of similar word vectors of the first word vector in each medical record in the database is used to determine the first The associated medical records of the word vector are the records of the first It should be understood that since each word vector is also its corresponding similar word vector, the associated medical records of each word vector also include the medical records with the word vector itself. Similar word vectors of the word vector (including the The average number of word vectors itself) is obtained Frequency index of word vectors: ; in, Indicates Frequency index of word vectors; Indicates The first word vector The first of the related medical records The number of similar word vectors of this word vector; Indicates The total number of medical records associated with the word vector; Represents the normalization function.
[0039] In the above formula, by obtaining the The average number of similar word vectors of the first word vector in the associated medical records is obtained. The frequency of use of the word vector is used to evaluate the The usage of word vectors in the associated medical records.
[0040] At the same time, obtain the The proportion of medical records associated with this word vector in all medical records in the database, and based on this word vector The frequency index and quantity ratio of the word vectors are used to determine the Use screening indicators for word vectors: ; in, Indicates Use screening indicators for seed word vectors; Indicates Frequency index of word vectors; Indicates The total number of medical records associated with the word vector; Represents the total number of all medical records in the database.
[0041] In the above calculation formula, based on The average number of similar word vectors in the associated medical records and the The proportion of medical records associated with this word vector in all medical records in the database is used to determine the The usage filter index of the word vector. The higher the value of the usage filter index, the more The more likely the word vector is to belong to basic medical vocabulary, that is, relatively common medical vocabulary, the worse the effect of word vector differentiation will be; at the same time, the smaller the value of the screening index is, the more likely the word vector is to belong to basic medical vocabulary, that is, relatively common medical vocabulary. The more likely a word vector is to belong to a rare word, the worse the corresponding word vector distinction will be.
[0042] At this point, through the above method, the usage screening index of each word vector can be obtained.
[0043] Step S300: Determine a tree relationship map based on the usage screening index of each word vector and the similarity between different word vectors, wherein the tree relationship map includes nodes corresponding to different word vectors and connections between different nodes.
[0044] In order to improve the retrieval efficiency, based on the usage screening indicators of each word vector determined above and combined with the similarities between different word vectors, a tree relationship graph can be constructed to classify all word vectors.
[0045] In the embodiment of the present invention, Figure 3 As shown, the tree relationship graph is determined, and the implementation steps include: Step S301: determining each node in the tree relationship graph, each of the nodes corresponding to a word vector; Step S302: Filter out candidate word vectors from all the word vectors according to the usage screening indicators of various word vectors; Step S303: determining the initial parent node word vectors in all the candidate word vectors according to the similarities between the different candidate word vectors, wherein the similarities between the different initial parent node word vectors are all less than a second set similarity threshold; Step S304: Based on the initial parent node word vector, the similarity between different word vectors, and the usage screening indicators corresponding to different word vectors, determine the connection between different nodes in the tree relationship graph.
[0046] For the above steps, as an example, each word vector is taken as a node in the tree relationship map, so that each node in the tree relationship map can be determined, and the number of nodes is equal to the number of word vector types. A first set value and a second set value are set in advance, and the usage of various word vectors is compared with the first set value and the second set value to classify and screen the usage of various word vectors to screen out candidate word vectors from all the word vectors, that is: the usage screening indicators of various word vectors are normalized to obtain a normalized usage screening indicator; the normalized usage screening indicator is judged in size, and the word vector corresponding to the normalized usage screening indicator greater than the first set value and less than the second set value is determined as a candidate word vector.
[0047] In an embodiment of the present invention, the maximum usage screening index among the usage screening indexes of various word vectors can be used to normalize the usage screening indexes of various word vectors, that is, the usage screening indexes of various word vectors are divided by the maximum usage screening index, and the quotient obtained is the normalized usage screening index of various word vectors. At the same time, the first setting value is set to 0.3, and the second setting value is set to 0.7, and the size of the normalized usage screening index of various word vectors is judged. If the normalized usage screening index is less than or equal to the first setting value, that is, the normalized usage screening index is between If the normalized filter index is greater than the first setting value and less than the second setting value, the normalized filter index is within the range of If the normalized filter index is greater than or equal to the second setting value, the normalized filter index is within the range of If the word vector is within the range of , it means that the corresponding word vector is frequently used. Considering that the frequently used word vectors may exist in all cases, they are usually not used for retrieval when performing medical information retrieval. For the word vectors that are not frequently used, they may make the constructed graph traversal incomplete when calculating the association. Therefore, the normalized use of the screening index is greater than the first set value and less than the second set value, that is, the normalized use of the screening index is located at All word vectors in the range are used as candidate word vectors to facilitate the subsequent determination of the initial parent node word vector.
[0048] Considering that when constructing a tree relationship graph, it is necessary to ensure that the similarity between the word vectors corresponding to the initial parent node is relatively poor, the initial parent node word vectors in all candidate word vectors are determined based on the similarity between different candidate word vectors, and the similarity between different initial parent node word vectors is less than the second set similarity threshold. Among them, the embodiment of the present invention also uses cosine similarity as a measure of the similarity between different initial parent node word vectors; the second set similarity threshold needs to be less than the value of the first set similarity threshold, and the embodiment of the present invention sets the value of the second set similarity threshold to 0.3.
[0049] On this basis, according to the similarity between all initial parent node word vectors and different word vectors, and combined with the corresponding usage screening indicators of different word vectors, the connections between different nodes in the tree relationship graph can be determined, thereby obtaining a complete tree relationship graph.
[0050] In an embodiment of the present invention, the connection between different nodes in the tree relationship graph is determined, and the implementation steps include: For any initial parent node word vector, connect the initial parent node word vector and the nodes corresponding to its similar word vectors in the tree relationship graph, and determine the next parent node word vector from all similar word vectors of the initial parent node word vector based on the use screening index of each similar word vector of the initial parent node word vector; Connecting the next parent node word vector with the node corresponding to the untraversed similar word vector in the tree relationship graph, and determining the next parent node word vector from all the untraversed similar word vectors of the next parent node word vector based on the use screening index of the untraversed similar word vector of the next parent node word vector; In the tree relationship graph, the next parent node word vector is connected to the node corresponding to its similar word vector that has not been traversed, and new parent node word vectors are continuously determined in the same manner as the above-mentioned determination of the next parent node word vector, and the new parent node word vector is connected to the node corresponding to its similar word vector that has not been traversed in the tree relationship graph until the convergence condition is met.
[0051] For the above steps, in the tree relationship diagram, for any initial parent node word vector, the node corresponding to the initial parent node word vector is used as the initial parent node, and the nodes corresponding to each similar word vector of the initial parent node word vector are used as each child node of the initial parent node, and the initial parent node and its each child node are connected.
[0052] Determine all usage screening indicators whose corresponding normalized values in the usage screening indicators of each similar word vector of the initial parent node word vector are greater than the first set value and less than the second set value, that is, All the usage screening indicators within the range are used to obtain the available usage screening indicators, and the maximum value among all the available usage screening indicators is determined, that is, the maximum usage screening indicator, and the similar word vector corresponding to the maximum usage screening indicator is used as the next parent node word vector of the initial parent node word vector, and the child node of the initial parent node corresponding to the next parent node word vector is used as the parent node of the next traversal, that is, the next parent node. The nodes corresponding to each similar word vector of the next parent node word vector that has not been traversed (the initial parent node vector belongs to the traversed similar word vector of the next parent node word vector) are used as the child nodes of the next parent node, and the next parent node is connected with each of its child nodes.
[0053] Determine all usage screening indicators whose corresponding normalized values among the usage screening indicators of each similar word vector that has not been traversed of the next parent node word vector are greater than the first set value and less than the second set value, that is, located at All the usage screening indicators within the range are used to obtain the available usage screening indicators, and the maximum value among all the available usage screening indicators is determined, that is, the maximum usage screening indicator, and the similar word vector corresponding to the maximum usage screening indicator is used as the next parent node word vector of the initial parent node word vector, and the child node of the next parent node corresponding to the next parent node word vector is used as the parent node of the next traversal, that is, the next parent node. The nodes corresponding to each similar word vector of the next parent node word vector that has not been traversed (the next parent node word vector belongs to the traversed similar word vector of the next parent node word vector) are used as the child nodes of the next parent node, and the next parent node is connected with each of its child nodes.
[0054] Determine all usage screening indicators whose corresponding normalized values in the usage screening indicators of each similar word vector that has not been traversed of the next parent node word vector are greater than the first set value and less than the second set value, that is, located at All the use screening indicators within the range are obtained to obtain the candidate use screening indicators, and the maximum value of all the candidate use screening indicators is determined. In the same way as the above determination of the next parent node word vector, the above steps are repeated to continuously determine the new parent node word vector, and the new parent node word vector is connected to the node corresponding to its untraversed similar word vector in the tree relationship graph until the convergence condition is met. The convergence condition means that the new parent node word vector does not have untraversed similar word vectors or no new connections are generated. Finally, a tree relationship graph is obtained.
[0055] Among them, the normalized value of the usage filtering indicator of each similar word vector of the above-mentioned parent node word vector refers to the normalized usage filtering indicator corresponding to the usage filtering indicator. Since the method of obtaining the normalized usage filtering indicator corresponding to the usage filtering indicator has been introduced in the above content, it will not be repeated here.
[0056] It should be understood that in the process of obtaining the above-mentioned tree-like relationship graph, since an initial parent node word vector can be a similar word vector of a new parent node word vector of another initial parent node word vector, the tree-like relationship graph finally obtained is a tree-like node relationship connection graph.
[0057] At this point, the tree relationship graph corresponding to all word vectors can be obtained.
[0058] Step S400: Based on the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes connected to it, all word vectors are divided into clusters centered around the core word vector.
[0059] Based on the tree relationship graph obtained in the above steps, by selecting the node with the most connections as the representative node in the current cluster, that is, the core node, and building an index through the core word vector corresponding to the representative node, data can be retrieved more quickly, and the relationship between word vectors can be better reflected when performing data mapping.
[0060] In the embodiment of the present invention, Figure 4 As shown in the figure, all word vectors are divided into clusters centered on the core word vector. The implementation steps include: Step S401: Determine the cumulative value of the similarity between each node in the tree relationship graph and the word vectors corresponding to other nodes connected to it, and obtain the cumulative similarity value; Step S402: The node corresponding to the normalized value of the similarity accumulation value greater than the set similarity accumulation threshold is regarded as a core node, and the word vector corresponding to the core node is determined as the core word vector.
[0061] Step S403: Based on the core word vector and the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes connected to it, all word vectors are divided into clusters centered on the core word vector.
[0062] For the above steps, as an example, for each node in the tree relationship graph (including all parent nodes and child nodes), determine the cumulative value of the similarity between the word vectors corresponding to each node and other nodes with which it is connected, and obtain the similarity cumulative value. Similarly, the similarity between the word vectors corresponding to each node and other nodes with which it is connected is measured by cosine similarity. The similarity cumulative value is normalized to obtain a normalized value of the similarity cumulative value. The normalization method can be reasonably selected as needed. The embodiment of the present invention uses the norm function to normalize the similarity cumulative value.
[0063] A similarity accumulation threshold is set in advance, and the normalized value of the similarity accumulation value corresponding to each node is compared with the set similarity accumulation threshold, and the node corresponding to the normalized similarity accumulation value greater than the set similarity accumulation threshold is taken as the core node, and the word vector corresponding to the core node is taken as the core word vector, which can replace each word vector corresponding to the node in the local area of a tree relationship graph. The similarity accumulation threshold can be reasonably set as needed. The embodiment of the present invention sets the value of the similarity accumulation threshold to 0.8.
[0064] After determining the core word vectors corresponding to each core node in the tree relationship graph, based on the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes connected to it, the entire tree relationship graph is segmented by obtaining the core nodes, and all nodes are divided into clusters corresponding to the nodes corresponding to the core word vectors, thereby achieving the segmentation of all original word vectors into clusters centered on the core word vector (one word vector can exist in multiple clusters).
[0065] In the embodiment of the present invention, all word vectors are divided into clusters centered on the core word vector, and the implementation steps include: Step S4031: determining the distance values between different nodes in the tree relationship graph based on the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes to which it is connected; Step S4032: Based on the distance value, determine the local reachability density of the core word vector corresponding to each core node in the tree relationship graph; Step S4033: Based on the local reachable density, all word vectors are divided into clusters centered on the core word vector.
[0066] For the above steps, as an example, based on the core nodes and their corresponding core word vectors in the tree relationship map determined above, the local reachable density is calculated with the core word vector as the center point. In the process of obtaining the local reachable density of the core word vector, the negative correlation mapping result of the cosine similarity of the word vectors corresponding to the two nodes with a connection in the tree relationship map is used as the distance between the corresponding word vectors, that is, the difference between the value 1 and the cosine similarity is used as the negative correlation mapping result of the cosine similarity. It should be noted that if there is no direct connection between the word vectors corresponding to the two nodes in the tree relationship map, but one node can reach another node through multiple intermediate nodes, then the cumulative sum of the distances corresponding to all the connections on each connection path of the two nodes is determined as a distance between the two nodes. If there are multiple connection paths between the two nodes, multiple distances corresponding to the two nodes can be obtained, and the smallest distance is selected as the final distance between the two nodes. For example, there is no direct connection between node a and node d in the tree relationship map, but there is a connection between node a and node b, a connection between node b and node c, and a connection between node c and node d. Node b and node c are the intermediate nodes of node a and node d. The connection between node a and node b, the connection between node b and node c, and the connection between node c and node d constitute a connection path between node a and node d. Based on the distance between different nodes in the obtained tree relationship map, the local reachable density of the core word vector corresponding to each core node can be determined. Since the process of obtaining the local reachable density belongs to the prior art, it will not be repeated here.
[0067] A local reachable density threshold is set in advance, and the value of the local reachable density threshold can be reasonably set as needed. The present invention sets the value of the local reachable density threshold to 0.75, and compares the local reachable density of the core word vector corresponding to each core node in the tree relationship map with the local reachable density threshold. When the local reachable density is greater than the local reachable density threshold for the first time, the core word vector will correspond to a word vector range composed of the core word vector and other word vectors. All word vectors in the word vector range constitute a cluster centered on the core word vector. In this way, all word vectors can be divided into clusters centered on the core word vector, and each core word vector corresponds to a cluster.
[0068] Step S500: construct a retrieval index for each cluster based on the core word vector, and encrypt each medical record and the retrieval index in the database respectively, to obtain ciphertext data and encrypted retrieval index accordingly.
[0069] After determining each cluster class centered on each core word vector through the above steps, a retrieval index can be constructed for each cluster class based on the core word vector. The constructed retrieval index can be a tree structure, such as a B+ tree or an R tree. In an embodiment of the present invention, a hash table index method is used to construct a retrieval index for each cluster class, that is, by obtaining the hash value of each core word vector, and constructing a retrieval index for each cluster class based on the hash value. Since the specific implementation process of constructing a retrieval index for each cluster class based on the core word vector using the hash table index method belongs to the prior art, it will not be repeated here. At this point, the construction of the retrieval index for all medical records in the database is completed.
[0070] All medical records in the database are encrypted to obtain the ciphertext data of all medical records, which includes the word vector ciphertext corresponding to the word vector in each medical record in the database. At the same time, the constructed retrieval index is also encrypted to obtain an encrypted retrieval index. The encryption method used to encrypt all medical records in the database and the encryption method used to encrypt the constructed retrieval index can be reasonably selected according to needs, and will not be repeated here.
[0071] Step S600: Perform ciphertext retrieval of medical records based on the ciphertext data and the encrypted retrieval index.
[0072] When it is necessary to query the patient's medical information, the corresponding query authority is obtained on the user side. Different users, such as patients, doctors and hospitals, have different query authorities and can query patient information in different types of databases in the server. When the query authority of a certain database is obtained, the search information is entered on the user side, and the encrypted search of the medical record can be performed based on the ciphertext data and encrypted search index.
[0073] In the embodiment of the present invention, Figure 5 As shown, based on the ciphertext data and the encrypted search index, the ciphertext search of the medical records is performed, and the implementation steps include: Step S601: Obtain word vectors in the medical record retrieval data and generate ciphertext retrieval data; Step S602: Based on the ciphertext retrieval data, searching the encrypted retrieval index for a matching cluster class that matches the word vector in the medical record retrieval data; Step S603: Acquire search result data based on the ciphertext search data and the ciphertext data in the matching cluster.
[0074] For the above steps, as an example, the user end obtains each word vector in the medical record retrieval data in the same way as the word vector in each medical record in the database, and obtains the hash value of each word vector, uses the hash value as the ciphertext retrieval data of each word vector, and sends the ciphertext retrieval data to the server. Based on the ciphertext retrieval data of each word vector, the server searches the encrypted retrieval index for a matching cluster that matches the corresponding word vector of the medical record retrieval data, that is, the matching cluster that is closest to the corresponding word vector in the medical record retrieval data. When querying, the cosine similarity between the hash value of each word vector in the medical record retrieval data and the hash value in the encrypted retrieval index is calculated. The larger the cosine similarity, the closer the match is. The ciphertext retrieval data is matched with the word vector ciphertext in the matching cluster, that is, the hash value of the word vector. The matching is to calculate the cosine similarity of the two hash values. The larger the cosine similarity, the closer the match is. The medical record ciphertext where the first set number (such as 3) of the word vector ciphertexts that match the front is used as the retrieval data. In this way, the search data corresponding to each word vector in the medical record search data can be obtained, and these search data constitute the search result data of the medical record search data. The server sends the search result data of the medical record search data to the user end.
[0075] After receiving all the final search data of the medical record search data, the user terminal decrypts all the final search data and displays it on the display screen.
[0076] Based on the same inventive concept, Figure 6 As shown, an embodiment of the present invention further provides a medical information collaborative sharing device for chronic disease management, the device comprising: A similar word vector acquisition module is used to obtain the word vectors in each medical record in the database, and determine the similar word vectors of each word vector based on the similarity between different word vectors; A screening index acquisition module is used to determine the use screening index of each word vector based on the distribution of similar word vectors of each word vector in all medical records; A tree relationship map acquisition module is used to determine a tree relationship map according to the usage screening index of each word vector and the similarity between different word vectors, wherein the tree relationship map includes nodes corresponding to different word vectors and connections between different nodes; A cluster acquisition module, used to divide all word vectors into clusters centered on a core word vector according to the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes to which it is connected; An index building encryption module is used to build a retrieval index for each cluster based on the core word vector, and to encrypt each medical record and the retrieval index in the database, respectively, to obtain ciphertext data and encrypted retrieval indexes; A retrieval module is used to perform ciphertext retrieval of medical records based on the ciphertext data and the encrypted retrieval index.
[0077] It should be noted that the device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above.
[0078] Based on the same inventive concept, the embodiment of the present invention also provides a medical information collaborative sharing system for chronic disease management. Figure 7 As shown, the sharing system includes: a memory 701, a processor 702, and a computer program 703 stored in the memory 701 and running on the processor 702, wherein when the processor 702 executes the computer program 703, the system can execute any one of the medical information collaborative sharing methods for chronic disease management introduced above.
[0079] The embodiment of the present invention can divide the system into functional modules according to the above method example. For example, each functional module can be corresponded to, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. There may be other division methods in actual implementation.
[0080] Based on the same inventive concept, an embodiment of the present invention also provides a computer program product, which includes: computer program code, when the computer program code runs on a computer, enables the computer to execute any one of the medical information collaborative sharing methods for chronic disease management introduced above.
[0081] Based on the same inventive concept, an embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program code. When the computer program code runs on a computer, the computer executes any one of the medical information collaborative sharing methods for chronic disease management introduced above.
[0082] It should be noted that the above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for collaborative sharing of medical information for chronic disease management, characterized in that: The following steps are involved: Obtain the word vectors in each medical record in the database, and determine the similar word vectors of each word vector based on the similarity between different word vectors; Based on the distribution of similar word vectors of each word vector in all medical records, determine the usage screening index of each word vector; Determine a tree relationship map according to the usage screening index of each word vector and the similarity between different word vectors, wherein the tree relationship map includes nodes corresponding to different word vectors and connections between different nodes; According to the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes connected to it, all word vectors are divided into clusters centered on the core word vector; Constructing a retrieval index for each cluster based on the core word vector, and encrypting each medical record and the retrieval index in the database, respectively, to obtain ciphertext data and encrypted retrieval indexes; Based on the ciphertext data and the encrypted search index, a ciphertext search of the medical records is performed.
2. A method for collaborative sharing of medical information for chronic disease management according to claim 1, characterized in that: Determine similar word vectors for each word vector, including: The similarity between each word vector and any other word vector is obtained, and other word vectors whose similarity is greater than a first set similarity threshold are determined as similar word vectors corresponding to each word vector.
3. A method for collaborative sharing of medical information for chronic disease management according to claim 1, characterized in that: Determine the use of screening indicators for each word vector, including: Determine the associated medical records of each word vector in the database, where similar word vectors of the corresponding word vector exist in the associated medical records; Determine the average number of all similar word vectors for each word vector in the associated medical records, normalize the average number, and obtain a usage frequency index for each word vector; According to the usage frequency index of each word vector and the proportion of the number of medical records associated with each word vector in all medical records in the database, the usage screening index of each word vector is determined, and the usage frequency index and the proportion of the number are positively correlated with the usage screening index.
4. A method for collaborative sharing of medical information for chronic disease management according to claim 2, characterized in that: Determine the tree relationship diagram, including: Determine each node in the tree relationship graph, each of the nodes corresponds to a word vector; According to the usage screening indicators of various word vectors, candidate word vectors are screened out from all the word vectors; Determine the initial parent node word vectors in all the candidate word vectors according to the similarities between the different candidate word vectors, wherein the similarities between the different initial parent node word vectors are all less than a second set similarity threshold; Based on the initial parent node word vector, the similarity between different word vectors, and the usage screening indicators corresponding to different word vectors, the connection between different nodes in the tree relationship graph is determined.
5. A method for collaborative sharing of medical information for chronic disease management according to claim 4, characterized in that: Determining the connection between different nodes in the tree relationship graph includes: For any initial parent node word vector, connect the initial parent node word vector and the nodes corresponding to its similar word vectors in the tree relationship graph, and determine the next parent node word vector from all similar word vectors of the initial parent node word vector based on the use screening index of each similar word vector of the initial parent node word vector; Connecting the next parent node word vector with the node corresponding to the untraversed similar word vector in the tree relationship graph, and determining the next parent node word vector from all the untraversed similar word vectors of the next parent node word vector based on the use screening index of the untraversed similar word vector of the next parent node word vector; In the tree relationship graph, the next parent node word vector is connected to the node corresponding to its similar word vector that has not been traversed, and new parent node word vectors are continuously determined in the same manner as the above-mentioned determination of the next parent node word vector, and the new parent node word vector is connected to the node corresponding to its similar word vector that has not been traversed in the tree relationship graph until the convergence condition is met.
6. A method for collaborative sharing of medical information for chronic disease management according to claim 5, characterized in that: The steps for determining other parent node word vectors except the initial parent node vector include: Determine all usage screening indicators whose corresponding normalized values among the usage screening indicators corresponding to various other word vectors are greater than the first set value and less than the second set value to obtain a candidate usage screening indicator, and determine the maximum value among all the candidate usage screening indicators to obtain a maximum usage screening indicator; The similar word vector corresponding to the maximum usage screening index is used as the parent node word vector.
7. A method for collaborative sharing of medical information for chronic disease management according to claim 4, characterized in that: Filter out candidate word vectors from all the word vectors, including: Normalize the usage screening indicators of various word vectors to obtain normalized usage screening indicators; The normalized usage screening index is judged to be large or small, and the word vector corresponding to the normalized usage screening index being greater than the first set value and less than the second set value is determined as a candidate word vector.
8. A method for collaborative sharing of medical information for chronic disease management according to claim 2, characterized in that: Divide all word vectors into clusters centered on the core word vector, including: Determine the cumulative value of the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes to which it is connected, to obtain the cumulative similarity value; The nodes corresponding to the normalized values of the similarity accumulation values that are greater than the set similarity accumulation threshold are regarded as core nodes, and the word vectors corresponding to the core nodes are determined as core word vectors; Based on the core word vector and the similarity between the word vectors corresponding to each node in the tree relationship graph and other nodes connected thereto, all word vectors are divided into clusters centered around the core word vector.
9. A method for collaborative sharing of medical information for chronic disease management according to claim 8, characterized in that: Divide all word vectors into clusters centered on the core word vector, including: Determine the distance between different nodes in the tree relationship map based on the similarity between the word vectors corresponding to each node in the tree relationship map and other nodes to which it is connected; Based on the distance value, determining the local reachable density of the core word vector corresponding to each core node in the tree relationship graph; Based on the local reachable density, all word vectors are divided into clusters centered around the core word vector.
10. A method for collaborative sharing of medical information for chronic disease management according to claim 1, characterized in that: Based on the ciphertext data and the search index, a ciphertext search of the medical records is performed, including: Obtain word vectors from medical record retrieval data and generate ciphertext retrieval data; Based on the ciphertext retrieval data, querying the encrypted retrieval index for a matching cluster class that matches the word vector in the medical record retrieval data; Based on the ciphertext retrieval data and the ciphertext data in the matching cluster, retrieval result data is obtained.
Citation Information
Patent Citations
Multi-keyword ciphertext sorting retrieval method based on keyword grouping reverse indexes
CN111966778A
Medical text label identification method and system based on medical knowledge conceptual graph
CN114168751A
Dynamic searchable encryption method for massive high-dimensional medical data
CN115422432A
Cloud computing technology atlas construction method, device, system, equipment and medium
CN116257635A
Electronic medical record nested named entity identification method
CN118821778A