Data efficient retrieval method based on neural network
By employing a neural network-based efficient data retrieval method, utilizing knowledge distillation and medical terminology knowledge graphs to generate lightweight models, and combining symptom-sequential attention mechanisms and dynamic routing algorithms, the accuracy and efficiency issues of electronic health record data retrieval are resolved, enabling rapid and accurate case matching and treatment plan recommendations.
Patent Information
- Application Number
- CN202511088556.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing electronic health record data retrieval methods struggle to deeply explore semantic relationships and potential features, resulting in inaccurate and inefficient retrieval, especially when faced with massive amounts of data and complex disease types, failing to meet the requirements for efficient retrieval.
We employ a neural network-based efficient data retrieval method. Through knowledge distillation, we compress a pre-trained biomedical language model into a lightweight model and inject entity embeddings from a medical terminology knowledge graph. Combining a symptom-series attention mechanism and a dynamic routing algorithm, we generate feature vectors of past health records, construct an inverted index and an approximate nearest neighbor index, and dynamically update the model to adapt to newly added rare cases.
It improves the accuracy and efficiency of retrieval, enabling the rapid identification of similar cases and the provision of effective treatment plans. It adapts to dynamic changes in data, ensuring the timeliness and comprehensiveness of retrieval results.
Smart Images

Figure CN120929649A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data retrieval, specifically a method for efficient data retrieval based on neural networks. Background Technology
[0002] With the development of electronic health records, a large amount of electronic health record data has been accumulated. This data is not only an important reference for clinical decision-making but also the foundation for medical research and disease management. Currently, keyword matching is commonly used as a retrieval method. By inputting information such as the patient's symptoms and health status assessment, the database searches for records containing these keywords. However, keyword matching often only captures surface information in past health records, making it difficult to delve into the semantic relationships and potential features between these records. Furthermore, symptom descriptions in past health records may vary from person to person, containing synonyms, near-synonyms, or vague descriptions, further increasing the difficulty and inaccuracy of the search.
[0003] To overcome these limitations, some advanced retrieval technologies have been introduced into the medical field, improving retrieval accuracy and efficiency to some extent, but some problems still exist. For example, while text vector-based methods can capture semantic similarity between past health records, they are often limited by the dimension of the vectors and the calculation methods, failing to fully reflect the complexity and diversity of past health records. Machine learning-based methods, on the other hand, require large amounts of labeled data and computational resources, and place high demands on the model's generalization ability and robustness. Furthermore, with the continuous advancement of medical technology and the increasing number of diseases, the number and complexity of cases in past health record databases are also growing rapidly. Existing retrieval technologies often struggle to simultaneously meet the requirements of efficient retrieval, resulting in problems such as inaccurate retrieval and low efficiency in practical applications. Faced with a massive database of past health records, how to quickly and accurately retrieve historical cases highly similar to current cases has become a key issue in improving the quality and efficiency of medical services.
[0004] Therefore, there is an urgent need for an efficient retrieval method that uses neural networks and comprehensively considers past health record data to improve retrieval accuracy and efficiency. Summary of the Invention
[0005] (1) Technical problems to be solved
[0006] This invention provides a neural network-based efficient data retrieval method to address the problem that hospitals struggle to meet the requirements of efficient retrieval when faced with massive amounts of past health records. This results in the inability to deeply explore the potential features between past health records in practical applications, and the inability to ensure the accuracy and timeliness of retrieval results.
[0007] (2) Technical solution
[0008] To achieve the above objectives, the present invention provides a method for efficient data retrieval based on neural networks, the method comprising:
[0009] S1. The pre-trained biomedical language model is compressed into a Transformer architecture with a preset number of layers using the knowledge distillation method and denoted as a lightweight biomedical language model. Entity embeddings from a medical terminology knowledge graph are then injected to generate feature vectors of past health records.
[0010] S2. High-frequency cases are obtained from the feature vector of previous health records and stored in the inverted index cache layer based on the joint key value of symptom-health status assessment. Long-tail cases are obtained from the feature vector of previous health records and constructed into a secondary approximate nearest neighbor index after product quantization compression. The query semantic complexity score is obtained, and the query semantic complexity score is used to dynamically select the retrieval path through a dynamic routing algorithm to obtain preliminary retrieval results.
[0011] S3. When the cumulative number of newly added rare disease cases exceeds the preset threshold for the number of rare disease cases, the lightweight biomedical language model is updated, and the feature vectors of the previous health records of the newly added rare disease cases are seamlessly inserted into the secondary approximate nearest neighbor index using the streaming graph index expansion method.
[0012] S4. Obtain the similarity score between the query case and the candidate cases; obtain the treatment effectiveness feedback of historical cases in the preliminary search results; obtain the temporal information in the feature vector of past health records; introduce the symptom temporal attention mechanism into the preliminary search results based on the temporal information; adjust the similarity score according to the symptom temporal decay factor and symptom severity weight of the symptom temporal attention mechanism; and reorder the preliminary search results in combination with the treatment effectiveness feedback to obtain the final search results.
[0013] Furthermore, the method of compressing the pre-trained biomedical language model into a Transformer architecture with a preset number of layers using knowledge distillation and denoting it as a lightweight biomedical language model, and injecting entity embeddings from a medical terminology knowledge graph to generate feature vectors of past health records includes:
[0014] Entity embedding alignment of a medical terminology knowledge graph is performed using a knowledge distillation loss function, followed by cross-entropy loss. The knowledge distillation loss function is obtained by linearly combining the entity embedding alignment loss with the knowledge distillation loss function. The knowledge distillation loss function The calculation formula is:
[0015] .
[0016] in, These are preset loss weight coefficients used to balance the cross-entropy loss. The relative importance between the entity embedding alignment loss and the actual entity embedding alignment loss; A set of symptom entities extracted from a medical terminology knowledge graph; Indicates the entity embedding alignment loss; Entity embedding vectors generated for a pre-trained biomedical language model. Entity embedding vectors generated for a lightweight biomedical language model.
[0017] Furthermore, the method for obtaining long-tail cases from the feature vectors of past health records, performing product quantization compression, and then constructing a second-level approximate nearest neighbor index includes:
[0018] The product quantization employs bit-width selection and residual compensation. Bit-width selection involves dividing the feature vectors of long-tailed cases' past health records into a predetermined number of subspaces before scalar quantization, and then calculating the similarity values before and after scalar quantization and residual compensation. The similarity value The calculation formula is:
[0019] .
[0020] in, Scalar quantization of similarity weights; This represents the similarity value of the feature vectors of past health records after scalar quantization. For the pre-defined set of key symptoms The exact match results for all symptom terms are summed. and These represent the query cases and candidate cases within the preset set of key symptoms, respectively. The first in Quantitative values for each symptom term; This is a function for exact matching of symptom terms. It returns 1 when two symptom terms match exactly, and 0 otherwise.
[0021] When the similarity value after residual compensation is less than the preset similarity threshold, the bit width is adjusted and the similarity value after residual compensation is recalculated until the similarity value after residual compensation is greater than the preset similarity threshold, and then a secondary approximate nearest neighbor index is constructed.
[0022] Furthermore, the method of adjusting the similarity score based on the symptom temporal decay factor and symptom severity weight of the symptom temporal attention mechanism, and re-ranking the preliminary search results in conjunction with treatment effectiveness feedback to obtain the final search results includes:
[0023] The symptom time decay factor The calculation formula is:
[0024] .
[0025] in, This indicates the number of days between the onset of symptoms and the current time. To preset the maximum effective time window, This is the decay rate coefficient.
[0026] The initial weight matrix is obtained by performing a Hadamard product operation on the symptom time decay factor and the symptom severity weight. The initial matrix is then normalized using Softmax to generate an attention weight matrix. The similarity score is adjusted based on the attention weight matrix. The symptom severity weight is obtained by acquiring medical knowledge and expert experience data and performing data statistical analysis. The similarity score is obtained by calculating the similarity between the query case and the candidate case using cosine similarity.
[0027] Furthermore, the method also includes:
[0028] The decay rate coefficient is determined by a pre-set clinical rule base. Associating and mapping with disease types; the preset clinical rule base is: the attenuation rate coefficient of cardiovascular diseases. =0.25±0.05, corresponding to symptoms of acute illnesses with high time sensitivity; decay rate coefficient of chronic diseases. =0.15±0.03, corresponding to persistent symptoms; attenuation rate coefficient for infectious diseases. =0.20±0.05, corresponding to symptoms of a rapidly developing and changing disease.
[0029] When multiple disease types are detected to coexist, the decay rate coefficient corresponding to each disease is obtained. The weighted average of the values.
[0030] Furthermore, the method for updating the lightweight biomedical language model when the cumulative number of newly detected rare disease cases exceeds a preset rare disease case threshold includes:
[0031] When monitoring new cases entering the case database, the health status assessment label and ICD code of the new cases are obtained; the cumulative number of rare disease cases is counted in real time based on the health status assessment label and ICD code; when the cumulative number of rare disease cases exceeds the preset cumulative number of cases threshold, the lightweight biomedical language model is updated through the elastic weight consolidation algorithm.
[0032] Furthermore, the method for obtaining a query semantic complexity score, and then dynamically selecting a retrieval path using a dynamic routing algorithm to obtain preliminary retrieval results includes:
[0033] The semantic complexity score is obtained by performing feature parsing on the input query information. The semantic complexity score The calculation formula is:
[0034] .
[0035] in, To query the number of medical entities contained in the information, The average semantic association between medical entities; Image feature dimension; This is a preset adjustment factor.
[0036] The query semantic complexity score is used to dynamically select the retrieval path through a dynamic routing algorithm: the retrieval path includes a cache layer index, a quantization layer index, or a full index deep retrieval.
[0037] Obtain the first semantic complexity score threshold and the second semantic complexity score threshold. When the semantic complexity score is less than the first semantic complexity score threshold, query the cache layer index. When the semantic complexity score is greater than or equal to the first semantic complexity score threshold and less than the second semantic complexity score threshold, query the cache layer index and the quantization layer index. When the semantic complexity score is greater than or equal to the second semantic complexity score threshold, trigger a full index deep search.
[0038] Furthermore, the aforementioned Entity embedding vectors generated for a pre-trained biomedical language model. Methods for generating entity embedding vectors for lightweight biomedical language models include:
[0039] During entity embedding alignment, symptom description information from a pre-trained biomedical language model and a lightweight biomedical language model is acquired and named entity recognition is performed. The identified entities are then mapped to knowledge graph nodes, and neighboring node information is aggregated using a graph attention network to generate entity embedding vectors. and .
[0040] (3) Beneficial effects
[0041] Compared with the prior art, the beneficial effects of the present invention are:
[0042] 1. A pre-trained biomedical language model is compressed into a lightweight model using knowledge distillation, and entity embedding from a medical terminology knowledge graph is combined to effectively generate feature vectors for past health records. This not only reduces computational complexity and improves retrieval speed but also enhances the model's semantic understanding capabilities by introducing a medical terminology knowledge graph, thereby improving retrieval accuracy. Furthermore, dynamically selecting the retrieval path based on query semantic complexity and re-ranking the retrieval results by incorporating symptom-sequence attention mechanisms and treatment effectiveness feedback further improves the relevance and practicality of the retrieval results.
[0043] 2. When the number of newly added rare disease cases exceeds a preset threshold, the lightweight biomedical language model is automatically updated, and a streaming graph index expansion method is used to seamlessly insert the new cases into the index, ensuring the timeliness and comprehensiveness of the retrieval. For long-tail cases and complex queries, efficient processing and fast response are achieved through product quantization compression and dynamic routing algorithms. Attached Figure Description
[0044] Figure 1 This is a flowchart of a neural network-based efficient data retrieval method according to Embodiment 1 of the present invention. Detailed Implementation
[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0046] Before providing examples, it's necessary to explain the application scenarios of this invention. This invention is a neural network-based efficient data retrieval method. In the medical field, the vast amounts of historical health record data accumulated by medical institutions can help doctors assess clinical health conditions. For instance, when doctors encounter complex cases, they can quickly find similar historical cases through retrieval, providing a reference for health assessment. However, current retrieval methods are slow when processing large-scale historical health record data, failing to meet doctors' immediate needs and easily overlooking cases with similar symptoms but different descriptions. Furthermore, with the continuous emergence of new cases and changes in disease characteristics, traditional retrieval systems cannot update their internal models and index structures in a timely manner, leading to a decline in the timeliness and accuracy of retrieval results. In summary, the neural network-based efficient data retrieval scheme can solve the challenges faced in retrieving historical health record data, improving retrieval efficiency and accuracy, and adapting to dynamic data changes.
[0047] Example 1: As Figure 1As shown, this embodiment provides a method for efficient data retrieval based on neural networks, the method including:
[0048] S1. The pre-trained biomedical language model is compressed into a Transformer architecture with a preset number of layers using knowledge distillation, and denoted as a lightweight biomedical language model. Entity embeddings from a medical terminology knowledge graph are then injected to generate feature vectors from past health records. In this embodiment, the preset number of layers in the Transformer architecture is 6 layers, and the feature vectors from past health records balance semantic depth and computational efficiency. Knowledge distillation is a model compression technique that uses a teacher model (i.e., the pre-trained biomedical language model) to guide the learning of a student model (i.e., the lightweight biomedical language model). This reduces the computational complexity and memory usage of the model without significantly sacrificing performance. Simultaneously with model compression, entity embeddings from a medical terminology knowledge graph are injected. A knowledge graph is a structured knowledge base containing a large number of entities, attributes, and relationships. In this embodiment, the medical terminology knowledge graph is used to provide additional semantic information to enhance the lightweight biomedical language model's understanding of past health record text. Through entity embedding, lightweight biomedical language models can map key information such as symptoms and diseases from past health records to corresponding nodes in a knowledge graph, thereby leveraging the rich semantic information within the knowledge graph for reasoning and retrieval. For example, given the past health record text "The patient presented with persistent fever, cough, and difficulty breathing," the lightweight biomedical language model can identify symptom entities such as "fever," "cough," and "difficulty breathing," and map them to corresponding nodes in the knowledge graph, thus utilizing the semantic information in the knowledge graph to enhance the understanding and processing capabilities of past health record texts.
[0049] S2. High-frequency cases are extracted from the feature vectors of past health records and stored in an inverted index cache layer based on the symptom-health status assessment joint key value. Long-tail cases are extracted from the feature vectors of past health records and constructed into a secondary approximate nearest neighbor index after product quantization compression. The query semantic complexity score value is obtained, and the query semantic complexity score value is used to dynamically select the retrieval path for retrieval through a dynamic routing algorithm to obtain preliminary retrieval results. High-frequency cases refer to cases that occur frequently and have high retrieval value. The symptom-health status assessment joint key value combines symptom description and health status assessment results as index key values, which can more accurately locate related cases. Using inverted index technology, high-frequency cases are indexed by the symptom-health status assessment joint key value and stored in the cache layer to achieve sub-millisecond retrieval for common diseases (such as diabetes and hypertension). The inverted index allows for the rapid retrieval of relevant past health records based on symptom description and health status assessment results, thereby improving retrieval speed. For example, in the past health record database, "fever + cough + pneumonia" is a high-frequency case combination. By constructing an inverted index, all past health records containing these symptoms and health status assessments can be quickly found. Long-tail cases refer to cases with low frequency of occurrence. Due to their small number and wide distribution, they are difficult to retrieve efficiently using traditional indexing methods. Product quantization compression performs product quantization on the feature vectors of past health records for long-tail cases to reduce storage space and computational complexity while maintaining a certain level of retrieval accuracy. Based on product quantization compression, a secondary approximate nearest neighbor index is built to support efficient similarity search. This is suitable for handling high-dimensional feature vectors, i.e., feature vectors of past health records, and can significantly improve retrieval speed at the expense of some accuracy. For example, for a rare genetic disease, due to its small number of cases and complex symptoms, product quantization compression of its past health record feature vectors and the construction of a secondary approximate nearest neighbor index ensure that even if a completely matching case cannot be found during retrieval, similar cases can be found as references. Queries with high semantic complexity scores require deeper retrieval paths to obtain accurate results, while queries with low semantic complexity scores can obtain results quickly through shallower retrieval paths.
[0050] S3. When the cumulative number of newly added rare disease cases exceeds a preset threshold, the lightweight biomedical language model is updated, and the feature vectors of past health records for the newly added rare disease cases are seamlessly inserted into the secondary approximate nearest neighbor index using a streaming graph indexing expansion method. When the cumulative number of newly added rare disease cases reaches a certain threshold, it indicates that new biomedical knowledge needs to be integrated. The lightweight biomedical language model is retrained using this new case data (including symptom descriptions, health status assessment results, treatment plans, etc.) to better understand and process this new rare disease information. The streaming graph indexing expansion method is an efficient indexing expansion technique for handling large-scale data streams. It allows the feature vectors of newly added rare disease cases to be seamlessly inserted into the existing secondary approximate nearest neighbor index dynamically without affecting the stability of the existing index structure. At this point, this new case data can be immediately used for retrieval and analysis.
[0051] S4. Obtain the similarity score between the query case and candidate cases; obtain treatment effectiveness feedback from historical cases in the preliminary search results; obtain the temporal information in the feature vectors of past health records, introduce a symptom temporal attention mechanism into the preliminary search results based on the temporal information, adjust the similarity score according to the symptom temporal decay factor and symptom severity weight of the symptom temporal attention mechanism, and re-rank the preliminary search results in conjunction with the treatment effectiveness feedback to obtain the final search results. A similarity score list is obtained by calculating the similarity between the feature vectors of the query case and the feature vectors of candidate cases, representing the degree of similarity between the candidate cases and the query case. Treatment effectiveness feedback comes from doctors, patients, and clinical trials, indicating the effectiveness of specific treatment plans for similar cases. The feature vectors of past health records usually contain temporal information, such as the time sequence and duration of symptom onset, which is crucial for understanding the development and evolution of diseases. Based on temporal information, a symptom temporal attention mechanism can be introduced to dynamically adjust the importance of symptoms at different time points, thereby more accurately reflecting the development process of the disease. A symptom temporal decay factor is defined in the symptom temporal attention mechanism, indicating that the influence of early symptoms on the current condition gradually weakens over time. Simultaneously, considering the weighting of symptom severity—that is, the degree to which different symptoms contribute to the severity of the condition—these factors collectively affect the similarity score, thus better reflecting the actual progression of the disease. Re-ranking the results by incorporating treatment effectiveness feedback from historical cases ensures that the search results not only resemble the query case but also provide validated and effective treatment options. For example, a query case might be about a patient with "chronic obstructive pulmonary disease (COPD)" whose symptoms include cough, sputum production, and shortness of breath, and which have gradually worsened over the past year. By calculating the similarity score between the query case and candidate cases, a list of COPD-related candidate cases can be obtained. In the initial search results, some candidate cases include treatment effectiveness feedback, such as symptom relief after using specific inhalers. Extracting temporal information from the feature vectors of past health records reveals that the query case's symptoms have gradually worsened over the past year, particularly shortness of breath. Introducing a symptom temporal attention mechanism adjusts the importance of symptoms at different time points. For example, recent dyspnea symptoms are given higher weight, while early cough and sputum symptoms are given lower weight. Adjusting the similarity score based on symptom attenuation factors and symptom severity weights allows candidate cases more similar to the query case in symptom progression and severity to receive higher scores. Incorporating treatment effectiveness feedback, the initial search results are reordered; for example, candidate cases with effective treatment plans and similar symptom progression are ranked higher.
[0052] The method described above, which compresses a pre-trained biomedical language model into a Transformer architecture with a preset number of layers using knowledge distillation and denotes it as a lightweight biomedical language model, and injects entity embeddings from a medical terminology knowledge graph to generate feature vectors of past health records, includes:
[0053] Entity embedding alignment of a medical terminology knowledge graph is performed using a knowledge distillation loss function, followed by cross-entropy loss. The knowledge distillation loss function is obtained by linearly combining the entity embedding alignment loss with the knowledge distillation loss function. The knowledge distillation loss function The calculation formula is:
[0054] .
[0055] in, These are preset loss weight coefficients used to balance the cross-entropy loss. The relative importance between the entity embedding alignment loss and the actual entity embedding alignment loss; A set of symptom entities extracted from a medical terminology knowledge graph; Indicates the entity embedding alignment loss; Entity embedding vectors generated for a pre-trained biomedical language model. Entity embedding vectors generated for a lightweight biomedical language model.
[0056] Entity embedding alignment loss measures the difference or distance between entity embedding vectors generated by the lightweight biomedical language model and those generated by the pre-trained biomedical language model. Minimizing this loss ensures that the entity embedding vectors from the lightweight model are semantically consistent with those from the pre-trained model. Cross-entropy loss measures the difference between the probability distribution output by the lightweight model and the soft labels generated by the pre-trained model. Soft labels, representing the probability distribution output by the pre-trained model, provide more information than hard labels (i.e., one-hot encoded labels), helping the lightweight model learn more nuanced category discrimination. Preset loss weights control the trade-off between classification performance and entity embedding alignment in the lightweight model. For example, the loss weights are adjusted based on symptom entity types. Anatomical terms (such as "myocardium"): =0.6 (emphasis on entity alignment); descriptive symptoms (e.g., "dull pain"): =0.8 (emphasis on semantic distillation). Through The Transformer architecture with a predetermined number of layers is used to align the entity semantics of the knowledge graph.
[0057] The method for extracting long-tailed cases from the feature vectors of past health records, performing product quantization compression, and then constructing a second-level approximate nearest neighbor index includes:
[0058] The product quantization employs bit-width selection and residual compensation. Bit-width selection involves dividing the feature vectors of long-tailed cases' past health records into a predetermined number of subspaces before scalar quantization, and then calculating the similarity values before and after scalar quantization and residual compensation. The similarity value The calculation formula is:
[0059] .
[0060] in, Scalar quantization of similarity weights; This represents the similarity value of the feature vectors of past health records after scalar quantization. For the pre-defined set of key symptoms The exact match results for all symptom terms are summed. and These represent the query cases and candidate cases within the preset set of key symptoms, respectively. The first in Quantitative values for each symptom term; This is a function for exact matching of symptom terms. It returns 1 when two symptom terms match exactly, and 0 otherwise.
[0061] When the similarity value after residual compensation is less than the preset similarity threshold, the bit width is adjusted and the similarity value after residual compensation is recalculated until the similarity value after residual compensation is greater than the preset similarity threshold, and then a secondary approximate nearest neighbor index is constructed.
[0062] Product quantization compression is an effective vector compression technique suitable for storing and retrieving large-scale vector data. Historical health record feature vectors typically have high dimensionality and large data volumes; therefore, product quantization compression is widely used to compress historical health record data and accelerate the retrieval process. Bit width selection involves dividing the historical health record feature vector into a predetermined number of subspaces, each containing a portion of the feature dimensions. Scalar quantization is then performed on each subspace, using codewords from a codebook to represent the vectors within that subspace. Choosing an appropriate bit width reduces storage and computational overhead while maintaining a certain level of accuracy. After scalar quantization, the residual between the original historical health record feature vector and the scalar-quantized vector is calculated. Residual compensation techniques are used to adjust the similarity calculation, reducing information loss introduced during quantization. When the similarity after residual compensation is less than a preset similarity threshold, it indicates significant information loss during scalar quantization, leading to decreased retrieval performance.
[0063] The method for adjusting the similarity score based on the symptom temporal decay factor and symptom severity weight according to the symptom temporal attention mechanism, and re-ranking the preliminary search results in conjunction with treatment effectiveness feedback to obtain the final search results includes:
[0064] The symptom time decay factor The calculation formula is:
[0065] .
[0066] in, This indicates the number of days between the onset of symptoms and the current time. To preset the maximum effective time window, This is the decay rate coefficient.
[0067] The initial weight matrix is obtained by performing a Hadamard product operation on the symptom time decay factor and the symptom severity weight. The initial matrix is then normalized using Softmax to generate an attention weight matrix. The similarity score is adjusted based on the attention weight matrix. The symptom severity weight is obtained by acquiring medical knowledge and expert experience data and performing data statistical analysis. The similarity score is obtained by calculating the similarity between the query case and the candidate case using cosine similarity.
[0068] The symptom time decay factor measures the impact of the interval between symptom onset and the current time on symptom importance. Over time, some symptoms may become less relevant or important, thus requiring a decay mechanism to reflect this change. As the number of days between symptom onset and the current time increases, the symptom time decay factor gradually decreases, indicating a decrease in the timeliness of the symptom. Symptom severity weights are obtained by acquiring medical knowledge and expert experience data and using statistical analysis methods (such as regression analysis and expert scoring). The Hadamard product operation is element-wise multiplication, and the Softmax function can convert any real-valued vector into a probability distribution, ensuring that the sum of all weights is 1. An attention weight matrix is used to adjust the similarity score between the query case and the candidate cases. That is, the attention weights are weighted and summed with the original similarity score to obtain the adjusted similarity score. Assume a query case with symptoms including headache, fever, and cough, which appeared 1 day, 3 days, and 5 days ago, respectively. A preset maximum effective time window is used. The attenuation period is 7 days, and the attenuation rate coefficient is 0.1. The attenuation factor for headache is d(1)≈0.986; the attenuation factor for fever is d(3)≈0.958; and the attenuation factor for cough is d(5)≈0.931. The severity weights of symptoms are obtained: headache: 0.6; fever: 0.8; cough: 0.5. The initial weight matrix is [0.986*0.6,0.958*0.8,0.931*0.5]≈[0.592,0.766,0.466]. After Softmax normalization, the attention weight matrix is obtained as [0.32,0.42,0.26]. Assume that there are three candidate cases in the preliminary search results, and the original similarity scores with the query case are 0.7, 0.6, and 0.5, respectively. The similarity score adjusted using the attention weight matrix is: [0.32*0.7, 0.42*0.6, 0.26*0.5] = [0.224, 0.252, 0.13], indicating that the second candidate case has the best treatment effect. Therefore, the second candidate case is ranked first in the search results to obtain the final search result.
[0069] The method further includes:
[0070] The decay rate coefficient is determined by a pre-set clinical rule base. Associating and mapping with disease types; the preset clinical rule base is: the attenuation rate coefficient of cardiovascular diseases. =0.25±0.05, corresponding to symptoms of acute illnesses with high time sensitivity; decay rate coefficient of chronic diseases. =0.15±0.03, corresponding to persistent symptoms; attenuation rate coefficient for infectious diseases. =0.20±0.05, corresponding to symptoms of a rapidly developing and changing disease.
[0071] When multiple disease types are detected to coexist, the decay rate coefficient corresponding to each disease is obtained. The weighted average of the values.
[0072] A pre-defined clinical rule base sets reasonable attenuation rate coefficient ranges for different types of diseases. These values are derived from medical knowledge and clinical experience data. For example, cardiovascular diseases typically have acute onset and high symptom time sensitivity, thus their attenuation rate coefficients are set relatively high; chronic disease symptoms tend to be persistent and change slowly, thus their attenuation rate coefficients are set relatively low; infectious disease symptoms develop rapidly and may be accompanied by sudden deterioration, thus their attenuation rate coefficients are set between those of cardiovascular diseases and chronic diseases. However, in actual clinical practice, patients may have multiple diseases simultaneously. When multiple disease types are detected to coexist in a patient, the weighted average of the attenuation rate coefficients corresponding to each disease is obtained and used as the attenuation rate coefficient for the patient's symptoms. Assuming a query case where the patient has both cardiovascular disease (e.g., acute myocardial infarction) and infectious disease (e.g., pneumonia), the attenuation rate coefficients for each disease are determined as follows: According to the pre-defined clinical rule base, the attenuation rate coefficient for cardiovascular disease is 0.25 ± 0.05, so β_cardiovascular = 0.25 is selected. The attenuation rate coefficient for infectious disease is 0.20 ± 0.05, so β_infectious = 0.20 is selected. Assuming that cardiovascular disease and infectious disease have the same impact on the patient's overall health, the arithmetic mean is used as the weighted average. Therefore, the weighted average of the attenuation rate coefficient = (β_cardiovascular + β_infectious) / 2 = (0.25 + 0.20) / 2 = 0.225.
[0073] The method for updating the lightweight biomedical language model when the cumulative number of newly detected rare disease cases exceeds a preset threshold for rare disease cases includes:
[0074] When monitoring new cases entering the case database, the health status assessment label and ICD code of the new cases are obtained; the cumulative number of rare disease cases is counted in real time based on the health status assessment label and ICD code; when the cumulative number of rare disease cases exceeds the preset cumulative number of cases threshold, the lightweight biomedical language model is updated through the elastic weight consolidation algorithm.
[0075] Rare diseases, due to their low incidence and small number of cases, often pose challenges to the assessment and treatment of medical health status. Health status assessment labels are the results of doctors' assessments of the health status of new cases, while ICD codes are internationally recognized disease classification codes used to standardize disease names and classifications. Based on health status assessment labels and ICD codes, the cumulative number of cases for each rare disease is counted in real time. This involves traversing and matching historical data in the case database to determine which cases belong to rare diseases and to calculate the total number. A threshold for the number of rare disease cases is preset. When the cumulative number of cases for a certain rare disease exceeds the threshold, it is considered that there is sufficient case data to support the update of the lightweight biomedical language model. Once it is determined that the lightweight biomedical language model needs to be updated, the Elastic Weight Consolidation (EWC) algorithm is used to update the lightweight biomedical language model. The EWC algorithm is a technique for continuous learning of neural networks, capable of learning new knowledge without forgetting old knowledge. The loss function of the Elastic Weight Consolidation algorithm in this embodiment is... for: ;in, The cross-entropy loss function; This is the regularization coefficient, obtained through analysis of experimental data; For the first Diagonal elements of the Fisher information matrix for each historical case These are the current parameters for a lightweight biomedical language model. ...
[0076] The method for obtaining a query semantic complexity score, and then using the query semantic complexity score to dynamically select a retrieval path through a dynamic routing algorithm to obtain preliminary retrieval results includes:
[0077] The semantic complexity score is obtained by performing feature parsing on the input query information. The semantic complexity score The calculation formula is:
[0078] .
[0079] in, To query the number of medical entities contained in the information, The average semantic association between medical entities; Image feature dimension; The preset adjustment factor can control the weight of image feature dimensions in the semantic complexity score.
[0080] The query semantic complexity score is used to dynamically select the retrieval path through a dynamic routing algorithm: the retrieval path includes a cache layer index, a quantization layer index, or a full index deep retrieval.
[0081] Obtain the first semantic complexity score threshold and the second semantic complexity score threshold. When the semantic complexity score is less than the first semantic complexity score threshold, query the cache layer index. When the semantic complexity score is greater than or equal to the first semantic complexity score threshold and less than the second semantic complexity score threshold, query the cache layer index and the quantization layer index. When the semantic complexity score is greater than or equal to the second semantic complexity score threshold, trigger a full index deep search.
[0082] The semantic complexity score of a query is calculated by analyzing factors such as the number of medical entities, the semantic relevance between entities, and the dimensionality of image features. A query containing multiple complex symptoms and detailed health assessment information may have a high semantic complexity score, while a query containing only simple symptoms may have a low score. Medical entities can be disease names, drug names, examination items, etc.; the more entities, the more semantically complex the query. The average semantic relevance between entities reflects the strength of semantic connections between entities in the query; higher relevance indicates a more focused query, while lower relevance indicates the query involves multiple unrelated topics. If the query includes medical image information, the dimensionality of image features also affects the semantic complexity assessment. Higher-dimensional image features represent richer information and greater complexity. Cached indexes typically store frequently accessed or highly relevant past health records, allowing for rapid result retrieval. Quantized indexes improve retrieval speed while maintaining a certain level of accuracy by quantifying past health record information. Full indexes contain a complete representation of all past health record information; although slower, they return the most accurate results.
[0083] The Entity embedding vectors generated for a pre-trained biomedical language model. Methods for generating entity embedding vectors for lightweight biomedical language models include:
[0084] During entity embedding alignment, symptom description information from a pre-trained biomedical language model and a lightweight biomedical language model is acquired and named entity recognition is performed. The identified entities are then mapped to knowledge graph nodes, and neighboring node information is aggregated using a graph attention network to generate entity embedding vectors. and .
[0085] Symptom description information comes from the patient's past health records, the doctor's health assessment report, or the patient's self-description. Mapping relationships can associate identified entities with corresponding nodes in the knowledge graph, thereby leveraging the rich information in the knowledge graph to enhance entity representation. Once a mapping relationship is established between an entity and a knowledge graph node, a Graph Attention Network (GAT) is used to aggregate information from neighboring nodes. A GAT is a graph neural network based on an attention mechanism that dynamically assigns weights according to the relationships between nodes, effectively aggregating information from neighboring nodes. The GAT consists of two layers, each with eight attention heads, and an aggregation hop count of 2. For example, given the symptom description: "Patient has a cough and fever, suspected of having pneumonia," the following entities are identified: "cough," "fever," and "pneumonia." These entities are mapped to corresponding nodes in the knowledge graph. For example, "cough" and "fever" might map to symptom nodes, and "pneumonia" might map to disease nodes. For the entity "pneumonia," the GAT aggregates information related to pneumonia symptoms (such as cough and fever), treatment drugs, and examination items.
[0086] Finally, it should be noted that although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for efficient data retrieval based on neural networks, characterized in that, The method includes: The pre-trained biomedical language model is compressed into a Transformer architecture with a preset number of layers using the knowledge distillation method and denoted as a lightweight biomedical language model. Entity embeddings from a medical terminology knowledge graph are then injected to generate feature vectors of past health records. High-frequency cases are obtained from the feature vectors of past health records and stored in an inverted index cache layer based on the joint key value of symptom-health status assessment. Long-tail cases are obtained from the feature vectors of past health records and constructed into a secondary approximate nearest neighbor index after product quantization compression. The query semantic complexity score is obtained, and the query semantic complexity score is used to dynamically select the retrieval path for retrieval through a dynamic routing algorithm to obtain preliminary retrieval results. When the cumulative number of newly added rare disease cases exceeds the preset threshold for the number of rare disease cases, the lightweight biomedical language model is updated, and the feature vectors of the previous health records of the newly added rare disease cases are seamlessly inserted into the secondary approximate nearest neighbor index using the streaming graph index expansion method. Obtain the similarity score between the query case and the candidate cases; obtain the treatment effectiveness feedback of historical cases in the preliminary search results; obtain the temporal information in the feature vector of past health records; introduce the symptom temporal attention mechanism into the preliminary search results based on the temporal information; adjust the similarity score according to the symptom temporal decay factor and symptom severity weight of the symptom temporal attention mechanism; and re-rank the preliminary search results in combination with the treatment effectiveness feedback to obtain the final search results.
2. The efficient data retrieval method based on neural networks according to claim 1, characterized in that, The method described above, which compresses a pre-trained biomedical language model into a Transformer architecture with a preset number of layers using knowledge distillation and denotes it as a lightweight biomedical language model, and injects entity embeddings from a medical terminology knowledge graph to generate feature vectors of past health records, includes: Entity embedding alignment of a medical terminology knowledge graph is performed using a knowledge distillation loss function, followed by cross-entropy loss. The knowledge distillation loss function is obtained by linearly combining the entity embedding alignment loss with the knowledge distillation loss function. The knowledge distillation loss function The calculation formula is: ; in, These are preset loss weight coefficients used to balance the cross-entropy loss. The relative importance between the entity embedding alignment loss and the actual entity embedding alignment loss; A set of symptom entities extracted from a medical terminology knowledge graph; Indicates the entity embedding alignment loss; Entity embedding vectors generated for a pre-trained biomedical language model. Entity embedding vectors generated for a lightweight biomedical language model.
3. The efficient data retrieval method based on neural networks according to claim 1, characterized in that, The method for extracting long-tailed cases from the feature vectors of past health records, performing product quantization compression, and then constructing a second-level approximate nearest neighbor index includes: The product quantization employs bit-width selection and residual compensation. Bit-width selection involves dividing the feature vectors of long-tailed cases' past health records into a predetermined number of subspaces before scalar quantization, and then calculating the similarity values before and after scalar quantization and residual compensation. The similarity value The calculation formula is: ; in, Scalar quantization of similarity weights; This represents the similarity value of the feature vectors of past health records after scalar quantization. For the pre-defined set of key symptoms The exact match results for all symptom terms are summed. and These represent the query cases and candidate cases within the preset set of key symptoms, respectively. The first in Quantitative values for each symptom term; This is a function for exact matching of symptom terms. It returns 1 when two symptom terms match exactly, and 0 otherwise. When the similarity value after residual compensation is less than the preset similarity threshold, the bit width is adjusted and the similarity value after residual compensation is recalculated until the similarity value after residual compensation is greater than the preset similarity threshold, and then a secondary approximate nearest neighbor index is constructed.
4. The efficient data retrieval method based on neural networks according to claim 1, characterized in that, The method for adjusting the similarity score based on the symptom temporal decay factor and symptom severity weight according to the symptom temporal attention mechanism, and re-ranking the preliminary search results in conjunction with treatment effectiveness feedback to obtain the final search results includes: The symptom time decay factor The calculation formula is: ; in, This indicates the number of days between the onset of symptoms and the current time. To preset the maximum effective time window, This is the decay rate coefficient; The initial weight matrix is obtained by performing a Hadamard product operation on the symptom time decay factor and the symptom severity weight. The initial matrix is then normalized using Softmax to generate an attention weight matrix. The similarity score is adjusted based on the attention weight matrix. The symptom severity weight is obtained by acquiring medical knowledge and expert experience data and performing data statistical analysis. The similarity score is obtained by calculating the similarity between the query case and the candidate case using cosine similarity.
5. The efficient data retrieval method based on neural networks according to claim 4, characterized in that, The method further includes: The decay rate coefficient is determined by a pre-set clinical rule base. Associating and mapping with disease types; the preset clinical rule base is: the attenuation rate coefficient of cardiovascular diseases. =0.25±0.05, corresponding to symptoms of acute illnesses with high time sensitivity; decay rate coefficient of chronic diseases. =0.15±0.03, corresponding to persistent symptoms; attenuation rate coefficient for infectious diseases. =0.20±0.05, corresponding to symptoms of a rapidly developing and changing disease; When multiple disease types are detected to coexist, the decay rate coefficient corresponding to each disease is obtained. The weighted average of the values.
6. The efficient data retrieval method based on neural networks according to claim 1, characterized in that, The method for updating the lightweight biomedical language model when the cumulative number of newly detected rare disease cases exceeds a preset threshold for rare disease cases includes: When monitoring new cases entering the case database, the health status assessment label and ICD code of the new cases are obtained; the cumulative number of rare disease cases is counted in real time based on the health status assessment label and ICD code; when the cumulative number of rare disease cases exceeds the preset cumulative number of cases threshold, the lightweight biomedical language model is updated through the elastic weight consolidation algorithm.
7. The efficient data retrieval method based on neural networks according to claim 1, characterized in that, The method for obtaining a query semantic complexity score, and then using the query semantic complexity score to dynamically select a retrieval path through a dynamic routing algorithm to obtain preliminary retrieval results includes: The semantic complexity score is obtained by performing feature parsing on the input query information. The semantic complexity score The calculation formula is: ; in, To query the number of medical entities contained in the information, The average semantic association between medical entities; Image feature dimension; This is a preset adjustment factor; The query semantic complexity score is used to dynamically select the retrieval path through a dynamic routing algorithm: the retrieval path includes a cache layer index, a quantization layer index, or a full index deep search. Obtain the first semantic complexity score threshold and the second semantic complexity score threshold. When the semantic complexity score is less than the first semantic complexity score threshold, query the cache layer index. When the semantic complexity score is greater than or equal to the first semantic complexity score threshold and less than the second semantic complexity score threshold, query the cache layer index and the quantization layer index. When the semantic complexity score is greater than or equal to the second semantic complexity score threshold, trigger a full index deep search.
8. The efficient data retrieval method based on neural networks according to claim 2, characterized in that, The Entity embedding vectors generated for a pre-trained biomedical language model. Methods for generating entity embedding vectors for lightweight biomedical language models include: During entity embedding alignment, symptom description information from a pre-trained biomedical language model and a lightweight biomedical language model is acquired and named entity recognition is performed. The identified entities are then mapped to knowledge graph nodes, and neighboring node information is aggregated using a graph attention network to generate entity embedding vectors. and .
Citation Information
Cited By
Rare disease data retrieval method and system
CN121641317A