Intelligent Retrieval Methods and Related Equipment for Scientific and Technological Literature Based on Generative Artificial Intelligence
By performing multi-layer semantic annotation and graph construction on medical and scientific literature, combined with a dual encoder architecture and knowledge base parsing, we have achieved deep understanding and accurate retrieval of medical terms, solving the problem of insufficient semantic understanding in existing technologies and improving the accuracy and efficiency of retrieval.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2026-04-03
AI Technical Summary
Existing intelligent retrieval methods for scientific and technological literature based on generative artificial intelligence lack sufficient semantic understanding depth, making it unable to accurately understand the complex relationships between medical terms and clinical contexts. This results in a significant deviation between retrieval results and actual clinical needs, making it difficult to achieve accurate matching of clinical scenarios.
By performing multi-layer semantic annotation on medical and scientific literature, a dynamic correlation map and semantic mapping of symptoms and diseases are constructed, and a medical and scientific literature knowledge base containing semantic fingerprints is established. A dual encoder architecture is used to perform medical context parsing and multimodal feature extraction, and Boolean operation nested parsing and semantic alignment are performed to generate a unified retrieval vector. Combined with evidence-based retrieval and clinical scenario matching, target candidate literature is screened and knowledge-weighted fusion is performed to generate a recommendation report.
It improves the accuracy of medical and scientific literature retrieval, ensures that retrieval results match patients' clinical characteristics, and reduces the time and workload for medical staff to screen valuable literature.
Smart Images

Figure CN120687597B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and related equipment for intelligent retrieval of scientific and technological literature based on generative artificial intelligence. Background Technology
[0002] With the rapid development of medical research and the explosive growth in the number of medical and scientific literature, intelligent retrieval of scientific and scientific literature based on generative artificial intelligence has become an important tool for clinical decision-making and medical research. Traditional intelligent retrieval methods for scientific and scientific literature based on generative artificial intelligence mainly rely on keyword matching and simple semantic similarity calculation, which have obvious limitations when dealing with complex medical queries.
[0003] In existing technologies, intelligent scientific literature retrieval systems based on generative artificial intelligence generally suffer from insufficient semantic understanding depth, failing to accurately grasp the complex relationships and clinical contexts between medical terms. Existing retrieval methods struggle to precisely match specific patient clinical characteristics with research scenarios in the literature, resulting in significant discrepancies between retrieval results and actual clinical needs. Medical professionals must then spend considerable time sifting through massive amounts of search results to extract truly valuable information. Summary of the Invention
[0004] The main objective of this invention is to solve the technical problems of insufficient semantic understanding depth and inability to achieve accurate clinical scenario matching in existing intelligent retrieval methods for scientific and technological literature based on generative artificial intelligence.
[0005] The first aspect of this invention provides a method for intelligent retrieval of scientific and technological literature based on generative artificial intelligence, the method comprising:
[0006] Multi-layer semantic annotation is performed on medical science and technology literature, and a dynamic correlation graph of symptoms and diseases is constructed based on the annotated medical terms. Semantic mapping is also performed on the medical terms to obtain a medical science and technology literature knowledge base containing semantic fingerprints.
[0007] The system employs a dual-encoder architecture to perform medical context parsing and multimodal feature extraction on user queries. It also performs nested Boolean operations to parse the logical relationships in user queries and uses the symptom-disease dynamic association graph and semantic fingerprint for standardization transformation and semantic alignment to obtain a unified retrieval vector.
[0008] Based on the unified retrieval vector and the medical science and technology literature knowledge base, evidence classification retrieval and clinical scenario matching are performed respectively to obtain a preliminary candidate literature set. Then, fine-grained semantic recalculation is performed on the preliminary candidate literature set to obtain the target candidate literature set.
[0009] The prior retrieval probability is calculated based on the evaluation dimensions of the target candidate literature set, and the candidate literature set is then subjected to knowledge-weighted fusion based on the prior retrieval probability to obtain a medical and scientific literature recommendation report.
[0010] Optionally, in a first implementation of the first aspect of the present invention, the step of performing multi-layer semantic annotation processing on medical and scientific literature, constructing a symptom-disease dynamic association map based on the annotated medical terms, and performing semantic mapping on the medical terms to obtain a medical and scientific literature knowledge base containing semantic fingerprints includes:
[0011] Medical terminology is identified and processed in medical science and technology literature. The medical terms are then annotated with entities using a medical ontology knowledge base to obtain the medical terminology annotation results. The medical terms include disease names, symptom descriptions, drug names, and treatment plans.
[0012] The association strength weight is calculated based on the co-occurrence patterns, causal relationship descriptions, and time series information of symptom and disease terms in the medical terminology annotation results.
[0013] A graph structure is constructed using symptom and disease terms as graph nodes and association strength weights as edge weights to obtain a dynamic association graph of symptoms and diseases.
[0014] The medical terminology annotation results are subjected to semantic mapping processing. By establishing a three-layer correspondence between the professional terminology layer, the clinical description layer, and the patient expression layer, the semantic mapping results are obtained.
[0015] Based on the symptom-disease dynamic association map and semantic mapping results, feature extraction is performed on medical and scientific literature. The main disease categories, core symptom clusters, treatment plan types and patient population characteristics are extracted to generate semantic fingerprints, resulting in a medical and scientific literature knowledge base containing semantic fingerprints.
[0016] Optionally, in the second implementation of the first aspect of the present invention, the step of performing medical context parsing and multimodal feature extraction on the user query through a dual encoder architecture, performing nested Boolean operations on the logical relationships in the user query, and performing standardization transformation and semantic alignment using the symptom-disease dynamic association graph and semantic fingerprint to obtain a unified retrieval vector includes:
[0017] The ViT-B / 16 encoder in the dual-encoder architecture extracts patch-level fine-grained features from the image input in the user query, and the BERT encoder in the dual-encoder architecture extracts token-level semantic features from the text query in the user query, thus obtaining multimodal query features.
[0018] Based on the multimodal query features, the query type of the user query is determined, the query context features are obtained, and the logical relationships in the user query are processed by Boolean operation nested parsing to obtain a structured query logic representation.
[0019] The symptom-disease dynamic association graph is used to standardize the descriptions of patient symptoms queried by users, converting non-standard symptom expressions into standard medical terms to obtain standardized symptom features.
[0020] Based on the semantic fingerprint, semantic alignment processing is performed on multimodal query features, standardized symptom features, and structured query logic representation, and a unified retrieval vector is generated by fusing query context features.
[0021] Optionally, in a third implementation of the first aspect of the present invention, the step of performing evidence-based hierarchical retrieval and clinical scenario matching based on the unified retrieval vector and the medical science and technology literature knowledge base to obtain a preliminary candidate literature set, and performing fine-grained semantic recalculation on the preliminary candidate literature set to obtain a target candidate literature set includes:
[0022] Evidence grading retrieval is performed based on the unified retrieval vector and the medical science and technology literature knowledge base. By matching query complexity and literature evidence level, candidate literature for evidence grading is obtained.
[0023] Clinical scenarios are matched with the unified retrieval vector and the medical science and technology literature knowledge base. By matching patient characteristics with the clinical scenarios of the literature, candidate literature for clinical scenarios is obtained.
[0024] The candidate literature for evidence grading and candidate literature for clinical scenarios are merged, duplicate literature is removed and sorted by search score to obtain a preliminary candidate literature set.
[0025] The semantic matching score between each candidate document in the preliminary candidate document set and the unified retrieval vector is calculated using the ColBERTv2 model. The top M documents are then sorted by score to obtain the coarsely ranked candidate documents.
[0026] By extracting patch-level and token-level features from the coarsely ranked candidate documents, calculating the maximum similarity matching score between features, re-ranking and filtering the Top-K documents, a target candidate document set is obtained.
[0027] Optionally, in a fourth implementation of the first aspect of the present invention, the step of performing evidence-level retrieval based on the unified retrieval vector and the medical science and technology literature knowledge base, and obtaining candidate documents for evidence-level grading by matching query complexity and literature evidence level, includes:
[0028] The query complexity score is obtained by calculating the semantic depth index, symptom combination complexity index, and diagnostic reasoning hierarchy index of the unified retrieval vector.
[0029] Based on the query complexity score, the literature in the medical science and technology literature knowledge base is pre-labeled according to the evidence level of randomized controlled trials, systematic reviews, cohort studies, and case reports, resulting in a graded labeled literature database;
[0030] Based on the preset complexity threshold and evidence level mapping relationship, the query complexity score is matched with documents of each level in the graded labeled literature database, and the selection priority score of each document is calculated to obtain a priority score list.
[0031] Literature with a matching degree exceeding a preset threshold is selected based on the priority score list, and then sorted according to the citation frequency of the literature to obtain candidate literature for evidence classification.
[0032] Optionally, in a fifth implementation of the first aspect of the present invention, the step of performing clinical scenario matching based on the unified retrieval vector and the medical science and technology literature knowledge base, and obtaining candidate clinical scenario literature by matching patient characteristics with literature clinical scenarios, includes:
[0033] A structured patient feature vector is obtained by extracting age range, gender information, main symptoms, medical history and medication information from a unified retrieval vector using a named entity recognition algorithm.
[0034] By analyzing the research subject descriptions, inclusion and exclusion criteria, and clinical trial designs of each article in the medical science and technology literature knowledge base, a clinical scenario feature matrix containing applicable populations, disease stages, and treatment plans is constructed for the articles.
[0035] By calculating the cosine similarity between the structured patient feature vector and the clinical scenario feature matrix of the literature, a list of patient-literature similarity scores is obtained.
[0036] Based on the similarity score list and combined with the Bayesian confidence model, the clinical applicability credibility of each article is calculated. Articles with clinical applicability credibility exceeding the preset threshold are selected to obtain candidate articles for clinical scenarios.
[0037] Optionally, in a sixth implementation of the first aspect of the present invention, the step of calculating the prior retrieval probability based on the evaluation dimension of the target candidate document set and performing knowledge-weighted fusion on the candidate document set based on the prior retrieval probability to obtain a medical and scientific literature recommendation report includes:
[0038] The prior probability distribution is obtained by calculating the prior probability of each document in the target candidate document set through the multi-dimensional evaluation vector of the evaluation dimension and softmax normalization.
[0039] Based on the prior probability distribution, the target candidate document set is subjected to knowledge weighted fusion. Through prior probability weighting and feature concatenation operations, a weighted fusion knowledge representation is obtained.
[0040] The weighted fusion knowledge representation is used to generate recommended content, which is then converted into structured recommended text by a decoder to obtain preliminary recommended content.
[0041] The preliminary recommendations are then citation-marked, and citation tags and lists are added based on the content sources to obtain a medical and scientific literature recommendation report.
[0042] A second aspect of the present invention provides a scientific and technological literature intelligent retrieval device based on generative artificial intelligence, the device comprising:
[0043] The semantic annotation module is used to perform multi-layer semantic annotation processing on medical and scientific literature, construct a dynamic correlation map of symptoms and diseases based on the annotated medical terms, and perform semantic mapping on the medical terms to obtain a medical and scientific literature knowledge base containing semantic fingerprints.
[0044] The query parsing module is used to perform medical context parsing and multimodal feature extraction on user queries through a dual encoder architecture, perform nested Boolean operations on the logical relationships in the user query, and perform standardized transformation and semantic alignment using the symptom-disease dynamic association graph and semantic fingerprint to obtain a unified retrieval vector.
[0045] The intelligent retrieval module is used to perform evidence-level retrieval and clinical scenario matching based on the unified retrieval vector and the medical science and technology literature knowledge base, respectively, to obtain a preliminary candidate literature set, and to perform fine-grained semantic recalculation on the preliminary candidate literature set to obtain a target candidate literature set.
[0046] The fusion recommendation module is used to calculate the prior retrieval probability based on the evaluation dimensions of the target candidate document set and to perform knowledge-weighted fusion of the candidate document set based on the prior retrieval probability to obtain a medical and scientific literature recommendation report.
[0047] A third aspect of the present invention provides a scientific and technological literature intelligent retrieval device based on generative artificial intelligence, comprising: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a circuit; the at least one processor invokes the instructions in the memory to cause the scientific and technological literature intelligent retrieval device based on generative artificial intelligence to perform the steps of the above-described scientific and technological literature intelligent retrieval method based on generative artificial intelligence.
[0048] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the steps of the above-described intelligent retrieval method for scientific and technological literature based on generative artificial intelligence.
[0049] The aforementioned intelligent retrieval method and related equipment for scientific and technological literature based on generative artificial intelligence establishes a medical and technological literature knowledge base containing semantic fingerprints by performing multi-layer semantic annotation on medical and technological literature and constructing a dynamic symptom-disease association graph. It performs medical context analysis and multimodal feature extraction on user queries, combining Boolean operation nested parsing and semantic alignment processing to generate a unified retrieval vector. Based on the unified retrieval vector, it performs evidence-level retrieval and clinical scenario matching to obtain a preliminary candidate literature set, which is then optimized into a target candidate literature set through fine-grained semantic recomputation. The prior probability of retrieval is calculated according to the evaluation dimensions, and the candidate literature is subjected to knowledge-weighted fusion to generate a medical and technological literature recommendation report. This invention deeply understands the association relationships of medical terms through a dynamic symptom-disease association graph, and ensures that the results match the patient's clinical characteristics through evidence-level retrieval and clinical scenario matching, thereby improving the accuracy of literature retrieval.
[0050] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0051] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the first embodiment of the intelligent scientific literature retrieval method based on generative artificial intelligence in this invention.
[0053] Figure 2 This is a schematic diagram of one embodiment of the intelligent scientific literature retrieval device based on generative artificial intelligence in this invention.
[0054] Figure 3 This is a schematic diagram of one embodiment of the intelligent scientific literature retrieval device based on generative artificial intelligence in this invention. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0056] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0057] To facilitate understanding of this embodiment, a detailed description of the intelligent scientific literature retrieval method based on generative artificial intelligence disclosed in this embodiment of the invention will be provided first. For example... Figure 1 As shown, this method includes the following steps:
[0058] 101. Perform multi-layer semantic annotation processing on medical and scientific literature, construct a dynamic correlation map of symptoms and diseases based on the annotated medical terms, and perform semantic mapping on the medical terms to obtain a medical and scientific literature knowledge base containing semantic fingerprints.
[0059] In one embodiment of the present invention, the step of performing multi-layer semantic annotation processing on medical and scientific literature, constructing a dynamic symptom-disease association graph based on the annotated medical terms, and performing semantic mapping on the medical terms to obtain a medical and scientific literature knowledge base containing semantic fingerprints includes: performing medical professional terminology recognition processing on medical and scientific literature; performing entity annotation on medical terms through a medical ontology knowledge base to obtain medical terminology annotation results, wherein the medical terms include disease names, symptom descriptions, drug names, and treatment plans; calculating association strength weights based on the co-occurrence patterns, causal relationship descriptions, and time series information of symptom terms and disease terms in the medical terminology annotation results; constructing a graph structure using symptom terms and disease terms as graph nodes and association strength weights as edge weights to obtain a dynamic symptom-disease association graph; performing semantic mapping processing on the medical terminology annotation results, obtaining semantic mapping results by establishing a three-layer correspondence between a professional terminology layer, a clinical description layer, and a patient expression layer; and extracting features from the medical and scientific literature based on the dynamic symptom-disease association graph and semantic mapping results to generate semantic fingerprints by extracting main disease categories, core symptom clusters, treatment plan types, and patient population features to obtain a medical and scientific literature knowledge base containing semantic fingerprints.
[0060] Specifically, the multi-layer semantic annotation of medical and scientific literature employs a BiLSTM-CRF (Bidirectional Long Short-Term Memory Network-Conditional Random Field) model combined with a medical ontology knowledge base for term extraction and annotation. BiLSTM-CRF is a sequence labeling model where BiLSTM captures contextual information, while the CRF layer ensures global consistency of the labeled sequence. This combination is particularly suitable for handling the problems of ambiguous term boundaries and strong contextual dependencies in medical texts. The medical ontology knowledge base is centered on the internationally standardized Medical Subject Headings (MeSH), covering standard terminology systems for diseases, drugs, and genes, while also integrating authoritative medical terminology databases such as UMLS and SNOMED CT. During term recognition, the system first uses a pre-trained BioBERT model to represent word vectors in the medical and scientific literature, and then uses the BiLSTM-CRF model to identify key medical entities such as disease names, symptom descriptions, drug names, and treatment plans. The identified terms are mapped to standard concepts in the medical ontology database through string matching and semantic similarity calculation. Each term obtains a corresponding CUI code (Concept Unique Identifier) and semantic type label, forming a structured medical term annotation result.
[0061] Specifically, based on the medical terminology annotation results, the system uses the Pointwise Mutual Information (PMI) algorithm to calculate the association strength between symptom and disease terms. The PMI algorithm measures the degree of association by calculating the ratio of the probability of two terms occurring simultaneously to the product of their independent probabilities. This algorithm can effectively identify term pairs with genuine semantic associations while filtering out occasional co-occurrences. Co-occurrence pattern analysis is performed at three granular levels: sentence, paragraph, and document. The system uses a sliding window mechanism to calculate co-occurrence frequencies within different distance ranges, assigning higher weight to closer co-occurrences. The extraction of causal relationship descriptions employs a dependency parsing-based approach. The system uses the Stanford CoreNLP toolkit to construct the dependency parsing tree of sentences and then matches and identifies predefined causal relationship patterns. These patterns include explicit causal markers such as "cause," "lead to," and "due to," as well as implicit causal relationships in subject-verb-object structures. The time series information processing employs time expression recognition and standardization techniques. The system first uses regular expressions and dictionary matching to recognize time expressions, and then uses the TimeML standard to convert relative time into time points on the absolute time axis to construct a time series model of symptom development. This processing method can capture the order and time interval information of symptom appearance during disease progression.
[0062] Specifically, the symptom-disease dynamic association graph is constructed using the NetworkX graph computing framework. Each node in the graph contains attribute information such as term ID, standard name, semantic type, and clinical frequency. During graph construction, the system uses the Louvain community detection algorithm to identify clustering patterns of symptom-disease associations. The Louvain algorithm is a community detection method based on modularity optimization, which can automatically discover closely related node groups in the graph and combine closely related terms to form disease syndromes. The dynamic update mechanism of the graph is based on an incremental learning algorithm. When new literature is added, the system only needs to recalculate the local graph structure involving the new terms, without reconstructing the entire graph. This design significantly improves the system's scalability and real-time performance. Semantic mapping processing adopts a multi-level semantic equivalence network. The professional terminology layer uses MeSH descriptors as the standard representation, the clinical description layer is constructed by mining doctor description texts in electronic medical records, and the patient expression layer is built based on user expression data from online health forums and Q&A platforms. The mapping relationship between the three layers is established by combining Word2Vec word vector similarity calculation and manual verification. Word2Vec is a word embedding technology that can map words to a high-dimensional vector space. Words with similar meanings are close in the vector space. The system establishes cross-layer mapping relationships by calculating the cosine similarity between terms at different levels.
[0063] Specifically, semantic fingerprint generation employs Graph Convolutional Network (GCN) technology, taking the symptom-disease association graph as input and learning low-dimensional representation vectors of nodes through multi-layer graph convolution operations. GCN is a deep learning model specifically designed for graph-structured data, effectively aggregating neighbor information of nodes, making semantically related medical concepts closer in the vector space. This characteristic is particularly suitable for handling complex relationships between concepts in the medical field. The system performs graph convolution calculations on the disease, symptom, treatment, and patient feature nodes involved in each document, obtaining feature vectors of four different dimensions. The main disease category is determined by scoring the importance of disease nodes in the document. The score is calculated based on node degree centrality and the PageRank algorithm. Degree centrality reflects the connectivity of nodes in the graph, while the PageRank algorithm considers the importance propagation effect of nodes. Core symptom cluster identification uses the density clustering algorithm DBSCAN. This algorithm can automatically identify typical symptom combination patterns based on the distribution density of symptom nodes in the graph, without pre-specifying the number of clusters, making it particularly suitable for discovering irregularly shaped symptom groups. Treatment plan type extraction involves constructing a hierarchical treatment ontology, mapping specific treatment measures to upper-level treatment categories. The system uses hierarchical clustering to build a treatment classification system based on the similarity of treatment effects and mechanisms of action. Patient population feature extraction uses named entity recognition technology to extract demographic information from the methodological aspects of the literature, and then obtains feature distributions such as age distribution, gender ratio, and disease severity through statistical analysis. The four-dimensional feature vectors are concatenated and subjected to principal component analysis (PCA) dimensionality reduction to generate a 512-dimensional semantic fingerprint vector. PCA is a linear dimensionality reduction technique that can reduce vector dimensions while preserving key information, improving computational efficiency. The semantic fingerprint retains the core medical semantics of the literature while possessing good computational performance.
[0064] 102. The user query is analyzed for medical context and multimodal features are extracted through a dual encoder architecture. Boolean operations are used to parse the logical relationships in the user query. The symptom-disease dynamic association graph and semantic fingerprint are used for standardization and semantic alignment to obtain a unified retrieval vector.
[0065] In one embodiment of the present invention, the step of performing medical context parsing and multimodal feature extraction on user queries through a dual-encoder architecture, performing Boolean operation nested parsing on the logical relationships in user queries, and performing standardization transformation and semantic alignment using the symptom-disease dynamic association graph and semantic fingerprint to obtain a unified retrieval vector includes: extracting patch-level fine-grained features from the image input in the user query using the ViT-B / 16 encoder in the dual-encoder architecture, and extracting token-level semantic features from the text query in the user query using the BERT encoder in the dual-encoder architecture to obtain multimodal query features; determining the query type of the user query based on the multimodal query features to obtain query context features, and performing Boolean operation nested parsing on the logical relationships in the user query to obtain a structured query logic representation; standardizing the description of patient symptoms in the user query using the symptom-disease dynamic association graph, converting non-standard symptom expressions into standard medical terms to obtain standardized symptom features; and performing semantic alignment processing on the multimodal query features, standardized symptom features, and structured query logic representation based on the semantic fingerprint, and fusing the query context features to generate a unified retrieval vector.
[0066] Specifically, the medical context parsing and multimodal feature extraction of user queries employ a dual-encoder architecture, consisting of a ViT-B / 16 (Vision Transformer Base / 16) encoder and a BERT (Bidirectional Encoder Representations from Transformers) encoder. The two encoders handle visual and textual information respectively. The ViT-B / 16 encoder specifically processes the image input in user queries, including visual information such as medical images, photos of skin lesions, and examination report images. This encoder segments the input image into 16×16 pixel patches, with each patch treated as an independent visual unit. This patch segmentation method borrows the token concept from the Transformer architecture, transforming image processing into a sequence processing problem. Each image patch is transformed into a fixed-dimensional vector representation through a linear projection layer, and then positional encoding information is added to preserve spatial relationships. Through the self-attention mechanism of the multi-layer Transformer encoder, ViT-B / 16 can capture long-distance dependencies between different image patches, extracting high-level semantic information reflecting the features of medical images. Meanwhile, the BERT encoder processes the text portion of the user query, including natural language expressions such as symptom descriptions, medical history information, and treatment needs. BERT uses WordPiece segmentation to decompose the text into tokens, and each token is converted into a vector representation through a pre-trained word embedding layer. BERT's bidirectional encoding mechanism allows it to simultaneously consider the contextual information surrounding the tokens, a feature particularly important for understanding the complex semantic relationships in medical texts. After processing by multiple layers of Transformer encoders, BERT outputs token-level feature vectors containing rich semantic information, integrating multi-level information such as lexical semantics, syntactic structure, and contextual information.
[0067] Specifically, query type determination is based on multimodal query features for classification and identification. This process employs a multi-head classifier to accurately identify query intent. The classifier identifies different query types, such as diagnostic consultation, treatment plan query, drug information retrieval, and prognostic assessment, based on the patterns in the feature vectors. Each type corresponds to a different contextual processing strategy. Diagnostic consultation queries focus on the correlation analysis between symptoms and diseases, treatment plan queries emphasize the effectiveness evaluation of intervention measures, drug information retrieval focuses on the drug's mechanism of action and side effects, and prognostic assessment emphasizes disease development trends and risk factors. Query contextual features include structured information such as query type labels, urgency ratings, and professional field classifications. These features provide important basis for subsequent retrieval strategy adjustments. Boolean operation nested parsing uses a recursive descent parsing algorithm, which can handle complex logical expression structures. The parsing process first identifies logical connectors in the query, including basic logical operators such as AND, OR, and NOT, as well as priority relationships indicated by parentheses. The recursive descent parser constructs an abstract syntax tree according to operator priority. Each node in the tree represents a logical operation or operand, and the leaf nodes correspond to specific medical concepts or conditions. This parser supports complex logical expressions with multiple levels of nesting and can accurately understand compound query conditions such as "(diabetes AND hypertension) OR (heart disease AND NOT kidney disease)". The parsing results form a structured query logic representation that preserves the logical semantics of the original query while providing an operable data structure for subsequent retrieval execution.
[0068] Specifically, the standardization transformation process utilizes a symptom-disease dynamic association graph to standardize patients' symptom descriptions, addressing the retrieval challenges posed by the diversity of terminology in the medical field. Patients often use colloquial expressions when describing symptoms, such as "headache," "stomach ache," and "shortness of breath," which differ significantly from standard medical terminology. The standardization transformation employs a graph matching algorithm to find the standard term nodes in the symptom-disease dynamic association graph that best match the patient's description. The algorithm first calculates the semantic similarity between the patient's symptom description and each symptom node in the graph, based on a weighted combination of cosine distance and edit distance of word vectors. Then, it combines edge weight information in the graph, considering the strength of association between symptoms, to identify the optimal term mapping path. For example, the patient's description of "chest tightness and pain" is mapped to a combination of the standard terms "chest pain" and "shortness of breath," while "dizziness" corresponds to standard descriptors such as "vertigo" and "sweating." This transformation not only standardizes terminology but also supplements relevant symptom information not explicitly expressed by the patient through the association information in the graph. Standardized symptom features are represented in vector form, with each dimension corresponding to a standard medical term. The vector value reflects the importance and confidence level of the symptom in the user's query.
[0069] Specifically, the semantic alignment process employs a multimodal fusion transformer architecture, specifically designed to handle the alignment and fusion of heterogeneous features. The alignment process first projects multimodal query features, standardized symptom features, and structured query logic representations onto a unified semantic space. A learned transformation matrix maps features from different modalities to vector representations of the same dimension. Semantic fingerprints play a crucial role in this process, serving as prior knowledge in the medical field to guide the direction and target of feature alignment. Specifically, semantic fingerprints provide standard positions of medical concepts in the semantic space, and the alignment process aligns query features towards these standard positions, ensuring that semantically similar concepts are close in the vector space. The alignment algorithm uses an attention mechanism to calculate the association weights between different feature components. High-weighted associations represent highly semantically related feature components, requiring more attention during the fusion process. Query context features participate in the fusion process as conditional information, and different fusion strategies are used for different query types. Diagnostic queries focus more on the contribution of symptom features, treatment queries emphasize the importance of logical relationships, and drug queries stress the synergistic effect of multimodal information. The fusion transformer learns the complex interactions between features through a multi-head self-attention mechanism, while the cross-attention layer is responsible for integrating information from different modalities. The resulting unified retrieval vector is a high-dimensional dense vector that encodes the complete semantic information of the user query, including multiple dimensions such as textual semantics, visual features, logical structure, standardized symptoms, and contextual information. The unified retrieval vector possesses good semantic expressive power and computational efficiency, preserving the rich information of the original query while providing optimized data representation for subsequent similarity calculations and retrieval matching.
[0070] 103. Based on the unified retrieval vector and the medical science and technology literature knowledge base, evidence classification retrieval and clinical scenario matching are performed respectively to obtain a preliminary candidate literature set, and fine-grained semantic recalculation is performed on the preliminary candidate literature set to obtain a target candidate literature set.
[0071] In one embodiment of the present invention, the step of performing evidence-level retrieval and clinical scenario matching based on the unified retrieval vector and the medical science and technology literature knowledge base to obtain a preliminary candidate literature set, and performing fine-grained semantic recalculation on the preliminary candidate literature set to obtain a target candidate literature set includes: performing evidence-level retrieval based on the unified retrieval vector and the medical science and technology literature knowledge base, and obtaining evidence-level candidate literature by matching query complexity and literature evidence level; performing clinical scenario matching based on the unified retrieval vector and the medical science and technology literature knowledge base, and obtaining clinical scenario candidate literature by matching patient characteristics and literature clinical scenarios; merging the evidence-level candidate literature and clinical scenario candidate literature, removing duplicate literature, and sorting them by retrieval score to obtain a preliminary candidate literature set; calculating the semantic matching score between each candidate literature in the preliminary candidate literature set and the unified retrieval vector using the ColBERTv2 model, sorting and filtering the top M literatures by score to obtain coarsely ranked candidate literatures; extracting patch-level and token-level features of the coarsely ranked candidate literatures, calculating the maximum similarity matching score between features, re-sorting and filtering the Top-K literatures to obtain the target candidate literature set.
[0072] Specifically, the evidence-based retrieval system employs a dual mechanism of query complexity assessment and literature evidence level matching. This process first performs complexity analysis on a unified search vector, assessing the query's complexity by calculating the distribution patterns and association densities of different semantic components within the vector. The query complexity assessment algorithm analyzes dimensions such as the number of symptom combinations, the depth of disease associations, the nesting level of logical operations, and the degree of fusion of multimodal information within the vector. Simple queries involve single symptoms or clear diagnostic needs, while complex queries involve multiple symptom combinations, rare diseases, multi-system diseases, or complex treatment options. Based on a preset complexity threshold and evidence level mapping relationship, the algorithm matches the query complexity score with the literature evidence level in the medical science and technology literature knowledge base. Literature evidence levels are categorized according to evidence-based medicine standards into different levels, including randomized controlled trials, systematic reviews, cohort studies, case-control studies, and case reports, each corresponding to different evidence strength and clinical credibility. High-complexity queries prioritize systematic reviews and high-quality randomized controlled trials, as these studies possess stronger evidentiary power and broader applicability, providing reliable evidence support for complex clinical problems. Medium-complexity queries are matched with cohort studies and case-control studies, which can provide valuable clinical evidence under specific conditions. Low-complexity queries can obtain useful information from case reports and clinical guidelines. The matching process uses a weighted scoring mechanism, which comprehensively considers factors such as query complexity matching degree, literature quality rating, publication time and citation frequency to generate an evidence grade score for each candidate article, ultimately forming an evidence grade candidate literature set.
[0073] Specifically, clinical scenario matching achieves precise screening through in-depth analysis of the similarity between patient characteristics and clinical scenarios in literature. This process extracts key clinical characteristics of patients from a unified search vector, including multi-dimensional information such as age range, gender, main symptoms, past medical history, medication use, and disease severity. Feature extraction uses a multi-label classifier to analyze different dimensions of the vector, identifying the patient's specific clinical profile. Simultaneously, each article in the medical science and technology literature knowledge base pre-constructs a clinical scenario feature matrix, which describes key information such as the research subjects' demographic characteristics, inclusion and exclusion criteria, disease stage, and treatment plan. The clinical scenario feature matrix is automatically extracted from the methodological section of the literature using natural language processing technology, combined with manual annotation for quality control. The matching algorithm employs multi-dimensional similarity calculation, calculating the degree of matching between patient characteristics and clinical scenarios in literature across demographic features, disease features, and treatment background. Demographic matching considers the overlap of age distribution and the consistency of gender ratio; disease feature matching analyzes the similarity of symptom combinations, disease severity, and complications; and treatment background matching assesses the consistency of factors such as past treatment history, medication experience, and treatment response. The algorithm also incorporates a Bayesian inference mechanism to calculate the credibility of a patient's characteristics belonging to a specific clinical scenario based on the probability distribution of the patient's features. The matching results form a patient-document similarity matrix, where each element represents the matching strength between the corresponding patient's features and the clinical scenario of the document. Candidate documents for clinical scenarios are obtained through threshold screening and ranking.
[0074] Specifically, the candidate literature merging process employs a combined strategy of deduplication and ranking. First, candidate literature for evidence grading and candidate literature for clinical scenarios undergo joint deduplication. The deduplication algorithm performs precise matching based on unique identifiers such as DOI and PubMedID, while simultaneously using literature fingerprinting technology to identify duplicate literature with identical content but different identifiers. Literature fingerprints encode key information such as the title, abstract, and keywords of a document using a hash algorithm, generating a unique digital fingerprint. Identical or highly similar documents have the same or similar fingerprint values. The deduplicated literature set is then ranked according to a comprehensive search score, which integrates multiple evaluation dimensions such as evidence grading score, clinical scenario matching score, literature quality indicators, and timeliness weight. The score fusion uses a weighted linear combination model, with weight parameters dynamically adjusted based on query type and user preferences. Diagnostic queries are given higher weight for clinical scenario matching, review queries emphasize the importance of evidence grading, and treatment queries balance the contributions of both dimensions. The ranked literature set constitutes a preliminary candidate literature set, which ensures both the evidence quality of the literature and a high relevance to the patient's clinical situation.
[0075] Specifically, fine-grained semantic recomputation employs a two-stage ranking optimization strategy. The first stage uses the ColBERTv2 model for coarse ranking. ColBERTv2 is an efficient retrieval model based on BERT, which encodes queries and documents into multiple vector sequences instead of a single dense vector. Specifically, ColBERTv2 processes the unified retrieval vector and the text content of candidate documents separately using a BERT encoder, generating token-level vector representation sequences. Each token corresponds to a vector, and the vector sequences preserve fine-grained semantic information and positional relationships within the text. Similarity calculation uses a maximum similarity matching strategy. For each token vector in the query, the vector with the highest similarity among all token vectors in the document is matched, and then the maximum similarity values of all tokens are summed to obtain the overall matching score. This calculation method can capture the precise semantic correspondence between the query and the document, identifying semantic relevance even when lexical expressions are not entirely identical. The advantage of ColBERTv2 lies in its good balance between computational efficiency and retrieval accuracy. By pre-compiling document vectors and using efficient vector retrieval algorithms, it can quickly complete similarity calculations in large-scale document databases. The initial ranking results are sorted in descending order of matching scores, and the top M articles are selected for the fine ranking stage. The second stage of fine ranking employs deep matching technology using patch-level and token-level features. This technology extracts more detailed semantic features from the initial candidate articles. Patch-level feature extraction is applied to non-textual content such as charts, formulas, and structured data in the articles, converting this content into vector representations through image segmentation and feature encoding techniques. Token-level features perform deeper semantic analysis on the textual content of the articles, including entity relation extraction, semantic role labeling, and sentiment analysis. The maximum similarity matching calculation between features uses a bidirectional attention mechanism, calculating not only the matching degree from query features to document features but also the matching degree from document features to query features, ensuring the completeness and accuracy of semantic correspondence through bidirectional matching. The fine ranking algorithm comprehensively considers multiple factors such as patch matching score, token matching score, and global semantic consistency to generate the final ranking score. Based on a preset threshold, the Top-K articles are selected to form a target candidate article set, which represents the most relevant and highest-quality medical and scientific literature recommendations to the user's query.
[0076] Furthermore, the step of performing evidence-level retrieval based on the unified retrieval vector and the medical science and technology literature knowledge base, and obtaining evidence-level candidate literature by matching query complexity with literature evidence level, includes: obtaining a query complexity score by calculating the semantic depth index, symptom combination complexity index, and diagnostic reasoning level index of the unified retrieval vector; pre-labeling the literature in the medical science and technology literature knowledge base according to randomized controlled trials, systematic reviews, cohort studies, and case reports based on the query complexity score to obtain a graded labeled literature library; matching the query complexity score with literature of each level in the graded labeled literature library according to a preset complexity threshold and evidence level mapping relationship, calculating the selection priority score for each literature to obtain a priority score list; filtering literature with a matching degree exceeding a preset threshold based on the priority score list, and sorting them according to the literature citation frequency to obtain evidence-level candidate literature.
[0077] Specifically, the query complexity score is calculated by extracting three core dimensions of complexity indicators from the unified retrieval vector. This multi-dimensional evaluation method comprehensively reflects the complexity of medical queries and the required strength of evidence. The semantic depth indicator is calculated by analyzing the abstraction level and semantic association depth of medical concepts in the unified retrieval vector. This indicator uses a medical ontology hierarchy for evaluation, calculating the average and maximum depth of the concepts involved in the MeSH tree structure. Specifically, if the query involves basic concepts such as surface symptoms like "fever" and "headache," its semantic depth is shallow, while the semantic depth increases significantly when it involves complex pathophysiological mechanisms, molecular biological processes, or refined diagnostic classifications. The algorithm traverses the activated medical concept nodes in the vector, calculates the path length of each concept from the root node in the ontology hierarchy, and then uses a weighted average method to obtain the overall semantic depth value. The weight allocation is based on the activation strength of the concepts in the vector; concepts with higher activation strength contribute more to the final depth calculation. The symptom combination complexity indicator evaluates the association patterns and combination complexity between symptoms in the query. This indicator is calculated by analyzing the connectivity and clustering patterns in the symptom-disease dynamic association graph. The algorithm first identifies all symptom concepts contained in the vector, and then constructs subgraph structures of these symptoms in the association graph. The complexity of the subgraph is measured by graph theory metrics such as connectivity, clustering coefficient, and shortest path length. Highly interconnected symptom combinations represent complex clinical syndromes that require higher levels of medical evidence, while isolated or weakly associated symptom combinations are relatively simple. The diagnostic reasoning level metric reflects the complexity of medical reasoning involved in the query. This metric is calculated by analyzing the nesting depth of logical operations, the complexity of conditional dependencies, and the length of the reasoning chain. Simple direct symptom-disease correspondences have lower reasoning levels, while queries involving differential diagnosis, multi-step reasoning, and probabilistic inference have higher reasoning levels. The algorithm parses the encoded logical structure in the vector, calculates the nesting level of Boolean operations, the number of conditional branches, and the complexity of reasoning steps. The three dimensions of metrics are standardized and then weighted and fused to obtain a comprehensive query complexity score.
[0078] Specifically, the evidence level pre-labeling of the medical science and technology literature knowledge base employs an automatic classification technology based on document features. This process labels each document in the knowledge base with its evidence level, providing a standardized grading basis for subsequent matching. The pre-labeling algorithm first extracts key features of the documents, including multiple dimensions such as research design type, sample size, research methods, statistical analysis methods, and impact factor of the publishing journal. Research design type identification is achieved through a combination of keyword matching and machine learning classifiers. The algorithm searches for specific research design identifiers in the title, abstract, and methodology sections of the documents, such as "randomized controlled trial" (randomized controlled trial), "systematic review" (systematic review), "cohort study" (cohort study), and "case report" (case report). Simultaneously, a multi-class classifier is trained to automatically classify the document content. The classifier learns text feature patterns for different research types based on the pre-labeled training data. Sample size is extracted from the documents using digital entity recognition technology. The algorithm identifies numerical information describing the number of research subjects, such as numbers near words like "participants," "patients," and "subjects." The identification of statistical analysis methods is achieved through the construction of a statistical method dictionary and pattern matching. Common statistical methods include t-tests, chi-square tests, regression analysis, and survival analysis. Impact factor information of published journals is obtained by matching with a journal database; literature published in high-impact factor journals receives a higher quality rating. These features, after feature engineering, are input into a multilayer perceptron classifier. The classifier outputs a probability distribution for each article belonging to different levels of evidence, and the evidence level label is determined based on the highest probability. The hierarchical labeled literature database is organized according to the evidence pyramid structure of evidence-based medicine, with the highest level being systematic reviews and meta-analyses, followed by randomized controlled trials, then cohort studies and case-control studies, and the lowest level being case reports and expert opinions.
[0079] Specifically, the mapping relationship between complexity thresholds and evidence levels is pre-established based on evidence-based medicine principles and clinical practice experience. This mapping relationship reflects the different requirements for evidence strength for queries of varying complexity. Low-complexity queries involve typical symptoms or standard treatments for common diseases, and these queries can be adequately supported by evidence from case reports, clinical guidelines, or low-level studies. Medium-complexity queries involve combinations of multiple symptoms, differential diagnosis of diseases, or comparison of treatment options, requiring moderate-strength evidence from cohort studies or case-control studies. High-complexity queries involve rare diseases, complex syndromes, innovative treatments, or multi-factor interactions, requiring high-strength evidence from randomized controlled trials or systematic reviews. The mapping algorithm compares the query complexity score with a pre-defined threshold range to determine the complexity category of the query, and then selects the matching evidence level range based on the correspondence. The calculation of the priority score comprehensively considers multiple factors such as the degree of complexity matching, evidence level weight, and literature quality indicators. The degree of complexity matching is measured by calculating the overlap between the query complexity score and the applicable complexity range of the literature; a higher overlap indicates a better match. Evidence level weights are allocated according to the evidence strength levels of evidence-based medicine, with higher-level evidence receiving higher base weights. Literature quality indicators include factors such as journal impact factor, citation count, and publication date. Recently published high-quality literature receives additional quality points. Priority scores are calculated using a weighted linear combination model, and the score list is sorted from highest to lowest priority, providing a quantitative basis for the subsequent screening process.
[0080] Specifically, the selection of candidate literature for evidence grading employs a strategy combining threshold filtering and multidimensional ranking. This process ensures that the selected literature meets both the complexity requirements of the query and possesses high academic and clinical application value. Threshold filtering, based on a preset minimum priority score requirement, filters out literature with low match to the query. The threshold setting considers a balance between query type, user needs, and retrieval accuracy. Diagnostic queries use higher matching thresholds to ensure retrieval accuracy, while exploratory queries appropriately lower the threshold to improve retrieval coverage. Citation frequency ranking uses the academic influence of literature as an important ranking criterion; high citation frequency reflects the recognition of literature in the academic community and its practical application value. The algorithm obtains citation data for each literature, including total citations, average annual citations, h-index, and other indicators. These indicators are normalized and then fused with the priority score. The fusion algorithm uses a dynamic weight allocation mechanism: for newly published literature, the priority score has a higher weight, and the citation frequency has a lower weight; while for literature published a long time ago, the weight of the citation frequency gradually increases. The ranking also considers the timeliness of the literature. Medical knowledge is updated rapidly, and recently published literature has stronger clinical guidance significance. The algorithm adjusts the weight of earlier published literature through a time decay function. The final candidate literature set for evidence grading includes high-quality literature that highly matches the query complexity, has an appropriate level of evidence, and has significant academic influence. These literatures provide a reliable evidentiary basis for subsequent clinical scenario matching.
[0081] Furthermore, the process of matching clinical scenarios based on the unified retrieval vector and the medical science and technology literature knowledge base, by matching patient characteristics with the clinical scenarios of the literature, yields candidate clinical scenario literature. This includes: extracting age range, gender information, main symptoms, past medical history, and medication information from the unified retrieval vector using a named entity recognition algorithm to obtain a structured patient feature vector; constructing a clinical scenario feature matrix for each literature in the medical science and technology literature knowledge base, including the applicable population, disease stage, and treatment plan; obtaining a patient-literature similarity score list by calculating the cosine similarity between the structured patient feature vector and the literature clinical scenario feature matrix; and calculating the clinical applicability credibility of each literature based on the similarity score list and a Bayesian confidence model, and selecting literature with a clinical applicability credibility exceeding a preset threshold to obtain candidate clinical scenario literature.
[0082] Specifically, the construction of structured patient feature vectors employs deep learning-based named entity recognition (NAME) technology, which is specifically optimized for entity extraction tasks in the medical field. The NAME recognition algorithm first decodes the unified retrieval vector, converting the high-dimensional dense vector back into a parsable semantic representation. This process uses a reverse mapping technique, where a trained decoder restores the numerical representation in the vector space to medical concepts and descriptive text. Age range extraction is achieved by recognizing the time and numerical information encoded in the vector. The algorithm searches for age-related semantic patterns, including explicit numerical expressions such as "65 years old" and "45-50 years old," as well as descriptive age expressions such as "middle-aged," "elderly," and "teenager." These descriptive expressions are converted into specific numerical ranges using a pre-constructed age mapping dictionary. The algorithm also considers the ambiguity and cultural differences in different age expressions, using probability distributions to represent the uncertainty of age. Gender information extraction is based on gender-related linguistic identifiers and medical terminology. The algorithm identifies explicit gender-descriptive words and analyzes gender-related medical symptoms and disease patterns, such as gynecological diseases and prostate problems—gender-specific health issues. The main symptom extraction utilizes a medical symptom ontology for semantic matching. The algorithm compares the symptom semantic activation patterns in the vectors with a standard symptom terminology database to identify the core symptom manifestations in the patient's description. Symptom extraction includes not only the symptoms explicitly expressed by the patient but also infers related implicit symptoms through symptom association graphs. The extraction of past medical history focuses on historical information that significantly impacts current health status, such as chronic diseases, major surgical histories, and hereditary diseases. The algorithm identifies the chronological order and duration of disease occurrence through time series analysis. Medication history extraction includes information on current medications, past medications, and drug allergies. The algorithm obtains a complete medication profile through techniques such as drug name recognition, dosage extraction, and medication time analysis. This extracted information undergoes standardization and vectorization encoding to form a multi-dimensional structured patient feature vector. Each dimension of this vector corresponds to a specific patient feature attribute, and the vector value reflects the presence and importance of that feature.
[0083] Specifically, the construction of the clinical scenario feature matrix in the literature is achieved through in-depth analysis of the methodological aspects and research design information of medical and scientific literature. This process requires extracting structured clinical scenario descriptions from unstructured literature texts. The analysis of the study subject descriptions employs information extraction techniques, focusing on descriptions of the demographic, clinical, and inclusion criteria of study participants in the literature. The algorithm extracts a basic profile of the study population by identifying key descriptive phrases and statistical data, such as "mean age," "male-to-female ratio," "disease duration," and "disease severity." The analysis of inclusion and exclusion criteria combines rule extraction and semantic parsing. The algorithm identifies explicitly listed inclusion and exclusion criteria in the literature, which indicate the applicable patient population and inapplicable situations. The parsing of inclusion criteria involves multiple dimensions, including disease diagnostic criteria, symptom severity, age range, and gender requirements, while exclusion criteria include restrictions such as comorbidities, medication conflicts, and special physiological conditions. Clinical trial design analysis identifies the applicable scenarios of the study by recognizing information such as study type, intervention, control setup, and follow-up time. Different study designs correspond to different clinical application conditions and strengths of evidence. The characteristics of the applicable population are determined by comprehensively describing the research subjects and using inclusion and exclusion criteria. The algorithm calculates the distribution of the research population across various characteristic dimensions, such as the mean and standard deviation of age distribution, gender ratio, and the graded distribution of disease severity. Disease stage information is obtained by analyzing descriptions of disease progression, staging, and prognosis in the literature. Different disease stages correspond to different treatment strategies and prognostic expectations. Treatment regimen characteristics include detailed information such as the type of intervention, dosage, duration of treatment, and combination therapy. This information determines the applicability and clinical translational value of the research results. The clinical scenario feature matrix organizes this information in the form of a multidimensional array. The rows of the matrix correspond to different feature dimensions, and the columns correspond to different feature values or feature intervals. The matrix elements represent the values or probability distributions of the corresponding features in the literature.
[0084] Specifically, patient-document similarity calculation uses the cosine similarity algorithm to measure the degree of matching between the structured patient feature vector and the clinical scene feature matrix of the document. This calculation method can effectively handle the similarity comparison problem of high-dimensional sparse vectors. Cosine similarity measures the directional similarity by calculating the cosine value of the angle between two vectors. This method is not affected by the vector length and is particularly suitable for handling feature comparisons with different dimensions and value ranges. In the calculation process, the algorithm first converts the clinical scene feature matrix of the document into a vector representation with the same dimension as the patient feature vector. This conversion process uses feature alignment and interpolation techniques. For continuous features such as age and disease course, the algorithm calculates the degree of overlap between the patient feature values and the feature distribution of the document. The higher the degree of overlap, the greater the similarity. For discrete features such as gender and disease type, the algorithm uses exact matching or fuzzy matching to calculate similarity. The allocation of feature weights is based on the importance of the features to clinical decision-making. Core diagnostic features and key risk factors receive higher weights, while secondary demographic features have relatively lower weights. Similarity calculation also considers the hierarchical structure of features. For example, symptoms can be organized according to organ systems, severity, etc. The algorithm calculates similarity at different levels and performs weighted fusion. The calculation results form a patient-document similarity score list. Each element in the list corresponds to the degree of matching between a document and the current patient. The score ranges from 0 to 1, with a higher score indicating a better match.
[0085] Specifically, the application of the Bayesian confidence model further enhances the accuracy and reliability of clinical applicability assessment. This model calculates the clinical applicability confidence of each article for a specific patient by integrating prior knowledge and observational evidence. The prior probability of the Bayesian model is derived from statistical analysis of large-scale clinical data, reflecting the applicability distribution of different types of articles under general circumstances. The model considers the impact of factors such as study type, sample size, study quality, and publication time on applicability. Randomized controlled trials typically have a higher prior applicability probability, while case reports have a relatively lower prior probability. Observational evidence comes from patient-article similarity scores, which are used as inputs to the likelihood function to update the prior probability. The Bayesian update process uses the Monte Carlo sampling method, which can handle complex probability distributions and high-dimensional parameter spaces. The model also introduces an uncertainty quantification mechanism, not only calculating the confidence value of the point estimate but also providing a confidence interval to represent the degree of uncertainty in the estimate. The calculation of clinical applicability confidence comprehensively considers multiple factors such as patient characteristic matching, study quality assessment, and clinical translation difficulty. High confidence indicates that the research results of the article have strong clinical guidance value for current patients. Threshold screening is based on a balance between clinical safety and efficacy. Setting the threshold too low will include too many irrelevant studies, while setting it too high will miss valuable research. The final set of candidate clinical scenarios represents research that is highly relevant to patients' clinical situations and has strong clinical application value. These studies provide personalized evidence-based support for clinical decision-making.
[0086] 104. Calculate the prior retrieval probability based on the evaluation dimensions of the target candidate document set, and perform knowledge-weighted fusion on the candidate document set based on the prior retrieval probability to obtain a medical and scientific literature recommendation report.
[0087] In one embodiment of the present invention, the step of calculating the prior probability of retrieval based on the evaluation dimension of the target candidate document set and performing knowledge-weighted fusion on the candidate document set based on the prior probability of retrieval to obtain a medical and scientific literature recommendation report includes: calculating the prior probability of each document in the target candidate document set through multi-dimensional evaluation vectors of the evaluation dimension and softmax normalization to obtain a prior probability distribution; performing knowledge-weighted fusion on the target candidate document set based on the prior probability distribution, and obtaining a weighted fusion knowledge representation through prior probability weighting and feature concatenation operations; generating recommendation content from the weighted fusion knowledge representation, converting it into structured recommendation text through a decoder to obtain preliminary recommendation content; and adding citation tags and citation lists to the preliminary recommendation content based on the content source to obtain a medical and scientific literature recommendation report.
[0088] Specifically, the prior probability calculation of the target candidate literature set is based on a multi-dimensional evaluation system. This system comprehensively evaluates each article from multiple perspectives, including evidence quality, clinical relevance, timeliness, and impact. The construction of the multi-dimensional evaluation vector first extracts key quality indicators from each article, including evidence quality dimensions such as study design type, sample size, rigor of statistical methods, impact factor of the publishing journal, and peer review quality. Evidence quality assessment adopts the evaluation criteria of evidence-based medicine, with randomized controlled trials and systematic reviews receiving the highest quality scores, cohort studies and case-control studies receiving moderate scores, and case reports and expert opinions receiving lower scores. The clinical relevance dimension is calculated by analyzing the degree of matching between the literature content and the user query. This dimension considers factors such as disease relevance, symptom matching, treatment applicability, and patient population similarity. The timeliness dimension assesses the guiding value of the publication date of the article to current clinical practice. Medical knowledge is rapidly updated, and recently published articles have stronger timeliness value. The algorithm uses a time decay function to adjust the weights of earlier published articles. The impact dimension measures the influence of literature in the academic and clinical communities through indicators such as citation count, downloads, and social media reach. These dimensions' evaluation values are standardized and combined into a multidimensional evaluation vector, where each component represents the literature's performance level in the corresponding dimension. The Softmax normalization function transforms the multidimensional evaluation vector into a probability distribution. This function ensures that the sum of the prior probabilities of all literature equals 1 through exponential transformation and normalization operations, while maintaining the relative magnitudes of the probability values. The temperature parameter of the Softmax function controls the smoothness of the probability distribution; a higher temperature value produces a smoother probability distribution, while a lower temperature value produces a sharper distribution. The choice of the temperature parameter is adjusted based on the query type and application scenario. The prior probability distribution reflects the degree to which each piece of literature is selected and valued; literature with high prior probabilities receives greater weight and more attention in subsequent knowledge fusion processes.
[0089] Specifically, the knowledge-weighted fusion employs a priori probability-guided feature integration technique. This technique weights and combines the knowledge content from the target candidate literature set according to prior probabilities to form a unified knowledge representation. The weighted fusion process first performs deep feature extraction on each document, using a pre-trained BERT model for the medical field to encode the core content such as the title, abstract, and key conclusions, generating high-dimensional semantic feature vectors. These feature vectors capture key information such as the main medical concepts, research findings, and clinical significance of the documents. The priori probability weighting operation multiplies each document's feature vector by its corresponding prior probability value. The feature contribution of high-probability documents is amplified, while the feature contribution of low-probability documents is reduced. This weighting mechanism ensures that high-quality, highly relevant documents dominate the fusion result. The feature concatenation operation concatenates all weighted document feature vectors in a predetermined order to form a comprehensive feature representation containing information from all candidate documents. The concatenation process uses an attention mechanism to handle the interaction between features. The attention weights are dynamically calculated based on the importance and relevance of the features, with highly relevant feature segments receiving more attention allocation. The fusion algorithm also introduces a redundancy elimination mechanism, which identifies and merges semantically similar feature segments through similarity calculation to avoid interference from duplicate information in the fusion results. The weighted fusion knowledge representation is a high-dimensional dense vector that encodes the comprehensive knowledge content of the candidate document set, preserving the unique contributions of each document while highlighting the core viewpoints and important findings of high-quality documents.
[0090] Specifically, the recommended content generation employs a sequence-to-sequence generative model. This model, based on a Transformer architecture decoder, converts weighted fusion knowledge representations into human-readable structured recommendation text. The decoder uses an autoregressive generation approach, generating the words and sentences of the recommendation report one by one, guided and constrained by the weighted fusion knowledge representation. The generative model has been specifically trained on medical and scientific literature abstracting and report generation tasks, learning language patterns, terminology usage norms, and report structure standards in the medical field. The structured recommendation text is organized according to standard medical report formats, including standard sections such as background introduction, main findings, clinical significance, treatment recommendations, and precautions. The background introduction summarizes basic information about the relevant disease and the current research status, while the main findings section summarizes important research results and clinical evidence from candidate literature. The clinical significance section explains the guiding value of the research findings for clinical practice, and the treatment recommendations section provides specific diagnostic and treatment suggestions based on evidence-based evidence. During the generation process, the decoder dynamically monitors different parts of the weighted fusion knowledge representation through an attention mechanism to ensure that the generated content is consistent with and accurate with the source literature. The model also integrates medical knowledge constraints to avoid generating content that conflicts with medical common sense. The initial recommendations are highly readable and professional, providing healthcare professionals with clear and accurate literature reviews and clinical guidance.
[0091] Specifically, the citation annotation employs automated citation management and annotation technology, ensuring that every viewpoint and conclusion in the recommendation report is supported by clear literature sources. The citation annotation process first establishes a mapping relationship between the recommended content and source literature, identifying the source literature fragments corresponding to each statement in the recommendation report through text similarity calculation and semantic matching technology. The algorithm analyzes the semantic structure of the recommended content, identifying key statements that require citation support, such as research data, clinical findings, treatment effects, and side effect information. For each statement requiring citation, the algorithm searches for the most relevant supporting evidence in the candidate literature set, considering factors such as semantic similarity, content consistency, and strength of evidence in the matching process. The addition of citation markers follows standard academic citation formats, using numerical superscripts or bracket citations to insert citation markers into the recommended content. Citation markers not only identify the source of information but also provide a level of credibility for the evidence; citations of high-quality research are distinguished using special markers. The citation list is organized according to standard bibliographic format, including complete bibliographic information such as author information, article title, journal name, publication year, volume, issue, and page numbers. The citation list also provides a brief summary and key findings for each article, helping readers quickly understand the main content of the cited literature. The algorithm also checks the completeness and accuracy of citations, ensuring that each citation mark corresponds to a relevant bibliographic entry, and that the information in each bibliographic entry is accurate. The final medical and scientific literature recommendation report possesses a complete citation system, guaranteeing both the academic rigor of the content and providing readers with literature clues for further in-depth research. The recommendation report is presented in a standardized medical document format, including a clear chapter structure, standard citation format, and a complete list of references, providing reliable literature reviews and evidence-based support for clinical decision-making and academic research.
[0092] In this embodiment, a medical science and technology literature knowledge base containing semantic fingerprints is established by performing multi-layer semantic annotation on medical science and technology literature and constructing a dynamic symptom-disease association graph. Medical context analysis and multimodal feature extraction are performed on user queries, combined with nested Boolean operation parsing and semantic alignment processing to generate a unified retrieval vector. Based on the unified retrieval vector, evidence-based hierarchical retrieval and clinical scenario matching are performed to obtain a preliminary candidate literature set, which is then optimized into a target candidate literature set through fine-grained semantic recalculation. The prior probability of retrieval is calculated according to the evaluation dimensions, and the candidate literature is subjected to knowledge-weighted fusion to generate a medical science and technology literature recommendation report. This invention deeply understands the association relationships of medical terms through a dynamic symptom-disease association graph, and ensures that the results match the patient's clinical characteristics through evidence-based hierarchical retrieval and clinical scenario matching, thereby improving the accuracy of literature retrieval.
[0093] The above describes the intelligent scientific literature retrieval method based on generative artificial intelligence in the embodiments of the present invention. The following describes the intelligent scientific literature retrieval device based on generative artificial intelligence in the embodiments of the present invention. Please refer to [link to relevant documentation] for details on this intelligent scientific literature retrieval device. Figure 2 One embodiment of the intelligent scientific literature retrieval device based on generative artificial intelligence in this invention includes:
[0094] The semantic annotation module 201 is used to perform multi-layer semantic annotation processing on medical and scientific literature, construct a dynamic correlation map of symptoms and diseases based on the annotated medical terms, and perform semantic mapping on the medical terms to obtain a medical and scientific literature knowledge base containing semantic fingerprints.
[0095] The query parsing module 202 is used to perform medical context parsing and multimodal feature extraction on user queries through a dual encoder architecture, perform Boolean operation nested parsing on the logical relationships in user queries, and perform standardized transformation and semantic alignment using the symptom-disease dynamic association graph and semantic fingerprint to obtain a unified retrieval vector;
[0096] The intelligent retrieval module 203 is used to perform evidence-level retrieval and clinical scenario matching based on the unified retrieval vector and the medical science and technology literature knowledge base respectively to obtain a preliminary candidate literature set, and to perform fine-grained semantic recalculation on the preliminary candidate literature set to obtain a target candidate literature set.
[0097] The fusion recommendation module 204 is used to calculate the prior retrieval probability based on the evaluation dimension of the target candidate document set and perform knowledge-weighted fusion on the candidate document set based on the prior retrieval probability to obtain a medical and scientific literature recommendation report.
[0098] In this embodiment of the invention, the intelligent scientific literature retrieval device based on generative artificial intelligence operates the aforementioned intelligent scientific literature retrieval method based on generative artificial intelligence. This device establishes a medical scientific literature knowledge base containing semantic fingerprints by performing multi-layer semantic annotation on medical scientific literature and constructing a dynamic symptom-disease association graph. It performs medical context analysis and multimodal feature extraction on user queries, combining Boolean operation nested parsing and semantic alignment processing to generate a unified retrieval vector. Based on the unified retrieval vector, it performs evidence-level retrieval and clinical scenario matching to obtain a preliminary candidate literature set, which is then optimized into a target candidate literature set through fine-grained semantic recomputation. The device calculates the prior probability of retrieval based on evaluation dimensions, performs knowledge-weighted fusion of candidate literature, and generates a medical scientific literature recommendation report. This invention deeply understands the association relationships of medical terms through a dynamic symptom-disease association graph, and ensures that the results match the patient's clinical characteristics through evidence-level retrieval and clinical scenario matching, thereby improving the accuracy of literature retrieval.
[0099] above Figure 2 The intelligent scientific and technological literature retrieval device based on generative artificial intelligence in the embodiments of the present invention will be described in detail from the perspective of modular functional entities. The intelligent scientific and technological literature retrieval device based on generative artificial intelligence in the embodiments of the present invention will be described in detail from the perspective of hardware processing.
[0100] Figure 3 This is a schematic diagram of the structure of a generative artificial intelligence-based intelligent retrieval device for scientific and technological documents provided in an embodiment of the present invention. The generative artificial intelligence-based intelligent retrieval device 300 can vary considerably due to different configurations or performance. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors) and a memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) for storing application programs 333 or data 332. The memory 320 and storage media 330 can be temporary or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the generative artificial intelligence-based intelligent retrieval device 300. Furthermore, the processor 310 may be configured to communicate with the storage media 330, executing a series of instruction operations in the storage media 330 on the generative artificial intelligence-based intelligent retrieval device 300 to implement the steps of the aforementioned generative artificial intelligence-based intelligent retrieval method for scientific and technological documents.
[0101] The intelligent scientific literature retrieval device 300 based on generative artificial intelligence may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 3 The illustrated structure of the intelligent scientific and technological literature retrieval device based on generative artificial intelligence does not constitute a limitation on the intelligent scientific and technological literature retrieval device based on generative artificial intelligence provided by the present invention. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0102] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the intelligent retrieval method for scientific and technological documents based on generative artificial intelligence.
[0103] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0104] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0105] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent retrieval of scientific and technological literature based on generative artificial intelligence, characterized in that, The intelligent retrieval method for scientific and technological literature based on generative artificial intelligence includes: Multi-layer semantic annotation is performed on medical science and technology literature, and a dynamic correlation graph of symptoms and diseases is constructed based on the annotated medical terms. Semantic mapping is also performed on the medical terms to obtain a medical science and technology literature knowledge base containing semantic fingerprints. The system employs a dual-encoder architecture to perform medical context parsing and multimodal feature extraction on user queries. It also performs nested Boolean operations to parse the logical relationships in user queries and uses the symptom-disease dynamic association graph and semantic fingerprint for standardization transformation and semantic alignment to obtain a unified retrieval vector. A query complexity score is obtained by calculating the semantic depth index, symptom combination complexity index, and diagnostic reasoning hierarchy index of the unified retrieval vector. Based on the query complexity score, documents in the medical science and technology literature knowledge base are pre-labeled according to evidence level, categorized as randomized controlled trials, systematic reviews, cohort studies, and case reports, resulting in a graded labeled literature library. According to a preset mapping relationship between complexity thresholds and evidence levels, the query complexity score is matched with documents of each level in the graded labeled literature library, and a selection priority score is calculated for each document, resulting in a priority score list. Documents with a matching degree exceeding a preset threshold are selected based on the priority score list and sorted according to citation frequency, resulting in evidence-level candidate documents. The unified retrieval... The vector is matched with the medical science and technology literature knowledge base for clinical scenarios. By matching patient characteristics with the clinical scenarios of the literature, clinical scenario candidate literature is obtained. The evidence grading candidate literature and clinical scenario candidate literature are merged, duplicate literature is removed, and they are sorted by retrieval score to obtain a preliminary candidate literature set. The semantic matching score between each candidate literature in the preliminary candidate literature set and the unified retrieval vector is calculated using the ColBERTv2 model. The top M literatures are sorted by score to obtain coarsely ranked candidate literature. By extracting patch-level and token-level features from the coarsely ranked candidate literature, the maximum similarity matching score between features is calculated. The top-K literatures are re-ranked and selected to obtain the target candidate literature set. The prior retrieval probability is calculated based on the evaluation dimensions of the target candidate literature set, and the candidate literature set is then subjected to knowledge-weighted fusion based on the prior retrieval probability to obtain a medical and scientific literature recommendation report.
2. The intelligent scientific literature retrieval method based on generative artificial intelligence according to claim 1, characterized in that, The process of performing multi-layer semantic annotation on medical and scientific literature, constructing a dynamic symptom-disease association graph based on the annotated medical terms, and semantically mapping the medical terms to obtain a medical and scientific literature knowledge base containing semantic fingerprints includes: Medical terminology is identified and processed in medical science and technology literature. The medical terms are then annotated with entities using a medical ontology knowledge base to obtain the medical terminology annotation results. The medical terms include disease names, symptom descriptions, drug names, and treatment plans. The association strength weight is calculated based on the co-occurrence patterns, causal relationship descriptions, and time series information of symptom and disease terms in the medical terminology annotation results. A graph structure is constructed using symptom and disease terms as graph nodes and association strength weights as edge weights to obtain a dynamic association graph of symptoms and diseases. The medical terminology annotation results are subjected to semantic mapping processing. By establishing a three-layer correspondence between the professional terminology layer, the clinical description layer, and the patient expression layer, the semantic mapping results are obtained. Based on the symptom-disease dynamic association map and semantic mapping results, feature extraction is performed on medical and scientific literature. The main disease categories, core symptom clusters, treatment plan types and patient population characteristics are extracted to generate semantic fingerprints, resulting in a medical and scientific literature knowledge base containing semantic fingerprints.
3. The intelligent scientific literature retrieval method based on generative artificial intelligence according to claim 1, characterized in that, The process involves using a dual-encoder architecture to perform medical context parsing and multimodal feature extraction on user queries, performing nested Boolean operations on the logical relationships within the user queries, and standardizing and semantically aligning the results using the symptom-disease dynamic association graph and semantic fingerprint to obtain a unified retrieval vector, including: The ViT-B / 16 encoder in the dual-encoder architecture extracts patch-level fine-grained features from the image input in the user query, and the BERT encoder in the dual-encoder architecture extracts token-level semantic features from the text query in the user query, thus obtaining multimodal query features. Based on the multimodal query features, the query type of the user query is determined, the query context features are obtained, and the logical relationships in the user query are processed by Boolean operation nested parsing to obtain a structured query logic representation. The symptom-disease dynamic association graph is used to standardize the descriptions of patient symptoms queried by users, converting non-standard symptom expressions into standard medical terms to obtain standardized symptom features. Based on the semantic fingerprint, semantic alignment processing is performed on multimodal query features, standardized symptom features, and structured query logic representation, and a unified retrieval vector is generated by fusing query context features.
4. The intelligent scientific literature retrieval method based on generative artificial intelligence according to claim 1, characterized in that, The process of matching clinical scenarios based on the unified retrieval vector and the medical science and technology literature knowledge base, by matching patient characteristics with the clinical scenarios of the literature, yields candidate literature for clinical scenarios, including: A structured patient feature vector is obtained by extracting age range, gender information, main symptoms, medical history and medication information from a unified retrieval vector using a named entity recognition algorithm. By analyzing the research subject descriptions, inclusion and exclusion criteria, and clinical trial designs of each article in the medical science and technology literature knowledge base, a clinical scenario feature matrix containing applicable populations, disease stages, and treatment plans is constructed for the articles. By calculating the cosine similarity between the structured patient feature vector and the clinical scenario feature matrix of the literature, a list of patient-literature similarity scores is obtained. Based on the similarity score list and combined with the Bayesian confidence model, the clinical applicability credibility of each article is calculated. Articles with clinical applicability credibility exceeding the preset threshold are selected to obtain candidate articles for clinical scenarios.
5. The intelligent scientific literature retrieval method based on generative artificial intelligence according to claim 1, characterized in that, The step of calculating the prior retrieval probability based on the evaluation dimensions of the target candidate document set and performing knowledge-weighted fusion on the candidate document set based on the prior retrieval probability to obtain the medical and scientific literature recommendation report includes: The prior probability distribution is obtained by calculating the prior probability of each document in the target candidate document set through the multi-dimensional evaluation vector of the evaluation dimension and softmax normalization. Based on the prior probability distribution, the target candidate document set is subjected to knowledge weighted fusion. Through prior probability weighting and feature concatenation operations, a weighted fusion knowledge representation is obtained. The weighted fusion knowledge representation is used to generate recommended content, which is then converted into structured recommended text by a decoder to obtain preliminary recommended content. The preliminary recommendations are then citation-marked, and citation tags and lists are added based on the content sources to obtain a medical and scientific literature recommendation report.
6. A scientific and technological literature intelligent retrieval device based on generative artificial intelligence, characterized in that, The intelligent scientific literature retrieval device based on generative artificial intelligence includes: The semantic annotation module is used to perform multi-layer semantic annotation processing on medical and scientific literature, construct a dynamic correlation map of symptoms and diseases based on the annotated medical terms, and perform semantic mapping on the medical terms to obtain a medical and scientific literature knowledge base containing semantic fingerprints. The query parsing module is used to perform medical context parsing and multimodal feature extraction on user queries through a dual encoder architecture, perform nested Boolean operations on the logical relationships in the user query, and perform standardized transformation and semantic alignment using the symptom-disease dynamic association graph and semantic fingerprint to obtain a unified retrieval vector. The intelligent retrieval module calculates a query complexity score by evaluating the semantic depth, symptom combination complexity, and diagnostic reasoning hierarchy of a unified retrieval vector. Based on this score, it pre-labels documents in the medical science and technology literature knowledge base according to their evidence level (randomized controlled trials, systematic reviews, cohort studies, case reports), creating a tiered labeled literature library. Following a pre-defined mapping between complexity thresholds and evidence levels, it matches the query complexity score with documents at each level in the tiered labeled literature library, calculating a selection priority score for each document, resulting in a priority score list. Based on this priority score list, it filters documents with a matching degree exceeding a pre-defined threshold and sorts them by citation frequency, obtaining candidate documents for evidence tiering. The unified retrieval vector is matched with the medical science and technology literature knowledge base for clinical scenarios. By matching patient characteristics with the clinical scenarios of the literature, clinical scenario candidate literature is obtained. The evidence grading candidate literature and clinical scenario candidate literature are merged, duplicate literature is removed, and they are sorted by retrieval score to obtain a preliminary candidate literature set. The semantic matching score between each candidate literature in the preliminary candidate literature set and the unified retrieval vector is calculated using the ColBERTv2 model. The top M literatures are sorted by score to obtain coarsely ranked candidate literature. By extracting patch-level and token-level features from the coarsely ranked candidate literature, the maximum similarity matching score between features is calculated. The top-K literatures are re-ranked and selected to obtain the target candidate literature set. The fusion recommendation module is used to calculate the prior retrieval probability based on the evaluation dimensions of the target candidate document set and to perform knowledge-weighted fusion of the candidate document set based on the prior retrieval probability to obtain a medical and scientific literature recommendation report.
7. A scientific and technological literature intelligent retrieval device based on generative artificial intelligence, characterized in that, The intelligent scientific literature retrieval device based on generative artificial intelligence includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the intelligent scientific literature retrieval device based on generative artificial intelligence to perform the steps of the intelligent scientific literature retrieval method based on generative artificial intelligence as described in any one of claims 1-5.
8. A computer-readable storage medium storing instructions thereon, characterized in that, When the instruction is executed by the processor, it implements the steps of the intelligent scientific literature retrieval method based on generative artificial intelligence as described in any one of claims 1-5.
Citation Information
Patent Citations
Literature review generation method and device, computer equipment and storage medium
CN118939752A
Optimization method for index retrieval of digital library
CN119128045A