AI-driven intelligent retrieval analysis and traceability system for academic information in medical field
The AI-driven intelligent retrieval, analysis, and traceability system for medical academic information solves the problems of insufficient retrieval accuracy, format adaptability, and traceability capabilities of existing systems, achieving efficient and accurate processing and traceability of medical academic information, and improving research efficiency and the reliability of results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI YIWANG NETWORK TECH CO LTD
- Filing Date
- 2025-12-16
- Publication Date
- 2026-05-05
AI Technical Summary
Existing medical academic information processing systems suffer from insufficient retrieval accuracy, poor format compatibility, limited analytical depth, and weak traceability capabilities. This causes medical researchers to spend a significant amount of time screening information, verifying data authenticity, and integrating contradictory conclusions, severely encroaching on their core research time.
The system employs an AI-driven intelligent retrieval, analysis, and traceability system for academic information in the medical field. It includes an academic tracking and collaboration unit, an information format extraction unit, an in-depth interpretation gridded unit, and an academic traceability verification unit. Through multi-source information priority sorting, dynamic format adaptation, multi-dimensional analysis, and a full-link traceability mechanism, it accurately identifies medical terms, monitors format deviations, classifies and analyzes information by sub-domain, and traces the source of information.
It improves the accuracy of medical academic information retrieval and the completeness of format extraction, enhances the professional adaptability and academic rigor of analysis results, ensures the authenticity and traceability of data, and reduces the time cost for researchers in information screening and verification.
Smart Images

Figure CN121980004A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent processing technology for medical academic information, specifically relating to an AI-driven intelligent retrieval, analysis, and tracing system for medical academic information. Background Technology
[0002] With the globalization and interdisciplinary integration of medical research, academic information in the medical field is experiencing explosive growth. Its sources cover multiple channels, including core journals in Chinese and foreign languages, clinical trial registration platforms, medical conference proceedings, and research reports from scientific research institutions. Moreover, the data formats vary significantly, posing a serious challenge to medical researchers in efficiently obtaining reliable information.
[0003] Existing retrieval systems mostly rely on keyword matching or simple field filtering, which are not fully adapted to the characteristics of the medical field. For example, they cannot effectively identify the correspondence between disease aliases (such as "myocardial infarction" and "acute myocardial infarction") and drug generic names and brand names, resulting in a high proportion of redundant information. At the same time, the information priority ranking mostly refers to the impact factor of a single journal, without taking into account the timeliness differences of medical subfields (such as the need to prioritize recent literature in infectious disease studies and the need to consider classic results from afar) and the relevance of the user's research direction, causing high-value information to be buried.
[0004] Existing format extraction tools mostly use fixed rule templates. When the source of medical information (such as journal websites and databases) updates the typesetting format, the extraction rules need to be manually reconfigured, resulting in a strong response lag. Furthermore, they do not distinguish between core fields (such as clinical trial registration numbers, DOIs, and statistical methods) and ordinary auxiliary fields (such as author contact addresses) in medical academic information, leading to omissions or high error rates in the extraction of core fields, which directly affects the reliability of subsequent analysis.
[0005] Existing analysis systems often process medical academic information at the level of topic classification and keyword frequency statistics, without refining the analysis dimensions according to medical subfields (such as targeted therapy in oncology and interventional cardiovascular technology). In terms of model application, they often use general language models to directly process medical data without optimizing terminology coverage and analysis accuracy for subfields. Furthermore, when faced with contradictory conclusions from multiple models, they lack a scientific judgment mechanism that combines medical evidence levels, and the output results are difficult to meet the needs of medical research for a complete interpretation of the "experimental design-statistical data-clinical significance" chain.
[0006] Existing systems mostly store only the final analysis results or some intermediate data, without establishing a full-link data association from "collection-screening-format refinement-analysis". This makes it impossible to trace the original data source of a conclusion or the record of model parameter adjustments. Anomaly identification often relies on single-data-dimensional verification (such as format error detection) without combining medical common sense (such as the reasonable range of clinical indicators) and multi-source cross-comparison (such as the consistency of information of the same clinical trial on different platforms). This makes it difficult to verify the authenticity of academic information and increases the risk of medical research citing incorrect data.
[0007] The aforementioned problems force medical researchers to spend a significant amount of time sifting through information, verifying data authenticity, and reconciling conflicting conclusions, severely impacting core research time and hindering the efficiency and speed of medical research translation. Therefore, developing an academic information processing system that integrates AI technology, is deeply adapted to the characteristics of the medical field, and possesses capabilities for accurate retrieval, intelligent format adaptation, professional in-depth analysis, and end-to-end traceability has become an urgent need in the field of medical research informatization. Summary of the Invention
[0008] To address the aforementioned problems in existing technologies, this invention provides an AI-driven intelligent retrieval, analysis, and tracing system for academic information in the medical field. The objective of this invention can be achieved through the following technical solutions: The AI-driven intelligent retrieval, analysis and traceability system for academic information in the medical field is characterized by including: an academic tracking and collaboration unit, an information format extraction unit, an in-depth interpretation gridded unit and an academic traceability verification unit. The academic tracking and collaboration unit uses a multi-source medical information priority ranking algorithm to obtain the release characteristics of medical information sources and constructs an information priority quantification model based on impact factor weight, release time decay coefficient, and user interest area matching degree; it also obtains the original data of the medical information sources and generates a preliminary screening dataset based on a preset medical terminology database and format verification rules. The information format extraction unit is based on a dynamic format adaptation iterative algorithm to identify the format patterns of basic text fields in the preliminary screening dataset in a hierarchical manner, establish a hierarchical deviation early warning mechanism, and monitor the format deviation value and assign weights. When the deviation value reaches the preset field matching degree threshold, a rule iteration scheme is generated, and the rule validity is verified through a preset historical medical academic data sample library to generate data format features of multi-source medical information. The deep interpretation raster unit classifies the multi-source medical information according to medical sub-domains through an academic granularity weighted aggregation algorithm, and initiates corresponding multi-dimensional raster analysis for different sub-domains; in the weighted aggregation stage, it calculates the historical analysis accuracy and terminology coverage of the large language model for the current sub-domain and generates professionalism weights; when dealing with contradictory conclusions, it introduces a medical evidence level ranking mechanism and outputs academic granularity analysis results. The academic tracing and verification unit receives full-link data and, based on an AI-driven medical academic tracing and verification algorithm, performs multi-source comparison on the academic granularity analysis results to generate an academic tracing and verification report that includes a tracing link map, core data verification results, and anomalies.
[0009] Specifically, during the process of generating preliminary screening results by the academic tracking collaboration unit, the dynamic update mechanism of the medical terminology lexicon is as follows: after associating with authoritative medical terminology databases and releasing new terms or revising existing terminology expressions, a lexicon synchronization command is triggered to import the updated content into the medical terminology lexicon and complete field mapping; a medical semantic similarity matching model is constructed: based on a pre-trained medical language model, a medical corpus including disease alias correspondences, drug generic name and brand name mapping rules, and clinical indicator synonym expression cases are input, and a concept association logic unique to the medical field is generated through iterative training.
[0010] Specifically, in the initial construction phase, the information priority quantification model sets basic values for impact factor weights based on the differences in professional attributes of medical sub-field journals; for the publication time decay coefficient, it introduces the logic of medical information timeliness classification; and for the matching degree of user's field of interest, it constructs a user field profile through multi-dimensional data: collecting user click, collection, and download behaviors of the filtered results, associating the topic clustering results of the user's historical search keywords, and the research direction annotation information of the user's institution, and calculating the fit between the user's field and the literature topic through semantic matching.
[0011] Specifically, the implementation process of the hierarchical deviation early warning mechanism is as follows: First, sort out the functional attributes of the fields in medical academic information, classify the fields that are directly related to the verification of the authenticity of academic data as core fields, and classify the auxiliary explanatory information as ordinary fields; then, set weights according to the degree of influence of the fields on academic analysis, and in the subsequent format monitoring process, provide real-time early warning for the format deviation of the core fields, and provide summary early warning for the deviation of the ordinary fields according to a preset period.
[0012] Specifically, the process by which the information format extraction unit verifies the validity of rules using the preset historical medical academic data sample database is as follows: First, it selects literature samples covering different publication periods, different disciplines, and different journal levels from the preset historical medical academic data sample database; then, it applies the iterated format extraction rules to the literature samples and calculates the extraction accuracy and field completeness of each field; if the extraction accuracy and field completeness do not meet the preset academic analysis requirements, it returns to the adjustment iteration scheme.
[0013] Specifically, the process by which the deep interpretation rasterization unit classifies the multi-source medical information according to medical sub-domains is as follows: It calls the medical subject thesaurus, matches the keywords in the document with the standard terms in the thesaurus, and initially generates the primary medical domain to which the document belongs; based on the research direction description, experimental method keywords, and research object information in the full text of the document, it initiates a keyword clustering algorithm to further divide the primary medical domain into secondary sub-domains; it then associates the academic research database of the secondary sub-domains, extracts research directions and commonly used experimental methods, and generates a sub-domain feature list.
[0014] Specifically, the process of calculating the analytical capability of the large language model for the current subdomain using the deep interpretation rasterization unit is as follows: From the literature analysis verification reports published by the academic community, select literature analysis cases related to the current subdomain that have passed peer review; compare the historical output results of the large language model for the literature analysis cases with the standard conclusions in the verification reports, calculate the similarity, and generate the historical analysis accuracy rate; construct a specialized terminology database for the current subdomain, then extract the analysis output text of the large language model for the current subdomain literature, calculate the proportion of specialized terms contained in the text to the total number of terms in the specialized terminology database, and generate the terminology coverage rate; combine the evaluation results of the historical analysis accuracy rate and the terminology coverage rate to generate the professionalism weight of the large language model in the current subdomain.
[0015] Specifically, the implementation process of the medical evidence ranking mechanism introduced by the deep interpretation grid unit is as follows: first, the research type in the literature is identified, and the evidence type is generated by extracting the research design description in the literature; then, based on the medical evidence ranking mechanism, the corresponding level attributes are assigned to different types of evidence.
[0016] Specifically, the process of deeply interpreting the academic granularity analysis results output by the rasterized unit is as follows: First, the academic granularity analysis results are divided into a basic layer and a core layer according to the importance of the information. The basic layer information is: the title, author name and affiliation, publication time, and journal name are directly extracted from the format-refined literature data. The extraction process of the core layer information is: through a multi-dimensional rasterized analysis mechanism, the experimental design scheme, statistical analysis data, and clinical application suggestions in the literature are deconstructed. At the same time, the credibility rating of the core layer information is marked, and the rating criteria include the academic influence of the journal from which the literature is sourced, the compliance of the research method, and the completeness of the data.
[0017] Specifically, the academic tracing and verification unit manages the entire data chain process as follows: First, when the unit performs an operation, it records the operation process data in real time, including the operation execution subject, operation instruction content, and operation execution timestamp; then, it establishes a correlation index for the data of each link according to the data flow order, including the unique identifier of the data of the previous link, the generation basis of the data of the current link, and the time node of data transmission.
[0018] Specifically, the academic tracing verification unit integrates a tracing link visualization mechanism. The specific implementation process is as follows: a dynamic medical academic tracing knowledge graph is constructed based on the full-link data, and the unit operation data is transformed into graph nodes and related edges. The graph nodes include data collection source identifiers, screening rule IDs, format extraction parameter sets, and model call records. The related edges mark the data flow direction and transformation relationship, and differentiated visual identifiers are used to distinguish academic elements from operation process data.
[0019] Specifically, the academic source tracing verification unit employs a multi-dimensional intelligent medical identification mechanism to identify anomalies: in the multi-source comparison stage, comparison dimensions are divided according to the type of academic information to generate data anomalies; for conclusion statement information, semantic similarity analysis is used to compare the conclusion tendencies of different documents on the same research question to generate conclusion contradiction anomalies; for related information content, the consistency of information across platforms is checked to generate related information anomalies; for the anomalies, an anomaly confidence level is calculated through an anomaly confidence level assessment mechanism, and anomalies with an anomaly confidence level reaching a preset threshold are included in the academic source tracing verification report.
[0020] This invention, based on the deep integration of AI technology and the characteristics of the medical field, addresses the problems existing in current medical academic information processing, such as insufficient retrieval accuracy, poor format adaptability, limited analytical depth, and weak traceability capabilities. Through the synergistic effect of four core units, it produces multi-dimensional beneficial effects, as detailed below: The academic tracking and collaboration unit constructs a quantitative model through a multi-source medical information priority ranking algorithm. It dynamically adjusts the weight of influence factors and time decay coefficients based on differences in medical sub-domains, and optimizes matching parameters based on user domain profiles. This effectively solves the problems of "information overload" and "demand mismatch" in traditional retrieval. At the same time, the dynamically updated medical terminology database and semantic similarity matching model can accurately identify synonyms of disease aliases, drug names, and clinical indicators, avoiding the omission of effective information due to inconsistent terminology, improving the accuracy of preliminary screening results, and reducing the time cost for medical researchers in the information screening process.
[0021] The information format extraction unit features a dynamic format adaptation iterative algorithm and a hierarchical deviation early warning mechanism. This allows for differentiated monitoring of the varying importance of core fields (such as clinical trial registration numbers and DOIs) and ordinary fields in medical academic information. When the format of the information source changes, it can automatically generate iterative rules and verify them against a historical sample database, avoiding the lag and errors of traditional manual format adjustment rules. Compared to existing fixed format extraction schemes, this unit can adapt to the format differences of different journals and types of medical literature, improving the completeness and accuracy of multi-source information format extraction.
[0022] The system deeply interprets the gridded units and performs multi-dimensional analysis according to medical subfields. It dynamically adjusts weights by combining a subfield-specific terminology database with the model's historical accuracy, avoiding the "generalized interpretation" problem of general analysis models for medical subfields. At the same time, the medical evidence ranking mechanism classifies and screens conclusions according to general standards in the field, giving priority to conclusions corresponding to high-level evidence (such as randomized controlled trials and meta-analyses), effectively eliminating low-credibility and contradictory information. The output academic granularity analysis results (including experimental design, statistical data, and stratified clinical recommendations) can directly support the design of medical research protocols and the verification of conclusions, improving the professional adaptability and academic rigor of the analysis results.
[0023] The academic traceability and verification unit uses a full-link data association index and multi-source comparison mechanism to trace information from collection, screening, format extraction to in-depth interpretation. Combined with the visualization of dynamic knowledge graphs, it makes the data flow path and operation records intuitive and traceable. At the same time, the medical multi-dimensional anomaly identification mechanism and confidence assessment can locate problems such as data deviation and conclusion conflict, helping users to trace the root cause of anomalies and implement rectification. This solves the pain points of "link breakage" and "difficulty in locating anomalies" in existing traceability solutions, ensuring the authenticity and traceability of medical academic information. Attached Figure Description
[0024] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.
[0025] Figure 1This is a system architecture diagram for an AI-driven intelligent retrieval, analysis, and traceability system for academic information in the medical field. Figure 2 This is a schematic diagram of the stratification deviation early warning mechanism in this invention; Figure 3 This is a diagram of the backend management system interface in this invention. Detailed Implementation
[0026] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of example embodiments to those skilled in the art. Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this disclosure. The blocks shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices. The flowcharts shown in the drawings are merely illustrative and do not necessarily include all contents and operations / steps, nor do they necessarily need to be performed in the order described. For example, some operations / steps can be broken down, while others can be combined or partially combined. Therefore, the actual execution order may change depending on the actual situation.
[0027] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided.
[0028] Please see Figure 1-3 A medical academic information intelligent retrieval, analysis and traceability system based on AI, characterized by comprising: an academic tracking and collaboration unit, an information format extraction unit, an in-depth interpretation gridded unit and an academic traceability verification unit; The academic tracking and collaboration unit uses a multi-source medical information priority ranking algorithm to obtain the release characteristics of medical information sources and constructs an information priority quantification model based on impact factor weight, release time decay coefficient, and user interest area matching degree; it also obtains the original data of the medical information sources and generates a preliminary screening dataset based on a preset medical terminology database and format verification rules. The information format extraction unit is based on a dynamic format adaptation iterative algorithm to identify the format patterns of basic text fields in the preliminary screening dataset in a hierarchical manner, establish a hierarchical deviation early warning mechanism, and monitor the format deviation value and assign weights. When the deviation value reaches the preset field matching degree threshold, a rule iteration scheme is generated, and the rule validity is verified through a preset historical medical academic data sample library to generate data format features of multi-source medical information. The deep interpretation raster unit classifies the multi-source medical information according to medical sub-domains through an academic granularity weighted aggregation algorithm, and initiates corresponding multi-dimensional raster analysis for different sub-domains; in the weighted aggregation stage, it calculates the historical analysis accuracy and terminology coverage of the large language model for the current sub-domain and generates professionalism weights; when dealing with contradictory conclusions, it introduces a medical evidence level ranking mechanism and outputs academic granularity analysis results. The academic tracing and verification unit receives full-link data and, based on an AI-driven medical academic tracing and verification algorithm, performs multi-source comparison on the academic granularity analysis results to generate an academic tracing and verification report that includes a tracing link map, core data verification results, and anomalies.
[0029] The TIANNA Academic Tracking-Robot Collaboration System is an independently developed artificial intelligence system that intelligently coordinates the collaborative work of three levels of terminals to achieve academic tracking and real-time dynamic retrieval of medical information sources. It can flexibly adapt to the characteristics of multi-source and high-frequency updates of medical academic information.
[0030] The functions and collaboration logic of the three-level terminals are as follows: Task Assignment Terminal (TA): As the core of system collaboration, it is responsible for coordinating and managing data processing terminals and data acquisition terminals, and dynamically assigning tasks. Its built-in intelligent module can learn the release characteristics of different medical information sources in real time (such as the update cycle of core journals, the release patterns of special topic literature, etc.), and adjust the acquisition frequency of data acquisition terminals accordingly; at the same time, it dynamically balances the task load of data processing terminals based on the real-time acquisition volume of data acquisition terminals to avoid overloading a single terminal. The system defaults to deploying at least two task assignment terminals running in parallel to ensure overall operational stability, and supports expanding or shrinking the terminal cluster at any time according to actual business needs, making deployment highly convenient.
[0031] Data Processing Terminal (DA): Receives the raw data output from the data acquisition terminal and uses a machine learning-trained intelligent judgment module to identify and precisely filter the data. Combining a medical terminology database and format verification rules, it removes redundant and erroneous information (such as non-target field literature and incorrectly formatted abstract data), extracts valid academic information, and feeds it back to the task allocation terminal, providing a high-quality data foundation for subsequent information format refinement.
[0032] Data Acquisition Terminal (DC): Responsible for accessing designated medical information sources (such as PubMed, CNKI Medical Sub-database, etc.) via the network and collecting raw data. Its configured intelligent web page verification module can quickly complete the access verification of the information source, ensuring smooth acquisition links. The acquisition process strictly follows the instructions of the task allocation terminal, and submits the raw data to the designated data processing terminal to ensure the orderly flow of data.
[0033] The Data Format Learning and Data Extraction System (DLDE) is an independently developed artificial intelligence system for the precise extraction of medical academic data. It fundamentally solves the problems of low content extraction accuracy and poor format compatibility caused by the differences in formats of multi-source information and the "contamination" of invalid information when searching online using existing LLM (such as ChatGPT and DeepSeek).
[0034] Its core functional implementation logic is as follows: Through its self-learning module, DLDE can autonomously grasp the data format characteristics of different medical information sources (such as official websites of different journals and academic conference platforms), accurately extract specified academic elements—including key fields such as document title, abstract, author and affiliated institution, illustration annotation information, DOI number, and publication time—and effectively remove incidental information unrelated to academic content (such as pop-up advertising text and irrelevant recommendation link descriptions), thus avoiding information "contamination."
[0035] When the data format of the medical information source changes (such as the adjustment of field positions on the journal's official website or the update of the bibliographic format), DLDE's intelligent module will automatically trigger a relearning process to generate new information extraction rules based on the changed data samples. This eliminates the need for extensive manual revisions, improves the system's ability to adapt to format changes, and ensures long-term stable output of accurate structured medical academic data.
[0036] The Academic Gridding System (AGS) is a deep dive into how it uses a multi-version large language model (LLM) as its core and an AI agent as its interaction medium to perform multi-dimensional automated analysis of structured medical academic data refined by DLDE. Through "grid" operations, it gradually refines the granularity of academic data, providing underlying data support for the subsequent construction of a "clinical think tank".
[0037] Multi-model collaboration ensures analytical accuracy: To avoid the analytical bias of a single LLM, AGS integrates multiple LLMs (including different versions of ChatGPT, DeepSeek, etc.) and initiates parallel judgment of multiple models on the same medical academic data. It calculates the historical analysis accuracy and terminology coverage of the large language model for the current subdomain, performs weighted aggregation of the output results of each model, and eliminates contradictory conclusions through collaborative judgment (such as different interpretations of the same clinical trial conclusion by different models), and finally generates highly reliable rasterized analysis results.
[0038] AI Agent optimizes the interactive experience: The system encapsulates operational functions in the form of AI Agents. Medical professionals can trigger raster analysis through natural language interactive commands (such as "extract the experimental grouping scheme of a certain oncology literature" or "analyze the statistical method compliance of a certain drug clinical trial") without complicated operations. This adapts to the usage habits of medical researchers and improves the efficiency of academic analysis.
[0039] Grid-based refinement of academic dimensions: During the analysis process, AGS breaks down academic data dimensions according to the characteristics of medical sub-fields (such as oncology and cardiology)—including research design type (randomized controlled trial / cohort study), experimental sample characteristics, statistical analysis methods, core conclusions and clinical application value, etc., and gradually refines the granularity to the level of "experimental indicator values - interpretation of clinical significance", providing refined data support for subsequent academic traceability verification and clinical application.
[0040] Specifically, during the process of generating preliminary screening results by the academic tracking collaboration unit, the dynamic update mechanism of the medical terminology lexicon is as follows: after associating with authoritative medical terminology databases and releasing new terms or revising existing terminology expressions, a lexicon synchronization command is triggered to import the updated content into the medical terminology lexicon and complete field mapping; a medical semantic similarity matching model is constructed: based on a pre-trained medical language model, a medical corpus including disease alias correspondences, drug generic name and brand name mapping rules, and clinical indicator synonym expression cases are input, and a concept association logic unique to the medical field is generated through iterative training.
[0041] Specifically, in the initial construction phase, the information priority quantification model sets basic values for impact factor weights based on the differences in professional attributes of medical sub-field journals; for the publication time decay coefficient, it introduces the logic of medical information timeliness classification; and for the matching degree of user's field of interest, it constructs a user field profile through multi-dimensional data: collecting user click, collection, and download behaviors of the filtered results, associating the topic clustering results of the user's historical search keywords, and the research direction annotation information of the user's institution, and calculating the fit between the user's field and the literature topic through semantic matching.
[0042] Calculate the priority quantification value P for a single article. This quantification value calculation needs to consider the characteristics of the medical field and introduces three core parameters to construct the formula: firstly, the impact factor weight. It dynamically assigns values based on the characteristics of the medical subdomain, such as the clinical medicine subdomain. Basic medical subfield adopts To adapt to the differentiated needs of different fields for journal authority; secondly, the publication time decay coefficient. Values are assigned based on the timeliness of medical information, such as emergency medicine and infectious diseases, which have high timeliness requirements. In the fields of chronic disease research and medical history, etc. ,and First, ensure that high-value recent literature receives priority attention; second, ensure the relevance of the content to users' areas of interest. The range of values is A value of 1 indicates a perfect match between the literature's topic and the user's research field, while 0 indicates no relevance. Furthermore, the formula must also include the baseline value of the literature's impact factor. (Taken from the latest impact factor of the journal to which the literature belongs), publication time difference of the literature (The difference between the current time and the publication time of the document, in months), and the weight of the relevance of the user's area of interest. Need and satisfy This ensures that the weights are normalized. The final priority calculation formula is: in This is a time decay term; the smaller T is (the newer the literature), the larger this term value is, which meets the special requirements of timeliness for medical information; subdomain differences are addressed through... Specific value assignments, for example, could be set in the subfield of infectious disease studies. Emphasizing a balance between timeliness and journal authority, chronic disease research is designed with... This further highlights the journal's influence.
[0043] To ensure that the prioritization model aligns with users' actual research needs, it is necessary to collect data on the actual application rate of high-priority literature by users. The calculation method is "the number of high-priority documents cited or downloaded by the user / the total number of high-priority documents recommended by the system"; a preset application rate threshold is used. ,like This indicates a discrepancy between the model's recommendations and user needs, and requires adjustment accordingly. Adjust weights (This is a weighting adjustment coefficient used to control the adjustment magnitude), by increasing... Enhance the impact of user domain relevance on priority, and narrow the gap between the model and actual needs.
[0044] Application output: The Task Assignment Terminal (TA) sends acquisition instructions to the Data Acquisition Terminal (DC) based on the ranking of P values. The TA selects the N documents with the highest P values (N is the amount of work that is dynamically adjusted according to the system load) for priority acquisition, ensuring that high-value documents are processed first. At the same time, the load of the Data Processing Terminal (DA) is adjusted according to the distribution characteristics of P values, and documents with high P values are assigned to DA terminals with high processing efficiency, thereby improving the overall data processing efficiency.
[0045] Specifically, the implementation process of the hierarchical deviation early warning mechanism is as follows: First, sort out the functional attributes of the fields in medical academic information, classify the fields that are directly related to the verification of the authenticity of academic data as core fields, and classify the auxiliary explanatory information as ordinary fields; then, set weights according to the degree of influence of the fields on academic analysis, and in the subsequent format monitoring process, provide real-time early warning for the format deviation of the core fields, and provide summary early warning for the deviation of the ordinary fields according to a preset period.
[0046] Specifically, the process by which the information format extraction unit verifies the validity of rules using the preset historical medical academic data sample database is as follows: First, it selects literature samples covering different publication periods, different disciplines, and different journal levels from the preset historical medical academic data sample database; then, it applies the iterated format extraction rules to the literature samples and calculates the extraction accuracy and field completeness of each field; if the extraction accuracy and field completeness do not meet the preset academic analysis requirements, it returns to the adjustment iteration scheme.
[0047] Format bias calculation and iterative process Single-field deviation calculation: To accurately measure the format extraction deviation of different fields, the field weights must first be defined. —Fields directly related to the verification of the authenticity of academic data (such as DOI and clinical trial registration number) are designated as core fields and assigned values. Supplementary information such as author's affiliation and funding projects is divided into ordinary fields and assigned values. ,and This ensures that deviations in core fields receive focused attention. Calculations must be compared with the actual extracted field values. (Field content extracted by DLDE from the information source) and standard field values (Content taken from compliant fields in a historical medical academic data sample database), calculate the local bias of a single field using the formula. : in, This represents the field matching score, with a value range of [0,1]. The lower the matching score, the higher the accuracy. The larger the value, the more intuitively it reflects the degree of format deviation of a single field.
[0048] Global Deviation Local deviation for all fields The weighted sum must include the total number of fields to be extracted during calculation. The formula is: Preset field matching threshold (like (This indicates that a 90% match rate is the compliance standard). This indicates that the global matching degree is lower than The format extraction has significant deviations, requiring rule iteration to be triggered; if If the format extraction is deemed compliant, the process can proceed to the next in-depth analysis stage.
[0049] After generating new formatting rules, M literature samples need to be randomly selected from the historical medical academic data sample database for validation. Validation metrics include extraction accuracy (A) and field completeness (C). Extraction accuracy (A) is calculated as "total number of correctly extracted fields / total number of extracted fields," measuring the correctness of field extraction. Field completeness (C) is calculated as "total number of extracted required fields / total number of required fields in the sample," ensuring that core academic information is not missing. A preset accuracy threshold is set. and integrity threshold ,like If the new rule meets the requirements for medical academic data processing, it can be officially implemented; otherwise, the rule parameters need to be adjusted until the verification is successful.
[0050] Specifically, the process by which the deep interpretation rasterization unit classifies the multi-source medical information according to medical sub-domains is as follows: It calls the medical subject thesaurus, matches the keywords in the document with the standard terms in the thesaurus, and initially generates the primary medical domain to which the document belongs; based on the research direction description, experimental method keywords, and research object information in the full text of the document, it initiates a keyword clustering algorithm to further divide the primary medical domain into secondary sub-domains; it then associates the academic research database of the secondary sub-domains, extracts research directions and commonly used experimental methods, and generates a sub-domain feature list.
[0051] Specifically, the process of calculating the analytical capability of the large language model for the current subdomain using the deep interpretation rasterization unit is as follows: From the literature analysis verification reports published by the academic community, select literature analysis cases related to the current subdomain that have passed peer review; compare the historical output results of the large language model for the literature analysis cases with the standard conclusions in the verification reports, calculate the similarity, and generate the historical analysis accuracy rate; construct a specialized terminology database for the current subdomain, then extract the analysis output text of the large language model for the current subdomain literature, calculate the proportion of specialized terms contained in the text to the total number of terms in the specialized terminology database, and generate the terminology coverage rate; combine the evaluation results of the historical analysis accuracy rate and the terminology coverage rate to generate the professionalism weight of the large language model in the current subdomain.
[0052] Specifically, the implementation process of the medical evidence ranking mechanism introduced by the deep interpretation grid unit is as follows: first, the research type in the literature is identified, and the evidence type is generated by extracting the research design description in the literature; then, based on the medical evidence ranking mechanism, the corresponding level attributes are assigned to different types of evidence.
[0053] Specifically, the process of deeply interpreting the academic granularity analysis results output by the rasterized unit is as follows: First, the academic granularity analysis results are divided into a basic layer and a core layer according to the importance of the information. The basic layer information is: the title, author name and affiliation, publication time, and journal name are directly extracted from the format-refined literature data. The extraction process of the core layer information is: through a multi-dimensional rasterized analysis mechanism, the experimental design scheme, statistical analysis data, and clinical application suggestions in the literature are deconstructed. At the same time, the credibility rating of the core layer information is marked, and the rating criteria include the academic influence of the journal from which the literature is sourced, the compliance of the research method, and the completeness of the data.
[0054] Multi-model result aggregation process Calculate the professionalism weight of the model : This is used to quantify the analytical capabilities of different large language model LLMs for current medical subfields, where the weights of medical-specific LLMs are denoted as follows: The weights of the general enhanced LLM are denoted as Two core metrics need to be incorporated into the calculation: one is the accuracy of the model's historical analysis. The value range is [0,1], and the calculation method is "number of correct analyses in the model's history / total number of analyses", reflecting the reliability of the model's past analyses; the second is the model terminology coverage. The value range is [0,1], and the calculation method is "the number of subdomain-specific terms contained in the model output / the total number of subdomain-specific terminologies", reflecting the model's mastery of medical terminology. An accuracy weighting coefficient is also set. (Values range from [0,1]), which is related to the coverage weighting coefficient. This is used to balance the influence of the two types of indicators. The calculation formula is: Assume that the output of the medical-specific LLM in the single-model analysis results is: The output of the general-purpose enhanced LLM is Weighted aggregation was performed based on the model's level of expertise. Preliminary results... The calculation formula is: Preset result deviation threshold ,like Interpretive bias (such as numerical differences in the "clinical efficacy rate of a certain drug"). This indicates that the model output has high consistency. It can be used directly as an intermediate result; if there is a deviation If so, the dispute needs to be resolved by entering the evidence level determination stage.
[0055] Handling contradictory conclusions (based on evidence level); introducing evidence level weights According to the commonly used evidence grading standards in the medical field, randomized controlled trials have the highest level of evidence and are assigned a value of [value missing]. The second step is queue research, and the assignment is... The case report level is the lowest, and is assigned a value of Let the analytical results supported by different pieces of evidence be as follows: Case reports support the final conclusion. The calculation formula is: Academic granular output: The analysis is broken down into three layers: "Basic Information Layer - Core Data Layer - Conclusion Interpretation Layer." The Basic Information Layer includes information such as the literature title, authors, and affiliated institutions, which do not require in-depth analysis. The Core Data Layer covers experimental design protocols and statistical analysis data (such as sample size, statistical methods, and key indicator values). The Conclusion Interpretation Layer analyzes the clinical application value. Each layer is associated with a corresponding... , The calculation process generates traceable, detailed academic results, providing a clear data framework for subsequent source tracing and verification.
[0056] Specifically, the academic tracing and verification unit manages the entire data chain process as follows: First, when the unit performs an operation, it records the operation process data in real time, including the operation execution subject, operation instruction content, and operation execution timestamp; then, it establishes a correlation index for the data of each link according to the data flow order, including the unique identifier of the data of the previous link, the generation basis of the data of the current link, and the time node of data transmission.
[0057] Specifically, the academic tracing verification unit integrates a tracing link visualization mechanism. The specific implementation process is as follows: a dynamic medical academic tracing knowledge graph is constructed based on the full-link data, and the unit operation data is transformed into graph nodes and related edges. The graph nodes include data collection source identifiers, screening rule IDs, format extraction parameter sets, and model call records. The related edges mark the data flow direction and transformation relationship, and differentiated visual identifiers are used to distinguish academic elements from operation process data.
[0058] Specifically, the academic source tracing verification unit employs a multi-dimensional intelligent medical identification mechanism to identify anomalies: in the multi-source comparison stage, comparison dimensions are divided according to the type of academic information to generate data anomalies; for conclusion statement information, semantic similarity analysis is used to compare the conclusion tendencies of different documents on the same research question to generate conclusion contradiction anomalies; for related information content, the consistency of information across platforms is checked to generate related information anomalies; for the anomalies, an anomaly confidence level is calculated through an anomaly confidence level assessment mechanism, and anomalies with an anomaly confidence level reaching a preset threshold are included in the academic source tracing verification report.
[0059] In this embodiment, the outlier calculation process Step 1: Single-Dimensional Anomaly Detection. Anomaly candidates are identified from three core dimensions: data, conclusions, and related information. Data Dimension: Introducing Data Deviation Values This measures the degree of difference between the data to be verified and the authoritative database data, and the calculation formula is as follows: The data value to be verified. (Standard values from authoritative medical databases), preset data deviation threshold. Mark them as candidate data anomalies; Conclusion dimension: Semantic similarity is used The semantic matching degree between the conclusion to be verified and the conclusion of the same source literature is quantified, with a value range of [0,1], and a preset semantic similarity threshold is set. This indicates that the conclusion conflicts with mainstream research and is marked as a candidate point of conclusion abnormality; Related information dimension: Through the consistency of related information The system determines the matching degree between the information to be verified (such as author institution, clinical trial registration number) and cross-platform information, with a value range of [0,1] and a preset consistency threshold. These are marked as candidate points for abnormal related information.
[0060] Step 2: Anomaly Confidence Calculation. To comprehensively determine the degree of anomaly, an anomaly confidence score is introduced. (The value range is [0,1]), its calculation needs to take into account the outliers in three dimensions and set the dimension weights. (Data Dimensions) (Conclusion dimension) (Related information dimensions), and (For example, in medical settings, priority is given to data accuracy, so...) The final formula is: In the formula, Convert the data deviation values into an indicator that is positively correlated with the confidence level. Similarly, ensuring that the more significant the anomaly in a certain dimension, the better for... The greater the contribution.
[0061] Step 3: Anomaly identification. Preset anomaly confidence threshold. (like ,like This indicates sufficient evidence of anomalies, and the candidate point is confirmed as a formal anomaly and included in the academic source tracing and verification report; if If the result is not found, it is considered "suspected abnormality" and requires further confirmation through manual review by medical experts to avoid misjudgment that could affect the reliability of academic data.
[0062] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. An AI-driven intelligent retrieval, analysis, and traceability system for academic information in the medical field, characterized in that: include: The unit includes an academic tracking and collaboration unit, an information format extraction unit, an in-depth interpretation of the gridded unit, and an academic source verification unit. The academic tracking and collaboration unit uses a multi-source medical information priority ranking algorithm to obtain the release characteristics of medical information sources and construct an information priority quantification model. The original data from the medical information source is obtained, and a preliminary screening dataset is generated based on a preset medical terminology database and format verification rules. The information format extraction unit is based on a dynamic format adaptation iterative algorithm to identify the format patterns of basic text fields in the preliminary screening dataset in a hierarchical manner, establish a hierarchical deviation early warning mechanism, and monitor the format deviation value and assign weights. When the deviation value reaches the preset field matching degree threshold, a rule iteration scheme is generated, and the rule validity is verified through a preset historical medical academic data sample library to generate data format features of multi-source medical information. The deep interpretation raster unit classifies the multi-source medical information according to medical sub-domains through an academic granularity weighted aggregation algorithm, and initiates corresponding multi-dimensional raster analysis for different sub-domains; in the weighted aggregation stage, it calculates the historical analysis accuracy and terminology coverage of the large language model for the current sub-domain and generates professionalism weights; when dealing with contradictory conclusions, it introduces a medical evidence level ranking mechanism and outputs academic granularity analysis results. The academic tracing and verification unit receives full-link data and, based on an AI-driven medical academic tracing and verification algorithm, performs multi-source comparison on the academic granularity analysis results to generate an academic tracing and verification report that includes a tracing link map, core data verification results, and anomalies.
2. The system according to claim 1, characterized in that, The process of generating the preliminary screening dataset by the academic tracking collaboration unit specifically includes: the preset medical terminology lexicon is associated with an authoritative medical terminology database. After adding new terms or revising existing terminology expressions, a lexicon synchronization command is triggered to import the updated content into the preset medical terminology lexicon and complete the field mapping; the construction of the medical semantic similarity matching model is based on a pre-trained medical language model. The input includes a medical corpus containing disease alias correspondences, drug generic name and brand name mapping rules, and clinical indicator synonym expression cases. Through iterative training, a concept association logic unique to the medical field is generated to produce the preliminary screening dataset.
3. The system according to claim 1, characterized in that, The information priority quantification model specifically includes: in the initial construction stage, for the impact factor weight, a basic value is set in combination with the differences in professional attributes of medical sub-field journals; for the publication time decay coefficient, a medical information timeliness classification logic is introduced; and for the matching degree of user's field of interest, a user field profile is constructed through multi-dimensional data: collecting user click, collection, and download behavior of the screening results, associating the topic clustering results of the user's historical search keywords, and the research direction annotation information of the user's institution, and calculating the fit between the user's field and the literature topic through semantic matching.
4. The system according to claim 1, characterized in that, The implementation process of the hierarchical deviation early warning mechanism is as follows: First, sort out the functional attributes of the fields in medical academic information, classify the fields related to the verification of the authenticity of academic data as core fields, and classify the auxiliary explanatory information as ordinary fields; then, set weights according to the degree of influence of the fields on academic analysis, and in the subsequent format monitoring process, provide real-time early warning for the format deviation of the core fields, and summarize and provide early warning for the deviation of the ordinary fields according to a preset period.
5. The system according to claim 1, characterized in that, The process by which the information format extraction unit verifies the validity of rules using a preset historical medical academic data sample library is as follows: First, it selects literature samples covering different publication periods, different disciplines, and different journal levels from the preset historical medical academic data sample library; then, it applies the iterated format extraction rules to the literature samples and calculates the extraction accuracy and field completeness of each field; if the extraction accuracy and field completeness do not meet the preset academic analysis requirements, it returns to the adjustment iteration scheme.
6. The system according to claim 1, characterized in that, The process by which the deep interpretation rasterization unit classifies the multi-source medical information according to medical sub-domains is as follows: It calls the medical subject thesaurus, matches the keywords in the document with the standard terms in the thesaurus, and initially generates the primary medical domain to which the document belongs; based on the research direction description, experimental method keywords, and research object information in the full text of the document, it initiates a keyword clustering algorithm to further divide the primary medical domain into secondary sub-domains; it then associates the academic research database of the secondary sub-domains, extracts research directions and commonly used experimental methods, and generates a sub-domain feature list.
7. The system according to claim 1, characterized in that, The process of calculating the analytical capability of the large language model for the current subdomain using the deep interpretation raster unit is as follows: From the literature analysis verification reports published by the academic community, select literature analysis cases related to the current subdomain that have passed peer review; compare the historical output results of the large language model for the literature analysis cases with the standard conclusions in the verification reports, calculate the similarity, and generate the historical analysis accuracy rate; construct a specialized terminology database for the current subdomain, then extract the analysis output text of the large language model for the current subdomain literature, calculate the proportion of specialized terms contained in the text to the total number of terms in the specialized terminology database, and generate the terminology coverage rate; combine the evaluation results of the historical analysis accuracy rate and the terminology coverage rate to generate the professionalism weight of the large language model in the current subdomain.
8. The system according to claim 1, characterized in that, The implementation process of the medical evidence ranking mechanism introduced by the deep interpretation gridded unit is as follows: First, the research type in the literature is identified, and the evidence type is generated by extracting the research design description in the literature; then, based on the medical evidence ranking mechanism, the corresponding level attributes are assigned to different types of evidence.
9. The system according to claim 1, characterized in that, The process of deeply interpreting the academic granularity analysis results output by the rasterized unit is as follows: the academic granularity analysis results are divided into a basic layer and a core layer according to the importance of information. The basic layer information is directly extracted from the format-refined literature data, including the title, author name and affiliation, publication time, and journal name. The core layer information is extracted through a multi-dimensional rasterized analysis mechanism, which breaks down the experimental design, statistical analysis data, and clinical application suggestions in the literature. At the same time, the core layer information is labeled with a credibility rating, and the rating criteria include the academic influence of the journal from which the literature is sourced, the compliance of the research methods, and the completeness of the data.
10. The system according to claim 1, characterized in that, The academic tracing and verification unit manages the entire data chain process as follows: First, when the unit performs an operation, it records the operation process data in real time, including the operation execution subject, operation instruction content, and operation execution timestamp; then, it establishes a correlation index for the data of each link according to the data flow order, including the unique identifier of the data of the previous link, the generation basis of the data of the current link, and the time node of data transmission.
11. The system according to claim 1, characterized in that, The academic tracing verification unit integrates a tracing link visualization mechanism. The specific implementation process is as follows: a dynamic medical academic tracing knowledge graph is constructed based on the full-link data. The unit operation data is transformed into graph nodes and related edges. The graph nodes include data collection source identifiers, screening rule IDs, format extraction parameter sets, and model call records. The related edges mark the data flow direction and transformation relationship. Differentiated visual identifiers are used to distinguish academic elements from operation process data.
12. The system according to claim 1, characterized in that, The process by which the academic source tracing and verification unit discovers anomalies specifically includes: in the multi-source comparison stage, dividing the comparison dimensions according to the type of academic information to generate data anomalies; for conclusion-statement information, comparing the conclusion tendencies of different documents on the same research question through semantic similarity analysis to generate conclusion contradiction anomalies; for related information content, verifying the consistency of information across platforms to generate related information anomalies; for the anomalies, calculating the anomaly confidence level through an anomaly confidence level assessment mechanism, and including anomalies with anomaly confidence levels reaching a preset threshold in the academic source tracing and verification report.