Oncology science popularization system and method based on large language model
By constructing a tumor science popularization system based on a large language model, the problems of fragmented, obscure, and outdated knowledge in the existing tumor science popularization system have been solved. It has achieved the integration and personalized generation of multi-source knowledge, and improved the systematicness and reliability of tumor science popularization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-03-27
AI Technical Summary
The existing cancer science popularization system lacks a unified knowledge organization framework, uses obscure and difficult-to-understand professional terminology, cannot keep up with medical progress in terms of update speed, lacks personalized interaction, and cannot effectively improve the public's level of cancer prevention and treatment knowledge.
A large language model-based oncology science popularization system is constructed, including a knowledge acquisition and semantic normalization module, an intelligent large-scale model training module, and an intelligent science popularization generation module. By automatically collecting multi-source structured and unstructured knowledge, an oncology semantic knowledge graph is constructed, an oncology-specific large language model is trained, personalized science popularization content is generated, and the accuracy and timeliness of the content are ensured through dynamic updates and a credible review mechanism.
It integrates multi-source knowledge and standardizes semantics, generates authoritative and easy-to-understand popular science content, supports personalized interaction, responds promptly to medical advancements, and improves the systematicness and reliability of oncology popular science.
Smart Images

Figure CN121745274A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and medical information processing technology, and in particular to a tumor science popularization system and method based on a large language model. Background Technology
[0002] In recent years, with the continuous rise in cancer incidence, the public's demand for cancer prevention and treatment knowledge has been increasing. However, the existing cancer science popularization system faces multiple challenges: medical knowledge is scattered across different sources such as research papers, clinical guidelines, news reports, and online communities, lacking a unified knowledge organization framework; professional terminology and academic expressions make key medical information difficult for the general public to understand; the update speed of popular science content cannot keep up with the rapid iteration of clinical research and guidelines; and existing Q&A platforms lack the ability to dynamically perceive and respond to users' cognitive level, focus, and health status.
[0003] Common solutions currently available include online health information platforms, medical knowledge bases, and intelligent question-and-answer systems. These systems aggregate medical news through web pages, build structured medical databases, or provide information services using retrieval-based question-and-answer technology. Some systems also attempt to use templated response mechanisms and keyword matching methods to respond to user queries.
[0004] These existing technologies have significant shortcomings. Information integration is limited; knowledge from different sources remains fragmented, failing to form a complete knowledge system. Content generation relies on fixed templates, making it difficult to adapt to the cognitive characteristics and information needs of different users. Knowledge updates depend on manual maintenance, causing content to lag behind medical advancements. The system lacks the ability to quantitatively assess users' cognitive levels, thus failing to achieve truly personalized interaction.
[0005] In response to the aforementioned technological shortcomings, there is an urgent need to build a technological system capable of integrating multi-source oncology knowledge, achieving unified semantic modeling, and supporting intelligent generation and dynamic updating of popular science content. Such a system needs to possess knowledge fusion, semantic understanding, personalized generation, and continuous evolution capabilities to effectively improve the public's knowledge level in areas such as cancer prevention, screening, diagnosis, treatment, and rehabilitation. Summary of the Invention
[0006] The technical problem to be solved by this invention is to address the shortcomings of existing technologies, specifically by providing a tumor science popularization system and method based on a large language model, as detailed below: 1) In a first aspect, the present invention provides a tumor science popularization system based on a large language model, the specific technical solution of which is as follows: It includes: a knowledge acquisition and semantic normalization module, an intelligent large-scale model training module, and an intelligent science popularization generation module; The knowledge acquisition and semantic normalization module is used to automatically collect structured and unstructured knowledge related to oncology from multiple authoritative data sources and construct an oncology semantic knowledge graph. The intelligent large-scale model training module is used to: train a large language model specifically for oncology based on medical corpus and oncology semantic knowledge graph; The intelligent science popularization generation module is used to generate oncology science popularization content based on the user's input questions, topics of interest, and the user's cognitive level, and then provide it to the user.
[0007] The beneficial effects of the oncology science popularization system based on a large language model provided by this invention are as follows: The knowledge acquisition and semantic normalization module automatically collects structured and unstructured oncology-related knowledge from multiple authoritative data sources and constructs an oncology semantic knowledge graph. This achieves the integration and semantic standardization of multi-source knowledge, eliminates information fragmentation, and establishes a unified knowledge framework. The intelligent large-scale model training module trains a dedicated oncology language model based on medical corpora and the oncology semantic knowledge graph. This enables the model to deeply master oncology knowledge and generate authoritative, accurate, and easily understandable popular science content, effectively overcoming the problem of obscure academic language. The intelligent popular science generation module dynamically generates hierarchical and personalized oncology popular science content based on user-input questions, topics of interest, and user cognitive level using the dedicated oncology language model. This ensures a high degree of matching between the output content and the user's knowledge background and needs, providing an intelligent interactive experience. Overall, through modular collaboration, the system achieves continuous knowledge integration and content optimization, enabling timely responses to medical advancements, improving the efficiency and quality of oncology knowledge dissemination, and helping the public better understand cancer prevention, screening, diagnosis, treatment, and rehabilitation.
[0008] Based on the above scheme, the tumor science popularization system based on a large language model of the present invention can be further improved as follows.
[0009] Furthermore, the knowledge acquisition and semantic normalization module is specifically used to: automatically collect structured and unstructured knowledge related to oncology from multiple authoritative data sources, and identify tumor types, pathological mechanisms, treatment plans, side effects and prognostic entities through named entity recognition and relation extraction technology, and construct an oncology semantic knowledge graph.
[0010] The beneficial effects of adopting the above-mentioned further solutions are as follows: In the field of oncology popularization, there is currently a serious problem of knowledge fragmentation. Cancer prevention and treatment information is scattered across research papers, guidelines, news, and online communities, lacking a unified knowledge framework and semantic standards; the content has poor comprehensibility, with academic language being obscure and difficult to understand; updates are lagging, with popular science content often lagging behind clinical research and guideline updates; and there is a lack of personalization and intelligent interaction, as existing knowledge-based question-and-answer platforms struggle to dynamically adjust content based on users' cognitive levels. The knowledge acquisition and semantic normalization module automatically collects structured and unstructured oncology-related knowledge from multiple authoritative data sources, achieving comprehensive coverage and efficient integration of multi-source knowledge and solving the problems of information dispersion and heterogeneity. By using named entity recognition and relation extraction technologies to identify entities related to tumor types, pathological mechanisms, treatment plans, side effects, and prognosis, the module constructs an oncology semantic knowledge graph, transforming fragmented knowledge into a structured semantic network, establishing a unified knowledge representation framework, and eliminating semantic inconsistencies. This provides a standardized and authoritative knowledge base for subsequent processing, ensuring the accuracy and completeness of the knowledge, supporting the dynamic maintenance and updating of the knowledge, and thus effectively improving the systematicness, reliability and accessibility of oncology science popularization.
[0011] Furthermore, it also includes a knowledge dynamic update module, which is used for: By monitoring the latest oncology papers and clinical treatment guidelines for oncology, the oncology semantic knowledge graph and oncology-specific large language model are automatically updated.
[0012] The beneficial effects of adopting the above-mentioned further solutions are as follows: In the field of oncology popularization, there is currently a problem of lagging knowledge updates. Updates to clinical research findings and treatment guidelines are often difficult to reflect in popular science content in a timely manner, resulting in a gap between the information obtained by the public and the latest medical advancements. By monitoring the latest oncology papers and clinical treatment guidelines, the system achieves automatic updates to the oncology semantic knowledge graph and the oncology-specific large language model. This dynamic update mechanism ensures that the system can continuously capture cutting-edge medical knowledge and promptly integrate new treatment plans, diagnostic criteria, and clinical evidence into the knowledge system. The oncology semantic knowledge graph maintains its integrity and timeliness through incremental updates, while the oncology-specific large language model continuously optimizes its knowledge reserves and generation capabilities through periodic fine-tuning. This automatic update feature effectively solves the problem of traditional popular science content lagging behind clinical developments, enabling the system to consistently provide popular science content that conforms to the latest medical consensus, ensuring the accuracy and cutting-edge nature of the output information, and significantly improving the reliability and practical value of oncology popular science services.
[0013] Furthermore, it also includes a trusted audit and knowledge traceability module, which is used for: Each piece of oncology science popularization content generated by the intelligent science popularization generation module is labeled with the source of knowledge, and receives review data from artificial medicine experts and model error correction feedback data.
[0014] The beneficial effects of adopting the above-mentioned further solutions are as follows: In the field of oncology science popularization, existing systems generally face the problem of insufficient content credibility. The generated content lacks clear source traceability, making it difficult for users to verify the accuracy and authority of the information. Furthermore, the lack of real-time review and error correction mechanisms by medical experts results in erroneous content not being corrected in a timely manner. The credibility review and knowledge traceability module adds knowledge source annotations, such as guideline versions and literature identifiers, to each piece of oncology science popularization content generated by the intelligent science popularization generation module, achieving transparency and traceability of the content and ensuring that each piece of information has clear authoritative basis. The module receives review data from human medical experts and model error correction feedback data, enabling experts to professionally evaluate and correct the generated content. This feedback data is used by the system to optimize the oncology-specific large language model, correct potential errors, and improve content accuracy. This mechanism effectively solves the problems of low content credibility and delayed error correction, enhances the reliability of the system and user trust, and ensures the professionalism and sustainable improvement capabilities of oncology science popularization services.
[0015] 2) Secondly, the present invention also provides a method for popularizing oncology knowledge based on a large language model, the specific technical solution of which is as follows: Structured and unstructured knowledge related to oncology are automatically collected from multiple authoritative data sources to construct an oncology semantic knowledge graph; Training a large language model specifically for oncology based on medical corpus and oncology semantic knowledge graph; Based on the user's input questions, topics of interest, and level of understanding, oncology-specific large language model is used to generate popular science content on oncology and provide it to the user.
[0016] Based on the above scheme, the tumor science popularization method based on a large language model of the present invention can be further improved as follows.
[0017] Furthermore, structured and unstructured knowledge related to oncology was automatically collected from multiple authoritative data sources to construct an oncology semantic knowledge graph, including: We automatically collect structured and unstructured knowledge related to oncology from multiple authoritative data sources, and use named entity recognition and relation extraction techniques to identify tumor types, pathological mechanisms, treatment plans, side effects and prognoses to construct an oncology semantic knowledge graph.
[0018] Furthermore, it also includes: By monitoring the latest oncology papers and clinical treatment guidelines for oncology, the oncology semantic knowledge graph and oncology-specific large language model are automatically updated.
[0019] Furthermore, it also includes: Each piece of oncology science popularization content generated is labeled with the knowledge source and receives review data from artificial medicine experts and model error correction feedback data.
[0020] 3) In a third aspect, the present invention also provides an electronic device, the electronic device including a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor, so as to enable the electronic device to implement any of the above-mentioned methods for popularizing oncology based on a large language model.
[0021] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, it implements any of the above-mentioned methods for popularizing oncology based on a large language model.
[0022] It should be noted that the beneficial effects of the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below: Figure 1 This is a schematic diagram of the structure of a tumor science popularization system based on a large language model according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for popularizing oncology based on a large language model, according to an embodiment of the present invention. Detailed Implementation
[0024] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0025] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.
[0026] like Figure 1 As shown in the figure, an embodiment of the present invention provides a tumor science popularization system based on a large language model, comprising: a knowledge acquisition and semantic normalization module, a knowledge large model training module, and an intelligent science popularization generation module; The knowledge acquisition and semantic normalization module is used to automatically collect structured and unstructured knowledge related to oncology from multiple authoritative data sources, and construct an oncology semantic knowledge graph. Specifically: Structured and unstructured knowledge related to oncology is automatically collected from multiple authoritative data sources. Named entity recognition and relation extraction techniques are used to identify entities related to tumor type, pathological mechanism, treatment plan, side effects, and prognosis, constructing an oncology semantic knowledge graph. The specific implementation process is as follows: The system automatically collects structured and unstructured oncology-related knowledge from multiple authoritative data sources. These sources include international oncology guidelines such as the American National Comprehensive Cancer Network guidelines and the Chinese Society of Clinical Oncology guidelines; research paper databases such as PubMed and CNKI; clinical databases such as hospital electronic health record systems; and statistical reports published by the National Cancer Center. Structured knowledge refers to data stored in predefined formats, such as database tables, XML files, or JSON-formatted knowledge bases. This data contains clearly defined fields and values, such as disease codes, treatment protocols, and patient statistics. Unstructured knowledge refers to text content without a fixed format, such as full-text medical literature, clinical notes, popular science articles, and online forum discussions. The collection process is implemented through a configured web crawler and application programming interface (API), with regular scheduling and incremental updates ensuring that new knowledge is captured promptly.
[0027] The system applies named entity recognition (NER) and relation extraction techniques to process the collected knowledge. NER employs a deep learning-based sequence labeling model to segment and tokenize the text, assigning entity labels to each token. The label set includes categories such as tumor type, pathological mechanism, treatment plan, side effects, and prognosis. For example, in the input text, the model identifies "lung cancer" as a tumor type entity, "EGFR mutation" as a pathological mechanism entity, "chemotherapy" as a treatment plan entity, "nausea" as a side effect entity, and "five-year survival rate" as a prognosis entity. Relation extraction uses a neural network classification model, taking entity pairs in a sentence as input, analyzing context and syntactic dependencies, and outputting the semantic relationship type between entities. For example, the "treatment" relation connects the treatment plan entity and the tumor type entity, and the "cause" relation connects the treatment plan entity and the side effect entity.
[0028] Then, the system constructs an oncology semantic knowledge graph using the extracted entities and relations. The knowledge graph is represented by a graph structure, where nodes represent entities and edges represent relations. The system converts entities and relations into RDF triples; for example, a triple consists of the subject "chemotherapy," the predicate "treatment," and the object "lung cancer." The construction process includes entity parsing and knowledge fusion, merging identical entities from different data sources and resolving semantic conflicts. The knowledge graph is stored in a graph database and integrates hierarchical classifications such as tumor types encoded according to the International Classification of Diseases (ICD) oncology album, ultimately forming a unified, queryable semantic network.
[0029] The term "multiple authoritative data sources" refers to a series of information sources widely recognized by the medical community. These include international oncology guidelines such as the American National Comprehensive Cancer Network guidelines and the Chinese Society of Clinical Oncology guidelines; research paper databases such as PubMed and CNKI; clinical databases such as hospital electronic health record systems; and cancer statistical reports published by government agencies such as the annual report of the National Cancer Center. These data sources provide the latest and most reliable oncology knowledge, ensuring the authority and timeliness of the information collected by the system, and laying the foundation for subsequent processing.
[0030] Structured knowledge refers to data organized in a standardized format that facilitates computer parsing and processing. Structured knowledge includes tabular data in databases, standardized medical codes such as the International Classification of Diseases, 10th Revision (ICD-10), treatment protocol documents, records in drug databases, and clinical trial registration information. This data typically contains clearly defined fields and values, such as patient age, tumor stage, treatment methods, and efficacy assessments, and can be directly used for knowledge graph construction without complex parsing.
[0031] Unstructured knowledge refers to text content without a fixed format, which requires natural language processing (NLP) technology for parsing. Unstructured knowledge includes full-text medical research papers, clinicians' notes, patient education materials, popular science articles, social media posts, and discussions on online health forums. These texts contain rich medical information, but need to be transformed into a structured form through extraction techniques before they can be effectively utilized by the system.
[0032] Tumor types refer to cancer categories classified based on histological, anatomical location, and molecular characteristics. Tumor types include specific disease entities such as lung cancer, breast cancer, colorectal cancer, and leukemia. In named entity recognition, tumor type entities are identified as the names or codes of specific cancers, used to represent different disease categories in the knowledge graph, supporting precise knowledge organization and querying.
[0033] Pathological mechanisms refer to the biological processes involved in the occurrence and development of cancer. These mechanisms include descriptions of gene mutations, abnormal signaling pathways, uncontrolled cell proliferation, and metastasis. In named entity recognition, pathological mechanism entities are identified as terms such as EGFR mutation or angiogenesis, used to explain the etiology and progression of cancer and enhance the semantic depth of the knowledge graph.
[0034] In this context, "treatment plan" refers to medical interventions used to treat cancer. Treatment plans include specific methods such as surgery, chemotherapy, radiotherapy, targeted therapy, immunotherapy, and palliative care. In named entity recognition, treatment plan entities are identified as names such as lobectomy or PD-1 inhibitors to represent available treatment options and are associated with other entities in the knowledge graph.
[0035] Side effects refer to any adverse reactions or side effects that may occur during treatment. These include symptoms such as nausea, vomiting, hair loss, bone marrow suppression, and cardiotoxicity. In named entity recognition, the "side effect" entity is identified as a term describing these side effects, used to assess the safety and risks of treatment, and to establish a connection between the side effect and the treatment plan within a knowledge graph.
[0036] In this context, prognostic entities refer to factors related to disease outcomes. Prognostic entities include concepts such as survival rate, recurrence risk, quality of life indicators, and prognostic prediction models. In named entity recognition, prognostic entities are identified as terms such as five-year survival rate or recurrence score, used to provide predictions of disease progression and support decision support functions within the knowledge graph.
[0037] In another feasible approach, the process of constructing an oncology semantic knowledge graph is as follows: 1) The system automatically collects structured and unstructured oncology-related knowledge from multiple authoritative data sources and performs multi-source data fusion processing. A distributed web crawler is deployed to periodically access oncology guideline publishing platforms, research paper databases, clinical data warehouses, and the National Cancer Center's statistical system to collect structured knowledge such as database tables and JSON-formatted records, as well as unstructured knowledge such as full-text medical literature and clinical text notes. During the collection process, data cleaning techniques are used to remove duplicate and invalid content, and semantic mapping methods are used to convert heterogeneous data into a unified format, achieving the integration and standardized storage of multi-source knowledge and providing a consistent data foundation for subsequent processing. Specifically: First, a distributed web crawler is deployed, consisting of multiple crawler nodes, each responsible for accessing specific authoritative data sources. These authoritative data sources include oncology guideline publishing platforms such as the American National Comprehensive Cancer Network (ANCN) and the Chinese Society of Clinical Oncology (CSCO) guideline websites; research paper databases such as PubMed and IEEE Xplore; clinical data warehouses such as hospital electronic health record systems; and national cancer center statistical systems such as the National Cancer Institute (NCI) database. The crawler is configured with a scheduled task to execute collection tasks periodically, such as performing a full collection once a day at midnight and triggering incremental collection when data source updates are detected. For structured knowledge, the crawler directly obtains database tables and JSON-formatted records through application programming interfaces (APIs) or database query interfaces, such as extracting patient treatment record tables from clinical data warehouses and downloading standardized recommendation protocol JSON files from guideline platforms. For unstructured knowledge, the crawler uses web scraping techniques to obtain full-text medical literature and clinical text notes, such as downloading PDF documents from research paper databases and exporting doctor's notes text from hospital systems. During the collection process, the crawler handles various network protocols and data formats to ensure complete capture of the target knowledge.
[0038] The collected raw data then enters the data cleaning stage. The system employs rule-based and machine learning-based data cleaning techniques to remove duplicate content and invalid information. Duplicate content detection uses a hash function to calculate the fingerprint of each data block, comparing fingerprint values to identify duplicate entries; for example, calculating the MD5 hash value for document abstracts. Invalid content filtering is achieved through a predefined rule base containing a list of keywords and grammatical patterns in the field of oncology, identifying and removing irrelevant advertising text, malformed data, and low-quality fragments. Data cleaning also includes standardization processes, such as unifying date formats to the ISO 8601 standard and converting character encoding to UTF-8, ensuring data consistency and processability. The cleaned data is temporarily stored in a buffer database for subsequent processing.
[0039] The system employs a semantic mapping method to convert heterogeneous data into a unified format. Semantic mapping is based on an ontology model that defines core concepts and relationships in the field of oncology; for example, mapping "lung cancer" to the standard medical term "malignant lung tumor." The mapping process uses natural language processing techniques, including word vector embedding and contextual analysis, to align terms from different data sources to corresponding nodes in the ontology model. For instance, the term "non-small cell lung cancer" extracted from research papers is mapped to the standard concept "NSCLC" through word vector similarity calculation, and the term "chemotherapy" extracted from clinical notes is mapped to the standard treatment regimen "chemotherapy" through semantic role labeling. The mapped data is then converted into a unified resource description framework format, where each entity and relationship is represented as a triple, such as subject "lung cancer," predicate "screening method," and object "low-dose spiral CT."
[0040] Finally, the system integrates and standardizes the storage of multi-source knowledge. The integration process uses data fusion algorithms to resolve conflicts and redundancies between different data sources. For example, when multiple guidelines have different recommendation levels for the same treatment plan, the system calculates a weighted score based on the authority and timeliness of the data source and selects the most reliable version. Standardized storage stores the fused knowledge in a central knowledge base, which adopts a graph database architecture to support efficient querying and updating. The storage format includes entity tables, relation tables, and metadata tables. Entity tables record the attributes of entities such as tumor type and pathological mechanism; relation tables record the semantic relationships between entities; and metadata tables record data sources and timestamps. The entire implementation process ensures that multi-source knowledge is integrated into a consistent and standardized data foundation, providing high-quality input for subsequent named entity recognition and relation extraction.
[0041] 2) The system applies deep learning-based named entity recognition and relation extraction techniques to identify entities related to tumor type, pathological mechanism, treatment plan, side effects, and prognosis. It uses pre-trained large language models, such as the medical-optimized Transformer architecture, to segment and label the fused text, accurately identifying entity boundaries and categories. Subsequently, a relation extraction model, based on contextual semantic analysis and attention mechanisms, parses the semantic relationships between entities, such as treatment correspondences, side effect triggering relationships, or prognostic impact relationships, outputting a structured list of entity relation triples.
[0042] The fused text is processed using a pre-trained large language model. This pre-trained model employs a medically optimized Transformer architecture, incorporating multi-layered self-attention mechanisms and feedforward neural networks, specifically fine-tuned on oncology corpora to enhance domain adaptability. The model receives text input after multi-source data fusion, performs word segmentation to divide the text into sub-word units or lexical units, and then performs sequence labeling. Sequence labeling uses a conditional random field layer as the output layer, assigning entity category labels to each lexical unit. The label set includes categories such as tumor type, pathological mechanism, treatment plan, side effects, and prognosis entities. Entity boundary detection is achieved by labeling start and internal labels in the sequence; for example, consecutive lexical units are labeled as the start and continuation parts of tumor type entities, thus accurately identifying the start and end positions of entities. Model training uses a cross-entropy loss function to optimize parameters. The loss function formula is: in, This represents the named entity recognition loss value. Indicates the number of training samples. Indicates the sequence length. Indicates the first The first sample The real label for each location This represents the probability distribution predicted by the model. Through iterative training, the model can accurately identify entity boundaries and categories in text, providing a foundation for subsequent relation extraction.
[0043] After named entity recognition, the system employs a relation extraction model to parse the semantic relationships between entities. The relation extraction model is based on an encoder-decoder architecture. The encoder uses the same medical domain-optimized Transformer architecture to encode the input text, generating context-aware vector representations. The decoder integrates a multi-head attention mechanism to calculate semantic relevance weights between entity pairs. For each entity pair, the model extracts text fragments within its context window and focuses on key semantic information through attention weights; for example, in "chemotherapy causes nausea," the attention mechanism strengthens the weight of the word "causes." Relation classification uses a softmax function to output relation type probabilities, including treatment-related relations, side-effect-causing relations, and prognostic-affecting relations. The relation extraction model is trained using a negative sampling strategy, generating negative examples from unrelated entity pairs. The loss function is defined as: in, This indicates the loss value extracted from the relation. Indicates the quantity of entity pairs. This represents the score of positive sample entity pairs. This represents the score of the negative sample entity pair. This represents the sigmoid function. The model outputs a structured list of entity-relation triples, where each triple is a subject entity, a relation type, and an object entity. For example, the subject entity is "chemotherapy," the relation type is "side effect," and the object entity is "nausea." The entire implementation process ensures the accuracy and consistency of entity and relation extraction, providing reliable structured data for constructing an oncology semantic knowledge graph.
[0044] 3) Knowledge verification and conflict resolution are performed on the identified entities and relationships to ensure the accuracy and consistency of knowledge. The system integrates a knowledge verification module, which calculates the confidence score for each entity by comparing the frequency of occurrence and contextual consistency of entities across different data sources. For relationship conflicts, decision rules based on evidence level and authority are adopted, such as prioritizing information from high-impact factor papers or the latest clinical guidelines while covering lower-priority sources, achieving reliable knowledge integration and error correction. The specific implementation process is as follows: First, a confidence score is calculated for each entity by comparing its frequency of occurrence and contextual consistency across different data sources. Frequency of occurrence counts the number of times an entity is mentioned in multiple authoritative data sources, while contextual consistency analyzes the semantic stability of the entity across different textual environments. The confidence score is calculated using the following formula: in, This represents the entity confidence score. Indicates the frequency of an entity's occurrence. Indicates the context consistency score. and Represents the weighting coefficients. Context consistency score. It is obtained by calculating the semantic vector similarity of entities in different documents, and the cosine similarity is used to measure the consistency of entity context representation.
[0045] For relationship conflicts, the system employs decision rules based on evidence level and authority for analysis. Evidence level is dynamically assigned based on data source type and publication time, with the latest clinical guidelines and high-quality randomized controlled trials receiving higher levels, and observational studies and case reports receiving lower levels. Authority assessment considers the academic influence of the publishing institution, with internationally renowned medical institutions and core journals receiving higher weight. The relationship conflict analysis formula is as follows: in, Indicates the final relationship adopted. Represents the set of candidate relations. Indicate the relationship corresponding to the level of evidence. Indicates the authority score of the relationship source. and This represents the weighting coefficient. The system prioritizes information from high-impact factor papers or the latest clinical guidelines, automatically overriding conflicting information from lower-priority sources.
[0046] The knowledge verification module performs knowledge integration and error correction. The integration process establishes entity mapping relationships, merging identical entities from different sources and resolving naming differences and terminology variations. Error correction is achieved through a confidence threshold; when the confidence score of an entity or relationship falls below a preset threshold, the system marks it as pending review and triggers a manual review process. The knowledge verification module also maintains version control records, tracking the evolution history of entities and relationships to ensure the traceability of knowledge updates. The entire implementation process involves multiple rounds of iterative verification, continuously optimizing confidence calculation parameters and decision rule weights to improve the accuracy and reliability of knowledge verification.
[0047] 4) Based on the validated entities and relationships, an oncology semantic knowledge graph is constructed, and graph optimization and maintenance are performed. The system utilizes graph database technology, using entities as nodes and relationships as edges to construct a directed weighted graph structure, where weights represent relationship strength or confidence levels. Graph embedding algorithms are applied to learn node vector representations, enhancing the graph's semantic query and reasoning capabilities. Periodic graph update operations, such as node merging and relationship adjustments, are performed to reflect new knowledge and maintain the graph's integrity and timeliness. Specifically: A directed weighted graph structure is established using graph database technology. The graph database employs a native graph storage engine, mapping each entity to a graph node. Node attributes include entity identifier, entity type, confidence score, and source information. Relationships are represented as directed edges connecting related entities, with edge attributes including relationship type, weight value, and source of evidence. The weight value represents the relationship strength or confidence level and is calculated using the following formula: in, Represents relation weights. This represents the entity confidence score. Represents the confidence score of the relationship. and This represents the adjustment coefficient. The graph database establishes an index structure to optimize the query efficiency of nodes and edges, supporting fast retrieval and path traversal operations based on attribute values.
[0048] Then, the system applies a graph embedding algorithm to learn node vector representations. Employing a random walk-based graph embedding method, multiple rounds of random walks are first performed on the knowledge graph to generate node sequence training samples. Then, a Skip-gram model is used to learn the node vector representations, with the objective function defined as: in, This represents the graph embedding loss function. Represents a set of nodes. Represents a node The set of neighboring nodes, Indicates at a given node Nodes appear under certain conditions The probability of similarity is calculated. The learned node vectors preserve the topological structure and semantic relationships of the graph, supporting similarity calculation and semantic reasoning.
[0049] The system establishes a regular graph update and maintenance mechanism. The update process includes node merging and relation adjustment operations. Node merging is based on semantic similarity calculation; a merge operation is performed when the vector similarity between two nodes exceeds a threshold. in, Represents a node and nodes similarity, and These represent the node embedding vectors. Relationship adjustment updates the edge weights and types based on newly acquired knowledge, removes outdated relationships, and adds newly discovered relationships.
[0050] Graph optimization also includes integrity checks and timeliness maintenance. Integrity checks verify the connectivity and consistency of the graph, identifying isolated nodes and contradictory relationships. Timeliness maintenance automatically marks and updates outdated information by setting knowledge expiration thresholds. The system periodically performs graph reconstruction operations, recalculating node embedding vectors to optimize the semantic representation and query performance of the graph. The entire implementation process ensures that the oncology semantic knowledge graph always remains accurate, complete, and timely.
[0051] The intelligent large-scale model training module is used to: train a large language model specifically for oncology based on medical corpus and oncology semantic knowledge graph; The oncology-specific large language model is adapted to the general large language model architecture for domain-specific application. It enhances the model's knowledge representation and generation capabilities in the oncology field by integrating medical corpora and an oncology semantic knowledge graph. The model structure adopts a variant of the Transformer architecture, specifically a multi-layer decoder structure containing an embedding layer, multiple Transformer blocks, and an output layer. The embedding layer converts the input text into a high-dimensional vector representation; each Transformer block consists of a multi-head self-attention sublayer and a feedforward neural network sublayer, and integrates layer normalization and residual connections to improve training stability; the output layer generates word probability distributions through linear transformation and a softmax function. To integrate the oncology semantic knowledge graph, a knowledge injection module is added after the embedding layer. This module uses a graph neural network to encode entities and relations in the knowledge graph and fuses the encoded knowledge vectors with the text representation, dynamically adjusting the knowledge weights through an attention mechanism.
[0052] The training process consists of two main phases: supervised fine-tuning and human feedback reinforcement learning. In the supervised fine-tuning phase, the model is trained using paired data constructed from medical corpora and an oncology semantic knowledge graph. The input includes oncology-related text and corresponding knowledge graph query results, and the output is the target popular science text. The loss function used is cross-entropy loss, with the following formula: in, Indicates the loss value. Indicates the number of training samples. Indicates the sequence length. Indicates the first The first sample One word, Represents conditional probability. Indicates the first Each sample corresponds to a knowledge graph embedding vector. Training optimizes model parameters using gradient descent, iteratively updating them to minimize the loss.
[0053] In the human feedback reinforcement learning phase, firstly, rating data from medical experts on the content generated by the model is collected to construct a feedback dataset; then, a reward model is trained. This model is based on a neural network structure, with the input being the model-generated text and reference text, and the output being a scalar reward value. The reward model is trained using mean squared error loss, with the formula as follows: in, This represents the loss of the reward model. Indicates the number of feedback samples. This represents the actual reward value given by human experts. This represents the reward value predicted by the reward model. Finally, the oncology-specific large language model is fine-tuned using a proximal policy optimization algorithm. The optimization objective is to maximize the expected reward, and the objective function formula is: in, Indicates the optimization objective. Indicates model parameters, Represent the status (input text and knowledge). Indicates an action (generating text). Represents the empirical distribution. Indicates the current strategy. This represents the reference strategy (the model after supervised fine-tuning). Represents the reward function, Represents the KL divergence coefficient. This represents the KL divergence. The training process involves multiple iterations to ensure the accuracy, understandability, and authority of the model's output.
[0054] The medical corpus refers to a collection of text data gathered from the medical field, specifically designed for training and optimizing large language models. This corpus includes clinical guidelines for oncology, full-text research papers, chapters from medical textbooks, patient education manuals, textual descriptions from hospital electronic health records, medical conference abstracts, and popular science materials published by public health institutions. These texts cover multiple aspects of oncology, such as disease diagnosis, treatment protocols, drug instructions, side effect management, and prognostic assessment, providing rich domain-specific language patterns and knowledge context to support the model in learning the correct usage of medical terminology and concepts.
[0055] In another feasible approach, the process of training a large language model specific to oncology includes: 1) Constructing a knowledge-enhanced multimodal training dataset. The system extracts oncology-related text sequences from medical corpora, including clinical guidelines, research papers, and patient education materials, while accessing entity nodes and relation edges in the oncology semantic knowledge graph. Entity linking technology is used to precisely match mentions in the text with corresponding entities in the knowledge graph, and graph coding algorithms are used to generate vector representations of the knowledge subgraph. The training sample set consists of text-knowledge pairs; each sample includes the input text, the corresponding knowledge graph embedding vector, and the target output text, ensuring data diversity and semantic richness, providing a structured foundation for model training. The specific implementation process is as follows: Text sequences related to oncology were collected from multiple authoritative medical data sources, including clinical guidelines from the American National Comprehensive Cancer Network, guidelines from the Chinese Society of Clinical Oncology, full-text research papers in the PubMed database, and patient education materials provided by medical institutions. The collected text sequences underwent data cleaning and standardization, removing irrelevant characters and formatting tags, unifying the text encoding format, and segmenting the documents into paragraphs suitable for model input. Metadata information for each text paragraph, including source identifier, publication time, and topic classification, was fully preserved and indexed.
[0056] The system performs an entity linking process, precisely matching medical concepts in the text with entity nodes in the oncology semantic knowledge graph. The entity linking technology employs a deep learning-based entity disambiguation method. First, a named entity recognition model is used to identify mentions of medical concepts in the text. Then, the semantic similarity between each mention and candidate entities in the knowledge graph is calculated. The similarity calculation combines contextual embedding vectors and entity attribute features, using the following formula to calculate the matching score: in, Indicates text reference With knowledge graph entities Match score, and These represent the context embedding vectors of mentions and entities, respectively. Represents the cosine similarity function. Indicates similarity based on entity attributes. and These are weighting coefficients. For each text mention, the system selects the entity with the highest matching score as the link result and records the confidence level.
[0057] After entity linking is completed, the system uses a graph encoding algorithm to generate a vector representation of the knowledge subgraph. For each text paragraph, the system extracts relevant entity nodes and relation edges from the oncology semantic knowledge graph to construct a local knowledge subgraph. The graph encoding algorithm employs a graph convolutional network, aggregating neighborhood information through a multi-layer message passing mechanism to learn an enhanced representation for each node. The specific formula for the graph convolution operation is as follows: in, Indicates the first Nodes in a layered network The representation vector, Represents a node The set of neighboring nodes, It is a normalization constant. It is a learnable weight matrix. It is a non-linear activation function. Ultimately, the system aggregates the node representations of the entire knowledge subgraph into a unified graph embedding vector through pooling operations.
[0058] Finally, the system constructs a training sample set consisting of text-knowledge pairs. Each training sample comprises three components: the original input text sequence, the corresponding knowledge graph embedding vector, and the target output text. The target output text consists of standard answers edited by medical experts or verified popular science content, ensuring the accuracy and standardization of the content. The system performs quality control and data balancing on the training samples, removing low-quality samples and ensuring a balanced distribution of samples across different tumor types and treatment plans. The completed training dataset is serialized and stored, and an efficient retrieval mechanism is established, providing a high-quality, multimodal training foundation for subsequent training of a large-scale language model specifically for oncology.
[0059] 2) Design a knowledge-aware Transformer model architecture. The model integrates a text encoder and a knowledge injection module. The text encoder uses multi-layer Transformer blocks to process the input sequence and generate context-aware text representations. The knowledge injection module uses a graph attention network to encode the structural information of the oncology semantic knowledge graph and dynamically fuses the text representation and knowledge representation through a cross-modal attention mechanism. The model output layer combines a language modeling head and a knowledge prediction head to support text generation and knowledge reasoning tasks. Parameter initialization is based on a pre-trained general medical language model to enhance domain adaptability.
[0060] The knowledge-aware Transformer model architecture employs an encoder-decoder structure, where the encoder comprises two parallel branches: a text encoder and a knowledge injection module. The text encoder is composed of multiple identical stacked Transformer blocks, each containing a multi-head self-attention layer and a feedforward neural network layer. The input text sequence is first converted into a vector representation through a word embedding layer, and then positional encoding information is added, calculated using the following formula: in, Indicates the position of a word in a sequence. Indicates a dimension index. The model dimension is represented. The text encoder generates a text representation matrix rich in contextual information through layers of transformation.
[0061] The knowledge injection module employs a graph attention network to encode the structural information of the oncology semantic knowledge graph. This module receives entity node features and relation edge information extracted from the knowledge graph and calculates the importance weights between nodes using a graph attention mechanism. The computation process of the graph attention layer is as follows: in, Represents a node initial characteristics, Represents a node The neighborhood group, It is a shared weight matrix. It is an attention vector. It is an activation function. This represents a vector concatenation operation. After processing by a multi-layer graph attention network, the structural information of the knowledge graph is encoded into a dense vector representation.
[0062] Next, the system dynamically fuses text representations and knowledge representations through a cross-modal attention mechanism. The cross-modal attention module uses the text representation as the query vector and the knowledge representation as the key-value pairs to calculate the association weights between the text and the knowledge. The specific calculation process is as follows: in, Text represents a matrix. and Knowledge representation matrix, This represents the dimension of the key vector. The output of cross-modal attention is a knowledge-enhanced text representation that retains the semantic information of the original text while incorporating the structured information of relevant knowledge.
[0063] The model's output layer employs a dual-head design, comprising a language modeling head and a knowledge prediction head. The language modeling head handles text generation, calculating the probability distribution of the next word through linear transformations and the softmax function. The knowledge prediction head handles knowledge inference.
[0064] Predict knowledge graph entities and relationships in the input text. Both heads share the underlying representation but use different parameter matrices for task-specific transformations.
[0065] During the parameter initialization phase, the system loads the weights of a pre-trained general medical language model as basic parameters. These pre-trained weights come from a language model trained on a large-scale medical corpus and already possess good medical language understanding capabilities. The model then undergoes further fine-tuning training, optimizing parameters through knowledge-enhanced supervised learning objectives. This allows the model to better adapt to the specific needs of the oncology field and improve its performance in tasks such as generating popular science content and knowledge reasoning related to oncology.
[0066] 3) Perform multi-objective joint optimization training. The training process simultaneously minimizes language modeling loss, knowledge consistency loss, and generation quality loss. Language modeling loss is based on an autoregressive next-word prediction task; knowledge consistency loss measures the semantic alignment of generated content with the oncology semantic knowledge graph; generation quality loss optimizes the fluency and accuracy of the content through adversarial training. The loss function is defined as a weighted sum: in, This represents the total loss value. Represents the language modeling loss, This indicates a loss of knowledge consistency. Indicates the generation of mass loss, , and This represents the task weight coefficients. Training uses an adaptive optimization algorithm to iteratively update the model parameters.
[0067] The language modeling loss module is based on an autoregressive next-word prediction task, calculating the difference between the model's output sequence and the target sequence. For a sequence of length... For the sequence, the language modeling loss is calculated using the cross-entropy loss function: in, This represents the language modeling loss value. Indicates the sequence length. Indicates the first The target word at each position, Indicates the preceding A sequence of words This represents the input knowledge graph information. This represents the conditional probability predicted by the model. This loss function ensures that the model fully utilizes the preceding context and relevant knowledge when generating each word.
[0068] The knowledge consistency loss module measures the semantic alignment between the generated content and the oncology semantic knowledge graph. First, entity references are extracted from the generated text and mapped to corresponding entities in the knowledge graph using entity linking techniques. Then, the difference between the semantic representation of the generated text and the representation of the relevant knowledge subgraph is calculated. in, This represents the knowledge consistency loss value. Indicates the number of entities extracted. Indicates the first in the generated text The semantic vector of an entity, This represents the vector representation of the corresponding entity in the knowledge graph. This represents the cosine similarity function. This loss term ensures that the content generated by the model is consistent with authoritative knowledge.
[0069] The generation quality loss module optimizes the fluency and accuracy of the content through adversarial training. The system trains a discriminator network that receives a text sequence and outputs the probability that it is a real sample. The generation quality loss is calculated based on the discriminator's rating of the generated content. in, This indicates the generated quality loss value. Indicates batch size, Indicates the first One generated text sample, Let represent the discriminator network. The discriminator is trained on both real and generated data, and its objective function is: The system combines the three loss terms into a weighted total loss: in, This represents the total loss value. , and This represents the task weight coefficient, whose optimal value is determined through grid search. The training process uses an adaptive optimization algorithm to iteratively update the model parameters. The optimizer employs the AdamW algorithm, with a decreasing learning rate and an initial value of [value missing]. The loss is reduced proportionally after a certain number of training steps. The model's performance is periodically evaluated on the validation set during training, and the training strategy is adjusted based on the validation loss to ensure that all three loss terms are adequately optimized. The entire training process continues until the total loss value converges, ultimately resulting in a large-scale language model specifically designed for oncology, possessing excellent language generation capabilities, knowledge accuracy, and content quality.
[0070] The intelligent science popularization generation module is used to generate oncology science popularization content based on the user's input questions, topics of interest, and the user's cognitive level, and then provide it to the user.
[0071] A user's cognitive level is calculated using a multi-dimensional quantitative model based on user-provided profiles and historical interaction data automatically collected by the system. Profiles include the user's self-reported occupational category, highest educational background, and medical-related work experience; historical interaction data includes the complexity of questions asked in past conversations, the depth of popular science content read, and feedback behaviors on generated content, such as likes or corrections.
[0072] The user's cognitive level score is calculated using the following formula: in, This represents a user's cognitive level score, ranging from zero to one. Indicates educational background score; Indicates career-related scores; Indicates historical interaction scores; , and This represents the weighting coefficient.
[0073] Historical Interaction Score The calculation formula is: in, This represents the current historical interaction score; This represents the attenuation factor, with a value of 0.1. Indicates the score of the current interactive session; This represents the historical interaction score calculated previously.
[0074] Score of the current interactive session Calculated using the following formula: in, Represents the problem complexity score; This represents the score indicating the depth of content reading.
[0075] Problem Complexity Score The calculation formula is: in, The number of technical terms in the question is indicated by matching the user's question text with an oncology terminology dictionary, which includes standard medical terms such as "adenocarcinoma," "immunohistochemistry," and "targeted therapy." This represents the preset maximum threshold for the number of terms, with a value of ten. Indicates the maximum depth of the parse tree; This represents the depth scaling factor, with a value of fifteen.
[0076] A parse tree is a tree-like representation of a user's query text obtained through dependency parsing. In this tree structure, each node represents a word, and edges represent grammatical dependencies between words. The maximum depth of a parse tree refers to the number of edges traversed from the root node to the farthest leaf node, reflecting the grammatical complexity of the sentence.
[0077] Content reading depth score Values are assigned based on user reading behavior: 0.2 for reading basic content, 0.5 for reading advanced content, and 0.8 for reading professional content.
[0078] The system divides users into three levels based on their cognitive level scores: low cognitive level corresponds to a score range of 0 to 0.4, medium cognitive level corresponds to a score range of 0.4 to 0.7, and high cognitive level corresponds to a score range of 0.7 to 1.
[0079] The system integrates the user's input question text, description of the topic of interest, and quantified cognitive level score into a structured prompt, which is then input into a large-scale language model specifically for oncology. This oncology-specific language model is based on a Transformer decoder architecture, integrating embedded representations from an oncology semantic knowledge graph. It dynamically fuses knowledge and generates text through a multi-head self-attention mechanism. The generation process first parses the key entities and semantic intent in the user's question, then retrieves relevant nodes and relationships from the oncology semantic knowledge graph, and finally adjusts the linguistic complexity and information depth of the generated content according to the user's cognitive level. For users with low cognitive levels, the model uses basic vocabulary and short sentence structures, avoiding specialized medical terminology; for users with medium cognitive levels, the model introduces basic medical terminology with brief explanations; for users with high cognitive levels, the model provides detailed descriptions of pathological mechanisms and references to clinical guidelines. After generating the content, the system pushes it to the user in real-time in a multimodal format via a web interface or mobile application, supporting user interaction and feedback. The system records data from each interaction to update the user's cognitive level score and optimize the model's generation strategy.
[0080] Optionally, the above technical solution also includes a knowledge dynamic update module, which is used for: By monitoring the latest oncology papers and clinical treatment guidelines, the oncology semantic knowledge graph and oncology-specific large language model are automatically updated. The specific implementation process is as follows: The latest oncology papers and clinical guidelines are monitored and collected from multiple predefined authoritative data sources. These authoritative data sources include internationally renowned academic databases such as PubMed and IEEE Xplore, oncology journal websites such as The Lancet Oncology and Journal of Clinical Oncology, official guideline publication platforms such as the American National Comprehensive Cancer Network Guidelines website and the Chinese Society of Clinical Oncology Guidelines website, as well as clinical trial registries. The monitoring process uses distributed web crawlers configured with specific rules to identify and download newly published full-text papers and guideline documents. The crawlers run regularly, for example, once a day, to ensure timely capture of updated content.
[0081] After acquiring new data, the system performs data preprocessing and parsing. For oncology papers, the system extracts the title, abstract, keywords, main text, and reference list; for oncology clinical treatment guidelines, the system extracts the version number, publishing institution, update date, recommendation level, and specific treatment suggestions. The parsed text data is input into the named entity recognition and relation extraction module. This module, based on a pre-trained deep learning model, identifies entities related to tumor type, pathological mechanism, treatment plan, side effects, and prognosis in the text, and extracts the semantic relationships between entities. For example, from a new paper, the system identifies the entities "immune checkpoint inhibitors" and "non-small cell lung cancer," and extracts the relation "treatment" connecting these two entities.
[0082] The system uses extracted entities and relationships to update the oncology semantic knowledge graph. The update process includes two sub-steps: entity linking and knowledge fusion. Entity linking matches newly identified entities with existing entities in the knowledge graph, calculating a matching score using string similarity and context embedding vectors, as shown in the formula: in, Represents a new entity With the old entity The matching score; This represents a string similarity function, calculated based on edit distance. and These represent the context embedding vectors of the new entity and the old entity, respectively; This represents the weighting coefficient, with a value of 0.6. This represents the cosine similarity function. If the matching score is below a threshold, the new entity is added to the knowledge graph; otherwise, the system performs knowledge fusion, merging duplicate entities and resolving semantic conflicts, such as prioritizing relationships from the latest guidelines based on evidence level. Knowledge graph updates are implemented through graph database transaction operations, ensuring data consistency and integrity.
[0083] After the oncology semantic knowledge graph is updated, the system triggers the update process of the oncology-specific large language model. The update employs an incremental fine-tuning strategy, using newly collected paper and guideline data as training corpus. The fine-tuning process first constructs a training dataset, including new texts and their corresponding knowledge graph embeddings, and then calculates the loss function for model fine-tuning, as shown in the formula: in, Indicates fine-tuning loss; Indicates the number of new training samples; Indicates the first The target output for each sample; Indicates the first The input text for each sample; Indicates the first The knowledge graph embedding vector corresponding to each sample; This represents the conditional probability, calculated using a large language model specific to oncology. Indicates model parameters; Represents the regularization term; This represents the regularization coefficient, with a value of 0.01. Fine-tuning uses the stochastic gradient descent algorithm to optimize model parameters, with the learning rate set to a low value, such as 0.0001, to prevent overfitting. After fine-tuning, the system validates the updated model, using a retained dataset to evaluate the accuracy and timeliness of the generated content. If the evaluation metrics decline, the system rolls back to the previous version.
[0084] The entire automated update process is coordinated by a workflow engine, recording logs and version information for each update, and supporting manual intervention to handle anomalies. The system periodically generates update reports summarizing the number of new entities, modified relationships, and model performance changes for administrator review. This approach ensures that the oncology semantic knowledge graph and the oncology-specific large language model continuously reflect the latest research progress and clinical practice, maintaining the authority and reliability of the system's output.
[0085] Optionally, the above technical solution also includes a trusted audit and knowledge traceability module, which is used for: Each piece of oncology science popularization content generated by the intelligent science popularization generation module is annotated with knowledge source labels, and receives review data from artificial medical experts and model error correction feedback data. The specific implementation process is as follows: When generating popular science content using a large language model specific to oncology, the system simultaneously records all knowledge sources used in the generation process. These knowledge sources include specific nodes and relationships in the oncology semantic knowledge graph, as well as the original data sources on which the generated content is based, such as specific chapters of clinical practice guidelines, titles and unique identifiers of digital objects in research papers, and version numbers of technical reports issued by official medical institutions.
[0086] The specific implementation of knowledge source labeling uses structured data format for storage. When each piece of oncology science popularization content is generated, the system creates a corresponding source labeling record, which includes a content identifier, a generation timestamp, and a list of source entries. Each source entry includes source type, source identifier, confidence score, and usage fragment information. The confidence score for source labeling information is calculated using the following formula: in, Indicates the confidence score; Indicates the level of authority of the source; Indicates the timeliness score of the source; This represents the relevance score between the content and its source. , and Indicates the weighting coefficient. Source authority level. The score is assigned based on the authority of the publishing institution: guidelines from internationally authoritative medical institutions are scored 1, articles from core journals are scored 0.8, and other sources are scored 0.6. Source timeliness score. Based on the release time, the formula is: ,in This indicates the time difference between the current time and the publication time. This represents the attenuation coefficient. The relevance score between the content and its source. It is obtained through semantic similarity calculation.
[0087] The system presents knowledge source annotations to users in two formats: one is a concise annotation format, which displays the guideline name and version number or paper title of the key source at the end of the popular science content; the other is a detailed annotation format, where users can click the expand button to view complete source information, including specific guideline recommendations, level of evidence, and reference links.
[0088] For receiving review data from medical experts and model error correction feedback, the system provides a specially designed medical expert review interface. This interface displays the oncology science popularization content to be reviewed, along with its corresponding knowledge source annotations, and offers various review tools. Medical experts can rate the content using a five-point scale: one point indicates serious errors, two points indicates numerous inaccuracies, three points indicates mostly accurate content with minor imprecisions, four points indicates accurate and clearly expressed content, and five points indicates completely accurate content with excellent educational value. Experts can also add annotations to specific content segments, pointing out specific problems and providing suggestions for improvement.
[0089] Model error correction feedback data is collected and processed through the following process: After medical experts submit their review results, the system stores the feedback data in the error correction feedback database. Each feedback record includes the original content identifier, reviewer identifier, review timestamp, overall score, segment-level annotations, and modification suggestions. The system periodically performs statistical analysis on this feedback data to calculate evaluation metrics for the model output quality. in, This indicates the model's output quality score; This indicates the total number of feedback received; Indicates the first The feedback rating; Indicates the number of items requiring significant modifications; Indicates the total number of items reviewed; This represents the penalty coefficient.
[0090] The system uses collected error correction feedback data to build a training dataset for model optimization. This dataset includes the original input, model output, expert-corrected versions, and explanations of the reasons for the corrections. The system regularly uses this dataset to perform targeted fine-tuning of the oncology-specific large language model, paying particular attention to areas where errors frequently occur, such as drug dosage descriptions, treatment indications, and side effect descriptions.
[0091] To achieve continuous improvement, the system also establishes a feedback loop. When the model is updated based on expert feedback, the system automatically selects similar questions previously marked as problematic, regenerates content, and sends the newly generated content back to medical experts for review to verify whether the issues have been resolved. This approach ensures the transparency and credibility of the knowledge sources for oncology popular science content, while continuously improving the accuracy and reliability of the oncology-specific large language model through ongoing review and feedback from medical experts.
[0092] The technical solution of the present invention will be further described through another embodiment.
[0093] The system in this embodiment achieves the following goals through multimodal data fusion, medical semantic knowledge graph construction, and intelligent text generation: establishing an authoritative and systematic oncology knowledge system; providing personalized and interactive oncology science Q&A and content generation services; and promoting an intelligent bridge between scientific research results and public health awareness. Specifically, the system will integrate multi-source oncology knowledge, achieve intelligent conversion from academic language to public language, and ensure the accuracy, timeliness, and comprehensibility of the content. It includes the following modules: 1) The knowledge acquisition and semantic normalization module automatically collects structured and unstructured oncology-related knowledge from multiple authoritative data sources. These authoritative data sources include oncology clinical guidelines, research paper databases, clinical databases, and reports from the National Cancer Center. Structured knowledge refers to data stored in a predefined format, such as database tables and standardized medical codes; unstructured knowledge refers to text content without a fixed format, such as full-text medical literature and clinical notes. Named entity recognition (NER) and relation extraction techniques are used to identify entities related to tumor types, pathological mechanisms, treatment plans, side effects, and prognosis. NER is used to extract specific entities from the text, while relation extraction is used to establish semantic relationships between entities. Based on the extraction results, an oncology semantic knowledge graph is constructed. This graph represents oncology knowledge nodes and relationships in a graph structure, achieving unified knowledge modeling and semantic standardization.
[0094] 2) The large-scale intelligent model training module trains a large-scale language model specifically for oncology based on medical corpora and an oncology semantic knowledge graph. The medical corpus refers to a collection of text data gathered from the medical field, including oncology guidelines, research papers, and patient education materials. The training process employs a hybrid training strategy, including supervised fine-tuning and human feedback reinforcement learning. Supervised fine-tuning uses paired data constructed from the medical corpus and knowledge graph to optimize model parameters; human feedback reinforcement learning further adjusts the model through ratings and feedback data from medical experts, ensuring that the model output is authoritative, credible, and interpretable.
[0095] 3) Intelligent Science Popularization Generation Module: Based on the user's input questions, topics of interest, and cognitive level, this module dynamically generates hierarchical and easily understandable oncology science popularization content using a large-scale language model specifically designed for oncology. The user's cognitive level is calculated using a quantitative model, assessing their knowledge comprehension ability based on their personal profile and historical interaction data. The generated content supports multilingual output and multimodal display, including text, charts, and animations, to cater to diverse user needs.
[0096] 4) The knowledge dynamic update module achieves automatic incremental updates of knowledge through paper monitoring and guideline update crawlers. The paper monitoring crawler regularly scans academic databases and journal websites, while the guideline update crawler tracks official guideline publication platforms. New knowledge is used to update the oncology semantic knowledge graph and periodically fine-tunes the oncology-specific large language model to maintain the cutting-edge nature and accuracy of the content.
[0097] 5) The Credibility Verification and Knowledge Source Traceability module adds knowledge source annotations to each piece of oncology science popularization content generated by the intelligent science popularization generation module, such as guideline versions and unique identifiers for digital objects of literature. These knowledge source annotations are stored in a structured format, including source type and confidence score. Simultaneously, the module supports human medical expert review and model error correction feedback. Experts can score and annotate content through a dedicated interface, and the feedback data is used to optimize the oncology-specific large language model.
[0098] Example 1: Generation of popular science articles on tumor prevention and treatment: Users input the question "How to detect lung cancer early?" through the system interface. After receiving the user's input, the intelligent science popularization generation module first parses the user's input, identifying the key entity "lung cancer" and the query intent "early detection methods". Next, the system accesses the oncology semantic knowledge graph and automatically retrieves relevant knowledge nodes, including entities such as "early screening methods for lung cancer", "low-dose spiral CT", and "high-risk population criteria" and their relationships.
[0099] During the content generation phase, the system dynamically generates tiered content based on users' cognitive level scores. For general users with cognitive level scores between 0 and 0.4, the system generates basic content, explaining the importance and methods of early lung cancer screening in easy-to-understand language, avoiding professional medical terminology, and keeping the content length under 200 words. For primary care physicians or medical students with cognitive level scores between 0.4 and 0.7, the system generates advanced content, providing detailed guideline-level screening strategies and evidence levels, introducing necessary professional terminology with concise explanations, and extending the content length to 500 words. Simultaneously, the system also generates educational graphic content, providing illustrated explanations through multimodal display capabilities, including diagrams of early screening processes and charts depicting high-risk population characteristics, facilitating dissemination on social media.
[0100] Throughout the generation process, the credibility audit and knowledge traceability module adds knowledge source annotations to each version of the content. The annotation information includes the name of the cited guideline, the version number, and the unique digital object identifier of the relevant research literature, ensuring the credibility and traceability of the content.
[0101] Example 2: Dynamic Knowledge Update: When the American National Comprehensive Cancer Network or the Chinese Society of Clinical Oncology releases new oncology treatment guidelines, the knowledge dynamic update module immediately initiates the update process. The system monitors the official websites of these authoritative institutions using a pre-configured guideline update crawler, automatically detecting and downloading the latest guideline documents. After downloading, the system parses the guideline documents, extracting updated treatment recommendations, newly added drug information, and modified clinical pathways. Named entity recognition and relation extraction technologies then process this content, identifying new tumor types, pathological mechanisms, treatment plans, side effects, and prognostic entities, and establishing semantic relationships between entities. This new knowledge is used to update the oncology semantic knowledge graph, integrating new entities with the existing graph through entity linking and knowledge fusion technologies.
[0102] After the knowledge graph is updated, the system triggers a fine-tuning process for the oncology-specific large language model. Fine-tuning uses newly collected guideline content as training corpus and employs an incremental learning strategy to incorporate new diagnostic and treatment information while maintaining existing knowledge. During fine-tuning, the system monitors model performance through a loss function to ensure the accuracy and stability of knowledge updates. The entire update process is typically completed within 24 hours, effectively maintaining the real-time and cutting-edge nature of the system's knowledge.
[0103] This invention is the first to perform semantically unified modeling of multi-source oncology knowledge, achieving intelligent mapping from academic language to public language through an oncology semantic knowledge graph. This process involves the application of named entity recognition and relation extraction technologies to ensure the consistency and interpretability of the knowledge structure. Based on the user's health literacy, professional level, and information needs, the system automatically adjusts the depth and linguistic complexity of the generated content by quantifying the user's cognitive level. The oncology-specific large language model dynamically adjusts its output according to the cognitive level score, from basic explanations to professional descriptions, meeting the needs of users at different levels. Combining text, images, and knowledge graph embedding, the oncology-specific large language model achieves highly interpretable knowledge generation. Images include pathological images and CT scan illustrations, and knowledge graph embedding provides semantic context, enhancing the accuracy and visualization of the generated content. The system has the ability to automatically track the latest clinical research and guideline changes, maintaining the authority and timeliness of the content through a dynamic knowledge update module. Knowledge source annotation and human medical expert review ensure the credibility of the generated content, while model error correction feedback supports continuous system optimization. The beneficial effects are as follows: The system provides the public with credible, authoritative, and personalized cancer prevention and treatment knowledge services. Through an intelligent science popularization generation module, the system offers tailored content based on users' cognitive levels, making complex medical knowledge easy to understand. A dynamic knowledge update module ensures that the popular science content always reflects the latest medical advancements, while a credibility verification and knowledge traceability module guarantees the accuracy and reliability of the information. This service model can significantly improve the public's knowledge of cancer prevention and treatment, promoting the formation of healthy behaviors. The system can assist doctors in patient education, enhancing the healthcare experience. Doctors can use the system to quickly generate popular science materials that match patients' cognitive levels, effectively explaining complex treatment plans and prognoses. The system's multimodal display functions, including text, charts, and animations, help patients understand medical information more intuitively. This support helps improve doctor-patient communication efficiency and enhances patient treatment adherence and satisfaction. The system achieves intelligent integration of scientific research knowledge with public understanding. By automatically collecting and analyzing the latest scientific papers, the system can promptly transform cutting-edge research results into easily understandable popular science content. This transformation not only accelerates the dissemination of scientific knowledge but also promotes public understanding and acceptance of medical advancements. Meanwhile, the system's knowledge traceability function ensures the accurate dissemination of research findings and maintains the rigor of science communication. The system is capable of interfacing with electronic medical records and health management platforms, enabling the construction of an intelligent health knowledge ecosystem. Through integration with electronic medical record systems, the system can generate personalized health guidance based on the patient's specific condition; integration with health management platforms allows the system to provide continuous health monitoring and educational services. This ecological application model expands the system's service scope and provides technical support for building a comprehensive intelligent health management system.
[0104] like Figure 2As shown in the figure, an embodiment of the present invention provides a method for popularizing oncology based on a large language model, which includes the following steps: S1. Automatically collect structured and unstructured knowledge related to oncology from multiple authoritative data sources to construct an oncology semantic knowledge graph; S2. Training a large language model specifically for oncology based on medical corpus and oncology semantic knowledge graph; S3. Based on the user's input questions, topics of interest, and the user's level of understanding, generate popular science content on oncology using a large language model specifically for oncology, and provide it to the user.
[0105] Optionally, in the above technical solution, structured and unstructured knowledge related to oncology is automatically collected from multiple authoritative data sources to construct an oncology semantic knowledge graph, including: We automatically collect structured and unstructured knowledge related to oncology from multiple authoritative data sources, and use named entity recognition and relation extraction techniques to identify tumor types, pathological mechanisms, treatment plans, side effects and prognoses to construct an oncology semantic knowledge graph.
[0106] Optionally, the above technical solution also includes: By monitoring the latest oncology papers and clinical treatment guidelines for oncology, the oncology semantic knowledge graph and oncology-specific large language model are automatically updated.
[0107] Optionally, the above technical solution also includes: Each piece of oncology science popularization content generated is labeled with the knowledge source and receives review data from artificial medicine experts and model error correction feedback data.
[0108] In another embodiment, structured and unstructured knowledge is automatically collected from multiple authoritative data sources. Named entity recognition and relation extraction techniques are used to process the knowledge, constructing an oncology semantic knowledge graph. An oncology-specific large-scale language model is trained based on medical corpora and the knowledge graph, and a hybrid training strategy is employed to optimize the model. Personalized science popularization content is generated using the oncology-specific large-scale language model based on user input and cognitive level. The generated content is pushed to users in a multimodal format, and user interaction data is collected to improve the service. A dynamic knowledge update module maintains the freshness of the knowledge, and a trusted auditing module enables continuous improvement.
[0109] It should be noted that the beneficial effects of the oncology popularization method based on a large language model provided in the above embodiments are the same as those of the oncology popularization system based on a large language model, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the system embodiments, and will not be repeated here.
[0110] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned methods for popularizing oncology based on a large language model.
[0111] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-mentioned methods for popularizing oncology based on a large language model.
[0112] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.
[0113] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.
[0114] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A tumor science popularization system based on a large language model, characterized in that, include: The module includes knowledge acquisition and semantic normalization, a large-scale knowledge model training module, and an intelligent science popularization generation module. The knowledge acquisition and semantic normalization module is used to: automatically collect structured and unstructured knowledge related to oncology from multiple authoritative data sources, and construct an oncology semantic knowledge graph; The intelligent large-scale model training module is used to: train an oncology-specific large-scale language model based on medical corpus and the oncology semantic knowledge graph; The intelligent science popularization generation module is used to generate oncology science popularization content based on the user's input questions, topics of interest, and the user's cognitive level, and then provide it to the user.
2. The oncology science popularization system based on a large language model according to claim 1, characterized in that, The knowledge acquisition and semantic normalization module is specifically used to: automatically collect structured and unstructured knowledge related to oncology from multiple authoritative data sources, and identify tumor types, pathological mechanisms, treatment plans, side effects and prognostic entities through named entity recognition and relation extraction technology, and construct an oncology semantic knowledge graph.
3. A tumor science popularization system based on a large language model according to claim 1 or 2, characterized in that, It also includes a knowledge dynamic update module, which is used for: By monitoring the latest oncology papers and clinical treatment guidelines for oncology, the oncology semantic knowledge graph and the oncology-specific large language model are automatically updated.
4. A tumor science popularization system based on a large language model according to claim 1 or 2, characterized in that, It also includes a trusted audit and knowledge traceability module, which is used for: Each piece of tumor science popularization content generated by the intelligent science popularization generation module is labeled with the source of knowledge, and the module receives review data from artificial medical experts and model error correction feedback data.
5. A method for popularizing oncology knowledge based on a large language model, characterized in that, include: Structured and unstructured knowledge related to oncology are automatically collected from multiple authoritative data sources to construct an oncology semantic knowledge graph; A large language model specifically for oncology was trained based on medical corpus and the oncology semantic knowledge graph. Based on the user's input questions, topics of interest, and level of understanding, the system uses the oncology-specific large language model to generate popular science content on oncology and provides it to the user.
6. A method for popularizing oncology knowledge based on a large language model according to claim 5, characterized in that, Structured and unstructured knowledge related to oncology was automatically collected from multiple authoritative data sources to construct an oncology semantic knowledge graph, including: We automatically collect structured and unstructured knowledge related to oncology from multiple authoritative data sources, and use named entity recognition and relation extraction techniques to identify tumor types, pathological mechanisms, treatment plans, side effects and prognoses to construct an oncology semantic knowledge graph.
7. A method for popularizing oncology knowledge based on a large language model according to claim 5 or 6, characterized in that, Also includes: By monitoring the latest oncology papers and clinical treatment guidelines for oncology, the oncology semantic knowledge graph and the oncology-specific large language model are automatically updated.
8. A method for popularizing oncology knowledge based on a large language model according to claim 5 or 6, characterized in that, Also includes: Each piece of oncology science popularization content generated is labeled with the knowledge source and receives review data from artificial medicine experts and model error correction feedback data.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the tumor science popularization method based on any one of claims 5 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the tumor science popularization method based on a large language model as described in any one of claims 5 to 8.