A breast cancer popularization method and system based on a multi-modal knowledge graph

By constructing a multimodal knowledge graph and combining deep learning and multi-task learning, the problem of insufficient personalization and interactivity in breast cancer science popularization has been solved. This has enabled the generation of personalized, authoritative, and highly interactive science popularization content, improving user experience and information accuracy.

CN120299676BActive Publication Date: 2025-12-05BEIJING FUYU MEDICAL TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510369301.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-12-05
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

Existing methods for breast cancer education lack the ability to integrate personalized and multimodal information, and cannot provide authoritative and highly interactive educational content. Traditional systems are unable to meet the personalized needs of different groups and lack in-depth analysis of imaging examination results and pathological diagnosis reports.

Method used

A multimodal knowledge graph is constructed, breast cancer image samples and text data are collected and processed, specific features are extracted, mapping relationships are established, and personalized popular science content is generated by combining deep learning models and multi-task learning architecture. Security review and confidence judgment mechanisms are introduced.

Benefits of technology

It has achieved a significant improvement in personalized science popularization capabilities, greatly enhanced authority and credibility, significantly improved interactivity and user experience, and stronger applicability and scalability. It can dynamically update content based on user input and ensure the accuracy of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299676B_ABST
    Figure CN120299676B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical information analysis, in particular to a breast cancer disease popularization method and system based on a multi-modal knowledge graph. The method comprises the following steps: collecting and processing breast cancer image samples and text data; extracting breast cancer specific features from the breast cancer image samples; creating corresponding nodes for the breast cancer specific features in the knowledge graph, and establishing a mapping relationship with medical diagnosis information to form a multi-modal unified representation; associating medical text entities in the text data with the breast cancer specific feature nodes to form a multi-modal knowledge graph; obtaining user inputted to-be-diagnosed information, retrieving associated nodes related to the to-be-diagnosed information from the multi-modal knowledge graph, and outputting popularization content. The application can realize personalized, authoritative and interactive breast cancer popularization, and provide high-quality popularization services for patients and the public.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical information analysis, and in particular to a breast cancer disease popularization method and system based on a multi-modal knowledge graph. BACKGROUND

[0002] Breast cancer is one of the malignant tumors with high incidence and mortality among women worldwide. Despite the continuous improvement of medical level, there are still deficiencies in the cognition of breast cancer disease and the acquisition of popular science knowledge among the general public. Traditional breast cancer popularization methods mainly rely on static data or offline education, which has slow information updating speed and is difficult to meet the individual needs of different groups. In addition, existing online question and answer and information retrieval tools have great limitations in content authority, information source traceability, multi-modal information integration, and interactivity with users.

[0003] Specifically, traditional popularization mainly relies on text or simple illustrations, which lacks in-depth analysis of specific patient conditions (such as imaging examination results, pathological diagnosis reports, etc.), making it difficult for users to obtain personalized and interpretable knowledge support. Pure text retrieval engines or general question and answer systems often cannot guarantee the medical rigor and credibility of the results, and once the information queried by the user is misleading or missing, it may cause adverse consequences. At the same time, the lack of structured knowledge management tools such as knowledge graphs makes existing information scattered in different data sources, making it difficult to achieve organic integration and efficient retrieval of multi-modal information. Furthermore, traditional systems lack interactivity and context understanding ability, and cannot continuously adjust and refine popularization content according to user follow-up questions and feedback, thereby affecting user experience and knowledge absorption effect.

[0004] Therefore, there is an urgent need for a breast cancer popularization method and system that can integrate multi-modal information, provide personalized popularization, and have high authority and interactivity to solve the deficiencies in the prior art. SUMMARY

[0005] In view of one or more of the problems existing in the prior art, the first aspect of the present application provides a breast cancer disease popularization method based on a multi-modal knowledge graph, comprising:

[0006] Collecting and processing breast cancer image samples and text data;

[0007] Extracting breast cancer-specific features from the breast cancer image samples;

[0008] In the knowledge graph, creating corresponding nodes for the breast cancer-specific features and establishing a mapping relationship with the medical diagnosis information to form a multi-modal unified representation;

[0009] Associating the text data with the breast cancer-specific feature nodes to form a multi-modal knowledge graph;

[0010] Obtaining user inputted information to be diagnosed, retrieving associated nodes related to the information to be diagnosed from a multi-modal knowledge graph, and outputting popular science content.

[0011] Preferably, collecting and processing breast cancer image samples and text data specifically includes:

[0012] Performing a processing operation on the breast cancer image sample, the processing operation including at least one of denoising, normalization, and contrast enhancement;

[0013] Performing preliminary lesion positioning and information labeling on the processed breast cancer image sample, and recording in the breast cancer image sample metadata;

[0014] Aligning the text data with professional terms, and uniformly indexing the aligned text data and breast cancer image sample metadata.

[0015] Preferably, extracting breast cancer specific features from the breast cancer image sample specifically includes:

[0016] Using a deep learning model to analyze the breast cancer image sample and extract mass morphology, density, and microcalcification features;

[0017] Combining image bottom-level features with high-level attention mechanisms to form a comprehensive representation of local texture and global context.

[0018] Preferably, an example graph is introduced for breast cancer specific feature nodes, and when the user inputted information to be diagnosed is similar to the example graph, it will be displayed on the user interface for easy comparison.

[0019] Preferably, the user inputted information to be diagnosed is obtained, and while retrieving associated nodes related to the information to be diagnosed from a multi-modal knowledge graph, supplementary information from an external database is combined to achieve multi-modal RAG enhancement.

[0020] Preferably, the user inputted information to be diagnosed is obtained through human-computer interaction dialogue, and the user inputted information to be diagnosed is dynamically updated according to the context in multi-turn dialogue to achieve continuous answers to personalized questions and popular science supplements.

[0021] Preferably, the breast cancer disease popular science method based on a multi-modal knowledge graph further includes designing fine-tuning tasks to improve the credibility and medical adaptability of breast cancer popular science content, the fine-tuning tasks including at least:

[0022] Decomposing breast cancer popular science core needs into several sub-tasks and performing joint training, the sub-tasks including:

[0023] Automatically detecting and labeling suspected morphology, density, and microcalcification features;

[0024] According to the breast cancer image, pathological information provides personalized treatment plan popular science;

[0025] Provide popular science psychological comfort and nursing suggestions for users.

[0026] Preferably, during training, a joint loss function can be defined to balance medical professionalism and popular readability;

[0027] The joint loss function The following formula is used for calculation:

[0028]

[0029] Wherein, The medical professional task loss is used to measure the accuracy or cross-entropy loss of image recognition and pathological stage prediction professional tasks; The popular science text loss is used to measure the readability and relevance of the popular science text; α is the weight coefficient of the medical professional task loss; β is the weight coefficient of the popular science text loss; y k Indicates the true value label of the professional task; Indicates the probability predicted by the model; K indicates the total number of categories in the medical professional task; k indicates the category index, from 1 to K; x i Indicates the word or sentence to be predicted in the popular science generation; x context Indicates the context information when generating the text; M indicates the total number of words or sentences in the popular science text; P indicates the probability of generating a certain word or sentence.

[0030] Preferably, the breast cancer disease popular science method based on the multi-modal knowledge graph further includes safety review and confidence judgment, specifically including:

[0031] The confidence of the generated popular science content is evaluated, and it is decided whether further review or adjustment is needed according to the evaluation result;

[0032] For content with a confidence lower than a preset threshold, the "retrieval" or "prompt referral to a professional doctor" mechanism is automatically triggered to improve the accuracy of the content;

[0033] For some sensitive topics or high-risk information, configure medical expert review or automatic filtering strategy to ensure the accuracy and safety of the popular science information.

[0034] The second aspect of the present application provides a breast cancer disease popular science system based on a multi-modal knowledge graph, comprising:

[0035] A data acquisition and preprocessing module is used to acquire and process breast cancer image samples and text data;

[0036] The feature extraction and mapping module is connected with the data acquisition and preprocessing module, and is configured to extract breast cancer specific features from the processed breast cancer image samples, and create corresponding nodes in the knowledge graph and establish a mapping relationship with medical diagnosis information.

[0037] The multi-modal knowledge graph construction module is connected with the feature extraction and mapping module, and is configured to associate medical text entities in the text data with the breast cancer specific feature nodes to form a multi-modal knowledge graph.

[0038] The retrieval and generation module is connected with the multi-modal knowledge graph construction module, and is configured to obtain user inputted diagnosis information, retrieve associated nodes related to the diagnosis information from the multi-modal knowledge graph, and output popular science content.

[0039] Preferably, the breast cancer disease popular science system based on the multi-modal knowledge graph further comprises:

[0040] The multi-task learning module is connected with the feature extraction and mapping module and the multi-modal knowledge graph construction module, and is configured to jointly train breast cancer popular science core requirements through a multi-task learning architecture, including image lesion identification, treatment path interpretation, and psychological counseling and emotional support sub-tasks.

[0041] The joint loss function module is connected with the multi-task learning module, and is configured to define and optimize a joint loss function to balance medical professionalism and popular science readability.

[0042] The safety review and confidence judgment module is connected with the retrieval and generation module, and is configured to perform "retrieval" or "prompt referral to professional doctors" on low-confidence outputs in the generation stage, and configure medical expert review or automatic filtering strategies to ensure the accuracy and safety of popular science information.

[0043] The above one or more embodiments of the present application have at least the following beneficial effects:

[0044] Significant improvement of personalized popular science ability: by constructing a multi-modal knowledge graph and integrating breast cancer image samples and text data, personalized popular science content can be provided for users. The system can not only retrieve relevant associated nodes according to user inputted diagnosis information, but also combine supplementary information from external databases to achieve multi-modal RAG enhancement, thereby providing more accurate and comprehensive popular science information. This personalized popular science ability helps users better understand their own condition, reduces unnecessary anxiety and fear, and improves treatment compliance and effectiveness.

[0045] The authority and credibility are greatly improved: the present application divides the core needs of breast cancer popularization into several sub-tasks through a multi-task learning architecture, including image lesion identification, treatment path interpretation, and psychological counseling and emotional support. This multi-task learning mechanism not only improves the generalization ability of the model, but also balances the medical professionalism and popularization readability through a joint loss function, ensuring that the generated popularization content is both professional and easy to understand. In addition, the system also introduces a safety proofreading and confidence judgment mechanism to evaluate the confidence of the generated popularization content, and sets rules for "retrieval" or "prompt referral to professional doctors" for low-confidence output, further ensuring the authority and credibility of the popularization information.

[0046] The interaction and user experience are significantly enhanced: the system supports human-computer interaction dialogue, which can dynamically update the user input information to be diagnosed through multi-round dialogue, realize continuous answers and popularization supplements for personalized problems. This interactive design not only improves the user experience, but also makes the popularization content more suitable for the actual needs of users. At the same time, the system introduces example graphs for breast cancer specific feature nodes, which will be displayed on the user interface when the user input information to be diagnosed is similar to the example graph, making it easy for users to compare and further enhance the visualization and interactivity of the system.

[0047] The applicability and expansibility are stronger: the system design of the present application is modular, with close connections between modules and clear functions, making it easy to maintain and expand. For example, the data acquisition and preprocessing module, the feature extraction and mapping module, the multi-modal knowledge graph construction module, the retrieval and generation module, etc. can be independently optimized and upgraded to adapt to different scenarios and needs. In addition, the system also supports the fusion and retrieval of multi-modal information, which can adapt to different types of user input such as text, image, etc., and has wide applicability and expansibility. BRIEF DESCRIPTION OF DRAWINGS

[0048] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, which together with the embodiments of the present application, is used to explain the present application, and does not constitute a limitation on the present application. In the drawings:

[0049] Figure 1 is a flowchart of a breast cancer disease popularization method based on a multi-modal knowledge graph according to an embodiment of the present application;

[0050] Figure 2 is a fine-tuning task flowchart in a breast cancer disease popularization method based on a multi-modal knowledge graph according to an embodiment of the present application. DETAILED DESCRIPTION

[0051] Embodiments of this application will now be described in detail, examples of which are illustrated in the accompanying drawings. The components of the embodiments of this application described and shown in the drawings herein can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application.

[0052] Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0053] In the description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0054] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0055] The following will combine Figure 1 and Figure 2 The technical solutions of this application are clearly and completely described. Obviously, the described embodiments are only some embodiments of this application, not all embodiments.

[0056] Please see Figure 1 and Figure 2 This application provides a method for popularizing breast cancer knowledge based on multimodal knowledge graphs, including:

[0057] S100: Acquire and process breast cancer image samples and text data.

[0058] The present application first needs to collect a large number of breast cancer image samples and text data. Image samples include but are not limited to mammography, magnetic resonance imaging (MRI), ultrasound and other image data, and text data includes clinical guidelines, medical papers, medical records, diagnosis and treatment records, etc. These data sources are extensive and cover different stages, different subtypes, and different lesion morphologies of breast cancer cases to cover common lesions (such as lumps, microcalcifications, and structural distortion) and rare special lesions, thereby ensuring the comprehensiveness and diversity of the data.

[0059] The collected breast cancer image samples and text data need to be preprocessed to improve data quality and usability.

[0060] In some embodiments, the preprocessing step includes:

[0061] Denoising: using filtering algorithms to remove noise in breast cancer images to improve image clarity.

[0062] Normalization: normalize image data to a uniform intensity range to ensure data consistency.

[0063] Contrast enhancement: enhance the contrast of the image by histogram equalization and other methods to make the lesions more obvious.

[0064] For text data, professional terminology alignment is needed, and unified indexing with image metadata is needed. For example, associate medical terms in the text with lesion location, size, and other information in the image to facilitate subsequent multi-modal fusion.

[0065] In some embodiments, a traditional segmentation or detection algorithm Faster R-CNN is used to preliminarily locate the lesions in the preprocessed breast cancer images, label key information (ROI), and record it in the image metadata to lay the foundation for subsequent multi-modal alignment. In one specific example, this step includes:

[0066] Input the breast image into the CNN network to obtain the corresponding feature map;

[0067] Use RPN to generate candidate regions, which may contain lesions. RPN generates a series of candidate regions on the feature map through sliding windows, and generates a fixed-size feature vector for each region;

[0068] Scale the feature matrix of each candidate region to a 7x7 feature map through the ROIPooling layer, then flatten the feature map into a vector, and get the prediction result through a series of fully connected layers;

[0069] Classify and regress the candidate regions through the fully connected layer and the softmax function, and output the final detection result.

[0070] The detected lesion area is labeled to generate key information (ROI) and record it in the metadata. These information includes the location, size, shape, etc. of the lesion, providing a basis for subsequent multi-modal alignment.

[0071] Here are some explanations for the above terms:

[0072] Faster R-CNN: A high-efficiency target detection algorithm that realizes the end-to-end training and detection of candidate region generation, feature extraction, boundary box classification and boundary box regression through the region proposal network (RPN).

[0073] CNN: Convolutional Neural Network, a deep learning model that automatically extracts image features through convolution, pooling and fully connected layers, widely used in image classification, target detection and other tasks.

[0074] RPN: Region Proposal Network, the core component of Faster R-CNN, generates a series of candidate regions on the feature map through sliding window, and generates a fixed-size feature vector for each region.

[0075] ROIPooling: Region of Interest Pooling, an operation that maps candidate region feature matrices of different sizes to a fixed-size feature map, so that the subsequent fully connected layer can be processed.

[0076] Softmax function: A mathematical function that converts the numerical output of the model into a probability distribution, thereby realizing the multi-classification task.

[0077] S200, extract breast cancer-specific features from breast cancer image samples.

[0078] First, use a deep learning model to analyze the pre-processed breast cancer image samples and extract breast cancer-specific features. These features include but are not limited to mass morphology, density, microcalcification, etc. Then, combine the image bottom-level features with high-level attention mechanisms to form a comprehensive representation of local texture and global context. This fusion method can more comprehensively capture the key information in the image and improve the expression ability of the features.

[0079] In some embodiments, a visual Swin Transformer architecture specifically pre-trained for breast cancer imaging is employed, focusing on fine-grained extraction of mass morphology, density, and microcalcification features. Swin Transformer is a Transformer-based visual model that addresses the high computational complexity of traditional Transformer models in computer vision tasks by introducing a hierarchical architecture and sliding window mechanism. It is widely used in image classification, object detection, and segmentation tasks. Swin Transformer effectively addresses the computational problem of high-resolution images through hierarchical feature extraction and window self-attention mechanisms, while gradually reducing resolution to extract high-level semantic information and reduce computational load.

[0080] In some embodiments, combining image bottom-level features with high-level attention mechanisms to form a comprehensive representation of local texture and global context includes the following steps:

[0081] Local texture feature extraction: Convolutional Neural Networks (CNN) are used to extract local texture features of the image, such as the boundaries, shapes, and internal structure features of the mass.

[0082] Global context feature extraction: Through self-attention mechanisms, global context information of the image is captured, such as the density distribution and uniformity features of the mass.

[0083] Feature fusion: Local texture features and global context features are fused through feature concatenation or weighted fusion to form a comprehensive representation.

[0084] S300, create corresponding nodes for the breast cancer-specific features and establish a mapping relationship with the medical diagnosis information to form a multi-modal unified representation.

[0085] In some embodiments, the specific steps for creating corresponding nodes for different image features and establishing a mapping relationship with actual medical diagnosis information to form a multi-modal unified representation include:

[0086] Node creation: Create corresponding nodes for each breast cancer-specific feature (such as mass morphology, density, microcalcification, etc.). For example, create "microcalcification cluster" and "string-like mass" nodes.

[0087] Relationship establishment: Connect feature nodes and medical diagnosis information (such as BIRADS classification, HER2 status, TNM staging, etc.) through logical relationship edges. For example, establish "a certain pathological type is related to a certain imaging sign" or "a certain treatment plan is suitable for a certain stage" logical relationship to form a multi-modal unified representation.

[0088] The following is an explanation of some of the above-mentioned terms:

[0089] BIRADS classification: BIRADS (Breast Imaging Reporting and Data System) classification is part of the Breast Imaging Reporting and Data System, used to evaluate the malignancy of breast lesions, divided into 0 to 6 levels, where 0 level indicates the need for further evaluation, and 6 level indicates pathologically confirmed malignant lesions.

[0090] HER2 status: HER2 (Human Epidermal growth factor Receptor 2) status refers to the expression status of human epidermal growth factor receptor 2, which is an important biomarker in breast cancer diagnosis and treatment, used to determine the type and prognosis of breast cancer, and guide the selection of treatment plan.

[0091] TNM staging: TNM (Tumor, Node, Metastasis) staging is an internationally recognized tumor staging system that evaluates tumor size (T), lymph node metastasis (N), and distant metastasis (M) to determine tumor staging, guide treatment and prognosis evaluation.

[0092] S400, associate medical text entities in text data with breast cancer-specific feature nodes to form a multi-modal knowledge graph.

[0093] In the construction of multi-modal breast cancer knowledge graph, text entities (such as diagnosis and treatment guidelines, medical record information) and image feature nodes are linked to each other, representing logical relationships such as "a certain pathological type is related to a certain image sign" or "a certain treatment plan is suitable for a certain stage". The specific steps are as follows:

[0094] Entity recognition and relationship extraction: Use natural language processing technology to identify medical entities (such as pathological types, treatment plans) and their relationships from text. For example, extract "a certain pathological type is related to a certain image sign" or "a certain treatment plan is suitable for a certain stage" from diagnosis and treatment guidelines.

[0095] Multi-modal unified representation: Align image feature nodes and text information to form a multi-modal unified representation for comprehensive retrieval and generation in the knowledge graph. For example, align the mass feature in the image with the diagnostic information in the text to form a multi-modal unified representation.

[0096] In some embodiments, an example graph or lesion screenshot is introduced for each image feature node to facilitate subsequent visual retrieval and comparison.

[0097] S500, obtain user input of information to be diagnosed, retrieve associated nodes related to the information to be diagnosed from the multi-modal knowledge graph, and output popular science content.

[0098] User input: The user inputs information to be diagnosed, which can be image information, text information, or a combination of both. After the user inputs the text or uploads the image, the most similar concept nodes and literature passages are first retrieved in the multi-modal knowledge graph, and then supplemented with information from external databases (such as medical papers and guidelines) to achieve multi-modal RAG (Retrieval-Augmented Generation) enhancement. This process is dynamically updated according to the context in multi-turn dialogue, enabling continuous answers to personalized questions and supplementary popular science.

[0099] In some embodiments, the specific steps are as follows:

[0100] User input processing: The user input text or image is preprocessed to extract key information. For example, the user input image is preliminarily positioned for lesion, key information (ROI) is labeled, and it is recorded in the metadata.

[0101] Graph retrieval: The most similar concept nodes and literature passages are retrieved in the knowledge graph, supplemented with information from external databases to achieve multi-modal RAG enhancement. For example, medical papers and guidelines related to the image features input by the user are retrieved.

[0102] Dynamic updating and generation: In multi-turn dialogue, the retrieved information is dynamically updated according to the user's subsequent questions and feedback, and personalized popular science content is generated.

[0103] In some embodiments, refer to Figure 2 To improve the credibility and medical applicability of breast cancer popular science, the present application designs a special fine-tuning task, uses a multi-modal large model (such as Transformer) as the base model, and uses breast cancer related data to fine-tune the multi-modal large model, so that the model can better understand and process breast cancer related image and text data. During the fine-tuning process, multiple tasks are learned simultaneously, such as lesion identification and popular science content generation. Through multi-task learning, the model can share underlying features and complement each other in high-level semantics, improving the generalization ability and accuracy of the model. The fine-tuned model is combined with the knowledge graph to retrieve and enhance information using the information in the knowledge graph. And through the GraphRAG (Graph-based Retrieval Augmented Generation) technology, multi-modal retrieval enhancement is achieved.

[0104] The fine-tuning task of the present application decomposes the core requirements of breast cancer popular science into several sub-tasks through a multi-task learning architecture and performs joint training. Specifically, it includes the following aspects:

[0105] Image lesion identification:

[0106] Using deep learning models such as Faster R-CNN to analyze breast images, automatically detect and label suspected lesion areas. For example, the system can detect lumps and microcalcifications in mammography images and generate labeling information (ROI), providing a basis for subsequent diagnosis.

[0107] Treatment Path Interpretation:

[0108] Providing personalized treatment plan popularization based on image and pathology information: combining image features and pathology information, the system can generate personalized treatment plan suggestions. For example, based on the size, location and pathology type of the lump, the system can suggest further examination or treatment for the user.

[0109] Psychological counseling and emotional support:

[0110] Providing users with popular psychological comfort and nursing suggestions: through natural language processing technology, the system can generate psychological comfort and nursing suggestions to help users alleviate anxiety and fear. For example, the system can provide psychological support content about early symptoms of breast cancer, prevention measures and treatment process.

[0111] Through multi-task learning, share underlying features and complement each other in high-level semantics, so as to obtain better results in breast cancer image and text processing, popular interpretation, clinical plan explanation, etc. The specific steps are as follows:

[0112] Data preparation: Collect a large number of breast cancer image samples and text data, including mammography, MRI, ultrasound and other image data, as well as clinical guidelines, medical papers, case reports and other text data.

[0113] Feature extraction: Use deep learning models such as Swin Transformer to extract image features, and use natural language processing techniques to extract text features.

[0114] Multi-task training: Joint training of image lesion recognition, treatment path interpretation and psychological counseling and emotional support, sharing underlying features and complementing each other in high-level semantics.

[0115] Model optimization: Through joint loss function optimization, balance medical professionalism and popular readability, ensure that the generated popular content is both professional and easy to understand.

[0116] Further, in the training process, the following joint loss function can be defined To balance medical professionalism and popular readability:

[0117]

[0118] Where, Medical professional task loss, used to measure the accuracy or cross-entropy loss of image recognition, pathological staging prediction professional tasks; Popular science text loss, used to measure the readability and relevance of popular science text; α is the weight coefficient of medical professional task loss; β is the weight coefficient of popular science text loss; y k True value label of professional task; Model predicted probability; K represents the total number of categories in the medical professional task; k represents the category index, from 1 to K; x i Word or sentence to be predicted in popular science generation; x context Context information when generating text; M represents the total number of words or sentences in the popular science text; P represents the probability of generating a certain word or sentence.

[0119] Joint loss function The objective function that the model needs to minimize during training. It takes into account two aspects:

[0120] Medical professionalism: Ensure the accuracy of the model on the medical professional task.

[0121] Popular science readability: Ensure that the popular science text generated by the model is easy to understand and relevant.

[0122] By adjusting the values of α and β, you can control the importance the model places on these two goals during training. For example, if the value of α is large, the model will pay more attention to the accuracy of the medical professional task; if the value of β is large, the model will pay more attention to the readability of the popular science text.

[0123] In some embodiments, the breast cancer disease popular science method based on the multi-modal knowledge graph provided by the present application further includes safety proofreading and confidence judgment, specifically including:

[0124] Confidence evaluation:

[0125] Confidence evaluation of generated popular science content, according to the evaluation result to decide whether to need further audit or adjustment. Confidence evaluation is realized through the confidence scoring mechanism inside the model, for example, using the probability distribution or confidence score output by the model to evaluate the reliability of the content.

[0126] Retrieval mechanism:

[0127] For content with a confidence score below a preset threshold, automatically trigger the "retrieval" mechanism to retrieve relevant information from the multi-modal knowledge graph and external databases to improve the accuracy of the content. For example, the system can call the GraphRAG mechanism again to retrieve concept nodes and literature passages related to user input, and generate more accurate popular science content by combining new information.

[0128] Referral to a professional doctor:

[0129] For content with a confidence level below a preset threshold, the system can also prompt the user to refer to a professional doctor to ensure that the user receives the most accurate and reliable medical advice. For example, the system can generate the following prompt: "Based on the current information, it is recommended that you consult a professional doctor for more detailed diagnosis and treatment recommendations."

[0130] Sensitive topics and high-risk information review:

[0131] For certain sensitive topics or high-risk information, configure medical expert review or automatic filtering strategies to ensure the accuracy and safety of popular science information. For example, the system can automatically detect and filter out content containing sensitive information, such as information related to privacy or information that may cause panic.

[0132] In some embodiments, the breast cancer disease popular science method based on multi-modal knowledge graph provided by the present application also includes interactive deployment and scenario application, specifically as follows:

[0133] 1. Visual image interpretation

[0134] When the user uploads a breast image, the model first performs lesion detection and comparison with the knowledge graph node. If a similar example graph is found, the suspicious area is highlighted on the user interface, and a "high similarity to graph XX node / literature" note is given. The system outputs the matched image and text fusion, allowing users to read the text description while also visually observing the suspicious feature annotations in the image. For example, the user uploads a breast X-ray mammography image, and the system detects an irregular mass and finds a similar example graph in the knowledge graph. Then, the system highlights the area on the user interface and gives the note: "This area has a high similarity to the 'irregular mass' node in the graph, and further examination is recommended."

[0135] 2. Multi-round dialogue and dynamic update

[0136] The system retains the content of the previous rounds of dialogue and the retrieved knowledge nodes. If the user continues to ask questions such as "Is the mass suitable for breast-conserving surgery?" and "What should be paid attention to during postoperative rehabilitation?", the model can call the GraphRAG mechanism to search for matching information in the graph and external literature again, maintaining personalized and context-related responses. If the model has insufficient confidence in answering certain questions, it will pop up a safety reminder or additional "refer to the doctor's diagnosis" prompt to reduce the risk of incorrect or misleading answers. For example, the user continues to ask "Is the mass suitable for breast-conserving surgery?", and the system calls the GraphRAG mechanism to search the knowledge graph and external literature, generating the following content: "Based on the size and location of the mass, breast-conserving surgery is a viable option, but it needs to be combined with pathological examination results to determine. It is recommended that you consult a professional doctor for more detailed diagnosis and treatment recommendations."

[0137] 3. Clinical and popular science dual-scene unification

[0138] In the clinical assistance-oriented scenario, more professional terms and literature references are output; in the popular science-oriented scenario, more concise and readable text and image descriptions are output, making it easy for users to understand the core points. For example, in the clinical assistance scenario, the system can output: "According to the BIRADS classification, the lump is preliminarily evaluated as level 4, and further pathological examination is recommended." In the popular science scenario, the system can output: "A suspicious lump is detected in the image, and you are recommended to have regular breast examinations and consult a doctor for more information."

[0139] The second aspect of the present application provides a breast cancer disease popular science system based on a multi-modal knowledge graph, which includes a data acquisition and preprocessing module, a feature extraction and mapping module, a multi-modal knowledge graph construction module, a retrieval and generation module, etc.

[0140] The data acquisition and preprocessing module is used to acquire a large amount of breast cancer image samples and text data. The image samples include but are not limited to mammography, MRI, ultrasound, and other image data, and the text data includes clinical guidelines, medical papers, case reports, and medical records.

[0141] The data acquisition and preprocessing module is also used to preprocess the collected data, including denoising, normalization, and contrast enhancement, to improve data quality and usability. For text data, professional terms are aligned and indexed with image metadata (such as staging, pathological type, etc.).

[0142] The feature extraction and mapping module is connected to the data acquisition and preprocessing module and can use deep learning models (such as Swin Transformer) to analyze the preprocessed breast cancer image samples, extract breast cancer-specific features such as lump morphology, density, and microcalcification features, and combine breast cancer-specific bottom-layer features with high-level attention mechanisms to realize comprehensive representation of local texture and global context.

[0143] The feature extraction and mapping module can also create corresponding nodes for the extracted breast cancer-specific features and establish mapping relationships with medical diagnosis information (such as BIRADS classification, HER2 status, TNM staging), forming a multi-modal unified representation.

[0144] The multi-modal knowledge graph construction module is connected to the feature extraction and mapping module and is used to associate medical text entities in the text data with breast cancer-specific feature nodes to form a multi-modal knowledge graph.

[0145] The retrieval and generation module is connected with the multi-modal knowledge graph construction module, and is used to obtain the to-be-diagnosed information input by a user, retrieve the associated nodes related to the to-be-diagnosed information from the multi-modal knowledge graph, and output popular science content.

[0146] In some embodiments, the multi-modal knowledge graph-based breast cancer disease popular science system further comprises a multi-task learning module, a joint loss function module, and a safety review and confidence judgment module.

[0147] The multi-task learning module is connected with the feature extraction and mapping module and the multi-modal knowledge graph construction module, and is used to jointly train the breast cancer popular science core requirements through a multi-task learning architecture, including sub-tasks such as image lesion identification, treatment path interpretation, and psychological counseling and emotional support.

[0148] The joint loss function module is connected with the multi-task learning module, and is used to define and optimize a joint loss function to balance medical professionalism and popular science readability.

[0149] The safety review and confidence judgment module is connected with the retrieval and generation module, and is used to perform "retrieval" or "prompt referral to a professional doctor" on low-confidence outputs in the generation stage, and configure a medical expert review or automatic filtering strategy to ensure the accuracy and safety of the popular science information.

[0150] The following is a specific process description of the present application:

[0151] (1) Image preprocessing

[0152] The system first collects and preprocesses various breast cancer images and text data:

[0153] Denoising: Use Gaussian filtering and other techniques to remove noise interference in the image, making the image clearer.

[0154] Normalization: Adjust the image data to a uniform intensity range to ensure comparability between different images.

[0155] Contrast enhancement: Use histogram equalization and other methods to enhance the contrast of the image and highlight the lesion area.

[0156] For text data, it needs to be cleaned and aligned with professional terms, and indexed with image metadata. For example, associate medical terms in the text with lesion location, size, and other information in the image for subsequent multi-modal fusion.

[0157] Use the traditional segmentation or detection algorithm Faster R-CNN to preliminarily locate the lesions of the preprocessed breast cancer images, label the key information (ROI), and record it in the image metadata to lay the foundation for subsequent multi-modal alignment.

[0158] For example, an original mammography image may have high noise and low contrast due to limited shooting conditions. After the preprocessing steps described above, the lesion features such as mass and microcalcification in the image will be marked with information, making them more obvious and laying the foundation for subsequent feature extraction and analysis.

[0159] (2) Feature extraction

[0160] After preprocessing, the system uses a deep learning model (such as Swin Transformer) to extract features from the image:

[0161] Mass morphology feature extraction: Through a multi-scale window mechanism, the boundary, shape, and internal structure features of the mass are captured to determine whether it is regular.

[0162] Density feature extraction: Using a multi-layer perceptron (MLP) and self-attention mechanism, the density distribution and uniformity features of the mass are extracted to distinguish different density types.

[0163] Microcalcification feature extraction: Through high-resolution feature maps, the size, shape, and distribution features of microcalcification lesions are detected to identify their morphology.

[0164] Combine the low-level features of the image with high-level attention mechanisms to form a comprehensive representation of local textures and global context.

[0165] For example, the system may extract the morphological features of a mass as irregular, the density features as high density and uneven distribution, and the microcalcification features as cluster distribution. These features will serve as important evidence for subsequent analysis.

[0166] (3) Feature mapping

[0167] The system creates corresponding nodes for the extracted features in the knowledge graph and establishes a mapping relationship with medical diagnosis information:

[0168] Node creation: Create nodes for each feature, such as "irregular mass", "high-density mass", "microcalcification cluster", etc.

[0169] Relationship establishment: Associate feature nodes with medical diagnosis information (such as BIRADS classification, HER2 status, TNM staging).

[0170] For example, connect the "irregular mass" node with the "BIRADS 4" node, indicating that this feature is related to a higher risk of malignancy. This mapping relationship helps the system understand the clinical significance of the features.

[0171] (4) Knowledge graph construction

[0172] The system associates medical entities in text data with image feature nodes, forming a multi-modal knowledge graph:

[0173] Entity recognition and relationship extraction: Using natural language processing techniques to identify medical entities and their relationships from text.

[0174] Image feature node creation: Create nodes for each image feature and establish a mapping with medical diagnosis information.

[0175] Multi-modal unified representation: Align image features and text information to form a unified representation.

[0176] For example, extract relationships such as "a certain pathological type is related to a certain imaging sign" or "a certain treatment plan is suitable for a certain stage" from clinical guidelines, and connect these relationships with image feature nodes to form a comprehensive knowledge graph.

[0177] (5) Retrieval

[0178] The system retrieves nodes related to the user's input text or image from the multi-modal knowledge graph:

[0179] For the user's input text or image, retrieve the most similar concept nodes and literature passages in the knowledge graph, and combine the supplementary information from external databases to achieve multi-modal RAG enhancement.

[0180] For example, the user uploads a mammography image, and the system retrieves similar nodes such as "irregular mass" and "high-density mass" in the knowledge graph, and combines relevant literature to provide a basis for subsequent popular science content generation.

[0181] (6) Generation

[0182] The system generates personalized popular science content based on the retrieved nodes:

[0183] Content generation: Use the retrieved node information to generate popular science content containing images and text.

[0184] Human-computer interaction: Through multiple rounds of dialogue and dynamic updates, continuously adjust and refine the popular science content according to the user's follow-up questions and feedback.

[0185] For example, the system generates "an irregular mass with high density and microcalcification is detected in the image, further pathological examination is recommended to confirm the diagnosis." The user continues to ask "is the mass suitable for breast-conserving surgery?" The system calls the GraphRAG mechanism, retrieves the knowledge graph and external literature, and generates "according to the size and location of the mass, breast-conserving surgery is a feasible option, but needs to be combined with pathological examination results to determine. It is recommended that you consult a professional doctor to obtain more detailed diagnosis and treatment recommendations."

[0186] (7) Safety Review and Confidence Assessment

[0187] The system performs confidence assessment on the generated popular science content:

[0188] Confidence Calculation: Assess the reliability of the content based on the probability distribution or confidence score output by the model.

[0189] Threshold Judgment: For content with a confidence level below the preset threshold, automatically trigger the "retrieval again" or "prompt referral to professional doctors" mechanism to improve the accuracy of the content.

[0190] For example, if the system's confidence in a certain diagnosis suggestion is below 80%, it will automatically trigger a re-retrieval to retrieve relevant information from the multi-modal knowledge graph and external databases to improve the accuracy of the content. If it still cannot achieve sufficient confidence, it will prompt the user to refer to a professional doctor.

[0191] (8) Interactive Deployment and Scenario Application

[0192] The system provides visual image interpretation for users, supports multi-round dialogue and dynamic updating, and realizes the unification of clinical and popular science scenes:

[0193] Visual image interpretation: Highlight suspicious areas on the user interface and provide explanations so that users can directly see the lesion characteristics.

[0194] Multi-round dialogue and dynamic updating: Preserve the dialogue history and call the GraphRAG mechanism according to the user's subsequent questions to maintain the context relevance of the replies.

[0195] Clinical and popular science scenes are unified: According to user needs, output professional terminology and literature references or simple and easy-to-understand popular science content.

[0196] For example, in the clinical assistance scene, the system outputs "According to the BIRADS classification, the lump is preliminarily evaluated as level 4, and further pathological examination is recommended." In the popular science scene, the system outputs "A suspicious lump is detected in the image, and you are advised to have regular breast examinations and consult a doctor for more information."

[0197] It should be noted that the technical solutions in each embodiment of the present application can be combined with each other, but the basis for mutual combination is that it can be realized by ordinary skilled personnel in the art; when the combination of technical solutions contradicts each other or cannot be realized, it should be considered that the combination of technical solutions does not exist, i.e. it is not within the protection scope of the present application.

[0198] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the same; although the present application has been described in detail with reference to the foregoing examples, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A breast cancer disease popularization method based on a multi-modal knowledge graph, characterized in that, The method comprises the following steps: Collect and process breast cancer image samples and text data; Extract breast cancer-specific features from breast cancer image samples; In the knowledge graph, create corresponding nodes for the breast cancer-specific features and establish a mapping relationship with the medical diagnosis information to form a multi-modal unified representation; Associate the text data with the breast cancer-specific feature nodes to form a multi-modal knowledge graph; Obtain user inputted diagnosis information, retrieve associated nodes related to the diagnosis information from the multi-modal knowledge graph, and output popular science content; The popular science method further comprises designing a fine-tuning task to improve the credibility and medical adaptability of the breast cancer popular science content, and the fine-tuning task at least comprises: Decompose the breast cancer popular science core demand into several sub-tasks and perform joint training, the sub-tasks comprising: Automatically detect and label suspected morphology, density and micro-calcification features; Provide personalized treatment plan popular science according to breast cancer images and pathological information; Provide popular science psychological comfort and nursing suggestions for users; During the training process, a joint loss function is defined to balance the medical professionalism and popular science readability; The joint loss function is calculated using the following equation: wherein, is a medical professional task loss, used to measure the accuracy or cross-entropy loss of the professional task of image recognition, pathological staging prediction; is a popular science text loss, used to measure the score of the popular science text in terms of readability and relevance; is a weight coefficient of the medical professional task loss; is a weight coefficient of the popular science text loss; represents the true value label of the professional task; represents the probability predicted by the model; represents the total number of categories in the medical professional task; represents the category index, from 1 to represents the word or sentence to be predicted in the popular science generation; represents the context information when generating the text; represents the total number of words or sentences in the popular science text; represents the probability of generating a certain word or sentence; The breast cancer-specific features extracted from the breast cancer image samples specifically comprise: Using a deep learning model to analyze the breast cancer image samples to extract mass morphology, density and micro-calcification features; Combine image low-level features with high-level attention mechanisms to form a comprehensive representation of local texture and global context. 2.The breast cancer disease popularization method based on a multi-modal knowledge graph according to claim 1, characterized in that, Collecting and processing breast cancer image samples and text data specifically comprises: Performing processing operations on the breast cancer image samples, the processing operations at least comprising one of denoising, normalization, and contrast enhancement; Performing preliminary lesion positioning and information labeling on the processed breast cancer image samples and recording them in the breast cancer image sample metadata; Align the text data with professional terms and uniformly index the aligned text data and breast cancer image sample metadata. 3.The breast cancer disease popularization method based on a multi-modal knowledge graph according to claim 1, characterized in that, Introduce an example graph for the breast cancer-specific feature nodes, and when the user inputted diagnosis information is similar to the example graph, it will be displayed on the user interface to facilitate comparison. 4.The breast cancer disease popularization method based on a multi-modal knowledge graph according to claim 1, wherein, Obtain user inputted diagnosis information, retrieve associated nodes related to the diagnosis information from the multi-modal knowledge graph, and combine external database supplementary information to achieve multi-modal RAG enhancement. 5.The breast cancer disease popularization method based on a multi-modal knowledge graph according to claim 4, characterized in that, Obtain user inputted diagnosis information through human-computer interaction dialogue, and dynamically update the user inputted diagnosis information according to the context in multi-turn dialogue to achieve continuous answers to personalized questions and popular science supplements. 6.The breast cancer disease popularization method based on a multi-modal knowledge graph according to claim 1, wherein, It also includes safety review and confidence judgment, specifically including: Conduct confidence evaluation on the generated popular science content, and decide whether further review or adjustment is needed according to the evaluation results; For content with a confidence level below a preset threshold, automatically trigger the "retrieval" or "prompt referral to a professional doctor" mechanism to improve the accuracy of the content; For some sensitive topics or high-risk information, configure medical expert review or automatic filtering strategies to ensure the accuracy and safety of the popular science information.

7. A breast cancer disease popularization system based on a multi-modal knowledge graph, characterized in that, The method comprises the following steps: A data collection and preprocessing module is used to collect and process breast cancer image samples and text data; The feature extraction and mapping module is connected with the data acquisition and preprocessing module, and is configured to extract breast cancer specific features from the processed breast cancer image samples, and create corresponding nodes in the knowledge graph, and establish a mapping relationship with medical diagnosis information. The multi-modal knowledge graph construction module is connected with the feature extraction and mapping module, and is configured to associate medical text entities in the text data with breast cancer specific feature nodes to form a multi-modal knowledge graph. And The retrieval and generation module is connected with the multi-modal knowledge graph construction module, and is configured to obtain user inputted diagnosis information, retrieve associated nodes related to the diagnosis information from the multi-modal knowledge graph, and output popular science content. The popular science system further includes designing fine-tuning tasks to improve the credibility and medical adaptability of the breast cancer popular science content, and the fine-tuning tasks at least include: Decomposing breast cancer popular science core needs into several sub-tasks and performing joint training, the sub-tasks including: Automatically detecting and labeling suspected morphology, density and micro-calcification features; Providing individualized treatment plan popular science according to breast cancer images and pathological information; Providing popular science psychological comfort and nursing suggestions for users; During the training process, a joint loss function is defined to balance medical professionalism and popular science readability; The joint loss function is calculated using the following equation: wherein, is a medical professional task loss, used to measure the accuracy or cross-entropy loss of the image recognition, pathological stage prediction professional task; is a popular science text loss, used to measure the score of the popular science text in terms of readability and relevance; is a weight coefficient of the medical professional task loss; is a weight coefficient of the popular science text loss; represents the true value label of the professional task; represents the probability predicted by the model; represents the total number of categories in the medical professional task; represents the category index, from 1 to represents the word or sentence to be predicted in the popular science generation; represents the context information when generating the text; represents the total number of words or sentences in the popular science text; represents the probability of generating a certain word or sentence; Wherein, the breast cancer specific features extracted from the breast cancer image samples specifically include: Using a deep learning model to analyze the breast cancer image samples to extract tumor morphology, density and micro-calcification features; Combining image low-level features with high-level attention mechanisms to form a comprehensive representation of local texture and global context.

8. The breast cancer disease popularization system based on a multi-modal knowledge graph according to claim 7, characterized in that, Further comprising: The multi-task learning module is connected with the feature extraction and mapping module and the multi-modal knowledge graph construction module, and is configured to perform joint training of breast cancer popular science core needs through a multi-task learning architecture, including image lesion identification, treatment path interpretation, and psychological counseling and emotional support sub-tasks; The joint loss function module is connected with the multi-task learning module, and is configured to define and optimize a joint loss function to balance medical professionalism and popular science readability; And The safety review and confidence judgment module is connected with the retrieval and generation module, and is configured to perform "retrieval" or "prompt referral to professional doctors" on low-confidence outputs during the generation stage, and configure medical expert review or automatic filtering strategies to ensure the accuracy and safety of popular science information.

Citation Information

Patent Citations

  • Method for establishing medical image atlas based on image segmentation and convolutional neural network (CNN)

    CN108389614A

  • Breast X-ray image classification model training method based on implicit apparent learning

    CN111415741A

  • CT report generation processing method and device

    CN117995344A

  • Large medical model intelligent inquiry reasoning method and system

    CN119380966A