Breast cancer science popularization method and system based on multi-modal knowledge graph

By constructing a multimodal knowledge graph, integrating breast cancer images and text data, extracting specific features, and establishing mapping relationships, the problem of insufficient personalization and interaction in the existing breast cancer popularization methods is solved, and the generation of popular science content with personalization, authoritativeness and interaction is achieved.

CN120299676AActive Publication Date: 2025-07-11BEIJING FUYU MEDICAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510369301.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing breast cancer popular science methods lack the ability to integrate personalized and multimodal information, and cannot provide authoritative and interactive popular science content, which is difficult to meet the personalized needs of users.

Method used

Build a multimodal knowledge graph, collect and process breast cancer images and text data, extract specific features, establish mapping relationships, and combine deep learning models and natural language processing technology to achieve the generation and interactive dialogue of personalized popular science content.

Benefits of technology

实现了个性化科普能力的提升,提高了科普内容的权威性和交互性,增强了用户体验,适用性和扩展性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299676A_ABST
    Figure CN120299676A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical information analysis, in particular to a breast cancer disease science popularization method and system based on a multi-modal knowledge graph. The method comprises the following steps: acquiring and processing breast cancer image samples and text data; breast cancer specific features are extracted from the breast cancer image sample; in a knowledge graph, creating corresponding nodes for the breast cancer specific characteristics, and establishing a mapping relation with medical diagnosis information to form multi-modal unified representation; associating medical text entities in the text data with breast cancer specific feature nodes to form a multi-modal knowledge graph; and obtaining to-be-diagnosed information input by a user, retrieving the associated nodes related to the to-be-diagnosed information from the multi-modal knowledge graph, and outputting popular science content. According to the application, individuation, authority and interactivity of breast cancer science popularization can be realized, and high-quality science popularization service is provided for patients and the public.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of medical information analysis, and particularly to a breast cancer disease popularization method and system based on a multimodal knowledge graph. Background Art

[0002] Breast cancer is one of the malignant tumors with relatively high incidence and mortality rates among women globally. Despite the continuous improvement of medical standards, there are still deficiencies among the majority of patients and the general public in terms of breast cancer disease awareness and access to popular science knowledge. Traditional breast cancer popularization methods mainly rely on static materials or offline education, with slow information update speeds and difficulty in meeting the personalized needs of different groups. In addition, existing online Q&A and information retrieval tools have significant limitations in terms of content authority, traceability of information sources, integration of multimodal information, and interactivity with users.

[0003] Specifically, traditional popularization mainly uses text or simple diagrams, lacking in-depth analysis of patients' specific situations (such as imaging examination results, pathological diagnosis reports, etc.), resulting in users having difficulty obtaining personalized and interpretable knowledge support. Pure text retrieval engines or general Q&A systems often cannot guarantee the medical rigor and credibility of the results. Once the information queried by users is misleading or missing, it may instead bring adverse consequences. At the same time, the lack of structured knowledge management tools such as knowledge graphs makes existing information scattered across different data sources, making it difficult to achieve the organic integration and efficient retrieval of multimodal information. Moreover, traditional systems lack interactivity and context understanding capabilities and cannot continuously adjust and refine popularization content based on users' subsequent questions and feedback, thus affecting the user experience and knowledge absorption effect.

[0004] Therefore, there is an urgent need for a breast cancer popularization method and system that can integrate multimodal information, provide personalized popularization, and have high authority and interactivity to address the deficiencies in the existing technology. Summary of the Invention

[0005] In view of one or more of the problems existing in the prior art, a first aspect of this application provides a breast cancer disease popularization method based on a multimodal knowledge graph, including:

[0006] Collect and process breast cancer image samples and text data;

[0007] Extract breast cancer-specific features from breast cancer image samples;

[0008] In the knowledge graph, create corresponding nodes for the breast cancer-specific features and establish a mapping relationship with medical diagnosis information to form a multimodal unified representation;

[0009] Associate the text data with the breast cancer-specific feature nodes to form a multimodal knowledge graph;

[0010] Obtain the information to be diagnosed input by the user, retrieve the associated nodes related to the information to be diagnosed from the multi-modal knowledge graph, and output popular science content.

[0011] Preferably, collecting and processing breast cancer image samples and text data specifically includes:

[0012] Perform processing operations on the breast cancer image samples, and the processing operations include at least one of denoising, normalization, and contrast enhancement;

[0013] Perform preliminary lesion localization and information annotation on the processed breast cancer image samples, and record them in the data of the breast cancer image sample metadata;

[0014] Align the text data with professional terms, and perform unified indexing on the aligned text data and breast cancer image sample metadata.

[0015] Preferably, extracting breast cancer specific features from breast cancer image samples specifically includes:

[0016] Use a deep learning model to analyze breast cancer image samples, and extract features of mass morphology, density, and microcalcifications;

[0017] Combine the underlying image features with the high-level attention mechanism to form a comprehensive representation of local texture and global context.

[0018] Preferably, introduce example graphs for breast cancer specific feature nodes, and when the information to be diagnosed input by the user is similar to the example graphs, they will be displayed on the user interface for convenient comparison.

[0019] Preferably, when obtaining the information to be diagnosed input by the user and retrieving the associated nodes related to the information to be diagnosed from the multi-modal knowledge graph, combine the supplementary information in the external database to achieve multi-modal RAG enhancement.

[0020] Preferably, obtain the information to be diagnosed input by the user through the way of human-computer interaction dialogue, and dynamically update the information to be diagnosed input by the user according to the context in multi-round conversations, so as to achieve continuous answering of personalized questions and popular science supplementation.

[0021] Preferably, the breast cancer disease popular science method based on the multi-modal knowledge graph further includes designing fine-tuning tasks to improve the credibility and medical adaptability of breast cancer popular science content, and the fine-tuning tasks include at least:

[0022] Decompose the core requirements of breast cancer popular science into several sub-tasks and perform joint training, and the sub-tasks include:

[0023] Automatically detect and label suspected morphology, density, and microcalcification features;

[0024] Provide popular science on personalized treatment plans based on breast cancer imaging and pathological information;

[0025] Provide popular science-based psychological comfort and nursing suggestions for users.

[0026] Preferably, during the training process, a combined loss function can be defined to balance medical professionalism and popular science readability;

[0027] The combined loss function is calculated using the following formula:

[0028]

[0029] where is the medical professional task loss, used to measure the accuracy or cross-entropy loss of the image recognition and pathological staging prediction professional tasks; is the popular science text loss, used to measure the score of the popular science text in terms of readability and relevance; α is the weight coefficient of the medical professional task loss; β is the weight coefficient of the popular science text loss; y k represents the true label of the professional task; represents the probability predicted by the model; K represents the total number of categories in the medical professional task; k represents the category index, ranging from 1 to K; x i represents the word or sentence to be predicted in the popular science generation; x context represents the context information when generating the text; M represents the total number of words or sentences in the popular science text; P represents the probability of generating a certain word or sentence.

[0030] Preferably, the breast cancer disease popular science method based on the multi-modal knowledge graph further includes safety review and confidence judgment, specifically including:

[0031] Evaluate the confidence of the generated popular science content, and decide whether further review or adjustment is needed according to the evaluation results;

[0032] For content with a confidence level lower than the preset threshold, automatically trigger the "re-retrieval" or "prompt referral to a professional doctor" mechanism to improve the accuracy of the content;

[0033] Configure medical expert review or automatic filtering strategies for certain sensitive topics or high-risk information to ensure the accuracy and safety of the popular science information.

[0034] The second aspect of the present application provides a breast cancer disease popular science system based on a multi-modal knowledge graph, including:

[0035] A data collection and preprocessing module, used to collect and process breast cancer image samples and text data;

[0036] A feature extraction and mapping module, connected to the data acquisition and preprocessing module, is used to extract breast cancer-specific features from the processed breast cancer image samples, create corresponding nodes in the knowledge graph, and establish a mapping relationship with medical diagnosis information;

[0037] A multi-modal knowledge graph construction module, connected to the feature extraction and mapping module, is used to associate medical text entities in the text data with breast cancer-specific feature nodes to form a multi-modal knowledge graph; and

[0038] A retrieval and generation module, connected to the multi-modal knowledge graph construction module, is used to obtain the information to be diagnosed input by the user, retrieve relevant associated nodes from the multi-modal knowledge graph, and output popular science content.

[0039] Preferably, the breast cancer disease popular science system based on the multi-modal knowledge graph further includes:

[0040] A multi-task learning module, connected to the feature extraction and mapping module and the multi-modal knowledge graph construction module, is used to jointly train the core requirements of breast cancer popular science through a multi-task learning architecture, including sub-tasks such as imaging lesion recognition, treatment path interpretation, and psychological counseling and emotional support;

[0041] A joint loss function module, connected to the multi-task learning module, is used to define and optimize the joint loss function to balance medical professionalism and popular science readability; and

[0042] A safety review and confidence judgment module, connected to the retrieval and generation module, is used to perform "re-retrieval" or "prompt referral to a professional doctor" on low-confidence outputs during the generation stage, and configure medical expert review or automatic filtering strategies to ensure the accuracy and safety of popular science information.

[0043] One or more of the above embodiments of the present application have at least the following beneficial effects:

[0044] The personalized popular science ability is significantly improved: By constructing a multi-modal knowledge graph and integrating breast cancer image samples and text data, personalized popular science content can be provided for users. The system can not only retrieve relevant associated nodes according to the information to be diagnosed input by the user, but also combine supplementary information from external databases to achieve multi-modal RAG enhancement, thereby providing more accurate and comprehensive popular science information. This personalized popular science ability helps users better understand their own conditions, reduce unnecessary anxiety and fear, and improve the compliance and effectiveness of treatment.

[0045] The authority and credibility are greatly improved: This application uses a multi-task learning architecture to decompose the core needs of breast cancer popular science into several sub-tasks and conduct joint training, including image lesion identification, treatment path interpretation, psychological counseling and emotional support. This multi-task learning mechanism not only improves the generalization ability of the model, but also balances the medical professionalism and popular science readability through a joint loss function, ensuring that the generated popular science content is both professional and easy to understand. In addition, the system also introduces a security review and confidence judgment mechanism to conduct confidence assessments on the generated popular science content, and sets "re-search" or "prompt referral to a professional doctor" rules for low-confidence outputs, further ensuring the authority and credibility of popular science information.

[0046] Significantly enhanced interactivity and user experience: The system supports human-computer interactive dialogue, and can dynamically update the information to be diagnosed input by the user through multi-round dialogue, achieving continuous answers to personalized questions and popular science supplements. This interactive design not only improves the user experience, but also makes the popular science content more in line with the actual needs of users. At the same time, the system introduces example graphs for breast cancer-specific feature nodes. When the information to be diagnosed input by the user is similar to the example graph, it will be displayed on the user interface to facilitate user comparison, further enhancing the visualization and interactivity of the system.

[0047] Greater applicability and scalability: The system design of this application is modular, with close connections between modules, clear functions, and easy maintenance and expansion. For example, the data acquisition and preprocessing module, feature extraction and mapping module, multimodal knowledge graph construction module, retrieval and generation module, etc. can be independently optimized and upgraded to adapt to different scenarios and needs. In addition, the system also supports the fusion and retrieval of multimodal information, and can adapt to different types of user input, such as text, images, etc., with wide applicability and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings are used to provide a further understanding of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings:

[0049] Figure 1 It is a flow chart of a breast cancer disease popular science method based on a multimodal knowledge graph provided in an embodiment of the present application;

[0050] Figure 2 It is a flowchart of the fine-tuning task in the breast cancer disease popular science method based on the multimodal knowledge graph provided in the embodiment of the present application. DETAILED DESCRIPTION

[0051] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings. Components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application.

[0052] All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the scope of protection of the present application.

[0053] In the description of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0054] In the description of the present application, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0055] The following will be combined with Figure 1 and Figure 2 to clearly and completely describe the technical solution of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all embodiments.

[0056] Please refer to Figure 1 and Figure 2 , a breast cancer disease popularization method based on a multimodal knowledge graph provided by the present application includes:

[0057] S100. Collect and process breast cancer image samples and text data.

[0058] This application first needs to collect a large number of breast cancer image samples and text data. The image samples include, but are not limited to, mammography, Magnetic Resonance Imaging (MRI), ultrasound and other imaging materials, while the text data includes clinical guidelines, medical papers, medical record reports, diagnosis and treatment records, etc. These data sources are extensive, covering breast cancer cases with different stages, different subtypes, and different lesion morphologies, covering common lesions (such as masses, microcalcifications, architectural distortions) and rare special lesions, thus ensuring the comprehensiveness and diversity of the data.

[0059] The collected breast cancer image samples and text data need to be preprocessed to improve the data quality and usability.

[0060] In some embodiments, the preprocessing steps include:

[0061] Denoising: Use a filtering algorithm to remove the noise in the breast cancer images and improve the image clarity.

[0062] Normalization: Normalize the image data to a unified intensity range to ensure the consistency of the data.

[0063] Contrast enhancement: Enhance the contrast of the images through methods such as histogram equalization to make the lesions more obvious.

[0064] For the text data, professional term alignment is required and unified indexing is performed with the image metadata. For example, associate the medical terms in the text with the lesion location, size and other information in the images for subsequent multimodal fusion.

[0065] In some embodiments, use the traditional segmentation or detection algorithm Faster R-CNN to perform preliminary lesion localization on the preprocessed breast cancer images, label the key information (ROI), and record it in the image metadata to lay the foundation for subsequent multimodal alignment. In a specific example, this step includes:

[0066] Input the breast image into the CNN network to obtain the corresponding feature map;

[0067] Use the RPN to generate candidate regions that may contain lesions. The RPN generates a series of candidate regions on the feature map by means of a sliding window and generates a fixed-size feature vector for each region;

[0068] Scale the feature matrix of each candidate region to a 7×7-sized feature map through the ROIPooling layer, then flatten the feature map into a vector, and obtain the prediction result through a series of fully connected layers;

[0069] Classify and regress the candidate regions through the fully connected layer and the softmax function to output the final detection result.

[0070] Label the detected lesion area, generate key information (ROI), and record it in the metadata. This information includes the location, size, shape, etc. of the lesion, providing a basis for subsequent multi-modal alignment.

[0071] The following are the explanations of the above-mentioned terms:

[0072] Faster R-CNN: An efficient object detection algorithm that realizes the end-to-end training and detection of generating candidate regions, feature extraction, bounding box classification, and bounding box regression through the Region Proposal Network (RPN).

[0073] CNN: Convolutional Neural Network, a deep learning model that automatically extracts image features through operations such as convolution, pooling, and fully connected layers, and is widely used in tasks such as image classification and object detection.

[0074] RPN: Region Proposal Network, the core component of Faster R-CNN, which generates a series of candidate regions on the feature map by sliding windows and generates a fixed-size feature vector for each region.

[0075] ROIPooling: Region of Interest Pooling, an operation that maps candidate region feature matrices of different sizes to a fixed-size feature map for subsequent processing by fully connected layers.

[0076] softmax function: A mathematical function used to convert the numerical values output by the model into a probability distribution, thereby realizing multi-classification tasks.

[0077] S200. Extract breast cancer-specific features from breast cancer imaging samples.

[0078] First, use a deep learning model to analyze the preprocessed breast cancer imaging samples and extract breast cancer-specific features. These features include, but are not limited to, mass morphology, density, microcalcifications, etc. Then, combine the underlying image features with the high-level attention mechanism to form a comprehensive representation of local texture and global context. This fusion method can capture key information in the image more comprehensively and improve the expression ability of features.

[0079] In some embodiments, a vision Swin Transformer architecture pre-trained specifically for breast cancer imaging is adopted, focusing on the fine-grained extraction of mass morphology, density, and microcalcification features. Swin Transformer is a vision model based on Transformer. By introducing a hierarchical architecture and a sliding window mechanism, it solves the high computational complexity problem of traditional Transformer models in computer vision tasks and is widely used in tasks such as image classification, object detection, and segmentation. Through hierarchical feature extraction and window self-attention mechanism, Swin Transformer effectively solves the computational problem of high-resolution images, can gradually reduce the resolution, extract high-level semantic information, and reduce the amount of computation.

[0080] In some embodiments, the specific steps of combining the underlying features of the image with the high-level attention mechanism to form a comprehensive representation of local texture and global context include:

[0081] Local texture feature extraction: Use a convolutional neural network (CNN) to extract the local texture features of the image, such as the boundary, shape, and internal structure features of the mass.

[0082] Global context feature extraction: Through the self-attention mechanism, capture the global context information of the image, such as the density distribution and uniformity features of the mass.

[0083] Feature fusion: Combine the local texture features and global context features through feature concatenation or weighted fusion to form a comprehensive representation.

[0084] S300. Create corresponding nodes for the breast cancer-specific features and establish a mapping relationship with the medical diagnosis information to form a multi-modal unified representation.

[0085] In some embodiments, the specific steps of creating corresponding nodes for different image features and establishing a mapping relationship with the actual medical diagnosis information to form a multi-modal unified representation include:

[0086] Node creation: Create corresponding nodes for each breast cancer-specific feature (such as mass morphology, density, microcalcification, etc.). For example, create nodes such as "microcalcification cluster" and "stranded mass".

[0087] Relationship establishment: Connect the feature nodes with the medical diagnosis information (such as BIRADS classification, HER2 status, TNM stage, etc.) through logical relationship edges. For example, establish logical relationships such as "a certain pathological type is related to a certain imaging sign" or "a certain treatment plan is applicable to a certain stage" to form a multi-modal unified representation.

[0088] The following are the explanations of some of the above terms:

[0089] BIRADS Classification: BIRADS (Breast Imaging Reporting and Data System) classification is part of the breast imaging reporting and data system, used to evaluate the malignancy of breast lesions, divided into grades 0 to 6, where grade 0 indicates the need for further evaluation and grade 6 indicates a pathologically confirmed malignant lesion.

[0090] HER2 Status: HER2 (Human Epidermal growth factor Receptor 2) status refers to the expression status of human epidermal growth factor receptor 2, which is an important biomarker in the diagnosis and treatment of breast cancer, used to determine the type and prognosis of breast cancer and guide the selection of treatment regimens.

[0091] TNM Staging: TNM (Tumor, Node, Metastasis) staging is an internationally accepted tumor staging system that determines the stage of a tumor by evaluating the size of the tumor (T), the presence of lymph node metastasis (N), and distant metastasis (M), guiding treatment and prognosis assessment.

[0092] S400. Associate the medical text entities in the text data with the breast cancer-specific feature nodes to form a multimodal knowledge graph.

[0093] When constructing a multimodal breast cancer knowledge graph, text entities (such as clinical practice guidelines, medical record information) are linked to image feature nodes, representing logical relationships such as "a certain pathological type is related to a certain imaging sign" or "a certain treatment regimen is applicable to a certain stage", etc. The specific steps are as follows:

[0094] Entity Recognition and Relationship Extraction: Use natural language processing techniques to identify medical entities (such as pathological types, treatment regimens) and their relationships from the text. For example, extract logical relationships such as "a certain pathological type is related to a certain imaging sign" or "a certain treatment regimen is applicable to a certain stage" from clinical practice guidelines.

[0095] Multimodal Unified Representation: Align the image feature nodes and text information to form a multimodal unified representation for comprehensive retrieval and generation in the knowledge graph. For example, align the mass features in the image with the diagnostic information in the text to form a multimodal unified representation.

[0096] In some embodiments, an example diagram or lesion screenshot is introduced for each image feature node to facilitate subsequent visual retrieval and comparison.

[0097] S500. Obtain the user-input diagnostic information to be diagnosed, retrieve the associated nodes related to the diagnostic information from the multimodal knowledge graph, and output popular science content.

[0098] The information to be diagnosed input by the user is image information, text information, or a combination of both. After the user inputs text or uploads an image, the most similar concept nodes and literature paragraphs are first retrieved from the multi-modal knowledge graph, and then supplemented with information from external databases (such as medical papers and guidelines) to achieve multi-modal Retrieval-Augmented Generation (RAG) enhancement. This process is dynamically updated according to the context in multi-turn conversations to achieve continuous answering of personalized questions and popular science supplementation.

[0099] In some embodiments, the specific steps are as follows:

[0100] User input processing: Preprocess the text or image input by the user to extract key information. For example, perform preliminary lesion localization on the user-input image, label the key information (ROI), and record it in the metadata.

[0101] Knowledge graph retrieval: Retrieve the most similar concept nodes and literature paragraphs from the knowledge graph, and combine the supplementary information from external databases to achieve multi-modal RAG enhancement. For example, retrieve the diagnosis and treatment guidelines and medical papers related to the image features input by the user.

[0102] Dynamic update and generation: In multi-turn conversations, dynamically update the retrieved information according to the user's subsequent questions and feedback, and generate personalized popular science content.

[0103] In some embodiments, referring to Figure 2 , in order to improve the credibility and medical adaptability of breast cancer popular science, this application designs a special fine-tuning task. Using a multi-modal large model (such as Transformer) as the basic model, the multi-modal large model is fine-tuned using breast cancer-related data so that the model can better understand and process breast cancer-related image and text data. During the fine-tuning process, multiple tasks are learned simultaneously, such as lesion recognition and popular science content generation. Through multi-task learning, the model can share underlying features and complement each other in high-level semantics, improving the generalization ability and accuracy of the model. Combine the fine-tuned model with the knowledge graph, and use the information in the knowledge graph for retrieval and enhancement. And through GraphRAG (Graph-based Retrieval-Augmented Generation) technology, multi-modal retrieval enhancement is achieved.

[0104] The fine-tuning task of this application decomposes the core requirements of breast cancer popular science into several sub-tasks through a multi-task learning architecture and conducts joint training. Specifically, it includes the following aspects:

[0105] Image lesion recognition:

[0106] Analyze breast images using a deep learning model (such as Faster R-CNN) to automatically detect and label suspected lesion areas. For example, the system can detect masses and microcalcifications in mammogram images and generate annotation information (ROI), providing a basis for subsequent diagnosis.

[0107] Interpretation of treatment path:

[0108] Provide popular science on personalized treatment plans based on image and pathological information: Combining image features and pathological information, the system can generate personalized treatment plan suggestions. For example, based on the size, location, and pathological type of the mass, the system can recommend further examinations or treatments to the user.

[0109] Psychological counseling and emotional support:

[0110] Provide popular science-based psychological comfort and nursing suggestions for users: Through natural language processing technology, the system can generate psychological comfort and nursing suggestions to help users relieve anxiety and fear. For example, the system can provide psychological support content on early symptoms, prevention measures, and treatment processes of breast cancer.

[0111] Through multi-task learning, share underlying features and complement each other in high-level semantics, thus achieving better results in breast cancer image and text processing, popular science interpretation, clinical plan explanation, etc. The specific steps are as follows:

[0112] Data preparation: Collect a large number of breast cancer image samples and text data, including mammogram, MRI, ultrasound and other imaging materials, as well as text data such as clinical guidelines, medical papers, and case reports.

[0113] Feature extraction: Use a deep learning model (such as Swin Transformer) to extract image features and extract text features through natural language processing technology.

[0114] Multi-task training: Jointly train sub-tasks such as image lesion recognition, treatment path interpretation, and psychological counseling and emotional support, share underlying features and complement each other in high-level semantics.

[0115] Model optimization: Optimize the model through a joint loss function, balance medical professionalism and popular science readability, and ensure that the generated popular science content is both professional and easy to understand.

[0116] Furthermore, during the training process, the following joint loss function can be defined To balance medical professionalism and popular science readability:

[0117]

[0118] Among them, It is the loss for medical professional tasks, used to measure the accuracy or cross-entropy loss of professional tasks such as image recognition and pathological stage prediction; It is the loss for popular science texts, used to measure the scores of popular science texts in readability and relevance; α is the weight coefficient of the loss for medical professional tasks; β is the weight coefficient of the loss for popular science texts; y k represents the true value label of the professional task; represents the probability predicted by the model; K represents the total number of categories in the medical professional task; k represents the category index, ranging from 1 to K; x i represents the word or sentence to be predicted in popular science generation; x context represents the context information when generating text; M represents the total number of words or sentences in the popular science text; P represents the probability of generating a certain word or sentence.

[0119] Joint loss function is the objective function that the model needs to minimize during the training process. It comprehensively considers two aspects:

[0120] Medical professionalism: Ensure the accuracy of the model in medical professional tasks.

[0121] Popular science readability: Ensure that the popular science text generated by the model is easy to understand and relevant.

[0122] By adjusting the values of α and β, the degree of emphasis on these two objectives by the model during the training process can be controlled. For example, if the value of α is larger, the model will pay more attention to the accuracy of medical professional tasks; if the value of β is larger, the model will pay more attention to the readability of popular science texts.

[0123] In some embodiments, the breast cancer disease popular science method based on the multi-modal knowledge graph provided by this application further includes safety review and confidence judgment, specifically including:

[0124] Confidence evaluation:

[0125] Conduct confidence evaluation on the generated popular science content, and decide whether further review or adjustment is needed according to the evaluation results. Confidence evaluation is achieved through the confidence scoring mechanism inside the model. For example, the reliability of the content is evaluated by using the probability distribution or confidence score output by the model.

[0126] Re-retrieval mechanism:

[0127] For content with a confidence level lower than the preset threshold, the "re-retrieval" mechanism is automatically triggered to re-retrieve relevant information from the multi-modal knowledge graph and external databases to improve the accuracy of the content. For example, the system can call the GraphRAG mechanism again to retrieve concept nodes and literature paragraphs related to the user input, and generate more accurate popular science content by combining the new information.

[0128] Prompt for referral to a professional doctor:

[0129] For content with a confidence level lower than the preset threshold, the system can also prompt the user to refer to a professional doctor to ensure that the user obtains the most accurate and reliable medical advice. For example, the system can generate the following prompt: "Based on the current information, it is recommended that you consult a professional doctor for more detailed diagnosis and treatment advice."

[0130] Review of sensitive topics and high-risk information:

[0131] For certain sensitive topics or high-risk information, configure medical expert review or automatic filtering strategies to ensure the accuracy and security of popular science information. For example, the system can automatically detect and filter out content containing sensitive information, such as information related to privacy or information that may cause panic.

[0132] In some embodiments, the breast cancer disease popular science method based on the multimodal knowledge graph provided by the present application further includes interactive deployment and scenario-based application, as follows:

[0133] 1. Visual interpretation of images

[0134] When the user uploads a breast image, the model first performs lesion detection and compares it with the nodes of the knowledge graph. If a similar example image is found, the suspicious area is highlighted on the user interface, and an explanation such as "high similarity to node XX in the graph / reference literature" is given. The system fuses the matched image and text and outputs it, enabling the user to intuitively see the markings of the suspicious features in the image while reading the text description. For example, when the user uploads a mammogram, the system detects an irregular mass, finds a similar example image in the knowledge graph, then highlights the area on the user interface and gives the following explanation: "This area has a high similarity to the 'irregular mass' node in the graph. Further examination is recommended."

[0135] 2. Multi-round conversation and dynamic update

[0136] The system retains the content of the previous few rounds of conversations and the retrieved knowledge nodes. If the user continues to ask questions such as "Is breast-conserving surgery suitable for this mass?" or "What should be noted during postoperative rehabilitation?", the model can call the GraphRAG mechanism to search for matching information in the graph and external literature again to maintain personalized and context-related responses. If the model does not have sufficient confidence in certain questions, a safety reminder will be popped up or a prompt such as "Refer to the diagnosis of a clinician" will be added to reduce the outflow of incorrect or misleading answers. For example, when the user continues to ask "Is breast-conserving surgery suitable for this mass?", the system calls the GraphRAG mechanism to retrieve the knowledge graph and external literature and generates the following content: "Based on the size and location of the mass, breast-conserving surgery is a viable option, but it needs to be determined in combination with the pathological examination results. It is recommended that you consult a professional doctor for more detailed diagnosis and treatment advice."

[0137] 3. Unification of Clinical and Popular Science Scenarios

[0138] When facing clinical assistance, more professional terms and literature citations are output; when facing public science popularization, more concise and highly readable text and image explanations are output, making it easy for users to understand the core points. For example, in the clinical assistance scenario, the system can output: "According to the BIRADS classification, this mass is initially evaluated as grade 4, and further pathological examination is recommended." In the public science popularization scenario, the system can output: "A suspicious mass is detected in the image. It is recommended that you have regular breast examinations and consult a doctor for more information."

[0139] The second aspect of this application provides a breast cancer disease popular science system based on a multi-modal knowledge graph, including: a data collection and preprocessing module, a feature extraction and mapping module, a multi-modal knowledge graph construction module, a retrieval and generation module, etc.

[0140] The data collection and preprocessing module is used to collect a large number of breast cancer image samples and text data. The image samples include, but are not limited to, mammography, MRI, ultrasound and other imaging materials, and the text data includes clinical guidelines, medical papers, case reports, diagnosis and treatment records, etc.

[0141] The data collection and preprocessing module is also used to preprocess the collected data, including denoising, normalization and contrast enhancement, to improve the data quality and usability. For text data, professional term alignment is performed and unified indexing is carried out with image metadata (such as stage, pathological type, etc.).

[0142] The feature extraction and mapping module is connected to the data collection and preprocessing module, and can use a deep learning model (such as Swin Transformer) to analyze the preprocessed breast cancer image samples, extract breast cancer-specific features, such as mass morphology, density and microcalcification features. And by combining the breast cancer-specific underlying features with the high-level attention mechanism, a comprehensive representation of local texture and global context is achieved.

[0143] The feature extraction and mapping module can also create corresponding nodes for the extracted breast cancer-specific features and establish mapping relationships with medical diagnosis information (such as BIRADS classification, HER2 status, TNM stage) to form a multi-modal unified representation.

[0144] The multi-modal knowledge graph construction module is connected to the feature extraction and mapping module, and is used to associate medical text entities in the text data with breast cancer-specific feature nodes to form a multi-modal knowledge graph;

[0145] The retrieval and generation module is connected to the multimodal knowledge graph construction module, and is used to obtain the information to be diagnosed input by the user, retrieve the associated nodes related to the information to be diagnosed from the multimodal knowledge graph, and output popular science content.

[0146] In some embodiments, the breast cancer disease popular science system based on the multimodal knowledge graph further includes a multi-task learning module, a joint loss function module, and a security review and confidence judgment module.

[0147] The multi-task learning module is connected to the feature extraction and mapping module and the multimodal knowledge graph construction module, and is used to jointly train the core requirements of breast cancer popular science through a multi-task learning architecture, including subtasks such as imaging lesion recognition, treatment path interpretation, and psychological counseling and emotional support;

[0148] The joint loss function module is connected to the multi-task learning module, and is used to define and optimize the joint loss function to balance medical professionalism and popular science readability;

[0149] The security review and confidence judgment module is connected to the retrieval and generation module, and is used to perform "re-retrieval" or "prompt referral to a professional doctor" on low-confidence outputs during the generation stage, and configure medical expert review or automatic filtering strategies to ensure the accuracy and security of popular science information.

[0150] The following is a specific process description of this application:

[0151] (1) Image preprocessing

[0152] The system first collects and preprocesses various breast cancer images and text data:

[0153] Denoising: Using techniques such as Gaussian filtering to remove noise interference in the image and make the image clearer.

[0154] Normalization: Adjust the image data to a unified intensity range to ensure comparability between different images.

[0155] Contrast enhancement: Enhance the contrast of the image through methods such as histogram equalization to highlight the lesion area.

[0156] For text data, it needs to be cleaned, aligned with professional terms, and unified indexed with image metadata. For example, associate medical terms in the text with information such as the location and size of lesions in the image for subsequent multimodal fusion.

[0157] Use the traditional segmentation or detection algorithm Faster R-CNN to perform preliminary lesion localization on the preprocessed breast cancer images, mark key information (ROI), and record it in the image metadata to lay the foundation for subsequent multimodal alignment.

[0158] For example, due to limitations in shooting conditions, an original mammogram may have a lot of noise and low contrast. After the above preprocessing steps, the lesion features such as masses and microcalcifications in the image will be marked with information and thus become more obvious, laying a foundation for subsequent feature extraction and analysis.

[0159] (2) Feature extraction

[0160] After preprocessing, the system uses a deep learning model (such as Swin Transformer) to extract features from the image:

[0161] Extraction of mass morphological features: Through a multi-scale window mechanism, capture the boundary, shape, and internal structure features of the mass to determine whether it is regular.

[0162] Extraction of density features: Use a multi-layer perceptron (MLP) and self-attention mechanism to extract the density distribution and uniformity features of the mass to distinguish different density types.

[0163] Extraction of microcalcification features: Through a high-resolution feature map, detect the size, shape, and distribution features of microcalcification foci to identify their morphology.

[0164] Combine the underlying features of the image with the high-level attention mechanism to form a comprehensive representation of local texture and global context.

[0165] For example, the system may extract that the morphological feature of a mass is irregular, the density feature is high density and uneven distribution, and the microcalcification feature is clustered distribution. These features will serve as important bases for subsequent analysis.

[0166] (3) Feature mapping

[0167] The system creates corresponding nodes for the extracted features in the knowledge graph and establishes a mapping relationship with medical diagnosis information:

[0168] Node creation: Create nodes for each feature, such as "irregular mass", "high-density mass", "microcalcification cluster", etc.

[0169] Relationship establishment: Associate the feature nodes with medical diagnosis information (such as BIRADS classification, HER2 status, TNM staging).

[0170] For example, connect the "irregular mass" node with the "BIRADS 4" node, indicating that this feature is related to a higher malignant risk. This mapping relationship helps the system understand the clinical significance of the features.

[0171] (4) Knowledge graph construction

[0172] The system associates medical entities in text data with imaging feature nodes to form a multi-modal knowledge graph:

[0173] Entity recognition and relationship extraction: Using natural language processing techniques, identify medical entities and their relationships from the text.

[0174] Imaging feature node creation: Create nodes for each imaging feature and establish a mapping with medical diagnosis information.

[0175] Multi-modal unified representation: Align imaging features and text information to form a unified representation.

[0176] For example, extract relationships such as "a certain pathological type is related to a certain imaging sign" or "a certain treatment plan is applicable to a certain stage" from clinical guidelines, and connect these relationships with imaging feature nodes to form a comprehensive knowledge graph.

[0177] (5) Retrieval

[0178] The system retrieves nodes related to the text or image input by the user from the multi-modal knowledge graph:

[0179] For the text or image input by the user, retrieve the most similar concept nodes and literature paragraphs in the knowledge graph, and combine the supplementary information from external databases to achieve multi-modal RAG enhancement.

[0180] For example, if the user uploads a mammogram, the system retrieves nodes such as "irregular mass" and "high-density mass" similar to it in the knowledge graph, and combines relevant literature to provide a basis for generating subsequent popular science content.

[0181] (6) Generation

[0182] The system generates personalized popular science content based on the retrieved nodes:

[0183] Content generation: Use the retrieved node information to generate popular science content containing images and text.

[0184] Human-computer interaction: Through multi-round conversations and dynamic updates, continuously adjust and refine the popular science content according to the user's subsequent questions and feedback.

[0185] For example, the system generates "An irregular mass with a high density and microcalcifications is detected in the image. It is recommended to further perform a pathological examination to confirm the diagnosis." The user continues to ask "Is this mass suitable for breast-conserving surgery?" The system calls the GraphRAG mechanism, retrieves the knowledge graph and external literature, and generates "According to the size and location of the mass, breast-conserving surgery is a feasible option, but it needs to be determined in combination with the pathological examination results. It is recommended that you consult a professional doctor for more detailed diagnosis and treatment advice."

[0186] (7) Safety Review and Confidence Judgment

[0187] The system evaluates the confidence of the generated popular science content:

[0188] Confidence calculation: Evaluate the reliability of the content according to the probability distribution or confidence score output by the model.

[0189] Threshold judgment: For content with a confidence level lower than the preset threshold, automatically trigger the "re - retrieval" or "prompt to refer to a professional doctor" mechanism to improve the accuracy of the content.

[0190] For example, if the confidence of the system in a certain diagnostic suggestion is lower than 80%, it will automatically trigger re - retrieval to retrieve relevant information from the multi - modal knowledge graph and external databases to improve the accuracy of the content. If sufficient confidence still cannot be achieved, the user will be prompted to refer to a professional doctor.

[0191] (8) Interactive Deployment and Scenario - based Application

[0192] The system provides users with visual image interpretation, supports multi - round conversations and dynamic updates, and realizes the unification of clinical and popular science scenarios:

[0193] Visual image interpretation: Highlight the suspicious areas on the user interface and give explanations so that users can intuitively see the lesion characteristics.

[0194] Multi - round conversations and dynamic updates: Retain the conversation history, call the GraphRAG mechanism according to the user's subsequent questions, and maintain the context relevance of the replies.

[0195] Unification of clinical and popular science scenarios: According to the user's needs, output professional terms and literature citations or simple and easy - to - understand popular science content.

[0196] For example, in the clinical assistance scenario, the system outputs "According to the BIRADS classification, this mass is initially evaluated as grade 4, and further pathological examination is recommended." In the public popular science scenario, the system outputs "A suspicious mass is detected in the image. It is recommended that you have regular breast examinations and consult a doctor for more information."

[0197] It should be noted that the technical solutions in each embodiment of this application can be combined with each other, but the basis for combination is that those skilled in the art can implement it; when the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist, that is, it does not belong to the protection scope of this application either.

[0198] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A breast cancer disease popular science method based on a multi-modal knowledge graph, characterized in that, Including: Collect and process breast cancer image samples and text data; Extract breast cancer specific features from breast cancer image samples; In the knowledge graph, create corresponding nodes for the breast cancer specific features and establish a mapping relationship with medical diagnosis information to form a multi-modal unified representation; Associate the text data with the breast cancer specific feature nodes to form a multi-modal knowledge graph; Obtain the information to be diagnosed input by the user, retrieve the associated nodes related to the information to be diagnosed from the multi-modal knowledge graph, and output popular science content.

2. The breast cancer disease popular science method based on the multi-modal knowledge graph according to claim 1, wherein Collecting and processing breast cancer image samples and text data specifically includes: Perform processing operations on the breast cancer image samples, and the processing operations at least include one of denoising, normalization, and contrast enhancement; Perform preliminary lesion localization and information annotation on the processed breast cancer image samples and record them in the metadata of the breast cancer image samples; Align the text data with professional terms and perform unified indexing on the aligned text data and the breast cancer image sample metadata.

3. The breast cancer disease popular science method based on the multi-modal knowledge graph according to claim 1, characterized in that Extracting breast cancer specific features from breast cancer image samples specifically includes: Use a deep learning model to analyze breast cancer image samples and extract features such as mass morphology, density, and microcalcifications; Combine the underlying image features with the high-level attention mechanism to form a comprehensive representation of local texture and global context.

4. The method for popularizing breast cancer disease based on a multi-modal knowledge graph according to claim 1, characterized in that, Introduce example diagrams for the breast cancer specific feature nodes, and when the information to be diagnosed input by the user is similar to the example diagrams, they will be displayed on the user interface for easy comparison.

5. The breast cancer disease popular science method based on a multi-modal knowledge graph according to claim 1, wherein, When obtaining the information to be diagnosed input by the user and retrieving the associated nodes related to the information to be diagnosed from the multi-modal knowledge graph, combine the supplementary information in the external database to achieve multi-modal RAG enhancement.

6. The method for popularizing breast cancer disease based on a multi-modal knowledge graph according to claim 5, characterized in that, Obtain the information to be diagnosed input by the user through a human-computer interaction dialogue, and dynamically update the information to be diagnosed input by the user according to the context in a multi-round dialogue to achieve continuous answering of personalized questions and popular science supplementation.

7. The breast cancer disease popular science method based on the multi-modal knowledge graph according to claim 1, characterized in that, It also includes designing fine-tuning tasks to improve the credibility and medical adaptability of breast cancer popular science content, and the fine-tuning tasks at least include: Decompose the core requirements of breast cancer popular science into several sub-tasks and conduct joint training, and the sub-tasks include: Automatically detect and label suspected morphology, density, and microcalcification features; Provide popular science on personalized treatment plans based on breast cancer images and pathological information; Provide popular science-based psychological comfort and nursing suggestions for users.

8. The method for popularizing breast cancer disease based on a multimodal knowledge graph according to claim 7, wherein During the training process, a joint loss function can be defined to balance medical professionalism and popular science readability; The combined loss function is calculated using the following formula: Among them, is the medical professional task loss, which is used to measure the accuracy rate or cross-entropy loss of the professional tasks of image recognition and pathological stage prediction; is the science popularization text loss, which is used to measure the scores of science popularization texts in readability and relevance; α is the weight coefficient of the medical professional task loss; β is the weight coefficient of the science popularization text loss; y k represents the true value label of the professional task; represents the probability predicted by the model; K represents the total number of categories in the medical professional task; k represents the category index, ranging from 1 to K; x i represents the word or sentence to be predicted in science popularization generation; x context represents the context information when generating text; M represents the total number of words or sentences in the science popularization text; P represents the probability of generating a certain word or sentence.

9. The breast cancer disease popular science method based on a multi-modal knowledge graph according to claim 1, wherein It also includes safety review and confidence judgment, specifically including: Evaluate the confidence of the generated popular science content, and decide whether further review or adjustment is needed according to the evaluation results; For content with a confidence level lower than the preset threshold, automatically trigger the "re-retrieval" or "prompt referral to a professional doctor" mechanism to improve the accuracy of the content; For certain sensitive topics or high-risk information, configure medical expert review or automatic filtering strategies to ensure the accuracy and safety of popular science information.

10. A breast cancer disease popular science system based on a multi-modal knowledge graph, characterized in that, Including: A data collection and preprocessing module for collecting and processing breast cancer image samples and text data; A feature extraction and mapping module, connected to the data collection and preprocessing module, is used to extract breast cancer-specific features from the processed breast cancer image samples, create corresponding nodes in the knowledge graph, and establish a mapping relationship with medical diagnosis information; A multi-modal knowledge graph construction module, connected to the feature extraction and mapping module, is used to associate medical text entities in the text data with breast cancer-specific feature nodes to form a multi-modal knowledge graph; and A retrieval and generation module, connected to the multi-modal knowledge graph construction module, is used to obtain the information to be diagnosed input by the user, retrieve related associated nodes from the multi-modal knowledge graph, and output popular science content.

11. The breast cancer disease popular science system based on the multi-modal knowledge graph according to claim 10, characterized in that, It further includes: A multi-task learning module, connected to the feature extraction and mapping module and the multi-modal knowledge graph construction module, is used to jointly train the core requirements of breast cancer popular science through a multi-task learning architecture, including sub-tasks such as imaging lesion recognition, treatment path interpretation, and psychological counseling and emotional support; A joint loss function module, connected to the multi-task learning module, is used to define and optimize the joint loss function to balance medical professionalism and popular science readability; and A security review and confidence judgment module, connected to the retrieval and generation module, is used to perform "re-retrieval" or "prompt referral to a professional doctor" on low-confidence outputs during the generation stage, and configure medical expert review or automatic filtering strategies to ensure the accuracy and security of popular science information.

Citation Information

Patent Citations

  • Method for establishing medical image atlas based on image segmentation and convolutional neural network (CNN)

    CN108389614A

  • Breast X-ray image classification model training method based on implicit apparent learning

    CN111415741A

  • Breast cancer diagnosis knowledge graph construction method and system based on ultrasonic examination report

    CN115101158A

  • Diabetes auxiliary diagnosis system, text processing method and map construction method

    CN116110570A

  • CT report generation processing method and device

    CN117995344A

Cited By

  • Intelligent environmental protection propaganda interaction system and method based on VI identification

    CN120510007A

  • Intelligent environmental protection propaganda interaction system and method based on VI recognition

    CN120510007B

  • Method and system for constructing reasoning agent for intelligent medical treatment guidance

    CN120525065A

  • A method and system for constructing a reasoning agent for intelligent medical guidance

    CN120525065B