Breast cancer axillary lymph node metastasis prediction method and device based on cross-modal data

Through cross-modal data prediction methods, combined with pathological images, imaging sequences and genomic features, the problem of preoperative evaluation of axillary lymph node metastasis of breast cancer was solved, non-invasive evaluation was achieved, and patient pain and medical expenses were reduced.

CN120525878BActive Publication Date: 2025-10-10SUN YAT SEN MEMORIAL HOSPITAL SUN YAT SEN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511017944.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-10
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

The existing technology lacks effective non-invasive methods to evaluate axillary lymph node metastasis in breast cancer patients before surgery, resulting in invasive examinations that increase patient pain and medical costs.

Method used

By combining pathological images, imaging sequences and genomic features, a cross-modal data prediction method is used to extract features of tumor cells and stromal regions, and semantic alignment and fusion are performed to predict the metastasis status of sentinel lymph nodes and non-sentinel lymph nodes.

Benefits of technology

It enables accurate preoperative assessment of axillary lymph node metastasis, provides a reliable basis for subsequent surgical planning and treatment decisions, reduces the necessity of invasive examinations, and reduces patient risks and medical costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525878B_ABST
    Figure CN120525878B_ABST
Patent Text Reader

Abstract

The application discloses a breast cancer axillary lymph node metastasis prediction method and device based on cross-modal data, which comprises the following steps: determining tumor cell regions and tumor interstitial regions in a pathological image according to a plurality of sub-images of the pathological image, and determining tumor omics features and interstitial omics features; obtaining effective sub-images according to the tumor cell regions and the tumor interstitial regions in the plurality of sub-images, and obtaining pathological features of the effective sub-images; determining axillary regions and primary lesion regions according to an image sequence of a detection object, and obtaining image features of the axillary regions and the primary lesion regions; mapping the pathological features and the image features to a unified semantic space through a semantic alignment model to obtain semantic features, and fusing the semantic features through an attention model to obtain first multi-modal features; extracting features from gene data of the detection object to obtain genomics features; and predicting a lymph node metastasis condition according to the first multi-modal features, the genomics features, the tumor omics features and the interstitial omics features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method and device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data. Background Art

[0002] Breast cancer is the malignant tumor with the highest incidence and tumor-related mortality rate among women worldwide. Breast cancer seriously threatens women's physical and mental health and places a heavy mental and financial burden on individuals and their families. The status of axillary lymph nodes in breast cancer patients is not only closely related to the development of surgical plans and overall treatment strategies, but is also an important indicator for assessing the prognosis of breast cancer patients. Axillary lymph node metastasis of breast cancer often indicates that tumor cells have spread from local invasion to the lymphatic system, increasing the risk of distant organ metastasis and reducing the survival rate of patients with axillary lymph node metastasis.

[0003] The sentinel lymph node (SLN) is the first site of metastasis from a primary tumor to the axilla. In routine clinical practice, axillary lymph node dissection (ALND) is often recommended if the SLN is positive. Sentinel lymph node biopsy (SLNB) has been incorporated into breast cancer diagnosis and treatment guidelines, improving the accuracy of axillary lymph node status assessment. SLNB and ALND are the two most important surgical procedures for evaluating axillary lymph node status. However, both procedures are invasive and can lead to various complications, such as edema of the affected upper limb and motor or sensory impairment. Although SLNB is less invasive than ALND and has a lower incidence of related complications, it still imposes a significant physical and mental burden on breast cancer patients. Furthermore, SLNB increases the duration of surgical anesthesia, increasing anesthetic risks and unnecessary medical costs. Therefore, seeking more accurate diagnosis and treatment strategies for axillary lymph node metastasis to optimize the "step-down" treatment model of the axilla (i.e., avoiding unnecessary SLNB or ALND as much as possible without increasing the patient's risk of axillary recurrence) can minimize the patient's medical risks and reduce medical costs.

[0004] SLNB is the standard surgical procedure for evaluating patients with clinically node-negative breast cancer. In traditional clinical practice, if SLNB confirms SLN metastasis, ALND is performed; if SLN metastasis is absent, ALND can be omitted. Multiple clinical studies have demonstrated the potential for ALND omission in patients with a low SLN burden (1–2 SLN metastases); however, patients with a high SLN burden (≥3 SLN metastases) still require ALND. Furthermore, only a subset of patients with a positive SLNB have non-sentinel lymph node (NSLN) metastases. Therefore, patients with a high SLN burden or NSLN metastases can proceed directly to ALND, thus reducing the waiting time for surgery. However, currently, there is no effective method for preoperative assessment of SLN and NSLN metastasis, other than pathological examination after SLNB and ALND. Therefore, a method for preoperative assessment of SLN and NSLN metastasis is urgently needed to provide a reliable basis for subsequent surgical planning and clinical treatment decisions. Summary of the Invention

[0005] The present application provides a method and apparatus for predicting axillary lymph node metastasis of breast cancer based on cross-modal data. By combining features of pathological images, image sequences, and genomics, a multi-dimensional comprehensive prediction is performed from macroscopic, microscopic, and genetic perspectives to accurately predict the metastatic status and number of sentinel lymph nodes of breast cancer, as well as the metastatic status of non-sentinel lymph nodes. This enables a non-invasive preoperative assessment of a patient's axillary lymph node metastasis, resolving the problem in the prior art of requiring invasive examinations to detect axillary lymph node metastasis. This provides a reliable basis for the formulation of subsequent surgical plans and the selection of subsequent clinical diagnosis and treatment decisions.

[0006] In a first aspect, the present application provides a method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data, comprising:

[0007] Dividing a pathological image of a detection object into a plurality of sub-images, determining a tumor cell region and a tumor interstitial region in the pathological image according to the plurality of sub-images, and determining tumoromic features and interstitial features according to the tumor cell region and the tumor interstitial region;

[0008] Acquire a valid sub-image containing tumor cells or tumor stroma from the multiple sub-images according to the tumor cell region and the tumor stroma region, and acquire pathological features of the valid sub-image;

[0009] determining a corresponding axillary region and a primary lesion region according to an image sequence of the detected object, and obtaining image features of the axillary region and the primary lesion region;

[0010] Mapping each pathological feature and each imaging feature to a unified semantic space through a semantic alignment model to obtain a plurality of semantic features, and fusing the plurality of semantic features through an attention model to obtain a first multimodal feature;

[0011] Extracting features from the genetic data of the test subject to obtain genomic features;

[0012] The metastasis status and number of sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes of the subject are predicted based on the first multimodal feature, the genomic feature, the tumoromic feature, and the interstitial feature.

[0013] In a second aspect, the present application provides a device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data, comprising:

[0014] a first feature acquisition module configured to divide a pathological image of a detection object into a plurality of sub-images, determine a tumor cell region and a tumor interstitial region in the pathological image according to the plurality of sub-images, and determine a tumoromic feature and an interstitial feature according to the tumor cell region and the tumor interstitial region;

[0015] a second feature acquisition module configured to acquire, from the plurality of sub-images, a valid sub-image containing tumor cells or tumor stroma based on the tumor cell region and the tumor stroma region, and acquire pathological features of the valid sub-image;

[0016] a third feature acquisition module configured to determine the corresponding axillary region and primary lesion region according to the image sequence of the detected object, and acquire image features of the axillary region and the primary lesion region;

[0017] a first feature fusion module configured to map each pathological feature and each imaging feature to a unified semantic space using a semantic alignment model to obtain a plurality of semantic features, and to fuse the plurality of semantic features using an attention model to obtain a first multimodal feature;

[0018] a fourth feature acquisition module, configured to extract features from the genetic data of the test subject to obtain genomic features;

[0019] A metastasis prediction module is configured to predict the metastasis status and number of metastases of the sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes of the detected subject based on the first multimodal feature, the genomic feature, the tumoromic feature, and the interstitial feature.

[0020] In a third aspect, the present application provides a device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data, comprising:

[0021] One or more processors; a storage device storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data as described in the first aspect.

[0022] In a fourth aspect, the present application provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data as described in the first aspect.

[0023] In this application, tumor cell areas and tumor interstitial areas are detected in multiple sub-images of the pathological image, and thus tumor omics features and interstitial omics features are determined based on the tumor cell areas and tumor interstitial areas. The tumor omics features and interstitial omics features reflect the tumor's invasive ability and the tumor's ability to break through the basement membrane and enter the lymphatic vessels, respectively. Effective sub-images are determined in the pathological image through the tumor cell area and the tumor interstitial area, and the pathological features of the microscopic tissue are mined using the effective sub-images. The image features of the macroscopic area are mined through the axillary area and the primary lesion area of ​​the image sequence. The pathological features of the microscopic tissue are semantically aligned and fused with the image features of the macroscopic area, obtaining the first multimodal feature that reflects the tumor and its microenvironment from multiple dimensions such as macro and micro, morphology and function. Feature extraction of genetic data is performed to obtain genomic features, which can reflect the invasiveness and metastatic potential of the tumor. The fusion of the first multimodal features, genomic features, tumor omics features and interstitial omics features provides more comprehensive and accurate feature information for the prediction of lymph node metastasis status. Combining these feature information can accurately predict the metastasis status of lymph nodes, and realize the preoperative assessment of the patient's axillary lymph node metastasis through non-invasive methods, providing a reliable basis for the formulation of subsequent surgical plans and the selection of subsequent clinical diagnosis and treatment decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a flowchart of a method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data provided in an embodiment of the present application;

[0025] Figure 2 This is a flowchart of determining the tumor cell region and the tumor stroma region based on the first segmentation model provided in an embodiment of the present application;

[0026] Figure 3 This is a flowchart of obtaining a valid sub-image containing tumor cells or tumor stroma and corresponding pathological features provided by an embodiment of the present application;

[0027] Figure 4 is a schematic diagram of the pathological image processing flow provided in an embodiment of the present application;

[0028] Figure 5 This is a flowchart of determining the axillary area and the primary lesion area and extracting corresponding image features provided by an embodiment of the present application;

[0029] Figure 6 is a schematic diagram of the image sequence processing flow provided by an embodiment of the present application;

[0030] Figure 7 This is a flow chart of the fusion of pathological features and imaging features provided in an embodiment of the present application;

[0031] Figure 8 This is a flow chart for predicting lymph node metastasis provided in an embodiment of the present application;

[0032] Figure 9 Schematic diagram of the processing flow of multi-omics features provided in the embodiments of the present application;

[0033] Figure 10 This is a flowchart of training a semantic alignment model, an attention model, and a multi-head classification model provided by an embodiment of the present application;

[0034] Figure 11 Schematic diagram of a device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data provided in an embodiment of the present application;

[0035] Figure 12 Schematic diagram of the structure of a device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] To further clarify the objectives, technical solutions, and advantages of this application, specific embodiments of the present application are described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are intended only to illustrate this application and are not intended to limit it. It should also be noted that, for ease of description, the drawings only illustrate portions relevant to this application, not all of them. Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the various operations (or steps) as sequential processes, many of the operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process may terminate upon completion of its operations, but may also include additional steps not shown in the accompanying drawings. The process may correspond to a method, function, procedure, subroutine, subprogram, or the like.

[0037] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0038] Axillary lymph node dissection (ALND) and sentinel lymph node biopsy (SLNB) are commonly performed to assess axillary lymph node metastasis. However, both SLNB and ALND are invasive procedures and can lead to complications such as upper extremity edema and motor or sensory impairment. Traditionally, if SLNB confirms the presence of SLN metastases, ALND is performed; if no SLN metastases are present, ALND can be omitted. Multiple clinical studies have demonstrated the potential for ALND omission in patients with a low SLN burden (1–2 SLN metastases); however, patients with a high SLN burden (≥3 SLN metastases) still require ALND. Furthermore, only a subset of patients with a positive SLNB have non-sentinel lymph node (NSLN) metastases. Therefore, patients with a high SLN burden or NSLN metastases can proceed directly to ALND, thus reducing the waiting time for surgery by avoiding SLNB. Therefore, it is important to seek more precise diagnostic and treatment strategies based on axillary lymph node metastasis to optimize the "step-down" management of the axilla. This approach can avoid unnecessary SLNB or ALND as much as possible without increasing the risk of axillary recurrence, thereby minimizing patient medical risks and reducing medical costs. However, currently, apart from pathological examination after SLNB and ALND, there is no effective method for preoperative assessment of SLN and NSLN metastasis.

[0039] To address the above issues, this embodiment provides a method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data. By combining the features of pathological images, image sequences, and genomics, a multi-dimensional comprehensive prediction is performed from macroscopic, microscopic, and genetic perspectives to accurately predict the metastatic status and number of sentinel lymph nodes of breast cancer, as well as the metastatic status of non-sentinel lymph nodes. This enables a non-invasive preoperative assessment of a patient's axillary lymph node metastasis, providing a reliable basis for the formulation of subsequent surgical plans and the selection of subsequent clinical diagnosis and treatment decisions.

[0040] The method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data provided in this embodiment can be performed by a device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data. The device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data can be implemented through software and / or hardware. The device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data can be composed of two or more physical entities, or it can be composed of a single physical entity. For example, the device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data can be an intelligent terminal with strong processing power, such as a computer or server. The server can be implemented as an independent server or a server cluster consisting of multiple servers.

[0041] The device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data is installed with at least one type of operating system. The device can install at least one application based on the operating system. The application can be a native application of the operating system or an application downloaded from a third-party device or server. In this embodiment, the device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data is installed with at least one application that can execute the method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data.

[0042] For ease of understanding, this embodiment is described by taking a server as an example of a subject that executes the method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data.

[0043] Figure 1 A flowchart of a method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data is provided in an embodiment of the present application. Figure 1 The method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data specifically includes:

[0044] S110 , dividing the pathological image of the detection object into multiple sub-images, determining the tumor cell region and the tumor interstitial region in the pathological image based on the multiple sub-images, and determining the tumoromic features and the interstitial features based on the tumor cell region and the tumor interstitial region.

[0045] The test subjects are patients undergoing preoperative evaluation for axillary lymph node metastasis of breast cancer. The pathological images are breast biopsy images of the test subjects. The acquisition process involves obtaining a breast biopsy tissue sample from the test subject, preparing pathological sections from the tissue sample, staining the sections with H&E (hematoxylin and eosin), and then scanning the stained sections with a digital pathology scanner to obtain a high-resolution image. The scanned image is used as the pathological image of the test subject.

[0046] For example, a pathology image can be divided into multiple sub-images of uniform size. The pathology image can be cropped using a sliding window of a preset size, with each resulting image being used as a sub-image of the pathology image. Once the sliding window has traversed the entire pathology image, the multiple sub-images corresponding to the pathology image can be obtained. The size and step size of the sliding window can be set based on actual conditions.

[0047] After dividing the pathological image into multiple sub-images, the sub-images can be used to identify tumor cells and tumor stroma through deep learning algorithms, determine the tumor cells and tumor stroma in each sub-image, and summarize the tumor cells and tumor stroma in each sub-image to obtain the tumor cell area and tumor stroma area in the pathological image.

[0048] It is understandable that tumor cells have significant differences from normal cells in terms of size, shape, nuclear-to-cytoplasmic ratio, and nuclear morphology. High nuclear-to-cytoplasmic ratios, irregular nuclear morphology, and prominent nucleoli are often characteristics of malignant tumor cells. These abnormal morphologies determine the invasive and metastatic capabilities of tumor cells. When tumor cells have a more aggressive morphology, the likelihood of axillary lymph node metastasis also increases. Therefore, this embodiment predicts axillary lymph node metastasis by obtaining characteristics such as the morphology, location, and spatial distribution of tumor cells. Furthermore, new blood vessels in the tumor stroma are important pathways for tumor cells to enter the blood circulation and metastasize to the axillary lymph nodes. The degree of proliferation of the fibrous and collagenous stroma in the tumor stroma and the degree of lymphocyte infiltration in the lymphoid stroma can reflect tumor cells. These stroma reflect how tumor cells break through the basement membrane and enter the lymphatic vessels. Therefore, this embodiment predicts axillary lymph node metastasis by obtaining characteristics such as the morphology, location, and spatial distribution of the tumor stroma. This embodiment comprehensively considers the multi-dimensional characteristic information of tumor cells and tumor stroma in pathological images to comprehensively analyze the biological behavior and metastatic potential of tumor cells from the perspective of pathological omics, which is conducive to improving the accuracy of predicting the axillary lymph node metastasis status of breast cancer.

[0049] When using deep learning algorithms to detect tumor cells and tumor stroma in sub-images, a pre-trained object detection model can be used to detect tumor cells and tumor stroma in the sub-images. Furthermore, a pre-trained segmentation model can be used to segment the tumor cells and tumor stroma in the sub-images. Subsequently, the tumor cell and tumor stroma regions in the pathology image are determined based on the tumor cells and tumor stroma in the sub-images.

[0050] In one embodiment, the tumor cell region and the tumor interstitial region can be segmented in the sub-image based on the first segmentation model, and the tumor cell region and the tumor interstitial region of the pathological image can be generated based on the tumor cell region and the tumor interstitial region of the multiple sub-images. The first segmentation model includes a first backbone network and a first segmentation head. Specifically, Figure 2 This is a flow chart of determining the tumor cell region and tumor stroma region based on the first segmentation model provided in the embodiment of the present application. Figure 2 As shown, the step of determining the tumor cell region and the tumor interstitial region based on the first segmentation model specifically includes S1101-S1103:

[0051] S1101. Extract pathological features of sub-images through a first backbone network.

[0052] Exemplarily, a sub-image is input into the first backbone network of the first segmentation model. The first backbone network extracts grayscale features, texture features, shape features, and geometric features of the sub-image as pathological features. The first backbone network can be a neural network capable of extracting deep image features from images, such as ResNet, VGGNet, DenseNet, or MobileNet. Alternatively, when the first segmentation model uses a U-net model, the first backbone network can be the encoder within the U-net model. This embodiment describes the pathological feature extraction process for the sub-image using the encoder within the U-net model as an example. Specifically, the sub-image is input into the encoder. The first convolutional block of the encoder extracts a preliminary feature map, which is then pooled to compress the spatial dimension of the feature map. After pooling, the feature map is passed to the next deeper convolutional block for further feature extraction, followed by another pooling operation. Each pooling operation doubles the number of channels in the feature map and halves the spatial dimension. The feature map output by the deepest layer of the encoder is input into the first segmentation head as the pathological feature of the sub-image.

[0053] S1102 : Segment the tumor cell region and / or tumor interstitial region in the sub-image based on the pathological features of the sub-image by using the first segmentation head.

[0054] Exemplarily, the first backbone network inputs the extracted pathological features into the first segmentation head of the first segmentation model. The first segmentation head performs semantic segmentation on the sub-image based on the corresponding input pathological features to output region masks for the tumor cell region and / or tumor interstitial region. The corresponding tumor cell region and / or tumor interstitial region are segmented in the sub-image based on the region masks. The first segmentation head can be a neural network, such as a fully convolutional network, a fully connected conditional random field, or a dilated convolution, that can identify pixel semantics based on image features to perform semantic segmentation. Alternatively, when the first segmentation model employs a U-net model, the first segmentation head can be a decoder in the U-net model. This embodiment describes the segmentation process of the tumor cell region and / or tumor interstitial region in the sub-image, using the decoder in the U-net model as an example. Specifically, the feature map output by the deepest layer of the encoder is input to the decoder. The decoder upsamples the feature map to restore the spatial dimension of the feature map. The upsampled feature map is then concatenated with the corresponding feature map of the same spatial dimension from the encoder via skip connections. The concatenated feature map is further fused and features are extracted using convolutional layers and activation functions. After multiple upsampling and feature fusion, the decoder generates a feature map of the same size as the sub-image. This feature map then passes through a 1×1 two-dimensional convolutional layer and a softmax function, outputting the probability distribution of each pixel belonging to the tumor cell region and the tumor stroma region. Pixels with a probability of belonging to the tumor cell region greater than a preset probability threshold are pooled to form a region mask for the tumor cell region, while pixels with a probability of belonging to the tumor stroma region greater than a preset probability threshold are pooled to form a region mask for the tumor stroma region.

[0055] It should be noted that the sub-image may only contain tumor cell areas or tumor interstitial areas, or of course there may be no tumor cell areas or tumor interstitial areas. When the sub-image does not contain tumor cell areas and tumor interstitial areas, the first segmentation head does not output a mask image or outputs a mask image of a non-tumor area.

[0056] S1103 : Generate a tumor cell region and a tumor interstitial region in a pathological image based on the tumor cell region and the tumor interstitial region in the multiple sub-images.

[0057] Exemplarily, after determining the tumor interstitial region and tumor cell region in each sub-image, the tumor interstitial region and tumor cell region in each sub-image are mapped to the pathological image according to the position of each sub-image in the pathological image to obtain the tumor interstitial region and tumor cell region in the pathological image.

[0058] In this embodiment, the tumor interstitial region includes a fibrous interstitial region, a collagenous interstitial region and a lymphatic interstitial region. The tumor cell region, the fibrous interstitial region, the collagenous interstitial region and the lymphatic interstitial region can be segmented in the sub-image through the first segmentation model, and the tumor cell region, the fibrous interstitial region, the collagenous interstitial region and the lymphatic interstitial region in each sub-image are mapped to the pathological image to obtain the tumor cell region, the fibrous interstitial region, the collagenous interstitial region and the lymphatic interstitial region in the pathological image.

[0059] It should be noted that the first segmentation model only has the ability to accurately segment tumor cell regions and tumor interstitial regions in sub-images after being trained with multiple sample images labeled with tumor cell regions, fibrous interstitial regions, collagenous interstitial regions, and lymphoid interstitial regions. During the training of the first segmentation model, the sample image is input into the first segmentation model, and the softmax function of the last layer of the first segmentation model outputs the probability distribution of each pixel in the sample image. Based on the probability distribution of each pixel and the true area of ​​each pixel labeled in the sample image, the cross-entropy loss value is calculated, and the model parameters in the first segmentation model are back-propagated based on the cross-entropy loss value to complete one training. When the number of training times reaches the upper limit or the model parameters of the first segmentation model reach the convergence condition, the first segmentation model is confirmed to have completed training, and the trained first segmentation model can be deployed on the server for real-time segmentation of tumor cell regions and tumor interstitial regions in sub-images.

[0060] This embodiment uses a first segmentation model to segment the sub-images into tumor cell regions, fibrous interstitial regions, collagenous interstitial regions, and lymphoid interstitial regions. This allows the tumor cell regions, fibrous interstitial regions, collagenous interstitial regions, and lymphoid interstitial regions in each sub-image to be used to generate the tumor cell regions, fibrous interstitial regions, collagenous interstitial regions, and lymphoid interstitial regions in the pathology image. This avoids performing semantic segmentation directly on the larger pathology image, thereby effectively improving segmentation accuracy and efficiency. Furthermore, this embodiment segments the fibrous interstitial regions, collagenous interstitial regions, and lymphoid interstitial regions in the pathology image, allowing for the subsequent acquisition of richer interstitial features from different types of tumor interstitial regions, thereby more accurately assessing the migration and invasion capabilities of tumor cells and improving the accuracy of predicting axillary lymph node metastasis status in breast cancer.

[0061] After segmenting the tumor cell regions in the pathology image, the morphological and texture features of the corresponding tumor cells are extracted from the tumor cell regions. Morphological features can describe the shape, size, boundary, area, perimeter, shape complexity (such as Hawking Index and Fractal Dimension), and irregularity (such as circularity, aspect ratio, and ConvexHull Ratio) of the tumor cells. Texture features can be calculated using gray-level co-occurrence matrix (GLCM) and gray-level run-length matrix (GLRLM). After segmenting the fibrous, collagenous, and lymphoid regions in the pathology image, morphological and texture features of each tumor stroma region are extracted based on the fibrous, collagenous, and lymphoid regions. Furthermore, spatial relationship characteristics between tumor cells and stroma, such as the tumor-to-stroma ratio and the ratio of tumor cells to stroma, can be determined based on the tumor cell and tumor stroma regions. The morphological and texture characteristics of tumor cells, as well as the spatial relationship characteristics between tumor cells and tumor stroma, are used as tumoromic features, while the morphological and texture characteristics of the tumor stroma, as well as the spatial relationship characteristics between tumor cells and tumor stroma, are used as stromamic features. This example quantifies the shape, texture, and spatial distribution of tumor cells and tumor stroma in pathological images to extract objective, accurate, and highly repeatable tumor cell and tumor stroma features for predicting tumor status.

[0062] It is important to note that when lymph node metastasis is present, the arrangement of collagen fibers in the fibrous stroma changes from a regular pattern to a disordered distribution, and the density increases. The collagen content in tumor tissue is higher than in normal tissue and is unevenly distributed. When collagen accumulates in large quantities at the tumor margin, it may form a barrier, limiting the local spread of tumor cells, but it may also promote metastasis to distant lymph nodes. An increase in the number and diameter of lymphatic vessels within the tumor stroma often indicates a tumor with a high potential for lymph node metastasis. The morphological and textural features of the fibrous, collagenous, and lymph node stroma regions can all reflect and assess the migration and invasion capabilities of tumor cells. Extracting the morphological and textural features of the corresponding fibrous, collagenous, and lymph node stroma regions in pathological images can improve the accuracy of predicting axillary lymph node metastasis in breast cancer.

[0063] S120 , obtaining a valid sub-image containing tumor cells or tumor stroma from the multiple sub-images according to the tumor cell region and the tumor stroma region, and obtaining pathological features of the valid sub-image.

[0064] The pathological features of a valid sub-image are the grayscale, texture, shape, and geometric features extracted from the sub-image. Sub-images containing tumor cells or tumor stroma are only useful for predicting axillary lymph node metastasis if and only if the sub-image contains tumor cells or tumor stroma. Therefore, sub-images containing tumor cells or tumor stroma are considered valid sub-images.

[0065] The effective sub-image can intuitively reflect characteristics such as tumor cell size and density, as well as tumor stroma composition, content, and distribution at the histological level. Therefore, the pathological features of the effective sub-image can be used to assess lymph node metastasis from a microscopic tissue dimension. Specifically, the pathological features in the effective sub-image are microscopic tissue features extracted from a microscopic perspective, such as cells or stroma. This embodiment aims to semantically align and fuse the microscopic tissue features extracted from the effective sub-image with the macroscopic tumor features extracted from the image sequence, thereby obtaining multimodal features that reflect the tumor and its microenvironment from multiple dimensions, including macroscopic and microscopic, morphological, and functional.

[0066] Optionally, after valid sub-images are determined, they can be fed into a pre-trained feature extraction network to obtain corresponding pathological features. The feature extraction network can be a convolutional neural network. This feature extraction network can be trained together with a subsequent classification network used to predict metastasis status.

[0067] In addition, the features extracted from the effective sub-image by the first backbone network in the first segmentation model can be directly used as the pathological features of the effective sub-image. Figure 3 This is a flow chart of obtaining effective sub-images and corresponding pathological features provided by an embodiment of the present application. Figure 3 As shown, the step of obtaining the effective sub-image and the corresponding pathological features specifically includes S1201-S1202:

[0068] S1201 : When a sub-image includes a tumor cell region and / or a tumor stroma region, determine the sub-image as a valid sub-image.

[0069] Exemplarily, when the first segmentation head segments the tumor cell region and / or the tumor stroma region in the sub-image, it indicates that the sub-image contains image information of the tumor cells and / or the tumor stroma, and thus the sub-image is determined as a valid sub-image.

[0070] S1202: Use the pathological features of the sub-image extracted by the first backbone network as the pathological features of the corresponding valid sub-image.

[0071] Exemplarily, after determining that the sub-image is a valid sub-image, the pathological features output by the first backbone network are used as the pathological features of the valid sub-image.

[0072] In this embodiment, the pathological features extracted from the effective sub-image by the first backbone network are used as the pathological features of the effective sub-image. There is no need to retrain the corresponding feature extraction network to extract features from the effective sub-image, which saves the model construction cost and improves the feature acquisition efficiency.

[0073] In order to more intuitively understand the pathological characteristics of the effective sub-images and the acquisition process of tumor cell characteristics and tumor stroma characteristics, this embodiment uses Figure 4 The schematic diagram of the pathological image processing flow is described as follows. Figure 4 As shown, a pathology image is slid through a sliding window to obtain multiple sub-images. The sub-images are then input into a pre-trained first segmentation model. The first backbone network of the first segmentation model outputs the pathological features of the sub-images, and the first segmentation head outputs the tumor cell region, fibrous stroma region, collagenous stroma region, and lymphoid stroma region of each sub-image. Based on the tumor cell region, fibrous stroma region, collagenous stroma region, and lymphoid stroma region of each sub-image, the tumor cell region, fibrous stroma region, collagenous stroma region, and lymphoid stroma region in the pathology image are generated. Tumoromics features and stromaomics features are determined based on the tumor cell region, fibrous stroma region, collagenous stroma region, and lymphoid stroma region in the pathology image. Based on the segmented regions of each sub-image, the pathological features of valid sub-images containing tumor cell regions or tumor stroma regions are screened from the pathological features of the multiple sub-images. The first segmentation model is trained using sub-images corresponding to the segmentation of multiple sample pathology images labeled with tumor cell regions, fibrous stroma regions, collagenous stroma regions, and lymphoid stroma regions.

[0074] S130 : Determine the corresponding axillary region and primary lesion region according to the image sequence of the detected object, and obtain image features of the axillary region and the primary lesion region.

[0075] The image sequence is composed of multiple breast MRI images. Breast MRI images are medical images generated using magnetic resonance imaging (MRI) for breast examination. For example, a strong magnetic field and radio waves can be applied to the patient's breast and axilla to generate breast MRI images. Breast MRI images can simultaneously perform multi-parameter and multi-directional imaging of the primary breast cancer lesion and the axilla, producing a variety of images. For example, T1-weighted imaging (T1W) and T2-weighted imaging (T2W) can be generated using T1-weighted or T2-weighted imaging, diffusion-weighted imaging (DWI) can be generated using diffusion MRI, and dynamic contrast enhancement (DCE) can be generated using dynamic contrast enhancement. T2W imaging not only demonstrates the morphological characteristics of breast tumors but also assesses tumor composition and invasion characteristics through signal changes within and surrounding tissues. DCE can obtain enhanced images at multiple time phases, before, during, and after contrast agent injection. By monitoring the influx, diffusion, and clearance of contrast agent within lesions, parameters related to tissue microcirculatory function can be obtained, reflecting physiological functions such as microcirculatory status in breast cancer. Diffusion refers to the random movement of water molecules within tissues due to thermal energy (Brownian motion). However, obstacles such as cell membranes, organelles, and biomacromolecules affect the free movement of water molecules, resulting in different diffusion patterns between different tissues. Diffusion MRI, which does not require contrast agent injection, is currently the only technique capable of non-invasively examining the diffusion of water molecules within microscopic tissue structures. Specific motion-sensitive MRI sequences, such as DWI, can highlight this diffusion contrast. Based on this, T1W, T2W, DWI, and DCE can be used as image sequences. Preprocessing of these sequences, such as image denoising, can be performed, and spatial and intensity alignment of the individual images within the sequence can be performed to improve image quality.

[0076] After preprocessing the image sequence, a deep learning algorithm is used to identify the primary lesion and axillary regions in each image in the sequence, thereby obtaining corresponding image features from the primary lesion and axillary regions. These image features include first-order statistical features, morphological features, and texture features of the primary lesion and axillary regions.

[0077] Optionally, a target detection model or a segmentation model can be used to determine the corresponding primary lesion area and axillary area in each image. In the case of a large image size, direct detection or segmentation is difficult. The image can be reduced in size and then input into the target detection model or segmentation model for detection or segmentation to obtain the primary lesion area and axillary area, and the primary lesion area and axillary area are mapped to the unreduced image to obtain the primary lesion area and axillary area in the image. In addition, the image can be divided into multiple sub-images, and each sub-image is input into the target detection model or segmentation model for detection or segmentation to obtain the primary lesion area and axillary area in the sub-image. The corresponding primary lesion area and axillary area are mapped into the image according to the position of the sub-image in the image to obtain the primary lesion area and axillary area in the image.

[0078] It should be noted that the axillary region may include the left axillary region and the right axillary region. By obtaining the characteristics of the left axillary region and the right axillary region and performing difference comparison, the metastasis of the lymph nodes can be further accurately determined.

[0079] After acquiring the primary lesion and axillary regions in the image, a pre-trained feature extraction network can be used to extract image features from the primary lesion and axillary regions. This feature extraction network can be trained together with the classification network used to predict metastasis outcomes.

[0080] In addition, features output by the backbone network in the model used to detect or segment the primary lesion region and the axillary region in the image can also be used as image features. For example, when the second segmentation model is used to segment the axillary region and the primary lesion region in each image of the image sequence, the second segmentation model includes a second backbone network and a second segmentation head, and the second backbone network can be used to extract corresponding image features in the axillary region and the primary lesion region. Specifically, Figure 5 This is a flow chart of determining the axillary region and the primary lesion region and extracting corresponding image features provided by the embodiment of the present application. Figure 5 As shown, the steps of determining the axillary region and the primary lesion region and extracting corresponding image features specifically include S1301-S1303:

[0081] S1301. Extract image features of the image sequence through the second backbone network.

[0082] Exemplarily, the image images in the image sequence are input into the second backbone network of the second segmentation model, and the second backbone network extracts the first-order statistical features, texture features, morphological features, etc. of the image images as image features. Similarly, the second backbone network can be a neural network such as ResNet, VGGNet, DenseNet, MobileNet, etc. that has the function of extracting deep image features from images; or, when the second segmentation model adopts a U-net model, the second backbone network can be an encoder in the U-net model. The specific process of using the encoder as the second backbone network to extract image features can refer to the specific process of using the encoder as the first backbone network to extract the pathological features of the sub-image in step S1101, which will not be repeated here.

[0083] S1302 : Segment the axillary region and the primary lesion region in the image sequence based on image features using a second segmentation head.

[0084] Exemplarily, the second backbone network inputs the extracted pathological features into the second segmentation head of the second segmentation model. The second segmentation head performs semantic segmentation on the sub-image based on the corresponding input image features to output a regional mask of the primary lesion region and a mask of the axillary region in the image. The primary lesion region in the image is retained based on the regional mask of the primary lesion region and the pixel values ​​of the remaining pixels are cleared. The axillary region in the image is retained based on the mask of the axillary region and the pixel values ​​of the remaining pixels are cleared.

[0085] Similarly, the second segmentation head can be a neural network such as a fully convolutional network, a fully connected conditional random field, and a dilated convolution, which can recognize pixel semantics based on image features to perform semantic segmentation. Alternatively, when the second segmentation model adopts a U-net model, the second segmentation head can be a decoder in the U-net model. The specific process of using the decoder as the second segmentation head to segment the axillary region and the primary lesion region can refer to the process of using the decoder as the first segmentation head to segment the tumor cell region and / or tumor stroma region in step S1102, and will not be repeated here.

[0086] S1303. Extract image features of the axillary region and the primary lesion region through the second backbone network.

[0087] For example, an image containing only the primary lesion region is input into the second backbone network, and the first-order statistical features, texture features, and morphological features extracted from the primary lesion region by the second backbone network are used as the image features of the primary lesion region. An image containing only the axilla region is input into the second backbone network, and the first-order statistical features, texture features, and morphological features extracted from the axilla region by the second backbone network are used as the image features of the axilla region.

[0088] It can be understood that after steps S1301-S1303, the image features of the primary lesion region and the axillary region in T1W, T2W, DWI, and DCE can be extracted.

[0089] Optionally, in the process of training the second segmentation model, a sample image sequence of a breast cancer patient can be obtained, and the primary lesion region and the axillary region of each image in the sample image sequence are labeled to obtain the label information of the sample image sequence. The image in the sample image sequence is input into the second segmentation model, and the softmax function of the last layer of the second segmentation model outputs the probability distribution of each pixel point in the image. According to the probability distribution of each pixel point and the true region of each pixel point labeled in the image, the cross-entropy loss value is calculated, and the model parameters in the second segmentation model are optimized according to the cross-entropy loss value, so as to complete one training. When the number of training reaches the upper limit or the model parameters of the second segmentation model reach the convergence condition, it is confirmed that the training of the second segmentation model is completed, and the trained second segmentation model can be deployed on the server for real-time segmentation of the axillary region and the primary lesion region of the image sequence.

[0090] The second backbone network extracts the image features of the axillary region and the primary lesion region, so that the corresponding feature extraction network does not need to be retrained for feature extraction of the axillary region and the primary lesion region, thereby saving the model building cost and improving the feature acquisition efficiency.

[0091] In order to more intuitively understand the acquisition process of the image features of the axillary region and the primary lesion region, the embodiment is described with reference to the schematic diagram of the processing flow of the image sequence shown in Figure 6 As shown in Figure 6 , the T1W image, the T2W image, the DWI image, and the DCE image in the image sequence are obtained, and the T1W image, the T2W image, the DWI image, and the DCE image are preprocessed and then input into the second segmentation model trained in advance. The second segmentation model segments the axillary region and the primary lesion region in the T1W image, the T2W image, the DWI image, and the DCE image. The axillary region and the primary lesion region of the T1W image, the T2W image, the DWI image, and the DCE image are input into the second backbone network of the second segmentation model, respectively, to obtain the image features of the T1W axillary region, the T1W primary lesion region, the T2W axillary region, the T2W primary lesion region, the DWI axillary region, the DWI primary lesion region, the DCE axillary region, and the DCE primary lesion region. The second segmentation model is trained by the sample image sequence labeled with the primary lesion region and the axillary region.

[0092] S140. Mapping each pathological feature and each imaging feature to a unified semantic space through a semantic alignment model to obtain multiple semantic features, and fusing the multiple semantic features through an attention model to obtain a first multimodal feature.

[0093] The unified semantic space is a high-dimensional vector space, and the semantic features in the unified semantic space are features that express the metastasis status of sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes. In this embodiment, the semantic alignment model can convert image features and pathological features into features that can express the metastasis status of sentinel lymph nodes and non-sentinel lymph nodes, thereby achieving semantic unification of image features and pathological features. This allows for the subsequent cross-modal fusion of the semantically aligned pathological features and image features, breaking down the modal barriers between pathological and imaging data, so that the fused features can reflect the metastasis status of breast cancer axillary lymph node cells from both a microscopic tissue perspective and a macroscopic tumor perspective.

[0094] Afterwards, the attention model is used to fuse the pathological features and imaging features of each semantic space to obtain the first multimodal feature. The first multimodal feature fuses semantically unified pathological features and imaging features, which can reflect the metastasis of breast cancer axillary lymph node cells from the perspective of microscopic tissue and macroscopic tumor.

[0095] Optionally, the semantic alignment model includes a pathological feature encoder and an image feature encoder, which can be used to map the pathological features and image features to a unified semantic space to obtain corresponding semantic features. The attention model is then used to determine the corresponding feature weights of multiple semantic features in the unified semantic space and perform weighted fusion to highlight the most critical semantic features of the metastasis prediction task, so that the subsequent classification network used to predict lymph node metastasis can focus on the key semantic features for prediction, thereby improving the accuracy of metastasis prediction. Specifically, Figure 7 This is a flow chart of the fusion of pathological features and imaging features provided in the embodiment of this application. Figure 7 As shown, the steps of fusing pathological features and imaging features specifically include S1401-S1403:

[0096] S1401 : Encode the pathological feature of the effective sub-image by a pathological feature encoder to obtain a first semantic feature.

[0097] The pathology feature encoder, along with the image feature encoder, attention model, and subsequent classification network, are trained using a multi-instance learning algorithm. During training, the pathology feature encoder gradually learns the mapping relationship between pathology features and the expression characteristics of sentinel lymph node metastasis status and non-sentinel lymph node metastasis status. After training, the pathology feature encoder encodes the pathology features based on the learned mapping relationship to obtain a first semantic feature that describes the metastasis status of both sentinel lymph nodes and non-sentinel lymph nodes.

[0098] Optionally, the pathological feature encoder can be ResNet, which extracts the expression of pathological features regarding the metastasis status of sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes through a convolution layer, a batch normalization layer, an activation function, and a pooling layer, and finally outputs the first semantic feature.

[0099] S1402: Encode the image features of the axillary region or the primary lesion region using an image feature encoder to obtain a second semantic feature.

[0100] Similarly, the image feature encoder, pathology feature encoder, attention model, and subsequent classification network are trained using a multi-instance learning algorithm. During training, the image feature encoder gradually learns the mapping relationship between image features and the expression characteristics of sentinel lymph node metastasis status and non-sentinel lymph node metastasis status. After training is complete, the image feature encoder encodes the image features based on the learned mapping relationship to obtain a second semantic feature that can describe the metastasis status of both sentinel lymph nodes and non-sentinel lymph nodes.

[0101] Optionally, the image feature encoder may also be a ResNet, which extracts the expression of image features regarding the metastasis status of sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes through a convolutional layer, a batch normalization layer, an activation function, and a pooling layer, and finally outputs a second semantic feature.

[0102] S1403. Determine the feature weight of each first semantic feature and each second semantic feature through the attention model, and perform weighted fusion of each first semantic feature and each second semantic feature according to the feature weight to obtain a first multimodal feature.

[0103] For example, after mapping each imaging feature and pathological feature to a unified semantic space to obtain the corresponding first and second semantic features, the first and second semantic features become semantically comparable. The attention model can then be used to compare the criticality of each first and second semantic feature to the metastatic status of sentinel lymph nodes and non-sentinel lymph nodes, thereby determining the feature weights of each first and second semantic feature. Each first and second semantic feature is then multiplied by the corresponding feature weight to adjust the proportion of the first and second semantic features in the first multimodal feature. The first multimodal feature is then obtained by summing the products of each first and second semantic feature with the corresponding feature weight. Similarly, the attention model is trained with the pathological feature encoder, the imaging feature encoder, and the subsequent classification network using a multi-instance learning algorithm. During training, the attention model gradually learns the importance of the semantic features converted from the pathological features and the imaging features to the metastatic status of sentinel lymph nodes and non-sentinel lymph nodes, thereby giving more attention to key semantic features and improving the discriminability and effectiveness of the semantic features.

[0104] Optionally, the attention model can be an MLP, which inputs multiple first semantic features and multiple second semantic features into the input layer of the MLP, and the input layer transmits the received multiple semantic features to the hidden layer. The hidden layer processes the multiple semantic features and inputs them into the output layer. The output layer predicts a weight for each input semantic feature and uses the softmax function to convert the weight into a probability distribution so that the sum of the weights of each feature is one.

[0105] This embodiment uses a pathology feature encoder and an image feature encoder to map pathology features and image features into a unified semantic space for semantic alignment to obtain corresponding semantic features, thus avoiding semantic misalignment that affects the accuracy of transfer prediction. The attention model focuses on key semantic features to improve the accuracy of transfer prediction.

[0106] S150: Extract features from the genetic data of the test object to obtain genomic features.

[0107] Genomic data refers to data containing genetic information obtained using various molecular biology techniques, such as quantitative polymerase chain reaction (qPCR), single-cell RNA sequencing (scRNA-seq), and next-generation sequencing (NGS). Genomics can reveal the inherent heterogeneity of breast cancer and screen for potential target genes that are more sensitive and correlated with axillary lymph node metastasis. For example, single-cell RNA sequencing can sequence the transcriptome at the single-cell level, thereby exploring gene expression within individual cells. Single-cell RNA sequencing can distinguish different subpopulations of tumor cells, some of which highly express genes associated with cell migration and invasion. It can also identify cell signaling pathways that play a key role in lymph node metastasis and discover potential metastasis marker genes, providing deep insight into the heterogeneous characteristics of tumor cells. Therefore, this embodiment obtains genomic features from genetic data and integrates multiple combinations of features from imaging genomics, pathogenomics, and genomics to more comprehensively reflect information about the tumor microenvironment, thereby accurately predicting lymph node metastasis status.

[0108] For example, after obtaining the gene test results using various molecular biological techniques, the gene test results are converted into structured gene data, and key indicators are extracted from the structured gene data to obtain genomic features.

[0109] S160. Predicting the metastasis status and number of sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes of the subject based on the first multimodal feature, the genomic feature, the tumoromic feature, and the interstitial feature.

[0110] For example, a multi-task network can be used to fuse the first multimodal features, genomic features, tumoromic features, and stromal features, and the metastatic status and number of sentinel lymph nodes, as well as the metastatic status of non-sentinel lymph nodes, can be predicted based on the fused features. The metastatic status is categorized as either metastatic or non-metastatic, and the number of metastases is categorized as greater than or equal to 3 or less than 3.

[0111] Optionally, a multi-head classification model can be used as a multi-task network. The multi-head classification network includes three classification heads, which can predict the metastasis status and number of metastases of sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes respectively. Specifically, Figure 8 This is a flow chart for predicting lymph node metastasis provided in the embodiment of the present application. Figure 8 As shown, the steps of predicting lymph node metastasis specifically include S1601-S1602:

[0112] S1601. Concatenate the first multimodal feature, the genomic feature, the tumoromic feature, and the interstitial omics feature to obtain a second multimodal feature.

[0113] Exemplarily, the first multi-modal feature output by the attention model, the genomics features extracted from the gene data based on the key indicators, and the tumor and interstitial features extracted from the tumor cell region and the tumor interstitial region are spliced to obtain a second multi-modal feature. Among them, the first multi-modal feature can reflect the tumor and its microenvironment from multiple dimensions such as macro and micro, morphology and function, the genomics features can reflect the tumor invasiveness and metastatic potential, and the tumor and interstitial features reflect the tumor invasion ability and the ability of the tumor to break through the basement membrane into the lymphatic vessel. The second multi-modal feature that integrates the first multi-modal feature, the genomics feature, the tumor feature and the interstitial feature provides more comprehensive and accurate feature information for the metastasis state prediction of the lymph node, and improves the prediction accuracy of the subsequent multi-head classification model.

[0114] S1602, input the second multi-modal feature into the pre-trained multi-head classification model, and respectively predict the metastasis state of the sentinel lymph node, the metastasis number of the sentinel lymph node, and the metastasis state of the non-sentinel lymph node through each classification head of the multi-head classification model.

[0115] Exemplarily, the multi-head classification model includes a third backbone network, a first classification head, a second classification head and a third classification head, the third backbone network is used for convolution on the input second multi-modal feature, the first classification head is used for predicting the metastasis state of the sentinel lymph node, the second classification head is used for predicting the metastasis data of the sentinel lymph node, and the third classification head is used for predicting the metastasis state of the non-sentinel lymph node. It can be understood that the first classification head, the second classification head and the third classification head share a backbone network. The second multi-modal feature is input into the third backbone network of the multi-head classification model, and the feature vector output by the third backbone network after convolution on the second multi-modal feature is input into the first classification head, the second classification head and the third classification head respectively, to obtain the metastasis state of the sentinel lymph node output by the first classification head, the metastasis number of the sentinel lymph node output by the second classification head, and the metastasis state of the non-sentinel lymph node output by the third classification head.

[0116] Optionally, the first classification head, the second classification head, and the third classification head can be fully connected layers. Specifically, the input layer of the first classification head receives the feature vector output by the third backbone network, the hidden layer processes the feature vector and inputs it into the output layer, and the output layer compresses the output value between 0 and 1 through sigmoid to represent the predicted probability of the metastasis status of the sentinel lymph node being metastasis and non-metastasis. When the predicted probability of metastasis is greater than 0.5, the metastasis status of the sentinel lymph node is metastasis, and when the predicted probability of non-metastasis is greater than 0.5, the metastasis status of the sentinel lymph node is non-metastasis. The input layer of the second classification head receives the feature vector output by the third backbone network, the hidden layer processes the feature vector and inputs it into the output layer, and the output layer compresses the output value between 0 and 1 through sigmoid to represent the predicted probability of the number of metastases in the sentinel lymph node being less than 3 and not less than 3. When the predicted probability of less than 3 metastases is greater than 0.5, the number of metastases in the sentinel lymph node is less than 3. When the predicted probability of not less than 3 metastases is greater than 0.5, the number of metastases in the sentinel lymph node is not less than 3. The input layer of the third classification head receives the feature vector output by the third backbone network. The hidden layer processes the feature vector and then inputs it into the output layer. The output layer uses a sigmoid function to compress the output value between 0 and 1 to represent the predicted probability of non-metastasis status in non-sentinel lymph nodes. When the predicted probability of metastasis is greater than 0.5, the non-sentinel lymph node is considered metastatic. When the predicted probability of non-metastasis is greater than 0.5, the non-sentinel lymph node is considered non-metastatic.

[0117] This example predicts lymph node metastasis by integrating secondary modality features from radiomics, pathomic analysis, and genomics, highlighting the limitations of single-modality prediction and improving accuracy. A multi-headed classification model simultaneously predicts the metastatic status and number of sentinel lymph nodes, as well as the metastatic status of non-sentinel lymph nodes. This enables feature data sharing to improve prediction efficiency and differentiates the metastatic risk of different lymph nodes to provide a more detailed picture of lymph node metastasis, providing a reliable basis for subsequent surgical planning and clinical diagnosis and treatment decisions.

[0118] In order to more intuitively understand the prediction process of lymph node metastasis, this embodiment uses Figure 9 The schematic diagram of the multi-omics feature processing flow is described. Figure 9As shown, the pathological features of the effective sub-image are input into the pathological feature encoder, and the image features of the axillary region and the primary lesion region are input into the image feature encoder, thereby mapping the pathological features and image features to a unified semantic space and aligning them to obtain corresponding first semantic features and second semantic features. Each first semantic feature and each second semantic feature are input into the attention model, which performs weighted fusion on each first semantic feature and each second semantic feature to output a first multimodal feature. The first multimodal feature, genomic features, tumor omics features, and interstitial omics features are concatenated to obtain a second multimodal feature. The second multimodal feature is input into the third backbone network of the multi-head classification model, which performs convolution on the second multimodal feature. The feature vector obtained by the convolution is input into the first, second, and third classification heads. The first classification head predicts whether the subject has sentinel lymph node metastasis, the second classification head predicts whether the number of sentinel lymph node metastases is less than 3, and the third classification head predicts whether the subject has non-sentinel lymph node metastasis.

[0119] From the above content, it can be seen that the first segmentation model and the second segmentation model can be trained independently, and the semantic alignment model, attention model and multi-head classification model can be regarded as an end-to-end network model, so as to train the semantic alignment model, attention model and multi-head classification model as a whole. For example, Figure 10 This is a flowchart of the training semantic alignment model, attention model and multi-head classification model provided by the embodiment of this application. Figure 10 As shown, the steps of training the semantic alignment model, the attention model, and the multi-head classification model specifically include S210-S250:

[0120] S210 , obtaining a sample pathological image and a sample image sequence of the sample object, obtaining pathological sample features of the corresponding effective sub-image based on the sample pathological image, and obtaining image sample features of the corresponding axillary region and primary lesion region based on the sample image sequence.

[0121] The sample subjects are patients with confirmed lymph node metastasis or patients without lymph node metastasis, the sample pathology images are breast puncture pathology images of the sample subjects, and the sample image sequence is an image sequence of magnetic resonance imaging of the sample subjects.

[0122] A sample pathology image can be imaged through a preset sliding window to obtain multiple sub-images, and the tumor cell area and / or tumor interstitial area in the sub-image can be segmented through a pre-trained first segmentation model. The sub-image containing the tumor cell area or the tumor interstitial area is determined as a valid sub-image, and the features extracted by the first backbone network of the first segmentation model in the valid sub-image are used as the pathological sample features of the valid sub-image.

[0123] The axillary region and the primary lesion region of each image in the sample image sequence are segmented by a pre-trained second segmentation model, and the image sample features of the axillary region and the primary lesion region are extracted using the second backbone network of the second segmentation model.

[0124] S220 , converting each pathological sample feature and each image sample feature into a corresponding semantic sample feature according to the semantic alignment model, and performing weighted fusion on each semantic sample feature through the attention model to obtain a first multimodal sample feature.

[0125] For example, a pathological feature encoder is used to encode the pathological sample features to obtain corresponding first semantic sample features, and an image feature encoder is used to encode the image sample features to obtain corresponding second semantic sample features. The attention model is used to perform weighted fusion of each first semantic sample feature and each second semantic sample feature to obtain the first multimodal sample feature.

[0126] S230 , obtaining sample gene data of the sample object, determining corresponding genomic sample features based on the sample gene data, and obtaining tumor genomic sample features and interstitial genomic sample features according to the tumor cell region and tumor interstitial region of the sample pathology image.

[0127] Exemplarily, genomic sample features are extracted from the sample genetic data based on key indicators. The tumor cell region and tumor interstitial region in each sub-image of the sample pathology image are mapped to the sample pathology image to obtain the tumor cell region and tumor interstitial region in the sample pathology image. Tumoromics sample features and interstitial sample features are extracted based on the tumor cell region and tumor interstitial region in the sample pathology image.

[0128] S240: splicing the first multimodal sample features, genomic sample features, tumor omics sample features, and interstitial omics sample features, and inputting the features into a multi-head classification model to obtain a sample prediction result output by the multi-head classification model.

[0129] Exemplarily, the first multimodal sample features, genomic sample features, tumor omics sample features and interstitial omics sample features are spliced ​​together, and the spliced ​​features are input into the third backbone network of the multi-head classification model. The third backbone network convolves the input features to obtain feature vectors, and the feature vectors are respectively input into the first classification head, the second classification head and the third classification head. The first classification head, the second classification head and the third classification head predict based on the feature vectors and output the metastasis status of the sentinel lymph nodes, the number of metastases of the sentinel lymph nodes and the metastasis status of the non-sentinel lymph nodes, respectively. The output results of the first classification head, the second classification head and the third classification head are used as sample prediction results.

[0130] S250. Back-propagate the sample prediction results and the label information of the sample object to optimize the multi-head classification model, attention model and semantic alignment model.

[0131] The label information of the sample object includes the metastasis status of the sample object's sentinel lymph node, the number of metastases in the sentinel lymph node, and the metastasis status of non-sentinel lymph nodes. A first loss value is calculated based on the output of the first classification head and the metastasis status of the sentinel lymph node in the label information. A second loss value is calculated based on the output of the second classification head and the number of metastases in the sentinel lymph node in the label information. A third loss value is calculated based on the output of the third classification head and the metastasis status of the non-sentinel lymph nodes in the label information. A total loss value is calculated based on the first, second, and third loss values. Backpropagation is performed based on the total loss value to optimize the network parameters of the first, second, and third classification heads, the third backbone network, the attention model, the pathology feature encoder, and the image feature encoder, thereby completing one end-to-end model training. When the end-to-end model training reaches an upper limit or the model parameters reach convergence, the end-to-end model training is confirmed to be complete. The trained end-to-end model can be deployed on a server to work with the first and second segmentation models to predict axillary lymph node metastasis in breast cancer.

[0132] In summary, the embodiment of the present application provides a method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data, which detects tumor cell areas and tumor interstitial areas in multiple sub-images of the pathological image, thereby determining tumor omics features and interstitial omics features based on the tumor cell areas and tumor interstitial areas. The tumor omics features and interstitial omics features reflect the tumor's invasive ability and the tumor's ability to break through the basement membrane and enter the lymphatic vessels, respectively. Effective sub-images are determined in the pathological image through the tumor cell area and the tumor interstitial area, and the pathological features of the microscopic tissue are mined using the effective sub-images. The image features of the macroscopic area are mined through the axillary area and the primary lesion area of ​​the image sequence. The pathological features of the microscopic tissue are semantically aligned and fused with the image features of the macroscopic area, and the first multimodal features reflecting the tumor and its microenvironment from multiple dimensions such as macro and micro, morphology and function are obtained. Genomic features are obtained by feature extraction of genetic data, and genomic features can reflect the invasiveness and metastasis potential of the tumor. The fusion of the first multimodal features, genomic features, tumor omics features and interstitial omics features provides more comprehensive and accurate feature information for the prediction of lymph node metastasis status. Combining these feature information can accurately predict the metastasis status of lymph nodes, and realize the preoperative assessment of the patient's axillary lymph node metastasis through non-invasive methods, providing a reliable basis for the formulation of subsequent surgical plans and the selection of subsequent clinical diagnosis and treatment decisions.

[0133] Based on the above embodiments, Figure 11This is a schematic diagram of the structure of a device for predicting breast cancer axillary lymph node metastasis based on cross-modal data provided in an embodiment of the present application. Figure 11 The breast cancer axillary lymph node metastasis prediction device based on cross-modal data provided in this embodiment specifically includes: a first feature acquisition module 31, a second feature acquisition module 32, a third feature acquisition module 33, a first feature fusion module 34, a fourth feature acquisition module 35 and a metastasis prediction module 36.

[0134] The first feature acquisition module 31 is configured to divide the pathological image of the detection object into multiple sub-images, determine the tumor cell region and the tumor interstitial region in the pathological image based on the multiple sub-images, and determine the tumoromic features and the interstitial features based on the tumor cell region and the tumor interstitial region;

[0135] The second feature acquisition module 32 is configured to acquire, from the plurality of sub-images, effective sub-images containing tumor cells or tumor stroma based on the tumor cell region and the tumor stroma region, and acquire pathological features of the effective sub-images;

[0136] The third feature acquisition module 33 is configured to determine the corresponding axillary region and primary lesion region according to the image sequence of the detected object, and acquire image features of the axillary region and the primary lesion region;

[0137] A first feature fusion module 34 is configured to map each pathological feature and each imaging feature to a unified semantic space using a semantic alignment model to obtain a plurality of semantic features, and to fuse the plurality of semantic features using an attention model to obtain a first multimodal feature;

[0138] The fourth feature acquisition module 35 is configured to extract features from the genetic data of the test subject to obtain genomic features;

[0139] The metastasis prediction module 36 is configured to predict the metastasis status and number of metastases of the sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes of the subject based on the first multimodal feature, the genomic feature, the tumoromic feature, and the stromal feature.

[0140] On the basis of the above embodiment, the tumor interstitial region includes a fibrous interstitial region, a collagen interstitial region and a lymphoid interstitial region, and the tumor cell region and the tumor interstitial region are determined by a pre-trained first segmentation model, and the first segmentation model includes a first backbone network and a first segmentation head; accordingly, the first feature acquisition module 31 includes: a pathological feature extraction unit, configured to extract the pathological features of the sub-image through the first backbone network; a first region segmentation unit, configured to segment the tumor cell region and / or tumor interstitial region in the sub-image based on the pathological features of the sub-image through the first segmentation head; and a first region generation unit, configured to generate the tumor cell region and tumor interstitial region in the pathological image based on the tumor cell region and tumor interstitial region in multiple sub-images.

[0141] Based on the above embodiment, the second feature acquisition module 32 includes: an effective sub-image determination unit, configured to determine the sub-image as a valid sub-image when the sub-image includes a tumor cell region and / or a tumor interstitial region; and a pathological feature determination unit, configured to use the pathological features of the sub-image extracted by the first backbone network as the pathological features of the corresponding effective sub-image.

[0142] Based on the above embodiment, the axillary region includes a left axillary region and a right axillary region, and the axillary region and the primary lesion region are determined by a pre-trained second segmentation model, and the second segmentation model includes a second backbone network and a second segmentation head; accordingly, the third feature acquisition module 33 includes: a first image feature extraction unit, configured to extract image features of the image sequence through the second backbone network; a second region segmentation unit, configured to segment the axillary region and the primary lesion region in the image sequence based on the image features through the second segmentation head; and a second image feature extraction unit, configured to extract image features of the axillary region and the primary lesion region through the second backbone network.

[0143] Based on the above embodiment, the semantic alignment model includes a pathological feature encoder and an image feature encoder; accordingly, the first feature fusion module 34 includes: a first encoding unit, configured to encode the pathological features of the effective sub-image through the pathological feature encoder to obtain a first semantic feature; a second encoding unit, configured to encode the image features of the axillary area or the primary lesion area through the image feature encoder to obtain a second semantic feature; a weighted fusion unit, configured to determine the feature weights of each first semantic feature and each second semantic feature through an attention model, and weightedly fuse each first semantic feature and each second semantic feature according to the feature weights to obtain a first multimodal feature.

[0144] On the basis of the above-mentioned embodiments, the metastasis prediction module 36 comprises: a feature splicing unit configured to splice the first multi-modal feature, the genomic feature, the tumoromic feature and the stromomic feature to obtain a second multi-modal feature; and a metastasis prediction unit configured to input the second multi-modal feature into a pre-trained multi-head classification model, and predict the metastasis state of the sentinel lymph node, the metastasis number of the sentinel lymph node and the metastasis state of the non-sentinel lymph node of the detection object through each classification head of the multi-head classification model.

[0145] On the basis of the above-mentioned embodiments, the breast cancer axillary lymph node metastasis device based on cross-modal data further comprises a model training module, and the model training module comprises: a first sample feature acquisition unit configured to acquire a sample pathological image and a sample image sequence of a sample object, acquire a pathological sample feature of a corresponding effective sub-image based on the sample pathological image, and acquire an image sample feature of a corresponding axillary region and primary lesion region based on the sample image sequence; a sample feature fusion unit configured to convert each pathological sample feature and each image sample feature into a corresponding semantic sample feature according to a semantic alignment model, and obtain a first multi-modal sample feature by weighted fusion of each semantic sample feature through an attention model; a second sample feature acquisition unit configured to acquire sample gene data of the sample object, determine a corresponding genomic sample feature based on the sample gene data, and acquire a tumoromic sample feature and a stromomic sample feature according to a tumor cell region and a tumor stromal region of the sample pathological image; a sample prediction unit configured to input the first multi-modal sample feature, the genomic sample feature, the tumoromic sample feature and the stromomic sample feature after splicing into the multi-head classification model to obtain a sample prediction result output by the multi-head classification model; and a model optimization unit configured to perform back propagation on the sample prediction result and label information of the sample object to optimize the multi-head classification model, the attention model and the semantic alignment model.

[0146] As mentioned above, the embodiment of the present application provides a breast cancer axillary lymph node metastasis prediction device based on cross-modal data, which detects tumor cell areas and tumor interstitial areas in multiple sub-images of the pathological image, thereby determining tumor omics features and interstitial omics features based on the tumor cell areas and tumor interstitial areas. The tumor omics features and interstitial omics features reflect the tumor's invasive ability and the tumor's ability to break through the basement membrane and enter the lymphatic vessels, respectively. Effective sub-images are determined in the pathological image through the tumor cell area and the tumor interstitial area, and the pathological features of the microscopic tissue are mined using the effective sub-images. The image features of the macroscopic area are mined through the axillary area and the primary lesion area of ​​the image sequence. The pathological features of the microscopic tissue are semantically aligned and fused with the image features of the macroscopic area, obtaining the first multimodal features that reflect the tumor and its microenvironment from multiple dimensions such as macro and micro, morphology and function. Feature extraction of genetic data is performed to obtain genomic features, which can reflect the invasiveness and metastasis potential of the tumor. The fusion of the first multimodal features, genomic features, tumor omics features and interstitial omics features provides more comprehensive and accurate feature information for the prediction of lymph node metastasis status. Combining these feature information can accurately predict the metastasis status of lymph nodes, and realize the preoperative assessment of the patient's axillary lymph node metastasis through non-invasive methods, providing a reliable basis for the formulation of subsequent surgical plans and the selection of subsequent clinical diagnosis and treatment decisions.

[0147] The apparatus for predicting axillary lymph node metastasis of breast cancer based on cross-modal data provided in the embodiments of the present application can be used to execute the method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data provided in the above embodiments, and has corresponding functions and beneficial effects.

[0148] Figure 12 This is a schematic diagram of the structure of a device for predicting breast cancer axillary lymph node metastasis based on cross-modal data provided by an embodiment of the present application, with reference to Figure 12 The device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data includes: a processor 41, a memory 42, a communication device 43, an input device 44, and an output device 45. The number of processors 41 in the device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data can be one or more, and the number of memories 42 in the device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data can be one or more. The processor 41, memory 42, communication device 43, input device 44, and output device 45 of the device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data can be connected via a bus or other means.

[0149] Memory 42, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data in any embodiment of the present application (e.g., first feature acquisition module 31, second feature acquisition module 32, third feature acquisition module 33, first feature fusion module 34, fourth feature acquisition module 35, and metastasis prediction module 34). Memory 42 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on device usage. Furthermore, memory 42 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory may further include memory remotely located from the processor, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0150] The communication device 43 is used for data transmission.

[0151] The processor 41 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 42, that is, implements the above-mentioned method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data.

[0152] The input device 44 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 45 may include a display device such as a display screen.

[0153] The above-mentioned device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data can be used to execute the method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data provided in the above-mentioned embodiment, and has corresponding functions and beneficial effects.

[0154] The embodiment of the present application also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to execute a method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data. The method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data comprises: dividing a pathological image of a detected object into multiple sub-images, determining a tumor cell region and a tumor interstitial region in the pathological image based on the multiple sub-images, determining tumor omics features and interstitial omics features based on the tumor cell region and the tumor interstitial region; obtaining a valid sub-image containing tumor cells or tumor interstitial in the multiple sub-images based on the tumor cell region and the tumor interstitial region; image, obtain the pathological features of the effective sub-image; determine the corresponding axillary area and primary lesion area according to the image sequence of the detected object, and obtain the image features of the axillary area and the primary lesion area; map each pathological feature and each image feature to a unified semantic space through a semantic alignment model to obtain multiple semantic features, and fuse multiple semantic features through an attention model to obtain a first multimodal feature; extract features from the genetic data of the detected object to obtain genomic features; predict the metastasis status and number of metastases of the detected object's sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes based on the first multimodal feature, genomic features, tumor omics features and interstitial omics features.

[0155] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media, such as CD-ROMs, floppy disks, or tape drives; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. Storage media may also include other types of memory or combinations thereof. In addition, the storage medium may be located in the first computer system in which the program is executed, or it may be located in a different second computer system that is connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). The storage medium may store program instructions (e.g., embodied as a computer program) that can be executed by one or more processors.

[0156] Of course, the storage medium containing computer-executable instructions provided in an embodiment of the present application is not limited to the above-mentioned method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data, and can also execute related operations in the method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data provided in any embodiment of the present application.

[0157] The breast cancer axillary lymph node metastasis prediction device based on cross-modal data, the breast cancer axillary lymph node metastasis prediction system based on cross-modal data, the storage medium, and the breast cancer axillary lymph node metastasis prediction equipment based on cross-modal data provided in the above embodiments can perform the breast cancer axillary lymph node metastasis prediction method based on cross-modal data provided by any embodiment of the present application. Technical details not described in detail in the above embodiments can be referred to the breast cancer axillary lymph node metastasis prediction method based on cross-modal data provided by any embodiment of the present application.

[0158] The above is only the preferred embodiment of the present application and the technical principle used. The present application is not limited to the specific embodiments herein. Various obvious changes, readjustments and replacements that can be made by those skilled in the art will not deviate from the protection scope of the present application. Therefore, although the present application is described in more detail through the above embodiments, the present application is not limited to the above embodiments. More other equivalent embodiments can be included without deviating from the concept of the present application, and the scope of the present application is determined by the scope of the claims.

Claims

1. A method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data, characterized in that: include: Dividing a pathological image of a detection object into a plurality of sub-images, determining a tumor cell region and a tumor interstitial region in the pathological image based on the plurality of sub-images, and determining tumoromic features and interstitial features based on the tumor cell region and the tumor interstitial region, wherein the tumor interstitial region includes a fibrous interstitial region, a collagenous interstitial region, and a lymphoid interstitial region, the tumoromic features include morphological features and texture features of tumor cells and spatial relationship features between tumor cells and tumor interstitial regions, and the interstitial features include morphological features and texture features of the tumor interstitial region and spatial relationship features between tumor cells and tumor interstitial regions; Acquire a valid sub-image containing tumor cells or tumor stroma from the multiple sub-images according to the tumor cell region and the tumor stroma region, and acquire pathological features of the valid sub-image; determining a corresponding axillary region and a primary lesion region according to an image sequence of the detected object, and obtaining image features of the axillary region and the primary lesion region; Mapping each pathological feature and each imaging feature to a unified semantic space through a semantic alignment model to obtain a plurality of semantic features, and fusing the plurality of semantic features through an attention model to obtain a first multimodal feature; Extracting features from the genetic data of the test subject to obtain genomic features; The metastasis status and number of sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes of the subject are predicted based on the first multimodal feature, the genomic feature, the tumoromic feature, and the interstitial feature.

2. The method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data according to claim 1, characterized in that: The tumor cell region and the tumor stroma region are determined by a pre-trained first segmentation model, wherein the first segmentation model includes a first backbone network and a first segmentation head; Accordingly, determining the tumor cell region and the tumor interstitial region in the pathological image according to the multiple sub-images includes: extracting pathological features of the sub-image through the first backbone network; Segmenting, by the first segmentation head, a tumor cell region and / or a tumor interstitial region in the sub-image based on the pathological features of the sub-image; The tumor cell region and the tumor interstitial region in the pathological image are generated according to the tumor cell region and the tumor interstitial region in the plurality of sub-images.

3. The method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data according to claim 2, wherein: The step of obtaining a valid sub-image containing tumor cells or tumor stroma from the plurality of sub-images based on the tumor cell region and the tumor stroma region, and obtaining a pathological feature of the valid sub-image includes: In a case where the sub-image includes the tumor cell region and / or the tumor stroma region, determining the sub-image as a valid sub-image; The pathological features of the sub-image extracted by the first backbone network are used as the pathological features of the corresponding valid sub-image.

4. The method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data according to claim 1, wherein: The axillary region includes a left axillary region and a right axillary region, the axillary region and the primary lesion region are determined by a pre-trained second segmentation model, and the second segmentation model includes a second backbone network and a second segmentation head; Accordingly, determining the corresponding axillary region and primary lesion region according to the image sequence of the detected object and obtaining the image features of the axillary region and the primary lesion region include: extracting image features of the image sequence through the second backbone network; Segmenting the axillary region and the primary lesion region in the image sequence based on the image features by the second segmentation head; The image features of the axillary region and the primary lesion region are extracted through the second backbone network.

5. The method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data according to claim 1, wherein: The semantic alignment model includes a pathological feature encoder and an image feature encoder; Accordingly, the semantic alignment model is used to map each pathological feature and each imaging feature to a unified semantic space to obtain multiple semantic features, and the attention model is used to fuse the multiple semantic features to obtain a first multimodal feature, including: Encoding the pathological feature of the valid sub-image by the pathological feature encoder to obtain a first semantic feature; Encoding the image features of the axillary region or the primary lesion region by the image feature encoder to obtain a second semantic feature; The feature weight of each first semantic feature and each second semantic feature is determined through an attention model, and the first semantic features and the second semantic features are weightedly fused according to the feature weight to obtain a first multimodal feature.

6. The method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data according to claim 1, wherein: The predicting of the metastasis status and number of sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes of the subject based on the first multimodal feature, the genomic feature, the tumoromic feature, and the interstitial omics feature includes: splicing the first multimodal feature, the genomic feature, the tumoromic feature, and the interstitial feature to obtain a second multimodal feature; The second multimodal features are input into a pre-trained multi-head classification model, and the metastasis status of the sentinel lymph nodes, the number of metastases in the sentinel lymph nodes, and the metastasis status of non-sentinel lymph nodes of the detection object are predicted respectively by each classification head of the multi-head classification model.

7. The method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data according to claim 6, characterized in that: The training steps of the semantic alignment model, the attention model and the multi-head classification model include: Acquire a sample pathological image and a sample image sequence of a sample object, acquire pathological sample features of a corresponding effective sub-image based on the sample pathological image, and acquire image sample features of a corresponding axillary region and primary lesion region based on the sample image sequence; Converting each pathological sample feature and each image sample feature into a corresponding semantic sample feature according to the semantic alignment model, and performing weighted fusion of each semantic sample feature through the attention model to obtain a first multimodal sample feature; Obtaining sample genetic data of the sample object, determining corresponding genomic sample features based on the sample genetic data, and obtaining tumoromic sample features and interstitial sample features according to the tumor cell region and tumor interstitial region of the sample pathological image; splicing the first multimodal sample feature, the genomic sample feature, the tumoromics sample feature, and the interstitial omics sample feature, and inputting the combined features into the multi-head classification model to obtain a sample prediction result output by the multi-head classification model; The sample prediction results and the label information of the sample objects are back-propagated to optimize the multi-head classification model, the attention model and the semantic alignment model.

8. A device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data, characterized in that: include: a first feature acquisition module configured to divide a pathological image of a detection object into a plurality of sub-images, determine a tumor cell region and a tumor interstitial region in the pathological image based on the plurality of sub-images, and determine a tumoromic feature and an interstitial feature based on the tumor cell region and the tumor interstitial region, wherein the tumor interstitial region includes a fibrous interstitial region, a collagenous interstitial region, and a lymphoid interstitial region, wherein the tumoromic feature includes morphological and textural features of tumor cells and spatial relationship features between tumor cells and tumor interstitial regions, and wherein the interstitial feature includes morphological and textural features of the tumor interstitial region and spatial relationship features between tumor cells and tumor interstitial regions; and a second feature acquisition module configured to acquire, from the plurality of sub-images, a valid sub-image containing tumor cells or tumor stroma based on the tumor cell region and the tumor stroma region, and acquire pathological features of the valid sub-image; a third feature acquisition module configured to determine the corresponding axillary region and primary lesion region according to the image sequence of the detected object, and acquire image features of the axillary region and the primary lesion region; A first feature fusion module is configured to map each pathological feature and each imaging feature to a unified semantic space through a semantic alignment model to obtain a plurality of semantic features, and fuse the plurality of semantic features through an attention model to obtain a first multimodal feature; a fourth feature acquisition module, configured to extract features from the genetic data of the test subject to obtain genomic features; A metastasis prediction module is configured to predict the metastasis status and number of metastases of the sentinel lymph nodes and the metastasis status of non-sentinel lymph nodes of the detected subject based on the first multimodal feature, the genomic feature, the tumor omics feature and the interstitial omics feature.

9. A device for predicting axillary lymph node metastasis of breast cancer based on cross-modal data, characterized in that: include: one or more processors; A storage device stores one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data as described in any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that: When executed by a computer processor, the computer executable instructions are used to perform the method for predicting axillary lymph node metastasis of breast cancer based on cross-modal data according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent lung cancer metastasis prediction system based on GCAVE-GAN and multi-mode fusion

    CN118628462A

  • Thyroid nodule benign and malignant classification method based on clinical information, radiomics and gene detection

    CN120356530A