Bladder Cancer Lymph Node Metastasis Prediction Method, Device, Equipment and Storage Medium

By combining the characteristic fusion of urine cell sections and nuclear magnetic resonance images, the prediction of bladder cancer lymph node metastasis is solved, and a more accurate lymph node metastasis detection is achieved, reducing the rate of misdiagnosis.

CN120013952BActive Publication Date: 2025-07-22SUN YAT SEN MEMORIAL HOSPITAL SUN YAT SEN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510507958.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-22
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In the prior art, the accuracy of detecting lymph node metastasis by single radiology is low, and misdiagnosis is prone to missed diagnosis or misdiagnosis, resulting in insufficient treatment or excessive treatment of patients.

Method used

By combining the lesion cell characteristics of urine cell slice images and the lesion area characteristics of the MRI image, the mutual attention model is used for feature fusion, and a pre-trained classification network is input to predict the metastasis results of bladder cancer lymph nodes.

Benefits of technology

It improves the accuracy of lymph node metastasis detection of bladder cancer, reduces the probability of misdiagnosis or misdiagnosis, and avoids the situation where patients are under-treated or over-treated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013952B_ABST
    Figure CN120013952B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and storage medium for predicting lymph node metastasis of bladder cancer. In the present application, a plurality of sliding window images corresponding to urine cell section images are obtained, and a plurality of cross-sectional images corresponding to nuclear magnetic resonance images are obtained; a plurality of image pairs are obtained by traversing and combining the plurality of sliding window images and the plurality of cross-sectional images, and each image pair includes a sliding window image and a cross-sectional image; the lesion cell features of the sliding window image in the image pair are extracted, the lesion area features of the cross-sectional image in the image pair are extracted, and the lesion cell features and the lesion area features are fused through a mutual attention model to obtain fused features; the fused features of each image pair are spliced and then input into a pre-trained classification network to obtain the metastasis result of the lymph nodes of bladder cancer output by the classification network. By the above technical means, the problem of low accuracy in detecting lymph node metastasis by single radiology in the prior art is solved, and the probability of misdiagnosis or missed diagnosis is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a method, device, equipment and storage medium for predicting bladder cancer lymph node metastasis. Background Art

[0002] Worldwide, bladder cancer is a common cancer, and accurate lymph node metastasis is crucial for guiding the treatment and prognosis evaluation of bladder cancer patients. Approximately 25% of patients with muscle-invasive bladder cancer have lymph node metastasis, which can lead to poor prognosis. For patients with lymph node metastasis, more aggressive treatment methods are usually required, including preoperative neoadjuvant treatment, intraoperative extended pelvic lymph node dissection, etc., which can improve overall survival and disease-free survival.

[0003] In the prior art, radiology (such as CT and MRI) is usually used to detect preoperative lymph node metastasis of bladder cancer. However, CT and MRI mainly detect malignant lymph nodes based on the size of the lymph nodes, with low accuracy, and are prone to missed diagnosis or misdiagnosis problems, resulting in under-treatment or over-treatment of patients. Summary of the Invention

[0004] This application provides a method, device, equipment and storage medium for predicting bladder cancer lymph node metastasis, which jointly predicts the metastasis result of bladder cancer lymph nodes through the lesion cell characteristics of urine cytological section images and the lesion area characteristics of nuclear magnetic resonance images, solves the problem of low accuracy in detecting lymph node metastasis by a single radiology in the prior art, reduces the probability of misdiagnosis or missed diagnosis, and avoids the problem of under-treatment or over-treatment of patients.

[0005] In the first aspect, this application provides a method for predicting bladder cancer lymph node metastasis, including:

[0006] Obtain a plurality of sliding window images corresponding to the urine cytological section image, and obtain a plurality of cross-sectional images corresponding to the nuclear magnetic resonance image;

[0007] Traverse and combine the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, and each image pair includes one of the sliding window images and one of the cross-sectional images;

[0008] Extract the lesion cell characteristics of the sliding window image in the image pair, extract the lesion area characteristics of the cross-sectional image in the image pair, and fuse the lesion cell characteristics and the lesion area characteristics through a mutual attention model to obtain a fused feature;

[0009] After splicing the fused features of each image pair, input them into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph node output by the classification network.

[0010] In a second aspect, the present application provides a device for predicting bladder cancer lymph node metastasis, including:

[0011] An image acquisition module configured to acquire a plurality of sliding window images corresponding to a urine cell section image and a plurality of cross-sectional images corresponding to a nuclear magnetic resonance image;

[0012] An image combination module configured to traverse and combine the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, each of the image pairs including one of the sliding window images and one of the cross-sectional images;

[0013] A feature fusion module configured to extract lesion cell features of the sliding window image in the image pair, extract lesion area features of the cross-sectional image in the image pair, and fuse the lesion cell features and the lesion area features through a mutual attention model to obtain fusion features;

[0014] A metastasis prediction module configured to splice the fusion features of each of the image pairs and input the spliced features into a pre-trained classification network to obtain a metastasis result of the bladder cancer lymph nodes output by the classification network.

[0015] In a third aspect, the present application provides a device for predicting bladder cancer lymph node metastasis, including:

[0016] One or more processors; a storage device storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method for predicting bladder cancer lymph node metastasis as described in the first aspect.

[0017] In a fourth aspect, the present application provides a storage medium containing computer-executable instructions, which are used to execute the method for predicting bladder cancer lymph node metastasis as described in the first aspect when executed by a computer processor.

[0018] In this application, by obtaining multiple sliding window images corresponding to urine cell slice images and multiple cross-sectional images corresponding to nuclear magnetic resonance images; traversing and combining the multiple sliding window images and the multiple cross-sectional images to obtain multiple image pairs, each image pair including a sliding window image and a cross-sectional image; extracting the lesion cell features of the sliding window image in the image pair and extracting the lesion area features of the cross-sectional image in the image pair, and fusing the lesion cell features and the lesion area features through a mutual attention model to obtain fusion features; splicing the fusion features of each image pair and inputting them into a pre-trained classification network to obtain the metastasis result of bladder cancer lymph nodes output by the classification network. Through the above technical means, the correlation features between the lesion cell features of each sliding window image in the urine cell slice image and the lesion area features of each cross-sectional image in the nuclear magnetic resonance image can be extracted through the mutual attention model. Splicing the correlation features between each sliding window image and each cross-sectional image can obtain the correlation features between the urine cell slice image and the nuclear magnetic resonance image. Analyzing the correlation features between the urine cell slice image and the nuclear magnetic resonance image through the classification network to output the metastasis result of bladder cancer lymph nodes realizes multi-modal fusion detection of lymph node metastasis to accurately detect the metastasis result of bladder cancer lymph nodes, solves the problem of low accuracy in detecting lymph node metastasis by single radiology in the prior art, reduces the probability of misdiagnosis or missed diagnosis, and avoids the problems of under-treatment or over-treatment of patients. Description of the Drawings

[0019] Figure 1 is a flowchart of a method for predicting metastasis of bladder cancer lymph nodes provided by an embodiment of the present application;

[0020] Figure 2 is a flowchart of extracting the lesion cell features of the sliding window image provided by an embodiment of the present application;

[0021] Figure 3 is a flowchart of extracting the lesion area features of the cross-sectional image provided by an embodiment of the present application;

[0022] Figure 4 is a schematic diagram of the prediction process of the metastasis result of bladder cancer lymph nodes provided by an embodiment of the present application;

[0023] Figure 5 is a flowchart of the model training process provided by an embodiment of the present application;

[0024] Figure 6 is a flowchart of training the first feature extraction model and the second feature extraction model provided by an embodiment of the present application;

[0025] Figure 7 is a flowchart of training the first self-attention model, the second self-attention model, and the mutual attention model provided by an embodiment of the present application;

[0026] Figure 8 It is a schematic structural diagram of a bladder cancer lymph node metastasis prediction device provided by an embodiment of the present application;

[0027] Figure 9 It is a schematic structural diagram of a bladder cancer lymph node metastasis prediction device provided by an embodiment of the present application. Detailed implementation manners

[0028] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further describes the specific embodiments of the present application in detail with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. Additionally, it should be noted that for the convenience of description, only parts related to the present application are shown in the drawings rather than all the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. When the operations are completed, the process can be terminated, but there may also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0029] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0030] In a relatively related implementation, radiology (such as CT and MRI) is usually used to examine the preoperative lymph node metastasis of bladder cancer. Specifically, the bladder radiological images are collected through radiological detection techniques, the lymph nodes in the bladder radiological images are segmented through image segmentation techniques, the size of the lymph nodes is calculated based on the segmented lymph node pixels, and whether the lymph nodes are malignant is judged based on the size of the lymph nodes. However, detecting malignant lymph nodes based on the size of the lymph nodes has low accuracy and is prone to missed diagnosis or misdiagnosis problems, resulting in under-treatment or over-treatment of patients.

[0031] To solve the above problems, this embodiment provides a method for predicting bladder cancer lymph node metastasis, which can predict the metastasis results of bladder cancer lymph nodes by combining the diseased cell characteristics of urine cell slice images and the lesion area characteristics of magnetic resonance imaging, thereby reducing the probability of misdiagnosis or missed diagnosis and avoiding the problem of insufficient or excessive treatment of patients.

[0032] The bladder cancer lymph node metastasis prediction method provided in this embodiment can be performed by a bladder cancer lymph node metastasis prediction device, which can be implemented by software and / or hardware, and the bladder cancer lymph node metastasis prediction device can be composed of two or more physical entities, or can be composed of one physical entity. For example, the bladder cancer lymph node metastasis prediction device can be an intelligent terminal with strong processing capabilities such as a computer and a server. Among them, the server can be implemented by an independent server or a server cluster composed of multiple servers.

[0033] The bladder cancer lymph node metastasis prediction device is installed with at least one type of operating system, and the bladder cancer lymph node metastasis prediction device can install at least one application based on the operating system, and the application can be an application that comes with the operating system, or an application downloaded from a third-party device or server. In this embodiment, the bladder cancer lymph node metastasis prediction device is installed with at least an application that can execute the bladder cancer lymph node metastasis prediction method.

[0034] For ease of understanding, this embodiment is described by taking a server as an example of a subject that executes the method for predicting bladder cancer lymph node metastasis.

[0035] Figure 1 A flowchart of a method for predicting bladder cancer lymph node metastasis provided in an embodiment of the present application is given. Figure 1 The method for predicting bladder cancer lymph node metastasis specifically includes:

[0036] S110, acquiring a plurality of sliding window images corresponding to the urine cell slice image, and acquiring a plurality of cross-sectional images corresponding to the nuclear magnetic resonance image.

[0037] The urine cell slice image is a digital image obtained by scanning a patient's urothelial cell liquid-based slice with a digital scanner. Exemplarily, the urine cell slice image can be cropped with a sliding window of a preset size according to a preset step length, and the image obtained by each sliding window cropping is used as a sliding window image. Since the sliding window image is cropped by a sliding window of a fixed size, each sliding window image has the same size.

[0038] The nuclear magnetic resonance image is a three-dimensional image obtained by imaging the patient's bladder using the principle of nuclear magnetic resonance. The nuclear magnetic resonance image is a three-dimensional image formed by superimposing multiple cross-sectional images, and multiple cross-sectional images can be obtained in the nuclear magnetic resonance image.

[0039] S120. Traverse and combine multiple sliding window images and multiple cross-sectional images to obtain multiple image pairs, where each image pair includes a sliding window image and a cross-sectional image.

[0040] Exemplarily, pair up all the sliding window images cropped from the urine cell slice image with all the cross-sectional images obtained from the nuclear magnetic resonance imaging. Each sliding window image will be combined with each cross-sectional image to form a group of image pairs, and finally multiple groups of image pairs are obtained. For example, if N sliding window images are cropped from the urine cell slice image and M cross-sectional images are obtained from the nuclear magnetic resonance imaging, traversing and combining the N sliding window images and M cross-sectional images will result in N*M image pairs.

[0041] S130. Extract the lesion cell features of the sliding window image in the image pair, extract the lesion area features of the cross-sectional image in the image pair, and fuse the lesion cell features and the lesion area features through a mutual attention model to obtain the fused features.

[0042] Exemplarily, extract the lesion cell features of the sliding window image in each image pair and extract the lesion area features of the cross-sectional image in each image pair. Through a pre-trained mutual attention model, extract the correlation features between the lesion cell features and the lesion area features corresponding to each image pair, and finally output the fused features of the lesion cell features and the lesion area features corresponding to each image pair. It can be understood that the lesion cell features are the morphological features of the lesion cells in the corresponding sliding window image, including cell color, cell membrane shape, and nucleus area, etc. The lesion area features are the features such as the color, position, and shape of the lesion area (tumor mass) in the corresponding cross-sectional image. The lesion cell features are the feature information of the tumor mass at the microscopic level, and the lesion area features are the feature information of the tumor mass at the macroscopic level. For the tumor mass in the same area, relatively high-similarity lesion cell features and lesion area features can be extracted. When using the mutual attention model to find the correlation features between the lesion cell features and the lesion area features, it is equivalent to using the lesion cell features at the microscopic level to verify whether the detection of the lesion area at the macroscopic level is accurate, thereby achieving precise positioning of the tumor mass and laying a solid data foundation for accurately predicting the metastasis result based on the position information of the tumor mass.

[0043] Since each image pair contains duplicate sliding window images and cross-sectional images, if the features of the sliding window image and the cross-sectional image in each image pair are extracted, the features of the same sliding window image will be repeatedly extracted, affecting the feature extraction efficiency. In this regard, the features of each sliding window image and each cross-sectional image can be extracted, and finally, based on the sliding window image and the cross-sectional image in the image pair, the corresponding lesion cell features and lesion area features are fused to obtain the fused features of the image pair.

[0044] In one embodiment, the lesion cell features of the sliding window image and the lesion area features of the cross-sectional image can be extracted through a convolutional neural network or an attention network.

[0045] In another embodiment, the lesion cell features of the sliding window image can be jointly extracted through a convolutional neural network and an attention network. Optionally, Figure 2 is a flowchart for extracting the lesion cell features of the sliding window image provided by an embodiment of the present application. As Figure 2 shown, the steps for extracting the lesion cell features of the sliding window image specifically include S1301 - S1302:

[0046] S1301: Input the sliding window image into a pre-trained first deep feature extraction model to obtain the first feature of the sliding window image.

[0047] Among them, the first deep feature extraction model is a convolutional neural network. Input the sliding window image into the first deep feature extraction model, and extract the morphological features of the lesion cells in the sliding window image as the first feature through the first deep feature model.

[0048] S1302: Input the first feature into a pre-trained first self-attention model to obtain the lesion cell features output by the first self-attention model.

[0049] Exemplarily, convert the first feature into a feature sequence and input it into the first self-attention model, and extract the self-attention features in the feature sequence as the lesion cell features through the first self-attention model.

[0050] In this embodiment, the lesion cell features of the sliding window image are extracted by combining a convolutional neural network and a self-attention network to effectively extract the local features and global features of the sliding window image, and improve the extraction accuracy of the lesion cell features.

[0051] Similarly, the lesion area features of the cross-sectional image can be jointly extracted through a convolutional neural network and an attention network. Optionally, Figure 3 is a flowchart for extracting the lesion area features of the cross-sectional image provided by an embodiment of the present application. As Figure 3 shown, the steps for extracting the lesion area features of the cross-sectional image specifically include S1303 - S1305:

[0052] S1303: Reduce the cross-sectional image to a preset size.

[0053] Exemplarily, the resolution of the cross-sectional image is relatively large. If a convolutional neural network is used to directly extract features from it, the required computing power is large and the operation time is long. Therefore, the cross-sectional image can be reduced to a preset size before extracting the corresponding lesion area features.

[0054] S1304. Input the scaled-down cross-sectional image into a pre-trained second-depth feature extraction model to obtain the second features of the cross-sectional image.

[0055] Among them, the second-depth feature extraction model is a convolutional neural network. Input the scaled-down cross-sectional image into the pre-trained second-depth feature extraction model, and extract features such as the color, position, and shape of the lesion area in the cross-sectional image through the second-depth feature extraction model as the second features.

[0056] S1305. Input the second features into a pre-trained second self-attention model to obtain the lesion area features output by the second self-attention model.

[0057] Exemplarily, convert the second features into a feature sequence and input it into the second self-attention model, and extract the self-attention features in the feature sequence through the second self-attention model as the lesion area features.

[0058] In this embodiment, the lesion area features of the cross-sectional image are extracted by combining a convolutional neural network and a self-attention network to effectively extract the local features and global features of the cross-sectional image, and improve the extraction accuracy of the lesion area features.

[0059] After extracting the lesion cell features of each sliding window image and the lesion area features of each cross-sectional image, the lesion cell features of a sliding window image and the lesion area features of a cross-sectional image can be spliced into a feature sequence. Input the spliced feature sequence into a pre-trained mutual attention model, and the mutual attention model determines the correlation weights between the lesion cell features and the lesion area features. Perform weighted summation of the correlation weights and the corresponding input feature sequence to obtain the fused features. It can be understood that the higher the feature similarity between the lesion cell features and the lesion area features, the greater the correlation weights. Therefore, the finally output fused features can be used to represent the association features between the lesion cell features and the lesion area features.

[0060] S140. Splice the fused features of each image pair and input them into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph nodes output by the classification network.

[0061] Exemplarily, splice the fused features of each image pair to obtain the spliced features, input the spliced features into a pre-trained classification network, and the classification network locates the position of the bladder cancer lymph nodes based on each fused feature in the spliced features. Finally, based on whether the position information of the bladder cancer lymph nodes moves, the metastasis result of whether the bladder cancer lymph nodes metastasize is output.

[0062] To more intuitively understand the process of predicting the metastasis result of bladder cancer lymph nodes, this embodiment takes Figure 4 the prediction process of the metastasis result of the bladder cancer lymph nodes shown as an example for detailed description. AsFigure 4 As shown, multiple sliding window images are obtained from urine cell section images, and multiple cross-sectional images are obtained from nuclear magnetic resonance images. The sliding window images are input into the first feature extraction model to obtain the first features output by the first feature extraction model, and the first features are input into the first self-attention model to obtain the lesion cell features output by the first self-attention model. The cross-sectional images are input into the second feature extraction model to obtain the second features output by the second feature extraction model, and the second features are input into the second self-attention model to obtain the lesion area features output by the second self-attention model. The lesion cell features and the lesion area features are concatenated and then input into the mutual attention model to obtain the fused features output by the mutual attention model. The fused features corresponding to each lesion cell feature and each lesion area feature are concatenated to obtain the concatenated features, and the concatenated features are input into the classification network to obtain the metastasis result of whether the lymph nodes of bladder cancer have metastasized output by the classification network.

[0063] Optionally, referring to Figure 4 , a lesion cell localization model and a lesion area localization model can also be added. After the first self-attention model outputs the lesion cell features, the lesion cell features are input into the lesion cell localization model, and the lesion cell localization model detects and outputs the positions of the lesion cells in the sliding window images. Also, after the second self-attention model outputs the lesion area features, the lesion area features are input into the lesion area localization model, and the lesion area localization model detects and outputs the positions of the lesion areas in the cross-sectional images. The positions of the lesion cells in the sliding window images and the positions of the lesion areas in the cross-sectional images can be used to assist doctors in judging whether the prediction results of the classification network are accurate, providing quality assurance for clinical diagnosis.

[0064] In one embodiment, the first deep feature extraction model, the second deep feature extraction model, the first self-attention model, the second self-attention model, the mutual attention model, and the classification network can form a network model. During the training phase, the network model is trained with a large number of sample images so that the network model has the function of predicting the metastasis result of the lymph nodes of bladder cancer.

[0065] In another embodiment, the training processes of the first deep feature extraction model, the second deep feature extraction model, the first self-attention model, the second self-attention model, the mutual attention model, and the classification network can be divided into three training phases. In the first training phase, the first deep feature extraction model and the second deep feature extraction model are first trained. In the second training phase, the first self-attention model, the second self-attention model, and the mutual attention model are then trained. Finally, the classification network is trained. Optionally, Figure 5 is the flowchart of the model training process provided by the embodiments of the present application. As Figure 5 shown, the steps of this model training process specifically include S210 - S260:

[0066] S210. Divide the urine cell slice sample images of each patient into multiple sliding window sample images, and obtain multiple cross-sectional sample images corresponding to the magnetic resonance sample images of each patient.

[0067] Exemplarily, obtain the urine cell slice images and magnetic resonance images of multiple patients as urine cell slice sample images and magnetic resonance sample images respectively. Cut out multiple sliding window sample images from each urine cell slice sample image through a sliding window, and take each cross-sectional image in each magnetic resonance image as the corresponding cross-sectional sample image.

[0068] S220. Based on the multiple sliding window sample images and multiple cross-sectional sample images, train the first feature extraction model and the second feature extraction model through a contrastive learning algorithm.

[0069] Exemplarily, divide the multiple sliding window sample images and multiple cross-sectional sample images into positive sample data and negative sample data, and train the first feature extraction model and the second feature extraction model through a contrastive learning algorithm based on the positive sample data and negative sample data.

[0070] Optionally, Figure 6 is the flowchart for training the first feature extraction model and the second feature extraction model provided by the embodiments of the present application. As Figure 6 shown, the steps for training the first feature extraction model and the second feature extraction model specifically include S2201 - S2205:

[0071] S2201. Traverse and combine the sliding window sample images and cross-sectional sample images of the same patient to obtain first training sample data, and randomly combine the sliding window sample images and cross-sectional sample images of different patients to obtain second training sample data.

[0072] Exemplarily, traverse and combine the multiple sliding window sample images corresponding to the urine cell slice sample images of the same patient and the multiple cross-sectional sample images corresponding to the magnetic resonance sample images to obtain multiple first image pairs, and take the multiple first image pairs as the first training sample data. Each first training sample data includes a sliding window sample image and a cross-sectional sample image belonging to the same patient. Randomly combine the sliding window sample images and cross-sectional sample images of different patients to obtain multiple second image pairs, and take the multiple second image pairs as the second training sample data.

[0073] S2202. Input the sliding window sample images and cross-sectional sample images in the first training sample data into the first classification model and the second classification model respectively to obtain the cell category and region category output by the first classification model and the second classification model. The first classification model includes a first feature extraction model and a first classification network, and the second classification model includes a second feature extraction model and a second classification network.

[0074] Exemplarily, input the sliding window sample images in the first training sample data into the first classification model. Extract the shape features of the diseased cells in the sliding window sample images as the first sample features through the first feature extraction model in the first classification model. Based on the first sample features, identify the cell category in the sliding window sample images as diseased cells or non-diseased cells through the first classification network in the first classification model. Input the cross-sectional sample images in the first training sample data into the second classification model. Extract the features of the lesion regions in the cross-sectional sample images as the second sample features through the second feature extraction model in the second classification model. Based on the second sample features, identify the difference category in the cross-sectional sample images as lesion regions or non-lesion regions through the second classification network in the second classification model.

[0075] S2203. When the cell categories and region categories corresponding to the sliding window sample images and cross-sectional sample images in the first training sample data are accurately identified as diseased cells and lesion regions, determine the first training sample data as true paired data, and determine the remaining first training sample data and the second training sample data as false paired data.

[0076] Exemplarily, the urine cell section sample images are labeled with corresponding diseased cell detection frames. The diseased cell detection frames mark the diseased cells in the urine cell section sample images. It is possible to determine whether there are diseased cells in the sliding window sample images based on the pixel coordinates of the diseased cell detection frames labeled on the urine cell section sample images. If the pixel coordinates of the diseased cell detection frames overlap with the pixel coordinates of the sliding window sample images, it is determined that there are diseased cells in the sliding window sample images; otherwise, it is determined that there are no diseased cells in the sliding window sample images. After the first classification model outputs the cell category of the sliding window sample images, if the cell category is diseased cells and there are indeed diseased cells in the sliding window sample images, it is determined that the first classification model accurately identifies the presence of diseased cells in the sliding window sample images.

[0077] Each cross-sectional sample image of the nuclear magnetic resonance sample images is labeled with a corresponding lesion region detection frame. After the second classification model outputs the region category of the cross-sectional sample images, if the region category is a lesion region and the cross-sectional sample images are labeled with lesion region detection frames, it is determined that the second classification model accurately identifies the presence of lesion regions in the cross-sectional sample images.

[0078] Only when the first classification model accurately identifies the presence of diseased cells in the sliding window sample images and the second classification model accurately identifies the presence of lesion regions in the cross-sectional sample images, can the first training sample data be determined as true paired data, thereby indicating that the sliding window sample images and cross-sectional sample images in the first training sample data are highly correlated positive sample data.

[0079] When the first classification model misidentifies the sliding window sample image and the cross-sectional sample image, or there are no diseased cells in the sliding window sample image or no lesion area in the cross-sectional sample image, the first training sample data is determined as mispaired data, indicating that the sliding window sample image and the cross-sectional sample image in the first training sample data are negative sample data with low correlation. For the second training sample data, the sliding sample image and the cross-sectional sample image belong to different patients, and their correlation is even lower. Therefore, the second training sample data is also regarded as mispaired data.

[0080] S2204. Maximize the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the true paired data, and minimize the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the mispaired data. The first sample feature and the second sample feature are respectively extracted by the first feature extraction model and the second feature extraction model.

[0081] Exemplarily, calculate the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the true paired data, and maximize this cosine similarity through a loss function. Calculate the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the mispaired data, and minimize this cosine similarity through a loss function.

[0082] S2205. Determine the symmetric cross-entropy loss value based on the cosine similarity of the true paired data and the cosine similarity of the mispaired data, and optimize the first classification model and the second classification model based on the symmetric cross-entropy loss value.

[0083] Exemplarily, determine the symmetric cross-entropy loss value based on the cross-entropy loss function and the cosine similarity of each true paired data and the cosine similarity of the mispaired data. Based on the symmetric cross-entropy loss value, respectively perform backpropagation to optimize the model parameters of the first classification network and the first feature extraction model in the first classification model and the second classification network and the second feature extraction model in the second classification model.

[0084] The process of steps S2201 - S2205 is to train the first classification model and the second classification model once. After training the first classification model and the second classification model once, the first training sample data is input into the first classification model and the second classification model again. Based on the cell categories and region categories output by the first classification model and the second classification model, the true paired data and the false paired data are re - screened. Based on the newly screened true paired data and false paired data, the symmetric cross - entropy loss value is calculated. Based on the symmetric cross - entropy loss value, the first classification model and the second classification model are optimized again. When the first classification model and the second classification model meet the preset convergence conditions, it is confirmed that the training of the first classification model and the second classification model is completed. The first feature extraction model and the second feature extraction model obtained after the training is completed are used for actual transfer result prediction and subsequent model training.

[0085] S230. Extract the first sample features of each sliding window sample image through the trained first feature extraction model, and extract the second sample features of each cross - section sample image through the trained second feature extraction model.

[0086] Exemplarily, each sliding window sample image is input into the trained first feature extraction model to obtain the first sample features output by the first feature extraction model, and each cross - section sample image is input into the trained second feature extraction model to obtain the second sample features output by the second feature extraction model, so as to be used for training the first self - attention model, the second self - attention model and the mutual attention model in the second training stage based on the first sample features and the second sample features.

[0087] S240. Based on the first sample features and the second sample features, train the first self - attention model, the second self - attention model and the mutual attention model through a segmentation training algorithm.

[0088] Exemplarily, the implementation process of the segmentation training algorithm is to perform cell segmentation and region segmentation on the first sample features and the second sample features based on a network model composed of the first self - attention model, the second self - attention model, the mutual attention model and the segmentation model, calculate the loss value based on the segmentation result, and optimize the model parameters of the network model based on the loss value.

[0089] Optionally, Figure 7 is a flowchart for training the first self - attention model, the second self - attention model, and the mutual attention model provided by an embodiment of the present application. As Figure 7 shown, the steps for training the first self - attention model, the second self - attention model, and the mutual attention model specifically include S2401 - S2404:

[0090] S2401. Input the first sample features and the second sample features corresponding to the first training sample data into the first self-attention model and the second self-attention model respectively, to obtain the lesion cell sample features output by the first self-attention model and the lesion area sample features output by the second self-attention model.

[0091] Exemplarily, separate the first sample features and the second sample features of the first training sample data, that is, the sliding window sample image and the cross-sectional sample image belonging to the same patient, and input them into the first self-attention model and the second self-attention model respectively. The first self-attention model extracts the self-attention features of the first sample features as the lesion cell sample features, and the second self-attention model extracts the self-attention features of the second sample features as the lesion area sample features.

[0092] S2402. Input the lesion cell sample features and the lesion area sample features into the mutual attention model to obtain the sample fusion features output by the mutual attention model.

[0093] Exemplarily, splice the lesion cell sample features and the lesion area sample features and input them into the mutual attention model. The mutual attention model extracts the correlation between the lesion cell sample features and the lesion area sample features, and performs weighted processing on the spliced features based on this correlation to obtain the sample fusion features.

[0094] S2403. Input the sample fusion features into the segmentation model to obtain the mask result output by the segmentation model. The mask result includes the lesion cell mask and the lesion area mask.

[0095] Exemplarily, input the sample fusion features into the segmentation model. The segmentation model simultaneously detects the lesion cells in the sliding window sample image and the lesion area in the cross-sectional sample image based on the sample fusion features, and outputs the lesion cell mask and the lesion area mask. The lesion cell mask contains each pixel point of the lesion cells, and the lesion area mask contains each pixel point of the lesion area.

[0096] S2404. Determine the loss value of each pixel point based on the mask result and the annotation information of the sliding window sample image and the cross-sectional sample image in the first training sample data. Determine the mean square error loss value based on the loss values of each pixel point, and optimize the first self-attention model, the second self-attention model, the mutual attention model and the segmentation model based on the mean square error loss value.

[0097] Exemplarily, the urine cell section sample image is labeled with a diseased cell detection frame. The pixel points of the diseased cells in the sliding window sample image are determined according to the pixel coordinates of the diseased cell detection frame. The loss value of each pixel point is calculated based on the pixel points of the diseased cells in the sliding window sample image and the pixel points in the diseased cell mask. The cross-sectional sample image is labeled with a lesion area detection frame. The pixel points of the lesion area in the cross-sectional sample image are determined according to the lesion area detection frame. The loss value of each pixel point is calculated based on the pixel points of the lesion area in the cross-sectional sample image and the pixel points in the lesion area mask. The average value is calculated based on the loss values of the above-mentioned respective pixel points, and the variance is calculated based on the loss values of the respective pixel points and the average value as the mean square error loss value. The segmentation model, the mutual attention model, the first self-attention model, and the second self-attention model are reversely optimized through the mean square error loss value.

[0098] Steps S2401 - S2404 are one training of the segmentation model, the mutual attention model, the first self-attention model, and the second self-attention model. After training the segmentation model, the mutual attention model, the first self-attention model, and the second self-attention model using the first sample features and the second sample features corresponding to all the first training sample data, the training ends. The mutual attention model, the first self-attention model, and the second self-attention model obtained after the training ends are used to predict the metastasis result and for subsequent training of the classification network in the third training stage.

[0099] S250. Extract the sample fusion features of the sliding window sample image and the cross-sectional sample image of the same patient through the trained first feature extraction model, second feature extraction model, first self-attention model, second self-attention model, and mutual attention model.

[0100] Exemplarily, only the classification network is trained in the third training stage so that the trained classification network has the function of predicting whether the lymph nodes metastasize. Therefore, the sample data used in the third training stage are the respective sliding window sample images and the respective cross-sectional sample images of the same patient. In the first training stage and the second training stage, the first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model, and the mutual attention model have been trained. The sample fusion features of each first training sample data obtained by traversing and combining the respective sliding window sample images and the respective cross-sectional sample images of the same patient can be directly extracted using the first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model, and the mutual attention model, so as to use the sample fusion features of each first training sample data as the sample data in the third training stage, which is beneficial to improving the training efficiency of the classification network.

[0101] Optionally, input the first sample features and the second sample features corresponding to the first training sample data into the trained first self-attention model and the second self-attention model respectively, to obtain the lesion cell sample features output by the first self-attention model and the lesion area sample features output by the second self-attention model; input the lesion cell sample features and the lesion area sample features into the trained mutual attention model, to obtain the sample fusion features output by the mutual attention model. For example, input the first sample features of the first training sample data extracted by the trained first feature extraction model into the trained first self-attention model and input the second sample features of the first training sample data extracted by the trained first feature extraction model into the trained second self-attention model, to obtain the lesion cell sample features output by the first self-attention model and the lesion area sample features output by the second self-attention model. Concatenate the lesion cell sample features and the lesion area sample features corresponding to the first training sample data and input them into the trained mutual attention model, to obtain the sample fusion features output by the mutual attention model.

[0102] S260. Concatenate the sample fusion features of each sample of the same patient to obtain a sample concatenation feature, and train a classification network based on the sample concatenation features of each patient.

[0103] Exemplarily, obtain the sample fusion features output by the trained mutual attention model, concatenate the sample fusion features of each sample of the same patient to obtain a sample concatenation feature, input the sample concatenation feature into the classification network, and optimize the model parameters of the classification network based on the classification results output by the classification network.

[0104] Optionally, concatenate the sample fusion features corresponding to the first training sample data of each sample of the same patient to obtain a sample concatenation feature; input the sample concatenation feature into the classification network to obtain the classification result output by the classification network; determine the loss value based on the classification result and the metastasis information of the corresponding patient, and optimize the classification network based on the loss value. For example, after the trained mutual attention model outputs the sample fusion features corresponding to the first training sample data of patient A, concatenate the sample fusion features corresponding to the first training sample data of patient A to obtain the sample concatenation feature of patient A, input the sample concatenation feature of patient A into the classification network to obtain the metastasis result of the bladder cancer lymph nodes of patient A output by the classification network. Determine the loss value according to the metastasis result of patient A output by the classification network and the actual metastasis result of patient A, and reversely optimize the model parameters of the classification network according to the loss value. After training the classification network based on the sample concatenation features of each patient, obtain the trained classification network.

[0105] After the classification network is also trained, the trained first feature extraction model, second feature extraction model, first self-attention model, second self-attention model, mutual attention model and classification network can be used to process the urine cell slice image and magnetic resonance image of the patient to predict whether the lymph nodes of the patient with bladder cancer have metastasized.

[0106] In summary, the method for predicting the metastasis of lymph nodes in bladder cancer provided by the embodiments of the present application obtains multiple sliding window images corresponding to the urine cell slice image and multiple cross-sectional images corresponding to the magnetic resonance image; traverses and combines the multiple sliding window images and multiple cross-sectional images to obtain multiple image pairs, and each image pair includes a sliding window image and a cross-sectional image; extracts the lesion cell features of the sliding window image in the image pair, extracts the lesion area features of the cross-sectional image in the image pair, and fuses the lesion cell features and the lesion area features through a mutual attention model to obtain fused features; splices the fused features of each image pair and inputs them into a pre-trained classification network to obtain the metastasis result of the lymph nodes of bladder cancer output by the classification network. Through the above technical means, the correlation features between the lesion cell features of each sliding window image in the urine cell slice image and the lesion area features of each cross-sectional image in the magnetic resonance image can be extracted through the mutual attention model. The correlation features between each sliding window image and each cross-sectional image can be spliced to obtain the correlation features between the urine cell slice image and the magnetic resonance image. The classification network analyzes the correlation features between the urine cell slice image and the magnetic resonance image to output the metastasis result of the lymph nodes of bladder cancer, realizing multi-modal fusion detection of lymph node metastasis to accurately detect the metastasis result of the lymph nodes of bladder cancer, solving the problem of low accuracy in detecting lymph node metastasis by single radiology in the prior art, reducing the probability of misdiagnosis or missed diagnosis, and avoiding the problems of under-treatment or over-treatment of patients.

[0107] Based on the above embodiments, Figure 8 This is a schematic structural diagram of a device for predicting the metastasis of lymph nodes in bladder cancer provided by the embodiments of the present application. Refer to Figure 8 This embodiment provides a device for predicting the metastasis of lymph nodes in bladder cancer, which specifically includes: an image acquisition module 31, an image combination module 32, a feature fusion module 33, and a metastasis prediction module 34.

[0108] Among them, the image acquisition module 31 is configured to obtain multiple sliding window images corresponding to the urine cell slice image and multiple cross-sectional images corresponding to the magnetic resonance image;

[0109] The image combination module 32 is configured to traverse and combine the multiple sliding window images and multiple cross-sectional images to obtain multiple image pairs, and each image pair includes a sliding window image and a cross-sectional image;

[0110] The feature fusion module 33 is configured to extract the lesion cell features of the sliding window images in the image pair, extract the lesion area features of the cross-sectional images in the image pair, and fuse the lesion cell features and the lesion area features through a mutual attention model to obtain fused features;

[0111] The metastasis prediction module 34 is configured to splice the fused features of each image pair and input them into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph nodes output by the classification network.

[0112] Based on the above embodiments, the feature fusion module 33 includes: a first feature extraction unit configured to input the sliding window image into a pre-trained first deep feature extraction model to obtain the first feature of the sliding window image; a first self-attention unit configured to input the first feature into a pre-trained first self-attention model to obtain the lesion cell features output by the first self-attention model.

[0113] Based on the above embodiments, the feature fusion module 33 includes: an image reduction unit configured to reduce the cross-sectional image to a preset size; a second feature extraction unit configured to input the reduced cross-sectional image into a pre-trained second deep feature extraction model to obtain the second feature of the cross-sectional image; a second self-attention unit configured to input the second feature into a pre-trained second self-attention model to obtain the lesion area features output by the second self-attention model.

[0114] Based on the above embodiments, the bladder cancer lymph node metastasis prediction device further includes a model training module, and the model training module includes: a sample image acquisition unit configured to train and divide the urine cell slice sample images of each patient into multiple sliding window sample images, and acquire multiple cross-sectional sample images corresponding to the nuclear magnetic resonance sample images of each patient; a first training unit configured to train the first feature extraction model and the second feature extraction model based on the multiple sliding window sample images and the multiple cross-sectional sample images through a contrast learning algorithm; a first sample feature acquisition unit configured to extract the first sample features of each sliding window sample image through the trained first feature extraction model, and extract the second sample features of each cross-sectional sample image through the trained second feature extraction model; a second training unit configured to train the first self-attention model, the second self-attention model, and the mutual attention model based on the first sample features and the second sample features through a segmentation training algorithm; a second sample feature acquisition unit configured to extract the sample fusion features of the sliding window sample images and the cross-sectional sample images of the same patient through the trained first feature extraction model, second feature extraction model, first self-attention model, second self-attention model, and mutual attention model; a third training unit configured to splice the sample fusion features of the same patient to obtain a sample splicing feature, and train the classification network based on the sample splicing features of each patient.

[0115] Based on the above embodiments, the first training unit includes: a sample image combination sub-unit configured to traverse and combine the sliding window sample images and cross-sectional sample images of the same patient to obtain first training sample data, and randomly combine the sliding window sample images and cross-sectional sample images of different patients to obtain second training sample data; a first classification sub-unit configured to input the sliding window sample images and cross-sectional sample images in the first training sample data into a first classification model and a second classification model respectively, to obtain the cell category and region category output by the first classification model and the second classification model. The first classification model includes a first feature extraction model and a first classification network, and the second classification model includes a second feature extraction model and a second classification network; a sample data classification sub-unit configured to, when the cell category and region category corresponding to the sliding window sample images and cross-sectional sample images in the first training sample data are accurately identified as diseased cells and lesion regions, determine the first training sample data as true paired data, and determine the remaining first training sample data and second training sample data as false paired data; a cosine similarity determination sub-unit configured to maximize the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the true paired data, and minimize the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the false paired data. The first sample feature and the second sample feature are respectively extracted by the first feature extraction model and the second feature extraction model; a first optimization sub-unit configured to determine a symmetric cross-entropy loss value based on the cosine similarity of the true paired data and the cosine similarity of the false paired data, and optimize the first classification model and the second classification model based on the symmetric cross-entropy loss value.

[0116] Based on the above embodiments, the second training unit includes: a first self-attention feature extraction sub-unit configured to input the first sample feature and the second sample feature corresponding to the first training sample data into a first self-attention model and a second self-attention model respectively, to obtain the diseased cell sample feature output by the first self-attention model and the lesion region sample feature output by the second self-attention model; a second mutual-attention feature extraction sub-unit configured to input the diseased cell sample feature and the lesion region sample feature into a mutual-attention model, to obtain the sample fusion feature output by the mutual-attention model; a segmentation processing sub-unit configured to input the sample fusion feature into a segmentation model, to obtain the mask result output by the segmentation model. The mask result includes a diseased cell mask and a lesion region mask; a second optimization sub-unit configured to determine the loss value of each pixel point based on the mask result and the annotation information of the sliding window sample image and the cross-sectional sample image in the first training sample data, determine the mean square error loss value based on the loss values of each pixel point, and optimize the first self-attention model, the second self-attention model, the mutual-attention model and the segmentation model based on the mean square error loss value.

[0117] Based on the above embodiments, the second sample feature acquisition unit includes: a second self-attention feature extraction subunit, configured to input the first sample feature and the second sample feature corresponding to the first training sample data into the trained first self-attention model and the second self-attention model respectively, to obtain the lesion cell sample feature output by the first self-attention model and the lesion region sample feature output by the second self-attention model; a second mutual-attention feature extraction subunit, configured to input the lesion cell sample feature and the lesion region sample feature into the trained mutual-attention model, to obtain the sample fusion feature output by the mutual-attention model; correspondingly, the third training unit includes: a feature splicing subunit, configured to splice the sample fusion features corresponding to each first training sample data of the same patient to obtain a sample splicing feature; a second classification subunit, configured to input the sample splicing feature into a classification network, to obtain the classification result output by the classification network; a third optimization subunit, configured to determine a loss value based on the classification result and the metastasis information of the corresponding patient, and optimize the classification network based on the loss value.

[0118] As described above, the bladder cancer lymph node metastasis prediction device provided by the embodiments of the present application obtains a plurality of sliding window images corresponding to the urine cell section image, and obtains a plurality of cross-sectional images corresponding to the nuclear magnetic resonance image; traverses and combines the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, each image pair including a sliding window image and a cross-sectional image; extracts the lesion cell features of the sliding window image in the image pair, extracts the lesion region features of the cross-sectional image in the image pair, and fuses the lesion cell features and the lesion region features through a mutual-attention model to obtain a fusion feature; splices the fusion features of each image pair and inputs them into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph node output by the classification network. Through the above technical means, the correlation features between the lesion cell features of each sliding window image in the urine cell section image and the lesion region features of each cross-sectional image in the nuclear magnetic resonance image can be extracted through the mutual-attention model, and the correlation features between each sliding window image and each cross-sectional image can be spliced to obtain the correlation features between the urine cell section image and the nuclear magnetic resonance image. The classification network analyzes the correlation features between the urine cell section image and the nuclear magnetic resonance image to output the metastasis result of the bladder cancer lymph node, realizing multi-modal fusion detection of lymph node metastasis, accurately detecting the metastasis result of the bladder cancer lymph node, solving the problem of low accuracy in detecting lymph node metastasis by single radiology in the prior art, reducing the probability of misdiagnosis or missed diagnosis, and avoiding the problems of under-treatment or over-treatment of patients.

[0119] The bladder cancer lymph node metastasis prediction device provided by the embodiments of the present application can be used to execute the bladder cancer lymph node metastasis prediction method provided by the above embodiments, and has the corresponding functions and beneficial effects.

[0120] Figure 9 It is a schematic structural diagram of a bladder cancer lymph node metastasis prediction device provided by an embodiment of the present application. Refer to Figure 9 , the bladder cancer lymph node metastasis prediction device includes: a processor 41, a memory 42, a communication device 43, an input device 44, and an output device 45. The number of processors 41 in the bladder cancer lymph node metastasis prediction device can be one or more, and the number of memories 42 in the bladder cancer lymph node metastasis prediction device can be one or more. The processor 41, memory 42, communication device 43, input device 44, and output device 45 of the bladder cancer lymph node metastasis prediction device can be connected through a bus or other means.

[0121] The memory 42, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the bladder cancer lymph node metastasis prediction method of any embodiment of the present application (for example, the image acquisition module 31, image combination module 32, feature fusion module 33, and metastasis prediction module 34 in the bladder cancer lymph node metastasis prediction device). The memory 42 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device, etc. In addition, the memory 42 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory can further include a memory remotely set relative to the processor, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.

[0122] The communication device 43 is used for data transmission.

[0123] The processor 41 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 42, that is, implements the above-mentioned bladder cancer lymph node metastasis prediction method.

[0124] The input device 44 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function control of the device. The output device 45 can include a display device such as a display screen.

[0125] The above-provided bladder cancer lymph node metastasis prediction device can be used to execute the bladder cancer lymph node metastasis prediction method provided by the above embodiment, and has corresponding functions and beneficial effects.

[0126] The embodiments of the present application further provide a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a method for predicting bladder cancer lymph node metastasis when executed by a computer processor. The method for predicting bladder cancer lymph node metastasis includes: obtaining a plurality of sliding window images corresponding to urine cell section images, and obtaining a plurality of cross-sectional images corresponding to nuclear magnetic resonance images; traversing and combining the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, each image pair including a sliding window image and a cross-sectional image; extracting the lesion cell features of the sliding window image in the image pair, extracting the lesion area features of the cross-sectional image in the image pair, and fusing the lesion cell features and the lesion area features through a mutual attention model to obtain fused features; splicing the fused features of each image pair and inputting them into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph nodes output by the classification network.

[0127] Storage medium - Any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media such as CD-ROMs, floppy disks or magnetic tape devices; computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. The storage medium may also include other types of memory or combinations thereof. Additionally, the storage medium may be located in a first computer system in which the program is executed, or may be located in a different second computer system that is connected to the first computer system via a network (such as the Internet). The second computer system may provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media residing in different locations (such as in different computer systems connected via a network). The storage medium may store program instructions (such as specifically implemented as a computer program) executable by one or more processors.

[0128] Of course, for a storage medium containing computer-executable instructions provided by the embodiments of the present application, the computer-executable instructions are not limited to the method for predicting bladder cancer lymph node metastasis as described above, and may also execute related operations in the method for predicting bladder cancer lymph node metastasis provided by any embodiment of the present application.

[0129] The device for predicting bladder cancer lymph node metastasis, the system for predicting bladder cancer lymph node metastasis, the storage medium, and the device for predicting bladder cancer lymph node metastasis provided in the above embodiments may execute the method for predicting bladder cancer lymph node metastasis provided by any embodiment of the present application. For technical details not described in detail in the above embodiments, reference may be made to the method for predicting bladder cancer lymph node metastasis provided by any embodiment of the present application.

[0130] The above is only the preferred embodiment of the present application and the technical principles applied. The present application is not limited to the specific embodiments here. Various obvious changes, re-adjustments and substitutions that can be made by those skilled in the art will not depart from the protection scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, it may also include more other equivalent embodiments, and the scope of the present application is determined by the scope of the claims.

Claims

1. A method for predicting lymph node metastasis of bladder cancer, characterized in that, Including: Obtaining a plurality of sliding window images corresponding to urine cell section images, and obtaining a plurality of cross-sectional images corresponding to nuclear magnetic resonance images; Performing traversal combination on the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, each of the image pairs including one of the sliding window images and one of the cross-sectional images; Extracting the diseased cell features of the sliding window image in the image pair, extracting the lesion region features of the cross-sectional image in the image pair, and fusing the diseased cell features and the lesion region features through a mutual attention model to obtain fusion features; wherein, the diseased cell features are extracted by combining a first feature extraction model and a first self-attention model, the lesion region features are extracted by a second feature extraction model and a second self-attention model, and the training steps of the first feature extraction model and the second feature extraction model include: performing traversal combination on the sliding window sample images and cross-sectional sample images of the same patient to obtain first training sample data, and randomly combining the sliding window sample images and cross-sectional sample images of different patients to obtain second training sample data; respectively inputting the sliding window sample images and cross-sectional sample images in the first training sample data into a first classification model and a second classification model to obtain the cell category and region category output by the first classification model and the second classification model, the first classification model including a first feature extraction model and a first classification network, and the second classification model including a second feature extraction model and a second classification network; in the case where the cell category and region category corresponding to the sliding window sample image and cross-sectional sample image in the first training sample data are accurately identified as diseased cells and lesion regions, determining the first training sample data as true paired data, and determining the remaining first training sample data and the second training sample data as false paired data; maximizing the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the true paired data, and minimizing the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the false paired data, the first sample feature and the second sample feature being respectively extracted by the first feature extraction model and the second feature extraction model; determining a symmetric cross-entropy loss value according to the cosine similarity of the true paired data and the cosine similarity of the false paired data, and optimizing the first classification model and the second classification model based on the symmetric cross-entropy loss value; After splicing the fusion features of each of the image pairs, inputting them into a pre-trained classification network to obtain the metastasis result of bladder cancer lymph nodes output by the classification network.

2. The bladder cancer lymph node metastasis prediction method according to claim 1, wherein The extracting the diseased cell features of the sliding window image in the image pair includes: Inputting the sliding window image into a pre-trained first deep feature extraction model to obtain a first feature of the sliding window image; Inputting the first feature into a pre-trained first self-attention model to obtain the diseased cell features output by the first self-attention model.

3. The method for predicting lymph node metastasis of bladder cancer according to claim 2, wherein, Extracting the lesion area features of the cross-sectional image in the image pair includes: Reducing the cross-sectional image to a preset size; Inputting the reduced cross-sectional image into a pre-trained second depth feature extraction model to obtain the second feature of the cross-sectional image; Inputting the second feature into a pre-trained second self-attention model to obtain the lesion area features output by the second self-attention model.

4. The method for predicting bladder cancer lymph node metastasis according to claim 3, wherein The training steps of the first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model, the mutual attention model, and the classification network include: Correspondingly dividing the urine cell slice sample images of each patient into multiple sliding window sample images, and correspondingly obtaining multiple cross-sectional sample images from the nuclear magnetic resonance sample images of each patient; Training the first feature extraction model and the second feature extraction model based on the multiple sliding window sample images and the multiple cross-sectional sample images through a contrast learning algorithm; Extracting the first sample features of each sliding window sample image through the trained first feature extraction model, and extracting the second sample features of each cross-sectional sample image through the trained second feature extraction model; Training the first self-attention model, the second self-attention model, and the mutual attention model based on the first sample features and the second sample features through a segmentation training algorithm; Extracting the sample fusion features of the sliding window sample image and the cross-sectional sample image of the same patient through the trained first feature extraction model, second feature extraction model, first self-attention model, second self-attention model, and mutual attention model; Stitching the sample fusion features of the same patient to obtain a sample stitching feature, and training the classification network based on the sample stitching features of each patient.

5. The method for predicting lymph node metastasis of bladder cancer according to claim 4, wherein, Training the first self-attention model, the second self-attention model, and the mutual attention model based on the first sample features and the second sample features through a segmentation training algorithm includes: Respectively inputting the first sample features and the second sample features corresponding to the first training sample data into the first self-attention model and the second self-attention model to obtain the lesion cell sample features output by the first self-attention model and the lesion area sample features output by the second self-attention model; Inputting the lesion cell sample features and the lesion area sample features into the mutual attention model to obtain the sample fusion features output by the mutual attention model; Inputting the sample fusion features into a segmentation model to obtain the mask result output by the segmentation model, where the mask result includes a lesion cell mask and a lesion area mask; Determining the loss value of each pixel point based on the mask result and the annotation information of the sliding window sample image and the cross-sectional sample image in the first training sample data, determining the mean square error loss value based on the loss values of each pixel point, and optimizing the first self-attention model, the second self-attention model, the mutual attention model, and the segmentation model based on the mean square error loss value.

6. The bladder cancer lymph node metastasis prediction method according to claim 4, characterized in that Extracting the sample fusion features of the sliding window sample images and cross-sectional sample images of the same patient through the trained first feature extraction model, second feature extraction model, first self-attention model, second self-attention model, and mutual attention model, includes: Inputting the first sample feature and the second sample feature corresponding to the first training sample data into the trained first self-attention model and second self-attention model respectively, to obtain the lesion cell sample feature output by the first self-attention model and the lesion area sample feature output by the second self-attention model; Inputting the lesion cell sample feature and the lesion area sample feature into the trained mutual attention model, to obtain the sample fusion feature output by the mutual attention model; Correspondingly, splicing the sample fusion features of the same patient to obtain a sample splicing feature, and training the classification network based on the sample splicing features of each patient, includes: Splicing the sample fusion features corresponding to the first training sample data of the same patient to obtain a sample splicing feature; Inputting the sample splicing feature into the classification network, to obtain the classification result output by the classification network; Determining a loss value based on the classification result and the metastasis information of the corresponding patient, and optimizing the classification network based on the loss value.

7. A device for predicting lymph node metastasis of bladder cancer, characterized in that, Includes: An image acquisition module, configured to acquire a plurality of sliding window images corresponding to urine cell section images, and acquire a plurality of cross-sectional images corresponding to nuclear magnetic resonance images; An image combination module, configured to traverse and combine the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, each of the image pairs including one of the sliding window images and one of the cross-sectional images; A feature fusion module, configured to extract the lesion cell features of the sliding window images in the image pair, extract the lesion area features of the cross-sectional images in the image pair, and fuse the lesion cell features and the lesion area features through a mutual attention model to obtain fusion features; wherein, the lesion cell features are extracted by combining a first feature extraction model and a first self-attention model, the lesion area features are extracted by a second feature extraction model and a second self-attention model, and the bladder cancer lymph node metastasis prediction device further includes a first training unit, and the first training unit includes: a sample image combination sub-unit, configured to traverse and combine the sliding window sample images and cross-sectional sample images of the same patient to obtain first training sample data, and randomly combine the sliding window sample images and cross-sectional sample images of different patients to obtain second training sample data; a first classification sub-unit, configured to input the sliding window sample images and cross-sectional sample images in the first training sample data into a first classification model and a second classification model respectively, to obtain the cell category and region category output by the first classification model and the second classification model, the first classification model includes a first feature extraction model and a first classification network, and the second classification model includes a second feature extraction model and a second classification network; a sample data classification sub-unit, configured to determine the first training sample data as true paired data when the cell category and region category corresponding to the sliding window sample images and cross-sectional sample images in the first training sample data are accurately identified as lesion cells and lesion areas, and determine the remaining first training sample data and the second training sample data as false paired data; a cosine similarity determination sub-unit, configured to maximize the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the true paired data, and minimize the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the false paired data, the first sample feature and the second sample feature are respectively extracted by the first feature extraction model and the second feature extraction model; a first optimization sub-unit, configured to determine a symmetric cross-entropy loss value according to the cosine similarity of the true paired data and the cosine similarity of the false paired data, and optimize the first classification model and the second classification model based on the symmetric cross-entropy loss value; A metastasis prediction module, configured to splice the fusion features of each image pair and input them into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph node output by the classification network.

8. A device for predicting lymph node metastasis of bladder cancer, characterized in that, Comprising: One or more processors; A storage device, storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the bladder cancer lymph node metastasis prediction method according to any one of claims 1-6.

9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the bladder cancer lymph node metastasis prediction method according to any one of claims 1-6 when executed by a computer processor.

Citation Information

Patent Citations

  • Medical image tumor segmentation method involving cross-modal attention mechanism

    CN115512110A

  • Tumor risk prediction method and device and storage medium

    CN118919073A