Bladder cancer lymph node metastasis prediction method, device and equipment and storage medium
Through multimodal fusion detection of urine cell slice images and nuclear magnetic resonance images, the characteristics of lesion cells and lesion area are fused using the mutual attention model and input into the classification network for prediction, which solves the problem of low accuracy in the existing technology of single radiological detection of lymph node metastasis, and achieves more accurate prediction of lymph node metastasis in bladder cancer.
Patent Information
- Application Number
- CN202510507958.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the prior art, the accuracy of detecting lymph node metastasis by single radiology is low, and misdiagnosis is prone to missed diagnosis or misdiagnosis, resulting in under-treatment or over-treatment of patients.
The lesion cell characteristics of urine cell slice images and the lesion area characteristics of nuclear magnetic resonance images were jointly predicted, and the two characteristics were fused using the mutual attention model, and a pre-trained classification network was input to output the metastasis results of bladder cancer lymph nodes.
It improves the accuracy of lymph node metastasis prediction, reduces the probability of misdiagnosis or misdiagnosis, and avoids the problem of under-treatment or over-treatment in patients.
Smart Images

Figure CN120013952A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device, equipment and storage medium for predicting bladder cancer lymph node metastasis. Background Art
[0002] Bladder cancer is a common cancer worldwide, and accurate lymph node metastasis is crucial for guiding the treatment and prognosis assessment of bladder cancer patients. Approximately 25% of patients with muscle-invasive bladder cancer have lymph node metastasis, which can lead to a poor prognosis. Patients with lymph node metastasis usually require more aggressive treatment, including preoperative neoadjuvant therapy and expanded pelvic lymph node dissection during surgery, which can improve overall survival and disease-free survival.
[0003] In the existing technology, radiology (such as CT and MRI) is usually used to check preoperative lymph node metastasis of bladder cancer. However, CT and MRI mainly detect malignant lymph nodes based on the size of lymph nodes, which has low accuracy and is prone to missed diagnosis or misdiagnosis, resulting in insufficient or overtreatment of patients. Summary of the invention
[0004] The present application provides a method, device, equipment and storage medium for predicting bladder cancer lymph node metastasis, which can predict the metastasis results of bladder cancer lymph nodes by combining the diseased cell characteristics of urine cell slice images and the lesion area characteristics of magnetic resonance imaging, thereby solving the problem of low accuracy of lymph node metastasis detected by single radiology in the prior art, reducing the probability of misdiagnosis or missed diagnosis, and avoiding the problem of insufficient or excessive treatment of patients.
[0005] In a first aspect, the present application provides a method for predicting bladder cancer lymph node metastasis, comprising: Acquire multiple sliding window images corresponding to the urine cell slice image, and acquire multiple cross-sectional images corresponding to the nuclear magnetic resonance image; Traversing and combining the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, each of the image pairs comprising one sliding window image and one cross-sectional image; Extracting diseased cell features of the sliding window image in the image pair, extracting lesion region features of the cross-sectional image in the image pair, and fusing the diseased cell features and the lesion region features through a mutual attention model to obtain fused features; The fusion features of each of the image pairs are spliced and then input into a pre-trained classification network to obtain the metastasis result of bladder cancer lymph nodes output by the classification network.
[0006] In a second aspect, the present application provides a device for predicting bladder cancer lymph node metastasis, comprising: An image acquisition module is configured to acquire a plurality of sliding window images corresponding to the urine cell slice image and to acquire a plurality of cross-sectional images corresponding to the nuclear magnetic resonance image; An image combination module, configured to traverse and combine the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, each of the image pairs comprising one sliding window image and one cross-sectional image; a feature fusion module, configured to extract the diseased cell features of the sliding window image in the image pair, extract the lesion area features of the cross-sectional image in the image pair, and fuse the diseased cell features and the lesion area features through a mutual attention model to obtain a fusion feature; The metastasis prediction module is configured to splice the fusion features of each of the image pairs and input the spliced features into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph nodes output by the classification network.
[0007] In a third aspect, the present application provides a device for predicting bladder cancer lymph node metastasis, comprising: One or more processors; a storage device storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method for predicting bladder cancer lymph node metastasis as described in the first aspect.
[0008] In a fourth aspect, the present application provides a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute the method for predicting bladder cancer lymph node metastasis as described in the first aspect.
[0009] In the present application, multiple sliding window images corresponding to urine cell slice images are obtained to obtain multiple cross-sectional images corresponding to nuclear magnetic resonance images; multiple sliding window images and multiple cross-sectional images are traversed and combined to obtain multiple image pairs, each image pair includes a sliding window image and a cross-sectional image; the diseased cell features of the sliding window image in the image pair are extracted, the lesion area features of the cross-sectional image in the image pair are extracted, and the diseased cell features and the lesion area features are fused through a mutual attention model to obtain fused features; the fused features of each image pair are spliced and input into a pre-trained classification network to obtain the metastasis results of bladder cancer lymph nodes output by the classification network. Through the above-mentioned technical means, the correlation features between the diseased cell features of each sliding window image in the urine cell slice image and the lesion area features of each cross-sectional image in the magnetic resonance image can be extracted through the mutual attention model, and the correlation features between each sliding window image and each cross-sectional image can be spliced to obtain the correlation features between the urine cell slice image and the magnetic resonance image. The correlation features between the urine cell slice image and the magnetic resonance image are analyzed through the classification network to output the metastasis results of the bladder cancer lymph nodes, realizing multimodal fusion detection of lymph node metastasis to accurately detect the metastasis results of bladder cancer lymph nodes, solving the problem of low accuracy of lymph node metastasis detected by single radiology in the prior art, reducing the probability of misdiagnosis or missed diagnosis, and avoiding the problem of insufficient or excessive treatment of patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a flow chart of a method for predicting bladder cancer lymph node metastasis provided in an embodiment of the present application; Figure 2 is a flow chart of extracting diseased cell features of a sliding window image provided by an embodiment of the present application; Figure 3 is a flow chart of extracting lesion area features of a cross-sectional image provided by an embodiment of the present application; Figure 4 It is a schematic diagram of the prediction process of bladder cancer lymph node metastasis results provided in the examples of the present application; Figure 5 is a flow chart of the model training process provided in the embodiment of the present application; Figure 6 is a flowchart of training a first feature extraction model and a second feature extraction model provided in an embodiment of the present application; Figure 7 It is a flowchart of training a first self-attention model, a second self-attention model, and a mutual attention model provided in an embodiment of the present application; Figure 8 It is a schematic diagram of the structure of a bladder cancer lymph node metastasis prediction device provided in an embodiment of the present application; Fig. 9It is a structural schematic diagram of a bladder cancer lymph node metastasis prediction device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0011] In order to make the purpose, technical scheme and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for the convenience of description, only part of the present application is shown in the accompanying drawings, but not all of the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow chart describes each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc.
[0012] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0013] In a more relevant implementation, radiology (such as CT and MRI) is usually used to examine preoperative lymph node metastasis of bladder cancer. Specifically, bladder radiographic images are collected through radiological detection technology, and lymph nodes in bladder radiographic images are segmented through image segmentation technology. The size of the lymph nodes is calculated based on the segmented lymph node pixels, and whether it is a malignant lymph node is determined based on the size of the lymph nodes. However, the accuracy of detecting malignant lymph nodes based on the size of the lymph nodes is low, and it is easy to miss or misdiagnose, resulting in insufficient or excessive treatment of patients.
[0014] To solve the above problems, this embodiment provides a method for predicting bladder cancer lymph node metastasis, which can predict the metastasis results of bladder cancer lymph nodes by combining the diseased cell characteristics of urine cell slice images and the lesion area characteristics of magnetic resonance imaging, thereby reducing the probability of misdiagnosis or missed diagnosis and avoiding the problem of insufficient or excessive treatment of patients.
[0015] The bladder cancer lymph node metastasis prediction method provided in this embodiment can be performed by a bladder cancer lymph node metastasis prediction device, which can be implemented by software and / or hardware, and the bladder cancer lymph node metastasis prediction device can be composed of two or more physical entities, or can be composed of one physical entity. For example, the bladder cancer lymph node metastasis prediction device can be an intelligent terminal with strong processing capabilities such as a computer and a server. Among them, the server can be implemented by an independent server or a server cluster composed of multiple servers.
[0016] The bladder cancer lymph node metastasis prediction device is installed with at least one type of operating system, and the bladder cancer lymph node metastasis prediction device can install at least one application based on the operating system, and the application can be an application that comes with the operating system, or an application downloaded from a third-party device or server. In this embodiment, the bladder cancer lymph node metastasis prediction device is installed with at least an application that can execute the bladder cancer lymph node metastasis prediction method.
[0017] For ease of understanding, this embodiment is described by taking a server as an example of a subject that executes the method for predicting bladder cancer lymph node metastasis.
[0018] Figure 1 A flowchart of a method for predicting bladder cancer lymph node metastasis provided in an embodiment of the present application is given. Figure 1 The method for predicting bladder cancer lymph node metastasis specifically includes: S110, acquiring a plurality of sliding window images corresponding to the urine cell slice image, and acquiring a plurality of cross-sectional images corresponding to the nuclear magnetic resonance image.
[0019] The urine cell slice image is a digital image obtained by scanning a patient's urothelial cell liquid-based slice with a digital scanner. Exemplarily, the urine cell slice image can be cropped with a sliding window of a preset size according to a preset step length, and the image obtained by each sliding window cropping is used as a sliding window image. Since the sliding window image is cropped by a sliding window of a fixed size, each sliding window image has the same size.
[0020] The nuclear magnetic resonance image is a three-dimensional image obtained by imaging the patient's bladder using the principle of nuclear magnetic resonance. The nuclear magnetic resonance image is a three-dimensional image formed by superimposing multiple cross-sectional images, and multiple cross-sectional images can be obtained in the nuclear magnetic resonance image.
[0021] S120 , traversing and combining the multiple sliding window images and the multiple cross-sectional images to obtain a plurality of image pairs, each image pair including a sliding window image and a cross-sectional image.
[0022] Exemplarily, all sliding window images cropped from the urine cell slice image and all cross-sectional images acquired from the nuclear magnetic resonance image are combined in pairs, and each sliding window image is combined with each cross-sectional image to form an image pair, and finally multiple sets of image pairs are obtained. For example, N sliding window images are cropped from the urine cell slice image, and M cross-sectional images are acquired from the nuclear magnetic resonance image. The N sliding window images and the M cross-sectional images are traversed and combined to obtain N*M image pairs.
[0023] S130, extracting the diseased cell features of the sliding window image in the image pair, extracting the lesion area features of the cross-sectional image in the image pair, and fusing the diseased cell features and the lesion area features through a mutual attention model to obtain a fusion feature.
[0024] Exemplarily, the lesion cell features of the sliding window image in each image pair are extracted, and the lesion area features of the cross-sectional image in each image pair are extracted. The correlation features between the lesion cell features and the lesion area features corresponding to each image pair are extracted through the pre-trained mutual attention model, and finally the fusion features of the lesion cell features and the lesion area features corresponding to each image pair are output. It can be understood that the lesion cell features are the morphological features of the lesion cells in the corresponding sliding window image, including cell color, cell membrane shape, and cell nuclear area. The lesion area features are the color, position, and shape of the lesion area (tumor mass) in the corresponding cross-sectional image. The lesion cell features are the feature information of the tumor mass at the microscopic level, and the lesion area features are the feature information of the tumor mass at the macroscopic level. For the tumor mass in the same area, lesion cell features and lesion area features with high similarity can be extracted. When searching for the correlation features between the lesion cell features and the lesion area features through the mutual attention model, it is equivalent to using the lesion cell features at the microscopic level to verify whether the lesion area detection at the macroscopic level is accurate, thereby realizing the accurate positioning of the tumor mass, and laying a solid data foundation for the subsequent accurate prediction of the metastasis results based on the position information of the tumor mass.
[0025] Since each image pair contains repeated sliding window images and cross-sectional images, if feature extraction is performed on the sliding window images and cross-sectional images of each image pair, the features of the same sliding window image will be repeatedly extracted, affecting the efficiency of feature extraction. In this regard, feature extraction can be performed on each sliding window image and each cross-sectional image, and finally, based on the sliding window images and cross-sectional images in the image pair, the corresponding diseased cell features and lesion area features are fused to obtain the fused features of the image pair.
[0026] In one embodiment, the diseased cell features of the sliding window image and the lesion area features of the cross-sectional image can be extracted by a convolutional neural network or an attention network.
[0027] In another embodiment, the features of diseased cells in the sliding window image can be extracted by jointly using a convolutional neural network and an attention network. Figure 2 FIG. 1 is a flow chart of extracting diseased cell features of a sliding window image provided by an embodiment of the present application. Figure 2 As shown, the step of extracting the diseased cell features of the sliding window image specifically includes S1301-S1302: S1301: Input the sliding window image into a pre-trained first deep feature extraction model to obtain a first feature of the sliding window image.
[0028] Among them, the first deep feature extraction model is a convolutional neural network. The sliding window image is input into the first deep feature extraction model, and the morphological features of the diseased cells in the sliding window image are extracted as the first features through the first deep feature model.
[0029] S1302. Input the first feature into a pre-trained first self-attention model to obtain the diseased cell feature output by the first self-attention model.
[0030] Exemplarily, the first feature is converted into a feature sequence and input into a first self-attention model, and the self-attention feature in the feature sequence is extracted by the first self-attention model as a diseased cell feature.
[0031] This embodiment extracts the diseased cell features of the sliding window image by combining a convolutional neural network and a self-attention network, so as to effectively extract the local features and global features of the sliding window image, thereby improving the extraction accuracy of the diseased cell features.
[0032] Similarly, the lesion area features of the cross-sectional image can be extracted by jointly using a convolutional neural network and an attention network. Figure 3 FIG. 1 is a flow chart of extracting the features of the lesion area of a cross-sectional image provided by an embodiment of the present application. Figure 3 As shown, the step of extracting the lesion area features of the cross-sectional image specifically includes S1303-S1305: S1303, reducing the cross-sectional image to a preset size.
[0033] For example, the resolution of the cross-sectional image is relatively large. If a convolutional neural network is used to directly extract its features, the computing power required is large and the operation time is long. In this regard, the cross-sectional image can be reduced to a preset size before the corresponding lesion area features are extracted.
[0034] S1304 , input the reduced cross-sectional image into a pre-trained second deep feature extraction model to obtain a second feature of the cross-sectional image.
[0035] Among them, the second deep feature extraction model is a convolutional neural network. The reduced cross-sectional image is input into the pre-trained second deep feature extraction model, and the second deep feature extraction model is used to extract features such as color, position, and shape of the lesion area in the cross-sectional image as the second feature.
[0036] S1305. Input the second feature into a pre-trained second self-attention model to obtain the lesion area feature output by the second self-attention model.
[0037] Exemplarily, the second feature is converted into a feature sequence and input into a second self-attention model, and the self-attention feature in the feature sequence is extracted by the second self-attention model as the lesion area feature.
[0038] This embodiment extracts the lesion area features of the cross-sectional image by combining the convolutional neural network and the self-attention network to effectively extract the local features and global features of the cross-sectional image, thereby improving the extraction accuracy of the lesion area features.
[0039] After extracting the lesion cell features of each sliding window image and the lesion area features of each cross-sectional image, the lesion cell features of a sliding window image and the lesion area features of a cross-sectional image can be spliced into a feature sequence. The spliced feature sequence is input into the pre-trained mutual attention model, and the mutual attention model determines the correlation weight between the lesion cell features and the lesion area features. The correlation weight is weighted and summed with the corresponding input feature sequence to obtain the fusion feature. It can be understood that the higher the feature similarity between the lesion cell features and the lesion area features, the greater the correlation weight, so the fusion feature finally output can be used to characterize the correlation features between the lesion cell features and the lesion area features.
[0040] S140, splicing the fusion features of each image pair and inputting them into a pre-trained classification network to obtain the metastasis result of bladder cancer lymph nodes output by the classification network.
[0041] Exemplarily, the fused features of each image pair are spliced to obtain a spliced feature, which is input into a pre-trained classification network. The classification network locates the position of the bladder cancer lymph nodes based on each fused feature in the spliced feature, and finally outputs a metastasis result of whether the bladder cancer lymph nodes have metastasized based on whether the position information of the bladder cancer lymph nodes has moved.
[0042] In order to more intuitively understand the process of predicting the metastasis results of bladder cancer lymph nodes, this example uses Figure 4 The prediction process of bladder cancer lymph node metastasis is shown as an example and described in detail. Figure 4As shown, multiple sliding window images are obtained in the urine cell slice image, and multiple cross-sectional images are obtained in the nuclear magnetic resonance image. The sliding window image is input into the first feature extraction model to obtain the first feature output by the first feature extraction model, and the first feature is input into the first self-attention model to obtain the lesion cell feature output by the first self-attention model. The cross-sectional image is input into the second feature extraction model to obtain the second feature output by the second feature extraction model, and the second feature is input into the second self-attention model to obtain the lesion area feature output by the second self-attention model. The lesion cell feature and the lesion area feature are spliced and input into the mutual attention model to obtain the fusion feature output by the mutual attention model, and each lesion cell feature and each lesion area feature The fusion feature corresponding to the feature is spliced to obtain the spliced feature, and the spliced feature is input into the classification network to obtain the classification network Output whether the bladder cancer lymph node has metastasized.
[0043] Optional, reference Figure 4 , a lesion cell localization model and a lesion region localization model can also be added. After the first self-attention model outputs the lesion cell features, the lesion cell features are input into the lesion cell localization model, and the lesion cell localization model detects and outputs the position of the lesion cell in the sliding window image. And, after the second self-attention model outputs the lesion region features, the lesion region features are input into the lesion region localization model, and the lesion region localization model detects and outputs the position of the lesion region in the cross-sectional image. The position of the lesion cell in the sliding window image and the position of the lesion region in the cross-sectional image can be used to assist doctors in judging whether the prediction results of the classification network are accurate, providing quality assurance for clinical diagnosis.
[0044] In one embodiment, the first deep feature extraction model, the second deep feature extraction model, the first self-attention model, the second self-attention model, the mutual attention model and the classification network can form a network model. During the training phase, the network model is trained with a large number of sample images so that the network model has the function of predicting the metastasis results of bladder cancer lymph nodes.
[0045] In another embodiment, the training process of the first deep feature extraction model, the second deep feature extraction model, the first self-attention model, the second self-attention model, the mutual attention model and the classification network can be divided into three training stages. In the first training stage, the first deep feature extraction model and the second deep feature extraction model are trained first, and in the second training stage, the first self-attention model, the second self-attention model and the mutual attention model are trained, and finally the classification network is trained. Optionally, Figure 5 Flowchart of the model training process provided by the embodiment of the present application. Figure 5 As shown, the steps of the model training process specifically include S210-S260: S210 , dividing the urine cell slice sample image of each patient into a plurality of sliding window sample images, and obtaining a plurality of cross-sectional sample images corresponding to the nuclear magnetic resonance sample image of each patient.
[0046] Exemplarily, urine cell slice images and nuclear magnetic resonance images of multiple patients are obtained as urine cell slice sample images and nuclear magnetic resonance sample images, respectively, and multiple sliding window sample images are cropped out from each urine cell slice sample image through a sliding window, and each cross-sectional image in each nuclear magnetic resonance image is used as the corresponding cross-sectional sample image.
[0047] S220 , based on the multiple sliding window sample images and the multiple cross-section sample images, train the first feature extraction model and the second feature extraction model through a contrast learning algorithm.
[0048] Exemplarily, multiple sliding window sample images and multiple cross-sectional sample images are divided into positive sample data and negative sample data, and based on the positive sample data and the negative sample data, a first feature extraction model and a second feature extraction model are trained by a contrast learning algorithm.
[0049] Optional, Figure 6 : is a flowchart of training the first feature extraction model and the second feature extraction model provided in the embodiment of the present application. Figure 6 As shown, the steps of training the first feature extraction model and the second feature extraction model specifically include S2201-S2205: S2201. Traversingly combining sliding window sample images and cross-sectional sample images of the same patient to obtain first training sample data, and randomly combining sliding window sample images and cross-sectional sample images of different patients to obtain second training sample data.
[0050] Exemplarily, multiple sliding window sample images corresponding to urine cell slice sample images of the same patient and multiple cross-sectional sample images corresponding to nuclear magnetic resonance sample images are traversed and combined to obtain multiple first image pairs, and the multiple first image pairs are used as first training sample data, each of which includes a sliding window sample image and a cross-sectional sample image belonging to the same patient. Sliding window sample images and cross-sectional sample images of different patients are randomly combined to obtain multiple second image pairs, and the multiple second image pairs are used as second training sample data.
[0051] S2202. Input the sliding window sample image and the cross-sectional sample image in the first training sample data into the first classification model and the second classification model respectively to obtain the cell category and the region category output by the first classification model and the second classification model. The first classification model includes a first feature extraction model and a first classification network, and the second classification model includes a second feature extraction model and a second classification network.
[0052] Exemplarily, the sliding window sample image in the first training sample data is input into the first classification model, the shape features of the diseased cells in the sliding window sample image are extracted as the first sample features by the first feature extraction model in the first classification model, and the cell category in the sliding window sample image is identified as a diseased cell or a non-lesioned cell based on the first sample features by the first classification network in the first classification model. The cross-sectional sample image in the first training sample data is input into the second classification model, the features of the lesion area in the cross-sectional sample image are extracted as the second sample features by the second feature extraction model in the second classification model, and the distinction category in the cross-sectional sample image is identified as a lesion area or a non-lesion area based on the second sample features by the second classification network in the second classification model.
[0053] S2203. When the cell categories and region categories corresponding to the sliding window sample images and the cross-sectional sample images in the first training sample data are accurately identified as diseased cells and lesion regions, the first training sample data are determined as true paired data, and the remaining first training sample data and second training sample data are determined as incorrect paired data.
[0054] Exemplarily, the urine cell slice sample image is annotated with a corresponding diseased cell detection frame, and the diseased cell detection frame marks the diseased cells in the urine cell slice sample image. Whether there are diseased cells in the sliding window sample image can be determined based on the pixel coordinates of the diseased cell detection frame annotated on the urine cell slice sample image. If the pixel coordinates of the diseased cell detection frame overlap with the pixel coordinates of the sliding window sample image, it is determined that there are diseased cells in the sliding window sample image, otherwise it is determined that there are no diseased cells in the sliding window sample image. After the first classification model outputs the cell category of the sliding window sample image, if the cell category is a diseased cell and the sliding window sample image does have diseased cells, it is determined that the first classification model accurately identifies the presence of diseased cells in the sliding window sample image.
[0055] Each cross-sectional sample image of the nuclear magnetic resonance sample image is annotated with a corresponding lesion area detection frame. After the second classification model outputs the area category of the cross-sectional sample image, if the area category is a lesion area and the cross-sectional sample image is annotated with a lesion area detection frame, it is determined that the second classification model accurately identifies the existence of a lesion area in the cross-sectional sample image.
[0056] The first training sample data can be determined as true paired data if and only if the first classification model accurately identifies the presence of diseased cells in the sliding window sample image and the second classification model accurately identifies the presence of lesion areas in the cross-sectional sample image, thereby characterizing that the sliding window sample image and the cross-sectional sample image in the first training sample data are highly correlated positive sample data.
[0057] When the first classification model misidentifies the sliding window sample image and the cross-sectional sample image, or there is no diseased cell in the sliding window sample image or there is no lesion area in the cross-sectional sample image, the first training sample data is determined as incorrectly paired data, thereby characterizing the sliding window sample image and the cross-sectional sample image in the first training sample data as negative sample data with low correlation. The sliding sample image and the cross-sectional sample image in the second training sample data belong to different patients, and their correlation is even lower, so the second training sample data is also regarded as incorrectly paired data.
[0058] S2204, maximizing the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the true paired data, and minimizing the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the false paired data, the first sample feature and the second sample feature are extracted by the first feature extraction model and the second feature extraction model respectively.
[0059] Exemplarily, the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the true paired data is calculated, and the cosine similarity is maximized by the loss function. The cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the false paired data is calculated, and the cosine similarity is minimized by the loss function.
[0060] S2205. Determine a symmetric cross entropy loss value according to the cosine similarity of the true paired data and the cosine similarity of the incorrect paired data, and optimize the first classification model and the second classification model based on the symmetric cross entropy loss value.
[0061] Exemplarily, based on the cross entropy loss function and the cosine similarity of each true paired data and the cosine similarity of the wrong paired data, a symmetric cross entropy loss value is determined. Based on the symmetric cross entropy loss value, the model parameters of the first classification network and the first feature extraction model in the first classification model and the second classification network and the second feature extraction model in the second classification model are reversely optimized respectively.
[0062] The process of steps S2201-S2205 is to train the first classification model and the second classification model once. After training the first classification model and the second classification model once, the first training sample data is input into the first classification model and the second classification model again, and the real paired data and the wrong paired data are re-screened based on the cell category and the region category output by the first classification model and the second classification. The symmetric cross entropy loss value is calculated based on the latest screened real paired data and the wrong paired data, and the first classification model and the second classification model are optimized again based on the symmetric cross entropy loss value. When the first classification model and the second classification model meet the preset convergence conditions, it is confirmed that the training of the first classification model and the second classification model is completed. The first feature extraction model and the second feature extraction model obtained after the training are used for actual transfer result prediction and model training in subsequent stages.
[0063] S230, extracting the first sample feature of each sliding window sample image by using the trained first feature extraction model, and extracting the second sample feature of each cross-section sample image by using the trained second feature extraction model.
[0064] Exemplarily, each sliding window sample image is input into a trained first feature extraction model to obtain a first sample feature output by the first feature extraction model, and each cross-sectional sample image is input into a trained second feature extraction model to obtain a second sample feature output by the second feature extraction model, so as to train the first self-attention model, the second self-attention model and the mutual attention model in the second training stage based on the first sample features and the second sample features.
[0065] S240. Based on the first sample features and the second sample features, train the first self-attention model, the second self-attention model and the mutual attention model through a segmentation training algorithm.
[0066] Exemplarily, the implementation process of the segmentation training algorithm is to perform cell segmentation and region segmentation on the first sample features and the second sample features based on a network model composed of a first self-attention model, a second self-attention model, a mutual attention model and a segmentation model, calculate a loss value based on the segmentation result, and optimize the model parameters of the network model based on the loss value.
[0067] Optional, Figure 7 Flowchart of training the first self-attention model, the second self-attention model, and the mutual attention model provided in the embodiment of the present application. Figure 7 As shown, the steps of training the first self-attention model, the second self-attention model, and the mutual attention model specifically include S2401-S2404: S2401. Input the first sample feature and the second sample feature corresponding to the first training sample data into the first self-attention model and the second self-attention model respectively to obtain the diseased cell sample feature output by the first self-attention model and the lesion area sample feature output by the second self-attention model.
[0068] Exemplarily, the first sample features and the second sample features of the first training sample data, i.e., the sliding window sample image and the cross-sectional sample image belonging to the same patient, are input into the first self-attention model and the second self-attention model separately. The first self-attention model extracts the self-attention features of the first sample features as the lesion cell sample features, and the second self-attention model extracts the self-attention features of the second sample features as the lesion area sample features.
[0069] S2402. Input the sample features of the diseased cells and the sample features of the lesion area into the mutual attention model to obtain the sample fusion features output by the mutual attention model.
[0070] Exemplarily, the diseased cell sample features and the lesion area sample features are spliced and input into the mutual attention model. The mutual attention model extracts the correlation between the diseased cell sample features and the lesion area sample features, and weights the spliced features based on the correlation to obtain the sample fusion features.
[0071] S2403. Input the sample fusion features into the segmentation model to obtain the mask result output by the segmentation model, where the mask result includes the diseased cell mask and the lesion area mask.
[0072] Exemplarily, the sample fusion feature is input into the segmentation model, and the segmentation model simultaneously detects the diseased cells in the sliding window sample image and the lesion area in the cross-sectional sample image based on the sample fusion feature, and outputs a diseased cell mask and a lesion area mask. The diseased cell mask includes each pixel point of the diseased cell, and the lesion area mask includes each pixel point of the lesion area.
[0073] S2404. Determine the loss value of each pixel based on the mask result and the annotation information of the sliding window sample image and the cross-sectional sample image in the first training sample data, determine the mean square error loss value based on the loss value of each pixel, and optimize the first self-attention model, the second self-attention model, the mutual attention model and the segmentation model based on the mean square error loss value.
[0074] Exemplarily, the urine cell slice sample image is annotated with a diseased cell detection frame, and the pixel points of the diseased cells in the sliding window sample image are determined according to the pixel coordinates of the diseased cell detection frame, and the loss value of each pixel point is calculated according to the pixel points of the diseased cells in the sliding window sample image and the pixel points in the diseased cell mask. The cross-sectional sample image is annotated with a lesion area detection frame, and the pixel points of the lesion area in the cross-sectional sample image are determined according to the lesion area detection frame, and the loss value of each pixel point is calculated according to the pixel points of the lesion area in the cross-sectional sample image and the pixel points in the lesion area mask. The average value is calculated based on the loss values of the above-mentioned individual pixels, and the variance is calculated based on the loss values of the individual pixels and the average value as the mean square error loss value, and the segmentation model, the mutual attention model, the first self-attention model, and the second self-attention model are reversely optimized through the mean square error loss value.
[0075] Steps S2401-S2404 are a training of the segmentation model, the mutual attention model, the first self-attention model and the second self-attention model. After the segmentation model, the mutual attention model, the first self-attention model and the second self-attention model are trained using the first sample features and the second sample features corresponding to all the first training sample data, the training ends. The mutual attention model, the first self-attention model and the second self-attention model obtained after the training are used to predict the transfer results and to train the classification network in the subsequent third training stage.
[0076] S250, extracting sample fusion features of the sliding window sample image and the cross-sectional sample image of the same patient through the trained first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model and the mutual attention model.
[0077] Exemplarily, the third training stage is only used to train the classification network so that the trained classification network has the function of predicting whether the lymph nodes are metastatic. Therefore, the sample data used in the third training stage are the sliding window sample images and the cross-sectional sample images of the same patient. The first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model and the mutual attention model have been trained in the first training stage and the second training stage. The first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model and the mutual attention model can be directly used to extract the sample fusion features of each first training sample data obtained by traversing and combining each sliding window sample image and each cross-sectional sample image of the same patient, so as to use the sample fusion features of each first training sample data as the sample data of the third training stage, which is conducive to improving the training efficiency of the classification network.
[0078] Optionally, the first sample feature and the second sample feature corresponding to the first training sample data are respectively input into the trained first self-attention model and the second self-attention model to obtain the sample feature of the lesion cell output by the first self-attention model and the sample feature of the lesion area output by the second self-attention model; the sample feature of the lesion cell and the sample feature of the lesion area are input into the trained mutual attention model to obtain the sample fusion feature output by the mutual attention model. For example, the first sample feature of the first training sample data extracted by the trained first feature extraction model is input into the trained first self-attention model, and the second sample feature of the first training sample data extracted by the trained first feature extraction model is input into the trained second self-attention model to obtain the sample feature of the lesion cell output by the first self-attention model and the sample feature of the lesion area output by the second self-attention model. The sample feature of the lesion cell and the sample feature of the lesion area corresponding to the first training sample data are spliced and input into the trained mutual attention model to obtain the sample fusion feature output by the mutual attention model.
[0079] S260, concatenating the sample fusion features of the same patient to obtain a sample concatenation feature, and training a classification network based on the sample concatenation features of each patient.
[0080] Exemplarily, each sample fusion feature output by the trained mutual attention model is obtained, each sample fusion feature of the same patient is spliced to obtain a sample splicing feature, the sample splicing feature is input into the classification network, and the model parameters of the classification network are optimized based on the classification results output by the classification network.
[0081] Optionally, the sample fusion features corresponding to each first training sample data of the same patient are spliced to obtain a sample splicing feature; the sample splicing feature is input into the classification network to obtain the classification result output by the classification network; the loss value is determined based on the classification result and the metastasis information of the corresponding patient, and the classification network is optimized based on the loss value. For example, after the trained mutual attention model outputs the sample fusion features corresponding to each first training sample data of patient A, the sample fusion features corresponding to each first training sample data of patient A are spliced to obtain the sample splicing feature of patient A, and the sample splicing feature of patient A is input into the classification network to obtain the metastasis result of bladder cancer lymph nodes of patient A output by the classification network. The loss value is determined based on the metastasis result of patient A output by the classification network and the actual metastasis result of patient A, and the model parameters of the classification network are reversely optimized based on the loss value. After the classification network is trained based on the sample splicing features of each patient, a trained classification network is obtained.
[0082] After the classification network is trained, the patient's urine cell slice images and magnetic resonance imaging can be processed based on the trained first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model, the mutual attention model and the classification network to predict whether the patient's bladder cancer lymph nodes have metastasized.
[0083] In summary, the bladder cancer lymph node metastasis prediction method provided in the embodiment of the present application obtains multiple sliding window images corresponding to the urine cell slice image and multiple cross-sectional images corresponding to the nuclear magnetic resonance image; traverses and combines the multiple sliding window images and the multiple cross-sectional images to obtain multiple image pairs, each image pair includes a sliding window image and a cross-sectional image; extracts the diseased cell features of the sliding window image in the image pair, extracts the lesion area features of the cross-sectional image in the image pair, and fuses the diseased cell features and the lesion area features through a mutual attention model to obtain a fused feature; splices the fused features of each image pair and inputs them into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph nodes output by the classification network. Through the above-mentioned technical means, the correlation features between the diseased cell features of each sliding window image in the urine cell slice image and the lesion area features of each cross-sectional image in the magnetic resonance image can be extracted through the mutual attention model, and the correlation features between each sliding window image and each cross-sectional image can be spliced to obtain the correlation features between the urine cell slice image and the magnetic resonance image. The correlation features between the urine cell slice image and the magnetic resonance image are analyzed through the classification network to output the metastasis results of the bladder cancer lymph nodes, realizing multimodal fusion detection of lymph node metastasis to accurately detect the metastasis results of bladder cancer lymph nodes, solving the problem of low accuracy of lymph node metastasis detected by single radiology in the prior art, reducing the probability of misdiagnosis or missed diagnosis, and avoiding the problem of insufficient or excessive treatment of patients.
[0084] Based on the above embodiments, Figure 8 This is a schematic diagram of the structure of a bladder cancer lymph node metastasis prediction device provided in an embodiment of the present application. Figure 8 The bladder cancer lymph node metastasis prediction device provided in this embodiment specifically includes: an image acquisition module 31, an image combination module 32, a feature fusion module 33 and a metastasis prediction module 34.
[0085] The image acquisition module 31 is configured to acquire a plurality of sliding window images corresponding to the urine cell slice image and acquire a plurality of cross-sectional images corresponding to the nuclear magnetic resonance image; An image combination module 32 is configured to traverse and combine the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, each image pair including a sliding window image and a cross-sectional image; The feature fusion module 33 is configured to extract the lesion cell features of the sliding window image in the image pair, extract the lesion area features of the cross-sectional image in the image pair, and fuse the lesion cell features and the lesion area features through a mutual attention model to obtain a fusion feature; The metastasis prediction module 34 is configured to splice the fusion features of each image pair and input them into a pre-trained classification network to obtain the metastasis result of bladder cancer lymph nodes output by the classification network.
[0086] Based on the above embodiment, the feature fusion module 33 includes: a first feature extraction unit, configured to input the sliding window image into a pre-trained first deep feature extraction model to obtain the first feature of the sliding window image; a first self-attention unit, configured to input the first feature into a pre-trained first self-attention model to obtain the diseased cell feature output by the first self-attention model.
[0087] Based on the above embodiment, the feature fusion module 33 includes: an image reduction unit, configured to reduce the cross-sectional image to a preset size; a second feature extraction unit, configured to input the reduced cross-sectional image into a pre-trained second deep feature extraction model to obtain a second feature of the cross-sectional image; a second self-attention unit, configured to input the second feature into a pre-trained second self-attention model to obtain a lesion area feature output by the second self-attention model.
[0088] On the basis of the above-mentioned embodiment, the bladder cancer lymph node metastasis prediction device further includes a model training module, which includes: a sample image acquisition unit, configured to train to divide the urine cell slice sample image of each patient into a plurality of sliding window sample images, and to obtain a plurality of cross-sectional sample images corresponding to the nuclear magnetic resonance sample image of each patient; a first training unit, configured to train a first feature extraction model and a second feature extraction model based on the plurality of sliding window sample images and the plurality of cross-sectional sample images through a comparative learning algorithm; a first sample feature acquisition unit, configured to extract a first sample feature of each sliding window sample image through the trained first feature extraction model, and to extract a first sample feature of each sliding window sample image through the trained second feature extraction model. The model extracts the second sample feature of each cross-sectional sample image; the second training unit is configured to train the first self-attention model, the second self-attention model and the mutual attention model through a segmentation training algorithm based on the first sample feature and the second sample feature; the second sample feature acquisition unit is configured to extract the sample fusion features of the sliding window sample image and the cross-sectional sample image of the same patient through the trained first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model and the mutual attention model; the third training unit is configured to splice the sample fusion features of the same patient to obtain the sample splicing features, and train the classification network based on the sample splicing features of each patient.
[0089] On the basis of the above embodiment, the first training unit includes: a sample image combination subunit, which is configured to traverse and combine the sliding window sample images and cross-sectional sample images of the same patient to obtain the first training sample data, and randomly combine the sliding window sample images and cross-sectional sample images of different patients to obtain the second training sample data; a first classification subunit, which is configured to input the sliding window sample images and cross-sectional sample images in the first training sample data into the first classification model and the second classification model respectively, to obtain the cell category and the region category output by the first classification model and the second classification model, the first classification model includes a first feature extraction model and a first classification network, and the second classification model includes a second feature extraction model and a second classification network; the sample data classification subunit is configured to classify the cell category and the region category corresponding to the sliding window sample images and the cross-sectional sample images in the first training sample data When the diseased cells and lesion areas are accurately identified, the first training sample data is determined as the true paired data, and the remaining first training sample data and second training sample data are determined as the wrong paired data; the cosine similarity determination subunit is configured to maximize the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the true paired data, and minimize the cosine similarity between the first sample feature of the sliding window sample image and the second sample feature of the cross-sectional sample image in the wrong paired data, and the first sample feature and the second sample feature are extracted by the first feature extraction model and the second feature extraction model respectively; the first optimization subunit is configured to determine the symmetric cross entropy loss value according to the cosine similarity of the true paired data and the cosine similarity of the wrong paired data, and optimize the first classification model and the second classification model based on the symmetric cross entropy loss value.
[0090] On the basis of the above embodiment, the second training unit includes: a first self-attention feature extraction subunit, which is configured to input the first sample feature and the second sample feature corresponding to the first training sample data into the first self-attention model and the second self-attention model respectively, to obtain the lesion cell sample feature output by the first self-attention model and the lesion area sample feature output by the second self-attention model; a second mutual attention feature extraction subunit, which is configured to input the lesion cell sample feature and the lesion area sample feature into the mutual attention model, to obtain the sample fusion feature output by the mutual attention model; a segmentation processing subunit, which is configured to input the sample fusion feature into the segmentation model, to obtain the mask result output by the segmentation model, and the mask result includes the lesion cell mask and the lesion area mask; a second optimization subunit, which is configured to determine the loss value of each pixel based on the mask result and the annotation information of the sliding window sample image and the cross-sectional sample image in the first training sample data, determine the mean square error loss value based on the loss value of each pixel, and optimize the first self-attention model, the second self-attention model, the mutual attention model and the segmentation model based on the mean square error loss value.
[0091] On the basis of the above embodiment, the second sample feature acquisition unit includes: a second self-attention feature extraction subunit, which is configured to input the first sample feature and the second sample feature corresponding to the first training sample data into the trained first self-attention model and the second self-attention model respectively, to obtain the lesion cell sample feature output by the first self-attention model and the lesion area sample feature output by the second self-attention model; a second mutual attention feature extraction subunit, which is configured to input the lesion cell sample feature and the lesion area sample feature into the trained mutual attention model, to obtain the sample fusion feature output by the mutual attention model; accordingly, the third training unit includes: a feature splicing subunit, which is configured to splice the sample fusion features corresponding to the first training sample data of the same patient to obtain the sample splicing feature; a second classification subunit, which is configured to input the sample splicing feature into the classification network, to obtain the classification result output by the classification network; a third optimization subunit, which is configured to determine the loss value based on the classification result and the transfer information of the corresponding patient, and optimize the classification network based on the loss value.
[0092] As described above, the bladder cancer lymph node metastasis prediction device provided in the embodiment of the present application obtains multiple sliding window images corresponding to the urine cell slice image and multiple cross-sectional images corresponding to the nuclear magnetic resonance image; traverses and combines the multiple sliding window images and the multiple cross-sectional images to obtain multiple image pairs, each image pair includes a sliding window image and a cross-sectional image; extracts the diseased cell features of the sliding window image in the image pair, extracts the lesion area features of the cross-sectional image in the image pair, and fuses the diseased cell features and the lesion area features through the mutual attention model to obtain a fused feature; splices the fused features of each image pair and inputs them into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph nodes output by the classification network. Through the above-mentioned technical means, the correlation features between the diseased cell features of each sliding window image in the urine cell slice image and the lesion area features of each cross-sectional image in the magnetic resonance image can be extracted through the mutual attention model, and the correlation features between each sliding window image and each cross-sectional image can be spliced to obtain the correlation features between the urine cell slice image and the magnetic resonance image. The correlation features between the urine cell slice image and the magnetic resonance image are analyzed through the classification network to output the metastasis results of the bladder cancer lymph nodes, realizing multimodal fusion detection of lymph node metastasis to accurately detect the metastasis results of bladder cancer lymph nodes, solving the problem of low accuracy of lymph node metastasis detected by single radiology in the prior art, reducing the probability of misdiagnosis or missed diagnosis, and avoiding the problem of insufficient or excessive treatment of patients.
[0093] The bladder cancer lymph node metastasis prediction device provided in the embodiment of the present application can be used to execute the bladder cancer lymph node metastasis prediction method provided in the above embodiment, and has corresponding functions and beneficial effects.
[0094] Fig. 9 is a schematic diagram of a bladder cancer lymph node metastasis prediction device provided in an embodiment of the present application, with reference to Fig. 9 The bladder cancer lymph node metastasis prediction device includes: a processor 41, a memory 42, a communication device 43, an input device 44, and an output device 45. The number of processors 41 in the bladder cancer lymph node metastasis prediction device can be one or more, and the number of memories 42 in the bladder cancer lymph node metastasis prediction device can be one or more. The processor 41, memory 42, communication device 43, input device 44, and output device 45 of the bladder cancer lymph node metastasis prediction device can be connected via a bus or other means.
[0095] The memory 42, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the bladder cancer lymph node metastasis prediction method of any embodiment of the present application (for example, the image acquisition module 31, the image combination module 32, the feature fusion module 33 and the metastasis prediction module 34 in the bladder cancer lymph node metastasis prediction device). The memory 42 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the device, etc. In addition, the memory 42 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0096] The communication device 43 is used for data transmission.
[0097] The processor 41 executes the software programs, instructions and modules stored in the memory 42 to perform various functional applications and data processing of the device, that is, to implement the above-mentioned bladder cancer lymph node metastasis prediction method.
[0098] The input device 44 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 45 may include a display device such as a display screen.
[0099] The above-mentioned bladder cancer lymph node metastasis prediction device can be used to execute the bladder cancer lymph node metastasis prediction method provided in the above-mentioned embodiment, and has corresponding functions and beneficial effects.
[0100] The embodiment of the present application also provides a storage medium containing computer executable instructions. When the computer executable instructions are executed by a computer processor, they are used to execute a method for predicting bladder cancer lymph node metastasis. The method for predicting bladder cancer lymph node metastasis includes: obtaining multiple sliding window images corresponding to urine cell slice images, and obtaining multiple cross-sectional images corresponding to nuclear magnetic resonance images; traversing and combining the multiple sliding window images and the multiple cross-sectional images to obtain multiple image pairs, each image pair includes a sliding window image and a cross-sectional image; extracting the diseased cell features of the sliding window image in the image pair, extracting the lesion area features of the cross-sectional image in the image pair, and fusing the diseased cell features and the lesion area features through a mutual attention model to obtain a fused feature; splicing the fused features of each image pair and inputting them into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph nodes output by the classification network.
[0101] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media, such as CD-ROM, floppy disk or tape device; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (such as hard disk or optical storage); registers or other similar types of memory elements, etc. Storage media may also include other types of memory or combinations thereof. In addition, the storage medium may be located in the first computer system in which the program is executed, or may be located in a different second computer system, which is connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media residing in different locations (for example, in different computer systems connected by a network). The storage medium may store program instructions (for example, embodied as a computer program) that can be executed by one or more processors.
[0102] Of course, the storage medium containing computer executable instructions provided in the embodiments of the present application is not limited to the above-mentioned method for predicting bladder cancer lymph node metastasis, and the computer executable instructions can also execute the related operations in the method for predicting bladder cancer lymph node metastasis provided in any embodiment of the present application.
[0103] The bladder cancer lymph node metastasis prediction device, bladder cancer lymph node metastasis prediction system, storage medium and bladder cancer lymph node metastasis prediction equipment provided in the above embodiments can execute the bladder cancer lymph node metastasis prediction method provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, please refer to the bladder cancer lymph node metastasis prediction method provided in any embodiment of the present application.
[0104] The above are only preferred embodiments of the present application and the technical principles used. The present application is not limited to the specific embodiments herein, and various obvious changes, readjustments and substitutions that can be made by those skilled in the art will not deviate from the protection scope of the present application. Therefore, although the present application is described in more detail through the above embodiments, the present application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the claims.
Claims
1. A method for predicting bladder cancer lymph node metastasis, characterized in that: include: Acquire multiple sliding window images corresponding to the urine cell slice image, and acquire multiple cross-sectional images corresponding to the nuclear magnetic resonance image; Traversing and combining the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, each of the image pairs comprising one sliding window image and one cross-sectional image; Extracting diseased cell features of the sliding window image in the image pair, extracting lesion region features of the cross-sectional image in the image pair, and fusing the diseased cell features and the lesion region features through a mutual attention model to obtain fused features; The fusion features of each of the image pairs are spliced and then input into a pre-trained classification network to obtain the metastasis result of bladder cancer lymph nodes output by the classification network.
2. The method for predicting bladder cancer lymph node metastasis according to claim 1, characterized in that: The extracting of the diseased cell features of the sliding window image in the image pair includes: Inputting the sliding window image into a pre-trained first deep feature extraction model to obtain a first feature of the sliding window image; The first feature is input into a pre-trained first self-attention model to obtain a diseased cell feature output by the first self-attention model.
3. The method for predicting bladder cancer lymph node metastasis according to claim 2, characterized in that: The extracting of the lesion area feature of the cross-sectional image in the image pair comprises: reducing the cross-sectional image to a preset size; Inputting the reduced cross-sectional image into a pre-trained second deep feature extraction model to obtain a second feature of the cross-sectional image; The second feature is input into a pre-trained second self-attention model to obtain a lesion area feature output by the second self-attention model.
4. The method for predicting bladder cancer lymph node metastasis according to claim 3, characterized in that: The training steps of the first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model, the mutual attention model and the classification network include: Divide the urine cell slice sample image of each patient into a plurality of sliding window sample images, and obtain a plurality of cross-sectional sample images corresponding to the nuclear magnetic resonance sample image of each patient; Based on a plurality of sliding window sample images and a plurality of cross-section sample images, training the first feature extraction model and the second feature extraction model by a contrastive learning algorithm; Extracting a first sample feature of each of the sliding window sample images by using a trained first feature extraction model, and extracting a second sample feature of each of the cross-sectional sample images by using a trained second feature extraction model; Based on the first sample feature and the second sample feature, training the first self-attention model, the second self-attention model, and the mutual attention model through a segmentation training algorithm; The sample fusion features of the sliding window sample image and the cross-sectional sample image of the same patient are extracted by using the trained first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model and the mutual attention model; The sample fusion features of the same patient are spliced to obtain sample splicing features, and the classification network is trained based on the sample splicing features of each patient.
5. The method for predicting bladder cancer lymph node metastasis according to claim 4, characterized in that: The method of training the first feature extraction model and the second feature extraction model by a contrast learning algorithm based on a plurality of sliding window sample images and a plurality of cross-section sample images includes: The sliding window sample images and cross-sectional sample images of the same patient are traversed and combined to obtain first training sample data, and the sliding window sample images and cross-sectional sample images of different patients are randomly combined to obtain second training sample data; Inputting the sliding window sample image and the cross-section sample image in the first training sample data into a first classification model and a second classification model respectively, obtaining the cell category and the region category output by the first classification model and the second classification model, wherein the first classification model includes a first feature extraction model and a first classification network, and the second classification model includes a second feature extraction model and a second classification network; In the case where the cell categories and region categories corresponding to the sliding window sample image and the cross-sectional sample image in the first training sample data are accurately identified as diseased cells and lesion regions, the first training sample data are determined as true paired data, and the remaining first training sample data and the second training sample data are determined as incorrect paired data; Maximizing the cosine similarity between a first sample feature of a sliding window sample image and a second sample feature of a cross-sectional sample image in the true paired data, and minimizing the cosine similarity between a first sample feature of a sliding window sample image and a second sample feature of a cross-sectional sample image in the false paired data, wherein the first sample feature and the second sample feature are extracted by the first feature extraction model and the second feature extraction model respectively; A symmetric cross entropy loss value is determined according to the cosine similarity of the true paired data and the cosine similarity of the false paired data, and the first classification model and the second classification model are optimized based on the symmetric cross entropy loss value.
6. The method for predicting bladder cancer lymph node metastasis according to claim 5, characterized in that: The training of the first self-attention model, the second self-attention model, and the mutual attention model by a segmentation training algorithm based on the first sample feature and the second sample feature includes: Inputting the first sample feature and the second sample feature corresponding to the first training sample data into the first self-attention model and the second self-attention model respectively, to obtain the lesion cell sample feature output by the first self-attention model and the lesion area sample feature output by the second self-attention model; Inputting the diseased cell sample features and the lesion area sample features into the mutual attention model to obtain sample fusion features output by the mutual attention model; Inputting the sample fusion features into a segmentation model to obtain a mask result output by the segmentation model, wherein the mask result includes a diseased cell mask and a lesion area mask; The loss value of each pixel is determined based on the mask result and the annotation information of the sliding window sample image and the cross-sectional sample image in the first training sample data, the mean square error loss value is determined based on the loss value of each pixel, and the first self-attention model, the second self-attention model, the mutual attention model and the segmentation model are optimized based on the mean square error loss value.
7. The method for predicting bladder cancer lymph node metastasis according to claim 5, characterized in that: The method extracts sample fusion features of the sliding window sample image and the cross-sectional sample image of the same patient by using the trained first feature extraction model, the second feature extraction model, the first self-attention model, the second self-attention model and the mutual attention model, including: Inputting the first sample feature and the second sample feature corresponding to the first training sample data into the trained first self-attention model and the second self-attention model respectively, to obtain the lesion cell sample feature output by the first self-attention model and the lesion area sample feature output by the second self-attention model; Inputting the diseased cell sample features and the lesion area sample features into a trained mutual attention model to obtain sample fusion features output by the mutual attention model; Accordingly, the sample fusion features of the same patient are spliced to obtain sample splicing features, and the classification network is trained based on the sample splicing features of each patient, including: The sample fusion features corresponding to the first training sample data of the same patient are spliced to obtain a sample splicing feature; Inputting the sample splicing features into the classification network to obtain a classification result output by the classification network; A loss value is determined based on the classification result and transfer information of the corresponding patient, and the classification network is optimized based on the loss value.
8. A device for predicting lymph node metastasis of bladder cancer, characterized in that: include: An image acquisition module is configured to acquire a plurality of sliding window images corresponding to the urine cell slice image and to acquire a plurality of cross-sectional images corresponding to the nuclear magnetic resonance image; An image combination module, configured to traverse and combine the plurality of sliding window images and the plurality of cross-sectional images to obtain a plurality of image pairs, each of the image pairs comprising one sliding window image and one cross-sectional image; a feature fusion module, configured to extract the diseased cell features of the sliding window image in the image pair, extract the lesion area features of the cross-sectional image in the image pair, and fuse the diseased cell features and the lesion area features through a mutual attention model to obtain a fusion feature; The metastasis prediction module is configured to splice the fusion features of each of the image pairs and input the spliced features into a pre-trained classification network to obtain the metastasis result of the bladder cancer lymph nodes output by the classification network.
9. A device for predicting lymph node metastasis of bladder cancer, characterized in that: include: one or more processors; A storage device stores one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the bladder cancer lymph node metastasis prediction method as described in any one of claims 1-7.
10. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions are used to execute the method for predicting bladder cancer lymph node metastasis as claimed in any one of claims 1 to 7 when executed by a computer processor.
Citation Information
Patent Citations
Endoscopic OCT image segmentation method and device for colorectal tumor, medium and product
CN115272283A
Medical image tumor segmentation method involving cross-modal attention mechanism
CN115512110A
Bidirectional interactive attention laparoscope image efficient segmentation method and device
CN118247500A
Scene classification model training method adaptable to multiple tasks and scene classification method
CN118247746A
Tumor risk prediction method and device and storage medium
CN118919073A