Processing method and device for predicting breast cancer classification by fusing multi-modal features
By constructing an encoder for multi-stage residual network and linear network, combining multimodal feature fusion and noise injection mechanisms, the problem of limited prediction performance in singlemodal feature fusion model is solved, and the accuracy of breast cancer classification is improved and the model robustness is enhanced.
Patent Information
- Application Number
- CN202510602811.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-26
AI Technical Summary
The existing breast cancer classification model is mainly based on single-modal features, resulting in insufficient feature dimensions, limited prediction performance, and difficulty in improving accuracy. Especially for young doctors, the rate of misjudgment and misjudgment is high.
The encoder is constructed using multi-level residual network and multi-level linear network, and the multi-level residual feature fusion mechanism, spatial attention fusion mechanism and noise injection mechanism are introduced to build a deep learning model for multi-modal feature fusion. Breast cancer classification is performed through imaging and genomic feature fusion, and model training with mask maps and maskless maps is combined to improve the robustness and generalization ability of the model.
Through multimodal feature fusion, the accuracy of breast cancer classification is improved, the rate of misjudgment and misjudgment is reduced, the robustness and noise resistance of the model are enhanced, and the recognition accuracy of young doctors is improved.
Smart Images

Figure CN120544841A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a processing method and device for fusing multimodal features to predict breast cancer classification. Background Art
[0002] Breast cancer classification essentially involves identifying breast tumor types, such as benign, malignant, and borderline. Malignant tumors are commonly referred to as breast cancer, while borderline tumors are known as precancerous adenocarcinomas. Conventionally, doctors make classification decisions based on a series of examinations and test data (such as breast ultrasound images and individual genetic sequencing data) based on their own medical experience. However, this conventional approach is limited by manual experience, and manual identification often results in low accuracy and high rates of misclassification and missed diagnosis, especially for medical students or young doctors with limited experience.
[0003] With the application and development of artificial intelligence technology in the medical field, some classification prediction models based on machine learning or deep learning models have also begun to be applied to the field of breast cancer classification. Using these classification prediction models as auxiliary identification tools can help improve the recognition accuracy of medical students and young doctors and reduce the rate of false positives and false negatives. However, research has found that most commonly used classification prediction models are based on single-modal features, such as predictions based on medical imaging features or genomic features. The drawback of single-modal prediction mechanisms is that the feature dimensionality is not rich enough and the feature expression is not comprehensive. This leads to limited prediction performance and makes it difficult to further improve the prediction accuracy after reaching a certain level. Summary of the Invention
[0004] The purpose of the present invention is to address the shortcomings of the existing technology and provide a processing method, device, electronic device and computer-readable storage medium for fusing multimodal features to predict breast cancer classification. The present invention first constructs a first encoder for imaging genomics feature encoding using a multi-level residual network as a feature extraction backbone, and introduces a multi-level residual feature fusion mechanism and a spatial attention fusion mechanism into the encoder; and constructs a second encoder for genomics feature encoding using a multi-level linear network as a feature extraction backbone, and introduces a noise injection mechanism into the encoder; and constructs a first prediction network for performing three-category prediction of benign, malignant, and borderline lesions of breast tumors based on a classifier model composed of a linear network and a softmax function; and constructs a breast cancer classification model based on the first and second encoders and the first prediction network for multimodal feature fusion of breast ultrasound imaging genomics features and genomic features, and for classifying and predicting benign, malignant, and borderline lesions of breast tumors based on the fused features; then, constructs a model training dataset through data acquisition, and performs two rounds of model training with and without mask images on the breast cancer classification model based on the dataset; and after the two rounds of model training, the breast cancer classification model is used for prediction application based on the data of the subjects. The present invention, on the one hand, improves feature richness by fusing multimodal features, thereby achieving the purpose of improving prediction accuracy and reducing false positive / missed positive rates; on the other hand, it enhances the robustness, noise resistance and generalization ability of the model by introducing a noise injection mechanism and a random deactivation layer in the second encoder.
[0005] To achieve the above objectives, a first aspect of an embodiment of the present invention provides a method for predicting breast cancer classification by fusing multimodal features, the method comprising:
[0006] A first encoder for imaging omics feature encoding is constructed using a multi-level residual network as the feature extraction backbone, and a multi-level residual feature fusion mechanism and a spatial attention fusion mechanism are introduced into the encoder; a second encoder for genomic feature encoding is constructed using a multi-level linear network as the feature extraction backbone, and a noise injection mechanism is introduced into the encoder; and a first prediction network for three-category prediction of benign, malignant and borderline lesions of breast tumors is constructed based on a classifier model composed of a linear network and a softmax function; and a deep learning model for multimodal feature fusion of breast ultrasound imaging omics features and genomic features and classification prediction of benign, malignant and borderline lesions of breast tumors based on the fused features is constructed based on the first and second encoders and the first prediction network, which is recorded as a breast cancer classification model;
[0007] Constructing a model training data set as a corresponding first data set; firstly performing a first round of mask image model training on the breast cancer classification model based on the first data set, and then performing a second round of non-mask image model training on the breast cancer classification model based on the first data set;
[0008] After two rounds of model training, the subject data input by the user is received; the breast cancer classification model performs prediction based on the subject data to obtain a corresponding subject prediction vector and feeds it back to the current user; the subject data includes breast ultrasound images, lesion area masks, and genomic feature matrices, and the lesion area masks are optional data; the subject prediction vector includes three prediction probabilities, namely, benign prediction probability, malignant prediction probability, and borderline lesion prediction probability.
[0009] Preferably, the first encoder is used to perform radiomics feature encoding processing based on the breast ultrasound image and the lesion area mask image input by the encoder and output a corresponding image feature vector;
[0010] The breast ultrasound image is a breast B-ultrasound image, a breast Doppler ultrasound image or a breast ultrasound elastic imaging; the image shape of the breast ultrasound image is C P ×H P ×W P , H P 、W P are the image height and width, C P is the pixel feature dimension of the image;
[0011] The image shape when the lesion area mask is not empty is 1×H P ×W P The lesion area mask is a binary mask obtained by outlining the lesion area on the corresponding breast ultrasound image; the image size of the lesion area mask is consistent with the breast ultrasound image, and the pixel feature dimension of the lesion area mask is 1; each pixel feature of the lesion area mask is a 0 / 1 binary feature, if it is 0, it indicates that the current pixel position is a non-lesion position, and if it is 1, it indicates that the current pixel position is a lesion position;
[0012] The image feature vector is a one-dimensional feature vector with a shape of 1×L1; L1 is a preset vector length threshold; the vector length threshold L1 defaults to 512; the vector eigenvalue of the image feature vector is constrained to be within the value range [0,1];
[0013] The first input end of the first encoder is used to receive the breast ultrasound image input by the encoder, and the second input end is used to receive the lesion area mask image input by the encoder; the output end of the first encoder is used to output the corresponding image feature vector;
[0014] The first encoder includes a 7×7 convolutional layer, a maximum pooling layer, a first-level residual network, a second-level residual network, a third-level residual network, an upsampling layer, a multi-level feature fusion layer, a spatial attention network, an attention feature fusion layer, and a feature mapping layer;
[0015] The input end of the 7×7 convolutional layer is connected to the first input end of the first encoder, and the output end is connected to the input end of the maximum pooling layer; the output end of the maximum pooling layer is connected to the input end of the first-level residual network; the output end of the first-level residual network is respectively connected to the input end of the second-level residual network and the first input end of the upsampling layer; the output end of the second-level residual network is respectively connected to the input end of the third-level residual network and the second input end of the upsampling layer; the output end of the third-level residual network is connected to the third input end of the upsampling layer; the output end of the upsampling layer is connected to the input end of the multi-level feature fusion layer; the output end of the multi-level feature fusion layer is respectively connected to the input end of the spatial attention network and the first input end of the attention feature fusion layer; the output end of the spatial attention network is connected to the second input end of the attention feature fusion layer; the third input end of the attention feature fusion layer is connected to the second input end of the first encoder, and the output end is connected to the input end of the feature mapping layer; the output end of the feature mapping layer is connected to the output end of the first encoder;
[0016] The convolution kernel size of the 7×7 convolution layer is 7×7; the 7×7 convolution layer is used to perform downsampling and feature extraction processing on the breast ultrasound image to obtain the corresponding first feature tensor and send it to the maximum pooling layer; the tensor shape of the first feature tensor is C1×H1×W1, where H1, W1, and C1 are the height, width, and feature dimension of the first feature tensor respectively; H1=H P / 2, W1=H P / 2, the preset dimension C1 defaults to 64;
[0017] The maximum pooling layer is used to downsample the first feature tensor to obtain a corresponding second feature tensor and send it to the first-level residual network; the tensor shape of the second feature tensor is C2×H2×W2, where H2, W2, and C2 are the height, width, and feature dimension of the second feature tensor respectively; H2=H1 / 2, W2=H1 / 2, and C2=C1; when the preset dimension C1 is 64, the feature dimension C2 is 64;
[0018] The first-level residual network is composed of three residual modules connected in sequence; the first-level residual network is used to downsample and extract features from the second feature tensor to obtain a corresponding third feature tensor, which is sent to the second-level residual network and the upsampling layer; the tensor shape of the third feature tensor is C3×H3×W3, where H3, W3, and C3 are the height, width, and feature dimension of the third feature tensor, respectively; H3=H2 / 2, W3=H2 / 2, and C3=C2; when the preset dimension C1 is 64, the feature dimension C3 is 64;
[0019] The second-level residual network is composed of four residual modules connected in sequence; the second-level residual network is used to downsample and extract features from the third feature tensor to obtain a corresponding fourth feature tensor, which is sent to the third-level residual network and the upsampling layer; the tensor shape of the fourth feature tensor is C4×H4×W4, where H4, W4, and C4 are the height, width, and feature dimension of the fourth feature tensor, respectively; H4=H3 / 2, W4=H3 / 2, and C4=2×C3; when the preset dimension C1 is 64, the feature dimension C4 is 128;
[0020] The third-level residual network is composed of six residual modules connected in sequence; the third-level residual network is used to downsample and extract features from the fourth feature tensor to obtain a corresponding fifth feature tensor and send it to the upsampling layer; the tensor shape of the fifth feature tensor is C5×H5×W5, where H5, W5, and C5 are the height, width, and feature dimension of the fifth feature tensor, respectively; H5=H4 / 2, W5=H4 / 2, and C5=2×C4; when the preset dimension C1 is 64, the feature dimension C5 is 256;
[0021] The upsampling layer is used to perform bilinear interpolation and a preset uniform height H std and uniform width W std The third, fourth and fifth feature tensors are upsampled respectively, and the height and width of the three tensors are all the same as the unified height H std and the uniform width W std The feature tensors of are recorded as the corresponding first, second and third level feature tensors and sent to the multi-level feature fusion layer; the first, second and third level feature tensors correspond one to one with the third, fourth and fifth feature tensors; the tensor shape of the first level feature tensor is C3×H LV1 ×W LV1 , H LV1 、W LV1 are the height and width of the first-level feature tensor respectively; the shape of the second-level feature tensor is C4×H LV2 ×W LV2 , H LV2、W LV2 are the height and width of the second-level feature tensor respectively; the tensor shape of the third-level feature tensor is C5×H LV3 ×W LV3 , H LV3 、W LV3 are the height and width of the third-level feature tensor respectively; H LV1 =H LV2 =H LV3 =H std , W LV1 =W LV2 =W LV3 =W std ;
[0022] The multi-level feature fusion layer is used to perform feature fusion processing on the first, second and third level feature tensors in a feature channel splicing manner to obtain the corresponding first fusion tensor and send it to the spatial attention network and the attention feature fusion layer; the tensor shape of the first fusion tensor is C6×H6×W6, where H6, W6 and C6 are the height, width and feature dimension of the first fusion tensor respectively; H6=H std , W6=W std , C6=C3+C4+C5; when the preset dimension C1 is 64, the feature dimension C6 is 448;
[0023] The spatial attention network first performs feature dimensionality reduction on the first fusion tensor through a 3×3 convolutional layer to obtain the corresponding first reduced dimensionality tensor; then performs nonlinear activation on the first reduced dimensionality tensor through the ReLU function to obtain the corresponding first activation tensor; then aggregates the feature channels of the first activation tensor through a 1×1 convolutional layer to obtain the corresponding second reduced dimensionality tensor; then normalizes the second reduced dimensionality tensor through the Sigmoid function and uses the processing result as the corresponding spatial attention weight matrix; finally, the normalized second reduced dimensionality tensor is sent to the attention feature fusion layer; the tensor shape of the first reduced dimensionality tensor is C7×H7×W7, where H7, W7, and C7 are the height, width, and feature dimension of the first reduced dimensionality tensor respectively, and H7=H std , W7=W std , C7=C6 / n; the preset dimensionality reduction multiple n is a positive integer greater than or equal to 2 and divisible by the feature dimension C6; the tensor shape of the first activation tensor is C7×H7×W7; the tensor shape of the second dimensionality reduction tensor is C8×H8×W8, where H8, W8, and C8 are the height, width, and feature dimension of the first dimensionality reduction tensor respectively, and H8=H std , W8=W std , C8=1; the matrix shape of the spatial attention weight matrix is H std ×Wstd ;
[0024] The attention feature fusion layer is used to identify whether the lesion area mask input by the encoder is empty; if the lesion area mask is empty, the spatial attention weight matrix is copied C6 times based on the feature dimension of the first fusion tensor to obtain a C6×H std ×W std The first weight tensor of ; if the lesion area mask map is not empty, then according to the bilinear interpolation algorithm and the unified height H std and the uniform width W std The lesion area mask image is downsampled to obtain a shape of H std ×W std The down-sampling mask map is obtained by performing a Hadamard product operation on the down-sampling mask map and the spatial attention weight matrix, and the operation result is used as the corresponding first weight matrix, and the first weight matrix is copied C6 times based on the feature dimension of the first fusion tensor to obtain a C6×H std ×W std The first weight tensor of the first fusion tensor is obtained by performing a Hadamard product operation on the first weight tensor and the first fusion tensor and sending the result of the operation as the corresponding second fusion tensor to the feature mapping layer; the tensor shape of the second fusion tensor is C8×H8×W8, where H8, W8, and C8 are the height, width, and feature dimension of the second fusion tensor respectively, and H8=H std , W8=W std , C8=C6;
[0025] The feature mapping layer first performs feature dimension upscaling on the second fusion tensor through a 1×1 convolution layer to obtain a corresponding first dimension upscaling tensor; then performs batch normalization processing on the first dimension upscaling tensor through a batch normalization layer to obtain a corresponding first normalized tensor; then performs nonlinear activation on the first normalized tensor through a ReLU function to obtain a corresponding second activation tensor; then performs global average pooling processing on the second activation tensor through a global average pooling layer and outputs the processing result as the corresponding image feature vector; the tensor shapes of the first dimension upscaling tensor, the first normalized tensor and the second activation tensor are all C9×H9×W9, where H9, W9 and C9 are the height, width and feature dimension of the first dimension upscaling tensor respectively, and H9=H std , W9=W std , C9=L1; the shape of the image feature vector is 1×L1, and its vector eigenvalue is constrained to be within the value range [0,1].
[0026] Preferably, the second encoder is used to perform genomic feature encoding processing according to the genomic feature matrix input by the encoder and output a corresponding gene feature vector;
[0027] Each row of the genomic feature matrix corresponds to a type of gene, and each column corresponds to a type of gene feature; the matrix shape of the genomic feature matrix is H G ×W G , H G 、W G are the height and width of the matrix, H G 、W G Match the total number of genes and the total number of gene features in the genomic feature matrix respectively;
[0028] The gene feature vector is a one-dimensional feature vector with a shape of 1×L2; L2 is a preset vector length threshold; the vector length threshold L2 defaults to 32; the vector eigenvalue of the gene feature vector is constrained to be within the value range [-1, 1];
[0029] The input end of the second encoder is used to receive the genomic feature matrix input by the encoder, and the output end is used to output the corresponding gene feature vector;
[0030] The second encoder includes a first linear layer, a first activation layer, a first normalization layer, a second linear layer, a second activation layer, a third linear layer, a noise injection layer, and a feature truncation layer;
[0031] The input end of the first linear layer is connected to the input end of the second encoder, and the output end is connected to the input end of the first activation layer; the output end of the first activation layer is connected to the input end of the first normalization layer; the output end of the first normalization layer is connected to the input end of the second linear layer; the output end of the second linear layer is connected to the input end of the second activation layer; the output end of the second activation layer is connected to the input end of the third linear layer; the output end of the third linear layer is connected to the input end of the noise injection layer; the output end of the noise injection layer is connected to the input end of the feature truncation layer; and the output end of the feature truncation layer is connected to the output end of the second encoder;
[0032] The first linear layer is used to flatten the genomic feature matrix into a matrix with a length of H G ×W G The one-dimensional vector is recorded as the first vector; and the first vector is fully connected to obtain the corresponding first fully connected vector and sent to the first activation layer;
[0033] The first activation layer is used to perform nonlinear activation on the first fully connected vector through a ReLU function to obtain a corresponding first activation vector and send it to the first normalization layer;
[0034] The first normalization layer is used to perform layer normalization processing on the first activation vector to obtain a corresponding first normalized vector and send it to the second linear layer;
[0035] The second linear layer is used to perform a full connection calculation on the first normalized vector to obtain a corresponding second fully connected vector and send it to the second activation layer;
[0036] The second activation layer is used to perform nonlinear activation on the second fully connected vector through a Tanh function to obtain a corresponding second activation vector and send it to the third linear layer; the vector eigenvalue of the second activation vector is constrained to be within the value range [-1, 1];
[0037] The third linear layer is used to perform a full connection calculation on the second activation vector to obtain a corresponding third fully connected vector and send it to the noise injection layer; the third fully connected vector is a one-dimensional feature vector with a shape of 1×L2;
[0038] The noise injection layer is used to construct a standard normal distribution noise vector with a vector length of L2 according to the model parameter ε, recorded as the first noise vector; and the first noise vector and the third fully connected vector are added to obtain the corresponding noise injection vector and sent to the feature truncation layer; the mean of the first noise vector is 0 and the variance is ε 2 ; The model parameter ε is a learnable model parameter, and its initial value is set to the preset initialization noise coefficient ε0;
[0039] The feature truncation layer is used to perform nonlinear activation on the noise injection vector through the Tanh function to obtain the corresponding gene feature vector and output it; the gene feature vector is a one-dimensional feature vector with a shape of 1×L2, and its vector eigenvalue is constrained within the value range [-1,1].
[0040] Preferably, the first prediction network is used to perform three-category prediction processing on benign, malignant and borderline lesions of breast tumors based on the fusion feature vector input by the network and output corresponding classification prediction vectors;
[0041] The fused feature vector is a one-dimensional feature vector with a shape of 1×(L1+L2); L1 and L2 are two preset vector length thresholds; the vector length threshold L1 defaults to 512, and the vector length threshold L2 defaults to 32; the first L1 vector eigenvalues of the fused feature vector are constrained to be within the value range [0,1], and the last L2 vector eigenvalues are constrained to be within the value range [-1,1];
[0042] The classification prediction vector includes three prediction probabilities, namely the benign prediction probability, the malignant prediction probability and the borderline lesion prediction probability;
[0043] The network input end of the first prediction network is used to receive the fusion feature vector input by the network, and the network output end is used to output the corresponding classification prediction vector;
[0044] The first prediction network includes a fourth linear layer, a second normalization layer, a third activation layer, a random inactivation layer, a fifth linear layer, and a Softmax function layer;
[0045] The input end of the fourth linear layer is connected to the input end of the first prediction network, and the output end is connected to the input end of the second normalization layer; the output end of the second normalization layer is connected to the input end of the third activation layer; the output end of the third activation layer is connected to the input end of the random deactivation layer; the output end of the random deactivation layer is connected to the input end of the fifth linear layer; the output end of the fifth linear layer is connected to the input end of the softmax function layer;
[0046] The fourth linear layer is used to perform a full connection calculation on the fused feature vector to obtain a corresponding fourth fully connected vector and send it to the second normalization layer;
[0047] The second normalization layer is used to perform batch normalization on the fourth fully connected vector to obtain a corresponding second normalized vector and send it to the third activation layer;
[0048] The third activation layer is used to perform nonlinear activation on the second normalized vector through a ReLU function to obtain a corresponding third activation vector and send it to the random inactivation layer;
[0049] The random deactivation layer is used to perform random deactivation processing on the vector features in the third activation vector in equal proportion according to a preset random deactivation ratio to obtain a corresponding local deactivation vector and send it to the fifth linear layer; the local deactivation vector has the same vector shape as the third activation vector, but some vector features in the local deactivation vector are set to preset deactivation feature values; the ratio of the total number of deactivation feature values of the local deactivation vector to the vector length meets the random deactivation ratio;
[0050] The fifth linear layer is used to perform a full connection calculation on the local deactivation vector to obtain a corresponding fifth fully connected vector and send it to the Softmax function layer; the vector length of the fifth fully connected vector is 3;
[0051] The Softmax function layer is used to perform three-category probability calculation based on the fifth fully connected vector through the Softmax function to obtain the corresponding benign prediction probability, the malignant prediction probability and the borderline lesion prediction probability to form the corresponding classification prediction vector and output it.
[0052] Preferably, the breast cancer classification model is used to classify and predict benign, malignant and borderline lesions of breast tumors based on the breast ultrasound image, the lesion area mask and the genomic feature matrix input into the model and output a corresponding classification prediction vector;
[0053] The model input of the breast cancer classification model includes a breast ultrasound image, a lesion area mask, and a genomic feature matrix, wherein the lesion area mask is an optional input; the model output of the breast cancer classification model is the classification prediction vector, and the classification prediction vector includes three prediction probabilities, namely, the benign prediction probability, the malignant prediction probability, and the borderline lesion prediction probability;
[0054] The first model input end of the breast cancer classification model is used to receive the breast ultrasound image and the lesion area mask, the second model input end is used to receive the genomic feature matrix, and the model output end is used to output the corresponding classification prediction vector;
[0055] The breast cancer classification model includes the first encoder, the second encoder, a multimodal feature fusion module and the first prediction network;
[0056] The encoder input of the first encoder is connected to the input of the first model, and the encoder output is connected to the first input of the multimodal feature fusion module; the encoder input of the second encoder is connected to the input of the second model, and the encoder output is connected to the second input of the multimodal feature fusion module; the output of the multimodal feature fusion module is connected to the network input of the first prediction network; and the network output of the first prediction network is connected to the output of the model;
[0057] The first encoder is used to perform radiomics feature encoding processing based on the breast ultrasound image and the lesion area mask image input by the encoder to obtain a corresponding image feature vector and send it to the multimodal feature fusion module; the image feature vector is a one-dimensional feature vector with a shape of 1×L1, and the vector length threshold L1 defaults to 512;
[0058] The second encoder is used to perform genomic feature encoding processing according to the genomic feature matrix input by the encoder to obtain a corresponding gene feature vector and send it to the multimodal feature fusion module; the gene feature vector is a one-dimensional feature vector with a shape of 1×L2, and the vector length threshold L2 defaults to 32;
[0059] The multimodal feature fusion module is used to sequentially concatenate the image feature vector and the gene feature vector to obtain a one-dimensional feature vector with a shape of 1×(L1+L2) as a corresponding fused feature vector; and send the fused feature vector to the first prediction network;
[0060] The first prediction network is used to perform three-category prediction processing on benign, malignant and borderline lesions of breast tumors according to the fusion feature vector and output the corresponding classification prediction vector.
[0061] Preferably, the first data set includes multiple first data records; each first data record corresponds to an acquisition object; the first data record includes a first training image, a first training mask image, a first training feature matrix and a first label vector; the first training image is the breast ultrasound image of the current acquisition object; the first training mask image is the lesion area mask image of the current acquisition object; the first training feature matrix is the genomic feature matrix of the current acquisition object; the first label vector includes three label probabilities, namely, benign label probability, malignant label probability and borderline lesion label probability, and one and only one of the three label probabilities is 1, and the other two label probabilities are 0; all the first training images of the first data set have the same image type, specifically breast B-ultrasound image, breast Doppler ultrasound image or breast ultrasound elastography.
[0062] Preferably, the model building training data set is recorded as the corresponding first data set, which specifically includes:
[0063] Recruiting multiple patients with benign breast tumors, malignant breast tumors, and borderline breast lesions to form a collection subject group; the collection subject group includes multiple collection subjects;
[0064] Select one from the three image types of breast B-ultrasound image, breast Doppler ultrasound image and breast ultrasound elastography as the corresponding current ultrasound image type; and use each of the acquisition objects as the corresponding current object; and perform a breast ultrasound examination on the current object based on the current ultrasound image type and use the breast ultrasound image obtained in this examination as a corresponding first training image; and perform a whole genome sequencing on the current object to obtain the corresponding current object gene sequence, and fill all matrix units of the genomic feature matrix with feature data based on the current object gene sequence and all gene types and all gene features specified by the genomic feature matrix, and use the filled genomic feature matrix as the corresponding first training feature matrix; and use the preset breast lesion delineation interface to sequence the current object on the first training image. The breast lesion area of the image is delineated and a lesion area mask is generated as the corresponding first training mask based on the delineation result; and the current object is identified; if the current object is a benign breast tumor patient, a first label vector is set in which the benign label probability is 1 and the malignant and borderline lesion label probabilities are 0; if the current object is a malignant breast tumor patient, a first label vector is set in which the malignant label probability is 1 and the benign and borderline lesion label probabilities are 0; if the current object is a borderline breast lesion patient, a first label vector is set in which the borderline lesion label probability is 1 and the benign and malignant label probabilities are 0; and a corresponding first data record is formed by the first training image, the first training mask, the first training feature matrix and the first label vector corresponding to the current object;
[0065] All the obtained first data records form the corresponding first data set.
[0066] Preferably, the first round of masked image model training is performed on the breast cancer classification model based on the first data set, and then the second round of unmasked image model training is performed on the breast cancer classification model based on the first data set, specifically including:
[0067] Step 81, set the training round to the first round;
[0068] Step 82: randomly split the first data set into two sub-data sets based on a preset first split ratio and record them as a first training set and a first evaluation set;
[0069] Wherein, both the first training set and the first evaluation set include a plurality of the first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio;
[0070] Step 83: extract the first first data record from the first training set as the corresponding current training record;
[0071] Step 84, identifying the training round; if the training round is the first round, inputting the first training image, the first training mask, and the first training feature matrix of the current training record as the current breast ultrasound image, the lesion area mask, and the genomic feature matrix into the breast cancer classification model for classification prediction processing, and using the classification prediction vector output from this processing as the corresponding first prediction vector; if the training round is the second round, inputting the first training image and the first training feature matrix of the current training record as the current breast ultrasound image and the genomic feature matrix, and setting the current lesion area mask to empty, and inputting the current breast ultrasound image, the lesion area mask, and the genomic feature matrix into the breast cancer classification model for classification prediction processing, and using the classification prediction vector output from this processing as the corresponding first prediction vector;
[0072] Step 85: Form a corresponding first prediction-label pair using the first prediction vector and the first label vector corresponding to the current training record; and subject the current first prediction-label pair to a preset first model loss function to calculate and obtain a corresponding first loss value;
[0073] Wherein, the first model loss function is implemented based on L1 loss function, L2 loss function or cross entropy loss function;
[0074] Step 86: Identify whether the first loss value satisfies a preset first loss value range. If the first loss value satisfies the first loss value range, identify whether the current training record is the last first data record in the first training set. If so, proceed to step 87. If not, use the next first data record in the first training set as the new current training record and return to step 84 to continue training. If the first loss value does not satisfy the first loss value range, perform a round of modulation on the model parameters of the breast cancer classification model in a direction that minimizes the first model loss function based on a preset first model optimizer. After this round of modulation, return to step 84 to continue training.
[0075] Wherein, the first model optimizer includes at least an Adam optimizer and an SGD optimizer;
[0076] Step 87, perform a round of traversal on all the first data records of the first evaluation set; and in this round of traversal, use the first data record currently traversed as the corresponding current evaluation record; and identify the training round; if the training round is the first round, then use the first training image, the first training mask map, and the first training feature matrix of the current evaluation record as the current breast ultrasound image, the lesion area mask map, and the genomic feature matrix to input the breast cancer classification model for classification prediction processing and use the classification prediction vector output by this processing as the corresponding second prediction vector; if the training round is the second round, then use the first training image, the first training mask map, and the first training feature matrix of the current evaluation record as the current breast ultrasound image, the lesion area mask map, and the genomic feature matrix to input the breast cancer classification model for classification prediction processing and use the classification prediction vector output by this processing as the corresponding second prediction vector; A training image and the first training feature matrix are used as the current breast ultrasound image and the genomic feature matrix, and the current lesion area mask is set to empty. The current breast ultrasound image, the lesion area mask, and the genomic feature matrix are input into the breast cancer classification model for classification prediction processing, and the classification prediction vector output from this processing is used as the corresponding second prediction vector; the second prediction vector corresponding to the current evaluation record and the first label vector form a corresponding second prediction-label pair; and at the end of this round of traversal, all the obtained second prediction-label pairs are brought into the preset first model evaluation function to calculate and obtain the corresponding first evaluation value;
[0077] Wherein, the first model evaluation function is implemented based on MAE function, MSE function or RMSE function;
[0078] Step 88, identifying whether the first evaluation value meets the preset first evaluation value range; if so, proceeding to step 89; if not, returning to step 83 to continue training;
[0079] Step 89, identify the training round; if the training round is the first round, reset the training round to the second round and return to step 82 for the next round of training; if the training round is the second round, stop training and confirm that the two rounds of model training are completed.
[0080] Preferably, the step of obtaining a corresponding predicted vector of the subject by the breast cancer classification model based on the subject data and feeding it back to the current user specifically includes:
[0081] Identify whether the lesion area mask exists in the subject data; if so, extract the corresponding breast ultrasound image, the lesion area mask and the genomic feature matrix from the subject data; if not, extract the corresponding breast ultrasound image and the genomic feature matrix from the subject data, and set a corresponding lesion area mask to be empty; and input the current breast ultrasound image, the lesion area mask and the genomic feature matrix into the breast cancer classification model for classification prediction processing, and feed back the classification prediction vector output by this processing as the corresponding subject prediction vector to the current user.
[0082] A second aspect of an embodiment of the present invention provides a device for implementing the processing method for predicting breast cancer classification by fusing multimodal features as described in the first aspect above, the device comprising: a model construction module, a model training module, and a model application module;
[0083] The model construction module is used to construct a first encoder for imaging omics feature encoding using a multi-level residual network as a feature extraction backbone, and introduce a multi-level residual feature fusion mechanism and a spatial attention fusion mechanism into the encoder; and to construct a second encoder for genomic feature encoding using a multi-level linear network as a feature extraction backbone, and introduce a noise injection mechanism into the encoder; and to construct a first prediction network for performing three-category prediction of benign, malignant and borderline lesions of breast tumors based on a classifier model composed of a linear network and a Softmax function; and to construct a deep learning model for performing multimodal feature fusion of breast ultrasound imaging omics features and genomic features and classifying and predicting benign, malignant and borderline lesions of breast tumors based on the first and second encoders and the first prediction network, which is recorded as a breast cancer classification model;
[0084] The model training module is used to construct a model training data set recorded as a corresponding first data set; and firstly perform a first round of masked image model training on the breast cancer classification model based on the first data set and then perform a second round of unmasked image model training on the breast cancer classification model based on the first data set;
[0085] The model application module is used to receive subject data input by the user after two rounds of model training are completed; the breast cancer classification model performs prediction based on the subject data to obtain a corresponding subject prediction vector and feeds it back to the current user; the subject data includes breast ultrasound images, lesion area masks, and genomic feature matrices, and the lesion area masks are optional data; the subject prediction vector includes three prediction probabilities, namely, benign prediction probability, malignant prediction probability, and borderline lesion prediction probability.
[0086] A third aspect of an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0087] The processor is configured to be coupled to the memory, read and execute instructions in the memory, so as to implement the method steps described in the first aspect above;
[0088] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
[0089] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a computer, the computer executes the instructions of the method described in the first aspect above.
[0090] Embodiments of the present invention provide a processing method, device, electronic device, and computer-readable storage medium for predicting breast cancer classification by fusing multimodal features. As can be seen from the above content, the embodiment of the present invention first constructs a first encoder for imaging genomics feature encoding using a multi-level residual network as the feature extraction backbone, and introduces a multi-level residual feature fusion mechanism and a spatial attention fusion mechanism into the encoder; and constructs a second encoder for genomics feature encoding using a multi-level linear network as the feature extraction backbone, and introduces a noise injection mechanism into the encoder; and constructs a first prediction network for three-category prediction of benign, malignant and borderline lesions of breast tumors based on a classifier model composed of a linear network and a softmax function; and constructs a breast cancer classification model based on the first, second encoders and the first prediction network for multimodal feature fusion of breast ultrasound imaging genomics features and genomic features and classification and prediction of benign, malignant and borderline lesions of breast tumors based on the fused features; then, a model training dataset is constructed through data collection, and two rounds of model training with / without mask images are performed on the breast cancer classification model based on the dataset; and after the two rounds of model training are completed, the breast cancer classification model is used for prediction application based on the subject data. On the one hand, the embodiments of the present invention improve feature richness and prediction accuracy and reduce the false positive / missed positive rate by introducing multimodal fusion features; on the other hand, by introducing a noise injection mechanism and a random deactivation layer in the second encoder, the robustness, noise resistance and generalization ability of the model are enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 A schematic diagram of a processing method for predicting breast cancer classification by fusing multimodal features provided in Example 1 of the present invention;
[0092] Figure 2 A schematic diagram of a module of a first encoder provided in Embodiment 1 of the present invention;
[0093] Figure 3 A schematic diagram of a module of a second encoder provided in the first embodiment of the present invention;
[0094] Figure 4 A schematic diagram of a module of a first prediction network provided in Example 1 of the present invention;
[0095] Figure 5 This is a module diagram of the breast cancer classification model provided in Example 1 of the present invention;
[0096] Figure 6 A module structure diagram of a processing device for predicting breast cancer classification by fusing multimodal features provided in Example 2 of the present invention;
[0097] Figure 7 This is a structural diagram of an electronic device provided in Example 3 of the present invention. DETAILED DESCRIPTION
[0098] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the embodiments described herein are merely some, rather than all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0099] The first embodiment of the present invention provides a processing method for predicting breast cancer classification by fusing multimodal features, such as Figure 1 A schematic diagram of a processing method for predicting breast cancer classification by fusing multimodal features provided in Example 1 of the present invention is shown. The method mainly includes the following steps:
[0100] Step 1: A first encoder for imaging genomics feature encoding is constructed with a multi-level residual network as the feature extraction backbone, and a multi-level residual feature fusion mechanism and a spatial attention fusion mechanism are introduced into the encoder; a second encoder for genomics feature encoding is constructed with a multi-level linear network as the feature extraction backbone, and a noise injection mechanism is introduced into the encoder; a first prediction network for three-category prediction of benign, malignant and borderline lesions of breast tumors is constructed based on a classifier model composed of a linear network and a Softmax function; and a deep learning model for multimodal feature fusion of breast ultrasound imaging genomics features and genomic features and classification prediction of benign, malignant and borderline lesions of breast tumors based on the first, second encoders and the first prediction network is constructed, which is recorded as a breast cancer classification model.
[0101] Here, the first encoder of the embodiment of the present invention is used to perform radiomics feature encoding processing based on the breast ultrasound image and the lesion area mask image input by the encoder and output the corresponding image feature vector, such as Figure 2 This is a module diagram of the first encoder provided in the first embodiment of the present invention; wherein the lesion area mask image is an optional input.
[0102] The breast ultrasound image of the embodiment of the present invention is a breast B-ultrasound image, a breast Doppler ultrasound image or a breast ultrasound elastic imaging; the image shape of the breast ultrasound image is C P ×H P ×W P , H P 、W P are the image height and width, C P is the pixel feature dimension of the image; H P 、W P They can be equal or unequal.
[0103] When the lesion area mask image of the embodiment of the present invention is not empty, its image shape is 1×H P ×W P ; The lesion area mask is a binary mask obtained by outlining the lesion area on the corresponding breast ultrasound image; the image size of the lesion area mask is consistent with the breast ultrasound image; the pixel feature dimension of the lesion area mask is 1, and each pixel feature is a 0 / 1 binary feature. If it is 0, it means that the current pixel position is not a lesion position, and if it is 1, it means that the current pixel position is a lesion position.
[0104] The image feature vector of the embodiment of the present invention is a one-dimensional feature vector with a shape of 1×L1; wherein L1 is a preset vector length threshold; the vector length threshold L1 defaults to 512; the vector eigenvalue of the image feature vector is constrained within the value range [0,1].
[0105] like Figure 2 As shown, the first input end of the first encoder is used to receive the breast ultrasound image input by the encoder, and the second input end is used to receive the lesion area mask image input by the encoder; the output end of the first encoder is used to output the corresponding image feature vector.
[0106] The internal component modules of the first encoder include: 7×7 convolution layer, maximum pooling layer, first-level residual network, second-level residual network, third-level residual network, upsampling layer, multi-level feature fusion layer, spatial attention network, attention feature fusion layer, and feature mapping layer.
[0107] The connection relationship between the internal components of the first encoder is as follows: the input of the 7×7 convolutional layer is connected to the first input of the first encoder, and the output is connected to the input of the maximum pooling layer; the output of the maximum pooling layer is connected to the input of the first-level residual network; the output of the first-level residual network is connected to the input of the second-level residual network and the first input of the upsampling layer respectively; the output of the second-level residual network is connected to the input of the third-level residual network and the second input of the upsampling layer respectively; the output of the third-level residual network is connected to the third input of the upsampling layer; the output of the upsampling layer is connected to the input of the multi-level feature fusion layer; the output of the multi-level feature fusion layer is connected to the input of the spatial attention network and the first input of the attention feature fusion layer respectively; the output of the spatial attention network is connected to the second input of the attention feature fusion layer; the third input of the attention feature fusion layer is connected to the second input of the first encoder, and the output is connected to the input of the feature mapping layer; the output of the feature mapping layer is connected to the output of the first encoder.
[0108] The functions of the components inside the first encoder are as follows.
[0109] 1) 7×7 convolutional layer:
[0110] The convolution kernel size of the 7×7 convolution layer in the embodiment of the present invention is 7×7; the 7×7 convolution layer is used to downsample and extract features from the breast ultrasound image to obtain a corresponding first feature tensor and send it to the maximum pooling layer;
[0111] Here, the function of the 7×7 convolution layer is consistent with that of the first layer of the ResNet34 residual network; the tensor shape of the first feature tensor is C1×H1×W1, where H1, W1, and C1 are the height, width, and feature dimension of the first feature tensor respectively; H1=H P / 2, W1=H P / 2, the default dimension C1 is 64.
[0112] 2) Maximum pooling layer:
[0113] The maximum pooling layer of the embodiment of the present invention is used to downsample the first feature tensor to obtain a corresponding second feature tensor and send it to the first-level residual network;
[0114] Here, the maximum pooling layer has the same function as the maximum pooling layer immediately following the first 7×7 convolutional layer in the ResNet34 residual network; the tensor shape of the second feature tensor is C2×H2×W2, where H2, W2, and C2 are the height, width, and feature dimension of the second feature tensor, respectively; H2=H1 / 2, W2=H1 / 2, C2=C1; when the preset dimension C1 is 64, the feature dimension C2 is 64.
[0115] 3) First-level residual network:
[0116] The first-level residual network of the embodiment of the present invention is composed of three residual modules connected in sequence, which correspond to the residual modules of the second layer of the ResNet34 residual network; the first-level residual network is used to downsample and extract the second feature tensor to obtain the corresponding third feature tensor and send it to the second-level residual network and the upsampling layer;
[0117] Among them, the tensor shape of the third feature tensor is C3×H3×W3, H3, W3, and C3 are the height, width, and feature dimension of the third feature tensor, respectively; H3=H2 / 2, W3=H2 / 2, C3=C2; when the preset dimension C1 is 64, the feature dimension C3 is 64.
[0118] 4) Second level residual network:
[0119] The second-level residual network of the embodiment of the present invention is composed of four residual modules connected in sequence, and these four residual modules correspond to the residual modules of the third layer of the ResNet34 residual network; the second-level residual network is used to downsample and extract the third feature tensor to obtain the corresponding fourth feature tensor and send it to the third-level residual network and the upsampling layer;
[0120] Among them, the tensor shape of the fourth feature tensor is C4×H4×W4, where H4, W4, and C4 are the height, width, and feature dimension of the fourth feature tensor, respectively; H4=H3 / 2, W4=H3 / 2, C4=2×C3; when the preset dimension C1 is 64, the feature dimension C4 is 128.
[0121] 5) Third-level residual network:
[0122] The third-level residual network of the embodiment of the present invention is composed of six standard residual modules connected in sequence. These six residual modules correspond to the residual modules of the fourth layer of the ResNet34 residual network. The third-level residual network is used to downsample and extract the fourth feature tensor to obtain the corresponding fifth feature tensor and send it to the upsampling layer.
[0123] Among them, the tensor shape of the fifth feature tensor is C5×H5×W5, where H5, W5, and C5 are the height, width, and feature dimension of the fifth feature tensor, respectively; H5=H4 / 2, W5=H4 / 2, C5=2×C4; when the preset dimension C1 is 64, the feature dimension C5 is 256.
[0124] 6) Upsampling layer:
[0125] The upsampling layer of the embodiment of the present invention is used to perform bilinear interpolation and a preset uniform height H. std and uniform width W stdThe third, fourth and fifth feature tensors are upsampled respectively, and the height and width of the three tensors are all the same height H std and uniform width W std The feature tensors of are recorded as the corresponding first, second and third level feature tensors and sent to the multi-level feature fusion layer;
[0126] Among them, the unified height H std and uniform width W std are two pre-set positive integers, which can be equal or unequal; the first, second and third level feature tensors correspond one to one with the third, fourth and fifth feature tensors; the shape of the first level feature tensor is C3×H LV1 ×W LV1 , H LV1 、W LV1 are the height and width of the first-level feature tensor respectively; the shape of the second-level feature tensor is C4×H LV2 ×W LV2 , H LV2 、W LV2 are the height and width of the second-level feature tensor respectively; the shape of the third-level feature tensor is C5×H LV3 ×W LV3 , H LV3 、W LV3 are the height and width of the third-level feature tensor respectively; H LV1 =H LV2 =H LV3 =H std , W LV1 =W LV2 =W LV3 =W std .
[0127] 7) Multi-level feature fusion layer:
[0128] The multi-level feature fusion layer of the embodiment of the present invention is used to perform feature fusion processing on the first, second and third level feature tensors in a feature channel splicing manner to obtain a corresponding first fused tensor and send it to the spatial attention network and the attention feature fusion layer;
[0129] The tensor shape of the first fused tensor is C6×H6×W6, where H6, W6, and C6 are the height, width, and feature dimension of the first fused tensor respectively; H6=H std , W6=W std , C6=C3+C4+C5; when the preset dimension C1 is 64, the feature dimension C6 is 448.
[0130] 8) Spatial Attention Network:
[0131] The spatial attention network of the embodiment of the present invention first performs feature dimensionality reduction on the first fusion tensor through a 3×3 convolution layer to obtain a corresponding first reduced dimensionality tensor; then performs nonlinear activation on the first reduced dimensionality tensor through a ReLU function to obtain a corresponding first activation tensor; then aggregates the feature channels of the first activation tensor through a 1×1 convolution layer to obtain a corresponding second reduced dimensionality tensor; then normalizes the second reduced dimensionality tensor through a Sigmoid function and uses the processing result as the corresponding spatial attention weight matrix; finally, the normalized second reduced dimensionality tensor is sent to the attention feature fusion layer;
[0132] Among them, the tensor shape of the first dimensionality reduction tensor is C7×H7×W7, H7, W7, and C7 are the height, width, and feature dimension of the first dimensionality reduction tensor respectively, and H7=H std , W7=W std , C7=C6 / n; the preset dimensionality reduction factor n is a positive integer greater than or equal to 2 and divisible by the feature dimension C6; the tensor shape of the first activation tensor is C7×H7×W7; the tensor shape of the second dimensionality reduction tensor is C8×H8×W8, where H8, W8, and C8 are the height, width, and feature dimension of the first dimensionality reduction tensor respectively, and H8=H std , W8=W std , C8=1; the matrix shape of the spatial attention weight matrix is H std ×W std .
[0133] 9) Attention feature fusion layer:
[0134] The attention feature fusion layer is used to identify whether the lesion mask input by the encoder is empty; if the lesion mask is empty, the spatial attention weight matrix is copied C6 times based on the feature dimension of the first fusion tensor to obtain a C6×H std ×W std The first weight tensor of ; If the lesion area mask is not empty, then the bilinear interpolation algorithm and the uniform height H std and uniform width W std The lesion area mask is downsampled to obtain a shape of H std ×W std The down-sampled mask map is obtained by performing a Hadamard product operation on the down-sampled mask map and the spatial attention weight matrix, and the result of the operation is used as the corresponding first weight matrix. Based on the feature dimension of the first fusion tensor, the first weight matrix is replicated C6 times to obtain a C6×H matrix. std ×W std and performing a Hadamard product operation on the first weight tensor and the first fusion tensor and sending the operation result as the corresponding second fusion tensor to the feature mapping layer;
[0135] Among them, the tensor shape of the second fused tensor is C8×H8×W8, H8, W8, and C8 are the height, width, and feature dimension of the second fused tensor respectively, and H8=H std , W8=W std , C8=C6.
[0136] 10) Feature Mapping Layer:
[0137] The feature mapping layer first performs feature dimension upscaling on the second fusion tensor through a 1×1 convolutional layer to obtain the corresponding first dimension upscaling tensor; then performs batch normalization (BN) processing on the first dimension upscaling tensor through a batch normalization layer to obtain the corresponding first normalized tensor; then performs nonlinear activation on the first normalized tensor through the ReLU function to obtain the corresponding second activation tensor; then performs global average pooling processing on the second activation tensor through a global average pooling layer and outputs the processing result as the corresponding image feature vector;
[0138] Among them, the tensor shapes of the first dimension-raising tensor, the first normalized tensor, and the second activation tensor are all C9×H9×W9, where H9, W9, and C9 are the height, width, and feature dimensions of the first dimension-raising tensor, respectively. H9=H std , W9=W std , C9=L1; the shape of the image feature vector is 1×L1, and its vector eigenvalue is constrained to be within the range [0,1].
[0139] The second encoder of the embodiment of the present invention is used to perform genomic feature encoding processing according to the genomic feature matrix input by the encoder and output the corresponding gene feature vector, such as Figure 3 This is a module diagram of the second encoder provided in Example 1 of the present invention.
[0140] Each row of the genomic feature matrix of the embodiment of the present invention corresponds to a type of gene, and each column corresponds to a type of gene feature; the matrix shape of the genomic feature matrix is H G ×W G , H G 、W G are the height and width of the matrix, H G 、W G Match the total number of genes and total number of gene features in the genomic feature matrix, respectively.
[0141] The gene feature vector of the embodiment of the present invention is a one-dimensional feature vector with a shape of 1×L2; L2 is a preset vector length threshold; the vector length threshold L2 defaults to 32; the vector eigenvalue of the gene feature vector is constrained within the value range [-1,1].
[0142] like Figure 3As shown, the input end of the second encoder is used to receive the genomic feature matrix input by the encoder, and the output end is used to output the corresponding gene feature vector.
[0143] The internal component modules of the second encoder include: a first linear layer, a first activation layer, a first normalization layer, a second linear layer, a second activation layer, a third linear layer, a noise injection layer, and a feature truncation layer.
[0144] The connection relationship between the internal components of the second encoder is as follows: the input end of the first linear layer is connected to the input end of the second encoder, and the output end is connected to the input end of the first activation layer; the output end of the first activation layer is connected to the input end of the first normalization layer; the output end of the first normalization layer is connected to the input end of the second linear layer; the output end of the second linear layer is connected to the input end of the second activation layer; the output end of the second activation layer is connected to the input end of the third linear layer; the output end of the third linear layer is connected to the input end of the noise injection layer; the output end of the noise injection layer is connected to the input end of the feature truncation layer; and the output end of the feature truncation layer is connected to the output end of the second encoder.
[0145] The functions of the components inside the second encoder are as follows:
[0146] 1) The first linear layer is used to flatten the genomic feature matrix into a matrix of length H G ×W G The one-dimensional vector is recorded as the first vector; and the first vector is fully connected to obtain the corresponding first fully connected vector and sent to the first activation layer;
[0147] 2) The first activation layer is used to perform nonlinear activation on the first fully connected vector through the ReLU function to obtain the corresponding first activation vector and send it to the first normalization layer;
[0148] 3) The first normalization layer is used to perform layer normalization (Layer Normalization, Layer Norm) processing on the first activation vector to obtain the corresponding first normalized vector and send it to the second linear layer;
[0149] 4) The second linear layer is used to perform a full connection calculation on the first normalized vector to obtain a corresponding second fully connected vector and send it to the second activation layer;
[0150] 5) The second activation layer is used to perform nonlinear activation on the second fully connected vector through the Tanh function to obtain the corresponding second activation vector and send it to the third linear layer; the vector eigenvalue of the second activation vector is constrained to be within the value range [-1, 1];
[0151] 6) The third linear layer is used to perform a full connection calculation on the second activation vector to obtain a corresponding third fully connected vector and send it to the noise injection layer;
[0152] Among them, the third fully connected vector is a one-dimensional feature vector with a shape of 1×L2;
[0153] 7) The noise injection layer is used to construct a standard normal distribution noise vector with a vector length of L2 according to the model parameter ε, recorded as the first noise vector; and the first noise vector and the third fully connected vector are added to obtain the corresponding noise injection vector and sent to the feature truncation layer;
[0154] Among them, the mean of the first noise vector is 0 and the variance is ε 2 ; The model parameter ε is a learnable model parameter, and its initial value is set to the preset initialization noise coefficient ε0;
[0155] 8) The feature truncation layer is used to perform nonlinear activation on the noise injection vector through the Tanh function to obtain the corresponding gene feature vector and output it;
[0156] As mentioned above, the gene feature vector is a one-dimensional feature vector with a shape of 1×L2, and its vector eigenvalue is constrained to be within the range [-1,1].
[0157] The first prediction network of the embodiment of the present invention is used to perform three-category prediction processing on benign, malignant and borderline lesions of breast tumors based on the fusion feature vector input by the network and output the corresponding classification prediction vector, such as Figure 4 This is a module diagram of the first prediction network provided in Example 1 of the present invention.
[0158] The fused feature vector of the embodiment of the present invention is a one-dimensional feature vector with a shape of 1×(L1+L2); the first L1 vector eigenvalues of the fused feature vector are constrained to be within the value range [0,1], and the last L2 vector eigenvalues are constrained to be within the value range [-1,1].
[0159] The classification prediction vector of the embodiment of the present invention includes three prediction probabilities, namely, benign prediction probability, malignant prediction probability, and borderline lesion prediction probability.
[0160] like Figure 4 As shown, the network input end of the first prediction network is used to receive the fusion feature vector input by the network, and the network output end is used to output the corresponding classification prediction vector.
[0161] The internal component modules of the first prediction network include: a fourth linear layer, a second normalization layer, a third activation layer, a random inactivation layer, a fifth linear layer, and a Softmax function layer.
[0162] The connection relationship between the internal components of the first prediction network is as follows: the input end of the fourth linear layer is connected to the input end of the first prediction network, and the output end is connected to the input end of the second normalization layer; the output end of the second normalization layer is connected to the input end of the third activation layer; the output end of the third activation layer is connected to the input end of the random deactivation layer; the output end of the random deactivation layer is connected to the input end of the fifth linear layer; and the output end of the fifth linear layer is connected to the input end of the Softmax function layer.
[0163] The functions of the components within the first prediction network are as follows:
[0164] 1) The fourth linear layer is used to perform a full connection calculation on the fused feature vector to obtain the corresponding fourth fully connected vector and send it to the second normalization layer;
[0165] Here, when L1 is 512 and L2 is 32, the vector length (L1+L2) of the fused feature vector is 544. The purpose of the fourth linear layer is actually to perform a short vector mapping on this fused feature vector; for example, mapping from 544 to 256, then the vector length of the fourth fully connected vector is 256;
[0166] 2) The second normalization layer is used to perform batch normalization on the fourth fully connected vector to obtain the corresponding second normalized vector and send it to the third activation layer;
[0167] Here, the vector length of the second normalized vector is consistent with the fourth fully connected vector;
[0168] 3) The third activation layer is used to perform nonlinear activation on the second normalized vector through the ReLU function to obtain the corresponding third activation vector and send it to the random inactivation layer;
[0169] Here, the vector length of the third activation vector is consistent with the second normalized vector;
[0170] 4) The random deactivation layer is used to perform random deactivation processing on the vector features in the third activation vector in equal proportion according to a preset random deactivation ratio to obtain a corresponding local deactivation vector and send it to the fifth linear layer;
[0171] The random deactivation ratio is a preset ratio parameter, such as 30%; the local deactivation vector has the same vector shape as the third activation vector, but some vector features in the local deactivation vector are set to preset deactivation feature values, and the ratio of the total number of deactivation feature values of the local deactivation vector to the vector length satisfies the random deactivation ratio; the deactivation feature value is a preset value and can be set to 0 by default;
[0172] 5) The fifth linear layer is used to perform full connection calculation on the local deactivation vector to obtain the corresponding fifth fully connected vector and send it to the Softmax function layer;
[0173] Among them, the vector length of the fifth fully connected vector is 3;
[0174] Here, the purpose of the fifth linear layer is actually to perform a short vector mapping on this local deactivation vector; for example, mapping from 256 to 3, the vector length of the fifth fully connected vector is 3;
[0175] 6) The Softmax function layer is used to perform three-category probability calculation based on the fifth fully connected vector through the Softmax function to obtain the corresponding benign prediction probability, malignant prediction probability and borderline lesion prediction probability to form the corresponding classification prediction vector and output it.
[0176] The breast cancer classification model of the embodiment of the present invention is used to classify and predict benign, malignant and borderline lesions of breast tumors based on the breast ultrasound image, lesion area mask and genomic feature matrix input by the model and output the corresponding classification prediction vector, such as Figure 5 This is a module diagram of the breast cancer classification model provided in Example 1 of the present invention.
[0177] The model inputs of the breast cancer classification model include breast ultrasound images, lesion area masks, and genomic feature matrices, where the lesion area masks are optional inputs. The model output of the breast cancer classification model is a classification prediction vector. As mentioned above, this classification prediction vector includes three prediction probabilities: benign prediction probability, malignant prediction probability, and borderline lesion prediction probability.
[0178] like Figure 5 As shown, the first model input end of the breast cancer classification model is used to receive breast ultrasound images and lesion area masks, the second model input end is used to receive genomic feature matrices, and the model output end is used to output corresponding classification prediction vectors.
[0179] The internal component modules of the breast cancer classification model include: a first encoder, a second encoder, a multimodal feature fusion module and a first prediction network.
[0180] The connection relationship between the internal components of the breast cancer classification model is as follows: the encoder input of the first encoder is connected to the first model input, and the encoder output is connected to the first input of the multimodal feature fusion module; the encoder input of the second encoder is connected to the second model input, and the encoder output is connected to the second input of the multimodal feature fusion module; the output of the multimodal feature fusion module is connected to the network input of the first prediction network; and the network output of the first prediction network is connected to the model output.
[0181] The functions of the components within the breast cancer classification model are as follows:
[0182] 1) As described above, the first encoder is used to perform radiomics feature encoding processing based on the breast ultrasound image and lesion area mask image input by the encoder to obtain the corresponding image feature vector and send it to the multimodal feature fusion module;
[0183] 2) As described above, the second encoder is used to perform genomic feature encoding processing based on the genomic feature matrix input by the encoder to obtain the corresponding gene feature vector and send it to the multimodal feature fusion module;
[0184] 3) The multimodal feature fusion module is used to sequentially concatenate the image feature vector and the gene feature vector to obtain a one-dimensional feature vector with a shape of 1×(L1+L2) as the corresponding fused feature vector; and send the fused feature vector to the first prediction network;
[0185] 4) As mentioned above, the first prediction network is used to perform three-category prediction processing on benign, malignant and borderline lesions of breast tumors based on the fused feature vector and output the corresponding classification prediction vector.
[0186] Step 2: construct a model training data set and record it as the corresponding first data set; firstly perform a first round of mask image model training on the breast cancer classification model based on the first data set, and then perform a second round of non-mask image model training on the breast cancer classification model based on the first data set;
[0187] Specifically comprising: step 21, constructing a model training data set and recording it as the corresponding first data set;
[0188] Among them, the first data set includes multiple first data records; each first data record corresponds to an acquisition object; the first data record includes a first training image, a first training mask image, a first training feature matrix and a first label vector; the first training image is a breast ultrasound image of the current acquisition object; the first training mask image is a lesion area mask image of the current acquisition object; the first training feature matrix is a genomic feature matrix of the current acquisition object; the first label vector includes three label probabilities, namely, a benign label probability, a malignant label probability and a borderline lesion label probability, and only one of the three label probabilities is 1, and the other two label probabilities are 0; all first training images of the first data set have the same image type, specifically, breast B-ultrasound image, breast Doppler ultrasound image or breast ultrasound elastography;
[0189] Specifically, the method includes: step 211, selecting a plurality of patients with benign breast tumors, malignant breast tumors, and borderline breast lesions to form a collection subject group;
[0190] Wherein, the collection object group includes multiple collection objects;
[0191] Step 212: Select one from the three image types of breast B-ultrasound image, breast Doppler ultrasound image, and breast ultrasound elastography as the corresponding current ultrasound image type; and use each collected object as the corresponding current object; and perform a breast ultrasound examination on the current object based on the current ultrasound image type, and use the breast ultrasound image obtained in this examination as a corresponding first training image; and perform a whole genome sequencing on the current object to obtain the corresponding current object gene sequence, and fill all matrix cells of the genomic feature matrix with feature data based on the current object gene sequence and all gene types and all gene features specified by the genomic feature matrix, and use the filled genomic feature matrix as the corresponding first training feature matrix; and use the preset breast lesion delineation interface in the first training image. The breast lesion area of the current object is outlined on the training image and a lesion area mask is generated based on the outline result as the corresponding first training mask; the current object is identified; if the current object is a benign breast tumor patient, a first label vector is set with a benign label probability of 1 and a malignant and borderline lesion label probability of 0; if the current object is a malignant breast tumor patient, a first label vector is set with a malignant label probability of 1 and a benign and borderline lesion label probability of 0; if the current object is a borderline breast lesion patient, a first label vector is set with a borderline lesion label probability of 1 and a benign and malignant label probability of 0; and a corresponding first data record is formed by the first training image corresponding to the current object, the first training mask, the first training feature matrix and the first label vector;
[0192] Here, the breast lesion delineation interface is a pre-set mask image processing interface, which can be a manual annotation interface, a machine annotation interface, or an input interface of another image semantic segmentation model;
[0193] Step 213: All the first data records obtained are combined into a corresponding first data set;
[0194] Step 22: firstly perform a first round of mask image model training on the breast cancer classification model based on the first data set, and then perform a second round of non-mask image model training on the breast cancer classification model based on the first data set;
[0195] Specifically comprising: step 221, setting the training round to the first round;
[0196] Step 222: randomly split the first data set into two sub-data sets based on a preset first split ratio and record them as a first training set and a first evaluation set;
[0197] Here, the first split ratio is a preset ratio parameter, such as 8:2; the first training set and the first evaluation set both include a plurality of first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first split ratio;
[0198] Step 223, extracting the first first data record of the first training set as the corresponding current training record;
[0199] Step 224: Identify the training round. If the training round is the first round, input the first training image, the first training mask, and the first training feature matrix of the current training record as the current breast ultrasound image, the lesion area mask, and the genomic feature matrix into the breast cancer classification model for classification prediction processing, and use the classification prediction vector output from this processing as the corresponding first prediction vector. If the training round is the second round, input the first training image and the first training feature matrix of the current training record as the current breast ultrasound image and the genomic feature matrix, set the current lesion area mask to empty, and input the current breast ultrasound image, the lesion area mask, and the genomic feature matrix into the breast cancer classification model for classification prediction processing, and use the classification prediction vector output from this processing as the corresponding first prediction vector.
[0200] Step 225: The first prediction vector and the first label vector corresponding to the current training record are used to form a corresponding first prediction-label pair; and the current first prediction-label pair is substituted into a preset first model loss function to calculate a corresponding first loss value.
[0201] The first model loss function is implemented based on the L1 loss function, the L2 loss function or the cross entropy loss function;
[0202] Step 226: Identify whether the first loss value satisfies a preset first loss value range. If the first loss value satisfies the first loss value range, identify whether the current training record is the last first data record of the first training set. If so, proceed to step 227. If not, use the next first data record of the first training set as the new current training record and return to step 224 to continue training. If the first loss value does not satisfy the first loss value range, perform a round of modulation on the model parameters of the breast cancer classification model based on a preset first model optimizer in a direction that minimizes the first model loss function. After this round of modulation, return to step 224 to continue training.
[0203] The first loss value range is a preset numerical range; the first model optimizer includes at least an Adam optimizer and an SGD optimizer;
[0204] Step 227, perform a round of traversal on all the first data records of the first evaluation set; and in this round of traversal, use the first data record currently traversed as the corresponding current evaluation record; and identify the training round; if the training round is the first round, use the first training image, the first training mask, and the first training feature matrix of the current evaluation record as the current breast ultrasound image, the lesion area mask, and the genomic feature matrix to input the breast cancer classification model for classification prediction processing and use the classification prediction vector output by this processing as the corresponding second prediction vector; if the training round is the second round, use the first training image, the first training mask, and the first training feature matrix of the current evaluation record as the current breast ultrasound image, the lesion area mask, and the genomic feature matrix to input the breast cancer classification model for classification prediction processing and use the classification prediction vector output by this processing as the corresponding second prediction vector; The training image and the first training feature matrix are used as the current breast ultrasound image and genomic feature matrix, and the current lesion area mask is set to empty. The current breast ultrasound image, lesion area mask, and genomic feature matrix are input into the breast cancer classification model for classification prediction processing, and the classification prediction vector output by this processing is used as the corresponding second prediction vector; and the second prediction vector corresponding to the current evaluation record and the first label vector are used to form a corresponding second prediction-label pair; and at the end of this round of traversal, all the obtained second prediction-label pairs are brought into the preset first model evaluation function to calculate and obtain the corresponding first evaluation value;
[0205] Wherein, the first model evaluation function is implemented based on the MAE function, the MSE function or the RMSE function;
[0206] Step 228: Identify whether the first evaluation value meets the preset first evaluation value range; if so, go to step 229; if not, return to step 223 to continue training;
[0207] Here, the first evaluation value range is a preset value range;
[0208] Step 229, identify the training round; if the training round is the first round, reset the training round to the second round and return to step 222 for the next round of training; if the training round is the second round, stop training and confirm that the two rounds of model training are completed.
[0209] Step 3: After the two rounds of model training are completed, the user's inputted subject data is received; the breast cancer classification model makes a prediction based on the subject data to obtain the corresponding subject prediction vector and feeds it back to the current user;
[0210] Specifically comprising: step 31, after two rounds of model training are completed, receiving the subject data input by the user;
[0211] The data of the subjects include breast ultrasound images, lesion area mask maps and genomic feature matrix. The lesion area mask map is optional data.
[0212] Step 32: The breast cancer classification model makes a prediction based on the subject's data to obtain a corresponding subject prediction vector and feeds it back to the current user;
[0213] The predicted vector of the subject includes three predicted probabilities, namely, benign prediction probability, malignant prediction probability and borderline lesion prediction probability;
[0214] Specifically, it includes: identifying whether there is a lesion area mask map in the subject's data; if so, extracting the corresponding breast ultrasound image, lesion area mask map and genomic feature matrix from the subject's data; if not, extracting the corresponding breast ultrasound image and genomic feature matrix from the subject's data, and setting a corresponding lesion area mask map to empty; and inputting the current breast ultrasound image, lesion area mask map and genomic feature matrix into the breast cancer classification model for classification prediction processing and feeding back the classification prediction vector output by this processing as the corresponding subject prediction vector to the current user.
[0215] Figure 6 This is a module structure diagram of a processing device for predicting breast cancer classification by integrating multimodal features provided in the second embodiment of the present invention. The device is a terminal device or server that implements the aforementioned method embodiment, or can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiment. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 6 As shown, the device includes: a model construction module 201, a model training module 202, and a model application module 203.
[0216] The model construction module 201 is used to construct a first encoder for imaging genomics feature encoding using a multi-level residual network as the feature extraction backbone, and introduce a multi-level residual feature fusion mechanism and a spatial attention fusion mechanism into the encoder; and to construct a second encoder for genomics feature encoding using a multi-level linear network as the feature extraction backbone, and introduce a noise injection mechanism into the encoder; and to construct a first prediction network for three-category prediction of benign, malignant and borderline lesions of breast tumors based on a classifier model composed of a linear network and a Softmax function; and to construct a deep learning model for multimodal feature fusion of breast ultrasound imaging genomics features and genomic features and classification prediction of benign, malignant and borderline lesions of breast tumors based on the first and second encoders and the first prediction network, which is recorded as a breast cancer classification model.
[0217] The model training module 202 is used to construct a model training data set recorded as the corresponding first data set; and first perform a first round of mask image model training on the breast cancer classification model based on the first data set and then perform a second round of non-mask image model training on the breast cancer classification model based on the first data set.
[0218] The model application module 203 is used to receive the subject data input by the user after two rounds of model training are completed; the breast cancer classification model predicts the subject data based on the subject data to obtain the corresponding subject prediction vector and feeds it back to the current user; the subject data includes breast ultrasound images, lesion area masks, and genomic feature matrices, with the lesion area masks being optional data; the subject prediction vector includes three prediction probabilities: benign prediction probability, malignant prediction probability, and borderline lesion prediction probability.
[0219] An embodiment of the present invention provides a processing device for predicting breast cancer classification by fusing multimodal features, which can execute the method steps in the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.
[0220] It should be noted that the division of the modules of the above devices is merely a division of logical functions. In actual implementation, they can be fully or partially integrated into a physical entity or physically separated. Furthermore, these modules can all be implemented in the form of software called by a processing element; or all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the model building module can be a separate processing element, or it can be integrated into a chip of the above device. In addition, it can be stored in the form of program code in the memory of the above device, and called by a processing element of the above device to perform the functions of the above-mentioned module. The implementation of other modules is similar. In addition, these modules can all or partly be integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed by hardware integrated logic circuits in the processor element or instructions in the form of software.
[0221] For example, the above modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, when a module is implemented by scheduling program code through a processing element, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0222] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the above method embodiments are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The above-mentioned computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above-mentioned computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.) means. The above-mentioned computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The above-mentioned available medium can be a magnetic medium (such as a floppy disk, hard disk, tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0223] Figure 7 This is a schematic diagram of the structure of an electronic device provided in the third embodiment of the present invention. The electronic device can be a terminal device or server that implements the method of the aforementioned embodiment, or it can be a terminal device or server that implements the method of the aforementioned embodiment connected to the aforementioned terminal device or server. Figure 7As shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver 303's transceiver actions. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the aforementioned embodiment method. Preferably, the electronic device involved in the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The above-mentioned communication port 306 is used for connecting and communicating between the electronic device and other peripherals.
[0224] exist Figure 7 The system bus 305 mentioned in the figure can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to realize communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include random access memory (RAM) and may also include non-volatile memory (Non-Volatil e Memory), such as at least one disk storage.
[0225] The above-mentioned processors can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0226] It should be noted that an embodiment of the present invention further provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the methods and processing procedures provided in the above embodiments.
[0227] Embodiments of the present invention provide a processing method, device, electronic device, and computer-readable storage medium for predicting breast cancer classification by fusing multimodal features. As can be seen from the above content, the embodiment of the present invention first constructs a first encoder for imaging genomics feature encoding using a multi-level residual network as the feature extraction backbone, and introduces a multi-level residual feature fusion mechanism and a spatial attention fusion mechanism into the encoder; and constructs a second encoder for genomics feature encoding using a multi-level linear network as the feature extraction backbone, and introduces a noise injection mechanism into the encoder; and constructs a first prediction network for three-category prediction of benign, malignant and borderline lesions of breast tumors based on a classifier model composed of a linear network and a softmax function; and constructs a breast cancer classification model based on the first, second encoders and the first prediction network for multimodal feature fusion of breast ultrasound imaging genomics features and genomic features and classification and prediction of benign, malignant and borderline lesions of breast tumors based on the fused features; then, a model training dataset is constructed through data collection, and two rounds of model training with / without mask images are performed on the breast cancer classification model based on the dataset; and after the two rounds of model training are completed, the breast cancer classification model is used for prediction application based on the subject data. On the one hand, the embodiments of the present invention improve feature richness and prediction accuracy and reduce the false positive / missed positive rate by introducing multimodal fusion features; on the other hand, by introducing a noise injection mechanism and a random deactivation layer in the second encoder, the robustness, noise resistance and generalization ability of the model are enhanced.
[0228] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0229] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for predicting breast cancer classification by integrating multimodal features, characterized in that: The method comprises: A first encoder for imaging omics feature encoding is constructed using a multi-level residual network as the feature extraction backbone, and a multi-level residual feature fusion mechanism and a spatial attention fusion mechanism are introduced into the encoder; a second encoder for genomic feature encoding is constructed using a multi-level linear network as the feature extraction backbone, and a noise injection mechanism is introduced into the encoder; and a first prediction network for three-category prediction of benign, malignant and borderline lesions of breast tumors is constructed based on a classifier model composed of a linear network and a softmax function; and a deep learning model for multimodal feature fusion of breast ultrasound imaging omics features and genomic features and classification prediction of benign, malignant and borderline lesions of breast tumors based on the fused features is constructed based on the first and second encoders and the first prediction network, which is recorded as a breast cancer classification model; Constructing a model training data set as a corresponding first data set; firstly performing a first round of mask image model training on the breast cancer classification model based on the first data set, and then performing a second round of non-mask image model training on the breast cancer classification model based on the first data set; After two rounds of model training, the subject data input by the user is received; the breast cancer classification model performs prediction based on the subject data to obtain a corresponding subject prediction vector and feeds it back to the current user; the subject data includes breast ultrasound images, lesion area masks, and genomic feature matrices, and the lesion area masks are optional data; the subject prediction vector includes three prediction probabilities, namely, benign prediction probability, malignant prediction probability, and borderline lesion prediction probability.
2. The method for predicting breast cancer classification by fusing multimodal features according to claim 1, characterized in that: The first encoder is used to perform radiomics feature encoding processing according to the breast ultrasound image and the lesion area mask image input by the encoder and output a corresponding image feature vector; The breast ultrasound image is a breast B-ultrasound image, a breast Doppler ultrasound image or a breast ultrasound elastic imaging; the image shape of the breast ultrasound image is C P ×H P ×W P , H P 、W P are the image height and width, C P is the pixel feature dimension of the image; The image shape when the lesion area mask is not empty is 1×H P ×W P The lesion area mask is a binary mask obtained by outlining the lesion area on the corresponding breast ultrasound image; the image size of the lesion area mask is consistent with the breast ultrasound image, and the pixel feature dimension of the lesion area mask is 1; each pixel feature of the lesion area mask is a 0 / 1 binary feature, if it is 0, it indicates that the current pixel position is a non-lesion position, and if it is 1, it indicates that the current pixel position is a lesion position; The image feature vector is a one-dimensional feature vector with a shape of 1×L1; L1 is a preset vector length threshold; the vector length threshold L1 defaults to 512; the vector eigenvalue of the image feature vector is constrained to be within the value range [0,1]; The first input end of the first encoder is used to receive the breast ultrasound image input by the encoder, and the second input end is used to receive the lesion area mask image input by the encoder; the output end of the first encoder is used to output the corresponding image feature vector; The first encoder includes a 7×7 convolutional layer, a maximum pooling layer, a first-level residual network, a second-level residual network, a third-level residual network, an upsampling layer, a multi-level feature fusion layer, a spatial attention network, an attention feature fusion layer, and a feature mapping layer; The input end of the 7×7 convolutional layer is connected to the first input end of the first encoder, and the output end is connected to the input end of the maximum pooling layer; the output end of the maximum pooling layer is connected to the input end of the first-level residual network; the output end of the first-level residual network is connected to the input end of the second-level residual network and the first input end of the upsampling layer respectively; The output end of the second-stage residual network is connected to the input end of the third-stage residual network and the second input end of the upsampling layer respectively; The output end of the third-level residual network is connected to the third input end of the upsampling layer; the output end of the upsampling layer is connected to the input end of the multi-level feature fusion layer; the output end of the multi-level feature fusion layer is respectively connected to the input end of the spatial attention network and the first input end of the attention feature fusion layer; the output end of the spatial attention network is connected to the second input end of the attention feature fusion layer; the third input end of the attention feature fusion layer is connected to the second input end of the first encoder, and the output end is connected to the input end of the feature mapping layer; the output end of the feature mapping layer is connected to the output end of the first encoder; The convolution kernel size of the 7×7 convolutional layer is 7×7; the 7×7 convolutional layer is used to downsample and extract features from the breast ultrasound image to obtain a corresponding first feature tensor and send it to the maximum pooling layer; the tensor shape of the first feature tensor is C1×H1×W1, where H1, W1, and C1 are the height, width, and feature dimension of the first feature tensor, respectively; H1=H P / 2, W1=H P / 2, the preset dimension C1 defaults to 64; The maximum pooling layer is used to downsample the first feature tensor to obtain a corresponding second feature tensor and send it to the first-level residual network; the tensor shape of the second feature tensor is C2×H2×W2, where H2, W2, and C2 are the height, width, and feature dimension of the second feature tensor respectively; H2=H1 / 2, W2=H1 / 2, and C2=C1; when the preset dimension C1 is 64, the feature dimension C2 is 64; The first-level residual network is composed of three residual modules connected in sequence; the first-level residual network is used to downsample and extract features from the second feature tensor to obtain a corresponding third feature tensor, which is sent to the second-level residual network and the upsampling layer; the tensor shape of the third feature tensor is C3×H3×W3, where H3, W3, and C3 are the height, width, and feature dimension of the third feature tensor, respectively; H3=H2 / 2, W3=H2 / 2, and C3=C2; when the preset dimension C1 is 64, the feature dimension C3 is 64; The second-level residual network is composed of four residual modules connected in sequence; the second-level residual network is used to downsample and extract features from the third feature tensor to obtain a corresponding fourth feature tensor, which is sent to the third-level residual network and the upsampling layer; the tensor shape of the fourth feature tensor is C4×H4×W4, where H4, W4, and C4 are the height, width, and feature dimension of the fourth feature tensor, respectively; H4=H3 / 2, W4=H3 / 2, and C4=2×C3; when the preset dimension C1 is 64, the feature dimension C4 is 128; The third-level residual network is composed of six residual modules connected in sequence; the third-level residual network is used to downsample and extract features from the fourth feature tensor to obtain a corresponding fifth feature tensor and send it to the upsampling layer; the tensor shape of the fifth feature tensor is C5×H5×W5, where H5, W5, and C5 are the height, width, and feature dimension of the fifth feature tensor, respectively; H5=H4 / 2, W5=H4 / 2, and C5=2×C4; when the preset dimension C1 is 64, the feature dimension C5 is 256; The upsampling layer is used to perform bilinear interpolation and a preset uniform height H std and uniform width W std The third, fourth and fifth feature tensors are upsampled respectively, and the height and width of the three tensors are all the same as the unified height H std and the uniform width W std The feature tensors of are recorded as the corresponding first, second and third level feature tensors and sent to the multi-level feature fusion layer; the first, second and third level feature tensors correspond one to one with the third, fourth and fifth feature tensors; the tensor shape of the first level feature tensor is C3×H LV1 ×W LV1 , H LV1 、W LV1 are the height and width of the first-level feature tensor respectively; the shape of the second-level feature tensor is C4×H LV2 ×W LV2 , H LV2 、W LV2 are the height and width of the second-level feature tensor respectively; the tensor shape of the third-level feature tensor is C5×H LV3 ×W LV3 , H LV3 、W LV3 are the height and width of the third-level feature tensor respectively; H LV1 =H LV2 =H LV3 =H std , W LV1 =W LV2 =W LV3 =W std ; The multi-level feature fusion layer is used to perform feature fusion processing on the first, second and third level feature tensors in a feature channel splicing manner to obtain the corresponding first fusion tensor and send it to the spatial attention network and the attention feature fusion layer; the tensor shape of the first fusion tensor is C6×H6×W6, where H6, W6 and C6 are the height, width and feature dimension of the first fusion tensor respectively; H6=H std , W6=W std , C6=C3+C4+C5; when the preset dimension C1 is 64, the feature dimension C6 is 448; The spatial attention network first performs feature dimensionality reduction on the first fusion tensor through a 3×3 convolutional layer to obtain a corresponding first reduced dimensionality tensor; then performs nonlinear activation on the first reduced dimensionality tensor through a ReLU function to obtain a corresponding first activated tensor; Then, a 1×1 convolutional layer is used to aggregate the feature channels of the first activation tensor to obtain the corresponding second reduced dimensionality tensor; the second reduced dimensionality tensor is then normalized using a Sigmoid function and the processing result is used as the corresponding spatial attention weight matrix; finally, the normalized second reduced dimensionality tensor is sent to the attention feature fusion layer; The tensor shape of the first dimensionality reduction tensor is C7×H7×W7, where H7, W7, and C7 are the height, width, and feature dimension of the first dimensionality reduction tensor, respectively. std , W7=W std , C7=C6 / n; the preset dimensionality reduction multiple n is a positive integer greater than or equal to 2 and divisible by the feature dimension C6; the tensor shape of the first activation tensor is C7×H7×W7; the tensor shape of the second dimensionality reduction tensor is C8×H8×W8, where H8, W8, and C8 are the height, width, and feature dimension of the first dimensionality reduction tensor respectively, and H8=H std , W8=W std , C8=1; the matrix shape of the spatial attention weight matrix is H std ×W std ; The attention feature fusion layer is used to identify whether the lesion area mask input by the encoder is empty; if the lesion area mask is empty, the spatial attention weight matrix is copied C6 times based on the feature dimension of the first fusion tensor to obtain a C6×H std ×W std The first weight tensor of ; if the lesion area mask map is not empty, then according to the bilinear interpolation algorithm and the unified height H std and the uniform width W std The lesion area mask image is downsampled to obtain a shape of H std ×W std The down-sampling mask map is obtained by performing a Hadamard product operation on the down-sampling mask map and the spatial attention weight matrix, and the operation result is used as the corresponding first weight matrix, and the first weight matrix is copied C6 times based on the feature dimension of the first fusion tensor to obtain a C6×H std ×W std The first weight tensor of the first fusion tensor is obtained by performing a Hadamard product operation on the first weight tensor and the first fusion tensor and sending the result of the operation as the corresponding second fusion tensor to the feature mapping layer; the tensor shape of the second fusion tensor is C8×H8×W8, where H8, W8, and C8 are the height, width, and feature dimension of the second fusion tensor respectively, and H8=H std , W8=W std , C8=C6; The feature mapping layer first performs feature dimension upscaling on the second fused tensor through a 1×1 convolutional layer to obtain a corresponding first dimension upscaling tensor; then performs batch normalization processing on the first dimension upscaling tensor through a batch normalization layer to obtain a corresponding first normalized tensor; then performs nonlinear activation on the first normalized tensor through a ReLU function to obtain a corresponding second activation tensor; then performs global average pooling processing on the second activation tensor through a global average pooling layer and outputs the processing result as the corresponding image feature vector; The tensor shapes of the first dimension-raising tensor, the first normalized tensor, and the second activation tensor are all C9×H9×W9, where H9, W9, and C9 are the height, width, and feature dimension of the first dimension-raising tensor, respectively, and H9=H std , W9=W std , C9=L1; the shape of the image feature vector is 1×L1, and its vector eigenvalue is constrained to be within the value range [0,1].
3. The method for predicting breast cancer classification by fusing multimodal features according to claim 1, characterized in that: The second encoder is used to perform genomic feature encoding processing according to the genomic feature matrix input by the encoder and output a corresponding gene feature vector; Each row of the genomic feature matrix corresponds to a type of gene, and each column corresponds to a type of gene feature; the matrix shape of the genomic feature matrix is H G ×W G , H G 、W G are the height and width of the matrix, H G 、W G Match the total number of genes and the total number of gene features in the genomic feature matrix respectively; The gene feature vector is a one-dimensional feature vector with a shape of 1×L2; L2 is a preset vector length threshold; the vector length threshold L2 defaults to 32; the vector eigenvalue of the gene feature vector is constrained to be within the value range [-1, 1]; The input end of the second encoder is used to receive the genomic feature matrix input by the encoder, and the output end is used to output the corresponding gene feature vector; The second encoder includes a first linear layer, a first activation layer, a first normalization layer, a second linear layer, a second activation layer, a third linear layer, a noise injection layer, and a feature truncation layer; The input end of the first linear layer is connected to the input end of the second encoder, and the output end is connected to the input end of the first activation layer; the output end of the first activation layer is connected to the input end of the first normalization layer; the output end of the first normalization layer is connected to the input end of the second linear layer; the output end of the second linear layer is connected to the input end of the second activation layer; the output end of the second activation layer is connected to the input end of the third linear layer; the output end of the third linear layer is connected to the input end of the noise injection layer; the output end of the noise injection layer is connected to the input end of the feature truncation layer; and the output end of the feature truncation layer is connected to the output end of the second encoder; The first linear layer is used to flatten the genomic feature matrix into a matrix with a length of H G ×W G The one-dimensional vector is recorded as the first vector; and the first vector is fully connected to obtain the corresponding first fully connected vector and sent to the first activation layer; The first activation layer is used to perform nonlinear activation on the first fully connected vector through a ReLU function to obtain a corresponding first activation vector and send it to the first normalization layer; The first normalization layer is used to perform layer normalization processing on the first activation vector to obtain a corresponding first normalized vector and send it to the second linear layer; The second linear layer is used to perform a full connection calculation on the first normalized vector to obtain a corresponding second fully connected vector and send it to the second activation layer; The second activation layer is used to perform nonlinear activation on the second fully connected vector through a Tanh function to obtain a corresponding second activation vector and send it to the third linear layer; the vector eigenvalue of the second activation vector is constrained to be within the value range [-1, 1]; The third linear layer is used to perform a full connection calculation on the second activation vector to obtain a corresponding third fully connected vector and send it to the noise injection layer; the third fully connected vector is a one-dimensional feature vector with a shape of 1×L2; The noise injection layer is used to construct a standard normal distribution noise vector with a vector length of L2 according to the model parameter ε, recorded as a first noise vector; and add the first noise vector and the third fully connected vector to obtain a corresponding noise injection vector and send it to the feature truncation layer; The mean of the first noise vector is 0 and the variance is ε 2 ; The model parameter ε is a learnable model parameter, and its initial value is set to the preset initialization noise coefficient ε0; The feature truncation layer is used to perform nonlinear activation on the noise injection vector through the Tanh function to obtain the corresponding gene feature vector and output it; the gene feature vector is a one-dimensional feature vector with a shape of 1×L2, and its vector eigenvalue is constrained within the value range [-1,1].
4. The method for predicting breast cancer classification by fusing multimodal features according to claim 1, characterized in that: The first prediction network is used to perform three-category prediction processing on benign, malignant and borderline lesions of breast tumors based on the fusion feature vector input by the network and output corresponding classification prediction vectors; The fused feature vector is a one-dimensional feature vector with a shape of 1×(L1+L2); L1 and L2 are two preset vector length thresholds; the vector length threshold L1 defaults to 512, and the vector length threshold L2 defaults to 32; the first L1 vector eigenvalues of the fused feature vector are constrained to be within the value range [0,1], and the last L2 vector eigenvalues are constrained to be within the value range [-1,1]; The classification prediction vector includes three prediction probabilities, namely the benign prediction probability, the malignant prediction probability and the borderline lesion prediction probability; The network input end of the first prediction network is used to receive the fusion feature vector input by the network, and the network output end is used to output the corresponding classification prediction vector; The first prediction network includes a fourth linear layer, a second normalization layer, a third activation layer, a random inactivation layer, a fifth linear layer, and a Softmax function layer; The input end of the fourth linear layer is connected to the input end of the first prediction network, and the output end is connected to the input end of the second normalization layer; the output end of the second normalization layer is connected to the input end of the third activation layer; the output end of the third activation layer is connected to the input end of the random deactivation layer; the output end of the random deactivation layer is connected to the input end of the fifth linear layer; the output end of the fifth linear layer is connected to the input end of the softmax function layer; The fourth linear layer is used to perform a full connection calculation on the fused feature vector to obtain a corresponding fourth fully connected vector and send it to the second normalization layer; The second normalization layer is used to perform batch normalization on the fourth fully connected vector to obtain a corresponding second normalized vector and send it to the third activation layer; The third activation layer is used to perform nonlinear activation on the second normalized vector through a ReLU function to obtain a corresponding third activation vector and send it to the random inactivation layer; The random deactivation layer is used to perform random deactivation processing on the vector features in the third activation vector in equal proportion according to a preset random deactivation ratio to obtain a corresponding local deactivation vector and send it to the fifth linear layer; the local deactivation vector has the same vector shape as the third activation vector, but some vector features in the local deactivation vector are set to preset deactivation feature values; the ratio of the total number of deactivation feature values of the local deactivation vector to the vector length meets the random deactivation ratio; The fifth linear layer is used to perform a full connection calculation on the local deactivation vector to obtain a corresponding fifth fully connected vector and send it to the Softmax function layer; the vector length of the fifth fully connected vector is 3; The Softmax function layer is used to perform three-category probability calculation based on the fifth fully connected vector through the Softmax function to obtain the corresponding benign prediction probability, the malignant prediction probability and the borderline lesion prediction probability to form the corresponding classification prediction vector and output it.
5. The method for predicting breast cancer classification by fusing multimodal features according to claim 1, characterized in that: The breast cancer classification model is used to classify and predict benign, malignant and borderline lesions of breast tumors based on the breast ultrasound image, the lesion area mask and the genomic feature matrix input by the model and output a corresponding classification prediction vector; The model input of the breast cancer classification model includes a breast ultrasound image, a lesion area mask, and a genomic feature matrix, wherein the lesion area mask is an optional input; the model output of the breast cancer classification model is the classification prediction vector, and the classification prediction vector includes three prediction probabilities, namely, the benign prediction probability, the malignant prediction probability, and the borderline lesion prediction probability; The first model input end of the breast cancer classification model is used to receive the breast ultrasound image and the lesion area mask, the second model input end is used to receive the genomic feature matrix, and the model output end is used to output the corresponding classification prediction vector; The breast cancer classification model includes the first encoder, the second encoder, a multimodal feature fusion module and the first prediction network; The encoder input of the first encoder is connected to the input of the first model, and the encoder output is connected to the first input of the multimodal feature fusion module; the encoder input of the second encoder is connected to the input of the second model, and the encoder output is connected to the second input of the multimodal feature fusion module; the output of the multimodal feature fusion module is connected to the network input of the first prediction network; and the network output of the first prediction network is connected to the output of the model; The first encoder is used to perform radiomics feature encoding processing based on the breast ultrasound image and the lesion area mask image input by the encoder to obtain a corresponding image feature vector and send it to the multimodal feature fusion module; the image feature vector is a one-dimensional feature vector with a shape of 1×L1, and the vector length threshold L1 defaults to 512; The second encoder is used to perform genomic feature encoding processing according to the genomic feature matrix input by the encoder to obtain a corresponding gene feature vector and send it to the multimodal feature fusion module; the gene feature vector is a one-dimensional feature vector with a shape of 1×L2, and the vector length threshold L2 defaults to 32; The multimodal feature fusion module is used to sequentially concatenate the image feature vector and the gene feature vector to obtain a one-dimensional feature vector with a shape of 1×(L1+L2) as a corresponding fused feature vector; and send the fused feature vector to the first prediction network; The first prediction network is used to perform three-category prediction processing on benign, malignant and borderline lesions of breast tumors according to the fusion feature vector and output the corresponding classification prediction vector.
6. The method for predicting breast cancer classification by fusing multimodal features according to claim 1, characterized in that: The first data set includes multiple first data records; each first data record corresponds to an acquisition object; the first data record includes a first training image, a first training mask image, a first training feature matrix and a first label vector; the first training image is the breast ultrasound image of the current acquisition object; the first training mask image is the lesion area mask image of the current acquisition object; the first training feature matrix is the genomic feature matrix of the current acquisition object; the first label vector includes three label probabilities, namely, benign label probability, malignant label probability and borderline lesion label probability, and only one of the three label probabilities is 1, and the other two label probabilities are 0; all the first training images of the first data set have the same image type, specifically breast B-ultrasound image, breast Doppler ultrasound image or breast ultrasound elastography.
7. The method for predicting breast cancer classification by fusing multimodal features according to claim 6, characterized in that: The model building training dataset is recorded as the corresponding first dataset, specifically including: Recruiting multiple patients with benign breast tumors, malignant breast tumors, and borderline breast lesions to form a collection subject group; the collection subject group includes multiple collection subjects; Select one from the three image types of breast B-ultrasound image, breast Doppler ultrasound image and breast ultrasound elastography as the corresponding current ultrasound image type; and use each of the acquisition objects as the corresponding current object; and perform a breast ultrasound examination on the current object based on the current ultrasound image type and use the breast ultrasound image obtained in this examination as a corresponding first training image; and perform a whole genome sequencing on the current object to obtain the corresponding current object gene sequence, and fill all matrix units of the genomic feature matrix with feature data based on the current object gene sequence and all gene types and all gene features specified by the genomic feature matrix, and use the filled genomic feature matrix as the corresponding first training feature matrix; and use the preset breast lesion delineation interface to sequence the current object on the first training image. The breast lesion area of the image is delineated and a lesion area mask is generated as the corresponding first training mask based on the delineation result; and the current object is identified; if the current object is a benign breast tumor patient, a first label vector is set in which the benign label probability is 1 and the malignant and borderline lesion label probabilities are 0; if the current object is a malignant breast tumor patient, a first label vector is set in which the malignant label probability is 1 and the benign and borderline lesion label probabilities are 0; if the current object is a borderline breast lesion patient, a first label vector is set in which the borderline lesion label probability is 1 and the benign and malignant label probabilities are 0; and a corresponding first data record is formed by the first training image, the first training mask, the first training feature matrix and the first label vector corresponding to the current object; All the obtained first data records form the corresponding first data set.
8. The method for predicting breast cancer classification by fusing multimodal features according to claim 6, characterized in that: The first step of performing a first round of masked image model training on the breast cancer classification model based on the first data set and then performing a second round of unmasked image model training on the breast cancer classification model based on the first data set specifically includes: Step 81, set the training round to the first round; Step 82: randomly split the first data set into two sub-data sets based on a preset first split ratio and record them as a first training set and a first evaluation set; Wherein, both the first training set and the first evaluation set include a plurality of the first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio; Step 83: extract the first first data record from the first training set as the corresponding current training record; Step 84, identifying the training round; if the training round is the first round, inputting the first training image, the first training mask, and the first training feature matrix of the current training record as the current breast ultrasound image, the lesion area mask, and the genomic feature matrix into the breast cancer classification model for classification prediction processing, and using the classification prediction vector output from this processing as the corresponding first prediction vector; if the training round is the second round, inputting the first training image and the first training feature matrix of the current training record as the current breast ultrasound image and the genomic feature matrix, and setting the current lesion area mask to empty, and inputting the current breast ultrasound image, the lesion area mask, and the genomic feature matrix into the breast cancer classification model for classification prediction processing, and using the classification prediction vector output from this processing as the corresponding first prediction vector; Step 85: Form a corresponding first prediction-label pair using the first prediction vector and the first label vector corresponding to the current training record; and subject the current first prediction-label pair to a preset first model loss function to calculate and obtain a corresponding first loss value; Wherein, the first model loss function is implemented based on L1 loss function, L2 loss function or cross entropy loss function; Step 86: Identify whether the first loss value satisfies a preset first loss value range. If the first loss value satisfies the first loss value range, identify whether the current training record is the last first data record in the first training set. If so, proceed to step 87. If not, use the next first data record in the first training set as the new current training record and return to step 84 to continue training. If the first loss value does not satisfy the first loss value range, perform a round of modulation on the model parameters of the breast cancer classification model in a direction that minimizes the first model loss function based on a preset first model optimizer. After this round of modulation, return to step 84 to continue training. Wherein, the first model optimizer includes at least an Adam optimizer and an SGD optimizer; Step 87, perform a round of traversal on all the first data records of the first evaluation set; and in this round of traversal, use the first data record currently traversed as the corresponding current evaluation record; and identify the training round; if the training round is the first round, then use the first training image, the first training mask map, and the first training feature matrix of the current evaluation record as the current breast ultrasound image, the lesion area mask map, and the genomic feature matrix to input the breast cancer classification model for classification prediction processing and use the classification prediction vector output by this processing as the corresponding second prediction vector; if the training round is the second round, then use the first training image, the first training mask map, and the first training feature matrix of the current evaluation record as the current breast ultrasound image, the lesion area mask map, and the genomic feature matrix to input the breast cancer classification model for classification prediction processing and use the classification prediction vector output by this processing as the corresponding second prediction vector; A training image and the first training feature matrix are used as the current breast ultrasound image and the genomic feature matrix, and the current lesion area mask is set to empty. The current breast ultrasound image, the lesion area mask, and the genomic feature matrix are input into the breast cancer classification model for classification prediction processing, and the classification prediction vector output from this processing is used as the corresponding second prediction vector; the second prediction vector corresponding to the current evaluation record and the first label vector form a corresponding second prediction-label pair; and at the end of this round of traversal, all the obtained second prediction-label pairs are brought into the preset first model evaluation function to calculate and obtain the corresponding first evaluation value; Wherein, the first model evaluation function is implemented based on MAE function, MSE function or RMSE function; Step 88, identifying whether the first evaluation value meets the preset first evaluation value range; if so, proceeding to step 89; if not, returning to step 83 to continue training; Step 89, identify the training round; if the training round is the first round, reset the training round to the second round and return to step 82 for the next round of training; if the training round is the second round, stop training and confirm that the two rounds of model training are completed.
9. The method for predicting breast cancer classification by fusing multimodal features according to claim 5, characterized in that: The breast cancer classification model predicts based on the subject data to obtain a corresponding subject prediction vector and feeds it back to the current user, specifically including: Identify whether the lesion area mask exists in the subject data; if so, extract the corresponding breast ultrasound image, the lesion area mask and the genomic feature matrix from the subject data; if not, extract the corresponding breast ultrasound image and the genomic feature matrix from the subject data, and set a corresponding lesion area mask to be empty; and input the current breast ultrasound image, the lesion area mask and the genomic feature matrix into the breast cancer classification model for classification prediction processing, and feed back the classification prediction vector output by this processing as the corresponding subject prediction vector to the current user.
10. A device for executing the processing method for predicting breast cancer classification by fusing multimodal features according to any one of claims 1 to 9, characterized in that: The device includes: a model construction module, a model training module, and a model application module; The model construction module is used to construct a first encoder for imaging omics feature encoding using a multi-level residual network as a feature extraction backbone, and introduce a multi-level residual feature fusion mechanism and a spatial attention fusion mechanism into the encoder; and to construct a second encoder for genomic feature encoding using a multi-level linear network as a feature extraction backbone, and introduce a noise injection mechanism into the encoder; and to construct a first prediction network for performing three-category prediction of benign, malignant and borderline lesions of breast tumors based on a classifier model composed of a linear network and a Softmax function; and to construct a deep learning model for performing multimodal feature fusion of breast ultrasound imaging omics features and genomic features and classifying and predicting benign, malignant and borderline lesions of breast tumors based on the first and second encoders and the first prediction network, which is recorded as a breast cancer classification model; The model training module is used to construct a model training data set recorded as a corresponding first data set; and firstly perform a first round of masked image model training on the breast cancer classification model based on the first data set and then perform a second round of unmasked image model training on the breast cancer classification model based on the first data set; The model application module is used to receive subject data input by the user after two rounds of model training are completed; the breast cancer classification model performs prediction based on the subject data to obtain a corresponding subject prediction vector and feeds it back to the current user; the subject data includes breast ultrasound images, lesion area masks, and genomic feature matrices, and the lesion area masks are optional data; the subject prediction vector includes three prediction probabilities, namely, benign prediction probability, malignant prediction probability, and borderline lesion prediction probability.
11. An electronic device, characterized in that: include: memory, processors, and transceivers; The processor is configured to be coupled to the memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 9; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image super-resolution reconstruction method based on residual module and attention mechanism
CN111179171A
Image processing method and system for identifying pneumonia in CT image
CN111724356A
Image target detection method and system based on convolutional neural network, and readable storage medium
CN113052006A
Breast lesion segmentation device, model training method and electronic equipment
CN114494230A
Two-dimensional ultrasonic medical image restoration method based on mask guidance
CN115147303A
Cited By
Breast cancer detection method and system based on morphological image and space transcriptome cross-graph collaborative learning
CN121582228A