Multi-mode breast volume ultrasonic focus grading method, medium and terminal

Through multimodal breast volume ultrasound technology, combined with visual Transformer and 2D CNN, the various characteristics of breast lesions are extracted and fused, and the accurate BI-RADS grading of breast lesions is achieved, solving the problem of low grading accuracy in the existing technology and improving the accuracy and efficiency of detection.

CN119963887APending Publication Date: 2025-05-09HUNAN PROVINCIAL TUMOR HOSPITAL

Patent Information

Application Number
CN202411961729.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing breast cancer lesion classification methods are highly subjective and poorly consistent, and the feature description is not comprehensive enough, resulting in low accuracy of lesion classification.

Method used

The multimodal breast volume ultrasound lesion grading method is used to obtain three-dimensional data through automatic volume ultrasound equipment, combine the visual Transformer model and 2D convolutional neural network to extract point cloud features, dynamic grayscale features and static elastic features, and feature fusion is performed through weighting algorithms or multimodal neural network structures, and finally BI-RADS grading of breast lesion lesions is performed through the neural network layer.

Benefits of technology

It improves the accuracy and consistency of breast lesion grading, reduces the impact of doctors' experience, is suitable for large-scale screening, and improves work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963887A_ABST
    Figure CN119963887A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of intelligent diagnosis, and relates to a multi-mode breast volume ultrasound focus grading method, medium and terminal, and the method comprises the steps: carrying out the full-coverage scanning of a target region through automatic volume ultrasound equipment, obtaining automatic volume ultrasound data, and carrying out the preprocessing; carrying out three-dimensional segmentation and point cloud feature extraction on the preprocessed data; processing the dynamic gray ultrasonic time sequence through a visual Transform model so as to extract dynamic gray features; the ultrasonic elastography image is processed through a 2D convolutional neural network, and static elastic features are extracted; fusing the point cloud features, the dynamic gray features and the static elastic features through a weighting algorithm or a multi-modal neural network structure; and processing the fused features through a neural network layer, inputting the processed features into a classification layer, and carrying out breast focus BI-RADS classification. The method is simple in process and convenient to operate, and the accuracy of breast focus grading is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent diagnosis technology, and in particular relates to a multimodal breast volume ultrasonic lesion grading method, medium and terminal. Background Art

[0002] Breast cancer is one of the most common malignant tumors in women. Early diagnosis is of great significance for improving cure rate and survival rate. At present, the diagnostic methods of breast cancer mainly include clinical examination, X-ray photography (molybdenum target), magnetic resonance imaging (MRI), ultrasound examination, etc. Among them, ultrasound examination has been widely used in breast cancer screening due to its non-invasive, low cost and strong repeatability. The existing automatic volume ultrasound technology can obtain three-dimensional images of breast tissue. By automatically controlling the ultrasound probe to perform a comprehensive scan of the breast, the transverse, sagittal and coronal images of the breast are obtained. This method mainly relies on the experience of doctors to identify and grade lesions. It is highly subjective, has poor consistency, and has low accuracy in lesion grading. Dynamic ultrasound grayscale imaging can record the echo characteristics of breast tissue at different time points, which is helpful to observe the dynamic changes of lesions. However, the existing feature extraction methods usually only focus on the data at a single time point and fail to make full use of time series information, resulting in incomplete feature description and low accuracy in lesion grading. Ultrasonic elastic imaging can reflect the hardness information of tissues and help distinguish benign from malignant lesions. The existing elastic feature extraction methods are often too simple and cannot fully describe the elastic characteristics of lesions.

[0003] The patent application with publication number CN118177994A provides a body surface marking system for primary lesions of breast cancer, including: body surface marking, the body surface marking is determined by doctors using marking ink to determine the location of the primary lesions of breast cancer, the three-dimensional scanning system (abus) scans the location of the body surface marking; the three-dimensional scanning system (abus) includes a data receiving and inputting system, the three-dimensional scanning system (abus) receives the scanned data of the primary lesions of breast cancer, the three-dimensional scanning system (abus) transmits the data obtained by scanning the primary lesions of breast cancer into a data processing system, the data processing system processes the data of the primary lesions of breast cancer, the data processing system matches the processed data of the primary lesions of breast cancer with a three-dimensional coordinate database, and the primary lesions of breast cancer matched by the three-dimensional coordinate database enter the digital model integration system to become a three-dimensional digital model. In this patent, breast lesions are all marked by multiple doctors, which has the same disadvantages as the prior art.

[0004] Therefore, how to provide a breast lesion grading method with high breast lesion grading accuracy is a problem that needs to be solved urgently by people in this technical field. Summary of the invention

[0005] In view of the deficiencies in the prior art, the object of the present invention is to provide a multimodal breast volume ultrasound lesion grading method to solve the problem of low accuracy of existing breast lesion grading; in addition, the present invention also provides a multimodal breast volume ultrasound lesion grading medium and terminal.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] In the first aspect, S10, use an automatic volume ultrasound device to perform a full coverage scan of the target area, obtain automatic volume ultrasound data and perform preprocessing, and perform three-dimensional segmentation and point cloud feature extraction on the preprocessed data;

[0008] S20, processing the dynamic grayscale ultrasound time series through the visual Transformer model to extract dynamic grayscale features;

[0009] S30, processing the ultrasound elastic imaging image through a 2D convolutional neural network to extract static elastic features;

[0010] S40, fusing the point cloud features, dynamic grayscale features, and static elastic features through a weighted algorithm or a multimodal neural network structure;

[0011] S50, processing the fused features through a neural network layer, and inputting the processed features into a classification layer to perform BI-RADS grading of breast lesions.

[0012] Furthermore, the specific steps of obtaining automatic volume ultrasound data and preprocessing in step S10 are as follows:

[0013] S101, obtaining cube data V (x, y, z) resolution of three-dimensional ultrasound data, where the data resolution is millimeter level in x, y, and z directions;

[0014] S102, using non-local mean filtering or three-dimensional Gaussian filtering to remove background noise, the filtering formula is:

[0015]

[0016] Among them, w(i,j,k) is the weight, which is related to the similarity between voxels;

[0017] S103, adjusting the image contrast to enhance the lesion boundary features;

[0018] S104, performing Z-score normalization on the image sequence data, assuming that the original data is V'(x, y, z), the normalization operation is expressed as:

[0019]

[0020] Among them, μ and δ are the pre-calculated mean and standard deviation. The data is normalized by Z-score standardization operation to satisfy the normal distribution, with a mean of 0 and a standard deviation of 1.

[0021] Furthermore, in step S10, a 3D U-Net network is used for three-dimensional segmentation. The 3D U-Net network includes an analysis path and a synthesis path. In the analysis path, each layer includes two 3×3×3 convolution operations, each convolution operation is followed by a ReLU activation function, and also includes a maximum pooling layer; in the synthesis path, each layer restores the resolution of the feature map through transposed convolution, and the upsampled feature map is spliced ​​with the feature map of the corresponding layer in the analysis path according to the channel dimension to integrate multi-level information. The spliced ​​feature map is again subjected to two 3×3×3 convolution operations to extract features, and then a probability distribution of each voxel is generated through a Softmax layer. According to the probability distribution output by the Softmax layer, a threshold is applied for binary segmentation, and the voxel position with a value of 1 is extracted from the segmentation result, which is converted into a three-dimensional coordinate point to generate point cloud data of the lesion area.

[0022] Furthermore, in step S10, a point cloud network is used to extract point cloud features. The point cloud data consists of a set of points in three-dimensional space. Each point is defined by its own coordinates. The point cloud is aligned through an input transformation network, and the aligned point cloud is sent to a shared multi-layer perceptron to extract local features from each point. A feature transformation network is used to improve the alignment of features, and local features are integrated through a maximum pooling layer to form a global feature vector for subsequent classification and segmentation tasks.

[0023] Furthermore, for classification tasks, the point cloud network passes the extracted global feature vector through the fully connected layer of the classification branch, and finally outputs the concept of each category. The classification branch includes several fully connected layers. The fully connected layer processes the global feature vector to generate the final classification result and identify the category to which the point cloud belongs; for segmentation tasks, the point cloud network concatenates the global feature vector with the local features of each point, and then passes the convolutional layer of the segmentation branch to output the segmentation result of each point. The segmentation branch includes several convolutional layers. The convolutional layer processes the concatenated feature vector to generate the segmentation label of each point and identify the category to which each point belongs.

[0024] Furthermore, the specific steps of step S20 are as follows:

[0025] S201, the visual Transformer model receives a dynamic grayscale ultrasound time series as input;

[0026] S202, decomposing the large image of the image block into small local regions, each block is flattened into a one-dimensional vector, and projected into a vector space of fixed dimension through a linear transformation;

[0027] S203, adding position codes to distinguish image blocks at different positions;

[0028] S204, analyzing the relationship between image blocks using a self-attention mechanism;

[0029] S205. Using the multi-head attention mechanism, the self-attention mechanism is independently operated in multiple subspaces, and information is processed in parallel in multiple subspaces to enhance the expressiveness of the model;

[0030] S206, extracting the local information of a single time point and the dynamic changes between different time points from the dynamic grayscale ultrasound time series through the self-attention mechanism and the multi-head attention mechanism;

[0031] S207, perform nonlinear transformation on the output of the self-attention layer through the feedforward network of the visual Transformer model to extract higher-level features;

[0032] S208, Visual Transformer model uses layer normalization and residual connections in self-attention layers and feed-forward networks;

[0033] S209, the visual Transformer model outputs dynamic grayscale features.

[0034] Furthermore, the specific steps of step S30 are as follows:

[0035] S301, preprocessing the ultrasound elastic imaging image, including normalization, cropping / filling and data enhancement;

[0036] S302, construct a 2D CNN model, including several convolutional layers, activation functions, pooling layers, fully connected layers and output layers;

[0037] S303, selecting a loss function and an optimizer;

[0038] S304, dividing the data set into a training set, a validation set and a test set, which are used for model training, hyperparameter adjustment and final evaluation respectively. The model training process includes forward propagation, loss calculation, back propagation and iterative optimization;

[0039] S305. Use the trained 2D CNN model to extract features from the new ultrasound elastography image.

[0040] Furthermore, the BI-RADS grading of breast lesions in step S50 includes the following levels:

[0041] BI-RADS 0: Incomplete assessment, further imaging examinations are required to complete the assessment; BI-RADS 1: Negative, no abnormal lesions were found; BI-RADS 2: Benign lesions; BI-RADS 3: Probable benign lesions, with a malignancy risk of less than 2%; BI-RADS 4: Suspected malignant lesions, with a malignancy risk between 2% and 95%; BI-RADS 5: Highly likely malignant lesions, with a malignancy risk greater than 95%; BI-RADS 6: Known malignant lesions;

[0042] The specific steps of the BI-RADS grading of breast lesions are as follows:

[0043] S501, fusing point cloud features, dynamic grayscale features, and static elastic features;

[0044] S502, the fused features are processed by a neural network layer to further extract features;

[0045] S503, the processed features are sent to the classification layer, the number of output nodes of the classification layer corresponds to the number of BI-RADS categories, and in the classification layer, each node represents the score of a category;

[0046] S504, the output of the classification layer is converted into a probability distribution through a Softmax function, so that the model outputs the probability of each BI-RADS category;

[0047] S505. According to the probability distribution output by the Softmax function, select the category with the highest probability as the prediction result of the model;

[0048] S506: Adjust the threshold of the prediction result and interpret the result.

[0049] In a second aspect, the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.

[0050] In a third aspect, the present invention further provides an electronic terminal, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes the method as described above.

[0051] Compared with the prior art, the multimodal breast volume ultrasound lesion grading method, medium and terminal provided by the present invention have at least the following beneficial effects:

[0052] At present, the subjectivity and consistency of breast cancer lesion detection are strong, the feature description is not comprehensive enough, and the accuracy of lesion classification is low. The present invention has a simple process and convenient operation. It obtains multiple features of breast tissue through multimodal ultrasound data, combines 3D-Unet model, visual Transformer model and 2D convolutional neural network (2D CNN), realizes three-dimensional segmentation of lesions, dynamic grayscale feature extraction and static elastic feature extraction, and then realizes accurate classification of breast lesions through feature fusion and intelligent grading model. The present invention comprehensively describes lesion characteristics through multimodal feature fusion, improves lesion detection accuracy; the automated grading method reduces the influence of doctor's experience, improves the consistency of diagnosis results, and reduces the subjective influence of doctors; improves efficiency, has a high degree of automation, is suitable for large-scale screening, and improves work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the scheme of the present invention, a brief introduction is given below to the figures required for use in the description of the embodiments. Obviously, the figures described below are some embodiments of the present invention. For ordinary technicians in this field, other figures can be obtained based on these figures without paying any creative work.

[0054] Figure 1 A flowchart of a multimodal breast volume ultrasound lesion grading method provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. The preferred embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present invention more thorough and comprehensive.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0057] The present invention provides a multimodal breast volume ultrasound lesion grading method, which is applied to the early diagnosis of breast cancer. The multimodal breast volume ultrasound lesion grading method comprises the following steps:

[0058] S10. Use the automatic volume ultrasound device to perform a full coverage scan of the target area, obtain the automatic volume ultrasound data and perform preprocessing, and perform three-dimensional segmentation and point cloud feature extraction on the preprocessed data; S20. Process the dynamic grayscale ultrasound time series through the visual Transformer model to extract dynamic grayscale features; S30. Process the ultrasound elastic imaging image through a 2D convolutional neural network to extract static elastic features; S40. Fusion the point cloud features, dynamic grayscale features and static elastic features through a weighted algorithm or a multimodal neural network structure; S50. Process the fused features through the neural network layer, and after processing, input them into the classification layer for BI-RADS grading of breast lesions.

[0059] The method of the present invention has a simple process and is convenient to operate, and effectively improves the accuracy of breast lesion grading.

[0060] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings.

[0061] The present invention provides a multimodal breast volume ultrasound lesion grading method, which is applied to the early diagnosis of breast cancer. Figure 1 As shown, in this embodiment, the multimodal breast volume ultrasound lesion grading method includes the following steps:

[0062] S10. Use automatic volumetric ultrasound equipment to perform full coverage scanning of the target area, obtain automatic volumetric ultrasound data and perform preprocessing, and perform three-dimensional segmentation and point cloud feature extraction on the preprocessed data, which is mainly used to reflect the three-dimensional structure, volume and morphological characteristics of breast tissue.

[0063] Specifically, automatic volumetric ultrasound data acquisition is a technology that uses conventional ultrasound equipment to obtain three-dimensional information of breast tissue. This technology automatically controls the ultrasound probe to perform a comprehensive scan of the breast, thereby obtaining images of the breast in the cross section, sagittal plane, and coronal plane. The automatic scanning of the probe reduces the error of human operation and improves the consistency and repeatability of the scan. Compared with traditional two-dimensional ultrasound imaging, automatic volumetric ultrasound can provide more detailed information on the internal structure of the breast, including the relationship between breast tissue, lesion area, and surrounding tissue. During the automatic volumetric ultrasound imaging process, the ultrasound probe emits high-frequency sound waves, which are absorbed or reflected to varying degrees by tissues of different densities when passing through the breast tissue. The reflected sound waves are received by the probe and converted into electrical signals, which are then processed by the ultrasound equipment to generate three-dimensional images of the breast. These images have high resolution and can clearly show the microscopic structure of the breast, providing doctors with rich diagnostic information. During the data acquisition process, the image may be preprocessed to improve the clarity and resolution of the image to ensure that subsequent three-dimensional segmentation and feature extraction can be performed based on high-quality data. These preprocessing steps may include noise removal, contrast enhancement, etc. to improve the diagnostic value of the image. Through automatic volumetric ultrasound data acquisition, doctors can more accurately assess breast lesions, providing important basis for intelligent grading and treatment of breast cancer.

[0064] Furthermore, in this embodiment, the specific steps of step S10 are as follows:

[0065] S101. Data acquisition: Use an automated volumetric ultrasound device (such as ABUS, Automated Breast Ultrasound) to perform a full coverage scan of the target area to obtain three-dimensional ultrasound data cube data V (x, y, z) resolution. The data resolution is usually at the millimeter level in the x, y, and z directions, which can clearly reflect the breast tissue structure.

[0066] S102, pre-processing: noise suppression: use non-local mean filtering (NLM) or three-dimensional Gaussian filtering to remove background noise. The filtering formula is:

[0067]

[0068] Among them, w(i,j,k) is the weight, which is related to the similarity between voxels.

[0069] S103, histogram equalization: adjust image contrast and enhance lesion boundary features.

[0070] S104, Normalization: Perform Z-score normalization on the image sequence data. The purpose of this step is to ensure that the data is comparable between different dimensions, thereby laying the foundation for subsequent deep learning model training. The data is normalized using the Z-score normalization operation. After normalization, the data eliminates the dimension, and the convergence speed is faster during training. Assuming that the original data is V'(x, y, z), the normalization operation can be expressed as:

[0071]

[0072] Among them, μ and δ are the pre-calculated mean and standard deviation. The data are normalized according to the Z-score standardization operation to satisfy the normal distribution, with a mean of 0 and a standard deviation of 1.

[0073] Furthermore, in this embodiment, a 3D U-Net network is used in step S10 for three-dimensional segmentation. The 3D U-Net network is mainly divided into two parts: an analysis path (encoder part): extracting high-level semantic features and gradually reducing the resolution of the feature map. A synthesis path (decoder part): restoring the resolution of the feature map and combining the features of the encoder to generate a segmentation result.

[0074] The analysis path includes: convolution layer. In the encoder part, each layer contains two 3×3×3 convolution operations, and the convolution is followed by a ReLU (Rectified Linear Unit) activation function to extract local features and enhance nonlinear expression capabilities.

[0075] Convolution formula:

[0076] y = ReLU(W*x+b);

[0077] Among them, x is the input feature map, W is the convolution kernel, * represents the convolution operation, b is the bias term, y is the output feature map, and the ReLU activation function is defined as:

[0078] ReLU(z)=max(0,z).

[0079] Max pooling layer,Max pooling is used for downsampling, reducing the resolution of feature maps while retaining key features. Assume that the shape of the input feature map is 3D data H×W×C, H: height, W: width, C: number of channels.

[0080] The pooling kernel size is k, the step size is s, and the shape of the output feature map is:

[0081]

[0082] Take the maximum value in each pooling kernel area:

[0083] Y i,j,c =max{Xp,q,c (p,q)∈region(i,j)};

[0084] Among them, region(i,j) is a local region.

[0085] The synthesis path includes: transposed convolution (up-sampling layer). In the decoder part, each layer restores the resolution of the feature map through transposed convolution (also called up-convolution). The formula is:

[0086] y up =ConvTranspose(z,W')+b';

[0087] Among them, z is the input feature map, W' is the transposed convolution kernel, b' is the bias term, and y up It is the upsampled feature map.

[0088] The splicing operation splices the upsampled feature map with the feature map of the corresponding layer of the encoder according to the channel dimension to integrate multi-level information.

[0089] Convolution layer, the concatenated feature map is again extracted through two 3×3×3 convolution operations, and the formula is the same as the analysis path.

[0090] Furthermore, in this embodiment, the probability distribution of each voxel is generated by the Softmax layer:

[0091]

[0092] Among them, z c is the score that the voxel belongs to category c, and P(cx) is the probability that the voxel belongs to category c.

[0093] Furthermore, in this embodiment, threshold segmentation: according to the Softmax output probability distribution, a threshold T is applied to perform binary segmentation:

[0094]

[0095] Among them, S(x) is the binary segmentation result, 1 represents the lesion area, and 0 represents the background.

[0096] Lesion extraction: Extract the voxel position with a value of 1 from the segmentation result, convert it into a three-dimensional coordinate point, and generate point cloud data of the lesion area for further analysis.

[0097] Furthermore, in this embodiment, a point cloud network is used to extract point cloud features in step S10. The point cloud network is a deep learning architecture specially designed to process unordered point cloud data. These data are composed of a set of points in three-dimensional space, and each point is defined by its coordinates (x, y, z). This network performs well in 3D object classification and segmentation tasks, and can learn global and local features from point sets.

[0098] Specifically, the input of the point cloud network is an unordered set of points, which is the lesion area of ​​3D-Unet in this paper. These points are distributed in three-dimensional space. In order to process these point cloud data, the network first aligns the point cloud through an input transformation network (InputTransform Net). This network consists of multiple convolutional layers and fully connected layers. Its purpose is to output a 3×3 transformation matrix T, which is used to perform affine transformation on the point coordinates in the point cloud. This alignment operation helps the network better recognize and understand the geometric structure of the point cloud because it reduces the irregularity and asymmetry of the data. The output T of the input transformation network can be expressed as:

[0099] T=f input (X);

[0100] Where X is the input point cloud data, with a shape of n×3, where n is the number of points, and each point is defined by 3 coordinates (x, y, z), and f input is a function of the input transformation network, which is learned through a series of convolutional layers and fully connected layers. The transformation matrix T is a 3×3 matrix used to perform an affine transformation on the point coordinates in the point cloud. The affine transformation transforms the input point cloud X into the aligned point cloud X' through matrix multiplication:

[0101] X' = ​​X·T;

[0102] Assume that the input point cloud X is an n×3 matrix, the specific form is as follows:

[0103]

[0104] The transformation matrix T is a 3×3 matrix used to perform affine transformation on the point coordinates in the point cloud. The specific form is as follows:

[0105]

[0106] The resulting matrix X' is still an n×3 matrix, where each row represents the three-dimensional coordinates (x', y', z') of an aligned point:

[0107]

[0108] The aligned point cloud X' is then fed into a shared multi-layer perceptron (Shared MLP), which is responsible for extracting local features from each point. The shared multi-layer perceptron consists of a series of convolutional layers that perform the same convolution operation on each point to capture the local geometric information of the point cloud. These local features provide the network with detailed information about the local area of ​​the point cloud, laying the foundation for further analysis.

[0109] Share the output h of the multilayer perceptron i It can be expressed as:

[0110] h i =f mlp (x i ');

[0111] Among them, x i ' is the coordinate of the i-th point after alignment.

[0112] In order to improve the alignment of features, the network also uses a feature transformation network (Feature Transform Net). This network is also composed of convolutional layers and fully connected layers. Its function is to output a 64×64 transformation matrix T', which is used to adjust the feature vectors to make them more regular and symmetrical in the feature space. This step further optimizes the expression of features, allowing the network to more accurately identify and distinguish different 3D structures.

[0113] The output T' of the feature transformation network can be expressed as:

[0114] T'=f feature (H);

[0115] The aligned features H' are affine transformed by the transformation matrix T':

[0116] H'=H·T';

[0117] After the local features are extracted and aligned, the network integrates these local features through the max pooling layer to form a global feature vector. The max pooling layer combines the local features into a single vector representing the global characteristics of the entire point cloud by selecting the maximum value of all point features. This global feature vector is the essence of the point cloud network's understanding of the entire 3D structure, and it is used for subsequent classification and segmentation tasks, such as identifying lesion types or defining lesion boundaries. The global feature vector g can be expressed as:

[0118]

[0119] For the classification task, the point cloud network extracts the global feature vector g through the fully connected layer of the classification branch and finally outputs the probability of each category. The classification branch consists of several fully connected layers, which process the global feature vector and generate the final classification result. In this way, the point cloud network can classify the entire point cloud and identify the category to which the point cloud belongs. The output p of the classification branch can be expressed as:

[0120] p=f classify (g);

[0121] Among them, f calssify It is a fully connected layer network of the classification branch.

[0122] For the segmentation task, the point cloud network combines the global feature vector g with the local feature h of each point i The concatenation is then passed through the convolutional layer of the segmentation branch to output the segmentation result of each point. The segmentation branch consists of several convolutional layers, which process the concatenated feature vectors to generate the segmentation label of each point. In this way, the point cloud network can perform fine-grained segmentation on each point in the point cloud and identify the category to which each point belongs. The concatenated feature vector f i It can be expressed as:

[0123] f i =concat(g,h i );

[0124] The output of the segmentation branch l i It can be expressed as:

[0125] l i =f segment (f i );

[0126] Among them, f segment It is the convolutional layer network of the segmentation branch.

[0127] S20. Process the dynamic grayscale ultrasound time series through the visual Transformer model to extract dynamic grayscale features.

[0128] Specifically, the Vision Transformer (ViT) has demonstrated its innovative application in computer vision tasks in processing dynamic grayscale ultrasound time series to extract dynamic grayscale features. This model, through the self-attention mechanism, can capture the time-varying features in ultrasound image sequences, which is crucial for understanding the dynamic behavior of lesion areas.

[0129] When applying the visual transformer to dynamic grayscale ultrasound time series, each frame in the time series is first segmented into multiple fixed-size image patches, which are converted into vectors and mapped into a high-dimensional feature space. Unlike natural image processing, these image patches contain not only spatial information but also changes in the time dimension, providing the model with rich time series data.

[0130] Position encoding also plays an important role in this process. It not only provides the spatial position information of the image block, but also encodes the timing information of each frame in the time series, allowing the model to consider the correlation between space and time at the same time.

[0131] The encoder of the visual Transformer consists of multiple identical layers, each of which includes a multi-head self-attention mechanism and a feedforward neural network. The multi-head self-attention mechanism enables the model to learn complex interactions between image patches, including spatial and temporal dependencies, thereby capturing dynamic grayscale features. The feedforward network further processes these features to extract higher-level representations that reflect the evolution of the lesion area over time.

[0132] The flexibility of the visual transformer lies in its ability to process ultrasound time series of different lengths, which makes it possible to analyze the lesion dynamics of different patients. Its advantage in processing long-distance dependencies makes it perform well in dynamic grayscale feature extraction tasks, especially in capturing subtle changes in lesion areas and predicting lesion development trends.

[0133] Although the Visual Transformer requires a large amount of labeled data during the training process, its potential in dynamic grayscale ultrasound time series analysis is enormous. Once trained, the Visual Transformer can achieve excellent performance in a variety of medical imaging analysis tasks, becoming an important development in the field of medical imaging. Although it may require higher computing resources and data than traditional methods, its advantage in capturing dynamic grayscale features makes it a research direction worth exploring.

[0134] Further, in this embodiment, the specific steps of step S20 are as follows:

[0135] S201, data input: Data input is the first step in the entire model processing, ensuring that the model can receive the correct data format. Dynamic grayscale ultrasound time series provides rich spatiotemporal information, which is essential for the detection and classification of breast lesions. Visual Transformer receives dynamic grayscale ultrasound time series as input. These time series consist of a series of continuous ultrasound image frames, each of which contains the grayscale information of breast tissue at a specific time point. Assume that there are T images, and the size of each image is H×W, where T is the length of the time series, that is, the number of image frames, H is the height of the image, and W is the width of the image.

[0136] S202. Image Coding: Image Block Segmentation: In the visual Transformer, each ultrasound image frame is first encoded into a series of image blocks (patches). These image blocks are local areas of the image. They are segmented and projected into a vector space of fixed dimensions so that the Transformer can process them. Assuming that the size of each patch is P×P, each image will be segmented into patches, where P is the height and width of each patch, and N is the number of patches per image. Image block segmentation decomposes large images into small local areas, allowing Transformer to efficiently process these local information, which not only reduces computational complexity but also allows the model to better focus on local features.

[0137] Linear embedding: Each patch is flattened into a one-dimensional vector and mapped to a fixed-dimensional vector space through a linear transformation. Assuming the dimension of the embedding vector is D, the embedding operation can be expressed as:

[0138] E i (t)=Linear(P i (t));

[0139] Among them, P i (t) is the i-th patch in the t-th image, E i (t) is the embedding vector corresponding to the i-th patch in the t-th image, and Linear is a linear transformation function, usually a fully connected layer. Linear embedding converts each patch into a vector of fixed dimension so that these vectors can be effectively operated and compared in the Transformer. This step ensures that the model can handle input images of different sizes.

[0140] S203, position encoding: Since the Transformer itself does not have the ability to process sequence data, it is necessary to add position encoding to provide the position information of the image block in the original image. The position encoding is added to the vector representation of the image block to retain the spatial information. The position encoding can be a learnable parameter or a fixed sine / cosine function. For example, a fixed position encoding can be expressed as:

[0141]

[0142] Among them, i is the index of the patch, i.e. the image block, k is the dimension index, PE(i,2k) is the even dimension of the position encoding, and PE(i,2k+1) is the odd dimension of the position encoding.

[0143] The number 10000 is a scaling factor that controls the periodicity of the sine and cosine functions. This scaling factor, together with the dimension index k and the model's dimension D, determines the frequency of the position encoding. Specifically, this factor affects the period of the sine and cosine functions in the position encoding, which affects how the model perceives the relative distances between image patches at different positions. Position encoding ensures that the model can distinguish between image patches at different positions, which is very important for capturing the spatial structure of the image. By adding position encoding, the model is able to process local features while maintaining awareness of the global structure.

[0144] S204, Self-attention mechanism: Visual Transformer uses the self-attention mechanism to analyze the relationship between image blocks. The self-attention mechanism allows the model to consider the information of all other image blocks when processing each image block, thereby capturing the dependencies between different regions in the image. The specific calculation formula of the self-attention mechanism is as follows:

[0145] Q=E i '(t)W Q , K=E i '(t)W K , V = E i '(t)W V

[0146]

[0147] Among them, Q is the query vector, K is the key vector, V is the value vector, and W Q , W K and W V is the weight matrix, d kis the dimension of the key vector, and Attention is the output of the attention mechanism. The self-attention mechanism enables the model to dynamically focus on different areas in the image and capture long-distance dependencies. By calculating the attention score, the model can determine which areas of information are more important for the current task, thereby improving the efficiency and accuracy of feature extraction.

[0148] S205, Multi-head Attention Mechanism: In order to enhance the expressive power of the model, the Visual Transformer adopts a multi-head attention mechanism, which means that the self-attention mechanism runs in parallel in multiple different representation subspaces, and each head learns different feature relationships. The calculation formula of multi-head attention is as follows:

[0149]

[0150] MultiHead(Q,K,V)=Concat(head1,head1,...,head h )W O

[0151] Among them, head j The output of the jth head, h is the number of heads, W O is the output weight matrix, is the weight matrix of the jth head. The multi-head attention mechanism enhances the expressiveness of the model by running the self-attention mechanism independently in multiple subspaces. Each head can focus on different feature relationships to capture more information. Finally, the outputs of these heads are spliced ​​together to form a comprehensive feature representation. In this way, the multi-head attention mechanism provides a powerful tool for the visual Transformer to process information in multiple subspaces in parallel, thereby improving the performance and accuracy of the model when processing visual tasks.

[0152] S206, Feature Extraction: Through the self-attention mechanism, the visual Transformer can extract dynamic grayscale features, which reflect the echo characteristics of breast tissue at different time points. These features include changes in grayscale values, texture information, and dynamic changes in possible lesion areas. The feature extraction process not only considers the local features of a single time point, but also captures the dynamic changes between different time points through the temporal self-attention mechanism. Feature extraction is one of the core tasks of the model. Through the self-attention mechanism and the multi-head attention mechanism, the model can extract rich features from the dynamic grayscale ultrasound time series. These features include not only the local information of a single time point, but also the dynamic changes between different time points, providing strong support for subsequent tasks.

[0153] S207, Feedforward Network: After the self-attention layer, the visual Transformer contains a feedforward network, which further processes the output of the self-attention layer to extract higher-level features. The structure of the feedforward network is usually a two-layer fully connected network with an activation function (such as ReLU) in the middle:

[0154] FFN(X)=ReLU(XW1+b1)W2+b2;

[0155] Among them, W1 and W2 are weight matrices, b1 and b2 are bias terms, and FFN(X) is the output of the feedforward network. The feedforward network further nonlinearly transforms the output of the self-attention layer to extract higher-level features. These features are more abstract and can better capture high-level semantic information in the image. The introduction of the feedforward network improves the expressiveness and generalization ability of the model.

[0156] S208, layer normalization and residual connection: In order to improve the stability of training and the performance of the model, the visual Transformer uses layer normalization and residual connection in the self-attention layer and feedforward network. These technologies help alleviate the gradient vanishing problem in deep networks and promote the flow of information. The specific operations are as follows:

[0157] X out =LayerNorm(X in +MultiHead(X in ))

[0158] X out =LayerNorm(X out +FFN(X out ));

[0159] Among them, X in is the input feature, X out is the output feature, and LayerNorm is the layer normalization function. By normalizing the input of each layer, the model training is ensured to be more stable. Layer normalization helps prevent gradient explosion and gradient vanishing problems, making the model easier to converge. Residual connections are by adding the input directly to the output. Residual connections help the model learn more complex functions. This structure can avoid information loss in deep networks and improve the performance of the model.

[0160] S209, output features: After being processed by multiple Transformer layers, the model outputs dynamic grayscale features. These features are then used for feature fusion and combined with features of other modalities (such as point cloud features and static elastic features) to perform the final BI-RADS grading of breast lesions.

[0161] S30. Process the ultrasonic elastic imaging image through a 2D convolutional neural network to extract static elastic features.

[0162] Specifically, in ultrasound elastography, 2D convolutional neural networks (2D CNNs) are used to extract static elastic features, which can reflect the hardness information of lesions and are important indicators for judging whether lesions are benign or malignant. 2D CNN processes ultrasound elastography images through its convolutional layers, automatically learns local features in the image, such as edges, textures, etc., and gradually extracts higher-level features as the network layers deepen. The architecture of 2D CNN usually includes multiple convolutional layers, activation functions, pooling layers, fully connected layers, and output layers. The convolutional layer performs local perception and feature extraction on the input image through a series of learnable convolutional kernels (or filters). Each convolutional kernel corresponds to a specific feature pattern, which is moved on the input image in a sliding window manner, and the dot product between the convolutional kernel and the local area of ​​the image is calculated to obtain a feature map. This process helps capture local features in the image, such as the edges and internal structures of the lesions. Activation functions, such as ReLU, are applied to the output of the convolutional layer to introduce nonlinearity, allowing the network to learn more complex features. The pooling layer, which usually uses maximum pooling or average pooling, is used to reduce the spatial dimension of the feature map, reduce the amount of computation, and improve the level of abstraction of the features. After a series of convolution and pooling operations, the feature map is flattened and input to the fully connected layer. The fully connected layer integrates the learned local features to form a global feature representation. The advantage of 2D CNN lies in its local perception and parameter sharing characteristics, which enables the network to process image data efficiently while reducing the risk of overfitting. In addition, multi-layer convolution and pooling operations enable CNN to process features at different levels and capture information at different scales in the image. In ultrasound elastic imaging, 2D CNN uses these characteristics to effectively extract static elastic features from images that are important for lesion classification.

[0163] Specifically, in this embodiment, the specific steps of step S30 are as follows:

[0164] S301, Data Preparation: Before training 2D CNN, the ultrasound elastography images need to be preprocessed to ensure that the model can learn and generalize better. The preprocessing steps include normalization, cropping / padding, and data enhancement.

[0165] Normalization: Normalizing all image pixel values ​​to between 0 and 1 helps speed up the training process and improve model performance.

[0166]

[0167] Among them, I is the original image and I' is the normalized image.

[0168] Cropping / Padding: Make sure all input images have the same size H×W. This step is to make the image suitable as input to the network. If the image size is larger than the target size, the center part can be cropped from it; if the image size is smaller than the target size, zeros or other background values ​​can be filled around the image.

[0169] Data augmentation: Increase the diversity of training data through operations such as rotation, flipping, and scaling to prevent overfitting. Specific operations include randomly rotating the image by a certain angle, flipping the image horizontally or vertically, randomly scaling the image, and randomly cutting part of the image.

[0170] S302, build 2D CNN model: 2D CNN model usually consists of multiple convolutional layers, activation functions, pooling layers, fully connected layers and output layers. These layers work together to extract useful features from the input image and make the final prediction.

[0171] Input layer: accepts preprocessed image data with shape (H, W, C), where H is the height, W is the width, and C is the number of channels (for grayscale images, C = 1; for RGB images, C = 3).

[0172] Convolution layer: contains multiple convolution kernels (or filters), each of which is responsible for detecting specific features in the input image. The convolution kernel slides on the input image, calculates the dot product with the local area of ​​the image, and generates a new feature map. The specific formula for the convolution operation is:

[0173] O = σ(W*I+b);

[0174] Among them, O is the output feature map, W is the convolution kernel, I is the input image, b is the bias term, * represents the convolution operation, and σ is the activation function (such as ReLU).

[0175] Activation function: The activation function introduces nonlinearity, allowing the network to learn more complex features. Common activation functions include ReLU, Sigmoid, and Tanh. The ReLU function changes all negative values ​​to zero and retains positive values:

[0176] ReLU(x)=max(0,x);

[0177] Pooling layer: Usually, maximum pooling or average pooling is used to reduce the spatial dimension of the feature map, reduce the amount of calculation, and improve the abstract level of the feature. Maximum pooling takes the maximum value in each pooling window, while average pooling takes the average value in each pooling window. The specific formula of the pooling operation is:

[0178] O pool =Pool(O);

[0179] Among them, O poolIt is the feature map after pooling, and Pool is the pooling operation.

[0180] Multi-layer stacking: By stacking multiple convolutional layers and pooling layers, the network can learn more complex features. Each layer extracts higher-level abstract features.

[0181] Fully connected layer: After the convolution layer and the pooling layer, one or more fully connected layers are usually added to perform high-level abstraction on the extracted features and finally output the feature vector. The fully connected layer flattens the output of the previous layer into a one-dimensional vector, performs a linear transformation through the weight matrix and the bias term, and then performs a nonlinear transformation through the activation function. The specific formula of the fully connected layer is:

[0182] O fc =σ(W fc Flatten(O pool )+b fc );

[0183] Among them, O fc is the output of the fully connected layer, W fc is the weight matrix of the fully connected layer, b fc is the bias term and Flatten is the flattening operation.

[0184] Output layer: Outputs the final prediction results according to the task requirements, such as classification or regression. For classification tasks, the softmax function is usually used to convert the output into a probability distribution:

[0185]

[0186] Among them, z i is the i-th score of the output of the fully connected layer, and the denominator is the exponential sum of all scores and is used for normalization to ensure that the sum of the probabilities of all classes is 1.

[0187] S303. Select loss function and optimizer: Loss function: Select a suitable loss function according to different tasks. For example, cross entropy loss can be selected for classification tasks.

[0188]

[0189] Among them, y i is the true label, p i is the predicted probability.

[0190] Optimizer: Commonly used optimizers include Adam, SGD, etc., which are responsible for updating the network weights to minimize the loss function.

[0191]

[0192] Among them, θ is the model parameter, η is the learning rate, is the gradient of the loss function with respect to the parameters.

[0193] S304, training model: divide the data set: divide the data set into training set, validation set and test set, which are used for model training, hyperparameter adjustment and final evaluation respectively.

[0194] Training process:

[0195] Forward propagation: input training data and get predicted output through the network.

[0196]

[0197] in, is the predicted output, x is the input data, f is the model function, and θ is the model parameter.

[0198] Calculate loss: Calculate the loss value based on the predicted output and the true label.

[0199]

[0200] Backpropagation: Update the network weights based on the loss value.

[0201]

[0202] Iterative optimization: Repeat the above steps until the model converges or reaches the predetermined number of training rounds.

[0203] S305, feature extraction: After the training is completed, the trained 2D CNN model can be used to extract features from new ultrasound elastic imaging images. In particular, features can be extracted from the last convolutional layer or a layer before the fully connected layer.

[0204] F = ExtractFeature(x;θ);

[0205] Among them, F is the feature vector extracted from the ultrasound elastography image, which can be used to combine the extracted features with other types of data to build a more accurate breast lesion BI-RADS grading task model.

[0206] S40, integrating the point cloud features, dynamic grayscale features and static elastic features through a weighted algorithm or a multimodal neural network structure.

[0207] Specifically, feature fusion refers to combining data features from different sources or different types to improve the performance of the model. In ultrasound elastic imaging, lesion point cloud features, dynamic grayscale features, and static elastic features provide different information. Through effective feature fusion methods, the complementary information of these features can be fully utilized to enhance the robustness and accuracy of classification.

[0208] The weighted algorithm is a simple and effective method to fuse features by giving different weights to different features. The choice of weights can be based on experience, cross-validation or other optimization methods.

[0209] Feature standardization: First, each feature is standardized to make it comparable.

[0210]

[0211] Among them, F i is the i-th class feature, μ i and σ i are the mean and standard deviation of the feature respectively.

[0212] Weighted fusion: The standardized features are weighted and summed according to the preset weights.

[0213]

[0214] Among them, F fused is the fused feature, w i is the weight of the i-th feature, and N is the number of features.

[0215] Weight optimization: weight w i It can be determined by cross-validation or other optimization methods (such as grid search, genetic algorithm, etc.) to maximize the classification performance.

[0216] Cross-validation: Perform multiple cross-validations on the training set to select the weight combination that gives the highest classification performance;

[0217] Grid search: try different weight combinations and select the best combination;

[0218] Genetic algorithm: Find the optimal weight combination by simulating natural selection and genetic mechanisms.

[0219] Multimodal neural network structure is a more complex but more powerful method that fuses different types of features by designing a specialized network structure. This method can automatically learn the relationship and complementary information between different features.

[0220] Feature extraction: First, different network structures are used to extract lesion point cloud features, dynamic grayscale features, and static elastic features respectively.

[0221] Lesion point cloud features: Point cloud features are extracted using a point cloud processing network.

[0222] Dynamic grayscale features: Use visual Transformer to extract dynamic grayscale features.

[0223] Static elastic features: Static elastic features can be extracted using 2D convolutional neural networks.

[0224] Feature fusion layer: Design a fusion layer to combine different types of features. Common fusion methods include concatenation, weighted sum, and attention mechanism.

[0225] Concatenation: directly concatenate different features into a high-dimensional feature vector.

[0226] F fused =[F1;F2;F3];

[0227] Weighted sum: sum different features according to their weights.

[0228]

[0229] Attention mechanism: Dynamically adjust the weights of different features through the attention mechanism.

[0230]

[0231] Among them, α i are the weights calculated by the attention mechanism.

[0232] Fully connected layer: The fused features are input into the fully connected layer for the final classification or regression task.

[0233] O=σ(W fc ·F fused +b fc );

[0234] Among them, O is the final output, W fc is the weight matrix of the fully connected layer, b fc is the bias term and σ is the activation function.

[0235] S50, processing the fused features through a neural network layer, and inputting the processed features into a classification layer to perform BI-RADS grading of breast lesions.

[0236] Specifically, there are 7 levels in BI-RADS classification, and the definition of each level is as follows:

[0237] BI-RADS 0: Incomplete assessment, further imaging examinations (such as ultrasound, MRI, etc.) are required to complete the assessment.

[0238] BI-RADS1: negative, no abnormal lesions were found.

[0239] BI-RADS2: Benign lesions, such as cysts, fibroadenomas, etc., do not require further treatment.

[0240] BI-RADS 3: Possibly benign lesions; short-term follow-up (e.g., 6-month review) is recommended; risk of malignancy is less than 2%.

[0241] BI-RADS 4: Suspicious for malignancy, requiring further biopsy or surgery, with a risk of malignancy between 2% and 95%. It is divided into three subcategories:

[0242] BI-RADS 4A: low suspicion for malignancy, risk of malignancy 2% to 10%.

[0243] BI-RADS 4B: moderate suspicion of malignancy, risk of malignancy 10% to 50%.

[0244] BI-RADS 4C: High suspicion of malignancy, risk of malignancy 50% to 95%.

[0245] BI-RADS 5: Highly likely malignant lesion, with a risk of malignancy greater than 95%, requiring immediate action (such as biopsy or surgery).

[0246] BI-RADS 6: Known malignant lesions, usually pathologically confirmed malignant tumors, require active treatment.

[0247] The model that integrates point cloud features, dynamic grayscale features, and static elastic features can provide support for this grading process. By analyzing these features to predict the BI-RADS classification of lesions, the accuracy and objectivity of grading can be improved and the interference of human factors can be reduced.

[0248] Specifically, in this embodiment, the specific steps of performing BI-RADS grading are as follows:

[0249] S501, feature fusion: point cloud features, dynamic grayscale features and static elastic features are fused through weighted algorithms or multimodal neural network structures to make full use of complementary information of different features.

[0250] S502, feature processing: The fused features are processed through a series of neural network layers, which may include fully connected layers, normalization layers, and activation function layers. The functions of these layers are to further extract features, reduce overfitting, and increase the nonlinear expression ability of the model.

[0251] S503, classification layer: The processed features are sent to the classification layer, which is usually one or more fully connected layers, and the number of its output nodes corresponds to the number of BI-RADS categories (7 categories). In the classification layer, each node represents the score of a category.

[0252] S504, Softmax function: The output of the classification layer is converted into a probability distribution through the softmax function, so that the model can output the probability of each BI-RADS category.

[0253] S505, prediction: According to the probability distribution output by the softmax function, the category with the highest probability is selected as the prediction result of the model. This prediction result represents the model's judgment on the BI-RADS grading of breast lesions.

[0254] S506, post-processing: In practical applications, the prediction results of the model may also need to be post-processed, such as threshold adjustment, result interpretation, etc., to meet the needs of clinical applications.

[0255] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, any one of the methods in the present embodiment is implemented.

[0256] An embodiment of the present invention further provides an electronic terminal, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes any one of the methods in this embodiment.

[0257] The computer-readable storage medium in this embodiment can be understood by ordinary technicians in this field: all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to the computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, the execution includes the steps of the above-mentioned method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk and other media that can store program codes.

[0258] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store computer programs, the communication interface is used to communicate, and the processor and the transceiver are used to run computer programs so that the electronic terminal executes each step of the above method.

[0259] The multimodal breast volume ultrasound lesion grading method, medium and terminal described in the above embodiments, compared with the prior art, at present, the subjectivity is strong and the consistency is poor in the process of breast cancer lesion detection, the feature description is not comprehensive enough, and the accuracy of lesion grading is low. The process of the present invention is simple and easy to operate. The multiple features of breast tissue are obtained through multimodal ultrasound data. The 3D-Unet model, the visual Transformer model and the 2D convolutional neural network (2D CNN) are combined to realize the three-dimensional segmentation of the lesion, the dynamic grayscale feature extraction and the static elastic feature extraction, and then the feature fusion and the intelligent grading model are used to realize the accurate grading of breast lesions. The present invention comprehensively describes the lesion characteristics through multimodal feature fusion, improves the accuracy of lesion detection; the automated grading method reduces the influence of doctor's experience, improves the consistency of diagnosis results, and reduces the subjective influence of doctors; improves efficiency, has a high degree of automation, is suitable for large-scale screening, and improves work efficiency.

[0260] Obviously, the embodiments described above are only preferred embodiments of the present invention, rather than all embodiments. The preferred embodiments of the present invention are shown in the accompanying drawings, but they do not limit the patent scope of the present invention. The present invention can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present invention more thorough and comprehensive. Although the present invention has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific embodiments, or to perform equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of the present invention, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present invention.

Claims

1. A multimodal breast volume ultrasound lesion grading method, characterized in that: The following steps are involved: S10, using an automatic volumetric ultrasound device to perform a full coverage scan of the target area, obtain automatic volumetric ultrasound data and perform preprocessing, and perform three-dimensional segmentation and point cloud feature extraction on the preprocessed data; S20, processing the dynamic grayscale ultrasound time series through the visual Transformer model to extract dynamic grayscale features; S30, processing the ultrasound elastic imaging image through a 2D convolutional neural network to extract static elastic features; S40, fusing the point cloud features, dynamic grayscale features, and static elastic features through a weighted algorithm or a multimodal neural network structure; S50, processing the fused features through a neural network layer, and inputting the processed features into a classification layer to perform BI-RADS grading of breast lesions.

2. A multimodal breast volume ultrasound lesion grading method according to claim 1, characterized in that: The specific steps of obtaining automatic volume ultrasound data and preprocessing in step S10 are as follows: S101, obtaining cube data V (x, y, z) resolution of three-dimensional ultrasound data, where the data resolution is millimeter level in x, y, z directions; S102, using non-local mean filtering or three-dimensional Gaussian filtering to remove background noise, the filtering formula is: Among them, w(i,j,k) is the weight, which is related to the similarity between voxels; S103, adjusting the image contrast to enhance the lesion boundary features; S104, performing Z-score normalization on the image sequence data, assuming that the original data is V'(x, y, z), the normalization operation is expressed as: Among them, μ and δ are the pre-calculated mean and standard deviation. The data is normalized by Z-score standardization operation to satisfy the normal distribution, with a mean of 0 and a standard deviation of 1.

3. A multimodal breast volume ultrasound lesion grading method according to claim 2, characterized in that: In step S10, a 3D U-Net network is used for three-dimensional segmentation. The 3D U-Net network includes an analysis path and a synthesis path. In the analysis path, each layer includes two 3×3×3 convolution operations, each convolution operation is followed by a ReLU activation function, and also includes a maximum pooling layer; in the synthesis path, each layer restores the resolution of the feature map through a transposed convolution, and the upsampled feature map is spliced ​​with the feature map of the corresponding layer in the analysis path according to the channel dimension to integrate multi-level information. The spliced ​​feature map is again subjected to two 3×3×3 convolution operations to extract features, and then a probability distribution of each voxel is generated through a Softmax layer. According to the probability distribution output by the Softmax layer, a threshold is applied for binary segmentation, and the voxel position with a value of 1 is extracted from the segmentation result, which is converted into a three-dimensional coordinate point to generate point cloud data of the lesion area.

4. A multimodal breast volume ultrasound lesion grading method according to claim 3, characterized in that: In step S10, a point cloud network is used to extract point cloud features. The point cloud data consists of a set of points in three-dimensional space. Each point is defined by its own coordinates. The point cloud is aligned through an input transformation network, and the aligned point cloud is sent to a shared multi-layer perceptron to extract local features from each point. A feature transformation network is used to improve the alignment of features, and local features are integrated through a maximum pooling layer to form a global feature vector for subsequent classification and segmentation tasks.

5. A multimodal breast volume ultrasound lesion grading method according to claim 4, characterized in that: For classification tasks, the point cloud network will extract the global feature vector through the fully connected layer of the classification branch, and finally output the concept of each category. The classification branch includes several fully connected layers. The fully connected layer processes the global feature vector to generate the final classification result and identify the category to which the point cloud belongs. For segmentation tasks, the point cloud network concatenates the global feature vector with the local features of each point, and then outputs the segmentation result of each point through the convolutional layer of the segmentation branch. The segmentation branch includes several convolutional layers. The convolutional layer processes the concatenated feature vector to generate the segmentation label of each point and identify the category to which each point belongs.

6. A multimodal breast volume ultrasound lesion grading method according to claim 1, characterized in that: The specific steps of step S20 are as follows: S201, the visual Transformer model receives a dynamic grayscale ultrasound time series as input; S202, decomposing the large image of the image block into small local regions, each block is flattened into a one-dimensional vector, and projected into a vector space of fixed dimension through a linear transformation; S203, adding position codes to distinguish image blocks at different positions; S204, analyzing the relationship between image blocks using a self-attention mechanism; S205. Using the multi-head attention mechanism, the self-attention mechanism is independently operated in multiple subspaces, and information is processed in parallel in multiple subspaces to enhance the expressiveness of the model; S206, extracting the local information of a single time point and the dynamic changes between different time points from the dynamic grayscale ultrasound time series through the self-attention mechanism and the multi-head attention mechanism; S207, perform nonlinear transformation on the output of the self-attention layer through the feedforward network of the visual Transformer model to extract higher-level features; S208, Visual Transformer model uses layer normalization and residual connections in self-attention layers and feed-forward networks; S209, the visual Transformer model outputs dynamic grayscale features.

7. A multimodal breast volume ultrasound lesion grading method according to claim 1, characterized in that: The specific steps of step S30 are as follows: S301, preprocessing the ultrasound elastic imaging image, including normalization, cropping / filling and data enhancement; S302, construct a 2D CNN model, including several convolutional layers, activation functions, pooling layers, fully connected layers and output layers; S303, selecting a loss function and an optimizer; S304, dividing the data set into a training set, a validation set and a test set, which are used for model training, hyperparameter adjustment and final evaluation respectively. The model training process includes forward propagation, loss calculation, back propagation and iterative optimization; S305. Use the trained 2D CNN model to extract features from the new ultrasound elastography image.

8. A multimodal breast volume ultrasound lesion grading method according to claim 1, characterized in that: The BI-RADS classification of breast lesions in step S50 includes the following levels: BI-RADS 0: Incomplete assessment, further imaging examinations are required to complete the assessment; BI-RADS 1: Negative, no abnormal lesions were found; BI-RADS 2: Benign lesions; BI-RADS 3: Probable benign lesions, with a malignancy risk of less than 2%; BI-RADS 4: Suspected malignant lesions, with a malignancy risk between 2% and 95%; BI-RADS 5: Highly likely malignant lesions, with a malignancy risk greater than 95%; BI-RADS 6: Known malignant lesions; The specific steps of the BI-RADS grading of breast lesions are as follows: S501, fusing point cloud features, dynamic grayscale features, and static elastic features; S502, the fused features are processed by a neural network layer to further extract features; S503, the processed features are sent to the classification layer, the number of output nodes of the classification layer corresponds to the number of BI-RADS categories, and in the classification layer, each node represents the score of a category; S504, the output of the classification layer is converted into a probability distribution through a Softmax function, so that the model outputs the probability of each BI-RADS category; S505. According to the probability distribution output by the Softmax function, select the category with the highest probability as the prediction result of the model; S506: Adjust the threshold of the prediction result and interpret the result.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

10. An electronic terminal, characterized in that: include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Body surface marking system for primary focus of breast cancer

    CN118177994A

Cited By

  • Ultrasound image processing device and method suitable for ultrasound-guided interventional surgical robot

    CN120381302A

  • Ultrasonic image processing device and method suitable for ultrasound-guided interventional surgery robot

    CN120381302B

  • Liver cancer focus segmentation method and system based on dynamic feature fusion

    CN120632797A

  • Breast ultrasonic automatic pressurization method and device based on image feedback

    CN121081018A

  • Breast ultrasound video processing method, device, equipment, medium and product

    CN122636618A