Intelligent breast lump diagnosis system based on combination of ultrasonic-molybdenum target multi-mode image and serum P16 protein detection
The intelligent diagnostic system that combines ultrasound-mammography multimodal imaging with serum P16 protein detection solves the problem of insufficient accuracy in early diagnosis of breast cancer in existing technologies, achieves efficient and reliable diagnosis of breast masses, improves the sensitivity and specificity of diagnosis, reduces the misdiagnosis rate, and enhances the clinical practicality of the system.
Patent Information
- Application Number
- CN202510700542.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing breast cancer diagnostic methods are insufficient in accuracy and sensitivity in diagnosing early breast cancer. In particular, ultrasound examination is insensitive to microcalcifications, and mammography has a low detection rate in dense breasts, resulting in a high misdiagnosis rate and an inability to effectively distinguish between benign and malignant breast lumps.
An intelligent diagnostic system based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection is adopted. Cross-modal features are extracted through deep learning and radiomics technology, and feature fusion is achieved through a dynamic weight strategy. The weights are adjusted in combination with the Bayesian probability integration formula. Cascade classifiers and multi-task Transformer networks are used for efficient classification. FocalLoss is introduced to address data imbalance problems. Dynamic optimization and federated learning frameworks are used for model iteration and privacy protection. Explainability technology and visualization modules are introduced to enhance the clinical practicality of the system.
The sensitivity and specificity of breast mass diagnosis were significantly improved, with an AUC of 0.94, which is better than single or dual-modality combinations, providing more reliable diagnostic basis, reducing the misdiagnosis rate, and enhancing the clinical practicality and credibility of the system.
Smart Images

Figure CN120585359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ultrasound and mammography detection, and in particular to an intelligent breast lump diagnosis system based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection. Background Art
[0002] With the development of medical imaging, diagnostic methods such as ultrasound and mammography have been increasingly used in the diagnosis of breast masses. Ultrasound, due to its non-invasive, safe, simple, effective, and reproducible nature, has become widely used in the diagnosis of breast diseases. Mammography, on the other hand, offers clear imaging, convenient and rapid testing, and low radiation exposure. It can accurately detect the shape, size, density, and properties of breast hyperplasia, lesions, masses, and calcifications. Furthermore, this study has found that serum P16 testing can serve as an important supplement to mammography and ultrasound in clinical practice, playing an increasingly important role in the diagnosis and treatment of breast diseases. Therefore, this study investigates the value of combining ultrasound and mammography with serum P16 autoantibodies in the diagnosis of benign and malignant breast masses. This study aims to improve the diagnostic accuracy of benign and malignant breast masses and provide reliable clinical data for the subsequent treatment of breast disease patients. This study is of great significance for the comprehensive advancement of diagnostic imaging technology in our province.
[0003] The incidence of breast diseases is currently increasing worldwide, significantly impacting women's physical and mental health. Breast cancer is a common malignant tumor in women, affecting over 1.3 million women worldwide each year, accounting for approximately 10.6% of all malignant tumors. Furthermore, breast cancer is highly aggressive, with no effective preventive measures. It can metastasize throughout the body even in its early stages, severely impacting women's physical and mental health. Studies have shown that the five-year survival rate for early-stage breast cancer exceeds 90%, while for late-stage patients, it is only around 20%. Early detection of breast tumors and differentiation between benign and malignant tumors are crucial for improving survival rates in breast cancer patients, crucial for selecting treatment options and ensuring patient outcomes. Therefore, improving the accuracy of breast cancer diagnosis remains a hot topic in this field.
[0004] The symptoms of early breast cancer are often not obvious, and are often dominated by local symptoms such as breast lumps, breast skin abnormalities, nipple discharge, and nipple or areola abnormalities. Because the manifestations are not obvious, they are very easy to be ignored. Breast lumps are the most common symptom of early breast cancer. Most breast cancers are painless lumps, and a few cases are accompanied by varying degrees of dull pain or tingling. At present, ultrasound, MRI, and mammography are often used as routine protocols for benign and malignant breast lumps in clinical diagnosis. In recent years, many studies have been conducted at home and abroad on the ultrasound diagnosis of benign and malignant breast lumps. One of the research reports used two-dimensional ultrasound to observe the morphology, boundaries, internal echoes, the presence of posterior attenuation and lateral acoustic shadowing of the mass, and used color Doppler blood flow imaging to observe the blood flow distribution inside the mass, and measured the blood flow resistance index (RI). The results showed that among 92 breast masses in 80 patients, 61 ultrasound-diagnosed breast cancer masses were consistent with pathological findings, while 3 were misdiagnosed as benign masses, for a diagnostic concordance rate of 95.3%. Another 27 ultrasound-diagnosed benign masses were consistent with pathological findings, with 1 misdiagnosed as malignant, for a diagnostic concordance rate of 96.4%. These results demonstrate that color Doppler ultrasound is highly valuable in differentiating benign from malignant breast masses, with a high concordance rate with pathology. Ultrasound diagnosis is simple and noninvasive, but color Doppler ultrasound has limitations, such as the tendency for sonograms to overlap, which can lead to diagnostic bias. While it has positive clinical value, it should still be used in conjunction with other diagnostic modalities. Summary of the Invention
[0005] The purpose of the present invention is to provide an intelligent diagnosis system for breast masses based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection. The system fuses the multimodal data of ultrasound imaging, mammography imaging, and serum P16 protein detection, and achieves optimal integration of cross-modal features through a dynamic weighting strategy. Experimental results show that the AUC (area under the curve) of the combined diagnosis scheme reaches 0.94, which is better than a single modality or a dual-modality combination, providing a more reliable diagnostic basis for clinical practice.
[0006] In order to achieve the above technical objectives and the above technical effects, the present invention is implemented through the following technical solutions:
[0007] An intelligent breast mass diagnosis system based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection, including:
[0008] The data acquisition module is responsible for the standardized collection of ultrasound, mammography and serum P16 protein data to ensure the compatibility and consistency of multi-source data. It is the basic input layer of the system.
[0009] The feature extraction and fusion module uses deep learning and imaging omics technology to achieve refined extraction and dynamic weight fusion of cross-modal features.
[0010] The intelligent diagnostic decision-making module achieves accurate classification through a hierarchical decision-making model and embeds clinical logical reasoning rules.
[0011] Dynamic optimization and visualization modules enable system self-iteration and transparency of the diagnostic process.
[0012] Beneficial effects of the present invention:
[0013] This study fuses multimodal data from ultrasound, mammography, and serum P16 protein testing, achieving optimal integration of cross-modal features through a dynamic weighting strategy. Ultrasound provides information on tumor morphology (e.g., edge irregularity and echogenicity uniformity) and hemodynamics (e.g., resistance index (RI) and Adler grade); mammography reveals features such as calcification density, cluster heterogeneity, and depth of structural distortion; and serum P16 protein testing provides molecular insights into malignancy risk, particularly valuable for the early identification of triple-negative breast cancer. A graph neural network (GNN) constructs a feature association graph, and dynamically adjusts weights using a Bayesian probability integration formula to ensure optimal integration of features from different modalities. In dense breasts, the mammography weight is increased to 0.45 to enhance calcification detection; in highly vascularized masses, the ultrasound feature weight is increased to improve the accuracy of hemodynamic analysis. This dynamic weighting strategy not only overcomes the limitations of single modalities (e.g., ultrasound's insensitivity to microcalcifications and mammography's low detection rate in dense breasts) but also significantly improves diagnostic sensitivity and specificity. Experimental results show that the AUC (area under the curve) of the combined diagnostic scheme reached 0.94, which is better than a single modality or dual-modality combination, providing a more reliable diagnostic basis for clinical practice.
[0014] The present invention uses deep learning and radiomics technology to extract high-dimensional features from complex medical images and achieves efficient classification through cascade classifiers and multi-task Transformer networks. Ultrasound images use the U-Net++ network to segment the mass boundary and calculate morphological indicators (such as circularity and posterior echo attenuation slope) and hemodynamic parameters (such as the angiogenesis index TAI); mammography images use a multi-scale feature pyramid network (MS-FPN) to analyze the heterogeneity of calcification clusters and the depth of structural distortion; serum P16 protein concentration is associated with clinical parameters to construct a logistic regression model. Feature fusion uses kernel principal component analysis (KPCA) to map multimodal features to a high-dimensional space, retaining the most discriminative information. The cascade classifier uses a hierarchical screening strategy to first use a first-level LightGBM model to quickly screen high-risk cases (based on L / S ratio, calcification density, and TAI×RI comprehensive score ≥6.5), and then uses a second-level multi-task Transformer network to simultaneously output BI-RADS grade and molecular subtype prediction. High P16 expression indicates triple-negative breast cancer. This combination of refined feature extraction and efficient classification not only significantly improves diagnostic accuracy but also reduces computational complexity, making it suitable for large-scale screening scenarios. Furthermore, the introduction of FocalLoss addresses data imbalance and further optimizes the model's performance on minority samples (such as malignant tumors).
[0015] The present invention realizes the continuous evolution of the model and multi-center data privacy protection through dynamic optimization and federated learning framework. The dynamic optimization module monitors the model performance in real time, dynamically adjusts the hyperparameters (such as the kernel width parameter γ and regularization parameter C of SVM) based on the Bayesian optimization method, and uses Monte Carlo Dropout simulation to evaluate controversial cases (σ<0.1 is the credible threshold) to ensure the reliability of the diagnosis results. The federated learning framework implements model iteration under the protection of multi-center data privacy, and adds Laplace noise through the differential privacy mechanism (ε=0.5) to ensure data security. The negative feedback closed loop automatically triggers the comparison of misdiagnosed case features and online learning, and updates the model parameters within 24 hours to ensure that the system maintains high diagnostic performance in practical applications. The diagnostic threshold T is dynamically adjusted in the screening scenario to improve sensitivity, and the threshold is increased in the diagnosis scenario to enhance specificity. This combination of dynamic optimization and federated learning not only significantly improves the generalization and continuous evolution capabilities of the model, but also provides safe and reliable technical support for multi-center collaboration, promoting the popularization and optimization of breast tumor diagnosis technology.
[0016] This invention significantly enhances the system's clinical practicality and decision-making support capabilities through interpretability technologies and visualization modules. Interpretability technologies (such as SHAP and LRP) visualize feature contributions, generating heat maps superimposed on image annotations of high-risk areas to help clinicians understand the model's decision-making logic. Microcalcification density contributes 28% to the diagnostic result, suggesting it is a key feature of malignant masses. The visualization module supports interactive weight adjustment within a 3D fusion view and dynamically optimizes diagnostic results based on clinical needs. The model performance visualization unit comprehensively displays the model's sensitivity and specificity by plotting receiver operating characteristic (ROC) curves and confusion matrices. The feature importance visualization unit displays the contribution of each modality's features through bar charts or radar plots, helping clinicians optimize data acquisition and model design. The user interaction logic unit provides clinicians with a flexible interface, supporting parameter adjustment and diagnostic feedback. By clicking a point on the ROC curve, clinicians can update the confusion matrix and diagnostic result chart in real time and optimize the model based on the feedback data. This combination of interpretability and visualization technologies not only significantly enhances the system's clinical practicality but also provides clinicians with intuitive and reliable decision support, improving diagnostic credibility and clinical value.
[0017] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0021] Example 1
[0022] The intelligent breast mass diagnosis system described in this embodiment, based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection, includes:
[0023] The data acquisition module integrates ultrasound, mammography, and serum P16 protein test data, using standardized protocols to ensure data consistency. Ultrasound images use a high-frequency probe to obtain two-dimensional grayscale images and hemodynamic parameters (Adler grade, resistance index RI), and an improved motion artifact suppression algorithm is used to improve image quality. Mammography images use adaptive histogram equalization (CLAHE) to enhance microcalcification contrast, and a deep denoising network (DnCNN) is used to eliminate quantum noise and quantify calcification density (≥5 cells / cm 2 Serum P16 detection uses electrochemiluminescence assay (ECLIA) with β-actin as internal reference calibration, setting ≥2.5 ng / mL as the positive threshold to ensure detection sensitivity (CV < 5%) and spatiotemporal synchronization with imaging data (interval ≤ 72 hours).
[0024] The feature extraction and fusion module extracts cross-modal features based on deep learning and radiomics technologies. Ultrasound images are segmented using the U-Net++ network to calculate morphological indicators (circularity <0.7, posterior echo attenuation slope ≥2.0dB / cm) and hemodynamic parameters (angiogenesis index (TAI)). Mammography images are analyzed using a multi-scale feature pyramid network (MS-FPN) to analyze calcification cluster heterogeneity (longitudinal standard deviation >15°) and depth of structural distortion (>5mm). A logistic regression model is constructed by correlating serum P16 concentrations with clinical parameters. Feature fusion utilizes a dynamic weighting strategy: a feature association graph is constructed using a graph neural network (GNN), and weights are dynamically adjusted using a Bayesian probability integration formula (e.g., the mammography weight is increased to 0.45 in dense breasts), achieving multi-dimensional information complementarity across anatomical, functional, and molecular dimensions.
[0025] The intelligent diagnostic decision-making module utilizes a cascade classifier to improve diagnostic efficiency. A primary LightGBM model rapidly screens high-risk cases (based on L / S ratio, calcification density, and a combined TAI × RI score ≥ 6.5), achieving an AUC of 0.94. A secondary multi-task Transformer network simultaneously outputs BI-RADS grading and molecular subtype prediction (e.g., high P16 expression suggests triple-negative breast cancer), incorporating FocalLoss to address data imbalance. A built-in clinical rule base (e.g., P16 positivity + microcalcification clusters + RI ≥ 0.75 mandates an upgrade to BI-RADS Category 5) is implemented, and Monte Carlo Dropout simulations are implemented for controversial cases (with a confidence threshold of σ < 0.1) to reduce the risk of misdiagnosis.
[0026] The dynamic optimization and visualization module uses interpretability techniques (SHAP and LRP) to visualize feature contributions (e.g., microcalcification density accounts for 28%), generating heat maps overlaid on high-risk image annotations. It also supports interactive weight adjustment within a 3D fusion view. A federated learning framework is used to implement model iteration while protecting multi-center data privacy (annual improvement ≥ 3%), and Laplace noise is added through a differential privacy mechanism (ε = 0.5). A negative feedback loop automatically triggers feature comparison and online learning for misdiagnosed cases, updating model parameters within 24 hours to ensure continuous system evolution.
[0027] In this embodiment, the data acquisition module includes an ultrasound image acquisition unit, a mammography image acquisition unit, and a serum P16 protein detection unit, specifically including:
[0028] Ultrasound image acquisition unit,
[0029] A high-frequency probe (frequency ≥ 7.5 MHz) was used to acquire a two-dimensional grayscale image of the breast mass and generate a pixel matrix I (x, y).
[0030] Image acquisition parameters include gain, depth, and focus. Parameter settings are dynamically adjusted using the following formula:
[0031] Gain: G = G min +(G max -G min )·(1-e -α·T )
[0032] Among them, G min and G max are the minimum and maximum gain values, α is the attenuation coefficient, and T is the tissue penetration depth.
[0033] Depth: D = D base +(1-e -β·S )
[0034] Among them, D base is the base depth, β is the depth adjustment factor, and S is the mass size.
[0035] Focus: F=F center +ΔF·sin(ω·t)
[0036] Among them, F center is the center position of the tumor, ΔF is the focus shift amplitude, ω is the frequency, and t is the time.
[0037] Gain: Dynamically adjusted according to tissue density, ranging from -20dB to +20dB.
[0038] Depth: Set to 4cm to 8cm depending on the location of the lump.
[0039] Focus: Set to the center of the mass to ensure clarity in the focused area.
[0040] In this embodiment, color Doppler flow imaging (CDFI) is used to obtain blood flow information and generate a blood flow velocity matrix V(x, y).
[0041] The calculation formula of resistance index RI is:
[0042] RI=(V max -V min ) / V max
[0043] Among them, V max and V min The peak systolic velocity and the minimum diastolic velocity are respectively.
[0044] The Adler classification is divided into 0-3 grades based on the richness of blood flow signals. The grading standards are as follows:
[0045] Adler=∑(w i ·N i )
[0046] Among them, w i is the weight coefficient, N i is the number of blood vessels.
[0047] In this embodiment, an improved motion artifact suppression algorithm is used in combination with Kalman filtering and adaptive motion compensation technology to eliminate artifacts caused by patient breathing or movement.
[0048] The algorithm formula is:
[0049] I corrected (x,y)=I(x,y)-K·(I(x,y)-I prev (x, y)
[0050] Where K is the Kalman gain coefficient, I prev (x, y) is the pixel value of the previous frame.
[0051] The calculation formula of Kalman gain coefficient K is:
[0052] K=P prev / (P prev +R)
[0053] Among them, P prev is the prediction error of the previous frame, and R is the observation noise variance.
[0054] The motion artifact suppression formula dynamically corrects pixel values through Kalman filtering to eliminate artifacts.
[0055] The Kalman gain coefficient formula dynamically adjusts the correction weight through the ratio of prediction error to observation noise.
[0056] Mammography image acquisition unit,
[0057] Adaptive histogram equalization was performed on mammography images to enhance the contrast of microcalcifications.
[0058] The CLAHE algorithm formula is:
[0059] I enhanced (x,y)=CLAHE(I(x,y),clip limit , tile size )
[0060] Among them, clip limit is the contrast limit parameter, tile size is the size of the local area; clip limit : Set to 2.0 to prevent over-enhancement of local areas; tile size : Set to 8×8 pixels to ensure the equalization effect of local areas.
[0061] clip limit Dynamically adjusted by the following formula:
[0062] clip limit =clip base ·(1+γ·(1-e -δ·D ))
[0063] Among them, clip base is the basic limit value, γ and δ are adjustment factors, and D is the calcification density.
[0064] The CLAHE formula enhances the contrast of microcalcifications through local region histogram equalization.
[0065] clip limit The formula dynamically adjusts the contrast limit value through an exponential function to prevent over-enhancement.
[0066] In this embodiment, DnCNN is used to eliminate quantum noise and improve the image signal-to-noise ratio (SNR≥30dB).
[0067] The network output formula is:
[0068] I denoised (x,y)=DnCNN(I(x,y),θ)
[0069] Among them, θ is the network parameter.
[0070] During network training, the mean square error (MSE) is used as the loss function:
[0071] MSE=1 / N·∑(I denoised (x, y)-I clean (x, y) 2
[0072] Where N is the total number of pixels, I clean (x, y) is a noise-free image.
[0073] The DnCNN formula removes quantum noise through a deep learning network and improves image quality.
[0074] The MSE formula optimizes the network parameters by calculating the difference between the denoised image and the noise-free image.
[0075] The calculation formula for calcification density is:
[0076] Density = N calc / Area
[0077] Among them, N calc is the number of calcification points, and Area is the area of the region of interest.
[0078] The structural distortion depth is extracted by edge detection algorithm (such as Canny operator) and its maximum depth D is calculated max .
[0079] The calcification density formula quantifies the richness of calcification by comparing the number of calcification points to the area of the region. The structural distortion depth formula uses an edge detection algorithm to extract the edge of the mass and calculate the maximum depth.
[0080] Serum P16 protein detection unit
[0081] The ECLIA technique was used to detect the level of P16 protein in serum and generate the concentration value C P16 .
[0082] The detection formula is:
[0083] C P16 =k·(I sample -I blank ) / (I standard -I blank )
[0084] Where k is the calibration coefficient, I sample is the light intensity signal of the sample, I blank is the light intensity signal of the blank control, I standard is the light intensity signal of the standard.
[0085] The ECLIA formula quantifies P16 protein concentration using the ratio of light intensity signals. β-actin is used as an internal reference to calibrate the assay results, ensuring data stability and accuracy.
[0086] The calibration formula is:
[0087] C calibrated =C P16 / C β-actin
[0088] Among them, C β-actin The internal reference calibration formula eliminates systematic errors by calculating the ratio of P16 protein concentration to the internal reference concentration.
[0089] The positive threshold is set at ≥2.5 ng / mL to ensure detection sensitivity (CV < 5%). The positive threshold formula is determined based on experimental data to ensure the reliability of the test results.
[0090] In this embodiment, the feature extraction and fusion module specifically includes:
[0091] Ultrasound image feature extraction unit. Ultrasound images provide morphological and hemodynamic information of breast masses and are an important basis for mass detection. In morphological feature extraction, edge irregularity (EI) is a key indicator. First, the Canny edge detection algorithm is used to extract the edge contour of the mass. Canny edge detection removes noise through Gaussian filtering, calculates image gradients, and uses a double threshold method to retain strong edges. Subsequently, the edge perimeter P and the area A are calculated. The edge irregularity formula is:
[0092] EI = P / (2·π·sqrt(A / π))
[0093] The larger the EI value, the more irregular the edge of the mass, indicating a higher possibility of malignancy. In the extraction of internal echo features, echo homogeneity (EH) and echo intensity (EI) are two important indicators. Echo homogeneity is calculated by calculating the image grayscale mean μ and variance σ 2 Get. Grayscale mean μ represents the average brightness of the image, variance σ 2 Indicates the discrete degree of pixel value. The echo uniformity formula is:
[0094] EH=1-(σ 2 / μ 2 )
[0095] The smaller the EH value, the more uneven the echo is, indicating that the internal structure of the tumor is complex and may be related to malignant tumors. The echo intensity is directly obtained by calculating the grayscale mean μ, and the formula is: EI = μ
[0096] In hemodynamic feature extraction, the resistance index (RI) and Adler grade are two key indicators. A larger RI value indicates higher blood flow resistance, indicating a higher likelihood of malignant tumors. Spectral analysis is performed using fast Fourier transform (FFT) to extract the frequency domain characteristics of blood flow velocity. The FFT formula is:
[0097] F(k)=∑(f(n)·e -2πikn / N )
[0098] Among them, f(n) is the time domain signal, F(k) is the frequency domain signal, and k is the frequency index. Adler classification is calculated by counting the number of blood vessels N in the tumor. vessel And weighted calculation is obtained, the formula is:
[0099] Adler=∑(w i ·N i )
[0100] Among them, w i =[0, 0.3, 0.5, 0.7], N i is the number of blood vessels.
[0101] The Adler grading formula quantifies the richness of blood flow signals by weighting the coefficient and the number of blood vessels. The Adler grading system is divided into 0-3 levels based on the richness of blood flow signals. The grading standards are as follows:
[0102] Level 0: No blood flow signal.
[0103] Grade 1: Small amount of blood flow signal, 1-2 blood vessels.
[0104] Grade 2: moderate blood flow signal, 3-4 vessels.
[0105] Level 3: Rich blood flow signals, ≥5 blood vessels.
[0106] The higher the Adler grade, the richer the blood vessels in the tumor, indicating a higher possibility of malignancy.
[0107] Mammography image feature extraction unit: Mammography provides calcification and structural information of breast masses, which is an important basis for mass detection. In calcification feature extraction, calcification density (CD) and calcification clustering (CC) are two key indicators. Calcification density is calculated by counting the number of calcification points N through the local peak detection algorithm. calc , and calculate the calcification density per unit area using the formula:
[0108] CD=Ncalc / A
[0109] The larger the CD value, the denser the calcification, indicating a higher possibility of malignant tumor. Local peak detection is achieved by sliding window method, calculating the local maximum value of the gray value within the window. The calcification concentration is calculated by calculating the average distance D between calcification points. ij The formula is:
[0110] CC=1 / (mean(D ij ))
[0111] A larger CC value indicates a higher degree of calcification aggregation, indicating a higher likelihood of malignant tumors. Structural Distortion Depth (SDD) and Texture Complexity (TC) are two key indicators in structural feature extraction. Structural Distortion Depth is extracted using the Canny edge detection algorithm to determine the maximum depth, using the formula:
[0112] SDD = max(edge(I(x, y)))
[0113] The larger the SDD value, the more severe the structural distortion, indicating a higher possibility of malignant tumors. Texture complexity is calculated by gray-level co-occurrence matrix entropy, the formula is:
[0114] TC = -∑(p(i, j)·log(p(i, j))) The larger the TC value, the more complex the texture, which indicates a higher possibility of a malignant tumor.
[0115] Serum P16 protein feature extraction unit. Serum P16 protein testing provides molecular biological information and is an important basis for tumor detection. P16 concentration is determined by electrochemiluminescence. Higher P16 concentrations indicate a higher likelihood of malignancy.
[0116] Feature fusion unit, the purpose of feature fusion is to integrate multimodal features into a unified high-dimensional feature space. First, all features are normalized so that their value range is unified to [0, 1]. The formula is:
[0117] X normalized =(XX min ) / (X max -X min )
[0118] Then, different weights are assigned according to the discriminability of the features, and the feature weight w is calculated using the LASSO regression algorithm. i , the formula is:
[0119] min(||Y-Xβ|| 2 +λ||β||1)
[0120] Among them, Y is the label, X is the feature matrix, β is the regression coefficient, and λ is the regularization parameter. The features corresponding to the non-zero coefficients are screened out through LASSO regression, and the weighted fusion features are calculated:
[0121] X fused =∑(w i ·X i )
[0122] Finally, kernel principal component analysis (KPCA) is used to map the fused features into a high-dimensional space using the Gaussian kernel function:
[0123] K(x,y)=exp(-||xy|| 2 / (2σ 2 )) where σ is the kernel width parameter. KPCA is used to extract the principal components and retain the most important feature information.
[0124] Feature optimization unit, the purpose of feature optimization is to remove redundant features and retain the most discriminative feature combination. First, the LASSO regression algorithm is used to perform sparse processing on the features and filter out the features corresponding to non-zero coefficients. Then, the genetic algorithm is used to optimize the feature combination. The objective function is:
[0125] Fitness=AUC+α·(1-FeatureRatio)
[0126] Where AUC is the classification performance, FeatureRatio is the ratio of the number of features to the total number of features, and α is the balance parameter. The optimal feature combination is searched iteratively through the genetic algorithm.
[0127] In this embodiment, the intelligent diagnosis and decision-making module includes:
[0128] A classification model unit used to distinguish benign from malignant breast lumps. This model uses a deep ensemble learning approach, combining the strengths of multiple base learners to improve the model's robustness and generalization capabilities. The base learners include support vector machines (SVMs), random forests (RFs), and deep neural networks (DNNs). The support vector machine implements nonlinear classification using a Gaussian kernel function, whose formula is:
[0129] K(x,y)=exp(-||xy|| 2 / (2·γ 2 ))
[0130] Where γ is the kernel width parameter, which is optimized by grid search. Random forest is based on the integrated decision tree, and the Gini index Gini(D) = 1-∑(p k 2 ) for feature segmentation, where p kis the proportion of samples in the kth class. The deep neural network extracts deep features through multi-layer nonlinear transformation and adopts the cross entropy loss function L = -∑(y true ·log(y pred )) is optimized, where y true is the true label, y pred is the prediction probability. The deep ensemble learning framework adopts a weighted voting strategy to fuse the prediction results of multiple base learners. The formula is:
[0131] y final =sign(∑(w i ·y i ))
[0132] where w i is the weight of the i-th base learner, which is calculated by the AUC value of the validation set. The formula is w i =AUC i / ∑(AUC j ).
[0133] The model training and optimization unit first standardizes the fusion features to make their mean 0 and variance 1. The formula is X std =(X-μ) / σ, where μ is the feature mean and σ is the feature standard deviation. The optimization goal of the support vector machine is to maximize the classification interval, and the formula is:
[0134] min(||w|| 2 / 2+C·∑(ξ i ))
[0135] Where w is the normal vector of the classification hyperplane, C is the regularization parameter, ξ i is a slack variable. Random forest constructs multiple decision trees through bootstrap sampling and integrates them using majority voting. The deep neural network is trained using the Adam optimizer, and its momentum update formula is:
[0136] m t =β1·m {t-1} +(1-β1)·g t
[0137] v t =β2·v {t-1} +(1-β2)·g t 2
[0138] where m t and v t is the first and second order momentum, g tis the gradient. Model optimization also includes using the Bayesian optimization method to search for the optimal combination of hyperparameters, and using the feature importance scores of random forests to screen key features. The formula is:
[0139] Importance(j)=∑(Gini(D)-Gini(D left )-Gini(D right ))
[0140] where Gini(D) is the impurity of the node, and Gini(D left ) and Gini(D right ) are the impurities of the sub-nodes after splitting.
[0141] The diagnostic decision logic unit determines the diagnostic result by calculating the malignancy probability P malignant =∑(w i ·P i ), where P i is the malignancy probability of the i-th base learner, and w i is the weight of the base learner. Set the threshold T (0.5) according to clinical needs. If P malignant ≥T, it is diagnosed as malignant; if P malignant <T, it is diagnosed as benign. The confidence of the diagnostic result is calculated by the formula Confidence=1-|P malignant -T|. The closer Confidence is to 1, the higher the credibility of the diagnostic result.
[0142] The model evaluation and validation unit evaluates the model performance using indicators such as AUC, sensitivity, and specificity. The sensitivity formula is Sensitivity=TP / (TP+FN), and the specificity formula is Specificity=TN / (TN+FP), where TP is the true positive, FN is the false negative, TN is the true negative, and FP is the false positive. Adopt K-fold cross-validation to ensure the generalization ability of the model. Randomly divide the dataset into K subsets, sequentially select one subset as the validation set, and the remaining K-1 subsets as the training set, repeat K times, and calculate the average value and standard deviation of the performance indicators. In addition, use an independent external dataset for validation to ensure the reliability of the model in practical applications.
[0143] The post-processing and decision optimization unit processes the prediction probability using the probability smoothing method. The formula is P smooth =(1 / N)·∑(P malignant(tk:t+k)), where N is the window size and k is the window radius. A diagnostic consistency check is also performed; if multiple consecutive inconsistent diagnostic results are found, a re-diagnosis is triggered. Decision optimization involves combining multi-expert consensus and dynamic threshold adjustment. The model output is compared with the diagnostic results of multiple experts, and the final diagnosis is determined by majority voting. The diagnostic threshold T is then adjusted based on clinical needs. In screening scenarios, the threshold is lowered to increase sensitivity, while in diagnosis scenarios, the threshold is raised to improve specificity.
[0144] In this embodiment, the dynamic optimization and visualization module includes:
[0145] The core function of the dynamic optimization module is to monitor model performance in real time and dynamically adjust model parameters and decision thresholds based on feedback data. It consists of three main components: performance evaluation, parameter optimization, and decision adjustment. First, performance evaluation quantifies model performance by calculating metrics such as AUC, sensitivity, and specificity in real time. AUC (Area Under Curve) is a key metric for evaluating the overall performance of a classifier. Its calculation formula is AUC = ∫(TPR(FPR))d(FPR), where TPR (True Positive Rate) is TPR = TP / (TP + FN), and FPR (False Positive Rate) is FPR = FP / (FP + TN). Sensitivity is used to assess the model's ability to identify malignant samples. The formula is Sensitivity = TP / (TP + FN), where TP represents true positives and FN represents false negatives. Specificity is used to assess the model's ability to identify benign samples. The formula is Specificity = TN / (TN + FP), where TN represents true negatives and FP represents false positives. By calculating these indicators in real time, the dynamic optimization module can comprehensively evaluate the current performance of the model and provide a basis for parameter adjustment.
[0146] In terms of parameter optimization, the dynamic optimization module adopts an advanced tuning method based on Bayesian optimization. The goal of Bayesian optimization is to predict the distribution of the objective function by constructing a Gaussian process model, so as to select the optimal parameter combination. Its implementation steps are as follows: The first step is to define the parameter search space, the kernel width parameter γ of SVM and the regularization parameter C. The second step is to construct a Gaussian process model to model the objective function. The formula is f(x)~GP(μ(x), k(x, x')), where μ(x) is the mean function and k(x, x') is the covariance function. The third step is to use the acquisition function (such as the expected improvement EI) to select the next sampling point. The formula is EI(x)=E[max(f(x)-f(x + ), 0)], where f(x +) is the current optimal value. In the fourth step, update the Gaussian process model and repeat the above steps until the optimal parameter combination is found. The Bayesian optimization method can find the optimal parameters within fewer iterations, thus significantly improving the model performance.
[0147] The decision adjustment module dynamically adjusts the diagnostic threshold T according to the clinical needs and changes in data distribution. The diagnostic threshold T is the key value for judging whether a sample is malignant or benign, and its adjustment logic is based on the ROC curve and clinical needs. The adjustment steps are as follows:
[0148] In the first step, calculate the ROC curve of the current model and determine the sensitivity and specificity at different thresholds. In the second step, select the optimal threshold according to the clinical needs, select a high-sensitivity threshold in the screening scenario, and select a high-specificity threshold in the confirmation scenario. In the third step, update the threshold T in real time and synchronously adjust the diagnostic logic. The formula is:
[0149] ifP malignant ≥T, diagnose as malignant;
[0150] ifP malignant <T, diagnose as benign;
[0151] By dynamically adjusting the threshold, the module can flexibly adapt to different clinical scenarios and maximize the diagnostic effect.
[0152] The visualization module uses advanced graphics and interactive technologies to intuitively display the model output, diagnostic results, and performance metrics to clinicians to assist them in making decisions. Visualizing the diagnostic results is one of the core functions of the visualization module, and its goal is to display the malignancy probability of the sample in the form of a heat map and a probability distribution map. The steps are as follows:
[0153] In the first step, generate a heat map of the malignancy probability of the sample, using a color gradient to represent the probability size. The formula is: Color = f(P malignant ), where f is the color mapping function.
[0154] In the second step, draw a probability distribution map to show the probability distribution of different classes. The formula is P benign = 1 - P malignant . Through these graphs, doctors can intuitively understand the credibility of the diagnostic results.
[0155] The model performance visualization unit comprehensively displays the performance of the model by drawing the ROC curve and the confusion matrix. The steps are as follows:
[0156] In the first step, draw the ROC curve, with the false positive rate FPR on the horizontal axis and the true positive rate TPR on the vertical axis. The formulas are FPR = FP / (FP + TN) and TPR = TP / (TP + FN).
[0157] The second step is to plot a confusion matrix, which displays the detailed classification results, including TP, FP, TN, and FN. These graphs allow doctors to comprehensively evaluate the model's performance and identify potential problems.
[0158] The feature importance visualization unit shows the contribution of each modality feature to the diagnosis result, helping doctors understand the decision logic of the model. The steps are as follows:
[0159] The first step is to calculate the feature importance score, using the random forest feature importance formula Importance(j) = ∑(Gini(D)-Gini(D left )-Gini(D right )), where Gini(D) is the impurity of the node, Gini(D left ) and Gini(D right ) is the impurity of the child node after segmentation.
[0160] The second step is to draw a bar chart or radar chart to show the importance of each feature. The formula is Importance(j)=f(feature j ), where f is a graph mapping function. Through these graphs, doctors can understand which features have the greatest impact on diagnostic results, thereby optimizing data collection and model design.
[0161] The user interaction logic unit provides doctors with a flexible operating interface that supports parameter adjustment and diagnostic feedback. The manual parameter adjustment function allows doctors to manually adjust model parameters, the SVM kernel width parameter γ and regularization parameter C, and the decision threshold T through the interactive interface. The adjusted parameters will update the model in real time and display the new diagnostic results and performance indicators. The diagnostic result feedback function allows doctors to provide feedback on the diagnostic results, mark misclassified samples, or adjust the diagnostic category. The feedback data will be used to further optimize the model. The formula is:
[0162] P malignant =f(feedback, P malignant )
[0163] Where f is the feedback adjustment function. The multi-view linkage function supports simultaneous updating of multiple views. Click a point on the ROC curve to update the confusion matrix and diagnosis result chart in real time. The formula is View = f(selected point ), where f is the view update function.
[0164] Example 2
[0165] Clinical data. A total of 100 patients with breast lumps admitted to our hospital were included. All patients underwent biopsy or postoperative pathological examination to obtain clear clinical pathological test results. The patients ranged in age from 26 to 70 years old, with an average age of (46.25±4.14) years. All patients underwent ultrasound, mammography, and serum P16 autoantibody testing before intervention. The results were compared with the pathological test results to clarify the sensitivity, specificity, and accuracy of the three diagnostic schemes combined with ultrasound combined with mammography. There was no statistically significant difference in the baseline data of the patients in the above groups, and they were comparable (P>0.05). All patients came to the hospital for breast lumps, underwent three diagnostic tests within 1 month, and obtained pathological test results. At the same time, the patients gave informed consent and signed an informed consent form for this study. Patients with mental illness were excluded.
[0166] Methods: All 100 patients underwent ultrasound, mammography, and serum P16 autoantibody testing, as follows:
[0167] Ultrasound examination: A Siemens Acuson S2000 ultrasound diagnostic system was used, with the probe frequency set between 9 and 13 MHz. The patient was placed in a supine position with both arms raised and slightly abducted to fully expose the breasts and axillae. A comprehensive radial scan of the breasts was performed, centered around the nipples.
[0168] Mammography: A fully digital mammography machine, model Senographe 2000D, manufactured by General Electric, USA, was used. The patient stood upright in front of the radiographing table, with the breast placed between the compression device and the table. Pressure was set to 7 to 10 N. The anode target plane was switched and X-ray exposure conditions were automatically selected. Conventional head-to-foot and lateral oblique views were taken, with bilateral comparison. Additional positions and localized compression and / or magnification were used as needed.
[0169] Serum P16 autoantibody testing: Fasting blood is collected from patients in the morning and centrifuged at 3000 rpm. The supernatant is stored at -20°C for testing. Serum P16 immunoglobulin (IgG) levels are measured using an ELISA. A linear antigen peptide is designed based on human leukocyte antigen type II, corresponding to the P16 protein and carrying an overlapping restricted epitope recognized by >90% of patients. A corn polypeptide antigen, which differs from the human antigen epitope, is used as a reference antigen for testing.
[0170] By comparing with the results of pathological examination, the sensitivity, specificity and accuracy of the combined application of the three diagnostic schemes and ultrasound combined with mammography were clarified.
[0171] Observation indicators
[0172] Sensitivity = number of true positive cases / (number of true positive cases + number of false negative cases) × 100.0%;
[0173] Specificity = number of true negative cases / (number of true negative cases + number of false positive cases) × 100.0%;
[0174] Accuracy = (number of true positive cases + number of true negative cases) / total number of cases × 100.0%.
[0175] To clarify the sensitivity, specificity, and accuracy of the diagnostic scheme of ultrasound and mammography combined with serum P16 autoantibody and ultrasound combined with mammography.
[0176] Statistical analysis: All data were processed using SPSS 23.0 statistical software. Measurement data were expressed as mean ± standard deviation (x ± s) and analyzed using the t test. Enumeration data were expressed as rate (%) and analyzed using the χ2 test. P < 0.05 indicated statistically significant differences between the groups.
[0177] Here are the results:
[0178]
[0179] In the dense breast (ACR density category C / D) subgroup (n=35), the combined diagnostic sensitivity reached 94.1% (16 / 17), significantly higher than the 85.7% (12 / 14) achieved by dual-modality. The dynamic weighting strategy increased the mammography weight to 0.45 in this scenario, contributing 32% to the calcification density feature and effectively identifying two cases of microcalcified ductal carcinoma that had been missed by ultrasound.
[0180] In patients with highly vascular masses (Adler grade ≥2, n=28), the combined diagnostic specificity reached 93.3% (14 / 15), superior to the dual-modality approach of 80.0% (12 / 15). The ultrasound hemodynamic feature weight was increased to 0.38, and three false-positive cases due to inflammatory pseudotumors were accurately excluded using a combined TAI×RI score (≥6.5). This study confirmed that the trimodal combined diagnostic approach achieved anatomical-functional-molecular complementarity across modalities through a dynamic weighting strategy, significantly outperforming traditional methods in both sensitivity and specificity (p<0.05). It has unique diagnostic value for early-stage triple-negative breast cancer (high P16 expression + microcalcifications) and occult lesions in dense breast tissue, providing more reliable decision support for clinical practice.
[0181] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. An intelligent diagnosis system for breast lumps based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection, characterized in that: include: The data acquisition module is responsible for the standardized collection of ultrasound, mammography, and serum P16 protein data to ensure the compatibility and consistency of multi-source data. It is the basic input layer of the system. The feature extraction and fusion module uses deep learning and radiomics technology to achieve refined cross-modal feature extraction and dynamic weight fusion; Intelligent diagnostic decision-making module, which achieves accurate classification through a hierarchical decision-making model and embeds clinical logic reasoning rules; Dynamic optimization and visualization modules enable system self-iteration and transparency of the diagnostic process.
2. The intelligent breast mass diagnosis system based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection according to claim 1, characterized in that: The data acquisition module integrates ultrasound, mammography, and serum P16 protein detection data, using standardized protocols to ensure data consistency. Ultrasound imaging uses a high-frequency probe to acquire two-dimensional grayscale images and hemodynamic parameters, and an improved motion artifact suppression algorithm is used to enhance image quality. Mammography uses adaptive histogram equalization to enhance microcalcification contrast and a deep denoising network to eliminate quantum noise and quantify calcification density and structural distortion characteristics. Serum P16 detection uses electrochemiluminescence, calibrated with β-actin as the internal reference, and a positive threshold of ≥2.5 ng / mL is set to ensure that detection sensitivity is synchronized with the temporal and spatial characteristics of the imaging data.
3. The intelligent breast mass diagnosis system based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection according to claim 1, characterized in that: The feature extraction and fusion module extracts cross-modal features based on deep learning and radiomics technology: Ultrasound images were segmented using the U-Net++ network to calculate morphological indicators and hemodynamic parameters. Mammography images were analyzed using a multi-scale feature pyramid network to analyze the heterogeneity of calcification clusters and the depth of structural distortion. A logistic regression model was constructed by correlating serum P16 concentration with clinical parameters. Feature fusion adopted a dynamic weighting strategy: a feature association graph was constructed through a graph neural network, and the weights were dynamically adjusted in combination with the Bayesian probability integration formula to achieve the complementarity of anatomical, functional, and molecular multidimensional information.
4. The intelligent breast mass diagnosis system based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection according to claim 1, characterized in that: The intelligent diagnostic decision-making module uses a cascade classifier to improve diagnostic efficiency: the first-level LightGBM model quickly screens high-risk cases; the second-level multi-task Transformer network simultaneously outputs BI-RADS grading and molecular subtype prediction, and introduces FocalLoss to solve the data imbalance problem; the embedded clinical rule library initiates Monte Carlo Dropout simulation for controversial cases to reduce the risk of misdiagnosis.
5. The intelligent breast mass diagnosis system based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection according to claim 1, characterized in that: Dynamic optimization and visualization module, which uses interpretability technology to visualize feature contributions and generate heat maps superimposed on imagery to mark high-risk areas; It supports interactive weight adjustment of 3D fusion views; adopts a federated learning framework to achieve model iteration under multi-center data privacy protection, and adds Laplace noise through a differential privacy mechanism; the negative feedback closed loop automatically triggers feature comparison and online learning of misdiagnosed cases, and updates model parameters within 24 hours to ensure continuous evolution of the system.
6. The intelligent breast mass diagnosis system based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection according to claim 2, wherein the data acquisition module includes an ultrasound image acquisition unit, a mammography image acquisition unit, and a serum P16 protein detection unit, and specifically includes: Ultrasound image acquisition unit, A high-frequency probe is used to obtain a two-dimensional grayscale image of the breast mass and generate a pixel matrix I(x, y); Image acquisition parameters include gain, depth, and focus, and parameter settings are dynamically adjusted using the following formula: Gain:G=G min +(G max -G min )·(1-e -α·T ) Among them, G min and G max are the minimum and maximum gain values, α is the attenuation coefficient, and T is the tissue penetration depth; Depth:D=D base +(1-e -β·S ) Among them, D base is the base depth, β is the depth adjustment factor, and S is the mass size; Focus:F=F center +ΔF·sin(ω·t) Among them, F center is the center position of the mass, ΔF is the focus shift amplitude, ω is the frequency, and t is the time; Gain: Dynamically adjusted according to tissue density, ranging from -20dB to +20dB; Depth: set to 4cm to 8cm according to the location of the mass; Focus: Set to the center of the mass to ensure clarity in the focus area; In this embodiment, color Doppler blood flow imaging is used to obtain blood flow information and generate a blood flow velocity matrix V(x, y); The calculation formula of resistance index RI is: RI=(V max -V min ) / V max Among them, V max and V min are the peak systolic velocity and the minimum diastolic velocity, respectively; An improved motion artifact suppression algorithm is used, combined with Kalman filtering and adaptive motion compensation technology to eliminate artifacts caused by patient breathing or movement; The algorithm formula is: I corrected (x,y)=I(x,y)-K·(I(x,y)-I prev (x,y)) Where K is the Kalman gain coefficient, I prev (x, y) is the pixel value of the previous frame; The calculation formula of Kalman gain coefficient K is: K=P prev / (P prev +R) Among them, P prev is the prediction error of the previous frame, R is the observation noise variance; The motion artifact suppression formula dynamically corrects pixel values through Kalman filtering to eliminate artifacts; The Kalman gain coefficient formula dynamically adjusts the correction weight through the ratio of prediction error to observation noise; Mammography image acquisition unit, Adaptive histogram equalization is performed on mammography images to enhance the contrast of microcalcifications; The CLAHE algorithm formula is: I enhanced (x,y)=CLAHE(I(x,y),clip limit ,tile size ) Among them, clip limit is the contrast limit parameter, tile size is the size of the local area; clip limit : Set to 2.0 to prevent over-enhancement of local areas; tile size : Set to 8×8 pixels to ensure the equalization effect of local areas; clip limit Dynamically adjusted by the following formula: clip limit =clip base ·(1+γ·(1-e -δ·D )) Among them, clip base is the basic limit value, γ and δ are adjustment factors, and D is the calcification density; The CLAHE formula enhances the contrast of microcalcifications by performing local region histogram equalization; clip limit The formula dynamically adjusts the contrast limit value through an exponential function to prevent over-enhancement; Use DnCNN to eliminate quantum noise and improve the image signal-to-noise ratio (SNR≥30dB); The network output formula is: I denoised (x,y)=DnCNN(I(x,y),θ) Among them, θ is the network parameter; During network training, the mean square error (MSE) is used as the loss function: MSE=1 / N·∑(I denoised (x,y)-I clean (x,y)) 2 Where N is the total number of pixels, I clean (x, y) is a noise-free image; The DnCNN formula removes quantum noise through a deep learning network to improve image quality; The MSE formula optimizes the network parameters by calculating the difference between the denoised image and the noise-free image; The calculation formula for calcification density is: Density=N calc / Area Among them, N calc is the number of calcification points, Area is the area of the region of interest; The structural distortion depth is extracted by edge detection algorithm, and its maximum depth D is calculated max ; The calcification density formula quantifies the richness of calcification by the ratio of the number of calcification points to the regional area; the structural distortion depth formula uses an edge detection algorithm to extract the edge of the mass and calculate the maximum depth; Serum P16 protein detection unit The ECLIA technique was used to detect the level of P16 protein in serum and generate the concentration value C P16 ; The detection formula is: C P16 =k·(I sample -I blank ) / (I standard -I blank ) Where k is the calibration coefficient, I sample is the light intensity signal of the sample, I blank is the light intensity signal of the blank control, I standard is the light intensity signal of the standard; The ECLIA formula quantifies the P16 protein concentration by the ratio of light intensity signals; β-actin is used as an internal reference to calibrate the test results to ensure data stability and accuracy; The calibration formula is: C calibrated =C P16 / C β-actin Among them, C β-actin is the concentration value of β-actin; the internal reference calibration formula eliminates the systematic error by the ratio of P16 protein concentration to the internal reference concentration; The positive threshold is set at ≥2.5 ng / mL to ensure detection sensitivity; the positive threshold formula is determined based on experimental data to ensure the reliability of the test results.
7. The intelligent breast mass diagnosis system based on ultrasound-mammography multimodal imaging combined with serum P16 protein detection according to claim 3, wherein the feature extraction and fusion module specifically comprises: The ultrasound image feature extraction unit uses the Canny edge detection algorithm to extract the edge contour of the mass. The Canny edge detection algorithm removes noise through Gaussian filtering, calculates the image gradient, and uses a double threshold method to retain strong edges. Then, calculate the edge perimeter P and the area A; The formula for edge irregularity is: EI = P / (2·π·sqrt(A / π)) The larger the EI value, the more irregular the edge of the mass, indicating a higher possibility of malignancy. In the extraction of internal echo features, echo uniformity and echo intensity are two important indicators. Echo uniformity is calculated by calculating the image grayscale mean μ and variance σ. 2 Get; gray mean μ represents the average brightness of the image, variance σ 2 Indicates the discrete degree of pixel value; the echo uniformity formula is: EH=1-(σ 2 / m 2 ) The smaller the EH value, the more uneven the echo, suggesting that the internal structure of the mass is complex and may be related to malignant masses. The echo intensity is directly obtained by calculating the grayscale mean μ, and the formula is: EI=μ In the extraction of hemodynamic characteristics, the resistance index and Adler grade are two key indicators. A larger RI value indicates higher blood flow resistance, indicating a higher possibility of malignant tumors. Spectral analysis is achieved through fast Fourier transform to extract the frequency domain characteristics of blood flow velocity. The FFT formula is: F(k)=∑(f(n)·e -2πikn / N ) Among them, f(n) is the time domain signal, F(k) is the frequency domain signal, and k is the frequency index; Adler classification is calculated by counting the number of blood vessels N in the tumor. vessel And weighted calculation is obtained, the formula is: Adler=∑(w i ·N i ) Among them, w i =[0, 0.3, 0.5, 0.7], N i is the number of blood vessels; The Adler grading formula quantifies the richness of blood flow signals by weighting the coefficient and the number of blood vessels. The Adler grading system is divided into 0-3 grades based on the richness of blood flow signals. The grading standards are as follows: Grade 0: no blood flow signal; Grade 1: small amount of blood flow signal, 1-2 blood vessels; Grade 2: moderate blood flow signal, 3-4 blood vessels; Level 3: rich blood flow signals, ≥5 blood vessels; The higher the Adler grade, the richer the blood vessels in the tumor, indicating a higher possibility of malignancy. Mammography image feature extraction unit: Mammography provides calcification and structural information of breast tumors, which is an important basis for tumor detection. In calcification feature extraction, calcification density and calcification aggregation are two key indicators. Calcification density is calculated by counting the number of calcification points N through the local peak detection algorithm. calc , and calculate the calcification density per unit area using the formula: CD=N calc / IN The larger the CD value, the denser the calcification, indicating a higher possibility of malignant tumors. Local peak detection is achieved by the sliding window method, calculating the local maximum value of the gray value within the window. The calcification concentration is calculated by calculating the average distance D between calcification points. ij The formula is: CC=1 / (mean(D ij )) The larger the CC value, the higher the calcification concentration, indicating a higher possibility of malignant tumors. In structural feature extraction, structural distortion depth and texture complexity are two key indicators. The structural distortion depth is extracted using the Canny edge detection algorithm to extract the maximum depth, and the formula is: SDD = max(edge(I(x, y))) The larger the SDD value, the more severe the structural distortion, indicating a higher possibility of malignant tumors. The texture complexity is calculated by the gray-level co-occurrence matrix entropy, and the formula is: TC = -∑(p(i, j)·log(p(i, j))) The larger the TC value, the more complex the texture, indicating a higher possibility of malignant tumor; Serum P16 protein feature extraction unit: Serum P16 protein detection provides molecular biological information and is an important basis for tumor detection. P16 concentration is determined by electrochemiluminescence. The higher the P16 concentration, the higher the possibility of malignant tumor. Feature fusion unit, the purpose of feature fusion is to integrate multimodal features into a unified high-dimensional feature space; first, all features are normalized so that their value range is unified to [0, 1], the formula is: X normalized =(X-X min ) / (X max -X min ) Then, different weights are assigned according to the discriminability of the features, and the feature weight w is calculated using the LASSO regression algorithm. i , the formula is: min(||Y-Xβ|| 2 +λ||β||1) Among them, Y is the label, X is the feature matrix, β is the regression coefficient, and λ is the regularization parameter. The features corresponding to the non-zero coefficients are screened out through LASSO regression, and the weighted fusion features are calculated: X fused =∑(w i ·X i ) Finally, kernel principal component analysis (KPCA) is used to map the fused features into a high-dimensional space using the Gaussian kernel function: K(x,y)=exp(-||xy|| 2 / (2σ 2 )) Where σ is the kernel width parameter; the principal components are extracted by KPCA to retain the most important feature information; Feature optimization unit, the purpose of feature optimization is to remove redundant features and retain the most discriminative feature combination; first, the LASSO regression algorithm is used to perform sparse processing on the features and filter out the features corresponding to non-zero coefficients; then, the genetic algorithm is used to optimize the feature combination, and the objective function is: Fitness=AUC+α·(1-Feature Ratio) Among them, AUC is the classification performance, Feature Ratio is the ratio of the number of features to the total number of features, and α is the balance parameter; the optimal feature combination is iteratively searched through the genetic algorithm.
Citation Information
Cited By
Emergency alarm system for power supply safety protection
CN121053764A