A Multimodal Fusion-Based Auxiliary Diagnostic System for Early Ovarian Cancer
By constructing a multimodal fusion-based early ovarian cancer auxiliary diagnostic system, and utilizing imaging and non-imaging feature extraction modules and attention mechanisms, the limitations of single-modal diagnosis are overcome, enabling accurate and interpretable risk assessment and optimized clinical decision-making for early ovarian cancer.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE THIRD AFFILIATED HOSPITAL OF XINJIANG MEDICAL UNIV
- Filing Date
- 2026-04-01
- Publication Date
- 2026-06-30
AI Technical Summary
Existing single-modality methods for ovarian cancer diagnosis have limitations. Serum markers such as CA125 and HE4 lack effective integration, and ultrasound imaging systems such as O-RADS have limited ability to differentiate early-stage ovarian cancer. Furthermore, the integration of multimodal information relies on physicians' subjective experience and lacks standardized strategies, thus failing to meet the needs for accurate diagnosis of early-stage ovarian cancer.
A multimodal fusion-based auxiliary diagnostic system for early ovarian cancer was constructed. Ultrasound images and serum biomarkers and clinical variables were processed by an image feature extraction module and a non-image feature extraction module, respectively. Attention mechanism was used for cross-modal fusion, feature weights were dynamically calculated, and finally, a risk probability value was output through a classifier.
It enables the collaborative use of multi-source data, overcomes the limitations of single-modality diagnosis, provides quantitative risk assessment basis, enhances the accuracy and interpretability of diagnosis, optimizes clinical decision-making, and reduces unnecessary interventions.
Smart Images

Figure CN122314342A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of early ovarian cancer auxiliary diagnosis technology, specifically to an early ovarian cancer auxiliary diagnosis system based on multimodal fusion. Background Technology
[0002] Ovarian cancer is the leading cause of death among malignant tumors of the female reproductive system, and early diagnosis is crucial for improving patient prognosis. Currently, commonly used clinical diagnostic methods mainly include serum tumor marker testing and ultrasound imaging assessment. Among serum markers, cancer antigen 125 (CA125) and human epididymal protein 4 (HE4) are the two most widely used indicators. However, CA125 has limited sensitivity in early ovarian cancer and is easily interfered with by benign lesions, leading to insufficient specificity. While HE4 has high specificity, its sensitivity is relatively low; the two exhibit a complementary "sensitivity-specificity" characteristic. In terms of ultrasound imaging, the Ovarian-Adnexal Reporting and Data System (O-RADS) launched by the American College of Radiology significantly improves the consistency and reproducibility of ultrasound diagnosis through a standardized ultrasound dictionary and five-level risk stratification. The area under the diagnostic curve for this system in the general population reaches 0.956. However, the aforementioned single-modality diagnostic methods all have inherent limitations, making it difficult to overcome the "ceiling effect" of accurate diagnosis of early ovarian cancer.
[0003] Specifically, the shortcomings of existing technologies are:
[0004] First, the diagnostic efficacy of a single serum marker or a single ultrasound modality has limitations. Although CA125 and HE4 are complementary, they lack an effective integration mechanism. The O-RADS system has limited ability to differentiate early ovarian cancer, especially for type 4 lesions, and the problem of insufficient specificity still exists.
[0005] Second, in clinical practice, the integration of multi-source information such as serological indicators, ultrasound imaging features and clinical variables relies on the subjective experience of physicians and lacks standardized and automated multimodal fusion strategies, making it difficult to fully leverage the synergistic and complementary value of multi-source data.
[0006] Third, existing diagnostic models are mostly based on advanced ovarian cancer cohorts. There is still a lack of intelligent diagnostic systems that target early lesions and integrate multimodal information such as imaging, serological and clinical variables, which cannot meet the clinical needs for accurate diagnosis of early ovarian cancer.
[0007] Therefore, there is an urgent need for an intelligent diagnostic system that can dynamically integrate multimodal data and automatically learn cross-modal correlation features to overcome the performance limitations of existing single diagnostic methods. Summary of the Invention
[0008] To address the technical problems in related technologies, this invention provides an early ovarian cancer auxiliary diagnostic system based on multimodal fusion.
[0009] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0010] A multimodal fusion-based auxiliary diagnostic system for early ovarian cancer, comprising:
[0011] The data acquisition module is used to acquire the original ultrasound images, serum marker test values, and clinical variable data of the subject to be diagnosed.
[0012] The image feature extraction module, connected to the data acquisition module, is used to preprocess the original ultrasound image and extract image depth features through a pre-trained image feature extraction network.
[0013] A non-image feature extraction module, connected to the data acquisition module, is used to fuse the serum biomarker detection values and the clinical variable data, and extract non-image features through a non-image feature coding network;
[0014] A multimodal fusion module is connected to the image feature extraction module and the non-image feature extraction module respectively. It is used to dynamically calculate the attention weights of the image depth features and the non-image features through the attention mechanism-based cross-modal fusion module, and generate fused features based on the attention weights.
[0015] The diagnostic output module, connected to the multimodal fusion module, is used to process the fused features through a classifier and output the risk probability value of the subject to be diagnosed having early ovarian cancer.
[0016] Optionally, the image feature extraction module includes:
[0017] An image preprocessing unit is used to adjust the original ultrasound image to a preset resolution and perform pixel value normalization processing.
[0018] The feature extraction unit has a built-in convolutional neural network, which is used to receive the preprocessed ultrasound image and output the feature vector generated by the global average pooling layer in the convolutional neural network as the image depth feature.
[0019] Optionally, the non-image feature extraction module includes:
[0020] The data stitching unit is used to stitch together the serum biomarker detection values and the clinical variable data to form a non-image input vector;
[0021] The feature encoding unit has a built-in fully connected network, which sequentially includes a first fully connected layer, a first batch normalization layer, a first activation function layer, a second fully connected layer, a second batch normalization layer, a second activation function layer, a third fully connected layer, a third batch normalization layer, and a third activation function layer. It is used to process the non-image input vector layer by layer and output the non-image features.
[0022] Optionally, the multimodal fusion module includes:
[0023] Linear mapping unit, used to transform the first linear transformation matrix The image depth features Mapping to the target dimension space yields the mapped image features. And through the second linear transformation matrix The non-image features Mapping to the target dimension space yields the mapped non-image features. ,in , ;
[0024] The attention calculation unit, connected to the linear mapping unit, is used to calculate attention weights. , In the formula, Let be the dimension value of the target dimension space. This is matrix multiplication, where T denotes matrix transpose;
[0025] The feature fusion unit, connected to the attention calculation unit, is used to generate fused features. .
[0026] Optionally, the multimodal fusion-based early ovarian cancer auxiliary diagnostic system further includes:
[0027] An interpretability analysis module, connected to the image feature extraction module and the diagnostic output module, is used to generate a heatmap corresponding to the original ultrasound image through a gradient-weighted class activation mapping algorithm, and to calculate the marginal contribution of the serum biomarker detection values and each feature in the clinical variable data to the risk probability value through the SHAP analysis method.
[0028] Optionally, the multimodal fusion-based early ovarian cancer auxiliary diagnostic system further includes:
[0029] The risk restratification module, connected to the data acquisition module, is used to acquire the ovarian-adnexal report and data system classification of the subject to be diagnosed. When the classification is O-RADS Category 4, the malignancy risk value is calculated using the simple rule risk model of the International Ovarian Tumor Analysis Group. When the malignancy risk value is less than a preset threshold, the risk level of the subject to be diagnosed is adjusted from O-RADS Category 4 to O-RADS Category 3.
[0030] Optionally, the multimodal fusion-based early ovarian cancer auxiliary diagnostic system further includes:
[0031] The model calibration evaluation module, connected to the diagnostic output module, is used to evaluate the calibration degree of the risk probability value through the Hosmer-Lemeshow goodness-of-fit test, and to calculate the clinical net benefit under different threshold probabilities through decision curve analysis, so as to determine the optimal threshold probability.
[0032] The formula for calculating the net clinical benefit is as follows:
[0033]
[0034] In the formula, TP represents the number of true positives, FP represents the number of false negatives, N represents the total sample size, and pt represents the threshold probability.
[0035] Optionally, the image feature extraction module further includes:
[0036] The model compression unit is used to perform lightweight processing on the pre-trained image feature extraction network to generate a compressed model. The compressed model has a smaller size than the original model and a higher inference speed.
[0037] Optionally, the image feature extraction module further includes:
[0038] The knowledge distillation unit is used to perform knowledge distillation using the pre-trained image feature extraction network as the teacher network and a convolutional neural network with fewer parameters as the student network to generate a lightweight student model. The inference latency of the lightweight student model is lower than that of the teacher network.
[0039] Optionally, the multimodal fusion-based early ovarian cancer auxiliary diagnostic system further includes:
[0040] The system integration interface is used to embed the image feature extraction module, the non-image feature extraction module, the multimodal fusion module, and the diagnostic output module into the hospital's image archiving and communication system. The system integration interface supports medical digital imaging and communication standards and is used to receive the original ultrasound image and output the risk probability value and the corresponding heat map.
[0041] Beneficial effects:
[0042] 1. Through the above technical solution, firstly, the system of the present invention processes three types of data through an image feature extraction module and a non-image feature extraction module respectively (wherein, ultrasound images provide morphological information, serum markers reflect the biological activity of tumors, and clinical variables reflect host factors), thereby enabling unified representation and collaborative utilization of multi-source heterogeneous data in the same system, providing a structured data foundation for subsequent cross-modal fusion.
[0043] Secondly, in this invention, the image feature extraction module preprocesses the original ultrasound image and extracts image depth features through a pre-trained image feature extraction network. This enables automated extraction of high-level semantic features from the ultrasound image. Simultaneously, the non-image feature extraction module fuses serum biomarker detection values and clinical variable data, and extracts non-image features through a non-image feature encoding network. Through layer-by-layer mapping of the encoding network, these heterogeneous structured data (serum biomarkers and clinical variables are structured data with varying numerical ranges and distribution characteristics) are uniformly encoded into feature vectors, achieving alignment with non-image data and creating conditions for cross-modal fusion. More importantly, this invention inputs image depth features and non-image features into an attention-based cross-modal fusion module, dynamically calculating the attention weights of the two modalities and generating fused features based on these weights. Thus, firstly, the attention mechanism enables the system to adaptively adjust the contribution ratio of image and non-image modalities in the fused features according to the specific circumstances of each diagnostic subject. For cases with typical ultrasound imaging features but atypical serum biomarkers, the system can assign higher weights to imaging features; for cases with atypical ultrasound findings but abnormal serological markers, the system can increase the weights of non-imaging features. This dynamic fusion strategy overcomes the limitations of traditional static fusion (such as feature splicing and weighted averaging) which uses fixed fusion weights for all samples. Secondly, cross-modal information complementarity is achieved. By calculating the correlation strength between imaging features and non-imaging features through an attention mechanism, the system can identify and strengthen the synergistic information between the two modalities, ensuring that the fused features retain both the morphological information of the images and integrate the biological activity information of the serological markers, achieving a complementary gain of "1+1>2".
[0044] Third, in this invention, the diagnostic output module processes the fused features through a classifier and outputs a risk probability value. In other words, this invention transforms the multimodal fusion result into a continuous risk probability value (typically between 0 and 1), providing a quantitative risk assessment basis for clinical decision-making. Compared to the five-level discrete classification (2, 3, 4, and 5 classes) of the traditional O-RADS system, the continuous probability value can more precisely distinguish the malignancy of different lesions within the same risk level, providing clinicians with flexible threshold adjustment space in different decision-making scenarios (such as selecting a low threshold in screening scenarios and a high threshold in diagnostic scenarios).
[0045] 2. Other beneficial effects or advantages of the present invention will be described in detail in conjunction with specific structures in specific embodiments. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In addition, it should be understood that the proportional relationship of each component in the drawings of this specification does not represent the proportional relationship in the actual material selection and design, but is only a schematic diagram of the structure or position, wherein:
[0047] Figure 1 This is a schematic diagram of the overall process provided by an exemplary embodiment of the present invention;
[0048] Figure 2 This is a schematic diagram of patient enrollment and process corresponding to an exemplary embodiment of the present invention;
[0049] Figure 3 This is a schematic diagram of decision curve analysis provided by an exemplary embodiment of the present invention. Detailed Implementation
[0050] To facilitate a clearer and more accurate understanding of the technical solutions of this invention by those skilled in the art, the following provides a more detailed description of the existing related technologies.
[0051] Ovarian cancer is the deadliest malignant tumor of the female reproductive system, and its prognosis is closely related to the clinical stage at diagnosis. The five-year survival rate for FIGO stage I patients is over 90%, while the five-year survival rate for stage III and IV patients is less than 30%. However, because the ovaries are located deep in the pelvic cavity, early lesions lack specific symptoms, and more than 70% of patients are diagnosed at an advanced stage. Therefore, improving the accuracy of early diagnosis of ovarian cancer is crucial to improving patient prognosis.
[0052] Currently, the commonly used auxiliary diagnostic methods for early ovarian cancer in clinical practice mainly include two categories: serum tumor marker detection and ultrasound imaging evaluation.
[0053] Among them, serum tumor marker detection is one of the most widely used methods for ovarian cancer screening, with cancer antigen 125 (CA125) and human epididymal protein 4 (HE4) being the most commonly used.
[0054] Since its introduction into clinical practice in 1981, CA125 has been a core indicator for the diagnosis of ovarian cancer. In clinical practice, a serum CA125 level exceeding 35 U / mL is generally considered positive, suggesting a possible malignancy. However, this indicator has significant limitations. For example, a 42-year-old premenopausal woman presented with a pelvic mass, and her serum CA125 level was 52 U / mL (above the upper limit of normal). Based on this result, the clinician might highly suspect malignancy and recommend surgical exploration. However, postoperative pathology revealed that the patient actually had an ovarian endometriotic cyst, a common benign lesion. In this case, the false-positive CA125 result directly led to unnecessary surgical intervention, causing surgical trauma and psychological burden for the patient. This phenomenon is particularly prominent in premenopausal women because CA125 can be elevated in various benign gynecological diseases such as endometriosis, pelvic inflammatory disease, and uterine fibroids, resulting in insufficient specificity.
[0055] HE4 is a novel biomarker for ovarian cancer that has garnered significant attention in recent years. Studies show that HE4 is almost not expressed in normal ovarian tissue, but is highly expressed in ovarian cancer tissue, and is less affected by benign lesions, thus exhibiting high specificity. For example, a 55-year-old postmenopausal woman was found to have an adnexal mass during a physical examination. Her serum HE4 level was 145 pmol / L (above the upper limit of normal for postmenopausal women, 140 pmol / L), while her CA125 level was 28 U / mL (within the normal range). Based on the positive HE4 result, the clinician recommended surgical exploration, and postoperative pathology confirmed high-grade serous carcinoma. In this case, the high specificity of HE4 avoided the risk of missed diagnosis that might have been due to a negative CA125 result. However, HE4 has limited sensitivity. For instance, a 60-year-old postmenopausal woman had an elevated serum CA125 level of 62 U / mL, but a HE4 level of 98 pmol / L (below the upper limit of normal, 140 pmol / L). Postoperative pathology diagnosed her with mucinous carcinoma. In this case, the false-negative result of HE4 led the clinician to make a diagnosis based solely on elevated CA125 levels, failing to obtain the supplementary information provided by HE4. HE4's inherent limitation lies in its low sensitivity to specific pathological types such as mucinous carcinoma and clear cell carcinoma.
[0056] CA125 and HE4 exhibit complementary "sensitivity-specificity" characteristics in diagnostic efficacy: CA125 has high sensitivity but is prone to false positives, while HE4 has high specificity but is prone to missing certain types of malignant lesions. Therefore, in clinical practice, both are often tested together to calculate the Ovarian Malignant Tumor Risk Algorithm (ROMA) index in order to improve diagnostic accuracy. However, even with combined use, the positive predictive value of the above-mentioned serum marker combination in early ovarian cancer is still insufficient to meet screening criteria, especially in premenopausal women, where the false positive rate remains high.
[0057] For ultrasound imaging assessment, transvaginal ultrasound is the first-line imaging method for evaluating adnexal masses, clearly displaying the size, shape, internal structure, and blood flow characteristics of the ovary. To standardize ultrasound descriptions and unify risk stratification, the American College of Radiology officially launched the Ovarian-Adnexal Reporting and Data System (O-RADS) in 2019, and released an updated version in 2022. This system classifies adnexal masses into categories 0 to 5, corresponding to different malignancy risk ranges: O-RADS 2 (malignancy risk <1%), O-RADS 3 (1%-10%), O-RADS 4 (10%-50%), and O-RADS 5 (≥50%).
[0058] The O-RADS system has significantly improved the consistency of ultrasound diagnosis in clinical applications. However, its diagnostic efficacy in early-stage ovarian cancer still faces challenges. For example, a 58-year-old postmenopausal woman presented with a 6.5cm multilocular cystic mass in the left adnexa on ultrasound. The cyst wall was smooth, with multiple septa visible within, no clearly defined solid components, and a blood flow score of 2. According to the 2022 O-RADS guidelines, this lesion was classified as O-RADS category 4 (intermediate malignancy risk). Following the guidelines' recommendations, the patient underwent surgical exploration, and postoperative pathological diagnosis revealed mucinous cystadenocarcinoma (early stage). In this case, the correct classification of O-RADS category 4 prevented a missed diagnosis. However, another patient's experience illustrates the limitations of O-RADS category 4 diagnosis. A 52-year-old postmenopausal woman presented with a 4.2cm cystic-solid mass in the right adnexa on ultrasound. The solid component was small and regular, with a blood flow score of 3, and was classified as O-RADS category 4. Based on this classification, the physician recommended surgery, but the postoperative pathology result was a theca cell tumor, a benign sex cord-stromal tumor. In this case, a false positive result for O-RADS category 4 led to unnecessary surgery.
[0059] The O-RADS category 4 covers a wide range of malignant risk (10%-50%), which often leads to decision-making dilemmas in clinical practice. For benign lesions with the lower risk, excessive surgical intervention may result in unnecessary trauma; for malignant lesions with the upper risk, conservative management may delay treatment. Furthermore, the O-RADS system presents diagnostic challenges for certain pathological types of early ovarian cancer. For example, a 45-year-old premenopausal woman presented with a large, multilocular cystic mass measuring 10.2 cm in diameter in the left adnexa on ultrasound. The mass contained no clearly defined internal components, only multiple cystic septa, and was classified as O-RADS category 3 (low malignant risk, 1%-10%). Based on this classification, the clinician recommended follow-up observation. However, a follow-up examination three months later showed a significant increase in mass size, and surgical pathology diagnosed mucinous carcinoma (advanced to stage IC). In this case, the O-RADS system's insufficient recognition of mucinous carcinoma led to a diagnostic delay, and the patient missed the optimal surgical opportunity.
[0060] In clinical practice, physicians typically need to make a comprehensive judgment based on serum biomarker test results, ultrasound imaging features, and patient clinical variables (such as age, menopausal status, and family history). However, the integration of this multimodal information is highly dependent on the physician's personal experience and subjective judgment, and lacks standardized integration strategies.
[0061] For example, a 50-year-old premenopausal woman presented with an ultrasound showing a 5.5cm cystic-solid mass in her right adnexa. The solid component was irregular, with a blood flow score of 4 and an O-RADS classification of 4. Her serum CA125 was 48 U / mL (elevated), and her HE4 was 68 pmol / L (below the upper limit of premenopausal normal 70 pmol / L). Faced with this case, different physicians might offer different treatment recommendations: one might believe that an O-RADS classification of 4 combined with elevated CA125 is sufficient to support surgical exploration; another might prefer conservative observation due to a normal HE4 level, suggesting short-term follow-up before making a final decision. This difference in decision-making reflects the subjectivity and uncertainty of multimodal information integration, which could lead some patients to miss the optimal treatment window or undergo unnecessary interventions.
[0062] In summary, the existing technology has the following problems:
[0063] First, single-modality diagnostic methods (serum markers or ultrasound) have inherent limitations. Although CA125 and HE4 are complementary, they lack an effective integration mechanism. The O-RADS system has limited ability to differentiate early ovarian cancer, especially for type 4 lesions, and the problem of insufficient specificity still exists.
[0064] Second, in clinical practice, the integration of multi-source information such as serological indicators, ultrasound imaging features and clinical variables relies on the physician's subjective experience and lacks standardized and automated multimodal fusion strategies, making it difficult to fully leverage the synergistic and complementary value of multi-source data.
[0065] Third, existing diagnostic models are mostly based on advanced ovarian cancer cohorts. There is still a lack of intelligent diagnostic systems that target early lesions and integrate multimodal information such as imaging, serological and clinical variables, which cannot meet the clinical needs for accurate diagnosis of early ovarian cancer.
[0066] Therefore, there is an urgent need for an intelligent diagnostic system that can dynamically integrate multimodal data and automatically learn cross-modal correlation features to overcome the performance limitations of existing single diagnostic methods.
[0067] In view of this, the present invention provides a novel solution: a multimodal fusion-based auxiliary diagnostic system for early ovarian cancer. The technical concept of this invention lies in constructing a multimodal fusion system based on an attention mechanism to deeply integrate three types of heterogeneous data: raw ultrasound images, serum biomarkers, and clinical variables. Specifically, the system automatically extracts deep morphological features from ultrasound images that are difficult for the human eye to recognize through an image feature extraction network, extracts structured information of serum biomarkers and clinical variables through a non-image feature encoding network, and introduces a channel-space dual attention mechanism to dynamically learn the contribution weights of different modal features to diagnostic decisions, achieving adaptive fusion of cross-modal information. Based on this, the system further integrates gradient-weighted class activation mapping and SHAP analysis modules to visualize and explain the fusion decision process, and optimizes the clinical management pathway for Category 4 lesions in the ovarian-adnexal reporting and data system through a risk restratification module. This technical concept breaks through the limitations of traditional single-modal diagnosis, replacing experience-dependent static combination methods with a data-driven dynamic fusion strategy, providing a systematic solution for accurate, interpretable, and intelligent auxiliary diagnosis of early ovarian cancer.
[0068] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.
[0069] According to a first aspect of the present invention, an early ovarian cancer auxiliary diagnostic system based on multimodal fusion is provided, comprising a data acquisition module, an image feature extraction module, a non-image feature extraction module, a multimodal fusion module, and a diagnostic output module.
[0070] The system comprises the following modules: a data acquisition module for acquiring raw ultrasound images, serum biomarker values, and clinical variable data of the subject to be diagnosed; an image feature extraction module connected to the data acquisition module for preprocessing the raw ultrasound images and extracting image depth features using a pre-trained image feature extraction network; a non-image feature extraction module connected to the data acquisition module for fusing serum biomarker values and clinical variable data and extracting non-image features using a non-image feature encoding network; a multimodal fusion module connected to both the image feature extraction and non-image feature extraction modules for dynamically calculating the attention weights of image depth features and non-image features using an attention-based cross-modal fusion module, and generating fused features based on these attention weights; and a diagnostic output module connected to the multimodal fusion module for processing the fused features using a classifier and outputting the risk probability value of the subject to be diagnosed having early-stage ovarian cancer.
[0071] Through the above technical solution, firstly, the system of the present invention processes three types of data through an image feature extraction module and a non-image feature extraction module (wherein, ultrasound images provide morphological information, serum markers reflect the biological activity of tumors, and clinical variables reflect host factors), thereby enabling unified representation and collaborative utilization of multi-source heterogeneous data in the same system, providing a structured data foundation for subsequent cross-modal fusion.
[0072] Secondly, in this invention, the image feature extraction module preprocesses the original ultrasound image and extracts image depth features through a pre-trained image feature extraction network. This enables automated extraction of high-level semantic features from the ultrasound image. Simultaneously, the non-image feature extraction module fuses serum biomarker detection values and clinical variable data, and extracts non-image features through a non-image feature encoding network. Through layer-by-layer mapping of the encoding network, these heterogeneous structured data (serum biomarkers and clinical variables are structured data with varying numerical ranges and distribution characteristics) are uniformly encoded into feature vectors, achieving alignment with non-image data and creating conditions for cross-modal fusion. More importantly, this invention inputs image depth features and non-image features into an attention-based cross-modal fusion module, dynamically calculating the attention weights of the two modalities and generating fused features based on these weights. Thus, firstly, the attention mechanism enables the system to adaptively adjust the contribution ratio of image and non-image modalities in the fused features according to the specific circumstances of each diagnostic subject. For cases with typical ultrasound imaging features but atypical serum biomarkers, the system can assign higher weights to imaging features; for cases with atypical ultrasound findings but abnormal serological markers, the system can increase the weights of non-imaging features. This dynamic fusion strategy overcomes the limitations of traditional static fusion (such as feature splicing and weighted averaging) which uses fixed fusion weights for all samples. Secondly, cross-modal information complementarity is achieved. By calculating the correlation strength between imaging features and non-imaging features through an attention mechanism, the system can identify and strengthen the synergistic information between the two modalities, ensuring that the fused features retain both the morphological information of the images and integrate the biological activity information of the serological markers, achieving a complementary gain of "1+1>2".
[0073] Third, in this invention, the diagnostic output module processes the fused features through a classifier and outputs a risk probability value. In other words, this invention transforms the multimodal fusion result into a continuous risk probability value (typically between 0 and 1), providing a quantitative risk assessment basis for clinical decision-making. Compared to the five-level discrete classification (2, 3, 4, and 5 classes) of the traditional O-RADS system, the continuous probability value can more precisely distinguish the malignancy of different lesions within the same risk level, providing clinicians with flexible threshold adjustment space in different decision-making scenarios (such as selecting a low threshold in screening scenarios and a high threshold in diagnostic scenarios).
[0074] In one embodiment of the present invention, the image feature extraction module of the present invention may include an image preprocessing unit and a feature extraction unit, wherein the image preprocessing unit is used to adjust the original ultrasound image to a preset resolution and perform pixel value normalization processing; the feature extraction unit has a built-in convolutional neural network, which is used to receive the preprocessed ultrasound image and output the feature vector generated after the global average pooling layer in the convolutional neural network as the image depth feature.
[0075] In this embodiment, firstly, the original ultrasound image is adjusted in resolution and normalized in pixel value by the image preprocessing unit, which can eliminate the differences in image scale and grayscale distribution caused by different devices and different acquisition parameters, thereby providing standardized input for feature extraction; secondly, through the composite scaling strategy of the EfficientNet-B4 convolutional neural network, the system of the present invention can output feature vectors from the global average pooling layer as image depth features while balancing the network depth, width and resolution. It retains the high-level semantic information of the image and has a moderate dimension, which is convenient for subsequent cross-modal fusion.
[0076] In one embodiment of the present invention, the non-image feature extraction module may include a data splicing unit and a feature encoding unit; the data splicing unit is used to splice serum biomarker detection values and clinical variable data to form a non-image input vector; the feature encoding unit has a built-in fully connected network, which sequentially includes a first fully connected layer, a first batch normalization layer, a first activation function layer, a second fully connected layer, a second batch normalization layer, a second activation function layer, a third fully connected layer, a third batch normalization layer, and a third activation function layer, for processing the non-image input vector layer by layer and outputting non-image features.
[0077] In this implementation, firstly, the serum biomarker detection values and clinical variable data are integrated into a unified non-image input vector through a data stitching unit, enabling standardized representation of structured data. Secondly, a three-layer fully connected network combined with batch normalization and activation functions can progressively map the original input vector to a higher-dimensional feature space, allowing the network to learn nonlinear combination relationships in non-image data. Finally, the layer-by-layer dimensionality reduction design (input → 128 dimensions → 64 dimensions → 32 dimensions) can reduce the feature dimension while retaining key information, reducing the computational burden on subsequent fusion modules.
[0078] In one embodiment of the present invention, the multimodal fusion module may include a linear mapping unit, an attention calculation unit, and a feature fusion unit. The linear mapping unit is used to perform a first linear transformation matrix... Image depth features Mapping to the target dimension space yields the mapped image features. And through the second linear transformation matrix Non-image features Mapping to the target dimension space yields the mapped non-image features. ,in , ;
[0079] The attention computation unit is connected to the linear mapping unit and is used to calculate attention weights. , In the formula, Let be the dimension value of the target dimension space. This is matrix multiplication, where T denotes matrix transpose;
[0080] The feature fusion unit is connected to the attention computation unit to generate fused features. .
[0081] In this implementation, firstly, the image depth features and non-image features are mapped to the same dimensional space by the linear mapping unit, thereby solving the problem of inconsistent dimensions and mismatched physical meanings of the two modal features and creating a comparable basis for attention calculation. Secondly, the attention calculation unit calculates the softmax-normalized dot product attention weights, enabling the system to dynamically evaluate the relative importance of the two modal features based on the specific circumstances of each sample. Finally, the feature fusion unit performs a weighted summation using the attention weights as coefficients, generating a fused feature that simultaneously contains information from both modalities, and the weight allocation adaptively changes with the input samples. In this way, true dynamic cross-modal fusion can be achieved.
[0082] In one embodiment of the present invention, the early ovarian cancer auxiliary diagnostic system based on multimodal fusion may further include an interpretability analysis module; the interpretability analysis module is connected to the image feature extraction module and the diagnostic output module, and is used to generate a heat map corresponding to the original ultrasound image through a gradient weighted class activation mapping algorithm, and to calculate the marginal contribution of each feature in the serum biomarker detection value and clinical variable data to the risk probability value through the SHAP analysis method.
[0083] In this implementation, firstly, a heatmap corresponding to the original ultrasound image is generated using a gradient-weighted class activation mapping algorithm. This visualizes the image regions that the image feature extraction module focuses on during decision-making, making the system's decision-making basis transparent and facilitating clinicians' verification of the model's rationale. Secondly, by calculating the marginal contribution of each input feature to the risk probability value using the SHAP analysis method, the influence of serum biomarkers and clinical variables in decision-making can be quantified, providing clinicians with a quantitative basis for understanding the model's judgment logic. Finally, the above two interpretation methods cover both image modalities and non-image modalities, thus ensuring comprehensive interpretability of the system's decisions.
[0084] In one embodiment of the present invention, the early ovarian cancer auxiliary diagnostic system based on multimodal fusion may further include a risk restratification module. The risk restratification module is connected to the data acquisition module and is used to acquire the ovarian-adnexal report and data system classification of the subject to be diagnosed. When the classification is O-RADS Category 4, the malignancy risk value is calculated using the simple rule risk model of the International Ovarian Tumor Analysis Group. When the malignancy risk value is less than a preset threshold, the risk level of the subject to be diagnosed is adjusted from O-RADS Category 4 to O-RADS Category 3.
[0085] In this implementation, for the clinical decision-making dilemma group in O-RADS Category 4, which has a wide range of malignant risk (10%-50%), the SRR model is introduced for secondary assessment to further refine the risk stratification. At the same time, by comparing preset thresholds, lesions assessed as low risk by SRR are downgraded from Category 4 to Category 3, freeing some benign lesions from unnecessary surgical exploration, thereby effectively reducing overtreatment. In addition, this re-stratification strategy can improve specificity to a certain extent without significantly reducing sensitivity, thus optimizing the clinical application pathway of the O-RADS system.
[0086] In one embodiment of the present invention, the early ovarian cancer auxiliary diagnostic system based on multimodal fusion may further include a model calibration and evaluation module; the model calibration and evaluation module is connected to the diagnostic output module and is used to evaluate the calibration degree of the risk probability value through the Hosmer-Lemeshow goodness-of-fit test, and to calculate the clinical net benefit under different threshold probabilities through decision curve analysis to determine the optimal threshold probability; wherein, the formula for calculating the clinical net benefit is:
[0087]
[0088] In the formula, TP represents the number of true positives, FP represents the number of false negatives, N represents the total sample size, and pt represents the threshold probability.
[0089] In this implementation, the Hosmer-Lemeshow goodness-of-fit test is used to assess the calibration of the risk probability values and verify the consistency between the model's output probability and the actual observation frequency, thereby ensuring the reliability of the model's predictions. Simultaneously, decision curve analysis is used to calculate the clinical net benefit under different threshold probabilities, quantifying the clinical advantages of using the system-guided decision-making approach compared to "full intervention" or "no intervention" strategies. Furthermore, the optimal threshold probability determined based on the net benefit curve provides a quantitative basis for threshold selection in different clinical scenarios (such as screening scenarios and preoperative assessment scenarios).
[0090] In one embodiment of the present invention, the image feature extraction module of the present invention may further include a model compression unit; the model compression unit is used to perform lightweight processing on the pre-trained image feature extraction network to generate a compressed model, the compressed model having a smaller size than the original model and a higher inference speed than the original model.
[0091] In this embodiment, the storage volume of the model is reduced through lightweight processing, enabling it to be deployed on clinical terminal devices with limited computing resources. At the same time, the inference speed of the compressed model is higher than that of the original model, which can meet the response speed requirements of real-time clinical diagnostic scenarios. In addition, the lightweighting of the model while maintaining the core diagnostic functions allows the system of the present invention to be embedded in the hospital image archiving and communication system (PACS).
[0092] In one embodiment of the present invention, the image feature extraction module of the present invention may further include a knowledge distillation unit; the knowledge distillation unit is used to perform knowledge distillation using a pre-trained image feature extraction network as a teacher network and a convolutional neural network with fewer parameters as a student network to generate a lightweight student model, the inference latency of which is lower than that of the teacher network.
[0093] In this embodiment, the learning ability of the teacher network is transferred to the student network with fewer parameters through knowledge distillation, which can significantly reduce the model complexity while retaining the core discriminative ability of the teacher network. At the same time, the inference latency of the student network is lower than that of the teacher network, making it suitable for clinical application scenarios with higher real-time requirements. In addition, the lightweight student network generated by knowledge distillation enables the system of the present invention to be deployed on different hardware platforms (such as portable ultrasound equipment and workstations in primary healthcare institutions).
[0094] In one embodiment of the present invention, the early ovarian cancer auxiliary diagnostic system based on multimodal fusion may further include a system integration interface; the system integration interface is used to embed the image feature extraction module, the non-image feature extraction module, the multimodal fusion module and the diagnostic output module into the hospital's image archiving and communication system, the system integration interface supports medical digital imaging and communication standards, and is used to receive raw ultrasound images and output risk probability values and corresponding heat maps.
[0095] In this embodiment, by supporting the DICOM standard, the system of the present invention can seamlessly interface with the hospital's existing image archiving and communication systems, receive raw ultrasound images and output risk probability values and heat maps without changing clinical workflows. At the same time, embedding the intelligent diagnostic module into the hospital information system enables the widespread application of artificial intelligence-assisted diagnostic functions in routine clinical work. In addition, the system's integrated interface allows the system of the present invention to directly utilize the hospital's existing hardware resources and data flow, reducing deployment costs and implementation barriers.
[0096] The present invention will now be described with reference to an exemplary embodiment.
[0097] 1.1. Collection of clinical data.
[0098] Based on previous epidemiological evidence and studies on ovarian cancer risk factors, the following clinical variables were systematically collected: (1) Demographic characteristics: age, height, weight (BMI = weight kg / height m²); (2) Menstrual and reproductive history: age at menarche, menopausal status (stratified according to the 2022 O-RADS guidelines: premenopausal, early menopause <5 years, late menopause ≥5 years), number of pregnancies, number of live births; (3) Family history: family history of ovarian or breast cancer in first-degree relatives (mother, sisters, daughters) (yes / no); (4) Clinical symptoms: a binary recording method was used, including abdominal distension, abdominal pain, irregular vaginal bleeding, gastrointestinal symptoms, etc. All clinical data were collected through the hospital's electronic medical record system, independently extracted and cross-checked by two researchers, and recorded using a standardized case report form (CRF).
[0099] 1.2 Serum sample collection and testing.
[0100] All subjects underwent fasting blood collection (5 mL) from their elbow vein in the morning before surgery, placed in vacuum blood collection tubes without anticoagulants. After standing at room temperature for 30 minutes, the blood samples were centrifuged at 3000 rpm for 10 minutes at 4°C to separate the serum. The serum was aliquoted into 1.5 mL cryovials, labeled, and immediately stored at -80°C until unified testing.
[0101] Serum CA125 and HE4 levels were detected using a Roche Cobas 8000 / e801 fully automated chemiluminescence immunoassay analyzer and accompanying reagents. The testing process strictly adhered to standard operating procedures, and each batch included high- and low-value quality controls (Liquichek™ tumor marker controls) to ensure an intra-batch coefficient of variation (CV) of <5% and an inter-batch CV of <8%. Laboratory personnel maintained blindness to the patients' clinical diagnoses and pathological results.
[0102] 1.3 Ultrasound image acquisition and O-RADS classification.
[0103] All ultrasound examinations were performed by two senior physicians from the gynecological ultrasound team. Before the examination, all physicians received specialized training on the O-RADS ultrasound dictionary and the 2022 guidelines. The examination protocol followed a standardized procedure: transvaginal ultrasound (probe frequency 5-9MHz) was the first choice. For lesions with a maximum diameter >10cm or patients with no history of sexual activity, transabdominal ultrasound (probe frequency 3-5MHz) was used as a supplement.
[0104] The following key images are retained according to the image acquisition specifications: (1) grayscale images of the largest transverse and longitudinal sections of the lesion (with and without measurement markers are saved respectively); (2) the image of the section with the richest blood flow in color Doppler mode; (3) close-up images of suspicious malignant signs (such as papillary protrusions, ascites, irregular thickening of the cyst wall, solid components, etc.). After the acquired images are reviewed and approved by senior physicians with more than 15 years of experience in gynecological ultrasound, they are uniformly stored in the hospital's PACS system.
[0105] Two ultrasound physicians independently classified the tumor according to the ACR 2022 O-RADS Ultrasound Dictionary and Guidelines. O-RADS classifications include: Category 0 (unable to be fully assessed), Category 1 (physiological, malignancy risk <1%), Category 2 (almost certainly benign, malignancy risk <1%), Category 3 (low malignancy risk, 1%-10%), Category 4 (intermediate malignancy risk, 10%-50%), and Category 5 (high malignancy risk, ≥50%). If there is disagreement regarding the classification results, a consensus is reached through negotiation; if a consensus still cannot be reached after negotiation, a third arbitrator physician makes the final decision.
[0106] 1.4 Data Preprocessing (see [link to documentation]) Figure 1 ).
[0107] Image preprocessing: The original ultrasound images were uniformly adjusted to a resolution of 256×256 pixels, and the pixel values were normalized to the [0,1] interval using the min-max normalization method. Experienced ultrasound physicians manually delineated lesion boundaries using annotation tools, extracting the lesion regions as model input. To expand the diversity of the training set, data augmentation strategies were applied: random rotation (±30°), horizontal flip (probability 0.5), random scaling (1.0-1.2), brightness perturbation (±0.25), and image distortion (probability 0.15).
[0108] Serum-clinical modality preprocessing: Continuous variables (age, BMI, age at menarche, CA125, HE4) were Z-score standardized; categorical variables (menopausal status, family history, clinical symptoms, O-RADS classification, etc.) were converted into binary vectors using one-hot encoding.
[0109] 1.5 Model Building.
[0110] This implementation proposes a multimodal attention fusion network (MAF-Net) based on an attention mechanism. The overall architecture consists of three core modules:
[0111] Image coding branch: EfficientNet-B4 is used as the backbone network. Based on ImageNet pre-training, a compound scaling strategy is used to balance the network depth, width, and resolution. The top fully connected classification layer is removed, and the feature vector after the global average pooling layer is retained, with an output dimension of 1792, which serves as the deep semantic representation of the ultrasound image.
[0112] Non-image coding branch: A three-layer fully connected network was designed, fusing serum biomarkers (CA125, HE4) and clinical features. Network structure: Input layer → 128 neurons (batch normalization + ReLU + Dropout 0.3) → 64 neurons (batch normalization + ReLU + Dropout 0.3) → 32 neurons (batch normalization + ReLU + Dropout 0.3).
[0113] Cross-modal attention fusion module: Introduces a channel-space dual attention mechanism to dynamically learn the weight contributions of features from different modalities. Fusion formula: F_fused = α·F_image + β·F_non-image, where α and β are learnable attention weight coefficients. The final fused features output the malignancy probability through a classification head (64→32→2), using the Sigmoid activation function.
[0114] 1.6 Model training strategy.
[0115] The training set employs five-fold cross-validation for hyperparameter optimization and model selection. The loss function uses binary cross-entropy combined with Focal Loss (γ=2.0) to address class imbalance. The optimizer is AdamW with an initial learning rate of 1e-4 and weight decay of 0.01. Learning rate scheduling uses cosine annealing, coupled with a warm-up mechanism (for the first 10 rounds). Regularization strategies include: Dropout rate of 0.3, label smoothing (ε=0.1), gradient clipping (max_norm=1.0), and early stopping (patience value for 10 rounds).
[0116] Training hardware environment: NVIDIA A100 (40GB VRAM) ×2, automatic mixed precision accelerated training, batch size 32, and a total training cycle limit of 200 rounds.
[0117] 1.7 Model interpretability analysis.
[0118] A dual visualization interpretation strategy is adopted: (1) Gradient-weighted class activation mapping (Grad-CAM++): a heat map is generated and superimposed on the original ultrasound image to show the key areas that the model focuses on when making decisions; (2) SHAP value analysis: based on the game theory Shapley value, the marginal contribution of each input feature to the individual prediction result is calculated, and the feature importance ranking and individualized decision map are generated.
[0119] 2.1 Study the baseline characteristics of the population.
[0120] Please see Figure 2 This invention consecutively included patients with adnexal masses who visited the gynecology department of a certain hospital and underwent surgical treatment between July 2023 and June 2025. All patients underwent standardized ultrasound examinations and serological tests before surgery, and obtained a clear pathological diagnosis after surgery. After screening according to inclusion and exclusion criteria, a total of 600 patients were included. A computer randomization function was used to divide the patients into a training set (420 cases) and an independent prospective validation set (180 cases) in a 7:3 ratio. The training set was used for model development, hyperparameter tuning, and five-fold cross-validation; the validation set remained completely blinded throughout the model training phase and was only unblinded during the final performance evaluation to ensure the level of research evidence.
[0121] 2.2 Distribution of pathological types.
[0122] Of the 600 cases of adnexal masses, 423 were benign (70.5%) and 177 were malignant (29.5%, including 38 borderline tumors). In the malignant group, 152 cases were FIGO stage I early-stage ovarian cancer, accounting for 85.9% of the malignant cases, consistent with the focus of this study on early-stage cancer. FIGO staging distribution: stage IA 48 cases (31.6%), stage IB 14 cases (9.2%), and stage IC 90 cases (59.2%).
[0123] 2.3 Single-modal baseline performance evaluation.
[0124] Before constructing the multimodal model, baseline models for each unimodal model were first established to quantify the independent contribution of each data source to the diagnosis of early ovarian cancer and to provide a performance benchmark for subsequent fusion models. All models were fine-tuned using five-fold cross-validation on the training set, and their final performance was evaluated on the independent validation set.
[0125] 2.3.1 Clinical-serological baseline model.
[0126] Based on clinical variables (age, menopausal status, clinical symptoms, BMI, age at menarche) and serological indicators (CA125, HE4), three classic machine learning models were constructed: Logistic Regression (LR), Support Vector Machine (SVM), Random Forest (RF), XGBoost, and Lightweight Gradient Boosting Machine (LightGBM). Grid search was used for hyperparameter tuning, and five-fold cross-validation was employed to select the optimal model.
[0127] Logistic regression: L2 regularization was used with a regularization coefficient C=0.1, and the solver was liblinear. Input features were standardized using Z-score.
[0128] Support Vector Machine: It adopts the radial basis function (RBF), gamma='scale', regularization parameter C=1.0, and class weights are 'balanced' to deal with class imbalance.
[0129] Random Forest: Number of trees n_estimators=200, maximum depth max_depth=10, minimum number of leaf samples min_samples_leaf=5, feature importance uses Gini impurity.
[0130] XGBoost: Learning rate eta=0.05, maximum depth max_depth=6, subsample=0.8, column sampling colsample_bytree=0.8, objective function binary: logistic, evaluation metric AUC.
[0131] LightGBM: learning rate=0.05, number of leaves num_leaves=31, feature sampling_fraction=0.8, data sampling_fraction=0.8, objective function binary.
[0132] The results show (see Table 1 below) that ensemble learning methods (RF, XGBoost, LightGBM) outperform single models (LR, SVM). LightGBM showed the best performance, with an AUC of 0.852 (95% CI: 0.811-0.893), sensitivity of 77.4%, and specificity of 82.7%. Random Forest had an AUC of 0.847 (95% CI: 0.806-0.888), and XGBoost had an AUC of 0.850 (95% CI: 0.809-0.891). Logistic Regression had an AUC of 0.821 (95% CI: 0.778-0.864), and SVM had an AUC of 0.835 (95% CI: 0.792-0.878).
[0133] When serum biomarkers are used alone, the AUC of CA125 is 0.753 (95% CI: 0.702-0.804), with a sensitivity of 67.9% and a specificity of 74.0% when the threshold is >35 U / mL; the AUC of HE4 is 0.781 (95% CI: 0.732-0.830), with a sensitivity of 71.7% and a specificity of 76.4% when the thresholds are >140 pmol / L (premenopausal) and >140 pmol / L (postmenopausal). The AUC of the combined biomarkers (Logistic regression) increases to 0.821, suggesting that the combined biomarkers have complementary value.
[0134] Table 1. Diagnostic performance of the clinical-serological baseline model (validation set)
[0135]
[0136] 2.3.2 Ultrasonic single-modal baseline model.
[0137] Based on ultrasound images, six deep learning models were constructed: ResNet50, ResNet101, DenseNet121, DenseNet169, InceptionV3, and the EfficientNet series (B0-B7). All models were initialized with ImageNet pre-trained weights and fine-tuned on the training set. Meanwhile, Structured O-RADS 2022 classification was used as the benchmark for traditional ultrasound evaluation.
[0138] Image preprocessing: The original DICOM images were uniformly adjusted to a resolution of 256×256 pixels, and the pixel values were normalized to the [0,1] interval using the min-max normalization method. Two senior ultrasound physicians manually delineated the lesion boundaries using annotation tools, and extracted the lesion regions as model input. To expand the diversity of the training set, online data augmentation strategies were applied: random rotation (±30°), horizontal flip (probability 0.5), vertical flip (probability 0.3), random scaling (0.8-1.2), brightness perturbation (±0.25), contrast perturbation (±0.2), Gaussian noise (σ=0.01), and image distortion (probability 0.15).
[0139] Training strategy: The optimizer used is AdamW, with an initial learning rate of 1e-4 and weight decay of 0.01. Cosine annealing is used for learning rate scheduling, combined with a warm-up mechanism (for the first 10 epochs). Binary cross-entropy is used as the loss function. Batch size is 32, training epochs are capped at 100, and an early stopping mechanism (15 epochs of patience) is implemented. Automatic mixed precision is used to accelerate training.
[0140] Model selection: Evaluate the performance of each model on the validation set, and select the optimal backbone network by comprehensively considering AUC, number of parameters, computational complexity and inference speed.
[0141] The results (see Table 2 below) show that EfficientNet-B4 has the best performance among the six deep models, with an AUC of 0.914 (95% CI: 0.882-0.946), sensitivity of 86.8% (46 / 53), specificity of 84.3% (107 / 127), and accuracy of 85.0% (153 / 180). The AUCs of ResNet50 and DenseNet121 are 0.897 and 0.905, respectively. Within the EfficientNet series, performance initially increases and then plateaus with increasing network depth (B0→B7), with B4 reaching the performance-efficiency balance point.
[0142] As a control, the O-RADS 2022 version (with ≥4 categories as positive) had an AUC of 0.895 (95% CI: 0.861-0.929), sensitivity of 90.6% (48 / 53), specificity of 78.0% (99 / 127), and accuracy of 81.7% (147 / 180). The deep learning model significantly outperformed O-RADS in specificity (84.3% vs 78.0%, McNemar test P=0.032), indicating that deep feature extraction based on original images can more accurately identify benign lesions and reduce false positives.
[0143] Table 2 Diagnostic performance of ultrasound single-modal model (validation set)
[0144]
[0145] 2.4 Comparison of multimodal fusion strategies.
[0146] Based on the EfficientNet-B4 image coding branch and clinical-serological coding branch, the performance differences of five multimodal fusion strategies are systematically compared:
[0147] (1) Early fusion: The original image features (processed by a shallow CNN) are concatenated with the clinical features at the input layer and then input into a single network. The image branch extracts 64-dimensional features through 3 convolutional layers, which are then concatenated with the clinical features (32-dimensional) and input into a fully connected network (128→64→2).
[0148] (2) Intermediate fusion: High-dimensional features are extracted separately and then concatenated in the hidden layer. The image branch extracts 1792-dimensional features using EfficientNet-B4, which are then compressed to 256-dimensionality by a dimensionality reduction layer; the clinical branch extracts 32-dimensional features using a three-layer fully connected layer (128→64→32). The two are then concatenated and input into the classification head (288→128→2).
[0149] (3) Late-stage fusion: Single-modal models are trained separately and weighted averages are performed at the decision layer. The output probability of the image branch is P_img, the output probability of the clinical branch is P_clin, and the final probability is P = w·P_img + (1-w)·P_clin. The weight w is optimized through cross-validation (optimal w=0.72).
[0150] (4) Gated fusion: A gating mechanism is introduced to dynamically adjust modal contributions. The gating network generates a 256-dimensional weight vector based on clinical features and performs element-wise weighting with image features: F_fused=F_img⊙σ(W_gate·F_clin+b_gate), where ⊙ represents the Hadamard product.
[0151] (5) Attention Fusion (MAF-Net): Introduces a channel-space dual attention mechanism to dynamically learn cross-modal weights.
[0152] MAF-Net detailed architecture:
[0153] Image encoding branch: EfficientNet-B4 backbone network, fine-tuned based on ImageNet pre-training. The top fully connected classification layer is removed, retaining the feature vectors after global average pooling, resulting in an output dimension of 1792. A self-attention module is added to calculate the dependencies between feature channels, enhancing the expression of key semantic features.
[0154] Non-image coding branch: A three-layer fully connected network fusing serum biomarkers (CA125, HE4) and clinical features (age, menopausal status, clinical symptoms, BMI). Network structure: Input layer (12-dimensional) → 128 neurons (batch normalization + ReLU + Dropout 0.3) → 64 neurons (batch normalization + ReLU + Dropout 0.3) → 32 neurons (batch normalization + ReLU + Dropout 0.3). Output: 32-dimensional feature vector.
[0155] Cross-modal attention module:
[0156] Channel attention: Global average pooling and global max pooling are performed on the image features, and channel weight vectors are generated through two fully connected layers (compression ratio r=16), which are then multiplied with the original image features.
[0157] Spatial attention: The channel-weighted image features are subjected to average pooling and max pooling along the channel axis, and after being stitched together, they are convolved by 7×7 to generate a spatial weight map, which is then multiplied with the features.
[0158] Cross-modal fusion: Image features (1792 dimensions) and non-image features (32 dimensions) are mapped to the same dimension (256 dimensions) through linear transformation. The attention score α=softmax(F_img·F_clin^T / √d) is calculated, and the fused feature F_fused=α·F_img +(1-α)·F_clin.
[0159] Classification head: The fused features are processed through two fully connected layers (256→64→2), and the sigmoid activation function is used to output the malignancy probability.
[0160] The results (see Table 3 below) show that the attention fusion strategy (MAF-Net) performed best among all fusion methods, with an AUC of 0.954 (95% CI: 0.931-0.977), significantly higher than early fusion (0.916, P=0.008), mid-stage fusion (0.928, P=0.021), late fusion (0.923, P=0.015), and gated fusion (0.942, P=0.038). MAF-Net had a sensitivity of 92.5% (49 / 53), a specificity of 89.8% (114 / 127), an accuracy of 90.6% (163 / 180), a positive predictive value of 79.0% (49 / 62), and a negative predictive value of 96.6% (114 / 118).
[0161] Table 3 Performance Comparison of Different Multimodal Fusion Strategies (Validation Set)
[0162]
[0163] 2.5 Comparison of ROC curves between MAF-Net, unimodal models, and clinical benchmarks.
[0164] The optimal multimodal model MAF-Net was compared head-to-head with the optimal single-modal model (EfficientNet-B4), the optimal clinical-serological model (LightGBM), the O-RADS 2022 guidelines, and the combination of serum biomarkers (CA125+HE4).
[0165] DeLong's test showed that the AUC of MAF-Net (0.954) was significantly higher than that of: serum CA125+HE4 combination (AUC=0.821, Z=4.82, P<0.001); LightGBM clinical model (AUC=0.852, Z=4.31, P<0.001); O-RADS 2022 guidelines (AUC=0.895, Z=2.76, P=0.006); and EfficientNet-B4 unimodal (AUC=0.914, Z=2.34, P=0.019).
[0166] 2.6 Decision Curve Analysis.
[0167] Decision curve analysis (DCA) was used to evaluate the clinical net benefit of the model at different threshold probabilities. Net benefit calculation formula:
[0168] Net Benefit = (TP / n) - (FP / n) × (pt / (1-pt)), where pt is the threshold probability. The DCA curve is shown (see [link]). Figure 3 MAF-Net showed significantly higher net benefit than the "all-treatment" and "no-treatment" strategies within a threshold probability range of 5%-35%, and outperformed the O-RADS guidelines, EfficientNet-B4 unimodal model, and LightGBM clinical model throughout the entire threshold range.
[0169] At the optimal threshold of 15% (the probability threshold corresponding to the Youden index), the net gains of each model are: MAF-Net: 0.218; EfficientNet-B4: 0.176; O-RADS 2022: 0.152; LightGBM: 0.124; CA125+HE4: 0.108.
[0170] The net benefit of MAF-Net, 0.218, means that for every 100 patients, model-guided decision-making can reduce unnecessary surgeries by approximately 22 compared to the "full surgical exploration" strategy, while maintaining a high detection rate for malignant lesions.
[0171] 2.7 Clinical decision threshold analysis.
[0172] Based on the probability distribution predicted by the validation set, clinical decision indicators under different thresholds were analyzed (as shown in Table 4). As the threshold increased from 0.05 to 0.30: sensitivity gradually decreased from 98.1% to 83.0%; specificity increased from 76.4% to 94.5%; positive predictive value increased from 63.4% to 83.0%; and negative predictive value decreased from 99.0% to 94.5%. In clinical practice, the threshold can be selected according to different scenarios: for screening scenarios, a low threshold (such as 0.10) can be selected to ensure high sensitivity (96.2%) and avoid missed diagnoses; for diagnosis scenarios, a high threshold (such as 0.25) can be selected to ensure high specificity (92.1%) and positive predictive value (81.0%) and reduce unnecessary interventions.
[0173] Table 4. Diagnostic performance of MAF-Net model at different thresholds (validation set)
[0174]
[0175] 2.8 Modal ablation.
[0176] Specific modalities were gradually removed, and performance changes were observed (see Table 5 below). Results showed that the imaging modality was the core contributor to model performance; after removing the imaging modality (retaining only serum and clinical data), the AUC decreased from 0.954 to 0.852 (Δ=-0.102, a relative decrease of 10.7%). Serum biomarkers contributed the second most, with the AUC decreasing to 0.932 after removal (Δ=-0.022, a relative decrease of 2.3%). Clinical variables contributed the least, with the AUC decreasing to 0.941 after removal (Δ=-0.013, a relative decrease of 1.4%). Full-modal fusion improved the AUC by 0.040 (a relative improvement of 4.4%) compared to the best single-modal imaging model (EfficientNet-B4, AUC=0.914).
[0177] Table 5. Results of Modal Ablation Experiments (Validation Set)
[0178]
[0179] 2.9 Attention module ablation.
[0180] The attention components in MAF-Net were gradually removed, and the contributions of each module were evaluated (see Table 6 below). The results show that: channel attention contribution AUC 0.009 (0.945→0.954); spatial attention contribution AUC 0.007 (0.947→0.954); cross-modal attention contribution AUC 0.012 (0.942→0.954); the full MAF-Net improved the mid-term fusion AUC by 0.026 (0.928→0.954) compared to the one without attention mechanism.
[0181] Table 6. Results of the attention module ablation experiment (validation set)
[0182]
[0183] 2.10 Data augmentation strategy ablation.
[0184] The impact of different data augmentation strategies on model performance was evaluated. The results showed that the complete data augmentation strategy improved the AUC by 0.021 (0.933→0.954) compared to no augmentation. Among the augmentation techniques, random rotation contributed the most (ΔAUC=0.008), followed by brightness perturbation (ΔAUC=0.005) and horizontal flipping (ΔAUC=0.004).
[0185] 3.1 Model compression and deployment.
[0186] To adapt to different deployment scenarios, MAF-Net models are compressed (see Table 7 below).
[0187] Table 7 Performance Comparison of Model Compression Versions
[0188]
[0189] The FP16 half-precision version maintains the same AUC (0.954) while halving the model size (38.6MB) and improving inference speed by 30% (9.8ms). The INT8 quantization version reduces the model size to 19.3MB, achieves an inference speed of 6.2ms, and slightly decreases the AUC by 0.003, making it suitable for resource-constrained edge deployment scenarios. The knowledge distillation version (student model EfficientNet-B2) reduces the number of parameters by 52%, achieves an inference speed of 10.1ms, and maintains an AUC of 0.948, demonstrating a good balance between performance and efficiency.
[0190] 3.2 Feasibility of integration with PACS system.
[0191] MAF-Net has the following features and supports embedding in hospital PACS systems: Model size: FP16 version 38.6MB, easily loaded into GPU memory. Inference speed: <15ms per instance, without affecting clinical workflow; Input / output: Accepts standard DICOM images, outputs probability values and heatmaps; Integration method: Outputs results via DICOM Structured Reports (SR) or HL7 interface; Deployment architecture: Supports local server deployment or cloud-edge collaboration.
[0192] In summary, firstly, this invention innovatively constructs a multimodal fusion network (MAF-Net) based on an attention mechanism, achieving excellent performance with an AUC of 0.954 on the independent validation set, with a sensitivity of 92.5%, specificity of 89.8%, and accuracy of 90.6%. This result is significantly better than the optimal single-modal model: compared with the ultrasound single-modal model (EfficientNet-B4, AUC=0.914), the AUC is improved by 0.040 (P=0.019); compared with the clinical-serological single-modal model (LightGBM, AUC=0.852), the AUC is improved by 0.102 (P<0.001); and compared with the O-RADS 2022 guidelines (AUC=0.895), the AUC is improved by 0.059 (P=0.006).
[0193] Second, the reason why multimodal fusion can significantly improve diagnostic efficacy lies in the pathophysiological basis of different modalities reflecting different aspects of tumor biology. Ultrasound imaging captures the morphological characteristics of tumors—solid components reflect densely proliferating areas of tumor cells, papillary projections suggest epithelial hyperplasia, and irregular infiltration of the cyst wall represents local invasive behavior. Serum biomarkers reflect the biological activity of tumors—high expression of HE4 is directly related to the proliferation and invasiveness of tumor cells, and elevated CA125 reflects more the tumor burden and the degree of peritoneal irritation. Clinical variables such as age and menopausal status reflect the influence of host factors on tumor development and progression. These three modalities are naturally complementary: morphological changes often precede significant increases in serum indicators, and some early-stage cancers with positive serological markers may have atypical ultrasound presentations. In this study, ablation experiments quantified the contribution of each modality: the imaging modality was the core contributor (AUC decreased by 0.102), followed by serum biomarkers (AUC decreased by 0.022), and clinical variables contributed the least (AUC decreased by 0.013). This quantitative result is highly consistent with the biological mechanism—morphological changes are direct evidence of tumor presence, while serological indicators and clinical factors are more indirect indications. Notably, multimodal fusion improved the AUC by 0.040 compared to the optimal single-modal imaging model. While this increase may seem small, it has significant clinical implications for early cancer diagnosis: it means that for every 100 patients, the multimodal model can correctly identify 4 more early-stage cancers than the single-modal imaging model, while reducing unnecessary interventions.
[0194] Third, compared with existing related technologies, the innovation of this invention lies not only in constructing a multimodal model, but also in employing an advanced attention fusion mechanism. Traditional fusion methods such as early fusion, mid-term fusion, and late-term fusion are all "hard fusions," that is, fixed weight splicing or weighting, which cannot dynamically adapt to the modal confidence of different samples. The channel-space dual attention mechanism of this invention achieves "soft fusion"—the model can dynamically calculate the weight contribution of different modal features according to the specific situation of each sample. The results show that the AUC (0.954) of attention fusion (MAF-Net) is significantly higher than that of early fusion (0.916), mid-term fusion (0.928), and late-term fusion (0.923), verifying the superiority of the dynamic weight allocation strategy. This mechanism has clear clinical rationale: for some cases with typical ultrasound features but low serum markers (such as some clear cell carcinomas), the model can assign higher weights to the imaging modality; for cases with atypical ultrasound manifestations but positive serological markers (such as some mucinous carcinomas), the model can increase the contribution of the serum modality. This "personalized instruction" decision-making approach is closer to the cognitive process of clinical experts who make comprehensive judgments based on multi-source information.
[0195] Overall, this invention successfully constructed a multimodal deep learning model, MAF-Net, based on an attention mechanism. By deeply fusing ultrasound images, serum biomarkers (CA125, HE4), and clinical variables, it achieved excellent diagnostic efficacy (AUC 0.954) in a prospective cohort of early-stage ovarian cancer, significantly outperforming the O-RADS guidelines and unimodal models. Interpretability analysis validated the consistency between model decisions and clinical understanding. Subgroup analysis clarified its unique value in O-RADS four-category predicament groups and seronegative patients. Decision curve analysis quantified its clinical utility in reducing unnecessary surgeries. This model combines high accuracy, interpretability, and deployment feasibility, providing an intelligent solution for the precise diagnosis of early-stage ovarian cancer and potentially promoting the homogenization and intelligentization of primary care screening.
[0196] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multimodal fusion-based auxiliary diagnostic system for early ovarian cancer, characterized in that, include: The data acquisition module is used to acquire the original ultrasound images, serum marker test values, and clinical variable data of the subject to be diagnosed. The image feature extraction module, connected to the data acquisition module, is used to preprocess the original ultrasound image and extract image depth features through a pre-trained image feature extraction network. A non-image feature extraction module, connected to the data acquisition module, is used to fuse the serum biomarker detection values and the clinical variable data, and extract non-image features through a non-image feature coding network; A multimodal fusion module is connected to the image feature extraction module and the non-image feature extraction module respectively. It is used to dynamically calculate the attention weights of the image depth features and the non-image features through the attention mechanism-based cross-modal fusion module, and generate fused features based on the attention weights. The diagnostic output module, connected to the multimodal fusion module, is used to process the fused features through a classifier and output the risk probability value of the subject to be diagnosed having early ovarian cancer.
2. The early ovarian cancer auxiliary diagnostic system based on multimodal fusion according to claim 1, characterized in that, The image feature extraction module includes: An image preprocessing unit is used to adjust the original ultrasound image to a preset resolution and perform pixel value normalization processing. The feature extraction unit has a built-in convolutional neural network, which is used to receive the preprocessed ultrasound image and output the feature vector generated by the global average pooling layer in the convolutional neural network as the image depth feature.
3. The early ovarian cancer auxiliary diagnostic system based on multimodal fusion according to claim 1, characterized in that, The non-image feature extraction module includes: The data stitching unit is used to stitch together the serum biomarker detection values and the clinical variable data to form a non-image input vector; The feature encoding unit has a built-in fully connected network, which sequentially includes a first fully connected layer, a first batch normalization layer, a first activation function layer, a second fully connected layer, a second batch normalization layer, a second activation function layer, a third fully connected layer, a third batch normalization layer, and a third activation function layer. It is used to process the non-image input vector layer by layer and output the non-image features.
4. The early ovarian cancer auxiliary diagnostic system based on multimodal fusion according to claim 1, characterized in that, The multimodal fusion module includes: Linear mapping unit, used to transform the first linear transformation matrix The image depth features Mapping to the target dimension space yields the mapped image features. And through the second linear transformation matrix The non-image features Mapping to the target dimension space yields the mapped non-image features. ,in , ; The attention calculation unit, connected to the linear mapping unit, is used to calculate attention weights. , In the formula, Let be the dimension value of the target dimension space. This is matrix multiplication, where T denotes matrix transpose; The feature fusion unit, connected to the attention calculation unit, is used to generate fused features. .
5. The early ovarian cancer auxiliary diagnostic system based on multimodal fusion according to claim 1, characterized in that, The multimodal fusion-based early ovarian cancer auxiliary diagnostic system also includes: An interpretability analysis module, connected to the image feature extraction module and the diagnostic output module, is used to generate a heatmap corresponding to the original ultrasound image through a gradient-weighted class activation mapping algorithm, and to calculate the marginal contribution of the serum biomarker detection values and each feature in the clinical variable data to the risk probability value through the SHAP analysis method.
6. The early ovarian cancer auxiliary diagnostic system based on multimodal fusion according to claim 1, characterized in that, The multimodal fusion-based early ovarian cancer auxiliary diagnostic system also includes: The risk restratification module, connected to the data acquisition module, is used to acquire the ovarian-adnexal report and data system classification of the subject to be diagnosed. When the classification is O-RADS Category 4, the malignancy risk value is calculated using the simple rule risk model of the International Ovarian Tumor Analysis Group. When the malignancy risk value is less than a preset threshold, the risk level of the subject to be diagnosed is adjusted from O-RADS Category 4 to O-RADS Category 3.
7. The early ovarian cancer auxiliary diagnostic system based on multimodal fusion according to claim 1, characterized in that, The multimodal fusion-based early ovarian cancer auxiliary diagnostic system also includes: The model calibration evaluation module, connected to the diagnostic output module, is used to evaluate the calibration degree of the risk probability value through the Hosmer-Lemeshow goodness-of-fit test, and to calculate the clinical net benefit under different threshold probabilities through decision curve analysis, so as to determine the optimal threshold probability. The formula for calculating the net clinical benefit is as follows: ; In the formula, TP represents the number of true positives, FP represents the number of false negatives, N represents the total sample size, and pt represents the threshold probability.
8. The early ovarian cancer auxiliary diagnostic system based on multimodal fusion according to claim 1, characterized in that, The image feature extraction module also includes: The model compression unit is used to perform lightweight processing on the pre-trained image feature extraction network to generate a compressed model. The compressed model has a smaller size than the original model and a higher inference speed.
9. The early ovarian cancer auxiliary diagnostic system based on multimodal fusion according to claim 1, characterized in that, The image feature extraction module also includes: The knowledge distillation unit is used to perform knowledge distillation using the pre-trained image feature extraction network as the teacher network and a convolutional neural network with fewer parameters as the student network to generate a lightweight student model. The inference latency of the lightweight student model is lower than that of the teacher network.
10. The early ovarian cancer auxiliary diagnostic system based on multimodal fusion according to claim 1, characterized in that, The multimodal fusion-based early ovarian cancer auxiliary diagnostic system also includes: The system integration interface is used to embed the image feature extraction module, the non-image feature extraction module, the multimodal fusion module, and the diagnostic output module into the hospital's image archiving and communication system. The system integration interface supports medical digital imaging and communication standards and is used to receive the original ultrasound image and output the risk probability value and the corresponding heat map.