Methods useful for determining the likelihood of endometriosis
By fusing MRI and TVUS data using an artificial intelligence model, the problem of low diagnostic accuracy for endometriosis in existing technologies has been solved. This provides a non-invasive and efficient diagnostic tool, improving diagnostic accuracy and reducing the waste of medical resources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Filing Date
- 2024-09-18
- Publication Date
- 2026-07-10
AI Technical Summary
Existing methods for diagnosing endometriosis using MRI and TVUS have low accuracy, and there is a lack of cost-effective, non-invasive alternatives to these diagnostic tools, leading to diagnostic delays and a waste of medical resources.
By employing an artificial intelligence model combined with MRI and TVUS data, and through cross-modal knowledge distillation and uncertainty estimation, MRI and TVUS modal information are fused to achieve multimodal analysis of signs of endometriosis, including detection of POD blockage, intestinal nodules, ovarian endometriotic cysts, and uterosacral ligament nodules.
It improves the accuracy of endometriosis diagnosis, provides a non-invasive and cost-effective diagnostic tool, reduces diagnostic delays and waste of medical resources, and avoids the need for surgery and anesthesia.
Smart Images

Figure CN122374843A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computer-implemented method and apparatus for analyzing images and video clips to detect at least one sign of endometriosis and thereby indicate the probability of diagnosing endometriosis. Background Technology
[0002] Endometriosis is a common gynecological condition affecting approximately 10% of women, girls, and those designated as female at birth. Pathologically, endometriosis refers to the growth of endometrial-like tissue outside the uterus, causing pain, prolonged menstrual bleeding, and infertility. Current treatment strategies do not cure endometriosis and primarily aim to alleviate clinical symptoms. Therefore, accurate early diagnosis of endometriosis is crucial for developing preventative treatment strategies and early intervention to slow disease progression.
[0003] Visualization of endometriosis during laparoscopy has historically been the only diagnostic method. This approach is costly, risky, and has limited accessibility given the high prevalence of endometriosis. In Australia, the average time to diagnosis is 6.4 years, resulting in impaired quality of life for many women with endometriosis, whose actual symptoms are often questioned by healthcare professionals and society. Endometriosis patients utilize healthcare resources more frequently in the 10 years prior to diagnosis than after. A cost-effective, non-invasive diagnostic test would reduce diagnostic delays and the time to initiate appropriate, targeted treatment, thereby lowering healthcare costs and potential surgical needs.
[0004] In most women and those designated as female at birth, menstrual fluid and shed endometrial tissue retrograde through the fallopian tubes and accumulate in the Pouch of Douglas. In approximately 10% of these women, endometrial tissue implants and develops into inflammatory endometriosis lesions, which grow and shed in response to hormonal changes during the menstrual cycle (Samson's theory). This can lead to inflammation, menstrual cramps, chronic pelvic pain, complicated pain syndrome, and infertility.
[0005] One method for diagnosing endometriosis is to detect the closure of the rectouterine pouch (POD). POD closure is a strong diagnostic marker for conditions such as endometriosis. The POD (also known as the rectouterine pouch or posterior pouch) is a narrow space located between the rectum and the posterior wall of the uterus.
[0006] POD closure can be visualized using magnetic resonance imaging (MRI) and / or transvaginal ultrasound (TVUS) scans. Our recent results indicate that TVUS is more accurate than MRI in classifying POD closure.
[0007] On MRI, POD closure can appear as pelvic nodules or plaque-like lesions and adhesions visible on T2-weighted images (see [link to MRI]). Figure 1C Positive POD blockade in the middle, with Figure 1A (Compared to negative cases in the middle).
[0008] When using TVUS, the dynamic "sliding sign" technique can identify POD closure when the cervix is gently pressed. A positive classification is indicated if the anterior rectal wall slides freely across the uterus; this is a normal finding. Figure 1B If movement is restricted or constrained, the classification is negative, indicating POD closure. Figure 1D ).
[0009] In clinical practice, the expertise required to diagnose endometriosis using TVUS and MRI is not widely available. The accuracy rate for diagnosing POD closure using TVUS is between 75% and 95%, while for MRI, the accuracy rate (endometriosis detection performed by human experts) is between 40% and 69%.
[0010] In current practice, patients typically undergo only one of these two examinations (MRI or TVUS), and especially in TVUS, the ability of sonographers to induce slippage varies. Because the accuracy of using MRI and TVUS independently is relatively low, a better alternative diagnostic tool is needed that can produce a diagnosis with higher independent accuracy or at least similar accuracy.
[0011] Any reference in this document to a patent document or other matter given as prior art shall not be construed as an admission that the document or matter is known, or that the information contained therein is common general knowledge on the priority date of any claim. Summary of the Invention
[0012] According to the present invention, apparatus, methods, computer programs, and computer-readable media are provided as described in the appended claims. Other features of the invention will be apparent from the dependent claims and the description below.
[0013] In one aspect, a computer-implemented method is provided for determining the likelihood of endometriosis, the method comprising: acquiring MRI or TVUS results of a subject; inputting the MRI or TVUS results into an artificial intelligence (AI) model and / or a learnable cross-modal knowledge distillation (LCKD) model; detecting at least one sign of endometriosis through the AI model and / or the LCKD model; and outputting a detection result of the at least one sign of endometriosis; wherein the at least one sign of endometriosis is at least one of a Douglas pouch (POD) closure, an intestinal nodule, an ovarian endometriotic cyst, and an uterosacral ligament (USL) nodule. Outputting a detection result of the at least one sign of endometriosis (e.g., a POD detection) results leads to a determination of the likelihood of endometriosis. Those skilled in the art will understand that the method may include acquiring MRI and / or TVUS results of a subject; inputting the MRI and / or TVUS results into an artificial intelligence (AI) model and / or a learnable cross-modal knowledge distillation (LCKD) model. Therefore, two imaging modalities can be used simultaneously in this method.
[0014] The AI model may include: extracting at least one MRI embedding from an endometriosis MRI dataset in an MRI stream using a magnetic resonance imaging (MRI) feature extractor; extracting at least one TVUS embedding from a TVUS dataset in a TVUS stream using a transvaginal ultrasound (TVUS) feature extractor; measuring uncertainty in the MRI stream; measuring uncertainty in the TVUS stream; performing cross-modal fusion of MRI and TVUS embeddings by multiplying at least one MRI embedding by the uncertainty measured in the TVUS stream and multiplying the at least one TVUS embedding by the uncertainty measured in the MRI stream; cascading the cross-modal fused MRI and TVUS embeddings; and detecting at least one sign of endometriosis.
[0015] The AI model may also include: pre-training an MRI feature extractor based on an MRI dataset; and / or pre-training a TVUS feature extractor based on a TVUS dataset.
[0016] The AI model may also include: matching MRI images with positive / negative POD closed labels in the endometriosis MRI dataset with TVUS videos with negative / positive slip signs in the TVUS dataset. The AI model may also include: matching MRI images with positive / negative endometriosis sign labels in the endometriosis MRI dataset with TVUS videos with negative / positive endometriosis signs in the TVUS dataset.
[0017] AI models may include passing cascaded, cross-modal fused MRI and TVUS embeddings to linear layers for prediction.
[0018] Endometriosis MRI dataset and TVUS dataset can be used unpaired for model optimization.
[0019] An endometriosis MRI dataset may include multiple MRI images, at least one of which may include a labeled POD-closed case. An endometriosis MRI dataset may include multiple MRI images, at least one of which includes at least one of a labeled POD-closed case, a labeled intestinal nodule case, a labeled ovarian endometriotic cyst case, and a labeled uterosacral ligament nodule case.
[0020] The TVUS dataset may include multiple sliding sign TVUS video clips, wherein at least one of the multiple TVUS video clips may include at least one of labeled POD closed cases, labeled intestinal nodule cases, labeled ovarian endometriotic cyst cases, and labeled uterosacral ligament nodule cases.
[0021] MRI datasets can be unlabeled and include multiple MRI images of the female pelvic region.
[0022] The AI model may include: measuring channel-level uncertainty in an MRI stream via a first random network prediction (RNP) module; and / or measuring channel-level uncertainty in a TVUS stream via a second RNP module.
[0023] Both the first and second RNP modules can include fixed-weight random networks and learnable predictive networks.
[0024] The AI model may include training the first and second RNP modules based on minimizing the loss function calculated from the difference between the outputs of the prediction network and the random network.
[0025] The AI model may include: outputting a single scalar for each mode to represent uncertainty through a first RNP module; and outputting a single scalar for each mode to represent uncertainty through a second RNP module.
[0026] Obtaining MRI or TVUS results from a subject may involve receiving MRI and / or TVUS results by a medical professional or the subject.
[0027] MRI results (e.g., MRI data) can be MRI images. TVUS results (e.g., TVUS data) can be TVUS video clips of the sliding sign.
[0028] The TVUS feature extractor can be a convolutional neural network.
[0029] MRI feature extractors can be 3D vision transformers.
[0030] Testing for at least one of the following—POD block, intestinal nodules, ovarian endometriotic cysts, and uterosacral ligament nodules—can be used to diagnose endometriosis.
[0031] LCKD models may include selecting teacher modalities through a validation process.
[0032] LCKD models may include distilling cross-modal knowledge between each pair of available teacher and student modalities using a loss function.
[0033] LCKD models may include training teacher and student modal pairs, where during training, the student modal simulates the behavior and predictions of the teacher modal.
[0034] On the other hand, a data processing apparatus is provided, which includes means for performing the above-described method.
[0035] On the other hand, a computer program is provided that includes instructions which, when executed by a computer, cause the computer to perform the methods described above.
[0036] On the other hand, a computer-readable storage medium is provided, which includes instructions that, when executed by a computer, cause the computer to perform the methods described above.
[0037] As those skilled in the art will appreciate, this technology can be embodied in a method, apparatus, computer program product, or computer-readable storage medium. Therefore, this technology can take the form of a completely hardware implementation, a completely software implementation, or an implementation combining software and hardware aspects.
[0038] Furthermore, this technology can take the form of a computer program product embedded in a computer-readable medium, on which computer-readable program code is embedded. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable medium can be, for example, but not limited to, any physical device or material capable of storing digital data, such as a hard disk, CD-ROM, flash drive, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing.
[0039] The computer program code used to perform the operations of this technology can be written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. Code components can be embodied as procedures, methods, or similar forms, and can include sub-components that can take the form of instructions or sequences of instructions at any level of abstraction, from direct machine instructions of the native instruction set to high-level compiled or interpreted language constructs.
[0040] Examples of this technology also provide a non-transient data carrier that carries code, which, when implemented on a processor, causes the processor to execute any of the methods described herein.
[0041] This technology can provide processor control code to implement the methods described above, for example, on a general-purpose computer system or a digital signal processor (DSP). The technology also provides processor control code carried by a carrier, which, when executed, implements any of the methods described above, particularly on a non-transient data carrier. This code can be provided on a carrier such as a disk, microprocessor, CD-ROM, or DVD-ROM; on a programmable memory such as non-volatile memory (e.g., flash memory) or read-only memory (firmware); or on a data carrier such as an optical or electrical signal carrier. The code (and / or data) used to implement the technical embodiments described herein can include source code, object code, or executable code in a conventional programming language (interpreted or compiled), such as Python, C, or assembly code; code for setting up or controlling an ASIC (Application-Specific Integrated Circuit) or FPGA (Field-Programmable Gate Array); or code for a hardware description language such as Verilog (RTM) or VHDL (Very High Speed Integrated Circuit Hardware Description Language). As those skilled in the art will appreciate, such code and / or data can be distributed among multiple coupled components that communicate with each other. The technology may include a controller that includes a microprocessor, working memory, and program memory coupled to one or more of the components of the device.
[0042] Those skilled in the art will also appreciate that all or part of the logical methods according to the present technical embodiments can be suitably implemented in a logic device including logic elements for performing the above-described method steps, and such logic elements may include components such as logic gates in a programmable logic array or application-specific integrated circuit. Such logical arrangements can also be further implemented in enabling elements for temporarily or permanently establishing logical structures in said array or circuit, using, for example, a virtual hardware description language, which can be stored and transmitted using a fixed or transmissible carrier medium.
[0043] This technology can be implemented in the form of a data carrier containing functional data, which includes a functional computer data structure. When the functional computer data structure is loaded into a computer system or network and operated by it, the computer system is able to perform all the steps of the above method.
[0044] The above methods can be executed, in whole or in part, on a device (i.e., an electronic device). The methods can be processed by a dedicated AI processor designed in a hardware architecture specifically for processing AI models. The AI model can be obtained through training. Here, "obtained through training" means training a basic AI model using multiple training datasets via a training algorithm to obtain predefined operating rules or an AI model configured to perform desired features (or objectives). The AI model may include multiple neural network layers.
[0045] As described above, this technology can be implemented using an AI model. AI-related functions can be executed via non-volatile memory, volatile memory, and a processor. The processor may include one or more processors. These processors control the processing of input data according to predefined operating rules or artificial intelligence (AI) models stored in the non-volatile memory and volatile memory. The predefined operating rules or AI models are provided through training or learning. Here, "provided through learning" means generating predefined operating rules or AI models with desired characteristics by applying a learning algorithm to multiple learning data sets. Learning can be performed within the device itself that performs the AI according to the implementation scheme, and / or can be implemented via a separate server / system. Attached Figure Description
[0046] Figure 1A An example of an MRI image with a negative POD closure (normal) is shown; Figure 1B An example of a TVUS image with a positive slip sign (normal) is shown; Figure 1C An example of an MRI image with a positive POD closure is shown; Figure 1D An example of a TVUS image with a negative slip sign is shown; Figure 2A The diagram shows the principle of MRI feature extractor pre-training; Figure 2B The diagram shows the principle of pre-training the TVUS feature extractor; Figure 2C A schematic diagram representing multimodal fusion with uncertain estimation is shown; Figure 3A flowchart illustrating a method for obtaining an AI model according to this technology is shown; Figure 4 A schematic diagram representing an exemplary learnable cross-modal knowledge distillation (LCKD) model is shown; Figure 5 A flowchart showing the method for obtaining the LCKD model is displayed; Figure 6 This demonstrates a method for detecting at least one sign of endometriosis according to the present technology; Figure 7 The results of case studies applying the techniques disclosed herein are shown; and Figure 8 A schematic diagram of a computing device suitable for implementing this technical method is shown. Detailed Implementation
[0047] In summary, this disclosure describes a computer-implemented method for the diagnosis of endometriosis that provides a solution to the modality missing problem by fusing TVUS and MRI modalities with estimation uncertainty using unpaired datasets. Advantageously, the method of this technique can use TVUS or MRI data to accurately classify at least one sign of endometriosis (e.g., POD closure) and thereby indicate an endometriosis diagnosis in a patient. The classification can be binary (yes / no).
[0048] The applicant recognizes that for the diagnosis of endometriosis based on signs of endometriosis (such as POD block detection), multimodal analysis combining TVUS and MRI is superior. MRI provides spatial information and allows for precise characterization of tissue, but it cannot capture dynamic information like TVUS. Therefore, combining MRI and TVUS allows for the identification of unique and complementary biomarkers, which can improve the diagnostic accuracy of signs of endometriosis and thus indicate the likelihood of endometriosis. However, patients typically only have either MRI or TVUS results.
[0049] In summary, this disclosure describes an artificial intelligence (AI) model trained using unpaired MRI and TVUS data, and preferably tested using only one of the modalities. In some embodiments, the accuracy of the results obtained by the AI model using MRI or TVUS results can be as high as the accuracy of using the best-performing modality (e.g., TVUS) alone. This bimodal AI model achieves this goal through training in two phases: 1) pre-training a TVUS feature extractor for the TVUS stream to detect signs of endometriosis (e.g., POD closure) from TVUS data, and pre-training an MRI feature extractor for the MRI stream using a large-scale unlabeled pelvic 3D MRI roll using a 3D masked autoencoder (MAE); and 2) combining the MRI and TVUS modalities through multimodal fusion with uncertainty estimation for each modality. Those skilled in the art will understand that an LCKD model can be used to improve the accuracy of this method.
[0050] Advantageously, this technique provides a multimodal model trained to combine knowledge from TVUS and MRI using unpaired data to detect signs of endometriosis (e.g., POD closure) even in the absence of a specific modality. The implementation of this technique employs a multimodal learning algorithm based on cross-modal random network prediction, which improves upon multimodal segmentation requiring all input modalities at test time to handle multimodal POD closure classification with missing modalities at test time. In cross-modal random network prediction, this is achieved by introducing channel-level uncertainty estimation, thereby encouraging the model to focus more on another modality when features from one modality are uncertain. A further benefit of this approach is that the diagnosis of endometriosis can be achieved non-invasively, i.e., without surgery or anesthesia.
[0051] The implementation of this technology utilizes a multimodal learning algorithm, namely Cross-Modal Random Network Prediction (CRNP, hereinafter referred to as the AI model), which improves upon multimodal segmentation requiring all input modalities at test time to handle multimodal classification of endometriosis signs (e.g., POD closure) with missing modalities at test time. It introduces channel-level uncertainty estimation to prioritize one modality when another is uncertain. In other examples, the LCKD model can be combined with Cross-Modal Random Network Prediction. Those skilled in the art will understand that both the LCKD model and the Cross-Modal Random Network Prediction model are artificial intelligence models, but for simplicity, the Cross-Modal Random Network Prediction model is referred to as the "AI model," and the LCKD model is referred to as the "LCKD model."
[0052] Those skilled in the art will understand that multimodal detection of signs of endometriosis (e.g., POD block) in this technology is important because, generally, multimodal methods can achieve better performance than their single-modal counterparts.
[0053] Those skilled in the art will also understand that the use of POD blockade is one method / indication for diagnosing endometriosis. There may be other indications that can be used to determine endometriosis. Those skilled in the art will understand that modifications can be made to the methods and systems of this disclosure to include other indications for determining endometriosis.
[0054] Signs of endometriosis can be at least one of the following: POD blockage, intestinal nodules, ovarian endometriotic cysts, and uterosacral ligament (USL) nodules. In other words, signs of endometriosis can be POD blockage, and / or intestinal nodules, and / or ovarian endometriotic cysts, and / or uterosacral ligament (USL) nodules. All of these signs of endometriosis can be identified by TVUS and MRI. Therefore, the AI model and LCKD model of this technology can be trained to detect at least one sign of endometriosis.
[0055] In the context of endometriosis, intestinal nodules refer to areas where endometrial-like tissue grows on or within the intestinal wall, forming small masses or lumps. On TVUS, intestinal nodules may appear as dark (hypoechoic) plaques, showing thickening or invasion of the serosa of the sigmoid colon, colon, or rectum, or protruding into the intestinal lumen like mushrooms. On MRI images, intestinal nodules may appear thickened or contracted due to fibrosis. Multifocal lesions (within 2 cm of the primary lesion) can be distinguished from multisite intestinal endometriosis (more than 2 cm apart), which is crucial for planning surgical resection.
[0056] Advantageously, using intestinal nodules to detect endometriosis can explain symptoms such as intestinal pain, constipation, diarrhea, nausea, and loss of appetite. Furthermore, it can reduce the time and resources wasted by gastroenterologists using colonoscopy and endoscopy to examine the gastrointestinal system when referral to gynecology is required. Additionally, it can avoid misdiagnosis of irritable bowel syndrome and allow for timely initiation of targeted endometriosis treatment.
[0057] Ovarian endometriotic cysts (also known as "chocolate cysts") are a special type of cyst that typically forms on the ovary or between the ovary and the peritoneal wall due to the repeated growth and breakdown of endometriotic tissue. These cysts are usually filled with old, dark blood. TVUS identifies endometriotic cysts with high sensitivity and specificity by detecting them as cysts with a uniform appearance and diffuse low-level echogenicity. Similarly, MRI can identify ovarian endometriotic cysts with high accuracy as ovarian cysts that appear as high signal (bright white) on T1-weighted images and heterogeneous low to intermediate signal (T2 shading) on T2-weighted images.
[0058] Advantageously, detecting endometriosis through ovarian endometriotic cysts allows for consideration of egg freezing. This is especially important before surgery, as surgical removal of the cysts may reduce the number of eggs.
[0059] The uterosacral ligament nodules (located posterior to the uterus) and / or the uterine cylinder (the attachment points of these ligaments in the lower posterior part of the uterus) are the most common sites of endometriosis. On MRI images, uterosacral ligament nodules typically appear as low signal intensity and nodular lesions on T2-weighted images, although thickening or contraction of the ligaments can raise suspicion. On T1-weighted images, high signal intensity white nodules can be identified. MRI is less sensitive than TVUS in detecting USL nodules.
[0060] Large nodules on the uterosacral ligaments can compress the ureters and lead to obstructive renal failure. These nodules can also cause inflammation, affecting bowel function and fertility. Therefore, advantageously, detecting endometriosis through uterosacral ligament nodules may influence decisions regarding medication to suppress menstruation, surgical procedures, and which symptoms patients should be aware of and monitor for (such as kidney and bladder pain and recurrent urinary tract infections).
[0061] Create an AI model for detecting at least one sign of endometriosis (e.g., POD blockage) and for endometriosis diagnosis.
[0062] Figure 2A The diagram shows the principle of pre-training an MRI feature extractor. Figure 2B The diagram shows the principle of pre-training the TVUS feature extractor. Figure 2C The diagram shows the principle of multimodal fusion with uncertain estimates. Figure 3 A flowchart illustrating a method for obtaining an AI model according to this technology is shown. The method for creating the AI model is described in... Figure 2A , 2B The AI model is schematically depicted in 2C. The AI model at least performs... Figure 3 The methods and steps described. Figure 6A method for detecting at least one sign of endometriosis (e.g., POD blockage) is shown, which incorporates an AI model of this technology.
[0063] Methods 300, 400, and 500 of this technology are computer-implemented. Methods 300, 400, and 500 are suitable for detecting at least one sign of endometriosis (e.g., POD block and / or intestinal nodules and / or ovarian endometriotic cysts and / or uterosacral ligament (USL) nodules) and thereby outputting a score indicating the probability of an endometriosis diagnosis. The score can be a "yes" or "no" judgment regarding endometriosis, or a probability score. Advantageously, methods 300, 400, and 500 can be performed outside the patient's body. Methods 300, 400, and 500 can be powered by a computer processor (e.g., ...). Figure 8 This can be achieved using one of the methods shown, and the processor can be integrated into diagnostic devices used in clinical settings, such as TVUS or MRI consoles.
[0064] In summary, the AI model created by method 300 consists of two main stages: 1) self-supervised pre-training of an MRI feature extractor on an unlabeled MRI dataset and training a sliding sign classifier using a TVUS dataset (Figure 2(A), (B)); and 2) fine-tuning of the TVUS and MRI feature extractors using unpaired TVUS and MRI data via a two-stream cross-modal network to produce an output as a probabilistic indicator of signs of endometriosis (e.g., POD closure and / or intestinal nodules and / or ovarian endometriotic cysts and / or uterosacral ligament (USL) nodules), which indicates an endometriosis diagnosis (Figure 2(C)). As shown in Figure 2(C), the method of this technique uses a bimodal model to estimate modal uncertainty so as to allow priority to be given to one modality when the representation of the other modality has high uncertainty.
[0065] The AI and LCKD models are described as being trained on a dataset where the sign of endometriosis is a closed POD (Posterior Ovary Diagnosis). However, it is understood that the AI and LCKD models can be trained (additionally or specifically) on other signs of endometriosis. Signs of endometriosis can be at least one of the following: closed POD, intestinal nodules, ovarian endometriotic cysts, and uterosacral ligament (USL) nodules. All of these signs of endometriosis can be identified by TVUS and MRI. Therefore, the AI and LCKD models of this technique can be trained to detect at least one sign of endometriosis.
[0066] In some examples, the AI model and / or LCKD model can be trained for the first sign of endometriosis (e.g., POD closure). Then, the AI model and / or LCKD model can be trained for the second sign of endometriosis (e.g., intestinal nodules). In some examples, as described in more detail below, the AI model and / or LCKD model can be trained simultaneously for two or more seed signs of endometriosis (e.g., POD closure and intestinal nodules).
[0067] This represents an MRI dataset for endometriosis, containing N T2 SPACE MRI images. (Where H, W, and D are height, width, and depth), and are labeled with N one-hot vectors representing the closed label of the POD. The endometriosis MRI dataset includes multiple T2-space MRI and / or T1-space MRI images. T2-space MRI images may be pelvic scans of women. The endometriosis MRI dataset includes multiple MRI images with identified POD closure cases. Those skilled in the art will understand that focusing the images in the endometriosis MRI dataset on the diagnostic area of POD closure near the uterus is preferred. The endometriosis MRI dataset includes 3D images. The endometriosis MRI dataset can be labeled by human annotators.
[0068] This refers to the TVUS dataset. The TVUS dataset includes multiple sliding sign TVUS results, containing data of size [size missing]. of M video segments of one frame And the corresponding sliding heat label The TVUS dataset includes multiple video clips of the "slipping sign." These video clips include multiple cases of negative slipping sign (POD closed). Each slipping sign video includes at least one frame. Each frame can be labeled. The TVUS dataset can be labeled by human annotators.
[0069] This represents the MRI dataset used for training. The MRI dataset includes multiple MRI scan images used for self-supervised pre-training, containing K T2 MRI volumes. , where K >> N and K >> M. The MRI dataset used for self-supervised pre-training may include T2 MRI volumes from women with endometriosis-like symptoms. However, those skilled in the art will understand that The dataset may include at least one MRI image that may contain signs of other diseases, and the examined area may not necessarily be aligned with the diagnostic area for endometriosis. MRI datasets for self-supervised pre-training. It is unlabeled. Preferably, the MRI dataset includes at least five thousand images.
[0070] The endometriosis MRI dataset and the slippery sign TVUS dataset are unpaired. The TVUS and endometriosis MRI datasets are matched only based on their POD closed labels. Accordingly, method 300 of this technique includes the step of constructing unpaired datasets. Specifically, by using each MRI volume... With randomly selected TVUS videos Matching is performed to construct the unpaired dataset, where In other words, MRI rolls with positive / negative POD closures were paired with TVUS videos with negative / positive slip signs.
[0071] Although not shown in the accompanying figures, method 300 may include acquiring an endometriosis MRI dataset. Method 300 may also include acquiring a TVUS dataset. Method 300 may also include acquiring an MRI dataset. The datasets may be obtained from one or more third-party providers. The AI model (and the LCKD model) may be pre-trained and trained outside the computing device on which the AI model is implemented. For example, a user may download and / or install the AI model on their computing device or on a TVUS or MRI console. Alternatively, all steps of methods 300, 400, and 500 may be implemented on the same computing device or TVUS / MRI console. For example, the AI model may be incorporated into a software product that can be distributed to end-user devices. The AI model may also be stored in the cloud. The end-user may access the AI model in the cloud to upload their TVUS or MRI results to the AI model, which provides indicators of the likelihood that the TVUS or MRI results indicate signs of endometriosis (e.g., POD closure), an indicator of endometriosis diagnosis. At least all of the above configurations applicable to the AI model also apply to the LCKD model.
[0072] During the pre-training phase, method 300 of this technique can use a dataset. To optimize the mask reconstruction of MRI volumes, and using the dataset To optimize the sliding sign classification of TVUS sequences.
[0073] Method 300 includes step S302: pre-training an MRI feature extractor.
[0074] For pre-training the MRI feature extractor, method 300 extends the self-supervised masked autoencoder (MAE) to 3D and trains it using an MRI dataset. The MRI dataset is an unlabeled T2 MRI dataset. Masked autoencoders (MAEs) are self-supervised learning methods for computer vision that mask random patches of an input image and reconstruct missing pixels. Advantageously, MAEs are scalable, efficient, and effective, and can achieve state-of-the-art results in a variety of image classification and transfer learning tasks.
[0075] The 3D-MAE network can be a 3D Vision Transformer (3D ViT), which is optimized to pass through, for example... Figure 2A The asymmetric encoder-decoder architecture shown reconstructs the masked MRI volume. More specifically, each 3D MRI image is divided into 8×8×8 non-overlapping cubes, and a portion of the cube volume (e.g., 50%) is randomly masked. Those skilled in the art will understand that each 3D MRI image can be divided into non-overlapping cubes of different sizes. Those skilled in the art will also understand that other similar Transformers can be used. For example, the SwinTransformer (Swin Transformer) can be used.
[0076] Then, the remaining cubes are input into the encoder. For example, the encoder could be a 3D ViT with 12 Transformer blocks, embedded via a linear projection of dimension E = 768, with added positional embedding. Those skilled in the art will understand that the encoder can have more or fewer than 12 Transformer blocks, and the linear projection of dimension E can be greater than or less than 768. These specific parameters can be selected or fine-tuned depending on the available GPUs and / or the size of the dataset.
[0077] Then, the shallow / lightweight decoder (For example, represented by a 3DViT with 8 Transformer blocks) Visible tokens from the encoder and learnable mask tokens are applied to reconstruct the volume. The encoder and decoder are optimized by minimizing the mean squared error (MSE) between the reconstructed volume and the original volume in voxel space, as follows:
[0078] (Equation 1)
[0079] in, K This refers to the size of the unlabeled MRI dataset. L This represents the total number of cubes that are masked. Indicates that the original volume was The voxel values of the mask cube, and Indicates by The corresponding generated reconstructed volume Voxel values. After pre-training, the encoder is used as an MRI feature extractor, thereby generating embeddings for the downstream POD closed classification task. .
[0080] 3D MAE networks can Train from scratch on the MRI dataset.
[0081] Understandably, step S302, which pre-trains the MRI feature extractor, may further include fine-tuning the MRI feature extractor for the region of interest (ROI).
[0082] When using MRI to detect signs of endometriosis, clinicians typically focus on the area surrounding the uterus within the pelvic cavity. Therefore, the region of interest can be a segment of the uterus. However, standard pelvic scans often include extra-pelvic areas such as bone, muscle, and adipose tissue. While these additional structures provide valuable background information for a complete anatomical assessment, they are not directly related to endometriosis and may add unnecessary noise and complexity to the imaging data. Therefore, cropping irrelevant areas from the raw MRI roll allows the model to focus on the most critical areas, advantageously reducing computational requirements and improving the efficiency and accuracy of the diagnostic tool.
[0083] To fine-tune the MRI feature extractor for regions of interest, a segmentation head (defined as...) is integrated after the encoder. ), and add a category header (defined as This utilizes a Visual Transformer (ViT) model that has been pre-trained using a Masked Autoencoder (MAE).
[0084] The segmentation model can be optimized during training by simultaneously minimizing the classification and segmentation loss functions, as described below:
[0085] (Equation 1B)
[0086] in, Represents cross-entropy loss, Represents the binary cross-entropy loss, where Ω represents the image grid. ( Indicates image location The segmentation label at the location, and Indicates in Segmentation prediction at the location.
[0087] For example, the ROIs for detecting POD closures and intestinal nodules are both located in the posterior uterus and rectal regions. Therefore, the ROI can be defined by applying a predefined padding around the predicted uterine segmentation mask. This padding amount can be empirically determined based on the maximum distance from the front of the uterus to the back of the rectum, and the maximum distance from the coccyx to the vaginal opening.
[0088] Method 300 includes step S304: pre-training a TVUS feature extractor.
[0089] In this example, the TVUS feature extractor The probabilistic simplex (RSM) is represented by a convolutional neural network (CNN). For example, this CNN could be a ResNet(2+1)D. In one example, a ResNet(2+1)D might contain 18 R(2+1) convolutional layers, with batch normalization performed after convolution and before Rectified Linear Unit (ReLU) activation. In some implementations, the ResNet(2+1)D can be pre-trained on the Kinetics-400 dataset and trained for multiple epochs on the TVUS dataset.
[0090] The TVUS feature extractor can be pre-trained to minimize the sliding characteristic TVUS dataset. The cross-entropy (CE) loss for mid-sliding feature annotation is shown below:
[0091] (Equation 2)
[0092] Following this pre-training, at least one TVUS embedding for each sample is extracted from the Adaptive Average Pooling layer via 18 R(2+1) convolutional layers, where To form vectors Therefore, the feature extractor for the TVUS stream generates at least one TVUS embedding for the downstream POD closed classification task. The feature extractor for the TVUS stream can be a CNN.
[0093] Adaptive average pooling may outperform traditional pooling operations in image classification tasks because it can reduce overfitting and improve the model's generalization ability.
[0094] Method 300 includes step S306: constructing an unpaired dataset. Specifically, by... With randomly selected TVUS videos Matching is performed to construct the unpaired dataset, where In other words, in S306, MRI volumes / images with positive / negative POD closures are paired with TVUS videos with corresponding negative / positive slip signs. Constructing an unpaired dataset may include matching MRI images with positive / negative POD closure labels from the endometriosis MRI dataset with randomly selected TVUS videos with corresponding negative / positive slip signs from the TVUS dataset. Therefore, advantageously, better POD classification performance can be achieved on the endometriosis MRI dataset (and vice versa) using the TVUS dataset. More generally, in S306, MRI volumes / images with positive / negative endometriosis signs are paired with TVUS videos with corresponding positive / negative endometriosis signs. Constructing an unpaired dataset may include matching MRI images with positive / negative endometriosis sign labels from the endometriosis MRI dataset with randomly selected TVUS videos with corresponding negative / positive endometriosis signs from the TVUS dataset.
[0095] Understandably, in some examples, more than one type of endometriosis sign may be combined and used during training. In other words, S306's construction of unpaired datasets may include multi-label mix-ups. For example, training data may include POD closed MRI images and TVUS clips (as described above), as well as intestinal nodule (BN) MRI images and TVUS clips.
[0096] Essentially, in some examples, an MRI roll simultaneously labeled with POD closure and intestinal nodules can be paired with two separate TVUS videos: one based on the POD label and the other based on the intestinal nodule label. Since TVUS datasets may not contain any samples with both POD closure and BN positive annotations, a multi-label mixing strategy can be used. This technique integrates two TVUS scans to create... This serves as the final input for training the TVUS encoder, as shown below:
[0097] (Equation 2B)
[0098] in, From the beta distribution Beta ( , Randomly obtained from ) It is a predefined hyperparameter that determines the mixing ratio.
[0099] Although TVUS and MRI may share the same labels, this technique allows for empirical mixing of labels. Therefore, advantageously, this strategy enables the model's TVUS stream to handle the classification of multiple signs using a single encoder, rather than creating a separate encoder for each sign. Advantageously, maintaining a shared encoder helps the TVUS encoder learn more general feature representations, while also significantly reducing the overall complexity of the model.
[0100] Once trained, the model can perform classification using any combination of available data modalities.
[0101] Method 300 may include: S306, using a learnable cross-modal knowledge distillation (LCKD) model. The use of an LCKD model enhances... Figure 3 The performance of the entire AI model 300 is described. However, those skilled in the art will understand that the AI model 300 does not necessarily require the use of the LCKD model to function. Those skilled in the art will understand that the LCKD model / method 500 can also be used alone (without performing the other steps in method 300) to identify POD closure or any other signs of endometriosis, and thus determine the probability of an endometriosis diagnosis.
[0102] The creation of LCKD models in Figure 4 and Figure 5 Further details are provided below.
[0103] Method 300 may include: S310, extracting at least one MRI embedding from an endometriosis MRI dataset in an MRI stream by an MRI feature extractor.
[0104] This step involves initializing the MRI feature extractor in the MRI stream. The MRI stream uses an MRI encoder from a pre-trained MRI feature extractor, where… .
[0105] Method 300 may include: S312, whereby a TVUS feature extractor extracts at least one TVUS embedding from the TVUS dataset in the TVUS stream.
[0106] This step involves initializing a TVUS feature extraction embedding extractor from a pre-trained TVUS dataset within the TVUS stream. The TVUS stream uses an embedding extractor derived from a pre-trained TVUS model, where... For example, useful features can be extracted as vectors from each TVUS video using R(2+1)D. Further computation can then be performed on the extracted vectors.
[0107] Those skilled in the art will understand that the multimodal analysis of endometriosis MRI and TVUS datasets aims to achieve accurate classification between the two modalities, even when one modality is missing from the input of the AI model. This missing modality analysis capability is extremely beneficial because endometriosis patients typically only have one of the two examination results (MRI or TVUS) available for testing. This also allows the training of this technique to use unpaired TVUS and MRI datasets, which are matched only based on their endometriosis sign labels (e.g., POD closed labels and / or any other endometriosis sign labels).
[0108] Method 300 may also include: S314, measuring uncertainty in the MRI flow.
[0109] Method 300 may also include: S316, measuring the uncertainty in the TVUS stream.
[0110] Measuring the uncertainty in each stream can be performed, for example, via a Random Network Prediction (RNP) module. Therefore, the method may include initializing at least two RNP modules to determine the uncertainty in the TVUS and MRI streams. Thus, the feature extractors for both the MRI and TVUS streams are followed by RNP modules, denoted as follows: and This is used to measure the uncertainty of each stream. A stream refers to a data stream that is analyzed and processed in real-time or near real-time according to the techniques disclosed herein. For example, a TVUS stream corresponds to a data stream that processes TVUS results, while an MRI stream corresponds to a data stream that processes MRI results.
[0111] It is represented as Fixed-weight random networks (similarly, have ), and learnable prediction networks (Similarly, ).
[0112] Advantageously, learnable predictive networks can have the same output space as fixed-weight random networks but with different structures, where the capacity of the learnable predictive network is smaller than that of the fixed-weight random network to prevent potentially trivial solutions. More specifically, the fixed-weight random network can be a neural network with three one-dimensional convolutional layers. The LeakyReLU activation function can be applied after the first and second convolutional layers. The learnable predictive network can be a neural network with two one-dimensional convolutional layers. The LeakyReLU activation function can be applied after the first convolutional layer.
[0113] The method of this technique may further include training the RNP module based on minimizing a loss function calculated from the difference between the outputs of the predictive network and the random network (within the same stream). The predictive network fits the outputs of the random network better for samples in dense regions of the feature space (i.e., low uncertainty); however, it fits poorly for samples in sparse regions (i.e., high uncertainty). By minimizing the mean squared error (MSE) between the outputs of the predictive network and the random network, the predictive network can learn the uncertainty measurement function.
[0114] RNP is trained based on minimizing a loss function calculated from the difference between the outputs of the predicted network and the random network, as shown below:
[0115] (Equation 3)
[0116] In Equation 3, embeddings in dense regions of the embedding space yield a smaller loss. For example, in a binary classification problem, considering that the embedding space is in a two-dimensional orthogonal coordinate system, easily classifiable samples should be densely clustered at the far ends of the first (1,1) and third (-1,-1) quadrants, while difficult-to-classify samples should be sparsely scattered at (0,0). Therefore, embeddings in dense regions are more advantageous for MRI. and The mean squared error (MSE) between them has low prediction uncertainty; for TVUS, in and The mean squared error between streams has low prediction uncertainty. On the other hand, embeddings in sparse regions (e.g., embeddings from blank input data representing missing modalities) have high prediction uncertainty. The uncertainty in each stream can be represented as a scalar value.
[0117] Method 300 further includes: S318, performing cross-modal fusion of MRI and TVUS embeddings. This step includes performing cross-modal fusion of MRI and TVUS embeddings by multiplying at least one MRI embedding by an uncertainty measured in the TVUS stream and at least one TVUS embedding by an uncertainty measured in the MRI stream.
[0118] In the multimodal fusion phase, this method deals with a multimodal classification task that relies on multiple data modalities (MRI and TVUS), where the domain shifts between these modalities are greater than those between different MRI modalities. Furthermore, the training data for the AI model is unpaired. Moreover, the testing of the AI model may preferably be performed in the absence of modalities. Therefore, method 300 of this technique may include generating a single scalar for each modality to represent modal uncertainty, i.e., generating for the MRI stream. Targeting TVUS stream generation It is calculated based on the output of the RNP module to represent the modal uncertainty, as shown below:
[0119] (Equation 4)
[0120] The method of this technology includes embedding MRI. With cross-modal scalar uncertainty Multiply and embed TVUS With cross-modal scalar uncertainty Multiplication. In other words, MRI embedding. TVUS Embedded Multiply by scalar uncertainty respectively and (See Figure 2(C)). This cross-modal fusion feature with MRI and TVUS modal uncertainties can be represented as:
[0121] (Equation 5)
[0122] The TVUS embedding can be a one-dimensional vector of length 512. The MRI embedding can be a one-dimensional vector of length 768. Each sample can have its own embedding, and there will be one MRI embedding corresponding to one TVUS embedding. In the linear projection layer, the final length of the vector can be 512.
[0123] For example, during the multimodal training phase, Method 300 can use a "5-fold cross-validation" strategy on the endometriosis MRI dataset (the same folding can be used for pre-training the TVUS classifier). For each fold, the multimodal fusion model can be trained for 100 epochs with a weight decay of 0.05, a batch size of 7, and a learning rate of 1e-4. The RNP module can be optimized using stochastic gradient descent (SGD) and a momentum of 0.9. During this multimodal training, the TVUS classifier can be kept either frozen (i.e., not fine-tuned during multimodal training) or unfrozen (i.e., fine-tuned during multimodal training). In some examples, the RNP module can be initialized with randomized parameters. An optimizer can be used during training to improve model performance by updating model parameters. The SGD optimizer adjusts parameters based on the calculated gradient of the loss function, allowing the model to gradually converge to the optimal solution. By iteratively updating parameters using optimization algorithms such as gradient descent, the optimizer helps the model find the optimal set of parameters that minimizes the loss and improves model accuracy and predictive power. Advantageously, using a frozen TVUS feature extractor can ensure that the model has good classification performance on TVUS and save training resources.
[0124] The method 300 of this technology also includes: S320, concatenating cross-modal fusion of MRI and TVUS embedding.
[0125] Method 300 may further include: S322, embedding the cascaded MRI and TVUS into at least one linear layer. This will produce a classification prediction of at least one sign of endometriosis (e.g., POD closure) indicating a diagnosis of endometriosis.
[0126] Cross-modal fusion features and Cascaded and passed to the final linear layer (represented as) The linear layer generates classification predictions.
[0127] The entire model can be optimized as follows:
[0128] (Equation 6)
[0129] During training, this method can iteratively optimize the equation from Equation 3. and from equation 6 This loss function summarizes the optimization strategy for the entire model. Ideally, the final step of optimization should minimize the cross-entropy so that all parts of the model are optimized during training.
[0130] The AI model is preferably tested using only one modality (MRI or TVUS input). During testing using only the MRI modality, input is provided to the TVUS stream. Similarly, during testing using only the TVUS modality, the MRI input was set to... Without TVUS, The value will be high because the TVUS modality is indeterminate, which leads to an amplification of the MRI embedding via Equation 5 (the opposite is true for missing MRI).
[0131] In some examples, when more than one sign of endometriosis (such as POD closure and intestinal nodules) is used simultaneously during training, at least one linear layer may include a multi-head attention (MHA) layer.
[0132] Considering the inherent connections and varying incidence rates among different signs of endometriosis, a multi-label classification module has been developed that leverages these relationships and effectively addresses the challenges of imbalanced learning. To capture the relationships between labels, a multi-head attention (MHA) layer can be employed.
[0133] In some examples, to implement the MHA layer, cross-modal fusion features can be processed first through two fully connected (FC) layers. The system generates initial classification outputs for each label. These outputs can be stacked and fed into an MHA layer, which identifies the correlations between labels. The results from the MHA can be destacked and passed through another set of classification layers to generate the multi-label probabilities needed to diagnose endometriosis. The multi-label classification module containing the MHA layer can be trained using a two-way loss (TWL) function, which has been shown to be effective in handling class imbalance in multi-label learning environments. The TWL function specifically addresses class imbalance in multi-label learning by distinguishing between different classes and samples, providing a more efficient solution to the imbalance problem compared to the traditional binary cross-entropy loss function.
[0134] In some examples, the TWL function can be minimized in the following ways:
[0135] (Equation 6B)
[0136] The following formula can be used in n One sample and C Optimize TWL functions on each label:
[0137] (Equation 6C)
[0138] Therefore, the entire model can be fine-tuned by minimizing the following unpaired multimodal multilabel (UMM) loss:
[0139] (Equation 6D)
[0140] in, It can be a large-scale female pelvic MRI dataset used for self-supervised pre-training. It could be an MRI dataset for endometriosis. It could be a TVUS sliding dataset. It could be the TVUS intestinal nodule dataset.
[0141] Method 300 further includes: S324, classifying signs of endometriosis (e.g., POD blockage). Thus, at least one sign of endometriosis, such as POD blockage, intestinal nodules, ovarian endometriotic cysts, and uterosacral ligament nodules, can be detected. This result can be output to the user of the AI model as an indication of signs of endometriosis (e.g., POD blockage and / or intestinal nodules and / or ovarian endometriotic cysts and / or uterosacral ligament nodules) and / or an indication of the likelihood of an endometriosis diagnosis.
[0142] LCKD model
[0143] The AI model 300 of this technology has the ability to diagnose endometriosis using medical imaging techniques, particularly magnetic resonance imaging (MRI) and transvaginal ultrasound (TVUS) modalities. To further enhance the performance of the AI model and address the challenge of modality loss during diagnosis, a learnable cross-modal knowledge distillation (LCKD) method / model 500 can be implemented during the creation of AI model 300, as detailed in the appendix. Figure 3 As described in step 308.
[0144] Figure 4 The framework for creating LCKD model 500 is illustrated in diagram form. Figure 5 A method for creating an LCKD model 500 is demonstrated. The LCKD model 500 can be used in the AI model 300 to improve the accuracy of the method 300. Therefore, those skilled in the art will understand that the LCKD model 500 can be combined with the AI model 300. Those skilled in the art will also understand that the LCKD model / method 500 can also be used alone (without performing the other steps in method 300) to identify at least one sign of endometriosis (e.g., POD closure and / or intestinal nodules and / or ovarian endometriotic cysts and / or uterosacral ligament nodules), thereby determining the probability of an endometriosis diagnosis.
[0145] The Learnable Cross-Modal Knowledge Distillation (LCKD) model 500 addresses the problem of missing modalities during training and testing by automatically identifying important modalities and extracting knowledge from them to learn parameters beneficial to all tasks while training against other modalities. In other words, LCKD is a method / model that uses knowledge learned from one modality to assist in the analysis of another, automatically detecting both "teacher" and "student" modalities. In the case of the current AI model 300, when a particular patient lacks one of the modalities (endometriosis MRI or TVUS), LCKD can be used to more effectively bridge this gap. Advantageously, by distilling cross-modal knowledge, the LCKD model can utilize features from available modalities to more accurately fill in missing information and further improve diagnostic accuracy.
[0146] In the current AI model, only two modalities are used: endometriosis MRI and TVUS. Therefore, the LCKD method 500 is configured to adapt to this specific setting. Those skilled in the art will understand that this LCKD model 500 can be adapted to handle more than two modalities.
[0147] In LCKD Model 500, the teacher network is trained on an available modality (MRI or TVUS, or both) to provide guidance and knowledge to a student network analyzing another modality. Advantageously, this setup allows the student network to learn from the knowledge of the teacher network, thereby enabling it to utilize cross-modal information.
[0148] Knowledge distillation (KD) involves transferring knowledge from a teacher network to a student network. The student network receives instruction from the teacher network during training and learns to imitate its behavior. Therefore, advantageously, this knowledge transfer facilitates the student network's use of available modalities to fill in missing modalities.
[0149] Creating an LCKD model 500 may include: S502, preprocessing the dataset. Preprocessing can be performed on endometriosis MRI and TVUS images to ensure compatibility. This step may involve normalization and / or resizing and / or data augmentation. For example, preprocessing of endometriosis MRI and TVUS datasets may be performed using techniques such as mirror manipulation and / or Gaussian noise techniques and / or brightness and contrast enhancement.
[0150] Creating an LCKD model 500 may include pre-training. This step corresponds to... Figure 3 The pre-training steps S302, S304, and S306 are described. Therefore, in the context of creating an LCKD model, the description of these steps is omitted. Those skilled in the art will understand that when an LCKD model is combined with an AI model, pre-training is performed only once and does not need to be repeated. The LCKD model can be used independently. Therefore, advantageously, the LCKD model does not require pre-training, but it can have at least one pre-trained encoder input as initialization. The encoder can be configured according to... Figure 3 Pre-training is performed as shown.
[0151] Creating the LCKD model 500 includes: S504, Teacher Election. The teacher election procedure is performed on the available modalities using either the endometriosis MRI or TVUS dataset to automatically detect the most suitable modality as the "teacher." Specifically, teacher election involves using a validation process. The validation process examines each individual modality and determines its performance on the validation set. Subsequently, the best-performing modality is selected as the teacher.
[0152] For each task k (in The application verification process selects the modality with the best performance as the teacher. In form, we have:
[0153] (Equation 7)
[0154] in, I Indexing different modalities, It is by A parameterized LCKD segmentation model, including encoder and decoder parameters. It is a function for calculating the Dice score.
[0155] This program can occur when processing all modes. Previously, among them express Data samples, and Indexed mode. Those skilled in the art will understand that... Figure 4 The LCKD framework in the previous model showed more than two modalities. However, in the current implementation, only two modalities (TVUS or MRI) exist.
[0156] The process of selecting a teacher involves choosing a modality that demonstrates good performance. In the current scenario, subjects often only have either MRI images or TVUS videos. Therefore, the teacher will be selected based on the modality available to the subject (TVUS or MRI). This is in... Figure 4 This is explained in the text, one of the modes Selected as a teacher As a student, and It is assumed to be missing.
[0157] Subsequently, each modality is encoded to output features separately. For available modalities (i.e. Knowledge distillation is performed between each pair of teacher and student modalities. However, for missing modalities... Its characteristics Through available features The missing modal features are generated during the process of generating these features.
[0158] The teacher is a modality with good performance. In methods for identifying endometriosis, the teacher will be a usable modality. Furthermore, a teacher election procedure is introduced to automatically select suitable teachers for different tasks.
[0159] In some examples or as an alternative, S504 teacher election may include electing teachers who are oriented towards the worst students.
[0160] In this step, the encoder associated with the modality that performs worst for a particular label acts as the student, learning from the teacher, who is the encoder of the modality that performs well for that label. Advantageously, this approach allows modalities to learn from each other, enhancing their strengths and compensating for their weaknesses.
[0161] A teacher modality can be elected by comparing the performance of the two modalities in each training iteration. For example, during the first training epoch, the teacher modality can be randomly selected. After the first epoch, the model's performance can be evaluated using a validation dataset. The validation dataset can be created by stratified random sampling of 20% of the training dataset. This evaluation measures the model's performance on different labels for each modality.
[0162] For example, if the signs of endometriosis are POD blockade and intestinal nodules, the procedure can be based on POD blockade using MRI data. and intestinal nodules The validation set, and the POD enclosure of TVUS data. and intestinal nodules The validation set is used to generate classification results based on the area under the receiver operating characteristic curve (AUC). By comparing these four metrics, the worst-performing student and their corresponding modality can be identified. Then, another modality can be designated as the teacher for subsequent training rounds. Specifically, in In each round, the worst-performing metric can be defined as:
[0163] (Equation 7B) The worst-performing mode in a round can be determined using the following formula:
[0164] (Equation 7C)
[0165] therefore, The teacher and student modalities for each round are as follows:
[0166] (Equation 7D)
[0167] During training, the teacher modality may not be static. Due to the inherent differences in learning complexity between each modality and label, the teacher modality may change dynamically. Advantageously, this approach allows two modalities to take turns acting as the teacher, thereby facilitating efficient mutual knowledge distillation.
[0168] Creating LCKD model 500 includes: S506, distilling cross-modal knowledge.
[0169] like Figure 4 As shown, in each mode Input to by After parameterizing the encoder, the features of each modality are obtained. As shown below: (
[0170] (Equation 8)
[0171] Cross-modal knowledge distillation (CKD) is defined by a loss function that makes features from all modalities approach the teacher modality in a pairwise manner across all tasks, as follows:
[0172] (Equation 9)
[0173] in, This represents the p-norm operation, and assumes a modality. Missing. Minimizing this loss pushes the model parameter values toward a point in the parameter space that maximizes the performance of all modalities across all tasks.
[0174] Because knowledge distillation exists between each teacher-student pair, the features of each modality in the feature space should approximate "real" features that can perform well across different tasks. Still assuming modalities... Missing, therefore missing features Available features generate:
[0175] (Equation 10)
[0176] For knowledge distillation (KD), we use a specific metric (L1 or L2 distance) to bridge the gap between features of two modalities. That is, once the model is trained, we can replace the missing modalities with the available modalities because their features are very similar.
[0177] Creating the LCKD model 500 includes: S508, training teacher-student pairs. The teacher-student pairs will be trained using available modalities (endometriosis MRI or TVUS, or both) and guided by the teacher network's knowledge. During training, the student network aims to mimic the behavior and predictions of the teacher network. Advantageously, this process enables the student network to utilize cross-modal knowledge and fill in features of missing modalities. The training process can be iterative. Similarly, the teacher election procedure can also be iterative.
[0178] All features encoded according to Equation 8 or generated according to Equation 10 are then concatenated and input into a system composed of... In the parameterized decoder, to predict:
[0179] (Equation 11)
[0180] in, It is the predicted result of the task.
[0181] The LCKD model is trained by minimizing the following objective function:
[0182] (Equation 12)
[0183] in, The objective function is the overall objective function (e.g., for signs of endometriosis such as POD block detection, using cross-entropy and Dice loss); α is the trade-off between the task objective and the cross-modal objective. If α=0, model performance will decrease; while when α>0, the results will improve. A suitable value for α could be 0.1.
[0184] The LCKD model can be tested based on the following steps: obtain all available image modalities in the input to generate features according to Equation 8, and generate features of the missing modalities according to Equation 10. These features are then provided to the decoder to predict the segmentation result according to Equation 11.
[0185] In the reasoning stage, such as Figure 4 or Figure 6 As shown, when the LCKD model is used in AI Model 300 or alone, when a particular patient lacks a modality (MRI or TVUS), the trained student network can combine cross-modal knowledge learned during training to generate predictions using the available modalities. This ensures that the diagnostic process can continue even when a modality is unavailable. Therefore, by utilizing cross-modal knowledge, the LCKD model can fill in missing modal features, thus ensuring diagnostic reliability even when a modality is unavailable.
[0186] The AI model using this technology is used to identify endometriosis.
[0187] Figure 6 A computer-implemented method 400 is presented to use an AI model and / or an LCKD model to detect at least one sign of endometriosis (POD closure and / or intestinal nodules and / or ovarian endometriotic cysts and / or uterosacral ligament nodules) and thereby determine endometriosis.
[0188] Method 400 includes: S408, acquiring the subject's MRI results or TVUS results. Those skilled in the art will also understand that this method is equally effective when S408 includes acquiring both the subject's MRI results and TVUS results. MRI results are MRI scan images of a specific patient. TVUS results are TVUS examination videos of a specific patient. Method 400 also includes: S410, inputting the MRI results and / or TVUS results into an AI model and / or an LCKD model. The AI model in... Figure 3 The LCKD model is described in [the text]. Figure 5 As described in the text.
[0189] Subjects may be patients. MRI results or TVUS results may be input by the user into the method / AI model / LCKD model. Those skilled in the art will understand that subjects may input their MRI results and / or TVUS results into the AI model and / or LCKD model. Users may be medical professionals. Alternatively, patients may be able to input MRI results and / or TVUS results into the method / AI model and / or LCKD model. MRI results may be T2 SPACE images. TVUS results may be videos of the slippage sign examination. Those skilled in the art will understand that any other similar format may be used, but MRI results are preferably in the same format as the endometriosis MRI dataset, and TVUS results are preferably in the same format as the TVUS dataset used to construct the AI model and / or LCKD model.
[0190] The subject's MRI and / or TVUS results can be obtained by the user directly uploading them to a computing device configured to implement the method / AI model and / or LCKD model of this technology. The subject's MRI and / or TVUS results can be obtained from the user via a communication network (e.g., the Internet, Bluetooth, etc.).
[0191] Those skilled in the art will understand that method 400 may include acquiring the subject's MRI results and / or TVUS results, and method 400 may also include: S410, inputting the MRI results and / or TVUS results into the AI model and / or LCKD model.
[0192] Method 400 further includes: S424, detecting at least one sign of endometriosis (e.g., POD block and / or intestinal nodules and / or ovarian endometriotic cysts and / or uterosacral ligament nodules). The at least one sign of endometriosis may be classified / detected by an AI model and / or an LCKD model.
[0193] Method 400 further includes: S426, outputting the results of the detection of signs of endometriosis. In the case of POD closure, a sign of endometriosis, POD closure is a strong indicator for diagnosing endometriosis. Therefore, method 400 of this technology can output a conclusion to the user regarding whether signs of endometriosis (e.g., POD closure) have been detected. Alternatively, method 400 of this technology can additionally or alternatively output a score indicating the probability of signs of endometriosis (e.g., POD closure). This score can be expressed as a percentage of confidence. The classification can be "yes" (e.g., POD closure detected) or "no" (e.g., POD closure not detected). These results can be displayed on a display of a device configured to implement the method and / or AI model and / or LCKD model of this technology. Alternatively, the results can be sent to a user-selected device via secure communication technologies (e.g., a private / password-protected email address, etc.).
[0194] Interaction between AI model and LCKD model
[0195] Those skilled in the art will understand that the AI model can be implemented independently without using the LCKD model. In other words, Figure 3 The method can be implemented without S308 (using the LCKD model). Signs of endometriosis (e.g., POD block and / or intestinal nodules and / or ovarian endometriotic cysts and / or uterosacral ligament nodules) can be detected simply by using Figure 3 The AI model is used for detection.
[0196] Those skilled in the art will understand that the LCKD model can be implemented independently without using an AI model. In other words, Figure 5 The method can be used alone. At least one sign of endometriosis (e.g., POD block and / or intestinal nodules and / or ovarian endometriotic cysts and / or uterosacral ligament nodules) can be detected by using only the method. Figure 5 The LCKD model is used for detection.
[0197] Those skilled in the art will also understand that LCKD models and AI models can be as follows: Figure 3 The diagram shows them combined. This means... Figure 3 The steps (involving building an AI model) can be compared with... Figure 5 The steps involved in creating an LCKD model are combined. Advantageously, this combination has yielded improved results in detecting at least one sign of endometriosis.
[0198] In this implementation, the LCKD model and the AI model can be combined in a stacked manner. For example, as Figure 3As shown, the AI model is stacked on top of the LCKD model. More specifically, there are two modalities: MRI images and TVUS videos. These two modalities are passed to the LCKD model to handle two cases: 1) if both modalities are available, the knowledge distillation process of the LCKD model is performed; 2) if one modality is missing, the missing features are predicted by the LCKD model. Subsequently, the features from both modalities (whether available or predicted) are used for uncertainty estimation in the AI model and ultimately for classification prediction.
[0199] Example 1: Endometriosis signs of POD blockage
[0200] In the case study, the datasets consisted of a female pelvic MRI dataset for self-supervised pre-training, an MRI endometriosis dataset, and a TVUS endometriosis dataset. The female pelvic MRI dataset contained T2 MRI volumes from 8984 women aged 18–45 years with endometriosis-like symptoms. These scans were performed at different locations using different protocols, and are highly likely to contain signs of other diseases, with the examined areas not necessarily aligned with the diagnostic area for endometriosis. The MRI endometriosis dataset consisted of 88 pelvic T2 SPACE MRI scans from women aged 18–45 years, performed at multiple clinics using standard protocols, including 19 cases of POD closure. These endometriosis MRIs focused on the diagnostic area of POD closure near the uterus. The TVUS endometriosis dataset contained 749 video clips of the "slipping sign," including 103 cases of negative slipping sign (POD closure).
[0201] In this case study, we pre-trained the 3D MAE network from scratch on the female pelvic MRI dataset for 195 epochs, using the AdamW / Adam optimizer, a batch size of 3, and a learning rate of 1e-3. During the multimodal fusion phase, we applied a linear layer after the encoder as an MRI feature extractor to adjust the output size (768) to be the same as the output size (512) of the TVUS encoder. For the ResNet(2+1)D network, we pre-trained on the Kinetics-400 dataset and trained it for 30 epochs on the TVUS endometriosis dataset, using the Adam / AdamW optimizer, a learning rate of 1e-5, and a batch size of 30. For the training of this TVUS classifier, we implemented a "5-fold cross-validation" strategy, dividing the dataset into 5 folds through stratified random sampling, using 4 folds for training and 1 fold for testing. During the multimodal training phase, we performed a “5-fold cross-validation” strategy on the endometriosis MRI dataset (the same folding was used for pre-training the TVUS classifier). For each fold, the multimodal fusion model was trained for 100 epochs using AdamW
[15] , a weight decay of 0.05, a batch size of 7, and a learning rate of 1e-4. The RNP module was optimized using SGD and a momentum of 0.9. During this multimodal training, we kept the TVUS classifier either frozen (i.e., not fine-tuned during multimodal training) or unfrozen (i.e., fine-tuned during multimodal training).
[0202] In this context, we also evaluated the area under the receiver operating characteristic (AUC) of the five-fold training and present its mean and standard deviation results. All experiments were performed using Python 3.9.7 with PyTorch 1.7.1. 3DMAE pre-training ran on three NVIDIA RTX 3090 GPUs with 24 GB of VRAM and took 98 hours. The multimodal model was trained on one NVIDIA RTX A6000 GPU, taking 4 hours per fold and requiring 47 GB of VRAM.
[0203] The results of the case study are as follows Figure 7 As shown. The AUC results for POD closed classification are shown in [the image / reference]. Figure 7The ResNet(2+1)D model (unimodal TVUS training and testing) achieved an AUC of 92.38% in the POD closed classification, slightly lower than AUC 96.5% due to the 5-fold cross-validation experiment. Due to the small dataset size, the 3D ViT (unimodal MRI training and testing) trained de novo on the endometriosis MRI dataset only achieved a low AUC of 65.03%. Self-supervised pre-training (PT) on a large-scale female pelvic MRI dataset based on 3D MAE significantly improved the ViT results to AUC 90% (unimodal MRI training and testing). The untrained multimodal fusion (MMF) had a low AUC of 63.48% (multimodal training, MRI testing with missing TVUS). Based on MAE pre-training, our multimodal learning model with a non-frozen (i.e., learnable) pre-trained TVUS classifier achieved an AUC of 91.79% (MRI testing with missing TVUS). Freezing the TVUS classifier during multimodal learning further improved the AUC of POD detection to 92.14% (MRI test with missing TVUS). For TVUS test with missing MRI, the POD detection results were slightly improved compared to single-modal TVUS training and testing, with an AUC of 92.57%.
[0204] Example 2: Signs of endometriosis include POD blockage and intestinal nodules.
[0205] In another case study, the dataset consisted of a female pelvic MRI endometriosis dataset and a TVUS endometriosis dataset. The MRI endometriosis dataset contained pelvic T2SPACE MRI scans of 171 women aged 18–45 years, including 63 scans showing varying degrees of POD closure and 11 scans containing intestinal nodules. Uterine segmentation was performed using ITK-SNAP with manual annotation. All 3D volume data were normalized to a right-anterior-superior (RAS) coordinate system to ensure anatomical alignment consistency, then resampled to a uniform 1×1×1 mm spacing using SimpleITK, and finally resized to 64×128×128. The TVUS sliding sign dataset contained 749 video clips, each approximately 10 seconds long, characterized by the "sliding sign," including 103 cases with a negative sliding sign (indicating POD closure). Additionally, the TVUS intestinal nodule dataset contained 519 intestinal TVUS scan video clips, each approximately 5 seconds long, with 35 cases showing intestinal nodules. During preprocessing, the top 70 pixels of each frame are automatically cropped to focus on a fan-shaped region. Each sequence is then uniformly temporally sampled for 32 frames, and the spatial resolution is adjusted to 112 × 112 using bilinear interpolation.
[0206] During the pre-training phase, the 3D MAE network utilizes an encoder consisting of 12 Transformer blocks and a 768-dimensional linear projection layer. The mask rate is set to... = 0.7. The decoder consists of 8 Transformer blocks. The entire network was trained from scratch on a female pelvic MRI dataset for a total of 195 epochs. The model parameters were fine-tuned using the AdamW optimizer (Loshchilov and Hutter, 2017), where 1 is set to 0.9. 2 is set to 0.95. Batch size is fixed at 3. The initial base learning rate is set to 1e-3 and gradually decreased following a linear scaling rule. = _ × / 256, as described in (He et al., 2022). For all subsequent training phases, a 5-fold cross-validation strategy was employed, dividing the endometriosis MRI dataset and two TVUS datasets into 5 folds using stratified random sampling, with 4 folds used for training and 1 fold reserved for testing. This initial partitioning remained consistent throughout all experiments. Furthermore, it was ensured that the same patient was not included in both the training and validation / test sets simultaneously.
[0207] During the segmentation training phase, in each fold, the encoder and decoder pre-trained with 3D MAE underwent 200 epochs of training using the AdamW optimizer with a learning rate of 1e-4 and a batch size of 3. The segmentation head and classification head had dimensions of 1536 and 768, respectively. Adaptive histogram equalization and random pruning were applied as data augmentation techniques during training.
[0208] In the multimodal, multi-label classification stage, settings for TVUS video mixing are implemented. = 4. The MRI feature extractor is a 3DMAE pre-trained encoder with additional linear layers to change the output dimension from 768 to 512 to fit the output of the TVUS encoder. The TVUS feature extractor uses a ResNet(2+1)D model with 18 R(2+1) convolutional layers. Batch normalization is used after each convolution and before the activation of the Rectified Linear Unit (ReLU). The network contains an adaptive average pooling layer and has been pre-trained on the Kinetics-400 dataset. For each fold, the model is trained for 200 epochs using the AdamW optimizer with weight decay set to 0.05. The batch size is configured to 16, and gradient accumulation is performed over 4 iterations to increase the effective batch size to 64. The initial learning rate is set to 1e-3. To simulate scenarios where modalities may be missing without requiring specific training modifications, one modality is randomly omitted in 50% of the iterations. The first set of fully connected (FC) layers has a dimension configuration of 64, while the multi-head attention (MHA) layer has 4 heads.
[0209] For MRI images, data augmentation techniques included random cropping, mirror flipping, contrast enhancement, simulating low resolution, Gaussian blurring, Gaussian noise, brightness adjustment, gamma correction, and multiplicative brightness adjustment. No data augmentation techniques were applied to the TVUS dataset.
[0210] In the five-fold cross-validation framework, AUCs for four tasks were calculated for the validation set corresponding to each fold: AUC for classifying closed PODs using MRI data (MRI_POD), AUC for classifying intestinal nodules using MRI data (MRI_BN), AUC for classifying closed PODs using TVUS data (TVUS_POD), and AUC for classifying intestinal nodules using TVUS data (TVUS_BN). For each task, AUC values were collected from the five folds of each round, and then the mean and standard deviation were calculated. This procedure was repeated in each round, generating a set of mean and standard deviation for each task in each round. Next, the AUC values for all four tasks in each round were calculated, followed by the mean and standard deviation of these four AUCs. The highest average AUC achieved across all rounds (AVG AUC) was determined as the model's performance metric, and the AUCs for each task were recorded in the round in which this highest average was achieved. For each round, the AUC values for all four tasks were determined, and then the mean and standard deviation of these AUCs were calculated. The classification results are shown in Table 1.
[0211] Table 1
[0212] The results show that, as a multimodal version of the LCKD method, Dynamic Mutual Cross-Modal Knowledge Distillation (DMKD) in this disclosure achieves the highest AUC on multiple diagnostic labels, and notably, it achieves the best average AUC (AVG) of 0.8481.
[0213] Figure 8 This is a schematic diagram of a computing device 800 that can be used to implement this technical method. (See attached diagram.) Figure 3 , 5 The technical methods described in Section 6 are computer-implemented methods. These methods can be implemented on a computing device, such as computing device 500.
[0214] Figure 8 Includes computing devices 800, which can be any suitable electronic device, such as a personal computer or computing device, laptop, tablet, smartphone, etc. It should be understood that this is a non-exhaustive and non-limiting list of example devices. Figure 6 Some components of computing device 500 are shown; it should be understood that other standard components not shown may also be present. The device includes at least one processor 502 coupled to memory 504. The at least one processor 502 may include one or more of the following: a microprocessor, a microcontroller, and an integrated circuit. Memory 504 may include, for example, volatile memory (such as random access memory (RAM)) used as temporary memory, and / or non-volatile memory (such as flash memory, read-only memory (ROM), or electrically erasable programmable ROM (EEPROM)) used to store data, programs, or instructions. Computing device 500 typically includes at least one user interface 506. Computing device 500 may include at least one display 508 for providing results and / or data generated during the methods described above. Device 500 may also include a machine learning module for training, storing, and executing [the machine learning process]. Figure 3 AI models and / or Figure 5 The LCKD model can also be implemented in a cloud computing system. Accordingly, the method of this technology can be provided to users as a web-based application (e.g., SaaS (Software as a Service)). The AI model can be downloaded to the user's device.
[0215] Various aspects of this technology are useful in improving the use of imaging techniques such as TVUS and / or MRI in the diagnosis of endometriosis, namely, by technically supplementing missing imaging modalities. This is beneficial in several ways. For example, POD closure is one of the most easily detectable signs of endometriosis on TVUS. However, on MRI, this sign is more difficult to detect visually due to fibrous tissue strands, blurred posterior uterine tissue margins, and the lack of fluid-filled space behind the uterus. The technology disclosed in this disclosure can improve the ability of MRI to detect POD closure to almost the same level as TVUS. During surgery, POD closure can hinder laparoscopic observation of posterior uterine lesions (the most common site), prolonging operative time and requiring surgeons with higher skill levels to remove endometriotic lesions and adhesions around the intestines and other organs. Identifying this sign is particularly useful for surgical triage, ensuring that surgery is performed by a surgeon with the appropriate skill level, allocating sufficient time, and requiring bowel preparation if adhesions are present. This can also affect fertility, as the fallopian tubes may become entangled or blocked due to adhesions, thus hindering the functional transport of eggs, sperm, and embryos.
[0216] For intestinal nodules, specialized ultrasound examination skills are required. Furthermore, the detection of this sign is limited by the detection range of the ultrasound probe; nodules in the upper intestine may be missed, while MRI may be more effective in detecting them. Improving the ability to detect intestinal nodules associated with endometriosis can explain symptoms such as intestinal pain, constipation, diarrhea, nausea, and loss of appetite. When referral to gynecology is necessary, this can also reduce the time and resources wasted by gastroenterologists using colonoscopy and endoscopy to examine the gastrointestinal system. This can prevent misdiagnosis as irritable bowel syndrome, allowing for timely initiation of targeted endometriosis treatment. More accurate detection also aids in surgical planning and triage.
[0217] Similarly, more accurate detection of signs of endometriosis, such as ovarian endometriotic cysts and / or uterosacral ligament (USL) nodules, can influence decisions regarding egg freezing and IVF (in vitro fertilization) treatment for ovarian endometriotic cysts, as well as decisions regarding medication to suppress menstruation, surgery, and monitoring for kidney and / or bladder pain for uterosacral ligament nodules.
[0218] Unless otherwise expressly stated, it will be understood from the following discussion that throughout the specification, discussions using terms such as “processing,” “calculation,” “accounting,” “determining,” “analysis,” or similar terms refer to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or convert data expressed as physical quantities (such as electronic quantities) into other data similarly expressed as physical quantities.
[0219] In a similar manner, the terms "controller" or "processor" can refer to any device or part of a device that processes electronic data (e.g., from registers and / or memory) to convert that electronic data into other electronic data, for example, that can be stored in registers and / or memory. "Computer," "computing machine," or "computing platform" can include one or more processors.
[0220] Throughout this specification, the terms "an embodiment," "some embodiments," or "an embodiment" refer to a specific feature, structure, or characteristic described in connection with that embodiment, which is included in at least one embodiment of this disclosure. Therefore, the phrases "in one embodiment," "in some embodiments," or "in an embodiment" appearing throughout the specification do not necessarily all refer to the same embodiment. Furthermore, specific features, structures, or characteristics may be combined in any suitable manner, as will be apparent to those skilled in the art from this disclosure, and may be present in one or more embodiments.
[0221] As used herein, unless otherwise stated, the use of ordinal adjectives such as “first,” “second,” “third,” etc., to describe common objects merely indicates reference to different instances of the same object and is not intended to imply that the objects being described must be in a given order in time, space, ranking, or any other way.
[0222] In the claims below and in the description herein, any of the terms “comprising,” “comprised of,” or “which comprises” are open-ended terms, meaning that they include the element / feature that follows it, but do not exclude other elements / features. Therefore, the term “comprising” as used in the claims should not be construed as limiting oneself to the means, elements, or steps listed thereafter. For example, the expression “device comprising A and B” should not be limited to a device consisting only of elements A and B. The terms “including” or “which includes / that includes” as used herein are also open-ended terms, similarly meaning that they include at least the element / feature following the term, but do not exclude other elements / features. Therefore, “including” and “comprising” have the same meaning.
[0223] It should be understood that in the foregoing description of embodiments of this disclosure, various features of this disclosure are sometimes combined in a single embodiment, drawing, or description thereof in order to simplify the disclosure and facilitate understanding of one or more of the various inventive aspects. However, this approach to disclosure should not be construed as reflecting an intention that the claims require more features than are expressly listed in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, each claim being itself a separate embodiment of this disclosure.
[0224] Furthermore, while some embodiments described herein include features that are also included in other embodiments but not others, different combinations of features from different embodiments are intended to fall within the scope of this disclosure and form different embodiments, as will be understood by those skilled in the art. For example, any claimed embodiment may be used in any combination of the following claims.
[0225] Many specific details are set forth in the description provided herein. However, it should be understood that embodiments of this disclosure can be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0226] It is also important to note that the term "coupled" as used in the claims should not be interpreted as limited to a direct connection. The terms "coupled" and "connected," and their derivatives, may be used. It should be understood that these terms are not intended to be synonymous. Therefore, the expression "device A coupled to device B" should not be limited to devices or systems where the output of device A is directly connected to the input of device B. It means that there is a path between the output of A and the input of B, which may include other devices or means. "Coupled" can refer to two or more elements in direct physical, electrical, or optical contact, or it can refer to two or more elements that are not in direct contact but still cooperate or interact with each other.
[0227] The embodiments described herein are intended to cover any adaptive modifications or variations of the invention. Although the invention has been described and explained in conjunction with specific embodiments, those skilled in the art will recognize that other embodiments falling within the scope of the invention can be readily conceived.
[0228] When any or all of the terms “comprise”, “comprises”, “comprised” or “comprising” are used in this specification (including the claims), they shall be construed as explicitly indicating the presence of the stated feature, integer, step or component, but not excluding the presence of one or more other features, integers, steps or components.
[0229] Those skilled in the art will understand that while the foregoing has described what is considered the best mode for carrying out this technology, as well as other modes where appropriate, this technology should not be limited to the specific configurations and methods disclosed in this preferred embodiment description. Those skilled in the art will recognize that this technology has a wide range of applications and that extensive modifications can be made to the embodiments without departing from any inventive concept as defined in the appended claims.
Claims
1. A computer-implemented method for determining the likelihood of endometriosis, the method comprising: Obtain the subject's magnetic resonance imaging (MRI) results or transvaginal ultrasound (TVUS) results; The MRI results or the TVUS results are input into an artificial intelligence (AI) model and / or a learnable cross-modal knowledge distillation (LCKD) model; and Detect at least one sign of endometriosis using the AI model and / or the LCKD model; and Output the results of the detection of at least one sign of endometriosis; The at least one sign of endometriosis is at least one of Douglas's pit (POD) closure, intestinal nodules, ovarian endometriotic cysts, and uterosacral ligament (USL) nodules.
2. The method as described in claim 1, in, The AI model includes: At least one MRI embedding is extracted from an endometriosis MRI dataset in an MRI stream using a magnetic resonance imaging (MRI) feature extractor. At least one TVUS embedding is extracted from the TVUS dataset in the TVUS stream using a transvaginal ultrasound (TVUS) feature extractor. Measuring the uncertainty in the MRI stream; Measure the uncertainty in the TVUS stream; Cross-modal fusion of the MRI embedding and the TVUS embedding is performed in the following manner: Multiply the at least one MRI embedding by the uncertainty measured in the TVUS stream, and Multiply the at least one TVUS embedding by the uncertainty measured in the MRI stream; Cascaded cross-modal fusion of the MRI embedding and the TVUS embedding; and Detect at least one of the signs of endometriosis.
3. The method of claim 2, wherein the AI model further comprises: The MRI feature extractor is pre-trained based on the MRI dataset; and / or The TVUS feature extractor is pre-trained based on the TVUS dataset.
4. The method of claim 2 or 3, wherein the AI model comprises: MRI images labeled with positive / negative signs of endometriosis from the endometriosis MRI dataset were matched with TVUS videos labeled with negative / positive signs of endometriosis from the TVUS dataset.
5. The method of any one of claims 2 to 4, wherein the AI model comprises: The cascaded, cross-modal fused MRI embedding and TVUS embedding are passed to the linear layer.
6. The method of claim 3, wherein The endometriosis MRI dataset and the TVUS dataset are unpaired.
7. The method of claim 3 or 6, wherein The endometriosis MRI dataset includes multiple MRI images, among which, At least one of the plurality of MRI images includes at least one of the following: a labeled POD closed case, a labeled intestinal nodule case, a labeled ovarian endometriotic cyst case, and a labeled uterosacral ligament nodule case.
8. The method of claim 3 and / or any one of claims 6 to 7, wherein The TVUS dataset includes multiple sliding TVUS video clips, among which, At least one of the plurality of TVUS video segments includes at least one of the following: a labeled POD closed case, a labeled intestinal nodule case, a labeled ovarian endometriotic cyst case, and a labeled uterosacral ligament nodule case.
9. The method of claim 3, wherein The MRI dataset is unlabeled and includes multiple MRI images of the female pelvic region.
10. The method of any one of claims 2 to 9, wherein the AI model comprises: Channel-level uncertainty in the MRI stream is measured using a first random network prediction (RNP) module; and / or The channel-level uncertainty in the TVUS stream is measured using the second RNP module.
11. The method of claim 10, wherein Both the first RNP module and the second RNP module include a fixed-weight random network and a learnable predictive network.
12. The method of claim 10 or 11, wherein the AI model comprises: The first RNP module and the second RNP module are trained by minimizing the loss function calculated from the difference between the output of the prediction network and the random network.
13. The method of any one of claims 10 to 12, wherein the AI model comprises: The first RNP module outputs a single scalar for each mode to represent the uncertainty. and The second RNP module outputs a single scalar for each mode to represent the uncertainty.
14. The method as described in any of the preceding claims, in, Obtaining the MRI results or TVUS results of the subject includes receiving the MRI results and / or TVUS results by a medical professional or the subject.
15. The method as described in any of the preceding claims, wherein The MRI results are MRI images; and in, The TVUS result is a sliding TVUS video clip.
16. The method according to any one of claims 2 to 15, wherein The TVUS feature extractor is a convolutional neural network.
17. The method according to any one of claims 2 to 16, wherein The MRI feature extractor is a 3D vision Transformer.
18. The method as described in any of the preceding claims, wherein The test detects at least one of POD blockade, intestinal nodules, ovarian endometriotic cysts, and uterosacral ligament nodules, and the test is used to diagnose endometriosis.
19. The method as described in any of the preceding claims, wherein The LCKD model includes selecting teacher modalities through a validation process.
20. The method of claim 19, wherein The LCKD model includes distilling cross-modal knowledge between each pair of available teacher and student modalities using a loss function.
21. The method of claim 20, wherein The LCKD model includes training teacher and student modal pairs, wherein... During training, the student modality simulates the behavior and predictions of the teacher modality.
22. A data processing apparatus comprising means for performing the method as described in any of the preceding claims.
23. A computer program comprising instructions that, when executed by a computer, cause the computer to perform the method as claimed in any one of claims 1 to 21.
24. A computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform the method as claimed in any one of claims 1 to 21.