Automated disease detection system
By using an automated detection system and image processing and fuzzy inference techniques, the problems of poor standardization and high false negative rate of the IFA method in disease screening have been solved, achieving efficient and accurate automated detection of IFA.
Patent Information
- Application Number
- CN202180049741.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-09
- Filing Date
- 2021-06-07
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2041-06-07
AI Technical Summary
While existing immunofluorescence assays (IFA) are highly sensitive in disease screening, they require interpretation by human experts, resulting in poor standardization and lack of scalability. Furthermore, existing scalable methods such as ELISA and qPCR have high false negative rates.
An automated detection system, combining image processing, convolutional neural networks (CNN), and fuzzy inference, is used to automatically identify cell patterns in IFA images. Disease detection is performed using probability index (PI) and fuzzy rules, reducing reliance on human evaluation.
It improves the scalability and accuracy of disease detection, achieves consistency with expert human pathologists' testing, reduces the false negative rate, and enhances the scalability and testing efficiency of IFA.
Smart Images

Figure CN116075835B_ABST
Abstract
Description
Technical Field
[0001] Examples of disease detection related to, for example, nasopharyngeal carcinoma (NPC) or other autoimmune diseases are disclosed. Background Technology
[0002] Nasopharyngeal carcinoma (NPC) is thought to be caused by the reactivation of Epstein-Barr virus (EBV) in the nasal epithelium. A characteristic feature of this reactivation is the expression of the EBV early antigen (EA) complex. For this reason, secretory IgA antibodies against the EA complex in patient serum are highly sensitive and specific biomarkers against NPC (see, for example, references [1] and [2]). Because EA is a large complex comprising multiple protein subunits (see, for example, reference [3]), expression of the entire natural EA complex in cell-based assays provides the broadest antigenic coverage and therefore the highest sensitivity for the detection of NPC (see, for example, references [4], [5], and [6]). This method, known as immunofluorescence assay (IFA), is the preferred method for screening NPC in high-risk individuals. IFA is also a preferred method for detecting other diseases, such as autoimmune diseases. Summary of the Invention
[0003] There are some challenges. For example, while IFA is the preferred method for disease screening in high-risk individuals, it unfortunately requires interpretation by human experts and is therefore poorly standardized and not scalable (see, for example, reference [3]). In short, routine IFA is more of an art than a science.
[0004] Recent attempts to increase the scalability of disease screening have focused on ELISA (see, for example, references [7], [8], and [9]) and qPCR (see, for example, reference
[10] ) targeting EBV DNA. In both cases, these scalable methods also have high false-negative rates, which is not ideal for screening (see, for example, references [4], [6], and
[11] ). In contrast, IFA has many advantages. Recent studies have shown that IFA detects new NPC cases in high-risk groups with 100% sensitivity (see, for example, reference [5]). In three out of five patients, IFA positivity precedes visual nasal endoscopy confirmation, suggesting that IFA has the potential to enable early disease detection. An added advantage is that the staining pattern associated with EA-positive (EA+) samples is easily distinguishable from false-positive patterns caused by autoantibodies (see, for example, reference
[12] ) and immune complexes (see, for example, reference
[13] ).
[0005] This disclosure does not replace IFA with a scalable but less efficient modality, but rather increases the scalability of IFA by reducing the need for human evaluation. Basic pattern recognition is known to be used to automate the quantification (also known as titer testing) of IFA signals for EA+ samples (see, for example, reference
[14] ). In this disclosure, an automated detection system is used to distinguish EA+ and EA- samples to a degree comparable to that of expert human evaluators. Attached Figure Description
[0006] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate various embodiments.
[0007] Figure 1 The illustration shows a typical pattern encountered at IFA.
[0008] Figure 2 The illustration shows a disease detection system (DDS) according to an embodiment.
[0009] Figure 3A and 3B The diagram illustrates the membership function according to an embodiment.
[0010] Figure 3C The illustration depicts fuzzy rules according to an embodiment.
[0011] Figure 4 The illustration shows a convolutional neural network (CNN) according to an embodiment.
[0012] Figure 5 This is a flowchart illustrating a process according to an embodiment.
[0013] Figure 6A and 6B The ROC curve is shown.
[0014] Figure 7 A and 7B show the distribution for EBV scores and EA+ indices.
[0015] Figure 7 C, 7D, 7E, and 7F show the EBV score and EA+ index plotted against standard data.
[0016] Figure 7 G and 7H show the defuzzified output values plotted against standard data.
[0017] Figure 8 The PI distribution is shown for a comparison of FI and DeLFI.
[0018] Figure 9 This is a flowchart illustrating a process according to an embodiment.
[0019] Figure 10The illustration shows a DDS according to some embodiments. Detailed Implementation
[0020] The detection of serological antibodies against Epstein-Barr virus proteins is considered the gold standard for screening NPC in high-risk groups. Among current detection methods, immunofluorescence assay (IFA) is the most sensitive. Given the high survival rate of early-stage asymptomatic patients compared to the poor prognosis of late-stage NPC, IFA has enormous life-saving potential for screening the general population. The advantages of IFA stem from its ability to identify and enumerate cell staining patterns. In particular, IFA excels in detecting low titers with weakly fluorescent positive patterns and in excluding false positive samples with bright but negative patterns. However, these advantages rely on highly trained IFA evaluators with suitable personality traits and endurance for microscopic work. Therefore, IFA is not scalable because it requires trained pathologists to visually interpret cell staining patterns.
[0021] Therefore, this disclosure provides an automated disease detection system (DDS) 200 (see [link]). Figure 2 This drawback is overcome. In one embodiment, the detection system 200 models the IFA evaluation process, achieving a high degree of consensus with expert human pathologists in recognizing cellular patterns indicative of NPCs. Therefore, the detection system 200 disclosed herein significantly improves the scalability and accuracy of disease detection.
[0022] In one embodiment, the detection system 200 is designed in collaboration with expert IFA reviewers. When the reviewer examines the IFA slide, three basic input variables need to be evaluated. First, is there a sufficient number of cells to make a decision? Next, what proportion of these cells (if any) are brighter than the baseline cells? Finally, do these brighter cells exhibit a pattern consistent with a positive test result? Figure 1 The illustration shows a typical pattern encountered at IFA. For example... Figure 1 As illustrated, negative patterns tend to have low baseline fluorescence, while positive patterns include patterns with a “spot,” “peripheral dot,” or “cytoplasmic” appearance.
[0023] Figure 2The illustration shows a detection system 200 according to an embodiment. The detection system 200 includes an image processor 202 that receives an IFA image 201 as input and generates a processed IFA image. The processed image is input to a cell detector 204 for detecting cells in the processed IFA image. For each detected cell, pixel information about the detected cell is provided to a probability index (PI) generator 206, which uses the pixel information to generate a PI value (also called a “pattern score”) for the detected cell. These PI values are input to an input variable generator (IVG) 208, which generates three input variables: the total number of detected cells (numCells), the “EA+ index”, and the “EBV score”.
[0024] IVG 208 can calculate numCells (i.e., the total number of visible cells) by counting the number of PI values output by generator 206. IVG 208 calculates the EA+ index by determining the total number of PI values above a certain threshold and dividing that by numCells. That is, cells with PI values above the threshold are defined as EA+ cells, and their proportion in the total cell population is the EA+ index. In one embodiment, the EBV score is the average of a set of PI values above the threshold.
[0025] The input variables are fed into the fuzzy inference (FI) system 210, which uses the input variables to distinguish between test negative samples and test positive samples. In one embodiment, membership functions are constructed to define low, medium, and high ranges for the three input variables (numCells, EA+ index, and EBV score), and negative, boundary, and positive ranges for the outcome. Figure 3A The illustrations depict these membership functions used in the "Fuzzy Inference (FI)" embodiment, and Figure 3B The diagram illustrates these membership functions of a "Deep Learning FI (DeLFI)" embodiment. For example... Figure 3A and 3B As shown, all three variables (numCells, EA+ index, and EBV score) are divided into three fuzzy sets: low, medium, and high. The low and high sets are modeled as S-shaped, while the medium set is modeled as Gaussian. The FI system 210 employs fuzzy rule 302 (e.g., ...). Figure 3C As shown, the input variables are mapped to the output function. These fuzzy rules can be quickly summarized as follows: if numCells is high and both the EA+ index and EBV score are high, then the output is "positive"; otherwise, if numCells is high and both the EA+ index and EBV score are low, then the output is "negative"; otherwise, the output is "boundary line" or "uncertain".
[0026] In the FI embodiment, the PI value is defined as the coefficient of variation (CV) of pixel intensity for each identified cell. In the DeLFI embodiment, a convolutional neural network (CNN) is used to compute the PI value for each cell to assign the PI value to each cell. That is, in the DeLFI embodiment, the PI value generator 206 includes a CNN, an example of which is... Figure 4 The diagram is shown in the image. In one embodiment, as shown in the image... Figure 4 As shown, CNN 400 takes a 150x150x3 image as input and consists of four convolutional blocks, each including a 3x3 convolution, ReLU activation, and Max pooling. The final block is fed into three reduced-size fully connected layers, terminating in a final output layer activated by a sigmoid function for binary classification. CNN 400 is trained using a dataset of 550EA+ and 550EA- cell images identified by IFA experts from the IFA image database.
[0027] 1. Research Design - Materials and Methods
[0028] Figure 5 The illustration shows a process 500 for comparing the performance of the FI and DeLFI embodiments of system 200 using titers from manually evaluated historical serum samples as standard data. To evaluate FI and DeLFI, two hundred and ten (210) historical serum samples with known titers were randomly selected (see [link to documentation]). Figure 5 (Step s502). All samples were processed according to the IFA protocol (step s504), and titers were then manually assigned by a blinded IFA expert evaluator (step s506). In parallel with this manual evaluation, IFA slides were also imaged with a single dilution and analyzed using FI and DeLFI (step s508). The titers from the manual evaluation were then used as standard data to evaluate the performance of both (step s510).
[0029] 1.1 Serum Samples
[0030] Two hundred and ten (210) anonymous serum samples from historical screenings were obtained from the WHO Collaborating Centre for Immunology Research and Training in Singapore. The distribution of the samples according to their assigned titers is shown below.
[0031]
[0032] All samples were processed according to the standard IFA protocol described below. IFA-processed wells with a serum dilution of 1:2.5 were imaged at 20x magnification and a fixed exposure time of 30ms on a Leica DM4500B microscope with a scientific CMOS image sensor camera (pco.edge 3.1).
[0033] 1.2. Immunofluorescence assay (IFA)
[0034] Indirect immunofluorescence assay (IFA) was used to measure serum titers of anti-EBV-EA IgA, as previously described (see, for example, references [4], [5], and
[14] ). Briefly, Raji and P3HR1 cells were cultured in flasks and then induced for 2 days in a 37°C CO2 incubator with sodium butyrate (3 mM) and Phorbol 12-myristate 13-acetate (20 pg / ml). The induced cells were then washed four times in PBS by centrifugation, discarding the supernatant and resuspending the clumps in PBS. The resulting cell suspension was then dispensed onto Teflon-coated slides using a multichannel pipette. The slides were allowed to air dry overnight on a workbench, then fixed in ice-cold acetone for 10 minutes and allowed to dry completely. The fixed slides were stored at -80°C until use. Fixed slides coated with Raji cells (EA) were incubated for 30 minutes with 10 μL of serum, which had been serially diluted in PBS (1:10, 1:40, 1:160, and 1:640). For EBV-EA IgA, a serum dilution of 1:5 was also tested. The slides were then rinsed and further incubated for 30 minutes with fluorescein-conjugated anti-human IgA rabbit antibody (SPDS Scientific, Singapore), and then evaluated under a Leica DM4500B fluorescence microscope using a scientific CMOS image sensor camera (pco.edge 3.1) to capture IFA images.
[0035] 1.3 Image Processor 202 and Cell Detector 204
[0036] Image processor 202 is used to process IFA image 201 to produce a processed IFA image. In one embodiment, image processor 202 first smooths IFA image 202 to produce a smoothed image. In this particular example, a median filter of size (5x5) is applied to smooth the image. Next, the background is suppressed. In this example, a reconstructed top hat is applied to the smoothed image to suppress the background. This operator is applied to a square structuring element (SE) of size 41 pixels. The resulting image is then thresholded using Rosin's method to separate the background region from the foreground region, thereby producing a binarized image. The binarized image is then thinned by applying a filling hole (i.e., a region of black pixels surrounded by white pixels in the binarized image) using morphological reconstruction, followed by a binarized opening of a disk-shaped SE with a diameter of 3 pixels. To remove non-cellular objects from the thinned binary image, an area-size filter (minimum size: 700, and maximum size: 7000) is applied to produce a filtered thinned binary image. Cell detector 204 uses a label-controlled watershed algorithm to separate foreground objects from the filtered image into individual cells, where local maxima of the geodesic distance map are used as seeds. As a post-processing step, non-circular regions are removed (threshold: 0.4).
[0037] 1.4 PI value generator 206 and EA+ cell detection
[0038] Following cell detection, two different techniques (i) coefficient of variation (CV) and (ii) deep learning (DL) are used to calculate a dimensionless exponent called the probability index (PI) for each detected cell. For the CV method, PI = σ / μ, where μ and σ are the mean and standard deviation of the pixel intensity of the detected cell. For the DL method, a CNN (e.g., CNN 400) is used to classify EA+ and EA- cells. The output layer of the CNN consists of a single neuron with a sigmoid activation function. This output is used as the PI value for the DL method. A higher PI value here indicates the similarity to the training set of EA+ cell images. The design and training of the CNN are described in Section 1.5 below.
[0039] A higher PI value indicates greater variability in fluorescence associated with EA+ staining. PI values for all detected cells were calculated using both CV and DL methods. A given detected cell, cell_i, is classified as EA+ if the PI value PI_i is greater than or equal to a threshold; otherwise, the cell is classified as EA-. For CV PI values, if the CV-PI value for a cell is greater than or equal to: (μ...) c +σ c If μ > μ, the cell is classified as EA+, where μ > ... c +σ crepresents the mean and standard deviation of the Poisson binomial distribution of the CV PI values {CV-PI_1, CV-PI_2, ..., CV-PI_N} for N detected cells. For the DL PI values (i.e., {DL-PI_1, DL-PI_2, ..., DL-PI_N}), if the DL-PI value for a cell is greater than 0.5, the cell is classified as EA+.
[0040] EBV score is obtained by the average PI value of EA+ cells. The EA+ index is the ratio of the number of EA+ cells to the total number of cells.
[0041] 1.5. Convolutional Neural Networks (Design and Training)
[0042] CNN 400 was designed for EA-positive and EA-negative cell classification. The CNN takes a 150x150x3 image as input. The CNN consists of four convolutional blocks, each comprising a 3x3 convolution-ReLU activation-max pooling operation with 32, 64, 128, and 128 kernels respectively. All four max pooling operations are performed with a stride of 2 to reduce the output dimension of the convolutional operations. The final convolutional block is connected to three fully connected layers of sizes 6272, 512, and 1. The final output layer is activated with a sigmoid activation function for binary classification. The CNN model was trained using a dataset of 550EA+ and 550EA- cell images identified by IFA experts.
[0043] 1.6 Fuzzy Inference (FI) System 210
[0044] The FI system 210 is used to distinguish between test negative and test positive samples based on three input variables (numCells, EA+ index, and EBV score, also known as “clear input” values). The FI system uses fuzzy set theory to map clear input values to clear output values. As is known in the art, the main components of the FI system are a fuzzer, fuzzy rules (also known as a fuzzy rule base), an inference engine, and a defuzzifier (see, for example, reference
[17] , Chapter 4, “Fuzzy Inference Systems”). An overview of the operation of the FI system is as follows.
[0045] 1: The fuzzer uses the membership function described above to map sharp input values and output variables to fuzzy values (0 to 1).
[0046] 2: Then, the fuzzy input is interpreted by a fuzzy rule base (in the form of IF-THEN rules), which describes how the FI system should make decisions on a set of inputs.
[0047] 3: The inference engine activates rules for a given set of inputs and finds the results of the rules by combining the rule strength and the output membership function. These results are then combined to obtain a fuzzy output.
[0048] 4: Then use a deblurrer to convert the blurred output into a clear output.
[0049] More specifically, each of the three input variables was modeled as three fuzzy sets with the following membership functions: “low” (S-shaped), “medium” (Gaussian), and “high” (S-shaped). Fuzzy rules were created in consultation with IFA experts. The output variable was similarly modeled as three fuzzy sets with the following membership functions: “negative” (S-shaped), “boundary” (Gaussian), and “positive” (S-shaped). Parameters for these input and output variables were estimated using an IFA image of a reference sample including 39 negative control samples and 132 positive control samples. By setting a threshold for positive, the defuzzified, sharp output values were then used for classification.
[0050] Receiver-operator characteristic (ROC) curves were generated by evaluating the sensitivity and specificity across the unfuzzified output range. Optimal performance was defined as the point that maximizes Youden's J(sensitivity + specificity - 1). Cohen's κ was calculated using the following: in, and
[0051] 1.7 Uncertainty Filter
[0052] When using an uncertainty filter for decision-making, we only evaluate performance using the defuzzified output value after excluding samples falling within the uncertainty window (which we define as ±0.05 of the defuzzified output threshold). The uncertainty window range used in this study is the clear output range described in Table 3.
[0053] Performance of FI and DeLFI
[0054] The performance of FI and DeLFI was evaluated using their ROC curves. Figure 6A and Figure 6B First, both sensitivity (sen) and specificity (spe) with respect to standard data are evaluated at different thresholds for the deblurred output (also known as the clear output) value. Only samples greater than or equal to the threshold are considered positive. This analysis is repeated after excluding samples falling within ±0.05 of the deblurred output threshold. In a real workflow, samples excluded by this "uncertainty filter" are manually evaluated by IFA experts.
[0055] In ROC analysis, when all samples were analyzed without an uncertainty filter, both FI (AUC = 0.944) and DeLFI (AUC = 0.0985) performed well. Figure 6AThis demonstrates that both FI and DeLFI are highly consistent with the evaluations made by our IFA experts. At their optimal cutoff values, both DeLFI (sen = 0.971, spe = 0.919) and FI (sen = 0.949, spe = 0.865) provide summary statistics with reasonable performance.
[0056] Table 1 (All Samples in FI)
[0057]
[0058] Table 2 (All DeLFI Samples)
[0059]
[0060]
[0061] Table 1 shows the performance metrics of FI at different clear output cutoff thresholds. The optimal cutoff (defined as the point that maximizes Youden's J(sensitivity + specificity - 1)) is 0.400, and the AUC is 0.944. Table 2 shows the performance metrics of DeLFI at different clear output cutoff thresholds. The optimal cutoff (defined as the point that maximizes Youden's J(sensitivity + specificity - 1)) is 0.750, and the AUC is 0.985.
[0062] This is also reflected in Cohen's kappa at these cutoffs, showing that both DeLFI (κ = 0.90) and FI (κ = 0.82) are almost identical to the manual evaluation (see, for example, reference
[15] ). When the uncertainty filter is active, the AUC for both FI (AUC = 0.950) and DeLFI (AUC = 0.990) increases, but DeLFI still outperforms FI (AUC = 0.990). Figure 6B As expected, the aggregate statistics at their optimal cutoff were also improved for both DeLFI (sens = 0.945, spe = 0.986) and FI (sen = 0.896, spe = 0.943).
[0063] Table 3 (FI uncertain samples minus)
[0064]
[0065] Table 4 (DeLFI uncertain samples minus)
[0066]
[0067] Table 3 shows the performance of FI (after subtracting uncertain samples). Uncertain samples are defined as samples included within a specific clear output range. After subtracting these uncertain samples, a performance metric is generated for each clear output range. The optimal clear output range (defined as the range that maximizes Youden's J(sensitivity + specificity - 1)) is [0.5–0.6], and the AUC is 0.950.
[0068] Table 4 shows the performance of DeLFI (after subtracting uncertain samples). Uncertain samples are defined as samples included within a specific clear output range. After subtracting these uncertain samples, a performance metric is generated for each clear output range. The optimal clear output range (defined as the range that maximizes Youden J(sensitivity + specificity - 1)) is [0.8–0.9], and the AUC is 0.990.
[0069] At these cutoff values, DeLFI produced 14 uncertain samples (6.7% of the total), which is comparable to, and no greater than, the 15 uncertain samples (7.1% of the total) for FI. Therefore, the superiority of DeLFI over FI is not attributable to the increase in uncertain samples excluded from the analysis. Cohen's kappa increases for both DeLFI (maxκ = 0.92) and FI (maxκ = 0.84), demonstrating that removing uncertain samples produces a small but significant improvement in consistency.
[0070] Although both FI and DeLFI perform robustly, DeLFI generally outperforms FI in every comparison. To understand the differences between FI and DeLFI, we plotted the distributions of EBV scores and EA+ indices for them. We found that the EBV score distributions are indeed significantly different for FI and DeLFI. Figure 7 A and Figure 7 B). A relatively high degree of overlap was observed between the positive and negative sample distributions for FI, while the distribution for DeLFI was biased towards the extremes of the EBV score range.
[0071] An interesting trend emerged when both the EBV score and the EA+ index were plotted against standard data. Figure 7 (CF). In the case of FI, the EBV score shows a monotonic increase across all standard data categories, but the EA+ index remains largely unchanged. In contrast, in the case of DeLFI, it is the EA+ index, not the EBV score, that is shown proportionally across all standard data categories. Therefore, in addition to assessing test positivity, both the EBV score of FI and the EA+ index of DeLFI can be used to infer the EA titer.
[0072] Since FI identifies nearly the same proportion of EA+ cells in all images, regardless of titer, the discriminative power of FI must already be provided by the EBV score. This contrasts with DeLFI, where both the EBV score and the EA+ index distinguish between negative and positive samples. This improved discrimination is reflected in DeLFI's clearer separation between negative and positive samples when the unblurred output values are plotted against standard data. Figure 7 (GH). This makes it easier to find the cutoff point for achieving high sensitivity and specificity. In contrast, FI exhibits a relatively long tail, which overlaps between negative and positive samples, explaining its slightly worse performance.
[0073] The EBV score and EA+ index are ultimately derived from the PI value, so we wanted to know how the PI values are distributed for FI and DeLFI given a typical IFA image. Here, we observe that the distribution summarizes what was observed with the EBV score distribution, where the positive and negative distributions for DeLFI are more significantly separated than for FI (see [link to EBV score distribution]). Figure 8 ). Figure 8 The PI distributions for FI and DeLFI are shown. The PI distributions for FI and DeLFI are shown for a single typical IFA image. The dashed lines depict the threshold above which cells are considered EA+. The distribution for FI shows a positively skewed unimodal distribution, while DeLFI produces a bimodal distribution clearly separated by the threshold. The above analysis leads us to conclude that the fundamental basis for DeLFI's superior performance over FI is that its latent PI value is more discriminative between negative and positive samples.
[0074] Figure 10 This is a block diagram of a disease detection system (DDS) 200 according to some embodiments. For example... Figure 10As shown, DDS 200 includes: a processing circuit (PC) 1002, which may include one or more processors (P) 1055 (e.g., a general-purpose microprocessor and / or one or more other processors, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc.); the processors may coexist in a single housing or a single data center, or may be geographically distributed (i.e., DDS 200 may be a distributed computing system); at least one network interface 1048, including a transmitter (Tx) 1045 and a receiver (Rx) 1047, for enabling DDS 200 to transmit and receive data from other nodes connected to a network 110 (e.g., an Internet Protocol (IP) network) to which the network interface 1048 is connected (directly or indirectly) (e.g., the network interface 1048 may be wirelessly connected to the network 110, in which case the network interface 1048 is connected to an antenna arrangement); and a storage unit (also referred to as a "data storage system") 1008, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where PC 1002 includes a programmable processor, a computer program product (CPP) 1041 may be provided. CPP 1041 includes a computer-readable medium (CRM) 1042 storing a computer program (CP) 1043 including computer-readable instructions (CRI) 1044. CRM 1042 may be a non-transient computer-readable medium, such as a magnetic medium (e.g., a hard disk), an optical medium, a memory device (e.g., random access memory, flash memory), etc. In some embodiments, the CRI 1044 of computer program 1043 is configured such that, when executed by PC 1002, the CRI causes DDS 200 to perform the steps described herein (e.g., the steps described herein with reference to the flowcharts). In other embodiments, DDS 200 may be configured to perform the steps described herein without requiring code. That is, for example, PC 1002 may consist solely of one or more ASICs. Therefore, the features of the embodiments described herein can be implemented in hardware and / or software.
[0075] Summary of various implementation examples
[0076] A1. A computer-implemented method for detecting diseases (900, see also) Figure 9The method includes: obtaining (s902) an immunofluorescence assay (IFA) image associated with a sample (e.g., a serum sample); processing (s904) the IFA image to produce a processed IFA image; detecting (s906) cells in the processed IFA image; determining (s908) numCells, where numCells is the total number of detected cells; for each detected cell, classifying (s910) the cells into a first type of cell (EA+ cells) or a second type of cell (EA- cells); calculating (s908) based on numCells and numCellsEA+. 912) Index value (EA+ index), where numCellsEA+ is the total number of detected cells classified as EA+ cells; calculate (s914) score value (EBV score); map numCells to the first set of fuzzy values using (s916) the first set of membership functions; map the EA+ index to the second set of fuzzy values using (s918) the second set of membership functions; map the EBV score to the third set of fuzzy values using (s920) the third set of membership functions; and classify the sample using (s922) the first set of fuzzy values, the second set of fuzzy values, the third set of fuzzy values, and fuzzy rules.
[0077] A2. The method according to embodiment A1, wherein the step of classifying the cells is performed by a convolutional neural network (CNN).
[0078] A3. The method according to embodiment A2, wherein, for each detected cell, the CNN determines a probability index (PI) value for the cell, and uses the PI value and a predetermined threshold to determine whether the cell should be classified as an EA+ cell.
[0079] A4. The method according to embodiment A3, wherein, as a result of determining that the PI value for a specific cell exceeds the threshold, the CNN classifies the specific cell as an EA+ cell.
[0080] A5. The method according to embodiment A1, wherein classifying the cell as a first type of cell (EA+ cell) or a second type of cell (EA- cell) includes: obtaining pixel information for the cell; using the pixel information to calculate a probability index (PI) value for the cell; and using the PI value and a predetermined threshold to determine whether the cell should be classified as an EA+ cell.
[0081] A6. The method according to embodiment A5, wherein the pixel information for the cell includes a set of pixel intensity values, wherein each pixel intensity value in the set of pixel intensity values indicates the intensity of a pixel corresponding to the cell.
[0082] A7. The method according to embodiment A6, wherein using the pixel information to calculate the PI value for the cell includes calculating: PI = σ / μ, where μ is the mean of the pixel intensity values and σ is the standard deviation of the pixel intensity values.
[0083] A8. The method according to any one of embodiments A3-A7, wherein the EBV score is the average of PI values greater than or equal to the predetermined threshold, or the EBV score is the average of PI values greater than the predetermined threshold.
[0084] B1. A computer program (1043) including instructions (1044) that, when executed by the processing circuitry (1002) of a disease detection system, cause the disease detection system (200) to perform the method according to any one of embodiments A1-A8.
[0085] B2. A carrier comprising the computer program according to embodiment B1, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer-readable storage medium.
[0086] C1. A disease detection system (200) adapted to perform the method according to any one of embodiments A1-A8.
[0087] D1. A disease detection system (200) comprising: a processing circuit (1002); and a memory (1042) containing instructions (1044) executable by the processing circuit, wherein the disease detection system (200) is operated to perform the method according to any one of embodiments A1-A8.
[0088] in conclusion
[0089] The advantages of IFA stem from its ability to identify and enumerate cell staining patterns. In particular, IFA excels at detecting low titers with weakly fluorescent positive patterns and excluding false positive samples with bright but negative patterns. These advantages are based on highly trained IFA evaluators with suitable personality traits and the stamina required for microscopy. Encapsulating such expertise in scalable computational models is key to providing IFA services at scale.
[0090] To build the computational model, we want to take advantage of the huge performance advances made by convolutional neural networks (CNNs) and address the problem that CNNs are black boxes that are uninterpretable in many respects. Such interpretability issues are a concern for applications involving or influencing medical decision-making. Hybrid systems are one way to mitigate this problem (see, for example, reference
[16] ). In this disclosure, an interpretable rule-based fuzzy framework incorporating CNN modules is shown to yield state-of-the-art results in both worlds. Fuzzy reasoning provides a broad framework by synthesizing control rules based on human experience, while the specific task of cell image recognition is performed by the CNN. Limiting the CNN to analyzing individual cell patterns reduces computational complexity compared to analyzing the entire image. Such an approach also allows for easy modification of cell pattern categories in the future.
[0091] Beyond improved scalability, the precise quantitative output based on single-sample dilutions is another major advantage of automated analysis. In manual methods, human evaluators must choose from five incremental dilutions at which positive patterns (if detected) are no longer visible to the naked eye. This can be particularly challenging near the decision boundary, where distinguishing between 1:10 positive and negative samples is often difficult. Since samples near this boundary are generally considered to have a higher error rate, we initially included an "uncertainty filter" to reference such samples for further human evaluation. Interestingly, however, DeLFI exceeded our expectations by actually matching human performance (AUC = 0.985, κ = 0.90) without requiring a filter. We attribute this performance to the fuzzy rules and the high quality of the training image dataset, both created in close consultation with our IFA experts. Notably, DeLFI's precise quantitative output allows for fine-tuning of the decision boundary for different clinical scenarios. For example, by selecting an appropriate, clear output cutoff, DeLFI can be used to maximize sensitivity (sen = 0.992, spe = 0.916) or specificity (sen = 0.945, spe = 0.986) (Table 4). Maximizing sensitivity is advantageous when screening for endemic NPCs in high-prevalence populations. Conversely, maximizing specificity will minimize false positives when screening for NPCs in non-prevalence populations. Such flexibility is not possible with this level of granularity achieved using manual IFA titers.
[0092] This disclosure discloses early steps that enable numerous laboratories running the same software model to achieve high performance, thereby improving the overall quality and reproducibility of IFA testing. This opens the door to accurate and scalable population screening of NPCs using IFA.
[0093] Although various embodiments have been described herein, it should be understood that they have been presented by way of example only and not limitation. Therefore, the breadth and scope of this disclosure should not be limited by any of the exemplary embodiments described above. Furthermore, unless otherwise indicated herein or clearly contradicted by the context, the elements described above are covered by this disclosure in all possible combinations of their variations.
[0094] Furthermore, although the process described above and illustrated in the accompanying figures is shown as a series of steps, this is done solely for illustrative purposes. Therefore, it should be expected that some steps may be added, some steps may be omitted, the order of the steps may be rearranged, and some steps may be performed in parallel.
[0095] References:
[0096] [1]Screening Test Review Committee.Report of the Screening TestReview Committee.(2019).
[0097] [2] Chan, SH et al. MOH Clinical Practice Guidelines 1 / 2010. (2010).
[0098] [3] Middeldorp, JMEpstein-Barr Virus-Specific Humoral ImmuneResponses in Health and Disease. Curr. Top. Microbiol. Immunol. 391, 289–323 (2015).
[0099] [4] Tay, JK et al. The Role of Epstein-Barr Virus DNA Load and Serology as Screening Tools for Nasopharyngeal Carcinoma. Otolaryngol. Head NeckSurg. 155, 274–280 (2016).
[0100] [5] Tay, J.K. et al. A comparison of EBV serology and serum cell-free DNA as screening tools for nasopharyngeal cancer: Results of the Singapore NPC screening cohort. Int. J. Cancer (2019) doi:10.1002 / ijc.32774.
[0101] [6] Chan, S.H. et al. Epstein Barr virus (EBV) antibodies in the diagnosis of NPC--comparison between IFA and two commercial ELISA kits. Singapore Med. J. 39, 263–265 (1998).
[0102] [7] Coghill, A.E. et al. Epstein-Barr virus serology as a potential screening marker for nasopharyngeal carcinoma among high-risk individuals from multiplex families in Taiwan. Cancer Epidemiol. Biomarkers Prev. 23, 1213–1219 (2014).
[0103] [8] Hutajulu, S.H. et al. Seroprevalence of IgA anti Epstein-Barr virus is high among family members of nasopharyngeal cancer patients and individuals presenting with chronic complaints in head and neck area. PLoS One 12, e0180683 (2017).
[0104] [9]Paramita, D.K., Fachiroh, J., Haryana, S.M. & Middeldorp, J.M. Evaluation of commercial EBV RecombLine assay for diagnosis of nasopharyngeal carcinoma. J. Clin. Virol. 42, 343–352 (2008).
[0105]
[10] Chan, K.C.A. et al. Analysis of Plasma Epstein–Barr Virus DNA to Screen for Nasopharyngeal Cancer. N. Engl. J. Med. 377, 513–522 (2017).
[0106]
[11] Nicholls, J.M. et al. Negative plasma Epstein-Barr virus DNA nasopharyngeal carcinoma in an endemic region and its influence on liquid biopsy screening programmes. Br. J. Cancer (2019) doi:10.1038 / s41416-019-0575-6.
[0107]
[12] Ricchiuti, V., Adams, J., Hardy, D.J., Katayev, A. & Fleming, J.K. Automated Processing and Evaluation of Anti-Nuclear Antibody Indirect Immunofluorescence Testing. Front. Immunol. 9, 927 (2018).
[0108]
[13] Horsfall, A.C., Venables, P.J., Mumford, P.A. & Maini, R.N. Interpretation of the Raji cell assay in sera containing anti-nuclear antibodies and immune complexes. Clin. Exp. Immunol. 44, 405–415 (1981).
[0109]
[14] Goh, S.M.P. et al. Increasing the accuracy and scalability of the Immunofluorescence Assay for Epstein Barr Virus by inferring continuous titers from a single sample dilution. J. Immunol. Methods 440, 35–40 (2017).
[0110]
[15] Landis, J.R. & Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 33, 159–174 (1977).
[0111]
[16] Bonanno, D., Nock, K., Smith, L., Elmore, P. & Petry, F. An approach to explainable deep learning using fuzzy inference. in Next-Generation Analyst V vol. 10207 102070D (International Society for Optics and Photonics, 2017).
[0112]
[17] Knapp, B., Fuzzy Sets and Pattern Recognition, available at URL; www.cs.princeton.edu / courses / archive / fall07 / cos436 / HIDDEN / Knapp / fuzzy.htm.
Claims
1. A computer-implemented method (900) for detecting a disease, the method comprising: obtaining an immunofluorescence test, IFA, image associated with a sample; processing the IFA image to produce a processed IFA image; detecting cells in the processed IFA image; determining numCells, where numCells is a total number of detected cells; for each detected cell, classifying the cell as a first type of cell, an EA+ cell, or a second type of cell, an EA- cell; calculating an index value, an EA+ index, based on numCells and numCellsEA+, where numCellsEA+ is a total number of detected cells classified as EA+ cells; calculating a score value, an EBV score; mapping numCells to a first set of fuzzy values using a first set of membership functions; mapping the EA+ index to a second set of fuzzy values using a second set of membership functions; mapping the EBV score to a third set of fuzzy values using a third set of membership functions; and classifying the sample using the first set of fuzzy values, the second set of fuzzy values, the third set of fuzzy values, and fuzzy rules.
2. The method of claim 1, wherein, The step of classifying the cells is performed by a convolutional neural network, CNN.
3. The method of claim 2, wherein, For each detected cell, the CNN determines a probability index, PI, value for the cell, and uses the PI value and a predetermined threshold to determine whether the cell should be classified as an EA+ cell.
4. The method of claim 3, wherein, As a result of determining that the PI value for a particular cell exceeds the threshold, the CNN classifies the particular cell as an EA+ cell.
5. The method of claim 1, wherein, Classifying the cells as a first type of cell, an EA+ cell, or a second type of cell, an EA- cell comprises: obtaining pixel information for the cell; using the pixel information to calculate a probability index, PI, value for the cell; and using the PI value and a predetermined threshold to determine whether the cell should be classified as an EA+ cell.
6. The method of claim 5, wherein, The pixel information for the cell comprises a set of pixel intensity values, where each pixel intensity value in the set of pixel intensity values indicates an intensity of a pixel corresponding to the cell.
7. The method of claim 6, wherein, Using the pixel information to calculate the PI value for the cell comprises calculating: PI = σ / μ, where μ is a mean of the pixel intensity values, and σ is a standard deviation of the pixel intensity values.
8. The method of any one of claims 3-7, wherein the EBV score is an average of the PI values that are greater than or equal to the predetermined threshold, or the EBV score is an average of the PI values that are greater than the predetermined threshold.
9. A computer program (1043) comprising instructions (1044) which, when executed by processing circuitry (1002) of a disease detection system, cause the disease detection system (200) to perform the method according to any one of claims 1-8.
10. A carrier containing the computer program of claim 9, wherein the carrier is one of an electronic signal, an optical signal, a radio frequency signal or a computer readable storage medium. The carrier is one of: an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (1042).
10. A disease detection system (200) comprising: processing circuitry (1002) configured to perform the method according to any one of claims 1-8; and a carrier (1041) configured to carry the processing circuitry (1002).
11. A disease detection system (200) configured to: obtain an immunofluorescence assay, IFA, image associated with a sample; process the IFA image to produce a processed IFA image; detect cells in the processed IFA image; determine numCells, where, numCells is a total number of detected cells; for each detected cell, classify the cell as a first type of cell, an EA+ cell, or a second type of cell, an EA- cell; compute an index value, an EA+ index, based on numCells and numCellsEA+, where numCellsEA+ is a total number of detected cells classified as EA+ cells; compute a score value, an EBV score; map numCells to a first set of fuzzy values using a first set of membership functions; map the EA+ index to a second set of fuzzy values using a second set of membership functions; map the EBV score to a third set of fuzzy values using a third set of membership functions; and classify the sample using the first set of fuzzy values, the second set of fuzzy values, the third set of fuzzy values, and fuzzy rules.
12. The disease detection system of claim 11, wherein, the disease detection system is further configured to perform the method of any one of claims 2-8.
13. A disease detection system (200) comprising: processing circuitry (1002); and a memory, the memory containing instructions (1044) executable by the processing circuitry, whereby the disease detection system (200) is configured to perform the method of any one of claims 1-8.
Citation Information
Patent Citations
Self-learning mechanism-base fast matching fuzzy reasoning method
CN105787563A
Full-automatic detection system for precision diagnosis of nasopharyngeal cancer
CN106841628A