CTC dynamic identification system and quantitative analysis method based on multi-mode AI
By using multimodal AI technology to identify and analyze CTCs, the problems of missed detection and insufficient dynamic analysis of EMT-type CTCs in existing technologies have been solved, achieving high-precision CTC detection and dynamic treatment response tracking, thus improving the accuracy of detection and clinical efficacy.
Patent Information
- Application Number
- CN202511136454.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-11-21
AI Technical Summary
Existing CTC detection technologies cannot effectively capture EMT-type CTCs and lack the ability to analyze the dynamic evolution of cell subtypes and treatment responses, resulting in poor consistency of detection results and insufficient clinical utility.
A CTC dynamic recognition system based on multimodal AI is adopted. The image processing module identifies cells in multimodal images, extracts morphological and multispectral fluorescent labeling features, fuses feature vectors and performs spatiotemporal modeling, and outputs dynamic recognition results. The quantitative analysis module performs quantitative analysis on the recognition results and generates a quantitative analysis report that includes cell category attributes, subtype classification and time-series evolution status.
It achieves high-precision identification of EMT-type CTCs, improves the ability to dynamically track cell subtype classification and treatment response, shortens the median decision-making time, constructs a closed loop of identification-quantification-decision, and improves the accuracy and timeliness of detection.
Smart Images

Figure CN120997830A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cell detection technology, specifically to a CTC dynamic identification system and quantitative analysis method based on multimodal AI. Background Technology
[0002] CTC (Circulating Tumor Cell) detection, as a core technology of liquid biopsy, has important value in early cancer screening, efficacy evaluation, and prognostic monitoring.
[0003] Current mainstream CTC detection methods primarily rely on manual microscopy and semi-automated platforms (such as the CellSearch system). Manual microscopy requires technicians to observe enriched cell samples field-by-field under a microscope, which is time-consuming, labor-intensive, and highly experience-dependent. Subjective differences between operators lead to inconsistent results, especially for EMT-type CTCs undergoing epithelial-mesenchymal transition (EMT). While existing semi-automated platforms (such as the CellSearch system) partially replace manual methods, the enrichment technology relied upon by EpCAM cannot capture EMT-type CTCs and only provides static CTC counts, lacking the ability to analyze the dynamic evolution of cell subtypes and treatment responses. More importantly, the current technical process is fragmented, ending after enrichment and counting, failing to establish a closed loop of "identification-quantification-decision," making it difficult for detection results to drive clinical intervention.
[0004] Therefore, there is an urgent need to build a complete "identification-quantification-decision" system to address the fundamental deficiencies of CTC testing in terms of accuracy, timeliness, and clinical efficacy. Summary of the Invention
[0005] In order to build a complete "identification-quantification-decision" system to address the fundamental deficiencies of CTC detection in terms of accuracy, timeliness, and clinical utility, this application provides a CTC dynamic identification system and quantitative analysis method based on multimodal AI.
[0006] The CTC dynamic recognition system and quantitative analysis method based on multimodal AI provided in this application adopt the following technical solution: A CTC dynamic recognition system based on multimodal AI includes: The image processing module is used to identify each cell in the multimodal image and extract the morphological features and multispectral fluorescent labeling features of each cell; The result recognition module is used to fuse morphological features and multispectral fluorescent labeling features to obtain a fused feature vector, and outputs the dynamic recognition result of CTC based on the fused feature vector; The quantitative analysis module is used to perform quantitative analysis on the dynamic identification results and output a quantitative analysis report.
[0007] Furthermore, the steps for identifying individual cells in a multimodal image include: Multimodal images are spatially partitioned to establish spatial partition entities for each cell; After associating the corresponding multimodal attribute features with each spatial partition entity, the cells in each spatial partition entity are identified.
[0008] Further steps for extracting the morphological features and multispectral fluorescent labeling features of each cell include: Obtain cell coordinate sets based on spatial partition entities; Morphological features were calculated by applying spatial geometric operations combined with cell coordinate sets; and, Multi-channel spectral analysis combined with cell coordinate sets was used to calculate the multi-spectral fluorescent labeling characteristics.
[0009] Furthermore, the steps for fusing morphological features and multispectral fluorescent labeling features to obtain the fused feature vector include: A first quality assessment index for evaluating morphological features; morphological weights generated based on the first quality assessment index; and... A second quality assessment index is used to evaluate the characteristics of multispectral fluorescent labels, and fluorescent label weights are generated based on the second quality assessment index. By combining morphological features and morphological weights, weighted morphological features are obtained; and by combining multispectral fluorescent labeling features and fluorescent labeling weights, weighted multispectral fluorescent labeling features are obtained. The weighted morphological features and the weighted multispectral fluorescent labeling features are combined to generate a fused feature vector.
[0010] Furthermore, the steps for outputting the dynamic recognition result of CTC based on the fused feature vector include: A spatiotemporal knowledge graph is constructed based on the fused feature vectors, and the evolution path of the spatiotemporal knowledge graph is constrained by a pre-set cell state transition rule base. Generate classification decision functions based on spatiotemporal knowledge graphs; The fused feature vector is dynamically analyzed using a classification decision function, and the dynamic recognition result is output.
[0011] This application also proposes a quantitative analysis method for CTC based on multimodal AI, including: Based on the input dynamic recognition results, perform quantitative analysis on the dynamic recognition results and output a quantitative analysis report.
[0012] Furthermore, the steps for quantitatively analyzing the dynamic identification results and outputting a quantitative analysis report include: The CTC category attributes, subtype classification, and time-series evolution status in the dynamic recognition results are converted into quantitative parameters; Quantitative parameters are calculated based on the time dimension to obtain statistical indicators; Integrate quantitative parameters and statistical indicators to generate quantitative analysis reports.
[0013] Furthermore, after calculating the quantitative parameters based on the time dimension and obtaining the statistical indicators, the process also includes: It receives evolution path identifiers from a preset cell state migration rule base and integrates these identifiers into statistical indicators.
[0014] Beneficial effects achieved: This application provides a CTC dynamic recognition system based on multimodal AI, comprising: an image processing module for identifying each cell in a multimodal image and extracting morphological features and multispectral fluorescent labeling features of each cell; a result recognition module for fusing morphological features and multispectral fluorescent labeling features to obtain a fused feature vector, and outputting the dynamic recognition result of CTC based on the fused feature vector; and a quantitative analysis module for performing quantitative analysis on the dynamic recognition result and outputting a quantitative analysis report.
[0015] In this application, the CTC dynamic recognition system based on multimodal AI identifies cells in multimodal images through an image processing module and simultaneously extracts morphological features and multispectral fluorescent labeling features, overcoming the limitation of traditional single-modal detection in missing EMT-type CTCs of epithelial-mesenchymal transition. The result recognition module integrates dual-modal features to generate a fused feature vector, and outputs a three-in-one dynamic recognition result based on spatiotemporal joint modeling, including cell category attributes, subtype classification, and time-series evolution status, thereby improving the accuracy of dynamic tracking of treatment response and subtype analysis. The quantitative analysis module calculates population evolution statistical indicators on the dynamic recognition results and generates a quantitative analysis report that integrates quantitative parameters and clinical decision-making suggestions, driving the improvement of treatment adjustment accuracy and the reduction of median decision-making time. Finally, it constructs a closed loop of the entire process of "recognition-quantification-decision-making", systematically solving the fundamental defects of CTC detection in terms of accuracy, timeliness, and clinical efficacy. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of a CTC dynamic recognition system based on multimodal AI according to this application; Figure 2 This is a schematic diagram of the process steps that the image processing module of this application can execute; Figure 3 This is a schematic diagram of the process steps that the result recognition module of this application can perform; Figure 4 This is a schematic diagram of the process steps that the quantitative analysis module of this application can perform; Figure 5 This is a schematic diagram of the complete testing process for CTC.
[0017] Explanation of icon numbers: 10. Image processing module; 20. Result recognition module; 30. Quantitative analysis module. Detailed Implementation
[0018] The following is in conjunction with the appendix Figure 1-5 This application will be described in further detail.
[0019] In the description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0020] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0021] This application discloses a CTC dynamic recognition system and quantitative analysis method based on multimodal AI.
[0022] Please refer to Figure 1 In one embodiment of this application, a CTC dynamic recognition system based on multimodal AI includes: Image processing module 10 is used to identify each cell in the multimodal image and extract the morphological features and multispectral fluorescent labeling features of each cell; result recognition module 20 is used to fuse morphological features and multispectral fluorescent labeling features to obtain a fused feature vector, and output the dynamic recognition result of CTC based on the fused feature vector; quantitative analysis module 30 is used to perform quantitative analysis on the dynamic recognition result and output a quantitative analysis report.
[0023] CTC detection, as a core technology of liquid biopsy, plays an irreplaceable role in early cancer screening, dynamic evaluation of treatment efficacy, and personalized treatment decisions. Currently, mainstream CTC detection methods primarily rely on manual microscopic examination and semi-automated platforms.
[0024] However, the above-mentioned approach has fundamental flaws: on the one hand, the EpCAM (Epithelial Cell Adhesion Molecule)-dependent enrichment mechanism cannot capture the EMT (Epithelial-Mesenchymal Transition) type CTC subpopulation that has undergone epithelial-mesenchymal transition, and manual microscopic examination is affected by subjective experience, resulting in poor consistency of results; on the other hand, the existing technical process stops at cell counting, lacks the ability to dynamically analyze cell subtype evolution and treatment response, and cannot generate quantitative decision reports that can drive clinical intervention, resulting in a serious disconnect between detection and treatment.
[0025] Therefore, based on the aforementioned problems, this embodiment proposes: performing cell-level analysis on the input multimodal image through the image processing module 10, identifying each cell entity using synchronous segmentation and feature extraction technology, and calculating its morphological features in parallel, including physical properties such as cell size, nucleocytoplasmic ratio, and nuclear deformity, as well as multispectral fluorescent labeling features, including multi-channel fluorescence intensity of epithelial markers, mesenchymal markers, and stem cell markers, thus overcoming the limitations of traditional single-modal detection. In this embodiment, the channel is a fluorescence channel.
[0026] Then, the result recognition module 20 integrates the above-mentioned dual-modal features based on a dynamic weighting strategy to generate a fusion feature vector with unified semantic expression. It then analyzes the cell phenotype migration and treatment response patterns through spatiotemporal joint modeling and outputs a three-in-one dynamic recognition result that includes cell category attributes, subtype classification, and time-series evolution status.
[0027] Finally, the quantitative analysis module 30 performs numerical conversion and population evolution statistical index calculation on the dynamic identification results, and integrates them to generate a quantitative analysis report that includes CTC proportion, subtype distribution vector, treatment response efficiency and clinical decision-making suggestions. This constructs a closed loop of the entire process of "feature extraction - dynamic identification - quantitative decision-making", systematically overcoming the fundamental defects in the accuracy, timeliness and clinical efficacy of CTC detection.
[0028] The specific implementation examples of each of the above modules are as follows: ① For details on the specific operations that the image processing module 10 can perform, please refer to [link / reference]. Figure 2 As shown, steps S10 to S14 are included: Step S10: Divide the multimodal image into image spaces and establish spatial partition entities for each cell.
[0029] In multimodal image processing, image space partitioning of multimodal images aims to divide the image into several independent spatial partition entities using segmentation algorithms such as watershed and U-Net deep learning. Each spatial partition entity corresponds to a potential cell region. Based on pixel-level alignment of morphological images and multispectral fluorescence images, it is ensured that the same spatial partition entity covers the complete physical boundary of a single cell and the corresponding molecular marker distribution, thereby constructing a spatial carrier based on cells. This allows each spatial partition entity to simultaneously bind morphological features and multispectral fluorescence marker features, thus solving the feature extraction errors caused by cell adhesion, blurred boundaries, or phenotypic heterogeneity in traditional single-modal segmentation. It also provides a spatially aligned multimodal attribute set for subsequent fusion feature vector generation, ultimately supporting the result recognition module 20 to achieve accurate evolution statistics of cell-level dynamic recognition and quantitative analysis module 30. This systematically overcomes the problem of insufficient accuracy and disconnect from clinical decision-making caused by the fragmentation of spatial information in existing technologies.
[0030] It should be noted that, in this embodiment, multimodal images refer to a collection of image data reflecting the multidimensional characteristics of cells, acquired simultaneously through different imaging techniques. Their core value lies in integrating complementary information to overcome the limitations of a single modality. Morphological images specifically refer to images of the physical structure of cells captured using optical microscopy techniques such as bright-field, phase-contrast, and differential interference microscopy, showcasing geometric and textural features such as cell size, shape, nucleocytoplasmic ratio, and pseudopodia morphology. Multispectral fluorescence images are images acquired after labeling molecular targets with fluorescent probes excited at specific wavelengths. Through multichannel imaging, protein expression or gene activity is quantified, achieving spatial localization of molecular phenotypes.
[0031] Step S11: After associating the corresponding multimodal attribute features with each spatial partition entity, identify each cell in each spatial partition entity.
[0032] First, based on affine transformation and scale-invariant feature transformation keypoint matching algorithms, morphological images and multispectral fluorescence images are spatially aligned at the pixel level to ensure that the coordinates of entities in the same spatial partition completely coincide in the bimodal image. Then, a mapping index table between spatial partition entities and the original multimodal data is established, associating each spatial partition entity with unprocessed original morphological pixel blocks and corresponding multispectral fluorescence pixel blocks within its coverage area, forming a binding relationship of "spatial partition entity coordinates → original data block". Finally, a spatial topological relationship is constructed using a Delaunay triangulation network to verify the continuity of the coverage area of adjacent spatial partition entities and the gradual change in fluorescence distribution, filtering out abnormal spatial partition entities caused by staining artifacts or tissue fragments. This results in the construction of a set of precisely bound multimodal original data blocks for each valid spatial partition entity, solving the spatial offset distortion problem caused by feature extraction followed by alignment in traditional processes.
[0033] After the original data blocks are bound, quantitative indicators are generated synchronously from the original data blocks of each spatial partition entity using feature extraction algorithms, such as morphological geometric analysis and fluorescence intensity statistics. This forms numerical feature vectors. Based on the unique identifier of the spatial partition entity, such as the hash value of the mask center coordinates, a key-value mapping relationship of "unique identifier of spatial partition entity → feature vector" is established. Then, the continuity of the feature space is verified by Delaunay triangulation topology, abnormal bindings are filtered out, and finally, the multimodal attribute feature set of each valid spatial partition entity is output, which is accurately associated with the spatial partition entity, thus completing the association between the spatial partition entity and the attribute features.
[0034] Next, based on the pre-trained classification decision function, the cell presence confidence score is calculated for the bound multimodal attribute features. When the confidence score exceeds a set threshold, the spatial partition entity is determined to be a CTC. At the same time, the topological relationship between entities across spatial partitions is combined to eliminate fragmentation or impurity interference. By integrating multimodal attribute features through a dynamic weighted fusion strategy, a cell-level feature fusion mechanism is constructed to solve the problems of missed detection and false detection in traditional identification, improve cell identification accuracy to >98%, and provide spatially aligned feature input for subsequent dynamic analysis, supporting the closed loop of temporal evolution and quantitative decision-making.
[0035] Step S12: Obtain the cell coordinate set based on the spatial partition entity.
[0036] After defining and validating the spatial partition entities, for each entity, its geometric center coordinates are extracted using the contour moment center point formula. Based on the unique identifier of each entity, a key-value mapping relationship of "unique identifier of spatial partition entity → centroid coordinates" is established. Simultaneously, the relative positions between spatial partition entities are recorded using the Delaunay triangulation adjacency matrix, ultimately forming a cell coordinate set containing absolute coordinates and relative topology. The aim is to provide a benchmark framework for cell dynamic tracking and spatial analysis: on the one hand, the centroid coordinate set supports cell migration trajectory mapping, such as changes in CTC position before and after chemotherapy, quantifying metastasis potential; on the other hand, the topological relationship matrix reveals the spatial structure of cell communities, such as the Delaunay triangulation density of CTC clusters aggregated in the blood vessel wall being >0.8, analyzing microenvironment interaction mechanisms. This provides spatial dimension input for temporal evolution modeling and quantitative decision-making, overcoming the spatial behavior analysis blind spot caused by missing coordinates in traditional static detection, and effectively improving the spatiotemporal resolution of CTC dynamic monitoring.
[0037] Step S13: Apply spatial geometric operations combined with cell coordinate sets to calculate morphological features.
[0038] Based on the contour point sequence provided by the cell coordinate set, let the contour point set be [(x1,y1),(x2,y2),...,(xn,yn)], calculate the cell area using Green's theorem integration: in, This represents the cell area.
[0039] The perimeter is generated by accumulating the Euclidean distances between adjacent contour points: in, The perimeter.
[0040] The nuclear area is calculated based on DAPI (Distributed Application Program Interface), and the cytoplasmic area is calculated by combining the total cell area. That is, the total cell area is subtracted from the projected area of the cell nucleus, and then the projected area of the cell nucleus is divided by the cytoplasmic area to obtain the nucleocytoplasmic ratio.
[0041] Simultaneously, by estimating discrete curvature, the curvature changes of contour points are analyzed, and the degree of nuclear deformity is quantified using the standard deviation of curvature, for example, 0.3. This transforms the spatial geometric attributes of cells into quantifiable and traceable morphological features, overcoming the subjectivity and inefficiency of traditional manual microscopy. It also provides mathematically verifiable morphological evidence for cell subtype classification and dynamic tracking, supporting subsequent multimodal fusion and quantitative decision-making closed loops. This systematically solves the problems of missed detection of EMT-type CTCs and misjudgment of treatment response caused by inaccurate morphological measurements in existing technologies.
[0042] It should be noted that cell area, perimeter, nucleocytoplasmic ratio, and nuclear deformity are morphological characteristics.
[0043] Step S14: Multi-channel spectral analysis combined with cell coordinate set is used to calculate the multi-spectral fluorescent labeling characteristics.
[0044] Based on the spatially partitioned entity-bound cell coordinate set [(x1,y1),(x2,y2),...,(xn,yn)], after locating the spatial region of the cell in the multispectral image, the pixel data of each fluorescence channel, such as the DAPI nuclear staining channel, the EpCAM-Cy3 epithelial marker channel, and the Vimentin-AF488 mesenchymal marker channel, are extracted sequentially. Through fluorescence intensity mean calculation, coefficient of variation analysis, and signal-to-noise ratio quantification, independent feature values of each fluorescence channel are generated, such as the EpCAM fluorescence intensity mean coefficient of variation.
[0045] Further, multi-fluorescence channel data is integrated to form fluorescent label feature vectors, and spatial continuity is verified by combining Delaunay triangulation topology to filter out abnormal staining regions. Finally, multispectral fluorescent label features are bound to each effective spatial partition entity. Its core purpose is to capture EMT-type CTCs that cannot be identified by single-fluorescence channel detection by quantifying the spatial heterogeneity of molecular expression and cross-fluorescence channel correlation, overcoming the false negative limitation of traditional EpCAM-dependent technology, and providing molecular phenotypic input for dynamic identification modules to support subsequent treatment response tracking and quantitative decision-making closed loop (such as generating high-risk early warning by combining the fluorescence signal intensity of drug resistance genes with morphological nuclear abnormality).
[0046] It should be noted that the average fluorescence intensity is calculated using Formula 1: ————Formula 1 in, This represents the mean fluorescence intensity. This represents the wavelength at spatial coordinates (xi, yi). The pixel fluorescence intensity value of the fluorescence channel, where N represents the total number of effective pixels within the cell region.
[0047] The coefficient of variation is calculated using formula 2: ————Formula 2 in, The coefficient of variation is 1. This represents the fluctuation range of fluorescence intensity values within the target area.
[0048] The signal-to-noise ratio is calculated using formula 3: ————Formula 3 in, For signal-to-noise ratio, Fluorescence intensity fluctuations in cell-free regions.
[0049] ② Regarding the specific operations that the result recognition module 20 can perform, please refer to [reference needed]. Figure 3 As shown, steps S20 to S25 are included: Step S20: Evaluate the first quality assessment index of morphological features and generate morphological weights based on the first quality assessment index.
[0050] Based on the contour point sequence provided by the cell coordinate set, the gradient magnitude at each contour point is calculated, and then the variance of the gradient magnitudes of all contour points is statistically analyzed. This variance value is the first quality assessment index. For example, a variance of 25 indicates a clear boundary, while a variance less than 10 indicates a blurred boundary. The technical logic is as follows: high gradient variance corresponds to sharp cell boundaries (such as CTCs without chemotherapy), and low gradient variance corresponds to blurred contours (such as fragmented cells after chemotherapy). This index objectively quantifies the reliability of morphological features, overcoming the subjective errors of manual judgment.
[0051] Based on this indicator, dynamic morphological weights are generated through the Sigmoid function. Finally, this weighted morphological feature is applied in multimodal fusion. Its core purpose is a triple closed loop: (1) Reliability-driven fusion: suppress feature distortion caused by blurred contours (such as reducing the misjudgment rate of fragmented cell area from 22% to 3%), and ensure that high-definition spatial partition entities dominate decision-making; (2) Dynamic anti-interference optimization: automatically reduce the contribution of low-quality features when cell morphology deteriorates during treatment, and avoid misjudgment of apoptosis response; (3) Quality control traceability enhancement: mark low-weight morphological features in the quantitative report (such as "contour variance = 7 weight 0.25, it is recommended to check the focus"), guide the calibration of microscope parameters, and break through the bottleneck of morphological feature distortion in dynamic treatment scenarios of fixed weight fusion.
[0052] It should be noted that the gradient magnitude is calculated based on Equation 4: ————Formula 4 in, For gradient magnitude, and These are the gradients in the x and y directions calculated using the Sobel operator.
[0053] The variance of the gradient magnitude at the contour points is calculated based on Equation 5: ————Formula 5 in, For variance, denoted as the average gradient magnitude, and N is the total number of contour points.
[0054] The dynamic morphological weights are calculated based on Formula 6: ————Formula 6 in, , where k is the slope adjustment coefficient, 15 is the manually set sharpness threshold, and e is the natural constant. When the slope parameter k=0.3, the weight increases by 0.2 for every 5 increase in variance. For example, when the high gradient variance is 22, the dynamic morphological weight is 0.85, while the dynamic morphological weight of the low gradient variance is reduced to 0.2.
[0055] Step S21: Evaluate the second quality assessment index of the multispectral fluorescent labeling characteristics, and generate fluorescent labeling weights based on the second quality assessment index.
[0056] For each fluorescence channel, the fluorescence intensity is first calculated in the target cell region based on Equation 1, while the standard deviation of background noise at the same wavelength is measured in a cell-free background region. The fluorescence intensity is then calculated in decibels based on Equation 3. This value, used as a second quality assessment indicator, quantifies the reliability of the signal, such as... A value greater than 20dB is considered a usable high-quality signal. A value within the 10~20dB range is considered a signal requiring review. A value less than 10dB is considered an invalid signal that requires re-staining.
[0057] Then the The values are mapped by the Sigmoid function and Formula 6 to generate dynamic fluorescent label weights. Finally, the weights are applied to weight the feature values of each fluorescent channel in the feature fusion stage. The core purpose of the triple closed loop is: (1) Reliability-driven fusion: suppress the interference of low signal-to-noise ratio fluorescent channels, such as when the weight of CD45 fluorescent channel SNR_λ=9dB is reduced to 0.1, so that the false detection rate of interstitial CTC is reduced from 18% to 3%; (2) Dynamic anti-attenuation optimization: automatically reduce the contribution of failed fluorescent channels when the fluorescence signal attenuates during treatment, such as when the weight of PD-L1 fluorescent channel SNR_λ decreases by 5dB per week, so as to reduce the weight by 0.3), to avoid distortion of drug resistance analysis; (3) Quality control traceability enhancement: mark the low weight fluorescent channel data in the quantitative report, such as "Vimentin fluorescent channel SNR_λ=8dB weight 0.15 recommended for re-staining", to guide the improvement of experimental procedures and break through the bottleneck of clinical decision failure in noisy scenarios of fixed weight fusion.
[0058] It should be noted that when the slope parameter in Equation 6 is 0.5, the weight increases by 0.2 for every 3dB increase in signal-to-noise ratio, for example, in the PD-L1 fluorescence channel. At a signal strength of 25 dB, the fluorescent labeling weight is 0.83. If treatment causes a signal attenuation of 5 dB... If the value is 20dB, the fluorescent labeling weight will adaptively decrease to 0.5.
[0059] Step S22: Combine morphological features and morphological weights to obtain weighted morphological features, and combine multispectral fluorescent labeling features and fluorescent labeling weights to obtain weighted multispectral fluorescent labeling features.
[0060] For each spatial partition entity, the morphological features bound to it are multiplied by the corresponding morphological weight to generate weighted morphological features. Similarly, when generating weighted fluorescence features by combining multispectral fluorescence label features and fluorescence label weights, the feature value of each fluorescence channel is multiplied by the corresponding fluorescence label weight to output the weighted multispectral fluorescence label features.
[0061] The above weighting process needs to be fully automated through matrix multiplication algorithm and feature dimension alignment needs to be verified. Its core objective is a triple closed loop: (1) Reliability-driven feature optimization: suppressing the contribution of low-quality features, such as the area weighting value of fragmented cells dropping to 24μm when the weight is 0.2. 2 (1) Avoid misjudging as CTC and reduce false positive rate; (2) Dynamic treatment adaptability: When the morphological weight decreases during chemotherapy, the apoptotic cell characteristics are automatically weakened, focusing on the analysis of surviving CTCs, effectively improving the timeliness of treatment response assessment; (3) Multimodal fusion anti-interference: High-weight features dominate decision-making, solving the problem of noise fluorescence channel contamination in traditional average fusion, effectively improving the detection rate of EMT-type CTCs, providing anti-interference feature input for dynamic identification module, supporting drug resistance clone evolution tracking and clinical precision decision-making, and breaking through the bottleneck of dynamic monitoring distortion caused by static fusion in existing technologies.
[0062] Step S23: Combine the weighted morphological features and the weighted multispectral fluorescent labeling features to generate a fused feature vector.
[0063] First, the weighted morphological features and the weighted multispectral fluorescent marker features are aligned according to the unique identifier of the spatial partition entity. The dimension is unified by zero-filling. Then, based on the current treatment stage, the bimodal weight ratio is dynamically adjusted, and a second weighting is performed on the spliced vector to finally output the fused feature vector.
[0064] Its dynamic weights are calculated in real time by the treatment phase decision-maker, and the nonlinear correlation between features is analyzed by the Transformer encoder to generate a fusion feature vector with unified semantic encoding.
[0065] The core objectives of this fusion mechanism are threefold breakthroughs: (1) Spatiotemporal biological analysis: quantifying morphological-molecular cross-modal associations to solve the problem of phenotypic migration that traditional single-modal methods cannot capture; (2) Adaptive treatment decision-making: dynamic weighted factors respond to the treatment phase to ensure that high-value features dominate the fusion, thus advancing the drug resistance warning time by 6 weeks; (3) Interpretable driving closed loop: the fusion vector is mapped to clinically readable semantics to generate a decision instruction of "weighted PD-L1>60 and nuclear malformation>0.6 → recommended combined targeting", supporting CTC liquid biopsy from static counting to dynamic precision intervention.
[0066] Step S24: Construct a spatiotemporal knowledge graph based on the fused feature vectors, constrain the evolution path of the spatiotemporal knowledge graph by using a preset cell state transition rule base, and generate a classification decision function based on the spatiotemporal knowledge graph.
[0067] First, the fused feature vectors are input into the node encoding layer of the graph neural network to generate initial node embeddings, and then the graph structure is constructed based on the adjacency relationships of the Delaunay triangulation of spatially partitioned entities.
[0068] Next, the features of neighboring nodes are aggregated through a message passing mechanism, and the graph node states containing spatial relationships are output. A preset cell state migration rule base is accessed simultaneously, and the evolution path identifiers in the preset cell state migration rule base are mapped to graph constraint edges. When the node state meets the rule conditions, virtual edges of the corresponding evolution path are automatically added, forcing the node state to evolve along the rule path.
[0069] By employing a graph convolutional network to aggregate node features in a hierarchical manner and encoding temporal state transitions through a temporal convolutional network, a three-dimensional dynamic attribute is output, namely, the probability distribution of category attribute, subtype classification, and temporal evolution state. This effectively solves the problem that static models cannot analyze the cell state transition mechanism.
[0070] For example, the elements 0.32, -1.4 to 0.8 contained in the 128-dimensional fusion feature vector based on the unique identifier 001 of the spatial partition entity can be converted into a 256-dimensional initial node embedding vector through the node encoding layer of the graph neural network, such as generating an initial node embedding vector containing elements 0.52, -0.3 to 1.2.
[0071] Based on the adjacency relationship of the Delaunay triangulation of spatial partition entities, when the centroid distance between the unique identifier 001 and the unique identifier 002 of the spatial partition entity is less than 20 micrometers, A[1][2]=1 is marked in the adjacency matrix to indicate the existence of a connecting edge.
[0072] Next, the preset cell state migration rule base is called. The threshold condition for the EMT evolution path identifier EMT002 is that the Vimentin embedding value is greater than 0.8 and the nuclear malformation degree is greater than 0.6. If the node state meets the conditions, the unique identifier 001 of the spatial partition entity is automatically added to the virtual edge of the EMT evolution path identifier EMT002 and a path weight coefficient of 0.7 is assigned. The message passing priority of this virtual edge is improved through the graph attention mechanism, so that the contribution of the EMT evolution path features is increased when the neighborhood features are aggregated.
[0073] Finally, a hierarchical graph convolutional network was used to aggregate node features with virtual edges, and a temporal convolutional network was used to encode the node state sequences S_1 to S_4 from weeks 1 to 4 of chemotherapy. The output dynamic attribute probability distribution, such as mesenchymal type probability of 0.92 and drug resistance increase of 0.3, was generated simultaneously. When the drug resistance increase exceeds 0.25, the instruction to switch osimertinib is triggered, and the EMT evolution path identifier EMT002 is associated to provide an expandable biological mechanism description.
[0074] Among them, the preset cell state migration rule base is a structured database that stores the state migration rules of cells under specific physiological or pathological conditions, such as EMT transformation and drug resistance escalation. Its core is to transform biological knowledge into computable decision logic through digital modeling.
[0075] Step S25: Dynamically analyze the attributes of the fused feature vector using a classification decision function, and output the dynamic recognition result.
[0076] The input spatially aligned fused feature vector is analyzed by the classification decision function based on the spatiotemporal joint modeling mechanism to determine its dynamic attributes. Specifically, the spatial correlation pattern in the feature vector is first analyzed through the graph neural network layer, and the temporal drift law of the feature is tracked through the temporal convolutional network layer. Then, the three-in-one dynamic recognition result is output: cell category attribute, subtype classification and temporal evolution state.
[0077] This process ensures the accuracy of dynamic attribute parsing by constraining spatial-temporal consistency through a joint loss function, thus avoiding the bottleneck of dynamic monitoring failure and clinical disconnect caused by static models.
[0078] ③ For specific operations that can be performed in the quantitative analysis module 30, please refer to [the relevant documentation]. Figure 4 As shown, steps S30 to S32 are included: Step S30: Convert the CTC category attributes, subtype classification, and time-series evolution status in the dynamic recognition results into quantitative parameters.
[0079] Binary encoding is performed on CTC category attributes, for example, setting CTC=1, white blood cells=2, and impurities=3. One-hot encoding vectors are used for subtype classification, for example, setting epithelial type=[1,0,0], mesenchymal type=[0,1,0], and stem cell-like type=[0,0,1]. The temporal evolution status is directly quantified by percentage values, for example, setting drug resistance +30% → value 0.3, metastatic potential increase → increase value 0.25, thereby generating quantitative parameters. For example, the unique identifier 001 of the spatial partition entity corresponds to [CTC category attribute=1, subtype classification=[0,1,0], temporal evolution status=0.3]. Then, a population statistical matrix is constructed based on the quantitative parameters, such as the average increase in drug resistance in the treatment cohort=0.15, and the proportion of mesenchymal type=0.65. The evolution trend slope is calculated through the ARIMA model, such as the weekly increase slope of drug resistance β=0.05. Finally, a structured report containing dynamic quantitative indicators is output, avoiding the decision delay and execution disconnect bottleneck caused by conventional text reports.
[0080] Step S31: Calculate the quantitative parameters based on the time dimension to obtain statistical indicators.
[0081] First, quantized parameter sequences of unique identifiers for entities within the same spatial partition are aligned according to treatment cycles. For example, the drug resistance parameter sequence for the unique identifier 001 of the spatial partition entity [Cycle 1=0.1, Cycle 2=0.15, Cycle 3=0.25, Cycle 4=0.3] is then processed using ARIMA (Autoregressive Integrated Moving Average). The model (autoregressive integral moving average) fits the evolutionary trend function, extracts the trend slope and acceleration from the evolutionary trend function, and simultaneously performs statistical analysis on the quantitative parameters of all spatial partition entities at the same time point, including mean (e.g., mean resistance on week 4 μ=0.28), standard deviation (σ=0.12), and coefficient of variation (CV=0.43), generating a time-population evolution matrix, such as [week 1: μ=0.12, σ=0.08; week 4: μ=0.28, σ=0.12]. Then, dynamic indicators are calculated through a sliding window, such as the weekly average increase in resistance Δμ=0.053 and the weekly decrease in the coefficient of variation of metastasis potential ΔCV=−0.1. Finally, a set of statistical indicators in the time dimension is output, namely "resistance trend slope β=0.067 → high risk, population resistance increase Δμ>0.05 → drug-resistant clonal expansion", which solves the bottleneck of clinical response misjudgment caused by the inability of conventional static indicators to capture time-series dynamics.
[0082] Step S32: Integrate quantitative parameters and statistical indicators to generate a quantitative analysis report.
[0083] Quantitative parameters and statistical indicators are input into the dynamic indicator matrix generation module. According to clinical decision tree rules, such as drug resistance β>0.06 or Δμ>0.05 triggering a high-risk alarm, the data is classified and organized to construct a structured reporting framework that includes core dynamic indicators, such as "interstitial type proportion = 65%↑", "drug resistance slope β = 0.067 → high risk", evolutionary heatmaps, such as the weekly increase heat distribution of drug resistance, and population heterogeneity analysis, such as coefficient of variation CV = 0.43 → polyclonal drug resistance.
[0084] Next, the threshold rule engine converts the values into clinically executable instructions. For example, when β>0.06, it generates a clinically executable instruction of "Accelerated increase in drug resistance → Recommend switching to osimertinib". Combined with the natural language generation module, it outputs a dual-version report that is both machine-readable and doctor-readable. Key indicators are linked to the original data traceability link. For example, when "β=0.067" is clicked, the time-series parameter of the unique identifier 001 of the spatial partition entity can be traced. This solves the bottleneck of decision delay and execution disconnect caused by the static abstraction of conventional text reports.
[0085] The process after step S31 also includes: Step S311: Receive evolution path identifiers from a preset cell state migration rule base and integrate the evolution path identifiers into statistical indicators.
[0086] The statistical indicators are compared in real time with preset conditions in the preset cell state migration rule base, such as "β>0.06 and Δμ>0.05 → match EMT transformation path ID_EMT002". When the statistical indicators meet the preset conditions, the call command of the corresponding evolution path identifier is triggered.
[0087] The evolutionary path identifier is integrated into the statistical indicators. By using the weighting coefficient associated with the identifier, such as the EMT path weighting factor W_path=0.7, the statistical indicators are weighted and corrected. For example, after correction, the drug resistance slope β'=β×W_path=0.067×0.7=0.0469. A new "evolutionary path ID" column is added to the time-indicator matrix, such as [week=4, β=0.0469, Δμ=0.08, path ID=EMT002]. It is synchronously linked to the path knowledge graph in the preset cell state migration rule base. For example, clicking ID_EMT002 can expand the biological explanation of "Vimentin↑ / E-cadherin↓→metastasis risk↑", which solves the problem that static indicators cannot analyze the cell state migration mechanism.
[0088] Specifically, refer to Figure 5 As shown, the complete process steps that can be executed are: Step 1: Register a unique username and password to log in to the system. After logging in, the username will be displayed on the main interface and linked to key operation records. Basic medical and patient information is uniformly entered and maintained through a centralized management module.
[0089] Step 2: Create a standardized, traceable digital testing task. Specifically, select the image source for storing the cells to be tested. At this time, a dialog box will pop up for "Select Patient Subject" and "Select Submitting Doctor". After selecting the entry corresponding to the current sample, start entering metadata and generating the necessary submission records for the final report.
[0090] Step 3: As Figures 2 to 4 The corresponding steps are described above and will not be repeated here.
[0091] Step 4: After selecting the sample to be reviewed from the record list, you will enter a three-column interface. The left side displays all suspected CTC slides in a grid of thumbnails; the central main area displays high-resolution cell images, supporting stepless zoom, panning, RGB value display, electronic caliper measurement, and manual labeling; the right side provides RGB single-channel and adjustable threshold binarized images to assist in diagnosis. Extended functions include: image source tracing (automatically locating cell positions in the original large image), secondary comprehensive review (correcting model missed / false detections), and labeled example images (typical positive cell insertion reports).
[0092] Step 5: After selecting the review record, a report preview will be automatically generated (including patient information, CTC test results, reviewer, etc.). The test annotations and clinical suggestions are editable and can be saved to the knowledge base. The final report emphasizes clinical decision-making authority—the system only displays numerical comparison results (e.g., positive / negative); the reviewer must comprehensively determine the report type based on clinical information. After confirmation, export the PDF report and record a log.
[0093] Meanwhile, after the verification is completed, the verified cells can be automatically saved and organized into a structure that conforms to the system format, which facilitates the subsequent iterative training and upgrading of the system model.
[0094] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A CTC dynamic recognition system based on multimodal AI, characterized in that, include: An image processing module is used to identify each cell in a multimodal image and extract the morphological features and multispectral fluorescent labeling features of each cell; The result recognition module is used to fuse the morphological features and the multispectral fluorescent labeling features to obtain a fused feature vector, and output the dynamic recognition result of the CTC based on the fused feature vector; The quantitative analysis module is used to perform quantitative analysis on the dynamic identification results and output a quantitative analysis report.
2. The CTC dynamic recognition system based on multimodal AI according to claim 1, characterized in that, The step of identifying individual cells in a multimodal image includes: The multimodal image is spatially divided to establish spatial partition entities for each cell; After associating each spatial partition entity with corresponding multimodal attribute features, each cell in each spatial partition entity is identified.
3. The CTC dynamic recognition system based on multimodal AI according to claim 2, characterized in that, The steps of extracting the morphological features and multispectral fluorescent labeling features of each cell include: Obtain the cell coordinate set based on the spatial partition entity; The morphological features are calculated by applying spatial geometric operations in conjunction with the cell coordinate set; and, The multispectral fluorescent labeling features were calculated by applying multichannel spectral analysis in conjunction with the cell coordinate set.
4. The CTC dynamic recognition system based on multimodal AI according to claim 1, characterized in that, The step of fusing the morphological features and the multispectral fluorescent labeling features to obtain the fused feature vector includes: A first quality assessment index for evaluating the morphological features; and morphological weights generated based on the first quality assessment index; and A second quality assessment index is used to evaluate the multispectral fluorescent labeling characteristics, and fluorescent labeling weights are generated based on the second quality assessment index; By combining the morphological features and the morphological weights, weighted morphological features are obtained; and by combining the multispectral fluorescent labeling features and the fluorescent labeling weights, weighted multispectral fluorescent labeling features are obtained. The weighted morphological features and the weighted multispectral fluorescent labeling features are combined to generate the fused feature vector.
5. The CTC dynamic recognition system based on multimodal AI according to claim 1, characterized in that, The step of outputting the dynamic recognition result of CTC based on the fused feature vector includes: A spatiotemporal knowledge graph is constructed based on the fused feature vectors, and the evolution path of the spatiotemporal knowledge graph is constrained by a preset cell state transition rule base. Generate classification decision functions based on spatiotemporal knowledge graphs; The classification decision function is used to perform dynamic attribute parsing on the fused feature vector, and the dynamic recognition result is output.
6. A quantitative analysis method for CTC based on multimodal AI, characterized in that, The CTC quantitative analysis method based on multimodal AI is applied to the quantitative analysis module of the CTC dynamic identification system as described in any one of claims 1 to 5, comprising: Based on the input dynamic recognition results, quantitative analysis is performed on the dynamic recognition results, and a quantitative analysis report is output.
7. The CTC quantitative analysis method based on multimodal AI according to claim 6, characterized in that, The step of quantitatively analyzing the dynamic identification results and outputting a quantitative analysis report includes: The CTC category attributes, subtype classifications, and time-series evolution states in the dynamic recognition results are converted into quantified parameters. The quantitative parameters are calculated based on the time dimension to obtain statistical indicators; The quantitative parameters and statistical indicators are integrated to generate the quantitative analysis report.
8. The CTC quantitative analysis method based on multimodal AI according to claim 7, characterized in that, After the step of calculating the quantification parameters based on the time dimension to obtain the statistical indicators, the method further includes: Receive evolution path identifiers from a preset cell state migration rule base, and integrate the evolution path identifiers into the statistical indicators.
Citation Information
Patent Citations
Fluorescence in situ hybridization (FISH) image parallel processing and analysis method
CN106296635A
CTC image recognition method and system based on artificial intelligence
CN111652095A
Remote sensing image cross enhancement method and system based on multi-source data
CN118628379A
Knowledge mining method and system for tumor field
CN119673479A
Detection of circulating tumor cells using imaging flow cytometry
US20080317325A1
Cited By
Multi-level caching method and system based on structure enhancement prediction and reinforcement learning
CN122220262A