A biological microenvironment image analysis and classification method

Through the processing of biological microenvironment images and machine learning model analysis, the problem of in-depth disclosure of cell functional states and interactions in the prior art is solved, and the refinement of functional state quantification and quantitative presentation of communication networks are realized, which improves the reliability and standardized application of analysis results.

CN120259303BActive Publication Date: 2025-09-02SHANGHAI XUNYUAN BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510737365.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-02
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The existing technology cannot deeply reveal the true functional status of cells in the biological microenvironment and the interaction between cells. The lack of quantitative and intuitive evaluation methods leads to poor repeatability of the analysis results and is difficult to standardize application in different studies or clinical practice.

Method used

By processing the biological microenvironment images, segmenting cell components, extracting multi-parameter initial features, using machine learning models to infer the functional substates of cells, compute the potential field map of biomolecules interaction, and generate biological microenvironment state descriptors, and classifying them in combination with a generative machine learning model.

Benefits of technology

It realizes the quantification of the refined functional state of cells in a specific environment, quantitatively displays the intercellular communication network, generates objective deviation characteristics, improves the accuracy and stability of the analysis results, and provides a basis for accurate diagnosis and personalized treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259303B_ABST
    Figure CN120259303B_ABST
Patent Text Reader

Abstract

The present invention provides a method for analyzing and classifying biological microenvironment images, which relates to the field of cell image analysis technology. The method includes processing an input biological microenvironment image, segmenting cell components, and extracting their multi-parameter initial features; using a first machine learning model to infer refined functional substates beyond traditional cell types based on the intrinsic characteristics of the cell and local microenvironment information; combining knowledge of biomolecular interactions to calculate and generate spatially resolved biomolecular interaction potential field maps to quantify the communication potential between cells; forming a comprehensive set of state descriptors; and inputting this descriptor set into a second machine learning classification model to generate accurate classification results for the biological microenvironment. The present invention realizes functional, quantitative, and interpretable analysis and classification of the biological microenvironment, significantly improving the depth, objectivity, and accuracy of the assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cell image analysis, and in particular to a biological microenvironment image analysis and classification method. Background Art

[0002] The biological microenvironment is a complex ecosystem composed of multiple cells and components such as the extracellular matrix. Its state plays a decisive role in the occurrence, development, and treatment response of major diseases such as cancer. Currently, analysis of the biological microenvironment mainly relies on histopathological staining and multiple imaging technologies, combined with computer vision for quantitative analysis.

[0003] However, existing technologies have the following major limitations:

[0004] Existing methods primarily focus on cell identification, counting, and morphological measurement, describing only the static composition of the microenvironment and failing to fully understand the true "functional state" of cells within their specific environment. Assessments of intercellular interactions often rely solely on analyzing spatial proximity, an overly simplistic approach that fails to quantitatively and intuitively capture the complex communication networks of specific strength and range mediated by specific ligand-receptor pairs and other biomolecules. Determination of the overall microenvironmental state, such as the qualitative characterization of a tumor microenvironment as "cold" or "hot," relies heavily on subjective experience and lacks unified, quantitative criteria. This results in poor reproducibility and hinders standardized application across diverse studies or clinical practice. Summary of the Invention

[0005] Technical problems solved

[0006] In view of the deficiencies of the prior art, the present invention provides a biological microenvironment image analysis and classification method, which solves the problems of the prior art.

[0007] Technical Solution

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: A biological microenvironment image analysis and classification method comprises the following steps:

[0009] Sp1. Process the biological microenvironment image, segment the cell components in the image, and extract the multi-parameter initial features of each cell component;

[0010] Sp2. Based on the multi-parameter initial features and the local microenvironment information of the cells, a first machine learning model is used to infer the refined functional substates of each cellular component, wherein the functional substates are characterized as states that go beyond basic cell types and are related to the biological behavior or potential of the cells;

[0011] Sp3. Based on the functional substates, spatial position information, and predefined biomolecular interaction knowledge of the cellular components, calculate and generate a spatially resolved biomolecular interaction potential field map covering the image analysis area, wherein the biomolecular interaction potential field map quantifies the potential or strength of specific biomolecule-mediated interactions occurring at each point in space;

[0012] Sp4. Extract and integrate at least two of the following sets of features to form a set of biological microenvironment state descriptors:

[0013] Sp4.1. Statistical or spatial distribution characteristics derived from inferred cellular functional substates;

[0014] Sp4.2. Quantitative features derived from the biomolecular interaction potential field map;

[0015] Sp4.3. The deviation features obtained by comparing the biological microenvironment image or its derivative representation with a preset reference microenvironment state are compared using a pre-trained generative machine learning model;

[0016] Sp5. Input the biological microenvironment state descriptor set into a pre-trained second machine learning classification model to generate a classification result for the biological microenvironment image.

[0017] Preferably, the first machine learning model for inferring functional substates in Sp2 is an unsupervised learning model or a self-supervised learning model, which identifies functional substates by learning the latent space representation of multi-parameter characteristics of cells.

[0018] Preferably, the biomolecular interaction knowledge in the Sp3 includes ligand-receptor pair information, and the calculation of the interaction potential field map includes the spatial proximity, expression level and functional substate of ligand-expressing cells and receptor-expressing cells.

[0019] Preferably, the pre-trained generative machine learning model in Sp4 is a generative adversarial network or a variational autoencoder, and the deviation feature includes a difference measure between the image to be analyzed and the corresponding reference state projection generated by the generative model, or a score of the image to be analyzed under the generative model discriminator.

[0020] Preferably, the method further comprises Sp6:

[0021] Based on the second machine learning classification model and the biological microenvironment state descriptor set, a computational attribution method is used to identify key features or feature combinations that have a significant impact on the classification results, and this identification result can be selectively used to guide the training of the second machine learning classification model or explain its output.

[0022] Preferably, the computational attribution method is configured to: identify the global importance of the biological microenvironment state descriptors that contribute significantly to the classification result;

[0023] The contribution of the biomolecular interaction potential field map in a specific spatial region is traced back to the precise spatial configuration and density threshold of one or more specific cellular functional substates that cause changes in the potential field intensity and pattern in that region, thereby revealing the spatially specific cellular-level functional synergistic or antagonistic pathways that drive the microenvironment classification state.

[0024] Preferably, the second machine learning classification model in Sp4 is a heterogeneous multi-branch deep neural network, and the heterogeneous multi-branch deep neural network includes:

[0025] At least one graph neural network branch is specifically used to process cell-level graph structure data constructed by the cell functional substates and their spatial adjacency relationships to learn cell community patterns and local cell niche characteristics;

[0026] at least one convolutional neural network branch specifically configured to process the spatially resolved biomolecular interaction potential field map and the spatial deviation feature map generated by the generative machine learning model to extract multi-scale spatial texture and intensity patterns thereof;

[0027] The deep feature vectors extracted by different branches are integrated through a learnable cross-modal attention fusion module, which can adaptively assign confidence weights to various features of different biological microenvironment samples or different spatial regions, and finally a classification head outputs the classification results.

[0028] Preferably, the classification result classifies the biological microenvironment image into one of a plurality of predefined "functional microenvironment prototypes", each prototype being defined and distinguished by the following set of integrated parameters:

[0029] Characteristic abundance profiles and spatial co-occurrence patterns of key cellular functional substates;

[0030] The characteristic intensity distribution and topology of the dominant biomolecular interaction potential field;

[0031] A multidimensional vector of deviations from a specific reference microenvironmental state, quantified by a deviation signature; prototypes are associated with the molecular mechanism response signature at the microenvironment-mediated level of a specific disease biological behavior or treatment.

[0032] Preferably, the output of Sp4 further includes generating and presenting at least one of the following computationally enhanced and contextually associated visual analysis maps:

[0033] The spatial distribution map of cell function substates shows the driving cell function substates or their spatial clustering areas that contribute most to the classification results by overlaying a significant heat map;

[0034] Comparative visualization maps, which display the biomolecular interaction potential field map of the current sample side by side or differentially with the average potential field map of the functional microenvironment prototype or the potential field map of the reference sample, highlighting the statistically significant differences in interaction patterns;

[0035] Based on the deviation features generated by the generative adversarial network, spatial back-projection is performed on the cellular or regional scale of the original image to generate a deviation map, indicating the structural or functional units that differ most significantly from the reference state and annotating the specific dimensions of their deviation.

[0036] Preferably, the biological microenvironment image is high-content imaging data with single-cell or subcellular resolution and containing at least 15 distinguishable molecular detection channels, including multiple immunofluorescence imaging, imaging mass spectrometry flow cytometry imaging and high-dimensional spatial omics imaging data; the multi-parameter information provided by a large number of molecular channels is the basis for the first machine learning model to reliably infer refined functional substates and execute them in Sp2, and to support the precise calculation of the biomolecular interaction potential field and execute it in Sp3.

[0037] Beneficial effects

[0038] The present invention provides a method for analyzing and classifying biological microenvironment images. It has the following beneficial effects:

[0039] 1. Through deep modeling, this invention can infer refined "functional substates" beyond cell types, accurately revealing the true biological behavior of cells in specific environments; quantify and visualize the "interaction potential field" mediated by specific molecules, intuitively showing the communication network and intensity between cells; and explore and identify the key factors that drive the formation of microenvironmental states.

[0040] 2. This invention incorporates a generative machine learning model that quantitatively compares the sample under analysis with a "healthy" or "ideal" reference state to generate an objective "deviation signature." This signature precisely quantifies the degree and direction of abnormality in the sample, completely freeing it from the constraints of subjective judgment. Based on this objective standard and integrating multidimensional functional information, this method achieves more accurate and stable classification results, providing a reliable technical basis for achieving precise diagnosis, prognosis, and guiding personalized treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a system function module diagram of the present invention;

[0042] Figure 2 It is a system flow chart of the present invention. DETAILED DESCRIPTION

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. Specific embodiment one:

[0045] like Figures 1 to 2 As shown, the entire system of a biological microenvironment image analysis and classification method includes a data and model library, a data input and preprocessing module, an image segmentation and feature extraction module, a cell function substate inference module, an interaction potential field calculation module, a deviation feature extraction module, a state descriptor integration and classification module, a computational attribution analysis module, and a result generation and visualization module.

[0046] A biological microenvironment image analysis and classification method comprises the following steps:

[0047] Sp1. Process the biological microenvironment image, segment the cell components in the image, and extract the multi-parameter initial features of each cell component;

[0048] The core technical feature of this step is to achieve high-precision and automated segmentation of various cell components in complex biological microenvironment images, mainly including cell nuclei, cytoplasm, cell membranes, and extracellular matrix areas that need to be analyzed in some cases. After the segmentation is completed, a comprehensive, multi-dimensional set of initial feature parameters will be extracted from each independently identified cell entity. These initial feature parameters not only cover the fine morphological information of the cell, such as size, shape, edge smoothness, etc., but also include its optical density or fluorescence intensity information under different imaging channels, as well as preliminary texture features that can reflect changes in subcellular structures. This series of operations lays a solid and comprehensive data foundation for the subsequent Sp2 step to carry out more complex cell refinement functional substate inference.

[0049] This step involves numerous crucial technical considerations. First, a standardized image preprocessing process is essential, ensuring the stability and comparability of subsequent analytical results. This requires effective mitigation of uneven brightness distribution, random and systematic noise, various imaging artifacts, and interference caused by tissue folding or bubbles, often found in images from different imaging devices, experimental batches, or sample preparation processes. A series of standardized preprocessing operations, such as background correction, flat-field correction, and denoising, significantly improve image quality. Second, given the inherently high heterogeneity within biological microenvironments, where cell size, morphology, and spatial density vary significantly across regions, the selection or design of segmentation algorithms must ensure they are flexible and can effectively handle this multi-scale variation. The algorithm should be able to accurately segment both sparsely distributed large cells and densely packed small cells within the field of view. Furthermore, the number of feature parameters extracted from each cell in the initial stage is enormous, reaching hundreds or even thousands of dimensions, inevitably including redundant information or features that are not particularly relevant to the specific analytical objective. Therefore, before entering the subsequent modeling steps, it is necessary to carefully consider the use of preliminary feature importance assessment methods, such as those based on statistical tests or built-in evaluation mechanisms of machine learning models, or the application of effective dimensionality reduction techniques, such as principal component analysis, to refine the feature set, reduce the computational burden, and improve model performance. At the same time, the entire image processing and feature extraction process should have good compatibility and scalability at the system design level, and be able to flexibly adapt to and efficiently process input data from different advanced imaging technologies. These technologies include but are not limited to multiplexed immunofluorescence imaging technologies such as cyclic immunofluorescence (CyCIF) or multiplexed ion beam imaging (MIBI), high-parameter imaging mass cytometry, and high-resolution hematoxylin and eosin-stained full-field digital scanning imaging that does not require specific labeling but can provide rich morphological information.

[0050] If the cell segmentation model used in this step, particularly one based on deep learning, is trained using a supervised learning paradigm, then high-quality training data collection is the cornerstone of model performance. This phase requires systematically assembling a broadly representative dataset of biological microenvironment images. These datasets should comprehensively cover the various cell types, tissue structures, pathological states, and potential technical variations anticipated in the research objectives. Image annotation is a core component of data collection and must be performed by pathologists with a strong professional background or by image analysts who have undergone rigorous professional training and demonstrated consistent interpretation. Annotation accuracy must reach pixel-level precision, requiring precise delineation or region filling for every nuclear and cytoplasmic boundary in the image, as well as for cell membrane structures discernible in certain high-resolution images. This provides the "gold standard" or "ground truth" data necessary for training the segmentation model. In these cases, to fully learn and achieve good convergence for complex deep learning models, the number of labeled data needs to reach hundreds to thousands of high-resolution raw image slices or fields of view, containing tens of thousands or even millions of cell instances to ensure model generalization.

[0051] Within this Sp1 step, the specific data processing process is rigorous and meticulous. First, the input data is the original, digitized biological microenvironment image without further processing. Common storage formats for these images include the general TIFF format, the SVS format for full-field digital pathology scans, or special formats such as ISYNTAX. After receiving the original image, a detailed preprocessing process is immediately launched. Intensity normalization in the preprocessing operation is an important means to eliminate the brightness differences between different images. The pixel intensity values ​​of each imaging channel are uniformly scaled to a floating point range of 0 to 1 or an integer range of 0 to 255 by linear stretching or compression. For more complex intensity drift situations, nonlinear normalization methods such as histogram matching based on reference images or global statistics are used. Noise suppression is a key link in preprocessing, and the choice of method needs to be based on the specific noise characteristics of the image. For common random noise, classic smoothing operators such as Gaussian filtering and median filtering can be used. For more complex noise patterns, methods such as non-local means filtering, which better preserve edge details, can be used. Alternatively, a more advanced strategy is to apply deep learning denoising models trained for specific imaging system noise patterns, such as the Noise2Void network based on self-supervised learning or the CARE network trained on paired noisy / denoised images. These deep models can effectively remove noise while maximally preserving the true biological structural information in the image. For brightfield microscopy images, especially hematoxylin and eosin-stained pathology images, color correction and white balance adjustment are required before quantitative analysis to eliminate color deviations caused by variations in light source color temperature or staining differences. A common and challenging issue for multichannel fluorescence images is cross-talk or fluorescence bleed-through caused by the partial overlap of emission spectra of different fluorescent dyes. This requires the use of precise algorithms, such as linear unmixing algorithms based on spectral characteristics or fluorescence signal compensation algorithms based on physical models of the dye combination.

[0052] In the image segmentation stage, accurate segmentation of cell nuclei is the cornerstone and primary task of the entire cell component identification process. Based on the U-Net architecture, its algorithm details include a symmetrical encoder-decoder structure. The encoder part is composed of a series of convolutional layers and maximum pooling layers alternating to extract higher-level abstract semantic features through layer-by-layer convolution, while the pooling layer gradually reduces the spatial resolution of the feature map. The decoder part gradually restores the spatial resolution of the feature map through upsampling operations, transposed convolution or bilinear interpolation, and compares it with the feature map of the corresponding layer of the encoder part through<x_bin_118> The decoder uses skip connections for concatenation or additive fusion. This skip connection mechanism enables the decoder to simultaneously utilize shallow, high-resolution detail features and deep, high-semantic features from the encoder, thereby accurately locating the object boundary. The activation function of neurons in the model generally uses the rectified linear unit (ReLU) or its improved LeakyReLU to accelerate model convergence, alleviate the vanishing gradient problem, and introduce nonlinearity.

[0053] The task of cytoplasm and cell membrane segmentation is performed after the cell nucleus is accurately segmented. If the original multi-channel fluorescence image contains a specific fluorescent marker channel for the cytoplasm or cell membrane, a deep learning semantic segmentation model similar to that for cell nucleus segmentation, U-Net, is applied again. At this time, the obtained cell nucleus segmentation mask image is used as additional input channel information, or as a spatial prior knowledge to guide the model to more accurately identify and segment the corresponding cytoplasm region or outline the cell membrane boundary. Another strategy is to use a morphological expansion algorithm based on the segmented cell nucleus when there is no specific cytoplasm / membrane marker or the marker signal is weak. Starting from the center point of each identified cell nucleus, a controlled synchronous region expansion or watershed expansion is performed outward, and the stopping boundary of the expansion is limited by the boundary signals of other neighboring cells, the estimated average cell size, or in some cases, the weak fluorescence intensity gradient of the cell membrane marker that can be used.

[0054] After completing the precise segmentation of all target cell components, the automated extraction of multi-parameter initial features begins. The extracted morphological features should comprehensively describe the geometric and shape characteristics of the cell, including: the area and perimeter of each cell nucleus, cytoplasm, and the entire cell; the lengths of the major and minor axes, which describe the degree of cell extension, and their ratio; the circularity or shape factor, which reflects the degree to which the cell is close to a circle; the eccentricity, which describes the degree of cell ellipse; the equivalent diameter of the cell; the ratio of the convex hull area, which describes whether the cell boundary is concave inward, to the area of ​​the cell itself, i.e., solidity; and a series of shape descriptors that can describe the complexity of the shape and are invariant to rotation, translation, and scaling. The extraction of optical density or fluorescence intensity features requires the precise quantification of multiple intensity statistics in each imaging channel for each segmented cell region, i.e., the nuclear region, cytoplasmic region, and cell membrane region. These statistics should include: the mean value of pixel intensity within the region, the median value (more robust to outliers), the standard deviation or coefficient of variation of intensity (reflecting intensity uniformity), the integrated sum of all pixel intensities (reflecting the total amount of markers), and specific percentiles that can reflect the skewness and tailing of the intensity distribution, the 25% quantile and the 75% quantile. The extraction of texture features aims to capture the subtle pattern differences of subcellular structures in the image and can be applied to the internal area of ​​the cell nucleus or the cytoplasm. Commonly used texture features include a series of statistics calculated based on the gray-level co-occurrence matrix, such as energy, contrast, correlation, entropy and local homogeneity; the spatial position characteristics of each identified cell in the image coordinate system, that is, the precise x and y coordinates of its center of mass (for three-dimensional imaging data, the z coordinate should also be included), must also be accurately recorded.

[0055] The final data output of this Sp1 step consists of two main components. The first is a cell segmentation mask image. In this image, each successfully segmented and identified individual cell is assigned a globally unique identifier, known as a cell ID. This mask image facilitates subsequent single-cell feature tracking, neighborhood analysis, and various visualization operations. The second component is the generation of a structured, detailed feature data table. This data table is stored in a comma-delimited CSV file format or a more specialized format such as HDF5, and can also be directly stored in a relational or non-relational database. Each row of the data table represents an individual cell instance in the image, and each column corresponds to the quantitative measurement of a specific multi-parameter initial feature extracted for that cell in the Sp1 step. In certain biological microenvironment analysis applications, where analysis of the structure and composition of the extracellular matrix is ​​equally important, this Sp1 step also outputs a mask image of the segmented and identified extracellular matrix region and its associated quantitative features.

[0056] Sp2. Based on the multi-parameter initial features and the local microenvironment information of the cell, a first machine learning model is used to infer the refined functional substates of each cellular component. Functional substates are characterized as states related to cellular biological behavior or potential that go beyond basic cell types.

[0057] The core technical feature of this step is the use of machine learning techniques, particularly unsupervised or self-supervised learning methods, to learn and automatically identify more refined and dynamic cellular functional states, namely functional substates, from the high-dimensional single-cell data produced by the Sp1 step, beyond the traditional definition of cell types based on a few markers or simple morphology. These functional substates can more deeply reveal the actual biological behavioral potential, activation state, differentiation stage, or stress response of cells in specific biological microenvironments. Another core feature is the integration of the cell's intrinsic characteristics, namely the comprehensive morphological parameters and multi-molecular expression profiles derived from Sp1, and information about the cell's immediate external environment, including the type, density, and functional state of its neighboring cells, making the inference of functional substates more physiologically relevant and context-dependent. In this way, this step is committed to achieving a deep characterization of cellular heterogeneity and dynamic plasticity in biological microenvironments.

[0058] The first machine learning model is a self-supervised model based on deep learning and requires training. Its data collection phase mainly relies on the large-scale single-cell feature data output by the Sp1 step. For the self-supervised learning paradigm, in order to subsequently cluster the learned cell embedding representations and functionally annotate the formed clusters, an independent reference cell dataset with partial cell types or rough functional annotations (based on pre-classification results of several key surface marker combinations) is also required, or the knowledge base of domain experts is relied upon for subsequent biological significance interpretation. For some weakly supervised learning methods, it is explicitly required that a portion of the cell data has such rough functional labels to guide the learning direction of the model.

[0059] In this Sp2 step, the detailed data processing flow is as follows:

[0060] The input data is the structured single-cell feature table output by the Sp1 step, in which each row represents a cell and the columns contain all the multi-parameter initial features of the cell; at the same time, the precise spatial coordinate information of each cell in the original image is also required to calculate the local microenvironment characteristics.

[0061] For each cell, local microenvironmental features are calculated based on its spatial coordinates, the preliminary types of other cells around it, and the cell features extracted by Sp1. These microenvironmental features include: the type composition ratio of the target cell's K nearest neighbor cells; the number, density, or proportion of cells of different types and specific functional subtypes within a specific radius centered on the target cell (the radius can be set as a multiple of the average cell diameter, with a default of 50 microns or 100 microns); and the Euclidean distance between the target cell and its nearest specific type of cell (including tumor cells and vascular endothelial cells).

[0062] After calculating and merging the intrinsic cell features and local microenvironmental features, the entire expanded feature set needs to be standardized. Z-score normalization is applied so that each feature has a mean of 0 and a standard deviation of 1, or all feature values ​​are linearly scaled to a similar numerical range, such as between 0 and 1, to eliminate the impact of different feature scale differences on subsequent distance calculations or model training. Before input into the machine learning model, feature selection or dimensionality reduction operations are selectively performed based on the specific model requirements and data characteristics. Principal component analysis is used to reduce feature dimensions while retaining the main variance information, allowing for visual exploration of high-dimensional cell data and assisting in the identification of potential cell cluster structures.

[0063] The first machine learning model for inferring refined functional substates of cells uses a self-supervised learning model to first learn an effective low-dimensional embedding representation of cell features and then cluster them in this embedding space. Deep generative models such as variational autoencoders are used to learn the latent space representation of cells.

[0064] The model details of VAE include: an encoder neural network (which is an MLP), which maps the input cell feature vector to the parameters of a normal distribution in the latent space, namely the mean vector and the logarithmic variance vector.

[0065] In the sampling step, a latent vector is randomly sampled from the parameterized normal distribution according to the mean and variance of the encoder output.

[0066] The decoder neural network (also an MLP) attempts to reconstruct the original cell feature vector from the sampled latent vector.

[0067] The loss function of VAE consists of two parts: one is the reconstruction loss, which is used to measure the difference between the reconstructed features output by the decoder and the original input features, using mean square error or binary cross entropy; the other part is the KL divergence, which acts as a regularization term to penalize the difference between the distribution of the latent space and a preset prior distribution (standard normal distribution), encouraging the latent space to have good structure and smoothness.

[0068] The acquisition of functional substates is similar to contrastive learning. After VAE training is completed, the cell latent vector (which is the mean vector of the distribution) generated by its encoder part is used for subsequent clustering analysis.

[0069] The final data output of this Sp2 step is an updated cell feature table. In this table, a new column of information is added for each cell, namely its assigned "functional substate ID" or specific substate label (including "activated macrophage substate M1 type", "exhausted CD8+ T cell substate", etc. These labels are obtained by manual annotation combined with biological knowledge after clustering is completed). In addition, to facilitate biological interpretation and verification, the average feature spectrum of each identified functional substate is also output, that is, the average or median value of all cells in the substate on each initial feature dimension, which helps to understand the key differences between different substates. At the same time, a visualization image of the functional substate in a two-dimensional or three-dimensional dimensionality reduction space such as UMAP or t-SNE is also generated to intuitively show the separation and relative relationship of different substates.

[0070] The first machine learning model for inferring functional substates in Sp2 is an unsupervised learning model or a self-supervised learning model, which identifies functional substates by learning the latent space representation of multi-parameter features of cells.

[0071] This includes selecting the dimensionality of the learned latent space. The dimensionality of the latent space is a key hyperparameter. A dimensionality that is too low causes cells of different functional substates to become mixed up in the embedding space, losing important discriminative information. A dimensionality that is too high, on the other hand, fails to effectively learn the intrinsic low-dimensional manifold structure of the data and may even introduce noise, resulting in poor subsequent clustering results. The optimal latent space dimensionality is determined through a series of experiments, evaluating the change in reconstruction error at different dimensionalities, the clarity and stability of downstream clustering results (using internal evaluation metrics such as the silhouette coefficient and the DBI index), and the performance of classification tasks on partially labeled data.

[0072] The results produced by self-supervised models, particularly clustering results, are sometimes sensitive to parameters such as model initialization, optimizer selection, and learning rate. In applications, measures should be taken to assess the stability of clustering results, such as by repeating training and comparing cluster consistency, or by adopting ensemble clustering. Furthermore, in-depth analysis and verification of the biological interpretability of the ultimately identified functional substates are crucial. Furthermore, if the analyzed cell data originate from different experimental batches, different samples, or different treatment conditions, batch effects can mask true biological differences. In such cases, it is necessary to explicitly consider removing or mitigating the impact of batch effects during feature extraction or model training. Data alignment algorithms should be used to correct the latent space, or regularization terms should be introduced into the loss function design for self-supervised learning to combat batch effects.

[0073] Application details of the self-supervised contrastive learning algorithm:

[0074] For multi-parameter cell feature data in tabular form, the choice of data augmentation strategy is crucial. An effective augmentation method should be able to create sufficiently diverse "views" of a cell while preserving its core identity (i.e., without changing its functional substate).

[0075] The encoder network architecture design requires a balance between expressiveness and computational efficiency. For biological microenvironment datasets containing millions of cells, each with hundreds of initial features, a multilayer perceptron with 2 to 4 hidden layers, each containing 128 to 512 neurons, is sufficient for the encoder network. Rectified linear units are used as the activation function between hidden layers, along with batch normalization layers to accelerate training and improve model stability and generalization.

[0076] The projection head network is a relatively simple, small MLP consisting of one hidden layer and one output layer. Its output dimension is typically set between 64 and 256. It further nonlinearly transforms the embedding vectors generated by the encoder into a space specifically used for computing the contrastive loss. The temperature parameter τ is a key parameter in the NT-Xent contrastive loss function. It controls the loss function's sensitivity to negative samples and, in other words, determines how strongly the model pushes against different cellular embedding representations. It requires careful adjustment based on the characteristics of the specific dataset, and its empirical value generally ranges from 0.05 to 0.5. The model is trained using large batch sizes, ranging from 512 to 4096 or even larger, to ensure that each positive sample is provided with a sufficient number and diversity of negative samples for comparison in each iteration. The optimizer is either Adam or AdamW, with an adaptive learning rate algorithm. The more stable LARS optimizer is used when training with very large batch sizes.

[0077] Model construction involves pre-training and fine-tuning phases. Ideally, self-supervised pre-training is performed on a broad and diverse dataset of cell features from a variety of tissue types, disease states, and even species. This pre-trained encoder learns more general and robust representations of cell features, capturing common patterns in cell state changes. For specific research questions or datasets, the encoder network is further fine-tuned using the target dataset based on this pre-trained model to better adapt it to the characteristics of the target data. After the encoder is trained or fine-tuned, it can be used to convert the raw high-dimensional feature vectors of all cells into more information-dense and semantically rich low-dimensional latent space embeddings. Applying graph clustering algorithms to these high-quality embeddings can uncover non-spherical, more complex cell cluster structures. The final number of clusters (i.e., the number of functional substates) is determined through a comprehensive evaluation of internal cluster evaluation metrics (such as the silhouette coefficient, Calinski-Harabasz index, and Davies-Bouldin index), combined with intuitive observation of the dimensionality reduction visualization results and the biological prior knowledge of domain experts.

[0078] Sp3. Calculate and generate a spatially resolved biomolecular interaction potential field map covering the image analysis area based on the functional substates and spatial location information of cellular components and predefined biomolecular interaction knowledge. The biomolecular interaction potential field map quantifies the potential or strength of specific biomolecule-mediated interactions at each point in space.

[0079] Construction and dynamic screening of a comprehensive and accurate biomolecular interaction knowledge base. The biomolecular interaction knowledge base contains known, experimentally verified ligand-receptor pairs, cytokine-receptor pairs, adhesion molecule pairs, etc., and it is best to be able to associate them with which cell types or in which cell functional states they are specifically and highly expressed. It is necessary to screen out the most relevant groups of interaction pairs from the huge knowledge base based on the biological background of the current study for subsequent potential field calculations. Secondly, it is necessary to combine the fluorescence intensity characteristics of each cell in the corresponding molecular marker channel extracted in the Sp1 step, or the definition of the functional substate itself contains a description of the expression status of these key communication molecules.

[0080] The detailed data processing flow of Sp3 is as follows:

[0081] The main input data include: the functional substate label output for each cell in step Sp2; the precise spatial coordinates (x, y, z) of each cell recorded in step Sp1; the raw expression profile data for each molecular marker channel extracted in step Sp1, particularly the expression levels of markers known to correspond to specific ligands and receptors; and a structured, predefined knowledge base of biomolecular interactions, which should at least contain the ligand name, receptor name, whether they form known interaction pairs, and characteristic information on which cell types or functional substates are highly or poorly expressed. During the data preparation phase, for each cell, combining its inferred functional substate with its raw molecular expression profile data, its potential to serve as a "source" cell for a specific ligand or a "sink" cell for a specific receptor is accurately determined. This involves setting a reasonable expression threshold for the ligand or receptor expression level of the cell, or determining whether the functional substate to which the cell belongs is known to be associated with high expression of the ligand / receptor.

[0082] The interaction potential field calculation process is performed independently for each specific ligand-receptor pair that the researcher is interested in. The first step is to identify the source cells and sink cells. Based on the results of the above data preparation, all cell populations in the image that significantly express the currently analyzed ligand L are screened and recorded as the source cell set At the same time, all cell populations that significantly express the corresponding receptor R were screened and recorded as sink cell collections The second step is to calculate the ligand potential field For each pixel point x in the image space (or for computational efficiency, calculated on a downsampled regular grid), its perceived pixels come from all surrounding source cells. The cumulative signal intensity sum of the released ligand L is mathematically modeled as follows: ; Represents the effective intensity or probability of the i-th source cell expressing or releasing ligand L. This value is directly quantified by the fluorescence intensity of the corresponding ligand marker in the cell and is further quantified by the functional substate to which it belongs. modulated (a certain “activated” functional substate has a higher ligand release efficiency). is the precise spatial location of the i-th source cell. is a spatial kernel function that describes the ligand L from the source cell After being emitted, the signal strength spreads in space or its ability to act attenuates with distance. is the characteristic parameter of the kernel function. For the Gaussian kernel function , Represents the effective radius or standard deviation of the ligand interaction. The specific form and parameters of the kernel function KL The choice of the receptor potential field should be based on an understanding of the biophysical properties of the ligand being studied, such as its molecular size, diffusion coefficient in tissue, and whether it is easily degraded or trapped by the matrix. More complex kernel functions can also be considered, such as anisotropic models that can simulate intercellular barrier effects or fluid dynamics within tissues. The third step is to calculate the receptor potential field. Similar to the calculation logic of the ligand potential field, for each pixel (or grid point) x in the image space, all the cells around it The integrated responsiveness or sum of the perceived strength of the expressed receptors R to ligand signals is modeled as: ; represents the effective intensity of receptor R expressed by the jth sink cell or its potential to bind to the ligand, which is also quantified by its corresponding fluorescence intensity and its functional isoform Modulated. is the spatial position of the jth stem cell. is a spatial kernel function that describes the spatial range in which receptor cells can effectively perceive ligand signals in their surrounding environment. is its corresponding characteristic perception radius. The fourth and final step is to calculate the final interaction potential field This potential field represents the overall potential for a specific LR interaction to occur at a spatial position x, which is integrated with the ligand accessibility at that position through a function f and receptor accessibility What we get: There are many forms of function f to reflect different biological interaction logics: the product model, PLR(x)=SL(x)⋅SR(x), which assumes that the contributions of ligand and receptor are independent and jointly determine the interaction strength. Or the "minimum restriction model", , which reflects the "barrel effect" that the strength of the interaction is limited by the weaker of the two. A model that is more consistent with the biochemical reaction kinetics is the Michaelis-Menten kinetic model based on the receptor saturation effect. , where Km is the Michaelis constant, representing the ligand concentration at which the interaction rate reaches half its maximum. In more advanced modeling, the function f is even a small, pre-trained neural network that takes the position and As input, the interaction potential is output by learning a nonlinear mapping relationship , thus being able to simulate more complex, synergistic or antagonistic interaction logics.

[0083] Fast Fourier transforms are used to convert spatial convolutions into frequency-domain products, which are then converted back to the spatial domain using an inverse FFT. This significantly improves the computational efficiency of large-scale potential field maps. For very large image data, a strategy is also employed to divide the image into blocks and parallelize the potential fields of each sub-region before stitching them together. Modern graphics processors, with their powerful parallel computing capabilities, are also well-suited for accelerating such computationally intensive pixel- or grid-based tasks.

[0084] The final data output of this Sp3 step is a series of raster image files. These files are stored in multi-layer TIFF format or other image formats that support multi-dimensional arrays, where each layer (or each independent file) accurately represents the interaction potential field map of a specific ligand-receptor pair (L-Rpair) in the entire image analysis area. The value of each pixel in the image directly reflects the potential size or expected strength value of the specific LR interaction occurring at that spatial location. Along with these potential field maps, a metadata file or record is also required to clearly describe the name of the specific ligand-receptor pair corresponding to each potential field layer and the key parameters used in the calculation process (kernel function type, action radius, expression threshold, etc.) to facilitate subsequent analysis, interpretation and result reproduction.

[0085] The biomolecular interaction knowledge in Sp3 includes ligand-receptor pair information, and the calculation of the interaction potential field map includes the spatial proximity, expression level and functional substate of ligand-expressing cells and receptor-expressing cells.

[0086] The potential field model is constructed by meticulously and quantitatively integrating multiple key biological factors, making it closer to real biological processes. These factors include: spatial proximity, that is, the physical distance between the ligand-expressing cells that send the signal and the receptor-expressing cells that receive the signal; expression level, that is, the actual expression abundance of the relevant ligand and receptor molecules on a single cell; and the functional substate of the cell, that is, the specific biological state of the cell inferred in the Sp2 step, because cells of different functional substates, even if they express the same level of ligand or receptor, have significant differences in their ability to release, present, or respond to these molecules.

[0087] Regarding the quantification and integration of expression levels, the original fluorescence intensity (or other equivalent expression measurement) of each cell in the corresponding ligand or receptor molecule marker channel extracted in the Sp1 step needs to be calibrated and normalized, and then used as the calculation of the cell as the signal source ( ) or Signal Sink ( ) is a core parameter of expression potential. An expression threshold needs to be set to distinguish between positive and negative expressing cells, or a continuous expression value needs to be used. Regarding the regulatory effect of functional substates on expression, the cellular functional substate information inferred in the Sp2 step needs to be correlated with the cell's ligand / receptor expression potential. A substate defined as "highly activated effector T cells" is known to secrete higher levels of a certain cytokine (ligand) or express a higher density of a certain chemokine receptor than the "resting T cell" substate. This regulatory relationship is achieved by assigning different expression efficiency weighting factors to different functional substates. These weighting factors are derived from published literature data, public expression profile databases, or indirectly learned through statistical analysis of the average expression profiles of cells of different functional substates in the current dataset. Regarding the modeling of spatial proximity, although the aforementioned kernel function K has implicitly included the influence of spatial proximity through its characteristic of attenuation with distance, in certain specific biological scenarios, when the study is limited to direct physical contact between cells or very short-range paracrine signals, more explicit distance constraints are also introduced into the model, or the form of the kernel function is adjusted to emphasize close-range effects.

[0088] The raw molecular expression data for each cell is tightly integrated with the functional substate labels assigned in the Sp2 step. Before calculating the interaction potential field for a specific ligand-receptor pair, a predefined biomolecular interaction knowledge base is retrieved to identify which known cell types or functional substates identified in the current dataset specifically express or highly express the ligand and receptor. This information is used to pre-screen potential source and sink cell populations, thereby narrowing the actual scope of subsequent potential field calculations and reducing noise introduced by nonspecific low-level expression. The generated potential field map is more focused on interactions with clear biological significance.

[0089] Sp4. Extract and integrate at least two of the following sets of features to form a set of biological microenvironment state descriptors.

[0090] Sp4 systematically extracts and organizes the multifaceted, multidimensional information generated previously into a comprehensive "descriptor set" that fully characterizes the overall state of the current biological microenvironment. It integrates two sets of core feature information from different analytical perspectives: one that directly reflects the functional heterogeneity of cell populations, and the other that quantifies the underlying communication networks between cells.

[0091] Sp4.1. Statistical or spatial distribution characteristics derived from inferred cellular functional substates;

[0092] From the overall picture of each cell's refined functional substates inferred by Sp2, quantitative features are extracted that can macroscopically describe the entire biological microenvironment image or the functional composition of cell populations within a specific region of interest. These features are designed to capture the relative abundance of cells in different functional substates, their dominant substates, substate diversity, and their spatial organization and distribution patterns, thereby depicting a holistic "portrait" of the biological microenvironment from the perspective of cellular function.

[0093] Statistical features: Global proportion / abundance of each functional substate: Calculate the percentage of cells with each identified functional substate in the image as a percentage of the total number of cells, or the cell density per unit area.

[0094] Functional substate diversity index: Shannon entropy or Simpson index are used as ecological diversity indicators to assess the richness and uniformity of cell functional substates in the image. High diversity indicates a more complex or unstable microenvironment.

[0095] Specific functional substate ratio: Calculate the ratio between some functional substates with specific biological significance, such as the ratio of pro-inflammatory substate cells to anti-inflammatory substate cells, or the ratio of effector immune cells to suppressive immune cells.

[0096] Spatial distribution characteristics: The spatial aggregation / discreteness of functional substates is used to assess whether the cells tend to be spatially clustered, randomly distributed, or evenly dispersed.

[0097] Spatial co-occurrence / mutual exclusion between functional substates: Analyze whether cells of different functional substates tend to co-occur or exclude each other in space.

[0098] Regional enrichment of functional substates: If the image is divided into multiple biologically relevant regions (tumor core, invasive margin, and adjacent normal tissue), the enrichment degree or specific distribution density of each functional substate in these different regions is calculated.

[0099] Community structure of cell functional substates: A community detection algorithm is run on a cell map constructed based on the spatial location and functional substates of cells to identify spatial cell communities composed of cells of a specific functional substate or a mixture of multiple functional substates in a specific pattern, and to extract characteristics such as the size, shape, and composition of these communities.

[0100] The data output of Sp4.1 is a numerical feature vector or a set of feature values ​​that represent the macroscopic characteristics of the current biological microenvironment image at the level of cellular functional substate composition and spatial organization. These features will form part of the set of descriptors for the biological microenvironment state.

[0101] Sp4.2. Quantitative Features from Biomolecular Interaction Potential Field Maps: Extract features that quantitatively describe the intensity, spatial extent, and patterns of potential cellular communication activities within a biological microenvironment from multiple spatially resolved biomolecular interaction potential field maps generated by Sp3. This compresses the complex, pixel-level potential field map information into a set of highly informative, model-friendly numerical descriptors, thereby quantifying the microenvironment's "communication landscape."

[0102] Specific data processing and feature extraction include:

[0103] Global intensity characteristics: For the potential field map of each ligand-receptor pair, its global statistics are calculated; average potential field intensity: the average value of the pixel values ​​of the entire potential field map, reflecting the overall basic level of the interaction.

[0104] Maximum / minimum potential field strength: reflects the peak and valley values ​​of the interaction potential.

[0105] Standard deviation / coefficient of variation of potential field strength: describes the spatial heterogeneity or fluctuation of interaction strength.

[0106] The total integral value of the potential field strength reflects the total potential of the interaction in the entire area.

[0107] High / low interaction area ratio: Set a threshold and calculate the ratio of the pixel area with potential field strength higher than a certain high threshold (representing strong interaction area) or lower than a certain low threshold (representing weak interaction area) to the total analysis area.

[0108] Spatial patterns and texture features:

[0109] Texture features of potential field maps: The potential field map is regarded as a grayscale image, and texture features derived from its gray-level co-occurrence matrix, such as contrast, correlation, energy, entropy, etc., are calculated to describe the complexity, uniformity, and directionality of the spatial distribution pattern of the interaction intensity.

[0110] Morphological features of the potential field map: The potential field map is threshold segmented to obtain a two-dimensional image of the high interaction area. The morphological features of these areas are then calculated, including the number of areas, average area, circularity, elongation, etc., to describe the spatial morphology of the interaction "hotspots".

[0111] Spatial autocorrelation of the potential field map: Calculate the Moran's I index or Geary's C coefficient to evaluate the degree of spatial aggregation of the interaction strength.

[0112] Characteristics of relationships between multiple potential field diagrams:

[0113] Pixel-level correlation between different interaction potential field maps: Pearson correlation coefficients or other correlation metrics are calculated between the potential field maps of different LR pairs to explore the spatial synergy or mutual exclusion relationship between different communication pathways.

[0114] Principal component analysis (PCA) dimensionality reduction: If there are a large number of potential field maps, flatten them and perform PCA to extract the main interaction patterns as features.

[0115] The data output of Sp4.2 is also a numerical feature vector or a set of eigenvalues ​​that characterizes the key characteristics of the current biological microenvironment image at the level of potential biomolecular communication between cells. These features will also be incorporated into the set of biological microenvironment state descriptors.

[0116] Sp4.3. Deviation signatures are obtained by comparing the biological microenvironment image or its derivative representation with a pre-set reference microenvironment state, and a pre-trained generative machine learning model performs the comparison. This introduces an analytical paradigm based on comparison with a "reference standard." Using a pre-trained generative machine learning model, the current biological microenvironment image to be analyzed is quantitatively compared with one or more pre-defined, biologically meaningful reference microenvironment states (such as "healthy / normal," "typical disease," and "drug-sensitive" states). The result of this comparison is not a simple similarity score, but rather a "deviation signature" vector that specifically describes how, in what manner, and to what extent the current microenvironment deviates from the reference state. This deviation signature provides a unique perspective for understanding abnormal microenvironmental states.

[0117] The output after Sp4 directly constitutes the deviation feature here. , using the generator trained in Sp4.2 that can map the input image to a reference state space , and obtain its projection in the reference space , or get its potential space representation zquery through the VAE encoder and reconstruct . Deviation eigenvector The calculation includes:

[0118] Image space deviation: directly calculated from the original input image Rather than being "normalized" or "projected" by the model to a reference state, The difference between the two images is calculated at the pixel level, and the difference image between the two images (or their specific channels) is calculated. , and then extract the statistical characteristics of the difference image, such as the mean, standard deviation, norm (L1 or L2 norm), or calculate the difference value of more advanced image similarity metrics such as the structural similarity index.

[0119] Feature space bias: If the generative model is operating on some derived feature representation of the image (a map of cell density or ECM structural parameters), the bias is calculated in this feature space.

[0120] Latent Space Bias: If you are using a generative model with an explicit encoder (such as a VAE or some types of GANs), computing the original input image Encoded vector in the model latent space , and the distance or difference between it and a set of typical latent vectors representing the reference state (such as the mean or cluster center of all sample encoding vectors in the reference dataset). Alternatively, calculate Encoding and its reconstructed image coding , if the model is trained to align to the reference state.

[0121] Discriminator output (specifically GAN): If a generative adversarial network is used, the trained discriminator It is a classifier that can distinguish whether the input data conforms to the reference state distribution. Input to the discriminator The output probability value or logit value itself is used as a measure A direct indicator of the degree of deviation from the reference state. These deviation measures from different sources together constitute the deviation feature vector , which quantifies the “degree of deviation” and “direction of deviation” of the current biological microenvironment state relative to the preset reference state from multiple perspectives.

[0122] The data output of this Sp4.3 is a numerical deviation feature vector, which will be added to the biological microenvironment state descriptor set.

[0123] The pre-trained generative machine learning models in Sp4 are generative adversarial networks or variational autoencoders. Bias features include a measure of the difference between the image being analyzed and the corresponding reference state projection generated by the generative model, or the score of the image being analyzed under the generative model's discriminator. Sp4.3 specifies the specific types of generative machine learning models used to compare and extract bias features, and details the sources and calculation methods of bias features.

[0124] Applications of Generative Adversarial Networks:

[0125] Model construction and training: Using the CycleGAN architecture, we train two generators, one that converts the input domain to the reference domain. , a function that converts the reference domain back to the input domain , and two corresponding discriminators. The loss function includes adversarial loss and cycle consistency loss. The dataset of the reference microenvironment state is used to define the target domain B.

[0126] Bias features extracted from GAN:

[0127] Conversion difference: convert the image to be analyzed Through the trained generator Convert to reference domain image Then calculate Its "ideal reference version" is defined by, or with Reverse conversion More directly, if there is a mapping from the reference domain to some normalized feature space, then the images to be analyzed are compared and reference domain images The projection into this feature space.

[0128] Discriminator scoring: The image to be analyzed (or its converted version , depends on the design of the discriminator) is input to the trained reference domain discriminator The scalar value it outputs (usually the probability or logit of being a “true reference sample”) can be directly used as a deviation feature. The lower the value, the further away from the reference state.

[0129] Applications of Variational Autoencoders (VAE):

[0130] Model construction and training: Train a VAE whose encoder maps the input microenvironment image (or its feature representation) to a latent space z, and the decoder reconstructs the image from z. If the goal is to compare with a reference state, the VAE can be trained on a dataset of the reference state to learn the latent manifold of the reference state.

[0131] Deviation features are extracted from VAE:

[0132] Reconstruction error: The image to be analyzed Input into the VAE trained on the reference dataset to obtain the reconstructed image .calculate and The reconstruction error between (such as mean square error or structural dissimilarity) and the larger the reconstruction error, the more likely it is that The pattern of the reference state is quite different from that of the control state.

[0133] Latent space position / distance: Mapping to latent vectors through encoders . It can be calculated The Mahalanobis distance or Euclidean distance from the center of the distribution formed by the reference data in the latent space, or its negative log-likelihood under the distribution, is used as a measure of the degree of deviation.

[0134] KL divergence deviation: If the prior distribution of the VAE is a standard normal distribution, and the encoder predicts a posterior distribution for each input , then the KL divergence between the posterior distribution and the prior distribution itself can also reflect the degree of “abnormality” of the input data.

[0135] These specifically defined bias features, whether it is the translation difference / discriminator score based on GAN or the reconstruction error / latent space distance based on VAE, provide unique information about the “atypicality” of the sample to the descriptor set.

[0136] Sp5. Input the biological microenvironment state descriptor set into the pre-trained second machine learning classification model to generate a classification result for the biological microenvironment image.

[0137] The second machine learning classification model makes the final decision on the comprehensive "biological microenvironment state descriptor set" formed by integrating Sp4.1 through Sp4.3, outputting a classification result for the original biological microenvironment image. This classification result is a predefined disease subtype, treatment response group, risk level, or other clinically or biologically meaningful category label. This is the endpoint for achieving the goal of automated analysis and classification for the entire method.

[0138] Data Collection (for training the second machine learning classification model): A dataset containing a large number of biological microenvironment samples is required. For each sample, its image is first processed using the aforementioned steps (Sp1 to Sp4.3) of the present method to extract a complete set of biological microenvironment state descriptors. Furthermore, each sample must have an accurate true class label determined by a gold standard method (e.g., clinical diagnosis, genomic typing, or prognosis or treatment response determined by long-term follow-up).

[0139] Data processing:

[0140] Input: A set of biological microenvironment state descriptors for each sample (a numerical feature vector).

[0141] Feature scaling / normalization: Uniformly scaling or normalizing all features in a descriptor set to have similar numerical ranges or statistical distributions is crucial for stable training of many machine learning models.

[0142] Train / validation / test split: Split the collected labeled dataset into independent training, validation, and test sets for model training, hyperparameter tuning, and final performance evaluation.

[0143] Dealing with class imbalance: If the number of samples in different classes varies greatly, it is necessary to adopt strategies such as oversampling the minority class, undersampling the majority class, or introducing class weights in the loss function.

[0144] The selected classification model is trained using the training set, and the model parameters are optimized by minimizing an appropriate loss function. A validation set is used to monitor the training process, perform early stopping to prevent overfitting, and perform hyperparameter tuning. The performance of the trained classification model is evaluated on an independent test set. Common evaluation metrics include accuracy, precision, recall, F1 score, area under the receiver operating characteristic curve, and area under the precision-recall curve.

[0145] For each input biological microenvironment image (after the above processing to obtain a set of descriptors), this outputs its final classification result, that is, a category label. It also outputs the model's confidence or probability score for the classification result.

[0146] The second machine learning classification model in Sp5 is a heterogeneous multi-branch deep neural network, which includes: at least one graph neural network branch, which is specifically used to process cell-level graph structure data constructed by cell functional substates and their spatial adjacency relationships to learn cell community patterns and local cell niche characteristics; at least one convolutional neural network branch, which is specifically used to process spatially resolved biomolecular interaction potential field maps and spatial deviation feature maps generated by generative machine learning models to extract their multi-scale spatial texture and intensity patterns; the deep feature vectors extracted by different branches are integrated through a learnable cross-modal attention fusion module, which can adaptively assign confidence weights to various features of different biological microenvironment samples or different spatial regions, and finally a classification head outputs the classification results.

[0147] Input data preparation and diversion:

[0148] Cell-level graph structure data: Based on the cellular functional substates inferred by Sp2 and their spatial coordinates obtained in Sp1, a cellular spatial relationship graph is constructed for each sample. The nodes of the graph are cells, and node attributes include their functional substate labels and some important single-cell features extracted by Sp1. Edges are defined based on spatial proximity relationships between cells or more complex rules. This graph data is input into the graph neural network branch.

[0149] Spatially resolved field map data: The biomolecular interaction potential field map generated by Sp3 and the spatial deviation feature map generated by the GAN in Sp4.3. These two-dimensional or three-dimensional image / field map data are input to the convolutional neural network branch.

[0150] Other vector features: global statistical features of Sp4.1, global quantitative features of potential field maps of Sp4.2, non-image form deviation features of Sp4.3, etc., are flattened and input into a standard multi-layer perceptron branch, or directly spliced ​​with the features extracted by the GNN / CNN branch before the fusion stage.

[0151] Graph Neural Network (GAT) branches: GAT aggregates information by assigning different weights to neighboring nodes through an attention mechanism. This makes it more suitable for processing graphs with high heterogeneity in node features and neighborhood structure. Multiple layers of GNNs are stacked to learn higher-order neighborhood information about a node. Node features are input and passed through several GNN layers (each layer includes feature transformation, neighborhood aggregation, and activation functions). Ultimately, an embedding vector is learned for each node, or an embedding representation of the entire graph is obtained through graph pooling operations (such as global average pooling and attention pooling). This method learns the spatial organization of cell populations at the functional substate level, such as the aggregation of cells in specific substates and how cells of different substates intersect to form specific local cellular niches.

[0152] Convolutional Neural Network (CNN) branches: Using a CNN architecture, multiple potential field maps or deviation feature maps are processed using a multi-input CNN, or using a CNN independently for each map to extract features before fusion. Through multiple layers of convolution, pooling, activation functions, and batch normalization, multi-scale spatial texture, intensity patterns, shape, and other features are extracted from the input 2D or 3D field map. Finally, a feature vector is obtained through global average pooling or flattening. The spatial intensity distribution of different communication pathways, the shape and range of hotspots, and the specific spatial patterns of microenvironmental deviations from the reference state are learned from the interaction potential field map. The specific spatial patterns of microenvironmental deviations from the reference state are learned from the deviation feature map.

[0153] Cross-modal attention fusion module: After the GNN and CNN branches each extract high-level feature vectors, these features from different modalities and dimensions need to be fused. The feature vectors from each branch are concatenated and then transformed and reduced in dimension through one or more fully connected layers. A self-attention mechanism is introduced to learn the importance of different feature dimensions.

[0154] Joint attention: A mechanism is designed to allow features from one modality to "focus" on relevant parts of features from another modality, and vice versa, thereby learning cross-modal correlations. Transformer-based fusion: The feature vectors of each branch are treated as part of the input sequence of the Transformer encoder or decoder, and deep fusion is performed using its self-attention and cross-attention mechanisms. The core of the attention module is to calculate attention weights, which determine the contribution of each original feature or feature source in the final integrated feature representation. These weights are learnable and can be dynamically adjusted based on the specific circumstances of the input sample, allowing the model to adaptively allocate higher attention to those features or spatial regions that are most discriminative for the classification of the current specific sample, thereby adapting to the high heterogeneity of the biological microenvironment. The fused feature vector is ultimately input into a classification head consisting of one or more fully connected layers, which outputs the probability of belonging to each predefined category through the Softmax activation function.

[0155] The classification results classify biological microenvironment images into one of multiple predefined "functional microenvironment prototypes", each of which is defined and distinguished by the following set of integrated parameters: characteristic abundance spectra and spatial co-occurrence patterns of a set of key cellular functional substates; characteristic intensity distribution and topological structure of a set of dominant biomolecular interaction potential fields; and multidimensional deviation vectors quantified by deviation characteristics relative to a specific reference microenvironment state; the prototype is associated with the molecular mechanism response characteristics at the level of specific disease biological behavior or treatment mediated by the microenvironment.

[0156] Each biological microenvironment image to be analyzed is mapped onto a predefined and carefully characterized "functional microenvironment prototype." These prototypes represent biologically distinguishable microenvironment states with specific functional tendencies and organizational patterns. More importantly, the definition of each functional microenvironment prototype is highly structured and multidimensional, and it must be accurately described and distinguished from each other through a set of integrated, quantitative parameters. This set of integrated parameters includes at least the following three levels:

[0157] At the cellular functional substate composition level, this is defined by the relative abundance profile of a set of key cellular functional substates (derived from Sp2) that characterize the archetype. This abundance profile indicates which functional substates are enriched and which are sparse within the archetype, as well as their relative proportions. Characteristic spatial co-occurrence patterns or spatial arrangement patterns of these key functional substates (e.g., whether a certain immune cell substate tends to cluster around a tumor cell substate, or whether two different stromal cell substates form a specific spatially interlaced structure) are also important criteria for defining the archetype. These spatial co-occurrence patterns can be quantified using spatial distribution features extracted from Sp4.1, such as the proximity index and colocalization coefficient of specific substate pairs.

[0158] The biomolecular interaction network level is defined by the characteristic intensity distribution patterns and network topology of the biomolecular interaction potential fields (derived from Sp3) that dominate the prototype. This includes indicating which specific ligand-receptor-mediated communication pathways are highly active in the prototype (corresponding to high intensity and wide range of the potential field graph) and which are inhibited. The topology of the potential field graph, as well as the shape, size, number, and connectivity of interaction "hotspots" (described by quantitative potential field features extracted from Sp4.2), also constitute the characteristics of this dimension of the prototype.

[0159] Relative deviation from the reference state: A multi-dimensional deviation vector is used to define the prototype's "position" or "difference" relative to one or more pre-set reference microenvironmental states (healthy physiological state, early stage of disease, or extreme state completely insensitive to treatment). This deviation vector is composed of deviation features obtained through comparison in Sp4.3 through the generative machine learning model. It quantifies the degree and direction of the difference between the prototype and the reference state in multiple dimensions.

[0160] Through the aforementioned steps of the present method, a complete set of state descriptors is extracted. Unsupervised or semi-supervised clustering algorithms (such as Gaussian mixture models or more advanced clustering methods capable of integrating multimodal information) are then applied within this high-dimensional descriptor space to identify naturally occurring, separable "prototype" clusters within the data. Further, in-depth analysis of the characteristic centers or representative samples of each cluster is performed to summarize and refine the integrated definition parameters at the three levels described above. Furthermore, it is necessary to establish a correspondence between these data-driven prototypes and known clinical pathological phenotypes or molecular biological features to assign them clear biological and clinical significance.

[0161] The output of this method further includes generating and presenting at least one of the following computationally enhanced and contextually associated visual analysis maps: specific visual analysis maps are of the following types:

[0162] The spatial distribution map of cell functional substates, overlaid with a significance heatmap, displays the identified driver cell functional substates or their spatial clusters that contribute most to the classification results. This visualization first assigns a color or symbol to each cell inferred in the Sp2 step based on its functional substate, and then renders it on the original high-resolution microenvironment image (or a simplified background version). This clearly demonstrates the precise location and distribution pattern of cells of different functional substates within the tissue space. Furthermore, it leverages the analysis results obtained by Sp5 and computational attribution methods. The driver cell functional substates identified as contributing most to the final microenvironment classification, or the spatial clusters of these driver substates, are superimposed on the cell functional substate distribution map as a significance heatmap. The color depth of the heatmap corresponds to the contribution or driving strength of each substate or region. This overlaid visualization allows users to clearly identify which specific functional cell types, at which specific spatial locations, are key factors in determining the current microenvironmental state.

[0163] Comparative visualizations display the current sample's biomolecular interaction potential map side by side or in a differentiated manner with the average potential map of a defined "functional microenvironment prototype" or a reference sample, highlighting statistically significant differences in interaction patterns. This visualization focuses on characterizing biomolecular interaction networks. It first extracts several important biomolecular interaction potential maps for the current sample, calculated in step Sp3. Then, based on the "functional microenvironment prototype" assigned to the current sample in step Sp5, the corresponding average potential map is retrieved from a pre-built prototype database, or the potential map of one or more representative reference samples is selected. Next, the potential map of the current sample is displayed side by side with the average potential map of the prototype or the reference sample for direct visual comparison. More advanced methods can be used to calculate a difference potential map between the two, highlighting regions of statistically significant differences in interaction strength or spatial pattern within the difference map using color coding or thresholding. This comparative visualization helps users understand how the communication network characteristics of the current sample conform to or deviate from the general pattern of the prototype to which it belongs, or what specific interaction changes exist compared to the reference state.

[0164] Based on the deviation features generated by the generative adversarial network, spatial back-projection is performed on the cellular or regional scale of the original image to generate a "deviation map", which indicates the structural or functional units that differ most significantly from the reference state and can annotate the specific dimensions of their deviation. The "deviation features" extracted in step Sp4.3, which are usually global or high-dimensional vectors, are remapped back to the spatial dimensions of the original image to generate an intuitive "deviation map". If the deviation feature itself is a pixel-level or regional-level difference map, it can be directly superimposed on the original image as a heat map. If the deviation feature is a global vector, it may be necessary to use the model's interpretability technology to identify which regions or cells in the original image contribute most to the deviation feature, and then highlight these high-contribution regions on the deviation map. The color or intensity of different regions on the map can indicate the degree of deviation from the reference state. Furthermore, if the deviation features are multidimensional, and each dimension corresponds to a specific deviation direction or pattern (deviation dimension 1 represents "excessive fibrosis" and deviation dimension 2 represents "inadequate immune cell infiltration"), these high-deviation areas can be interactively queried or automatically annotated on the deviation map, indicating in which specific dimensions they significantly deviate from the reference state. This visualization allows users to pinpoint the "most abnormal" or "unique" structural or functional units in the microenvironment and understand the specific nature of their abnormality.

[0165] The technical key points to realize these advanced visualization maps include: the need to develop efficient image rendering and overlay algorithms; ensuring that data from different sources (original images, segmentation masks, functional substate labels, potential field maps, attribution heat maps, etc.) can be accurately registered and fused in space; and providing a user-friendly interactive interface that allows users to perform operations such as zooming, panning, layer selection, threshold adjustment, and information query.

[0166] Biological microenvironment images are high-content imaging data with single-cell or subcellular resolution and contain at least 15 distinguishable molecular detection channels, including multiple immunofluorescence imaging, imaging mass spectrometry flow cytometry imaging, and high-dimensional spatial omics imaging data; the multi-parameter information provided by a large number of molecular channels is the basis for the first machine learning model to reliably infer refined functional substates and execute them in Sp2, as well as to support the precise calculation of biomolecular interaction potential fields and execute them in Sp3.

[0167] The inference of refined functional substates in step Sp2 relies on deep learning and pattern recognition of the complex combinatorial expression patterns of dozens of molecular markers for each cell (i.e., its high-dimensional "molecular fingerprint") and the resulting detailed morphological and microenvironmental features. This enables the first machine learning model to transcend traditional crude cell typing based on a few markers, reliably inferring more subtle and dynamic functional substates that are more closely linked to the cell's true biological behavior or potential. The greater the number of channels, the richer the feature dimensions that can be used to distinguish different substates, and the higher the accuracy and refinement of the inference results. Secondly, in step Sp3, the precise calculation of the biomolecular interaction potential field requires accurate quantification of the expression levels of specific ligands and receptors in individual cells (especially those classified into specific functional substates) to calculate their potential to act as signal sources or sinks (i.e., EL,i and ER,j in the aforementioned formula). Having a sufficient number of molecular channels allows us to simultaneously detect multiple ligands and their corresponding receptors that constitute different communication pathways, thereby constructing a more comprehensive, multi-pathway interaction potential field map and gaining a deeper understanding of the complex cellular communication networks in the biological microenvironment. If the number of molecular channels is too small, it will be impossible to fully capture the heterogeneity of cellular states and fully characterize the diverse intercellular interactions, which will severely limit the analytical depth and biological insights that can be achieved by the present method.

[0168] The operational steps of the entire technical solution are summarized as follows:

[0169] First start the system, then run:

[0170] Sp1: Image processing and initial feature extraction;

[0171] Sp2: Refined functional substate inference;

[0172] Sp3: calculation of interaction potential field map;

[0173] Sp4: construct state descriptors and classify;

[0174] Sp5: Computational attribution analysis;

[0175] After the process is completed, the final output result is obtained, and then the system is terminated. It should be noted that Sp5: Calculation attribution analysis can be adjusted according to the cost of the application environment and the configuration plan of the hardware equipment. This step can be skipped to directly output the result and terminate the system.

[0176] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a reference structure" does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0177] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A biological microenvironment image analysis and classification method, characterized by: The following steps are involved: Sp1. Process the biological microenvironment image, segment the cell components in the image, and extract the multi-parameter initial features of each cell component; Sp2. Based on the multi-parameter initial features and the local microenvironment information of the cells, a first machine learning model is used to infer the refined functional substates of each cellular component, wherein the functional substates are characterized as states that go beyond basic cell types and are related to the biological behavior or potential of the cells; Sp3. Based on the functional substates, spatial position information, and predefined biomolecular interaction knowledge of the cellular components, calculate and generate a spatially resolved biomolecular interaction potential field map covering the image analysis area, wherein the biomolecular interaction potential field map quantifies the potential or strength of specific biomolecule-mediated interactions occurring at each point in space; Sp4. Extract and integrate at least two of the following sets of features to form a set of biological microenvironment state descriptors: Sp4.

1. Statistical or spatial distribution characteristics derived from inferred cellular functional substates; Sp4.

2. Quantitative features derived from the biomolecular interaction potential field map; Sp4.

3. The deviation features obtained by comparing the biological microenvironment image or its derivative representation with a preset reference microenvironment state are compared using a pre-trained generative machine learning model; Sp5. Input the biological microenvironment state descriptor set into a pre-trained second machine learning classification model to generate a classification result for the biological microenvironment image.

2. A biological microenvironment image analysis and classification method according to claim 1, characterized in that: The first machine learning model for inferring functional substates in Sp2 is an unsupervised learning model or a self-supervised learning model, which identifies functional substates by learning the latent space representation of multi-parameter features of cells.

3. The biological microenvironment image analysis and classification method according to claim 1, characterized in that: The biomolecular interaction knowledge in the Sp3 includes ligand-receptor pair information, and the calculation of the interaction potential field map includes the spatial proximity, expression level and functional substate of ligand-expressing cells and receptor-expressing cells.

4. The biological microenvironment image analysis and classification method according to claim 1, characterized in that: The pre-trained generative machine learning model in Sp4 is a generative adversarial network or a variational autoencoder, and the deviation feature includes a difference measure between the image to be analyzed and the corresponding reference state projection generated by the generative model, or a score of the image to be analyzed under the discriminator of the generative model.

5. The biological microenvironment image analysis and classification method according to claim 1, characterized in that: The method further comprises Sp6: Based on the second machine learning classification model and the biological microenvironment state descriptor set, a computational attribution method is used to identify key features or feature combinations that have a significant impact on the classification results, and this identification result can be selectively used to guide the training of the second machine learning classification model or explain its output.

6. The biological microenvironment image analysis and classification method according to claim 5, characterized in that: The computational attribution method is configured to: identify the global importance of the biological microenvironment state descriptors that contribute significantly to the classification result; The contribution of the biomolecular interaction potential field map in a specific spatial region is traced back to the precise spatial configuration and density threshold of one or more specific cellular functional substates that cause changes in the potential field intensity and pattern in that region, thereby revealing the spatially specific cellular-level functional synergistic or antagonistic pathways that drive the microenvironment classification state.

7. The biological microenvironment image analysis and classification method according to claim 1, characterized in that: The second machine learning classification model in Sp4 is a heterogeneous multi-branch deep neural network, which includes: At least one graph neural network branch is specifically used to process cell-level graph structure data constructed by the cell functional substates and their spatial adjacency relationships to learn cell community patterns and local cell niche characteristics; at least one convolutional neural network branch specifically configured to process the spatially resolved biomolecular interaction potential field map and the spatial deviation feature map generated by the generative machine learning model to extract multi-scale spatial texture and intensity patterns thereof; The deep feature vectors extracted by different branches are integrated through a learnable cross-modal attention fusion module, which can adaptively assign confidence weights to various features of different biological microenvironment samples or different spatial regions, and finally a classification head outputs the classification results.

8. The biological microenvironment image analysis and classification method according to claim 1, characterized in that: The classification results categorize the biological microenvironment images into one of a number of predefined "functional microenvironment prototypes," each of which is defined and distinguished by a set of integrated parameters: Characteristic abundance profiles and spatial co-occurrence patterns of key cellular functional substates; The characteristic intensity distribution and topology of the dominant biomolecular interaction potential field; A multidimensional vector of deviations from a specific reference microenvironmental state, quantified by a deviation signature; prototypes are associated with the molecular mechanism response signature at the microenvironment-mediated level of a specific disease biological behavior or treatment.

9. The biological microenvironment image analysis and classification method according to claim 4, characterized in that: The output of Sp4 further includes generating and presenting at least one of the following computationally enhanced and contextually associated visual analysis maps: The spatial distribution map of cell function substates shows the driving cell function substates or their spatial clustering areas that contribute most to the classification results by overlaying a significant heat map; Comparative visualization maps, which display the biomolecular interaction potential field map of the current sample side by side or differentially with the average potential field map of the functional microenvironment prototype or the potential field map of the reference sample, highlighting the statistically significant differences in interaction patterns; Based on the deviation features generated by the generative adversarial network, spatial back-projection is performed on the cellular or regional scale of the original image to generate a deviation map, indicating the structural or functional units that differ most significantly from the reference state and annotating the specific dimensions of their deviation.

10. The biological microenvironment image analysis and classification method according to claim 1, characterized in that: The biological microenvironment image is high-content imaging data with single-cell or subcellular resolution and contains at least 15 distinguishable molecular detection channels, including multiple immunofluorescence imaging, imaging mass spectrometry flow cytometry imaging and high-dimensional spatial omics imaging data; the multi-parameter information provided by a large number of molecular channels is the basis for the first machine learning model to reliably infer refined functional substates and execute them in Sp2, and to support the precise calculation of biomolecular interaction potential fields and execute them in Sp3.

Citation Information

Patent Citations

  • Drug activity screening method and device

    CN118506911A

  • Cell mechanical phenotype analysis method and equipment based on deep learning

    CN119993284A