Biological microenvironment image analysis and classification method
Through the processing of biological microenvironment images and the application of machine learning models, the problem of in-depth disclosure of cell functional status and interactions in the prior art is solved, and quantitative and visualized biological microenvironment analysis is realized, which improves the accuracy and reliability of the analysis and supports accurate diagnosis and treatment.
Patent Information
- Application Number
- CN202510737365.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing technology cannot deeply reveal the true functional status of cells in the biological microenvironment and the interaction between cells, and the lack of quantitative and intuitive evaluation methods, resulting in poor repeatability of analysis results and difficulty in standardizing application.
By processing the biological microenvironment images, segmenting cell components, extracting multi-parameter initial features, using machine learning models to infer cell functional substates, generating biomolecular interaction potential field maps, combining generative machine learning models for classification, and integrating multi-dimensional features to generate accurate biological microenvironment classification results.
Quantitative analysis of the refined functional status of cells in a specific environment is realized, the intercellular communication network is intuitively displayed, objective deviation characteristics are generated, the depth and accuracy of the analysis are improved, and accurate diagnosis and personalized treatment are supported.
Smart Images

Figure CN120259303A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cell image analysis, and specifically to a method for analyzing and classifying biological microenvironment images. Background Art
[0002] The biological microenvironment is a complex ecosystem composed of various cells and extracellular matrix components, and its state plays a decisive role in the occurrence, development, and treatment response of major diseases such as cancer. At present, the analysis of the biological microenvironment mainly relies on histopathological staining and multiplex imaging techniques, combined with computer vision for quantitative analysis.
[0003] However, the existing technologies have the following main limitations:
[0004] The existing methods mainly perform cell recognition, counting, and morphological measurement, only describing the static composition of the microenvironment, and unable to deeply reveal the true "functional state" of cells in a specific environment; for the assessment of cell-cell interactions, most stay at the level of analyzing whether the cell spatial positions are adjacent, which is too simplistic and unable to quantitatively and intuitively display the complex communication network mediated by specific biomolecules such as ligand-receptor pairs with specific intensities and ranges of action. For the judgment of the overall state of the microenvironment, such as the "cold" and "hot" qualitative of the tumor microenvironment, it largely depends on the subjective experience of the observer, lacking a unified and quantitative evaluation standard, resulting in poor repeatability of the analysis results and being difficult to be standardized and applied in different studies or clinical practices. Summary of the Invention
[0005] Technical Problems to be Solved
[0006] Aiming at the deficiencies of the existing technologies, the present invention provides a method for analyzing and classifying biological microenvironment images, solving the problems of the existing technologies.
[0007] Technical Solutions
[0008] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for analyzing and classifying biological microenvironment images, comprising the following steps:
[0009] Sp1. Process the biological microenvironment image, segment the cell components in the image, and extract the multi-parameter initial features of each cell component;
[0010] Sp2. Based on the multi-parameter initial features and the local microenvironment information of the cells, use a first machine learning model to infer the refined functional sub-states of each cell component, where the functional sub-states are characterized as states related to cell biological behavior or potential beyond the basic cell types;
[0011] Sp3. Based on the functional sub-states of cell components, spatial location information, and predefined knowledge of biomolecular interactions, calculate and generate a spatially resolved biomolecular interaction potential field map covering the image analysis region, where the biomolecular interaction potential field map quantifies the potential or intensity of interactions mediated by specific biomolecules at each point in space;
[0012] Sp4. Extract and integrate at least two of the following sets of features to form a set of biological microenvironment state descriptors:
[0013] Sp4.1. Statistical or spatial distribution features derived from the inferred functional sub-states of cells;
[0014] Sp4.2. Quantitative features derived from the biomolecular interaction potential field map;
[0015] Sp4.3. Deviation features obtained by comparing the biological microenvironment image or its derived representation with a preset reference microenvironment state, and a pre-trained generative machine learning model performs the comparison;
[0016] Sp4.4. Input the set of biological microenvironment state descriptors into a pre-trained second machine learning classification model to generate a classification result for the biological microenvironment image.
[0017] Preferably, the first machine learning model for inferring functional sub-states in Sp2 is an unsupervised learning model or a self-supervised learning model, which identifies functional sub-states by learning the latent space representation of cell multi-parameter features.
[0018] Preferably, the biomolecular interaction knowledge in Sp3 includes ligand-receptor pair information, and the calculation of the interaction potential field map includes the spatial proximity, expression level, and functional sub-states of ligand-expressing cells and receptor-expressing cells.
[0019] Preferably, the pre-trained generative machine learning model in Sp4 is a generative adversarial network or a variational autoencoder, and the deviation features include the measure of the difference between the image to be analyzed and the projection of the corresponding reference state generated by the generative model, or the score of the image to be analyzed under the discriminator of the generative model.
[0020] Preferably, the method further includes Sp5:
[0021] Based on the second machine learning classification model and the set of biological microenvironment state descriptors, use a computational attribution method to identify key features or feature combinations that have a significant impact on the classification result, and optionally use this identification result to guide the training of the second machine learning classification model or to interpret its output.
[0022] Preferably, the calculation attribution method is configured to: identify the global importance of the biological microenvironment state descriptors that contribute significantly to the classification result;
[0023] Trace the contribution of the biomolecular interaction potential field map in a specific spatial region to the precise spatial configuration and density threshold of one or more specific cellular functional sub-states that cause changes in the potential field intensity and pattern in this region, thereby revealing the spatially specific cellular-level functional cooperation or antagonistic pathways that drive the classification state of the microenvironment.
[0024] Preferably, the second machine learning classification model in the Sp4 is a heterogeneous multi-branch deep neural network, and the heterogeneous multi-branch deep neural network includes:
[0025] At least one graph neural network branch, which is specifically used to process the cellular-level graph structure data constructed by the cellular functional sub-states and their spatial adjacency relationships to learn the cellular community patterns and local cellular niche characteristics;
[0026] At least one convolutional neural network branch, which is specifically used to process the spatially resolved biomolecular interaction potential field map and the spatial deviation feature map generated by the generative machine learning model to extract its multi-scale spatial texture and intensity pattern;
[0027] The deep feature vectors extracted by different branches are integrated through a learnable cross-modal attention fusion module, which can adaptively assign confidence weights to various features of different biological microenvironment samples or different spatial regions, and finally a classification result is output by a classification head.
[0028] Preferably, the classification result classifies the biological microenvironment image into one of multiple predefined "functional microenvironment prototypes", and each prototype is defined and distinguished by the following set of integration parameters:
[0029] The characteristic abundance spectrum and spatial co-occurrence pattern of key cellular functional sub-states;
[0030] The characteristic intensity distribution and topological structure of the dominant biomolecular interaction potential field;
[0031] A multi-dimensional deviation vector quantified by deviation features relative to a specific reference microenvironment state; the prototype is associated with the molecular mechanism response characteristics of specific disease biological behaviors or treatment levels mediated by the microenvironment.
[0032] Preferably, the output of the Sp4 further includes generating and presenting at least one of the following computationally enhanced and contextually relevant visual analysis maps:
[0033] Cell functional substate spatial distribution map, which shows the driving cell functional substate or its spatial aggregation region that contributes the most to the classification result by superimposing a significance heat map;
[0034] Comparison visualization map, which visually shows the biomolecular interaction potential field map of the current sample juxtaposed or differentiated with the average potential field map of the functional microenvironment prototype or the potential field map of the reference sample, and highlights the interaction pattern differences with statistical significance;
[0035] Based on the deviation features generated by the generative adversarial network, perform spatial backprojection at the cell or regional scale of the original image to generate a deviation map, indicating the structural or functional unit with the most significant difference from the reference state, and annotating the specific dimension of its deviation.
[0036] Preferably, the biomicroenvironment image is high-content imaging data with single-cell or subcellular resolution and containing at least 15 distinguishable molecular detection channels, including multiplex immunofluorescence imaging, imaging mass cytometry imaging, and high-dimensional spatial omics imaging data; the multi-parameter information provided by a high number of molecular channels is the basis for the first machine learning model to reliably infer refined functional substates and execute in Sp2, and support the accurate calculation of biomolecular interaction potential fields and execute in Sp3.
[0037] Beneficial effects
[0038] The present invention provides a method for analyzing and classifying biomicroenvironment images. It has the following beneficial effects:
[0039] 1. Through deep modeling, the present invention can infer refined "functional substates" beyond cell types, accurately reveal the true biological behaviors of cells in a specific environment; quantitatively and visually display the "interaction potential field" mediated by specific molecules, and intuitively show the communication network and intensity between cells; explore and identify the key factors driving the formation of microenvironment states.
[0040] 2. The present invention introduces a generative machine learning model, which generates objective "deviation features" by quantitatively comparing the sample to be analyzed with a "healthy" or "ideal" reference state. This feature accurately quantifies the degree and direction of the abnormality of the sample, completely getting rid of the bondage of subjective judgment. Based on this objective standard and integrating multi-dimensional functional information, the classification result of this method is more accurate and stable, providing a reliable technical basis for achieving accurate diagnosis, prognosis judgment, and guiding personalized treatment. Description of the drawings
[0041] Figure 1 It is the system functional module diagram of the present invention;
[0042] Figure 2 It is the system flow chart of the present invention. Detailed implementation mode
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Specific embodiment one:
[0045] As Figures 1 to 2 shown, the entire system composition of a method for biological microenvironment image analysis and classification includes a data and model library, a data input and preprocessing module, an image segmentation and feature extraction module, a cell functional substate inference module, an interaction potential field calculation module, a deviation feature extraction module, a state descriptor integration and classification module, a calculation attribution analysis module, and a result generation and visualization module.
[0046] A method for biological microenvironment image analysis and classification includes the following steps:
[0047] Sp1. Process the biological microenvironment image, segment the cell components in the image, and extract the multi-parameter initial features of each cell component.
[0048] The core technical feature of this step is to achieve high-precision and automated segmentation of various cell components in complex biological microenvironment images, mainly including cell nuclei, cytoplasm, cell membranes, and in some cases, extracellular matrix regions that need to be analyzed. After the segmentation is completed, a comprehensive and multi-dimensional set of initial feature parameters will be extracted from each independently identified cell entity. These initial feature parameters not only cover the fine morphological information of cells, such as size, shape, edge smoothness, etc., but also include their optical density or fluorescence intensity information under different imaging channels, as well as preliminary texture features that can reflect changes in subcellular structures. This series of operations lays a solid and comprehensive data foundation for the more complex cell refinement functional substate inference in the subsequent Sp2 step.
[0049] There are numerous and crucial technical points in the implementation of this step. First is the standardized process of image preprocessing, which is a prerequisite for ensuring the stability and comparability of subsequent analysis results. It is necessary to effectively address issues such as uneven luminance distribution, random and systematic noise, and various imaging artifacts in images from different imaging devices, different experimental batches, or different sample preparation processes, as well as interference caused by tissue folding or bubbles. Through a series of standardized preprocessing operations, such as background correction, flat-field correction, and denoising, the image quality is significantly improved. Secondly, given the inherently high heterogeneity within the biological microenvironment, where there are significant differences in cell size, morphology, and spatial distribution density in different regions, when selecting or designing a segmentation algorithm, it must be ensured that it can flexibly adapt to and effectively handle this multi-scale variation characteristic. The algorithm should be able to accurately segment both large-volume cells sparsely distributed and small-volume cells densely arranged in the field of view. Moreover, the number of feature parameters extracted from each cell in the initial stage is extremely large, with dimensions reaching hundreds or even thousands, and inevitably includes redundant information or features that are not highly relevant to the specific analysis objective. Therefore, before entering the subsequent modeling steps, it is necessary to carefully consider using preliminary feature importance assessment methods, such as those based on statistical tests or the built-in evaluation mechanisms of machine learning models, or applying effective dimensionality reduction techniques, such as principal component analysis, to refine the feature set, reduce the computational burden, and improve model performance. At the same time, the entire image processing and feature extraction process should have good compatibility and scalability at the system design level, being able to flexibly adapt to and efficiently process input data from different advanced imaging technologies, including but not limited to multiplexed immunofluorescence imaging techniques such as cyclic immunofluorescence (CyCIF) or multiplexed ion beam imaging (MIBI), high-parameter imaging mass cytometry, and high-resolution hematoxylin and eosin staining whole-slide digital scanning imaging that provides rich morphological information without specific labeling.
[0050] If the cell segmentation model in this step, especially the segmentation model based on deep learning, is trained using the supervised learning paradigm, then the collection of high-quality training data is the cornerstone of the model's performance. At this stage, a systematically assembled dataset of biological microenvironment images with broad representativeness is required, which should comprehensively cover all expected cell types, tissue structures, pathological states, and potential technical variations in the research target. The image annotation work is the core link of data collection and must be completed by pathologists with a deep professional background or image analysts who have received rigorous professional training and have the ability to make consistent interpretations. The annotation accuracy requirement reaches the pixel level, and for each nuclear boundary, cytoplasmic boundary, and even the cell membrane structure that can be identified in some high-resolution images in the image, precise contour drawing or area filling is required to form the "gold standard" or "ground truth" data necessary for training the segmentation model. In some cases, to meet the requirements for the complex deep learning model to fully learn and reach a good convergence state, the order of magnitude of the annotated data needs to reach hundreds to thousands of high-resolution original image slices or fields of view, and the total number of cell instances contained should reach tens of thousands or even millions to ensure the generalization ability of the model.
[0051] Inside this Sp1 step, the specific data processing flow is rigorous and meticulous. First of all, the input data are raw, unprocessed digitalized biological microenvironment images. Common storage formats of these images include the general TIFF format, the SVS format for whole-slide digital pathology scanning, or dedicated formats such as ISYNTAX. After receiving the raw images, a detailed preprocessing process is immediately launched. Intensity normalization in the preprocessing operation is an important means to eliminate the brightness differences between different images, and the pixel intensity values of each imaging channel are uniformly scaled to the floating-point range from 0 to 1 or the integer range from 0 to 255 through linear stretching or compression. For more complex intensity drift situations, non-linear normalization methods such as histogram matching based on reference images or global statistics are adopted. As a key link in preprocessing, the choice of noise suppression method depends on the specific noise characteristics of the image. For common random noise, classic smoothing operators such as Gaussian filtering and median filtering can be selected; for more complex noise patterns, methods such as non-local means filtering that can better preserve edge details can be used, or, the current more advanced strategy is to apply deep learning denoising models trained for the noise patterns of specific imaging systems, the Noise2Void network based on self-supervised learning or the CARE network trained based on paired noisy / denoised images. These deep models can effectively remove noise while maximizing the preservation of the true biological structure information in the images. For bright-field microscopic images, especially hematoxylin and eosin stained pathological images, color correction and white balance adjustment need to be performed before quantitative analysis to eliminate color deviations caused by changes in the light source color temperature or staining differences. For multi-channel fluorescence images, a common and tricky problem is the cross-talk effect or fluorescence leakage caused by partial overlap of the emission spectra of different fluorescent dyes, and precise algorithms must be used for processing, the linear unmixing algorithm based on spectral characteristics, or the fluorescence signal compensation algorithm based on the physical model of the dye combination.
[0052] In the image segmentation stage, the accurate segmentation of cell nuclei is the cornerstone and primary task of the entire cell component recognition process. Based on the U-Net architecture, its algorithm details include a symmetric encoder-decoder structure. The encoder part consists of a series of alternating convolutional layers and max-pooling layers, which extract higher-level abstract semantic features through layer-by-layer convolution. At the same time, the pooling layer gradually reduces the spatial resolution of the feature map. The decoder part restores the spatial resolution of the feature map step by step through upsampling operations, transposed convolution, or bilinear interpolation, and fuses it with the feature map of the corresponding level in the encoder part by concatenating or adding through <x_bin_118> connection. This skip connection mechanism enables the decoder to utilize both the shallow high-resolution detail features and the deep high-semantic features from the encoder, thus achieving precise localization of the target boundary. The activation function of neurons in the model generally uses the rectified linear unit ReLU or its improved version LeakyReLU to accelerate model convergence, alleviate the vanishing gradient problem, and introduce non-linearity.
[0053] The segmentation task of cytoplasm and cell membrane is carried out after the nuclei are accurately segmented. If the original multi-channel fluorescence image contains specific fluorescence marker channels for cytoplasm or cell membrane, the deep learning semantic segmentation model similar to that for nucleus segmentation, U-Net, is applied again. At this time, the obtained nucleus segmentation mask image is used as additional input channel information or as a kind of spatial prior knowledge to guide the model to more accurately identify and segment the corresponding cytoplasm region or outline the boundary of the cell membrane. Another strategy is to use a morphological expansion algorithm based on the segmented nuclei in the case of no specific property / membrane marker or weak marker signal. Starting from the center point of each identified nucleus, controlled regional synchronous expansion or watershed expansion is carried out outward, and the stopping boundary of the expansion is limited by the boundary signals of neighboring cells, the estimated average cell size, or in some cases, the weak fluorescence intensity gradient change of the available cell membrane marker.
[0054] After the precise segmentation of all target cell components is completed, it enters the stage of automated extraction of multi-parameter initial features. The morphological features extracted should comprehensively describe the geometric and shape characteristics of cells, specifically including: the area, perimeter of each cell nucleus, cytoplasm, and the whole cell; the lengths of the major axis and minor axis describing the degree of cell elongation, and their ratio; the circularity or shape factor reflecting the degree of the cell approaching a circle; the eccentricity describing the elliptical degree of the cell; the equivalent diameter of the cell; the ratio of the convex hull area to the cell's own area, i.e., the solidity, which describes whether the cell boundary is indented inward; and a series of shape descriptors that can describe shape complexity and are invariant to rotation, translation, and scaling. The extraction of optical density or fluorescence intensity features requires precise quantification of various intensity statistics within each imaging channel for each segmented cell region, namely the cell nucleus region, cytoplasm region, and cell membrane region. These statistics should include: the average value of pixel intensities within the region, the median value (more robust to outliers), the standard deviation or coefficient of variation of intensities (reflecting intensity uniformity), the integral sum of all pixel intensities (reflecting the total amount of markers), and specific percentiles, the 25th percentile and 75th percentile, that can reflect the skewness and tailing of the intensity distribution. The extraction of texture features aims to capture the subtle pattern differences of subcellular structures in the image and can be applied to the internal region of the cell nucleus or the cytoplasm region. Commonly used texture features include a series of statistics calculated based on the gray-level co-occurrence matrix, such as energy, contrast, correlation, entropy, and local homogeneity; the spatial position features of each identified cell in the image coordinate system, i.e., the precise x, y coordinates of its centroid (for three-dimensional imaging data, the z coordinate should also be included), which must also be accurately recorded.
[0055] The final data output of this Sp1 step mainly includes two parts. The first part is the mask image of cell segmentation. In such an image, each successfully segmented and identified independent cell is assigned a globally unique identity identifier, i.e., the cell ID. Such a mask image facilitates subsequent single-cell level feature tracking, proximity relationship analysis, and various visualization operations. The second part is to generate a structured and detailed feature data table. This data table is stored in the comma-separated CSV file format or more professional formats such as HDF5, or directly deposited into a relational or non-relational database. Each row of the data table represents an independent cell instance in the image, and each column corresponds to the quantified measurement value of a specific multi-parameter initial feature extracted for this cell in the Sp1 step. In some specific biological microenvironment analysis applications, if the analysis of the structure and composition of the extracellular matrix also occupies an equally important position, then this Sp1 step will additionally output the mask image of the segmented and identified extracellular matrix region and its related quantified features.
[0056] Sp2. Based on multi-parameter initial features and the local microenvironment information of cells, a first machine learning model is used to infer the refined functional sub-states of each cell component, and the functional sub-states are characterized as states related to cell biological behaviors or potentials that go beyond basic cell types;
[0057] The core technical feature of this step lies in using machine learning techniques, especially unsupervised or self-supervised learning methods, to learn from and automatically identify, from the high-dimensional single-cell data output in the Sp1 step, more refined and dynamic cell functional states beyond traditional cell types defined based on a few markers or simple morphology, that is, functional sub-states. These functional sub-states can more profoundly reveal the actual biological behavior potential, activation state, differentiation stage, or stress response of cells under specific biological microenvironments. Another core feature is the integration of the intrinsic features of cells, that is, the comprehensive morphological parameters and multi-molecular expression profiles from Sp1, as well as the immediate external environmental information of the cells, which includes the types, densities, and functional states of their neighboring cells, thus making the inference of functional sub-states more physiologically relevant and context-dependent. In this way, this step aims to achieve a deep characterization of cell heterogeneity and dynamic plasticity in the biological microenvironment.
[0058] The first machine learning model is a self-supervised model based on deep learning and needs to be trained. Its data collection phase mainly relies on the large-scale single-cell feature data output in the Sp1 step. For the self-supervised learning paradigm, in order to cluster the learned cell embedding representations and functionally annotate the formed clusters later, an independent reference cell data set with partial cell type or rough functional annotations (pre-classification results based on combinations of several key surface markers) is also required, or the knowledge base of domain experts is relied on for subsequent biological significance interpretation. For some weakly supervised learning methods, it is explicitly required that a part of the cell data has such rough functional labels to guide the learning direction of the model.
[0059] Within this Sp2 step, the detailed data processing flow is as follows:
[0060] The input data is the structured single-cell feature table output in the Sp1 step, where each row represents a cell and the columns contain all the multi-parameter initial features of the cell; at the same time, the precise spatial coordinate information of each cell in the original image is also required for calculating local microenvironment features.
[0061] For each cell, based on its spatial coordinates and the preliminary cell types and cell features extracted by Sp1 of other cells around it, the calculation of local microenvironment features is performed. These microenvironment features include: the proportional composition of the types of the K nearest neighbor cells of the target cell; within a specific radius range centered on the target cell (the radius can be set as a multiple of the average cell diameter, defaulting to 50 microns or 100 microns), the number, density of different types of cells, or the proportion of cells in a specific functional substate; the Euclidean distance between the target cell and its nearest cell of a certain specific type (including tumor cells, vascular endothelial cells);
[0062] After completing the calculation and combination of cell intrinsic features and local microenvironment features, it is necessary to perform standardization processing on the entire extended feature set. Apply Z-score standardization to make the mean of each feature 0 and the standard deviation 1, or linearly scale all feature values to a similar numerical range, such as between 0 and 1, to eliminate the impact of different feature scale differences on subsequent distance calculations or model training. Before inputting into the machine learning model, according to the requirements of the specific model and the characteristics of the data, feature selection or dimensionality reduction operations are also selectively performed. Use the principal component analysis method to reduce the feature dimensions while retaining the main variance information, and conduct visual exploration of high-dimensional cell data and assist in identifying potential cell cluster structures.
[0063] The first machine learning model for inferring the refined functional substate of cells uses a self-supervised learning model to first learn an effective low-dimensional embedded representation of cell features, and then perform clustering in this embedded space. A deep generative model such as a variational autoencoder is used to learn the latent space representation of cells.
[0064] The model details of the VAE include: an encoder neural network (which is an MLP), which maps the input cell feature vector to the parameters of a normal distribution in the latent space, that is, the mean vector and the log variance vector.
[0065] The sampling step, according to the mean and variance output by the encoder, randomly samples a latent vector from this parameterized normal distribution.
[0066] A decoder neural network (also an MLP), which attempts to reconstruct the original cell feature vector from the sampled latent vector.
[0067] The loss function of the VAE consists of two parts: one part is the reconstruction loss, which is used to measure the difference between the reconstructed features output by the decoder and the original input features, using the mean squared error or binary cross-entropy; the other part is the KL divergence, which serves as a regularization term to penalize the difference between the distribution in the latent space and a preset prior distribution (which is the standard normal distribution), and encourages the latent space to have a good structure and smoothness.
[0068] The acquisition of functional sub-states is similar to contrastive learning. After the VAE training is completed, the cell latent vectors (which are the mean vectors of the distribution) generated by its encoder part are used for subsequent clustering analysis.
[0069] The final data output of this Sp2 step is the updated cell feature table. In this table, a new column of information will be added for each cell, namely the "functional sub-state ID" or specific sub-state label it is assigned to (including "activated macrophage sub-state M1 type", "exhausted CD8+ T cell sub-state", etc., and these labels are manually annotated with biological knowledge after clustering). In addition, for the convenience of biological interpretation and verification, the average feature spectrum of each identified functional sub-state is also output, that is, the average or median value of all cells in this sub-state on each initial feature dimension, which helps to understand the key differential features between different sub-states. At the same time, visual images of the functional sub-states in two-dimensional or three-dimensional dimensionality reduction spaces such as UMAP or t-SNE are also generated to intuitively show the separation and relative relationships of different sub-states.
[0070] The first machine learning model for inferring functional sub-states in Sp2 is an unsupervised learning model or a self-supervised learning model, which identifies functional sub-states by learning the latent space representation of cell multi-parameter features.
[0071] This includes the selection of the dimension of the learned latent space. The dimension of the latent space is a key hyperparameter. If the dimension is too low, cells in different functional sub-states will be mixed up in the embedding space, losing important discriminative information; while if the dimension is too high, the intrinsic low-dimensional manifold structure of the data cannot be effectively learned, and even noise is introduced, resulting in poor subsequent clustering effects. The selection of the optimal latent space dimension needs to be determined through a series of experiments, evaluating the changes in the reconstruction error under different dimensions, the clarity and stability of the downstream clustering results (using internal evaluation metrics such as the silhouette coefficient and DBI index), and the performance of the classification task on the known partial label data, etc.
[0072] The results generated by self-supervised models, especially the clustering results, are sometimes sensitive to parameters such as the initialization parameters of the model, the optimizer selection, and the learning rate. In applications, measures need to be taken to evaluate the stability of their clustering results, either by repeating the training multiple times and comparing the clustering consistency, or by adopting the idea of ensemble clustering. At the same time, it is also crucial to conduct in-depth analysis and verification of the biological interpretability of the finally identified functional sub-states. Moreover, if the cell data analyzed comes from different experimental batches, different samples, or different treatment conditions, the batch effect will mask the true biological differences. In this case, it is necessary to explicitly consider removing or reducing the impact of the batch effect during the feature extraction stage or the model training stage, using data alignment algorithms to correct the latent space, or introducing a regularization term against the batch effect in the design of the self-supervised learning loss function.
[0073] Application details of the self-supervised contrastive learning algorithm:
[0074] For the multi-parameter feature data of cells in tabular form, the selection of data augmentation strategies is crucial. An effective augmentation method should be able to create sufficient diverse "views" for the cells while maintaining their core identity information (i.e., without changing the functional sub-states they should belong to).
[0075] The architecture design of the encoder network needs to balance the model's expressive power and computational efficiency. For a biological microenvironment dataset containing millions of cells, each with hundreds of initial features, a multi-layer perceptron with 2 to 4 hidden layers, each containing 128 to 512 neurons, is sufficient as the encoder network. The rectified linear unit is used as the activation function between hidden layers, and batch normalization layers are used in conjunction to accelerate the training process, improve the model's stability and generalization ability.
[0076] The projection head network is a relatively simple small MLP, consisting of a hidden layer and an output layer, with its output dimension generally set between 64 and 256. It further non-linearly transforms the embedding vectors generated by the encoder into a space specifically for calculating the contrastive loss. The temperature parameter τ is a key parameter in the NT-Xent contrastive loss function, which regulates the discrimination sensitivity of the loss function to negative samples, that is, determines the strength of the model to push apart the embedding representations of different cells, and needs to be carefully adjusted according to the characteristics of the specific dataset, and its empirical value range is generally between 0.05 and 0.5. The model is trained with a relatively large batch size, from 512 to 4096 or even larger, to ensure that each positive sample has a sufficient number and diversity of negative samples for contrast in each iteration. The optimizer is selected from one of the adaptive learning rate algorithms of Adam or AdamW, or the more stable LARS optimizer for ultra-large batch training.
[0077] Model construction includes a pre-training and a fine-tuning phase. Ideally, self-supervised pre-training is performed on as extensive and diverse a dataset of cellular features as possible from multiple different tissue types, different disease states, and even different species. The encoder obtained through such pre-training can learn more general and robust representations of cellular features, capturing the common laws of cellular state changes. For a specific research question or a specific dataset, based on this pre-trained model, the encoder network is further fine-tuned using the target dataset of the current study to better adapt to the characteristics of the target data. After completing the training or fine-tuning of the encoder, it can be used to convert the original high-dimensional feature vectors of all cells into low-dimensional latent space embedding vectors that are more information-dense and semantically richer. On these high-quality embedding vectors, a graph clustering algorithm is applied to discover non-spherical and more complexly shaped cell cluster structures. The determination of the final number of clusters (i.e., the number of functional sub-states) is carried out through a comprehensive evaluation of internal clustering evaluation metrics (such as silhouette coefficient, Calinski-Harabasz index, Davies-Bouldin index), combined with an intuitive observation of the dimensionality reduction visualization results, and the biological prior knowledge of domain experts.
[0078] Sp3. Based on the functional sub-states of cellular components, spatial location information, and predefined knowledge of biomolecular interactions, calculate and generate a spatially resolved biomolecular interaction potential field map covering the image analysis region. The biomolecular interaction potential field map quantifies the potential or intensity of interactions mediated by specific biomolecules at each point in space;
[0079] Construction and dynamic screening of a comprehensive and accurate biomolecular interaction knowledge base. The biomolecular interaction knowledge base contains known and experimentally verified ligand-receptor pairs, cytokine-receptor pairs, adhesion molecule pairs, etc., and preferably can be associated with the cell types in which they are highly specifically expressed or under what cellular functional states. It is necessary to screen out the most relevant several groups of interaction pairs from the huge knowledge base according to the biological background of the current study for subsequent potential field calculations. Secondly, it is necessary to combine the fluorescence intensity features of each cell extracted in the Sp1 step in the corresponding molecular marker channels, or the definition of functional sub-states itself contains a description of the expression status of these key communication molecules.
[0080] The detailed data processing flow in Sp3 is as follows:
[0081] The main input data include: the functional substate label output by the Sp2 step for each cell; the precise spatial coordinates (x, y, z) of each cell recorded in the Sp1 step; the original expression profile data of each cell in each molecular marker channel extracted by the Sp1 step, especially the expression of markers known to correspond to specific ligands and receptors; and a structured, predefined biomolecular interaction knowledge base, which should at least contain the ligand name, receptor name, whether they form a known interaction pair, and the characteristic information of high or low expression in which cell types or functional substates. In the data preparation stage, for each cell, combined with its inferred functional substate and original molecular expression profile data, accurately determine its potential as a "signal source" cell for a specific ligand or a "signal sink" cell for a specific receptor. This involves setting a reasonable expression threshold for the expression of the cell's ligand or receptor, or making a judgment based on whether the functional substate to which it belongs is known to be associated with the high expression of the ligand / receptor.
[0082] The interaction potential field is calculated independently for each specific ligand-receptor pair that is of interest to the researcher after screening: the first step is to identify the source and sink cells. Based on the results of the above data preparation, all cell populations in the image that significantly express the ligand L currently being analyzed are screened and recorded as the source cell set ; At the same time, all cell populations that significantly express the corresponding receptor R are screened out and recorded as the sink cell set The second step is to calculate the ligand potential field For each pixel x in the image space (or for computational efficiency, on a downsampled regular grid), the perceived The cumulative signal intensity sum of the released ligand L is mathematically modeled as follows: ; Represents the effective intensity or probability of the i-th source cell expressing or releasing ligand L. This value is directly quantified by the fluorescence intensity of the corresponding ligand marker in the cell and is further quantified by the functional substate to which it belongs. modulated (a certain “activated” functional substate has a higher ligand release efficiency). is the precise spatial position of the ith source cell. is a spatial kernel function that describes the ligand L from the source cell After being emitted, the signal strength spreads in space or its ability to act decays with distance. is the characteristic parameter of the kernel function. For the Gaussian kernel function , Represents the effective radius or standard deviation of the ligand effect. The specific form and parameters of the kernel function KL The selection should be based on an understanding of the biophysical properties of the ligand under study, factors such as its molecular size, diffusion coefficient in tissues, susceptibility to degradation or capture by the matrix, etc. More complex kernel functions are also considered, such as anisotropic models that can simulate the cell - cell barrier effect or hydrodynamic effects within tissues. The third step is to calculate the receptor potential field . Similar to the calculation logic of the ligand potential field, for each pixel (or lattice point) x in the image space, the combined response ability or the total perceived intensity of the ligand signal by all the cells around it expressing the receptor R is modeled as: ; represents the effective intensity of the receptor R expressed by the j - th sink cell or its binding potential with the ligand, which is also quantified by its corresponding fluorescence intensity and modulated by its functional sub - states . is the spatial position of the j - th stem cell. is a spatial kernel function that describes the spatial range within which the receptor cell can effectively sense the ligand signal in its surrounding environment, is its corresponding characteristic sensing radius. The fourth and final step is to calculate the final interaction potential field . This potential field represents the overall potential for a specific L - R interaction to occur at the spatial position x, which is obtained by integrating the ligand accessibility and receptor accessibility at that position through a function f: ; There are various forms for the selection of the function f to reflect different biological interaction logics: The product model is adopted, i.e., PLR(x)=SL(x)⋅SR(x), which assumes that the contributions of the ligand and the receptor are independent and jointly determine the interaction strength. Or the "minimum - constraint model", i.e., , which embodies the "barrel effect" where the interaction strength is limited by the weaker of the two. A model that is more in line with biochemical reaction kinetics is the Michaelis - Menten kinetic model based on receptor saturation effect, , where Km is the Michaelis constant, representing the ligand concentration at which the maximum interaction rate is half. In more advanced modeling, the function f is even a small, pre - trained neural network that takes the and at that position as inputs and outputs the interaction potential by learning the non - linear mapping relationship, thus being able to simulate more complex, cooperative or antagonistic interaction logics.
[0083] The convolution in the spatial domain is transformed into a product in the frequency domain using the Fast Fourier Transform, and then transformed back to the spatial domain through the inverse FFT, thus significantly improving the computational efficiency of large-scale potential field maps. For ultra-large-scale image data, a strategy of dividing the image into blocks and calculating the potential fields of each sub-region in parallel and then stitching them together is also adopted. Modern graphics processors are very suitable for accelerating such pixel- or lattice-based computationally intensive tasks due to their powerful parallel computing capabilities.
[0084] The final data output of this Sp3 step is a series of raster image files. These files are stored in a multi-layer TIFF format or other image formats that support multi-dimensional arrays, where each layer (or each independent file) precisely represents the interaction potential field map of a specific ligand-receptor pair (L-R pair) within the entire image analysis region. . The value of each pixel in the figure directly reflects the potential magnitude or expected intensity value of the occurrence of this specific L-R interaction at that spatial position. Along with these potential field maps, a metadata file or record is also required to clearly describe the name of the specific ligand-receptor pair corresponding to each potential field layer, as well as the key parameters (kernel function type, action radius, expression level threshold, etc.) used in the calculation process, for subsequent analysis, interpretation, and result reproduction.
[0085] The knowledge of biomolecular interactions in Sp3 includes ligand-receptor pair information, and the calculation of the interaction potential field map includes the spatial proximity, expression level, and functional sub-state of ligand-expressing cells and receptor-expressing cells.
[0086] Multiple key biological factors are carefully and quantitatively integrated, so that the constructed potential field model is closer to the real biological process. These factors specifically include: spatial proximity, that is, the physical distance between ligand-expressing cells that emit signals and receptor-expressing cells that receive signals; expression level, that is, the actual expression abundance of relevant ligand and receptor molecules on a single cell; the functional sub-state of the cell, that is, the specific biological state of the cell inferred in the Sp2 step, because cells in different functional sub-states have significant differences in their ability to release, present, or respond to these molecules even if they express the same level of ligand or receptor.
[0087] Regarding the quantification and integration of the expression level, it is necessary to calibrate and normalize the original fluorescence intensity (or other equivalent expression metrics) of each cell in the corresponding ligand or receptor molecule marker channel extracted in the Sp1 step, and then use it as the calculation for this cell as a signal source ( ) or signal sink ( )Core parameters of potential. An expression threshold needs to be set to distinguish between positive and negative expressing cells, or continuous expression values can be used. Regarding the regulatory effect of functional sub-states on expression, the cell functional sub-state information inferred in the Sp2 step needs to be associated with the ligand / receptor expression potential of the cells. A sub-state defined as "highly activated effector T cells" is known to secrete higher levels of a certain cytokine (ligand) or express a higher density of a certain chemokine receptor than the "resting T cell" sub-state. This regulatory relationship is achieved by setting different expression efficiency weight factors for different functional sub-states, which are derived from published literature data, public expression profile databases, or indirectly learned through statistical analysis of the average expression profiles of cells in different functional sub-states in the current dataset. Regarding the modeling of spatial proximity, although the influence of spatial proximity has been implicitly included in the aforementioned kernel function K through its decay with distance, in certain specific biological contexts, when the study is limited to direct physical contact between cells or very short-range paracrine signals, more explicit distance constraints are introduced into the model, or the form of the kernel function is adjusted to emphasize the short-range effect more.
[0088] Tightly combine the original molecular expression data of each cell with the functional sub-state labels assigned to it in the Sp2 step. Before calculating the interaction potential field of a specific ligand-receptor pair, retrieve from a predefined biomolecular interaction knowledge base which known cell types or functional sub-states identified in the current dataset specifically express or highly express the ligand and the receptor. Use this information to pre-screen the potential source cell and sink cell populations in advance, thereby narrowing the actual scope of the subsequent potential field calculation, reducing the noise introduced by non-specific low-level expression, and making the generated potential field map more focused on biologically meaningful interactions.
[0089] Sp4. Extract and integrate at least the following two sets of features to form a set of biological microenvironment state descriptors.
[0090] Sp4 systematically extracts and organizes a comprehensive "descriptor set" from the multi-level and multi-dimensional information generated previously, which can comprehensively characterize the overall state of the current biological microenvironment. Integrate two sets of core feature information from different analysis perspectives: one set is the features directly reflecting the functional heterogeneity of cell populations, and the other set is the features quantifying the potential communication network between cells.
[0091] Sp4.1. Statistical or spatial distribution features derived from the inferred cell functional sub-states;
[0092] From the overall situation of each cell refined functional sub-state inferred from Sp2, extract quantitative features that can macroscopically describe the entire biological microenvironment image or the functional composition of cell populations within a specific region of interest. These features aim to capture the relative abundance, dominant sub-states, sub-state diversity of cells in different functional sub-states, as well as their spatial organization and distribution patterns, so as to depict the overall "portrait" of the biological microenvironment from the perspective of cell function.
[0093] Statistical features: Global proportion / abundance of each functional sub-state: Calculate the percentage of the number of cells in each identified functional sub-state in the image accounting for the total number of cells, or the cell density per unit area.
[0094] Diversity index of functional sub-states: Use the Shannon entropy or Simpson index as ecological diversity indicators to evaluate the richness and evenness of cell functional sub-states in the image. High diversity means a more complex or unstable microenvironment.
[0095] Ratio of specific functional sub-states: Calculate the ratio between some functional sub-states with specific biological significance, such as the ratio of pro-inflammatory sub-state cells to anti-inflammatory sub-state cells, or the ratio of effector immune cells to inhibitory immune cells.
[0096] Spatial distribution features: Spatial aggregation / discreteness of functional sub-states, evaluate whether their cells tend to be aggregated, randomly distributed, or evenly dispersed in space.
[0097] Spatial co-occurrence / mutual exclusion relationship between functional sub-states: Analyze whether cells in different functional sub-states tend to co-occur or mutually exclude each other in space.
[0098] Regional enrichment degree of functional sub-states: If the image is divided into multiple biologically relevant regions (tumor core region, invasive margin region, adjacent normal tissue region), then calculate the enrichment degree or specific distribution density of each functional sub-state in these different regions.
[0099] Community structure of cell functional sub-states: On the cell graph constructed based on cell spatial positions and functional sub-states, run community detection algorithms to identify spatial cell communities composed of cells in specific functional sub-states or mixed by cells in multiple functional sub-states in a specific pattern, and extract features such as the size, shape, and composition of these communities.
[0100] The data output of this Sp4.1 is a numerical feature vector or a set of eigenvalues, representing the macroscopic characteristics of the current biological microenvironment image at the level of cell functional sub-state composition and spatial organization. These features will be part of the set of descriptors for the biological microenvironment state.
[0101] Sp4.2. Quantification features derived from the biomolecular interaction potential field maps; extract features from multiple spatially resolved biomolecular interaction potential field maps generated from Sp3 that can quantitatively describe the intensity, spatial range, and pattern of various potential cell communication activities in the biological microenvironment. Compress the complex, pixel-level potential field map information into a set of numerically descriptive features with high information content and easy to process by models, thereby quantifying the "communication landscape" of the microenvironment.
[0102] Specific data processing and feature extraction include:
[0103] Global intensity features: Calculate the global statistics for the potential field map of each ligand-receptor pair.
[0104] Average potential field intensity: The average value of the pixel values in the entire potential field image, reflecting the overall basic level of this interaction.
[0105] Maximum / minimum potential field intensity: Reflect the peak and valley values of the interaction potential.
[0106] Standard deviation / coefficient of variation of potential field intensity: Describe the spatial heterogeneity or fluctuation degree of the interaction intensity.
[0107] Total integral value of potential field intensity: Reflect the total potential of this interaction within the entire region.
[0108] Proportion of high / low interaction area: Set a threshold and calculate the proportion of the pixel area with potential field intensity higher than a certain high threshold (representing the strong interaction area) or lower than a certain low threshold (representing the weak interaction area) in the total analysis area.
[0109] Spatial pattern and texture features:
[0110] Texture features of the potential field map: Regard the potential field map as a grayscale image and calculate the texture features derived from its gray-level co-occurrence matrix, such as contrast, correlation, energy, entropy, etc., to describe the complexity, uniformity, and directionality of the spatial distribution pattern of the interaction intensity.
[0111] Morphological features of the potential field map: Perform threshold segmentation on the potential field map to obtain a two-dimensional image of the high interaction area, and then calculate the morphological features of these areas, including the number of regions, average area, circularity, elongation, etc., to describe the spatial morphology of the interaction "hot spots".
[0112] Spatial autocorrelation of the potential field map: Calculate Moran's I index or Geary's C coefficient to evaluate the degree of spatial aggregation of the interaction intensity.
[0113] Relationship features between multiple potential field maps:
[0114] Pixel-level correlation between different interaction potential field maps: Calculate the Pearson correlation coefficient or other correlation metrics between the potential field maps of different L-R pairs to explore the spatial cooperation or mutual exclusion relationships between different communication pathways.
[0115] Principal component analysis (PCA) dimensionality reduction: If there are a large number of potential field maps, flatten them and perform PCA to extract the main interaction patterns as features.
[0116] The data output of Sp4.2 is also a numerical feature vector or a set of eigenvalues, characterizing the key features of the current bio-microenvironment image at the level of potential biomolecular communication between cells. These features will also be incorporated into the set of bio-microenvironment state descriptors.
[0117] Sp4.3. Deviation features obtained by comparing the bio-microenvironment image or its derived representation with a preset reference microenvironment state are completed by a pre-trained generative machine learning model; an analysis paradigm based on comparison with a "reference standard" is introduced. Through a pre-trained generative machine learning model, the current bio-microenvironment image to be analyzed is quantitatively compared with one or more pre-defined reference microenvironment states with clear biological significance (such as "healthy / normal" state, "typical disease" state, "drug-sensitive" state, etc.). The result of the comparison is not simply a similarity score, but a "deviation feature" vector that can specifically describe in which aspects, in what way, and to what extent the current microenvironment deviates from the reference state. This deviation feature provides a unique perspective for understanding the abnormal state of the microenvironment.
[0118] The output after Sp4 directly constitutes the deviation feature here. For a new bio-microenvironment image to be analyzed , use the generator trained in Sp4.2 that can map the input image to a certain reference state space to obtain its projection in the reference space , or obtain its latent space representation zquery through the encoder of VAE and reconstruct . The calculation of the deviation feature vector includes:
[0119] Image space deviation: Directly calculate the difference between the original input image and the image after it is "normalized" or "projected" to the reference state by the model . This difference is calculated at the pixel level, calculate the difference image of the two images (or their specific channels) , and then extract the statistical characteristics of the difference image, such as mean, standard deviation, norm (L1 or L2 norm), or calculate the difference value of more advanced image similarity metrics such as the structural similarity index.
[0120] Feature space deviation: If the object of the generative model is a certain derivative feature representation of the image (cell density map or ECM structure parameter map), the deviation is calculated in this feature space.
[0121] Latent space deviation: If a generative model with an explicit encoder is used (such as VAE or certain types of GAN), calculate the original input image in the encoding vector of the model's latent space , and the distance or difference from a set of typical latent vectors representing the reference state (such as the mean or cluster center of the encoding vectors of all samples in the reference dataset). Alternatively, calculate of the encoding and its reconstructed image encoding if the model is trained to align with the reference state.
[0122] Discriminator output (specifically for GAN): If a generative adversarial network is used, its trained discriminator itself is a classifier that can distinguish whether the input data conforms to the reference state distribution. Therefore, input the image to be analyzed into the discriminator , and the probability value or logit value of its output itself serves as a direct indicator of the degree of deviation from the reference state. These deviation measures from different sources together constitute the deviation feature vector , which quantifies the "degree of deviation" and "direction of deviation" of the current biological microenvironment state relative to the preset reference state from multiple perspectives.
[0123] The data output of this Sp4.3 is a numerical deviation feature vector, which will be added to the set of biological microenvironment state descriptors.
[0124] The pre-trained generative machine learning model in Sp4 is a generative adversarial network or a variational autoencoder. The deviation features include the measure of the difference between the image to be analyzed and the corresponding reference state projection generated by the generative model, or the score of the image to be analyzed under the discriminator of the generative model. The specific type of the generative machine learning model used to compare and extract deviation features in Sp4.3, and refine the source and calculation method of the deviation features.
[0125] Application of generative adversarial network:
[0126] Model construction and training: Adopt the CycleGAN architecture and train the generators in two directions, one from the input domain to the reference domain , and one from the reference domain back to the input domain , and two corresponding discriminators. The loss function includes adversarial loss and cycle consistency loss. A dataset referring to the microenvironment state is used to define the target domain B.
[0127] Bias features are extracted from the GAN:
[0128] Transformation difference: The image to be analyzed is transformed into a reference domain image through the trained generator . Then calculate the difference between it and its "ideal reference version" by definition, or the inverse transformation compared. More directly, if there is a mapping from the reference domain to a certain normalized feature space, then compare the image to be analyzed and the reference domain image in the projection of this feature space.
[0129] Discriminator score: The image to be analyzed (or its transformed version , depending on the discriminator design) is input into the trained reference domain discriminator . The scalar value output by it (usually the probability or logit of being judged as a "real reference sample") can be directly used as the bias feature. The lower the value, the farther it deviates from the reference state.
[0130] Application of variational autoencoder (VAE):
[0131] Model construction and training: Train a VAE whose encoder maps the input microenvironment image (or its feature representation) to the latent space z, and the decoder reconstructs the image from z. If the goal is to compare with the reference state, the VAE can be trained on the dataset of the reference state to learn the latent manifold of the reference state.
[0132] Bias features are extracted from the VAE:
[0133] Reconstruction error: The image to be analyzed is input into the VAE trained on the reference dataset to obtain the reconstructed image . Calculate the reconstruction error between and (such as mean square error or structural dissimilarity). A larger reconstruction error usually means
[0134] a greater difference in patterns between and the reference state. . The The Mahalanobis distance or Euclidean distance from the distribution center formed by the reference data in the latent space, or its negative log-likelihood under this distribution, is used as a measure of the degree of deviation.
[0135] KL divergence deviation: If the prior distribution of the VAE is a standard normal distribution and the encoder predicts a posterior distribution for each input , then the KL divergence between this posterior distribution and the prior distribution can itself reflect the "abnormality" degree of the input data.
[0136] These specifically defined deviation features, whether based on the transformation difference / discriminator score of the GAN or the reconstruction error / latent space distance of the VAE, provide unique information about the "atypicality" of the samples for the descriptor set.
[0137] Sp4.4. Input the set of biological microenvironment state descriptors into a pre-trained second machine learning classification model to generate a classification result for the biological microenvironment image.
[0138] The second machine learning classification model makes a final decision and discrimination on the comprehensive "set of biological microenvironment state descriptors" formed by integrating Sp4.1 to Sp4.3, and thus outputs a classification result for the original biological microenvironment image. This classification result is a pre-defined disease subtype, treatment response group, risk level, or other category label with clinical or biological significance. This is the end point for achieving the goal of automated analysis and classification of the entire method.
[0139] Data collection (for training the second machine learning classification model): A dataset containing a large number of biological microenvironment samples is required. For each sample, first process its image through the aforementioned (Sp1 to Sp4.3) methods of the present invention to extract the complete "set of biological microenvironment state descriptors". At the same time, each sample must have an accurate true category label determined by a gold standard method (such as clinical diagnosis results, genomic typing, prognosis or treatment response determined by long-term follow-up).
[0140] Data processing:
[0141] Input: The set of biological microenvironment state descriptors for each sample (a numerical feature vector).
[0142] Feature scaling / normalization: Uniformly scale or normalize all features in the descriptor set so that they have a similar numerical range or statistical distribution, which is crucial for the stable training of many machine learning models.
[0143] Training set / validation set / test set division: Divide the collected labeled dataset into independent training sets, validation sets, and test sets for model training, hyperparameter tuning, and final performance evaluation.
[0144] Handling class imbalance: If there are significant differences in the number of samples in different classes, strategies such as oversampling the minority class, undersampling the majority class, or introducing class weights in the loss function need to be adopted.
[0145] Use the training set to train the selected classification model, and optimize the model parameters by minimizing an appropriate loss function. Use the validation set to monitor the training process, perform early stopping to prevent overfitting, and conduct hyperparameter tuning. Evaluate the performance of the trained classification model on an independent test set. Common evaluation metrics include accuracy, precision, recall, F1-score, area under the receiver operating characteristic curve, area under the precision-recall curve, etc.
[0146] For each input biological microenvironment image (after obtaining the descriptor set through the aforementioned processing), output its final classification result, which is a class label. Also output the confidence or probability score of the model for this classification result.
[0147] The second machine learning classification model in Sp4 is a heterogeneous multi-branch deep neural network. The heterogeneous multi-branch deep neural network includes: at least one graph neural network branch, which is specifically used to process the cell-level graph structure data constructed by cell functional sub-states and their spatial adjacency relationships to learn cell community patterns and local cell niche characteristics; at least one convolutional neural network branch, which is specifically used to process the spatially resolved biomolecular interaction potential field map and the spatial deviation feature map generated by the generative machine learning model to extract its multi-scale spatial texture and intensity patterns; the deep feature vectors extracted by different branches are integrated through a learnable cross-modal attention fusion module, which can adaptively assign confidence weights to various features of different biological microenvironment samples or different spatial regions, and finally a classification head outputs the classification result.
[0148] Input data preparation and splitting:
[0149] Cell-level graph structure data: Based on the cell functional sub-states inferred in Sp2 and their spatial coordinates obtained in Sp1, construct a cell spatial relationship graph for each sample. The nodes of the graph are cells, and the node attributes include their functional sub-state labels and some important single-cell features extracted by Sp1. The edges are defined according to the spatial proximity relationship between cells or more complex rules. This graph data is input into the graph neural network branch.
[0150] Spatially resolved field map data: The biomolecular interaction potential field map generated by Sp3 and the spatial deviation feature map generated by the GAN in Sp4.3. These two-dimensional or three-dimensional image / field map data are input into the convolutional neural network branch.
[0151] Other vector features: the global statistical features of Sp4.1, the global quantization features of the potential field map of Sp4.2, the non-image form deviation features of Sp4.3, etc., are flattened and input into a standard multi-layer perceptron branch, or directly concatenated with the features extracted by the GNN / CNN branch before the fusion stage.
[0152] Graph neural network branch: GAT aggregates information by assigning different weights to neighbor nodes through the attention mechanism, which is more suitable for processing graphs with large heterogeneity of node features and neighborhood structures. Stacking multiple layers of GNNs learns higher-order neighborhood information of nodes. Input node features pass through several layers of GNNs (each layer contains feature transformation, neighborhood aggregation, activation function), and finally learn an embedding vector for each node, or obtain an embedding representation of the entire graph through graph pooling operations (such as global average pooling, attention pooling). Learn the spatial organization pattern of cell populations at the functional sub-state level, such as the cell aggregation area of a specific sub-state and how different sub-state cells interleave to form a specific local cell niche.
[0153] Convolutional neural network branch: For multiple potential field maps or deviation feature maps, a CNN architecture is adopted. A CNN with multiple input channels is used, or CNNs are independently used for each map to extract features and then fused. Through multiple layers of convolution, pooling, activation function, and batch normalization operations, multi-scale spatial textures, intensity patterns, shapes, etc. are extracted from the input two-dimensional or three-dimensional field maps. Finally, a feature vector is obtained through global average pooling or a flattening layer. Learn the spatial intensity distribution law of different communication channels, the morphology and scope of hot spots from the interaction potential field map; learn the specific spatial pattern of the microenvironment deviating from the reference state from the deviation feature map.
[0154] Cross-modal attention fusion module: After the GNN branch and the CNN branch each extract high-level feature vectors, it is necessary to fuse these features from different modalities and different dimensions. The feature vectors of each branch are concatenated and then transformed and dimension-reduced through one or more fully connected layers, where a self-attention mechanism is introduced to learn the importance between different feature dimensions.
[0155] Co - attention: Design a mechanism that enables the features of one modality to "attend to" relevant parts of the features of another modality, and vice versa, thereby learning cross - modal correlations. Transformer - based fusion: Treat the feature vectors of each branch as part of the input sequence of a Transformer encoder or decoder, and utilize its self - attention and cross - attention mechanisms for deep fusion. The core of the attention module is to calculate attention weights, which determine the contribution of each original feature or feature source when finally forming the integrated feature representation. These weights are learnable and can be dynamically adjusted according to the specific situation of the input sample, enabling the model to adaptively allocate higher attention to those features or spatial regions that are most discriminative for the classification of the current specific sample to adapt to the high heterogeneity of the biological microenvironment. The fused feature vectors are finally input into a classification head composed of one or more fully - connected layers, and the probabilities belonging to each predefined class are output through the Softmax activation function.
[0156] The classification result classifies the biological microenvironment image into one of multiple predefined "functional microenvironment prototypes", and each prototype is defined and distinguished by the following set of integrated parameters: the characteristic abundance spectra and spatial co - occurrence patterns of a set of key cellular functional sub - states; the characteristic intensity distributions and topological structures of a set of dominant biomolecular interaction potential fields; and a multi - dimensional deviation vector quantified by deviation features relative to a specific reference microenvironment state; the prototype is associated with the characteristic of a specific disease biological behavior mediated by the microenvironment or the molecular mechanism response at the treatment level.
[0157] Map each biological microenvironment image to be analyzed onto a predefined and well - characterized "functional microenvironment prototype". These prototypes represent biologically distinguishable microenvironment states with specific functional tendencies and tissue patterns. More importantly, the definition of each functional microenvironment prototype is highly structured and multi - dimensional, and it must be precisely described and distinguished from each other through a set of integrated and quantitative parameters. This set of integrated parameters includes at least the following three aspects:
[0158] Cellular functional sub - state composition aspect: It is defined by the relative abundance spectra of a set of key cellular functional sub - states (derived from Sp2) that can characterize the prototype. The abundance spectrum indicates which functional sub - states are enriched, which are sparse, and their relative proportion relationships in this prototype. The characteristic co - occurrence patterns or spatial arrangement rules (whether a certain immune cell sub - state tends to aggregate around a tumor cell sub - state, or whether two different stromal cell sub - states form a specific spatial interleaved structure) of these key functional sub - states in space are also important bases for defining the prototype. These spatial co - occurrence patterns can be quantified by the spatial distribution features (such as the proximity index and co - localization coefficient of specific sub - state pairs) extracted from Sp4.1.
[0159] At the level of biomolecular interaction networks: It is defined by a characteristic intensity distribution pattern and network topology of a set of biomolecular interaction potential fields (derived from Sp3) that are dominant in this prototype. This includes indicating which specific ligand-receptor mediated communication pathways are highly active in this prototype (corresponding to high intensity and wide range in the potential field map), and which are inhibited. The topology of the potential field map, the shape, size, number, connectivity, etc. of the interaction "hot spots" (which can be described by the quantitative features of the potential field map extracted in Sp4.2) also jointly constitute the dimensional characteristics of this prototype.
[0160] At the level of relative deviation from the reference state: The "position" or "difference" of this prototype relative to one or more preset reference microenvironment states (healthy physiological state, early stage of disease, or extreme state that is completely insensitive to treatment) is defined by a multi-dimensional deviation vector. This deviation vector is composed of the deviation features obtained by comparing through a generative machine learning model in Sp4.3, which quantifies the degree and direction of the difference between this prototype and the reference state in multiple dimensions.
[0161] Extract the complete set of state descriptors through the foregoing steps of the method of the present invention, and then use unsupervised or semi-supervised clustering algorithms (such as Gaussian mixture models, or more advanced clustering methods that can integrate multi-modal information) in this high-dimensional descriptor space to identify the naturally occurring and separable "prototype" clusters in the data. Then, further analyze the characteristic centers or representative samples of each cluster, and summarize and refine the integrated definition parameters at the above three levels. In addition, it is necessary to establish the correspondence between these data-driven discovered prototypes and known clinicopathological phenotypes or molecular biological characteristics to endow them with clear biological and clinical significance.
[0162] The output of this method further includes generating and presenting at least one of the following computationally enhanced and contextually relevant visualization analysis maps: The specific visualization analysis maps are of the following types:
[0163] Cell functional substate spatial distribution map, on which the identified driving cell functional substate or its spatial aggregation region that contributes most to the classification result is displayed by superimposing a significance heat map; this visualization map first marks and renders each cell inferred in the Sp2 step on the original high-resolution biomicroenvironment image (or its simplified background version) with different colors or symbols according to its belonging functional substate, so as to clearly show the precise positions and distribution patterns of cells in different functional substates in the tissue space. Further, it utilizes the analysis results obtained by Sp5 and the computational attribution method. It superimposes and displays, in the form of a significance heat map, those driving cell functional substates identified as contributing most to the final biomicroenvironment classification result, or the aggregation regions formed by these driving substate cells in space, on the cell functional substate distribution map. The color depth of the heat map can correspond to the contribution degree or driving intensity of the substate or region. Such superimposed visualization enables users to clearly see at a glance which specific types of functional cells, at which specific spatial positions, constitute the key factors determining the current microenvironment state.
[0164] Comparison visualization map, which juxtaposes or differentially displays the biomolecular interaction potential field map of the current sample with the average potential field map of the determined "functional microenvironment prototype" or the potential field map of a reference sample, highlighting the interaction pattern differences with statistical significance; this visualization map focuses on presenting the characteristics of the biomolecular interaction network. It first extracts several important biomolecular interaction potential field maps calculated in the Sp3 step of the current sample to be analyzed. Then, according to the "functional microenvironment prototype" to which the current sample is classified in the Sp4.4 step, it retrieves the average potential field map corresponding to this prototype from the pre-constructed prototype database, or selects the potential field maps of one or more typical reference samples. Next, this map juxtaposes the potential field map of the current sample with the prototype average potential field map or the potential field map of the reference sample for direct visual comparison. More advancedly, it can calculate the differential potential field map between the two, and highlight the regions with statistically significant differences in interaction intensity or spatial pattern in the differential map through color coding or threshold segmentation. Such comparison visualization helps users understand how the communication network characteristics of the current sample conform to or deviate from the general pattern of its belonging prototype, or what specific interaction changes exist compared with the reference state.
[0165] Based on the deviation features generated by the generative adversarial network, perform spatial back-projection at the cellular or regional scale of the original image to generate a "deviation map", which indicates the structural or functional units with the most significant differences from the reference state and can annotate the specific dimensions of their deviations. Remap the "deviation features" extracted in step Sp4.3, which are usually in the form of global or high-dimensional vectors, back to the spatial dimensions of the original image, thereby generating an intuitive "deviation map". If the deviation features themselves are pixel-level or regional-level difference maps, they can be directly overlaid on the original image as heatmaps. If the deviation features are global vectors, it may be necessary to use the interpretability techniques of the model to identify which regions or cells in the original image contribute the most to this deviation feature, and then highlight these high-contribution regions on the deviation map. The colors or intensities of different regions on the map can indicate the degree of their deviation from the reference state. Furthermore, if the deviation features are multi-dimensional and each dimension corresponds to a specific deviation direction or pattern (deviation dimension 1 represents "excessive fibrosis", deviation dimension 2 represents "insufficient immune cell infiltration"), then interactive queries or automatic annotations can be performed on these highly deviated regions on the deviation map to indicate specifically in which dimensions they significantly deviate from the reference state. This visualization enables users to precisely locate the "most abnormal" or "distinctive" structural or functional units in the microenvironment and understand the specific nature of their abnormalities.
[0166] The technical key points for implementing these advanced visualization atlases include: the need to develop efficient image rendering and overlay algorithms; ensuring the precise spatial registration and fusion of data from different sources (original images, segmentation masks, functional sub-state labels, potential field maps, attribution heatmaps, etc.); and providing a user-friendly interactive interface that allows users to perform operations such as zooming, panning, layer selection, threshold adjustment, and information query.
[0167] Biomicroenvironment images are high-content imaging data with single-cell or subcellular resolution and containing at least 15 distinguishable molecular detection channels, including multiplex immunofluorescence imaging, imaging mass cytometry imaging, and high-dimensional spatial omics imaging data; the multi-parameter information provided by a high number of molecular channels is the basis for the first machine learning model to reliably infer refined functional sub-states and execute in Sp2, as well as support the accurate calculation of biomolecular interaction potential fields and execute in Sp3.
[0168] When inferring refined functional sub-states in the Sp2 step, it is precisely by relying on the complex combined expression patterns of these dozens of molecular markers for each cell (i.e., its high-dimensional "molecular fingerprint") and the derived fine morphological and microenvironmental characteristics for deep learning and pattern recognition that the first machine learning model can surpass the traditional rough cell typing based on a few markers, and thus reliably infer those more subtle and dynamic functional sub-states that are more closely related to the true biological behavior or potential of cells. The more channels there are, the richer the feature dimensions available for distinguishing different sub-states, and the higher the accuracy and fineness of the inference results. Secondly, when accurately calculating the biomolecular interaction potential field in the Sp3 step, the accurate quantification of the expression levels of specific ligand and receptor molecules on a single cell (especially a cell that has been classified into a specific functional sub-state) is a prerequisite for calculating its potential as a signal source or signal sink (i.e., EL,i and ER,j in the aforementioned formula). Having enough molecular channels allows us to simultaneously detect multiple ligands and their corresponding receptors that constitute different communication pathways, so as to be able to construct a more comprehensive, multi-path interaction potential field map, and thus more profoundly understand the complex cell communication network in the biological microenvironment. If the number of molecular channels is too small, it will not be possible to comprehensively capture the state heterogeneity of cells, nor fully characterize the diverse cell-cell interactions, which will severely limit the analysis depth and biological insight that the method of the present invention can achieve.
[0169] The operating steps of the entire technical solution are summarized as follows:
[0170] First, start the system, and then run in sequence:
[0171] Sp1: Image processing and initial feature extraction;
[0172] Sp2: Refined functional sub-state inference;
[0173] Sp3: Interaction potential field map calculation;
[0174] Sp4: Construct state descriptors and classify;
[0175] Sp5: Calculation attribution analysis;
[0176] After running the process to obtain the final output result, then end the system. It should be noted that Sp5: Calculation attribution analysis can be adjusted according to the cost of the application environment and the configuration scheme of the hardware device as the case may be, and this step can be skipped to directly output the result and end the system.
[0177] It should be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a reference structure" does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0178] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for biological microenvironment image analysis and classification, characterized in that: Including the following steps: Sp1. Process the biological microenvironment image, segment the cell components in the image, and extract the multi-parameter initial features of each cell component; Sp2. Based on the multi-parameter initial features and the local microenvironment information of the cells, use a first machine learning model to infer the refined functional sub-states of each cell component, where the functional sub-states are characterized as states related to cell biological behaviors or potentials that go beyond the basic cell types; Sp3. Based on the functional sub-states, spatial location information of the cell components, and predefined biomolecular interaction knowledge, calculate and generate a spatially resolved biomolecular interaction potential field map covering the image analysis region, where the biomolecular interaction potential field map quantifies the potential or intensity of the interaction mediated by specific biomolecules at each point in space; Sp4. Extract and integrate at least the following two sets of features to form a set of biological microenvironment state descriptors: Sp4.
1. Statistical or spatial distribution features derived from the inferred cell functional sub-states; Sp4.
2. Quantitative features derived from the biomolecular interaction potential field map; Sp4.
3. Deviation features obtained by comparing the biological microenvironment image or its derived representation with a preset reference microenvironment state, and a pre-trained generative machine learning model completes the comparison; Sp4.
4. Input the set of biological microenvironment state descriptors into a pre-trained second machine learning classification model to generate a classification result for the biological microenvironment image.
2. The method for analyzing and classifying biological microenvironment images according to claim 1, wherein The first machine learning model for inferring the functional sub-states in Sp2 is an unsupervised learning model or a self-supervised learning model, which identifies the functional sub-states by learning the latent space representation of cell multi-parameter features.
3. The method for analyzing and classifying biological microenvironment images according to claim 1, wherein The biomolecular interaction knowledge in Sp3 includes ligand-receptor pair information, and the calculation of the interaction potential field map includes the spatial proximity, expression level, and functional sub-states of ligand-expressing cells and receptor-expressing cells.
4. A method for biological microenvironment image analysis and classification according to claim 1, characterized in that The pre-trained generative machine learning model in Sp4 is a generative adversarial network or a variational autoencoder, and the deviation features include the measure of the difference between the image to be analyzed and the projection of the corresponding reference state generated by the generative model, or the score of the image to be analyzed under the discriminator of the generative model.
5. A method for analyzing and classifying biological microenvironment images according to claim 1, characterized in that, The method further includes Sp5: Based on the second machine learning classification model and the set of biological microenvironment state descriptors, use a computational attribution method to identify key features or combinations of features that have a significant impact on the classification result, and optionally use this identification result to guide the training of the second machine learning classification model or to interpret its output.
6. A method for analyzing and classifying biological microenvironment images according to claim 5, characterized in that The computational attribution method is configured to: identify the global importance of the biological microenvironment state descriptors that contribute significantly to the classification result; Trace the contribution of the biomolecular interaction potential field map in a specific spatial region to the precise spatial configuration and density threshold of one or more specific cell functional sub-states that cause changes in the potential field intensity and pattern in this region, thereby revealing the spatially specific cell-level functional cooperation or antagonistic pathways that drive the microenvironment classification state.
7. A method for analyzing and classifying biological microenvironment images according to claim 1, characterized in that, The second machine learning classification model in Sp4 is a heterogeneous multi-branch deep neural network, and the heterogeneous multi-branch deep neural network includes: At least one graph neural network branch, which is specifically used to process the cell-level graph structure data constructed by the cell functional sub-states and their spatial adjacency relationships to learn the cell community patterns and local cell niche characteristics; At least one convolutional neural network branch, which is specifically used to process the spatially resolved biomolecular interaction potential field map and the spatial deviation feature map generated by the generative machine learning model to extract its multi-scale spatial texture and intensity patterns; The deep feature vectors extracted by different branches are integrated through a learnable cross-modal attention fusion module, which can adaptively assign confidence weights to various features of different bio-microenvironment samples or different spatial regions, and finally a classification head outputs the classification result.
8. A method for biological microenvironment image analysis and classification according to claim 1, characterized in that, The classification result classifies the bio-microenvironment image into one of a plurality of predefined "functional microenvironment prototypes", and each prototype is defined and distinguished by the following set of integration parameters: The characteristic abundance spectrum and spatial co-occurrence pattern of key cell functional sub-states; The characteristic intensity distribution and topological structure of the dominant biomolecular interaction potential field; A multi-dimensional deviation vector quantified by deviation features relative to a specific reference microenvironment state; the prototype is associated with the molecular mechanism response characteristics of specific disease biological behaviors or treatment levels mediated by the microenvironment.
9. A method for biological microenvironment image analysis and classification according to claim 4, characterized in that The output of Sp4 further includes generating and presenting at least one of the following computationally enhanced and contextually relevant visualization analysis maps: A spatial distribution map of cell functional sub-states, which shows the driving cell functional sub-states or their spatial aggregation regions that contribute the most to the classification result by overlaying a significance heat map; A comparison visualization map, which visually shows the biomolecular interaction potential field map of the current sample juxtaposed or differentiated from the average potential field map of the functional microenvironment prototype or the potential field map of the reference sample, and highlights the interaction pattern differences with statistical significance; Based on the deviation features generated by the generative adversarial network, perform spatial back-projection at the cell or region scale of the original image to generate a deviation map, indicating the structural or functional units with the most significant differences from the reference state, and annotating the specific dimensions of their deviations.
10. A method for biological microenvironment image analysis and classification according to claim 1, characterized in that The bio-microenvironment image is high-content imaging data with single-cell or subcellular resolution and including at least 15 distinguishable molecular detection channels, including multiplex immunofluorescence imaging, imaging mass cytometry imaging, and high-dimensional spatial omics imaging data; the multi-parameter information provided by a high number of molecular channels is the basis for the first machine learning model to reliably infer refined functional sub-states and execute in Sp2, and to support the accurate calculation of the biomolecular interaction potential field and execute in Sp3.
Citation Information
Patent Citations
Systems and methods for characterizing cells and microenvironments
CN115885165A
Drug activity screening method and device
CN118506911A
Cell mechanical phenotype analysis method and equipment based on deep learning
CN119993284A
Spatial metric measurement methods for multiplex imaging and uses in tumor sample analysis
US20240402183A1
Cited By
Stem cell microscopic image feature extraction method and system
CN120526424A
Pathological section analysis device and method based on deep learning
CN121170353A
Immune cell state analysis system based on image processing technology
CN121280446A
Tumor immune microenvironment multi-cell recognition and spatial analysis system
CN122157256A
A tumor immune microenvironment multi-cell recognition and spatial analysis system
CN122157256B