Stress analysis method, system, device and medium for algal high-content images

By using a stress analysis method based on high-content images of algae, and employing quality screening, a dedicated segmentation model, and a binary classification model, we have achieved automated and accurate identification of algal stress states. This solves the problems of low efficiency and strong subjectivity in traditional methods, thereby improving the efficiency and accuracy of the analysis.

CN121564422BActive Publication Date: 2026-07-10ACADEMY OF MILITARY MEDICAL SCIENCES

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ACADEMY OF MILITARY MEDICAL SCIENCES
Filing Date
2025-11-28
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Traditional algal stress analysis methods rely on manual judgment, making it difficult to achieve automated, multi-dimensional stress state identification. Furthermore, existing methods cannot effectively extract subtle morphological changes and complex cellular structural features.

Method used

A stress analysis method based on high-content images of algae is adopted, including quality screening, algae-specific segmentation model, feature extraction, and binary stress analysis model. The stress status of algae is identified through preprocessing, segmentation, feature extraction, and automated judgment.

Benefits of technology

It enables automated and accurate identification of algal stress states, reduces human subjective bias, improves analysis efficiency and accuracy, and ensures the effectiveness and consistency of feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564422B_ABST
    Figure CN121564422B_ABST
Patent Text Reader

Abstract

The application relates to a stress analysis method, system, device and medium for algal high-content images. The method comprises the following steps: acquiring an algal high-content image set, and performing quality screening control processing on the algal high-content image set to obtain algal high-content images that pass quality control; inputting the algal high-content images that pass quality control into a pre-trained algal exclusive segmentation model based on cellpose-sam to obtain algal cell image segmentation masks; based on the algal cell image segmentation masks, performing feature extraction on the algal high-content images to obtain feature vectors for algal and internal structures thereof; and inputting the feature vectors into a pre-trained binary classification stress analysis model to obtain a stress analysis result. The method can realize automatic analysis of the stress state of algae in algal high-content images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, and in particular relates to a method, system, device and medium for stress analysis of high-content images of algae. Background Technology

[0002] With the development of algae research and environmental monitoring technologies, cell image analysis methods based on high-content imaging (HCI) technology have emerged. These methods achieve multi-channel and multi-dimensional information acquisition at the single-cell level, and feature high throughput, high resolution, and multi-parameter parallel analysis.

[0003] Traditional techniques for analyzing algal stress mainly rely on methods such as optical microscopy and manual counting, single-spectrum or fluorescence intensity measurement, and chemical index detection. These methods typically require manual sample selection and subjective judgment to complete the analysis, making it difficult to extract subtle morphological changes and complex cellular structural features from massive image data.

[0004] However, the above methods mostly remain at the stage of empirical thresholds or single-indicator judgment, and cannot achieve automated, multi-dimensional identification of stress states. Summary of the Invention

[0005] Therefore, it is necessary to provide a stress analysis method, system, device, and medium for high-content algal images that can achieve automated identification and visualization analysis of algal stress states, addressing the aforementioned technical problems.

[0006] In a first aspect, this application provides a stress analysis method for high-content algal images, including:

[0007] A high-content image set of algae was acquired, and the high-content image set of algae was subjected to quality screening and control processing to obtain high-content algae images that passed quality control.

[0008] High-content algae images that have passed quality control are input into a pre-trained algae-specific segmentation model based on cellpose-sam to obtain an algae cell image segmentation mask.

[0009] Based on algal cell image segmentation mask, feature extraction is performed on high-content algal images to obtain feature vectors oriented towards algae and their internal structure.

[0010] The feature vectors are input into a pre-trained binary classification stress analysis model to obtain stress analysis results.

[0011] In one embodiment, the high-content algae image set is subjected to quality screening and control processing to obtain high-content algae images that pass quality control, including:

[0012] Obtain the file metadata and pixel data of each original high-content algae image in the high-content algae image set;

[0013] Based on the preset expected image size range, file integrity is verified according to file metadata, and the original high-content algae image with a complete file is determined as the expected image;

[0014] Pixel statistics are performed on the pixel data of the desired image, and the desired image that meets the preset zero pixel ratio and saturated pixel ratio is identified as a high-content algae image that has passed quality control.

[0015] In one embodiment, the algae-specific segmentation model based on cellpose-sam obtains the algae cell image segmentation mask using the following method:

[0016] Standardized high-content algal images that have passed quality control are preprocessed to obtain standardized high-content algal images;

[0017] Probable algal cell regions are identified from standardized high-content algal images to obtain candidate algal cell regions.

[0018] Based on preset algae-specific parameters, the complete outline of algae cells is identified within the candidate region bounding box of algae cells to obtain the cell region mask; the algae-specific parameters include flow threshold and cell probability threshold;

[0019] The region of adherent cells within the cell region mask is determined based on the average area threshold of algal cells.

[0020] The watershed separation algorithm is used to classify cells in the region of adherent cells, and the separated cells are subjected to morphological closing operation to obtain an algal cell image segmentation mask.

[0021] In one embodiment, based on an algal cell image segmentation mask, feature extraction is performed on a high-content algal image to obtain a feature vector oriented towards algae and their internal structure, including:

[0022] Based on a preset filename regularization expression, the corresponding feature extraction channels are determined according to the filename type of the high-content algae image; the feature extraction channels include the bright field channel, chloroplast channel, DAPI channel, and cell mask channel;

[0023] The high-content algae images were sequentially processed by grayscale conversion, intensity standardization, target feature enhancement, and Gaussian filtering for noise reduction to obtain enhanced high-content algae images.

[0024] Based on algal cell image segmentation masks, a multi-level recognition strategy is used to identify cell nuclei, cell bodies, and chloroplasts in enhanced high-content algal images; the multi-level recognition strategy is determined by feature extraction channels.

[0025] Feature measurements were performed on each cell nucleus, cell body, and chloroplast in the feature extraction channel to obtain feature data; the feature data included size and shape features, intensity features, texture features, colocalization features, neighbor relationship features, overlap features, and image quality features;

[0026] The feature data is cleaned, and the cell nucleus-cell body association and chloroplast-cell body association are established based on the cleaned feature data to obtain feature vectors.

[0027] In one embodiment, feature measurements are performed on each cell nucleus, cell body, and chloroplast in the feature extraction channel to obtain feature data, including:

[0028] Based on the pixel spatial distribution of the algal cell image segmentation mask, the size and shape of each cell nucleus, cell body, and chloroplast are calculated to obtain size and shape features. The size and shape features include area, perimeter, eccentricity, compactness, extensibility, axial length, Euler number, compactness, and center coordinates.

[0029] The intensity features of each cell nucleus, cell body, and chloroplast are calculated based on the pixel gray values ​​corresponding to the enhanced high-content images of algae using feature extraction channels; the intensity features include average intensity, integral intensity, intensity standard deviation, maximum intensity, and minimum intensity;

[0030] The gray-level co-occurrence matrix is ​​obtained based on the gray-level pixel spatial distribution of the algal cell image segmentation mask. The texture features of the cell nucleus, cell body, and chloroplast are calculated based on the gray-level co-occurrence matrix. The texture features include the second moment of the angle, contrast, correlation, entropy, and inverse difference moment.

[0031] The feature extraction channels are paired to form channel pairs, and the pixel spatial overlap of each channel pair is calculated to obtain co-localization features. Co-localization features include linear correlation coefficient, rank correlation coefficient, overlap coefficient, and gray-level contribution overlap coefficient.

[0032] The spatial relationships of similar targets in the cell nucleus, cell body, and chloroplast are calculated separately to obtain the neighbor relationship characteristics. The neighbor relationship characteristics include nearest neighbor distance, number of neighbors, and average neighbor distance.

[0033] The overlap relationship between the cell nucleus, cell body and chloroplast is calculated based on the overlap ratio of the algal cell image segmentation mask, and the overlap feature is obtained.

[0034] Image quality features are obtained by performing full-image pixel statistics and calculating quantitative quality indicators for high-content algae images.

[0035] In one embodiment, the binary stress analysis model is obtained through the following method:

[0036] A training dataset was constructed using high-content images of unstressed algae and high-content images of algae under known stress. The training dataset includes feature vectors and corresponding stress labels.

[0037] Supervised learning methods were used to train a binary classification model on the training dataset to obtain the hyperparameters, random seeds, and training version information of the binary stress analysis model.

[0038] The discriminative ability of the binary stress analysis model is evaluated using assessment indicators, and the models whose discriminative ability meets the expectations are serialized and saved to obtain the binary stress analysis model.

[0039] In one embodiment, the method further includes:

[0040] Transform the results of the coercion analysis into textual conclusions;

[0041] When the stress analysis result indicates stress damage, an algal stress visualization map is generated based on the feature vectors. The algal stress visualization map is at least one of the following visualization forms: feature correlation network map, feature distribution heatmap, and principal component analysis (PCA) lithotripsy map.

[0042] Secondly, this application also provides a stress analysis system for high-content algal images, comprising:

[0043] The data module is used to acquire a high-content algae image set and perform quality screening and control processing on the high-content algae image set to obtain high-content algae images that meet the quality control requirements.

[0044] The image segmentation module is used to input high-content algae images that have passed quality control into a pre-trained algae-specific segmentation model based on cellpose-sam to obtain an algae cell image segmentation mask.

[0045] The feature extraction module is used to extract features from high-content algae images based on algae cell image segmentation masks, and obtain feature vectors oriented towards algae and their internal structures.

[0046] The stress analysis module is used to input feature vectors into a pre-trained binary classification stress analysis model to obtain stress analysis results.

[0047] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above-described methods for stress analysis of high-content images of algae.

[0048] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the above-described methods for stress analysis of high-content images of algae.

[0049] The aforementioned stress analysis methods, systems, equipment, and media for high-content algal images, through quality screening and control of the high-content algal image set, can eliminate low-quality images, ensuring that subsequent analysis is based on reliable data and reducing the interference of invalid data on the results. A pre-trained algal-specific segmentation model based on Cellpose-SAM is employed, which is more adapted to the morphological characteristics of algal cells than general segmentation models, and can generate segmentation masks more accurately, providing accurate target regions for feature extraction. Based on the segmentation mask, features of algae and their internal structures are extracted, ensuring that the extracted features are directly related to the algae themselves and key internal structures, avoiding background or impurity interference, and improving the effectiveness of the features. A pre-trained binary classification stress analysis model processes the feature vectors, replacing manual judgment, enabling rapid output of algal stress status, improving analysis efficiency, and reducing human subjective bias. Through a complete process of image quality control, cell segmentation, feature extraction, and stress assessment, the system achieves automated analysis of algal stress status in high-content images. It transforms the massive algal image data acquired by high-content imaging technology into quantifiable stress analysis results through a series of technical processes, solving the problems of low efficiency, strong subjectivity, and one-sided information from traditional manual observation, which makes it difficult to achieve efficient and accurate assessment of algal stress status. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a flowchart illustrating the stress analysis method for high-content algae images according to the present invention.

[0052] Figure 2 This is a flowchart illustrating the steps of step S102.

[0053] Figure 3 This is a flowchart illustrating the steps of step S103.

[0054] Figure 4 This is a structural diagram of the stress analysis system for high-content algae images of the present invention. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0056] In one embodiment, such as Figure 1 As shown, a stress analysis method for high-content algae images is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0057] S101. Obtain a high-content image set of algae, and perform quality screening and control processing on the high-content image set of algae to obtain high-content images of algae that meet the quality control requirements.

[0058] In illustrative terms, a high-content algal image set refers to a collection of images obtained by high-throughput imaging of algal samples using high-content imaging systems such as Operetta and ImageXpress. It includes one or more images from different imaging channels, such as bright field, fluorescently labeled chloroplast channels, and DAPI nuclear channels, and can cover samples from different algal species, different growth stages, or different stress levels.

[0059] Furthermore, low-quality images caused by imaging errors or sample contamination are removed through quality screening and control. For example, based on the image's metadata, such as file format, resolution, and imaging parameter records, the integrity of the file is verified, and images that are damaged or have abnormal formats are excluded. Secondly, the pixel data of the image is statistically analyzed to filter out images with large areas of missing background or overexposure, and finally, images that pass quality control are retained.

[0060] S102. Input the high-content algae images that have passed quality control into the pre-trained algae-specific segmentation model based on cellpose-sam to obtain the algae cell image segmentation mask.

[0061] The illustrative algae-specific segmentation model based on Cellpose-SAM is a deep learning model that integrates cell morphology priors with general image segmentation capabilities. Specifically, it is fine-tuned on the Cellpose-SAM basic model of Cellpose 4.0 using algae-specific samples containing algae images with different morphologies and stress states. This model retains Cellpose's sensitive recognition ability of biological cell edge contours, while combining SAM (Segment Anything Model)'s generalized recognition ability of targets in complex backgrounds. It is specifically optimized for the characteristics of algae cells, such as irregular morphology, easy adhesion, and internal structures, and can accurately distinguish algae cells from the background and internal cell structures, such as cell nuclei and chloroplasts.

[0062] Specifically, the segmentation process involves inputting high-content images of algae that have passed quality control into the model. The model then automatically identifies algal cell regions in the image based on learned algal morphological features and generates corresponding segmentation masks. These segmentation masks are single-channel images that use pixel value differences to mark cell regions, achieving pixel-level cell boundary delineation and defining a clear analysis range for subsequent feature extraction.

[0063] S103. Based on the algal cell image segmentation mask, feature extraction is performed on the high-content image of algae to obtain feature vectors oriented towards algae and their internal structure.

[0064] In a schematic manner, feature extraction utilizes segmentation masks to pinpoint the spatial location of algal cells and their internal structures. Combined with multi-channel pixel information from high-content images, such as grayscale values ​​and texture distribution, the morphological, functional, and spatial relationship features of the cells are quantified and ultimately integrated into a structured feature vector. Specifically, by identifying image channel types, such as bright field, chloroplast, and DAPI, the analysis objects—the entire cell, chloroplasts, and cell nucleus—are determined. Image preprocessing then enhances the feature signals. Furthermore, based on the segmentation mask, targeted algorithms are used to extract multi-dimensional features, including cell size and shape, fluorescence intensity, internal texture, and inter-structural colocalization. All features are integrated into a one-dimensional feature vector according to preset rules. Each vector corresponds to a comprehensive quantitative description of an algal cell or image region, providing data support for subsequent stress analysis.

[0065] S104. Input the feature vector into the pre-trained binary classification stress analysis model to obtain the stress analysis results.

[0066] The binary stress analysis model is a supervised learning-based machine learning model trained on algal samples under known stress conditions. It learns the difference patterns in feature vectors between normal and stressed algae. Illustratively, during model pre-training, feature vectors labeled with either no stress or stress are used as input. The model parameters are adjusted through optimization algorithms, enabling the model to output the corresponding stress state judgment based on newly input feature vectors.

[0067] Specifically, after inputting the feature vector into the pre-trained binary stress analysis model, the model compares the features in the vector with the difference patterns learned during the training phase, such as the decrease in cell area and the decrease in chloroplast intensity under stress, and calculates the probability that the sample belongs to the stress category. If the probability exceeds a preset threshold, the model outputs the stress result; otherwise, it outputs that the algae are not under stress, thus realizing the automatic judgment of the stress state of algae.

[0068] In the aforementioned stress analysis method for high-content algal images, the quality screening of the original image set eliminates images with incomplete files or abnormal pixels, directly improving the reliability of the input data and ensuring the accuracy of subsequent segmentation, feature extraction, and stress analysis results, thus avoiding analysis errors caused by low-quality data. The pre-trained cellpose-sam-based algal-specific segmentation model is optimized for algal cell characteristics, and compared to general models, it can more accurately identify algal cell boundaries and generate segmentation masks with higher accuracy, providing a clear target region definition for subsequent feature extraction and reducing feature distortion caused by segmentation errors. Based on the segmentation mask, features of algae and their internal structures can be extracted in a targeted manner, ensuring that features are directly related to the analysis target, avoiding interference from irrelevant background information, and improving the targeting and effectiveness of feature vectors. The pre-trained binary classification stress analysis model learns the feature patterns of known stress state samples, enabling rapid classification of newly input feature vectors and automated judgment of stress states. This significantly improves efficiency compared to manual analysis, and the model's judgments are highly consistent, reducing subjective differences among different analysts.

[0069] In one embodiment, the high-content algae image set is subjected to quality screening and control processing to obtain high-content algae images that pass quality control, including:

[0070] S11. Obtain the file metadata and pixel data of each original high-content algae image in the high-content algae image set.

[0071] File metadata consists of structured information automatically recorded by the high-content imaging system during image acquisition. It includes basic image attributes and acquisition parameters, such as file format, resolution, imaging channel identifiers, exposure time, and focal length. This metadata is typically embedded in the image file header and can be directly read and extracted using Python's PIL library, OpenCV library, or specialized image processing software. Pixel data is the core content of the image, represented as a two-dimensional pixel grayscale matrix or a multi-channel matrix. The grayscale value of each pixel corresponds to the light signal intensity at that location during imaging. Pixel data acquisition requires loading the image file using an image processing library and converting it into a computable numerical matrix. For example, OpenCV's cv2.imread() function can be used to read the image and convert it into a grayscale matrix, providing a data foundation for subsequent pixel statistics.

[0072] S12. Based on the preset expected image size range, perform file integrity verification according to file metadata, and determine the original algae high-content image with the verification result as the expected image.

[0073] Image size is a key indicator of file integrity. When the acquisition parameters of a high-content imaging system, such as objective lens model and sensor size, are fixed, the resolution of the output image is uniquely determined. If the image size deviates from the preset range, it usually means that the file was damaged during transmission, such as data loss leading to resolution truncation, or incorrect imaging system parameters or file format errors. If such images are used in subsequent analysis, they will cause the segmentation model to report errors due to size mismatch, or the feature extraction results to be distorted.

[0074] Specifically, based on the hardware specifications of the imaging system and the imaging parameters of the experimental design, a preset desired image size range is established, which can be set to 1024×1024 pixels or 2048×2048 pixels. The metadata resolution of each original image is compared with the preset size range. If the resolution matches perfectly, the image file is determined to be complete, with no data loss or parameter errors, and it is classified as a desired image. If the resolution deviates from the range, the file is determined to be incomplete and is directly rejected without proceeding to the subsequent pixel statistics stage, thus improving the screening efficiency.

[0075] S13. Perform pixel statistics on the pixel data of the desired image, and determine the desired image that meets the preset zero pixel ratio and saturated pixel ratio as a qualified high-content algae image.

[0076] While the images may pass file integrity checks, they may contain content-level invalidity issues. For example, insufficient sample loading in the well plates could result in large areas of no signal (high proportion of zero pixels), or overexposure could lead to large areas of signal saturation (high proportion of saturated pixels). The pixel information in such images typically shows zero-pixel areas lacking cells and saturated pixel areas losing cell details, failing to reflect the true state of algal cells. Further pixel statistics are needed for screening. The zero-pixel ratio refers to the proportion of pixels with a grayscale value of 0 out of the total number of pixels. Zero pixels usually correspond to areas without samples, such as blank areas at the edges of the well plates, or areas lacking imaging signal, such as areas not covered by samples. A preset threshold of 5% is recommended; if the ratio is higher than 5%, it indicates that the proportion of effective cell areas in the image is too low and has no analytical value. The saturated pixel ratio refers to the proportion of pixels with the maximum grayscale value out of the total number of pixels. Saturated pixels correspond to areas with excessively strong imaging signals, such as those caused by excessively long exposure times leading to loss of cell details. A preset threshold of 3% is typically recommended; if the ratio is higher than 3%, it will cause distortion in intensity features, such as chloroplast fluorescence intensity calculations, failing to reflect the true state of cells.

[0077] The pixel matrix of each desired image is traversed and counted to calculate the number of zero pixels and the number of saturated pixels. The corresponding ratios are then divided by the total number of pixels. If both ratios meet the preset threshold, the image content is deemed valid and classified as a high-content algae image that has passed quality control. If either ratio exceeds the threshold, the content is deemed invalid and is removed.

[0078] Optionally, a simple record can be made of the removed image files during the removal process to ensure subsequent image traceability and auditing.

[0079] In one embodiment, such as Figure 2 As shown, the algae-specific segmentation model based on cellpose-sam obtains the algae cell image segmentation mask using the following method:

[0080] S201. Standardize the high-content algae images that have passed quality control to obtain standardized high-content algae images.

[0081] To illustrate, quality-controlled images are uniformly adjusted to a resolution suitable for the segmentation model. If the original image size is too large, it is reduced using bilinear interpolation; if the size is too small, it is enlarged using interpolation, avoiding misidentification of cell morphology due to size mismatch. Furthermore, high-content images typically contain multiple channels. During preprocessing, the channel calibration information provided by the imaging system must ensure strict alignment of each channel in spatial coordinates to avoid confusion in subsequent structure recognition due to channel misalignment. Optionally, the channel that best reflects the algal cell outline, such as the bright-field channel or chloroplast channel, is prioritized as the main segmentation channel to reduce redundant information interference. Further, to address the issue of blurred edges in algal cell images, an adaptive histogram equalization algorithm is used to enhance the grayscale difference between cell edges and the background, making the cell membrane outline clearer. Simultaneously, Gaussian filtering is used to smooth the image, removing salt-and-pepper noise generated during imaging to prevent noise from being misjudged as cell details. The filtering radius can be set according to the cell size, such as 1-3 pixels.

[0082] S202. Perform possible algal cell region identification on the standardized high-content algal image to obtain candidate algal cell region boxes.

[0083] Furthermore, in high-content algal images, in addition to the target cells, there are often interferences such as culture medium impurities, air bubbles, or background textures. Relying on the general target recognition capability of SAM in the CellPose-SAM model, combined with the characteristics of algal samples trained in the early stage, such as the circular / elliptical shape of cells and specific gray-level ranges, a global scanning analysis is performed on the standardized images. Specifically, the global semantic features such as the regional gray-level distribution and shape complexity of the image are extracted by deep learning algorithms, and then compared with the typical features of algal cells in the training set to identify continuous regions whose gray-level distribution and morphology match the characteristics of algal cells. Obvious blank backgrounds and large impurity blocks are excluded, and the minimum bounding rectangle, i.e., the candidate region box, is drawn for each identified region. The region inside the box is the target range for subsequent fine segmentation.

[0084] The output of the candidate region box must meet the principle of high recall, that is, to cover all real cell regions as much as possible, allowing a small amount of background or impurities, and to avoid cell loss due to missed detection.

[0085] S203. Based on preset algae-specific parameters, identify the complete outline of algae cells within the candidate region box of algae cells to obtain the cell region mask; the algae-specific parameters include flow threshold and cell probability threshold.

[0086] Furthermore, by leveraging the cell morphology prior inference capabilities of the Cellpose module and combining it with algae-specific parameters, the complete outline of each cell is accurately delineated, generating a preliminary cell region mask. The algae-specific parameters are key parameters for tuning based on algae cell characteristics, such as easily blurred edges, high transparency, and irregular morphology. These include the flow threshold and the cellprob threshold. The flow threshold controls the determination of cell edge continuity. Algae cells often exhibit local edge breaks; setting the flow threshold to a low value, such as flow_threshold=0.4, allows the model to classify small gaps at the edges as continuous contours, avoiding the segmentation of a single cell into multiple fragments. The cellprob threshold controls the sensitivity of determining whether a region is a cell. Transparent algae cells have small grayscale differences from the background; setting the cellprob threshold to a negative value, such as cellprob_threshold=-0.5, reduces the model's dependence on high grayscale differences, preventing the misclassification of transparent cells as background.

[0087] For example, within the candidate region box, the model first calculates the motion vector of the pixel using CellPose's flow field prediction algorithm to simulate the direction of the cell edge. Combined with the flow threshold, continuous edge pixels are selected. Then, based on the cell probability threshold, regions with cell probabilities higher than the threshold are marked as cells. Finally, a binary cell region mask is generated, where regions with a pixel value of 1 in the mask represent cells and 0 represents the background.

[0088] S204. Determine the adherent cell region within the cell region mask based on the average area threshold of algal cells.

[0089] Algal cells tend to aggregate and form clusters of adherent cells, which may be identified as a connected region in the initial segmentation mask. Directly using this cluster for feature extraction can lead to problems such as excessively large cell areas and distorted morphological features. The average area threshold is an empirical value set based on the area statistics of a large number of normal algal cells. During model training, the area of ​​at least 1000 single, non-adherent algal cells can be measured, and their average and standard deviation can be calculated. The average area plus twice the standard deviation is used as the adhesion determination threshold. For example, if the average area of ​​Chlorella is 200 pixels and the standard deviation is 50 pixels, then the threshold is set to 300 pixels.

[0090] To illustrate, in the cell region mask, all connected regions are traversed, each connected region representing a cell that has been initially identified. The number of pixels in each region is calculated, i.e., the area. If the area of ​​a region exceeds the average area threshold, it is determined to be a region of adhered cells, presumably containing multiple overlapping cells; if the area is within the threshold range, it is determined to be a single cell region, requiring no further processing.

[0091] S205. The watershed separation algorithm is used to classify cells in the adherent cell region, and the separated cells are subjected to morphological closing operation to obtain an algal cell image segmentation mask.

[0092] Specifically, the distance transformation map of the adhered region is calculated, where the distance value of each pixel represents the shortest distance to the edge of the region, with the cell core region having the largest distance value, farthest from the edge. Seed points are set based on these distance values, i.e., pixels with the largest local distance values, each seed point corresponding to a potential single cell. Further, a watershed separation algorithm accumulates water starting from the seed points. When the water levels at different seed points meet, a dividing line, or watershed, is formed, ultimately segmenting the adhered region into multiple independent sub-regions, each corresponding to a separated cell. Further, the edges of the separated cells may have small gaps due to blurring of the original image or algorithm errors. These gaps can be filled using a dilation-erosion closing operation. The dilation operation first expands the cell edge to close the gap, and the erosion operation then shrinks the edge back to its original size, ultimately obtaining a cell mask with continuous and complete contours. In this mask, each algal cell is independently labeled, the background is 0, and the cell edges are complete and without adhesion.

[0093] In one embodiment, such as Figure 3 As shown, based on an algal cell image segmentation mask, feature extraction is performed on high-content algal images to obtain feature vectors oriented towards algae and their internal structures, including:

[0094] S301. Based on the preset filename regularization expression, determine the corresponding feature extraction channel according to the filename type of the high-content algae image; the feature extraction channels include the bright field channel, chloroplast channel, DAPI channel and cell mask channel.

[0095] Specifically, high-content images of algae typically contain multi-channel information. For example, the brightfield channel reflects the overall cell morphology, the chloroplast channel marks photosynthetic structures, and the DAPI channel shows the location of the cell nucleus. Different channels correspond to different analysis targets, such as the cell body, chloroplasts, and cell nucleus. Incorrect channel identification can lead to a mismatch between the feature extraction object and the target structure, directly affecting the effectiveness of the features. Illustratively, a pre-defined filename regularization expression is the core tool for automatic channel identification. It uses pre-defined string matching rules, such as regular expressions, to parse channel type information from image filenames. For example, filenames containing "brightfield" identify the brightfield channel, "chloroplast" identify the chloroplast fluorescence channel, "DAPI" identify the cell nucleus fluorescence channel, and "mask" identify the previously generated cell segmentation mask channel. By calling this regular expression using Python's `re` module, image filenames can be automatically scanned, matched, and labeled with the corresponding channel type for each file, thereby determining the type of features to be extracted for that channel.

[0096] S302. Perform grayscale conversion, intensity standardization, target feature enhancement, and Gaussian filtering noise reduction on the high-content algae image in sequence to obtain an enhanced high-content algae image.

[0097] Optionally, multi-channel color images can be converted to single-channel grayscale images to simplify data dimensions while preserving structural contour information. Furthermore, grayscale values ​​can be linearly mapped from their original range to the 0.0-1.0 interval, eliminating absolute intensity deviations caused by different imaging batches and making the intensity features of different images comparable. For specific structures, such as chloroplasts, a targeted enhancement algorithm is employed. For example, a Top-Hat transform is applied to the chloroplast channel image to highlight local brightness areas higher than the background and suppress uniform background; the Laplacian operator is used to enhance cell edges, strengthening the grayscale difference between the cell membrane and its surrounding environment. A Gaussian kernel with a preset radius is used to smooth the image, attenuating high-frequency noise while preserving low-frequency information of the cell structure, preventing noise from being misjudged as valid features.

[0098] S303. Based on the algal cell image segmentation mask, a multi-level recognition strategy is adopted to identify the cell nucleus, cell body and chloroplast in the enhanced algal high-content image; the multi-level recognition strategy is determined by the feature extraction channel.

[0099] For example, cell body recognition is based on an algal cell image segmentation mask, where the area with a pixel value of 255 in the mask is the cell body range; cell nucleus recognition can rely on the DAPI channel, utilizing the characteristic that DAPI dye specifically binds to nuclear DNA, making the cell nucleus appear bright in this channel. Within the cell body mask range, the Otsu global thresholding method is used to automatically determine the grayscale threshold, and the area above the threshold is marked as the cell nucleus; chloroplast recognition relies on the chloroplast channel, which contains chlorophyll autofluorescence or specific labeling. Similarly, within the cell body mask range, an adaptive thresholding method is used, and a local threshold is calculated through a sliding window to mark the high-brightness area as chloroplast.

[0100] S304. Perform feature measurements on each cell nucleus, cell body, and chloroplast in the feature extraction channel to obtain feature data; the feature data includes size and shape features, intensity features, texture features, colocalization features, neighbor relationship features, overlap features, and image quality features.

[0101] In a schematic manner, the feature measurement process relies on automated algorithms from tools such as CellProfiler. Based on the pixel information of identified structural regions such as cell bodies, nuclei, and chloroplasts and their corresponding channels, features are extracted according to the principle of matching structure type with channel characteristics. Specifically, feature measurement covers seven categories of indicators. Among them, size and shape features are calculated based on the spatial distribution of pixels in the structural region, reflecting the physical morphology of the structure; intensity features are calculated based on the pixel grayscale values ​​of the corresponding channels, reflecting the molecular content or activity of the structure; texture features are calculated based on the spatial distribution pattern of grayscale values ​​within the structural region, reflecting the microscopic distribution uniformity of the structure; co-localization features are calculated based on the spatial overlap of structures in different channels, reflecting the positional association between structures; neighbor relationship features are calculated based on the spatial positional relationship of structures of the same type, reflecting the group distribution pattern of structures; overlap features are calculated based on the area overlap ratio of different structural masks, reflecting the inclusion relationship of structures; and image quality features are calculated based on the statistical calculation of pixels across the entire image, providing a basis for verifying the reliability of feature measurement. These features together constitute a multi-dimensional descriptive system from individual to group, from morphology to function, ensuring that subtle changes caused by stress can be quantified and captured.

[0102] S305. Clean the feature data and establish the cell nucleus-cell body association and chloroplast-cell body association based on the cleaned feature data to obtain the feature vector.

[0103] The original feature data may contain noise or redundancy, and the attribution relationship of features of different structures must be clearly defined; otherwise, it will lead to errors in subsequent model analysis. As an illustration, the Z-score method is used to identify and remove extreme features caused by image noise or segmentation errors, such as a cell area of ​​10,000 pixels, which is far beyond the normal range. Furthermore, features with all median values ​​empty or standard deviations of 0 are deleted, such as texture features with the same value in all cells, lacking discriminative power. For a small number of missing features due to incomplete structure identification, the median of the same batch of samples is used to fill in the missing features, avoiding sample loss.

[0104] Furthermore, based on the spatial location of the structure, such as whether the coordinates of the cell nucleus center are within the cell body mask, a one-to-one association between the cell nucleus and cell body is established, meaning one cell body contains one cell nucleus. Based on the spatial inclusion relationship between chloroplasts and cell bodies, a one-to-many association between chloroplasts and cell bodies is established, meaning one cell body can contain multiple chloroplasts. Finally, all features of the same cell body, including its own morphological features, the features of its associated cell nucleus, and the summarized features of its associated chloroplasts, are integrated into a single record to form the feature vector of that cell. The output format of the feature vector is structured data, such as a CSV table, where each row represents one cell and each column represents one feature, providing standardized data for the input of subsequent binary classification models.

[0105] In one embodiment, feature measurements are performed on each cell nucleus, cell body, and chloroplast in the feature extraction channel to obtain feature data, including:

[0106] S21. Based on the pixel spatial distribution of the algal cell image segmentation mask, calculate the size and shape of each cell nucleus, cell body and chloroplast to obtain size and shape features; size and shape features include area, perimeter, eccentricity, compactness, extensibility, axial length, Euler number, compactness and center coordinates.

[0107] Size and shape characteristics are fundamental indicators for describing the physical morphology of algal cells and their internal structures. Their core value lies in capturing morphological anomalies caused by stress, such as cell swelling, nuclear deformation, and chloroplast fragmentation, by quantifying the spatial dimension and geometric shape of the structure. Illustratively, this can be achieved by analyzing the pixel coordinates and distribution patterns of the target region in a mask. Specifically, the total number of labeled pixels in the target mask for the nucleus, cell body, and chloroplasts is counted separately. Combined with the spatial resolution of the imaging system, this can be converted into the actual area, directly reflecting the size of the structure; for example, the area increases under stress. By tracing the contours of pixels at the mask edges and calculating the sum of Euclidean distances between adjacent edge pixels, the total length of the structure's edge, i.e., the perimeter, is reflected. For example, cell wall damage can lead to an abnormal increase in perimeter. The target region is fitted to a minimum circumscribed ellipse, and the ratio of the ellipse's major axis to its minor axis is calculated. The closer the ratio is to 1, the closer the structure is to a circle. For example, the eccentricity of a normal cell nucleus is about 1.2, but under stress it may increase to 1.8, resulting in elliptical deformation. Compactness is calculated by dividing the actual area of ​​the target by the area of ​​the smallest convex polygon, reflecting the fullness of the structure. For example, when chloroplasts break under stress, compactness decreases from 0.9 to 0.6. Extensibility is calculated by dividing the actual area of ​​the target by the area of ​​the smallest bounding rectangle, reflecting the filling efficiency of the structure within a rectangular space. For example, cell stretching reduces extensibility. The lengths of the major and minor axes of the fitted ellipse directly describe the size and stretching direction of the structure. The Euler number is calculated by subtracting the number of pores from the number of connected regions, reflecting the internal integrity of the structure. For example, when vacuoles appear in cells, the number of pores increases, and the Euler number decreases. Compactness is calculated based on 4π × area / perimeter². The closer the value is to 1, the closer the structure is to an ideal circle. For example, when stress causes cell edge wrinkling, compactness decreases. The average X and Y coordinates of all pixels in the target region are calculated to determine the spatial location of the structure, i.e., the center coordinates.

[0108] The combination of these features can comprehensively characterize the size, shape, and integrity of a structure, serving as the fundamental basis for distinguishing between normal and stressed states. For example, Scenedesmus under heavy metal stress often exhibits increased cell area, increased eccentricity, and decreased compactness; these changes can be precisely captured through quantitative differences in size and shape characteristics.

[0109] S22. The intensity features of each cell nucleus, cell body and chloroplast are calculated based on the pixel gray values ​​corresponding to the enhanced high-content image of algae based on the feature extraction channel; the intensity features include average intensity, integral intensity, intensity standard deviation, maximum intensity and minimum intensity.

[0110] Intensity features are indicators describing the intensity of light signals from a target structure in a specific imaging channel. Their core value lies in reflecting the biochemical functional state of the structure, such as the chlorophyll content of chloroplasts or the DNA staining intensity of cell nuclei, by quantifying the strength and distribution of fluorescence or bright-field signals. Stress often disrupts the biochemical balance of cells, such as chlorophyll degradation and DNA damage, leading to significant changes in intensity features. For example, the calculation of intensity features requires relating the target mask to the original pixel grayscale values ​​of the corresponding channel. Specifically, average intensity is calculated by averaging the grayscale values ​​of all pixels within the target mask area to reflect the overall brightness level of the structure; for example, a decrease in the average intensity of chloroplasts may indicate chlorophyll decomposition. Integrated intensity is calculated by summing the grayscale values ​​of all pixels within the target mask area, taking into account the combined effect of area and average intensity; for example, the low average intensity of large cells may reflect differences in total signal quantity through integrated intensity. Intensity standard deviation is calculated by the dispersion of pixel grayscale values ​​within the target area, reflecting the uniformity of signal distribution; for example, when stress causes chloroplast aggregation, the intensity standard deviation will increase due to increased local brightness differences. Maximum / minimum intensity is obtained by extracting the extreme gray values ​​within the target area, reflecting the local areas where the signal is strongest or weakest. For example, a decrease in the maximum intensity of the DAPI channel may indicate damage to nuclear DNA.

[0111] The intensity characteristics of the aforementioned different channels are specific. For example, the intensity characteristics of chloroplast channels are directly related to photosynthetic capacity, and their average and integrated intensities usually decrease significantly under stress. The intensity characteristics of DAPI channels reflect the integrity of the cell nucleus, and DNA breaks may lead to an increase in their intensity standard deviation. By combining the intensity characteristics of multiple channels, the stress state of algae can be distinguished at the functional level.

[0112] S23. Based on the gray-level pixel spatial distribution of the algal cell image segmentation mask, a gray-level co-occurrence matrix is ​​obtained. The texture features of the cell nucleus, cell body and chloroplast are calculated based on the gray-level co-occurrence matrix. The texture features include the second moment of the angle, contrast, correlation, entropy and inverse difference moment.

[0113] Texture features are indicators describing the spatial distribution of pixel grayscale values ​​within a target area. They capture microscopic structural changes imperceptible to the human eye, such as the density of chloroplasts and the roughness of cell membranes. These changes are often sensitive indicators of early stress responses; for example, heavy metal stress can cause chloroplasts to change from a uniform distribution to local aggregation. For instance, texture feature calculation relies on the Gray-Level Co-occurrence Matrix (GLCM). This matrix quantifies the spatial correlation of grayscale distribution by statistically analyzing the frequency of grayscale value combinations of adjacent pixels within specific directions, such as 0°, 45°, 90°, and 135°, and distances of 1-3 pixels. Specifically, the second angular moment is calculated as the sum of the squares of all elements in the GLCM; a larger value indicates a more uniform grayscale distribution. For example, the second angular moment of normal algal chloroplasts is relatively high, but decreases after stress aggregation. Contrast is the weighted sum of the squares of the distances between elements in the GLCM and their diagonal. A larger value indicates a more significant difference in grayscale between adjacent pixels. For example, high contrast between the edge and interior of a cell membrane reflects membrane structural integrity; membrane damage under stress will reduce contrast. Correlation is the covariance relationship between elements in a GLCM and their row and column means. A value closer to 1 indicates a more linear change in grayscale values; for example, the cell wall texture of normal algae has high correlation, while correlation decreases when stress causes texture disorder. Entropy is the information entropy of elements in a GLCM; a larger value indicates a more disordered grayscale distribution. For example, during cell apoptosis, internal structural damage increases entropy. Inverse moment is the weighted sum of the reciprocals of the distances between elements and their diagonals in a GLCM; a larger value indicates a smoother grayscale change. For example, the cytoplasmic region of healthy algae has a high inverse moment, which decreases under stress due to the appearance of vacuoles.

[0114] S24. Combine the feature extraction channels of each pair into channel pairs, and calculate the pixel spatial overlap of each channel pair to obtain co-localization features; co-localization features include linear correlation coefficient, rank correlation coefficient, overlap coefficient and gray-scale contribution overlap coefficient.

[0115] Colocalization features are indicators describing the degree of spatial overlap of target structures in different imaging channels. They reveal the positional relationships of key intracellular structures, such as the nucleus and chloroplasts, and chloroplasts and the cell membrane. These relationships often change under stress; for example, chloroplasts, normally distributed around the nucleus in normal cells, may deviate from this pattern under stress. For instance, calculating colocalization features requires pixel-level comparison of target regions in each pair of channels. Specifically, the Pearson correlation coefficient is obtained by calculating the linear correlation between the grayscale values ​​of pixels within the target regions of two channels. Its value ranges from -1 to 1; a value close to 1 indicates a consistent grayscale trend. For example, a high Pearson coefficient between chloroplasts and the nucleus indicates a close correlation between their distributions. The Spearman correlation coefficient is obtained by calculating the rank correlation between the grayscale values ​​of pixels in two channels. It does not depend on a linear relationship and is suitable for scenarios with non-normal grayscale distributions, such as colocalization analysis of low-brightness areas. The overlap coefficient is calculated by the ratio of the overlapping area of ​​the target regions of two channels to the area of ​​the smaller region. A value close to 1 indicates that the smaller region is almost completely located within the larger region. For example, chloroplasts are usually completely located within the cell body, and the overlap coefficient is close to 1. The coefficient decreases when stress causes chloroplast overflow. The gray-scale contribution overlap coefficient (Manders) is calculated by comparing the proportion of the gray-scale sum of the overlapping region of channel A with channel B to the total gray-scale sum of channel A (M1) and the proportion of the gray-scale sum of the overlapping region of channel B with channel A to the total gray-scale sum of channel B (M2). This reflects the functional correlation between the two structures. A high M1 indicates that the signal of channel A mainly comes from the region overlapping with channel B, suggesting functional synergy.

[0116] Colocalization features can reflect the internal organizational order of cells from the perspective of spatial relationships. For example, light stress may cause chloroplasts to change from being distributed around the cell nucleus to being aggregated at the cell edge. This can be quantified by features such as a decrease in the Pearson coefficient and a decrease in the overlap coefficient between the nucleus and chloroplasts, providing a basis for understanding the impact of stress on cell structure and organization.

[0117] S25. Calculate the spatial positional relationships of similar targets in the cell nucleus, cell body, and chloroplast respectively to obtain the neighbor relationship characteristics; the neighbor relationship characteristics include nearest neighbor distance, number of neighbors, and average neighbor distance.

[0118] Neighbor relationship features are indicators describing the spatial distribution patterns of similar structures, such as multiple cells or chloroplasts. They capture population-level distribution changes, such as cell density and structural aggregation, and correlate them with stress-induced growth inhibition or aggregation behavior. For example, algae may stop dividing under heavy metal stress, leading to increased intercellular spacing. For instance, the calculation of neighbor relationship features is based on the target's center coordinates and achieved through distance analysis. Specifically, for each target, the Euclidean distance between its center and the nearest other target of the same type is calculated to obtain the nearest neighbor distance, reflecting the local sparsity. For example, when stress causes some cells to die, the nearest neighbor distance of surviving cells will increase. A fixed radius is set with the target center as the center, and the number of similar targets within this radius is counted to obtain the neighbor number, reflecting the local aggregation degree. For example, the number of chloroplast neighbors in normal algae is stable, but aggregation under stress will increase the number. The average neighbor distance is calculated by averaging the center distances between the target and all neighboring targets, reflecting the overall uniformity of the surrounding neighbor distribution. For example, when stress causes localized chloroplast accumulation, the average neighbor distance will decrease.

[0119] Neighbor relationship features can supplement the deficiencies of individual features from a group perspective. For example, the morphology of a single cell may not show obvious abnormalities when observed alone, but if the nearest neighbor distance is found to be significantly greater than that of other cells through neighbor relationship features, it suggests that the cell may be in an isolated state in the early stage of stress.

[0120] S26. Based on the overlap ratio of the algal cell image segmentation mask, the overlap relationship between the cell nucleus, cell body and chloroplast is calculated, and the overlap feature is obtained.

[0121] Overlap features are indicators describing the spatial containment of different structural levels, such as the nucleus and cell body, or chloroplasts and cell bodies. They verify the normal organizational relationship of structures; for example, the nucleus should be completely contained within the cell body, and chloroplasts should not spill out of the cell. Stress can disrupt this relationship, such as cell rupture leading to chloroplast spillage. For example, overlap features are calculated based on the area overlap ratio of different structural masks. Specifically, the Jaccard coefficient is the ratio of the overlapping area of ​​two masks to the union area of ​​the two masks. The closer the value is to 1, the higher the degree of overlap. For example, the Jaccard coefficient for the nucleus and cell body is typically between 0.3 and 0.5; a value below 0.1 may indicate abnormal nuclear displacement. The Sorensen-Dice coefficient is calculated as 2 × the ratio of the overlapping area of ​​the two masks to the sum of their areas. Compared to the Jaccard coefficient, it places more emphasis on the contribution of the overlapping area and is more sensitive to the overlap of small structures such as the nucleus. When part of the nucleus spills out of the cell body, the Dice coefficient decreases more significantly than the Jaccard coefficient.

[0122] S27. Perform full-image pixel statistics and calculate quantitative quality indicators for high-content algae images to obtain image quality characteristics.

[0123] Image quality features are indicators describing the overall imaging quality of the original image, serving as the basis for verifying the reliability of feature measurements. Low-quality images can lead to distortion in features such as size and intensity. Therefore, it is necessary to screen effective samples through quality features or exclude samples with abnormal quality during model analysis. The calculation of image quality features is based on the statistical characteristics of all pixels in the image. For example, sharpness is calculated using the Laplacian operator to measure the gray-level gradient intensity of image edges. A larger value indicates clearer image edges. However, blurry focus can reduce sharpness, potentially resulting in smaller extracted cell perimeters. Saturation pixel ratio is the proportion of pixels with a statistically maximum gray-level value (e.g., 255) in the entire image. A value that is too high, such as >5%, indicates overexposure, which can lead to distortion of intensity features. Power spectral density is calculated by performing a Fourier transform on the image to analyze the energy distribution at different spatial frequencies. A high proportion of high-frequency components indicates rich image details, while a high proportion of low-frequency components suggests image blur.

[0124] In one embodiment, the binary stress analysis model is obtained through the following method:

[0125] S31. Construct a training dataset using high-content images of unstressed algae and high-content images of known stressed algae; the training dataset includes feature vectors and corresponding stress labels.

[0126] Indicatively, unstressed algae samples, i.e., algae in a normal growth state, such as algae cultured in standard culture media under suitable light / temperature conditions, serve as negative controls. Their feature vectors reflect the baseline pattern of health status. Algae samples under known stress, i.e., algae subjected to specific stresses through artificial intervention, such as heavy metals, high salinity, ultraviolet radiation, or nutrient deficiency, serve as positive controls. These positive controls need to cover different stress types and intensities, such as low / medium / high concentrations of heavy metals, and different stress durations, such as 24 / 48 / 72 hours of treatment, to ensure the model can identify diverse stress response patterns. Similarly, feature vectors are extracted for each sample and labeled with the corresponding stress label; for example, unstressed samples are labeled as 0, and known stress samples are labeled as 1.

[0127] S32. Use supervised learning methods to train a binary classification model on the training dataset to obtain the hyperparameters, random seeds, and training version information of the binary stress analysis model.

[0128] Furthermore, supervised learning algorithms enable the model to autonomously learn the difference patterns in feature vectors between unstressed and stressed samples from the training data. Optimization algorithms then adjust the model parameters, allowing the model to output the correct stress label based on the new input feature vector. For example, based on the characteristics of algal stress features, binary classification algorithms could be random forests, support vector machines (SVMs), or gradient boosting trees.

[0129] Model performance largely depends on hyperparameters, such as the number of trees in a random forest and the kernel parameters of an SVM. Different combinations of hyperparameters can be iteratively tested on the validation set using grid search or Bayesian optimization to find the optimal parameter configuration that improves model performance, such as the F1 score. During training, parameters are initialized with a random seed to ensure reproducible training results. Simultaneously, key information during training, such as hyperparameter values, random seed, training time, and validation set performance, is recorded to create training version information.

[0130] S33. Use evaluation indicators to assess the discriminative ability of the binary stress analysis model, and serialize and save the model whose discriminative ability meets the expectations to obtain the binary stress analysis model.

[0131] For example, accuracy is the proportion of correctly classified samples out of the total samples, reflecting the overall classification performance; precision is the proportion of samples predicted as stress that are actually stress, avoiding misclassification of normal samples as stress; recall is the proportion of samples that are actually stress that are correctly predicted, avoiding omission of stress samples; F1 score is the harmonic mean of precision and recall, balancing the contradiction between the two; AUC-ROC is the area under the ROC curve, reflecting the model's overall ability to distinguish between two classes of samples.

[0132] Evaluation must be conducted on an independent test set. If all metrics reach the preset thresholds, the model's discriminative ability is deemed to meet expectations. For models that meet expectations, the model parameters and structure are saved as binary files using serialization technology. The saved content includes the model itself, hyperparameters, random seeds, training version information, etc.

[0133] In one embodiment, the method further includes:

[0134] S41. Transform the results of the stress analysis into textual conclusions.

[0135] The raw output of a binary stress analysis model is typically a machine-recognizable label plus a probability value. For example, label 1 represents being under stress, and a probability value of 0.92 represents the confidence level. For instance, the model's output labels are mapped to intuitive expressions, such as label 1 corresponding to a stressed state for the algal sample and label 0 corresponding to a non-stressed state, directly answering the core question of whether stress exists. The accompanying probability value indicates the reliability of the conclusion, such as a confidence level of 92%, letting the user know the certainty of the result. Furthermore, 1-3 features most strongly associated with stress are extracted from the feature vector, such as average chloroplast intensity and cell eccentricity. This can be determined by the feature importance output by the model, briefly explaining the abnormal behavior of this feature.

[0136] S42. When the stress analysis result indicates stress damage, generate an algal stress visualization map based on the feature vector; the algal stress visualization map is at least one of the following visualization forms: feature correlation network diagram, feature distribution heatmap, and principal component analysis (PCA) lithotripsy diagram.

[0137] Eigenvectors are abstract data containing hundreds of dimensions. Visualization can transform high-dimensional data into an intuitive graphical language, helping users uncover the patterns of stress effects, verify the rationality of model conclusions, and even discover new stress response characteristics. For example, feature correlation network diagrams are used to display the strength of associations between different features, revealing patterns of coordinated changes in features caused by stress (e.g., cell area and eccentricity often increase synchronously under stress, reflecting cell swelling and deformation). Eigenvectors can be used to calculate Pearson or Spearman correlation coefficients between all features. Nodes represent features, with node size corresponding to feature importance; for example, features with higher contributions to the model have larger nodes. Edges represent correlations, with edge thickness corresponding to the absolute value of the correlation coefficient; positive correlations are shown in red, and negative correlations in blue. Feature distribution heatmaps are used to display the numerical distribution of key features in different samples. They can compare the differences in feature patterns between the stress group and the control group in batches, quickly locate the features most significantly affected by stress, select 10-20 features that are most strongly associated with stress from the feature vector, and standardize the feature values ​​to eliminate differences in dimensions. The samples are represented by rows and arranged in groups according to control group - low stress - medium stress - high stress. The key features are represented by columns, and the feature value magnitude is represented by color gradient. Principal Component Analysis (PCA) charts are used to reduce high-dimensional feature vectors to 2-3 dimensional space, visually displaying the clustering trend of samples, such as whether the stress group and control group are clearly separated, and whether different stress intensities show gradient clustering. The scree plot is used to explain the basis for dimensionality reduction, and the PCA result plot is used to display the clustering effect. For example, in a PCA scree plot, the horizontal axis represents the principal components (PC1, PC2, PC3, etc.), and the vertical axis represents the variance contribution rate of the principal component, that is, the proportion of feature variation that the principal component can explain. Usually, the top 2-3 principal components with a cumulative variance contribution rate of 70%-80% are selected for subsequent analysis. For example, if the cumulative contribution rate of the top 3 principal components is 82%, then these 3 principal components are used to construct a 3D PCA result plot to ensure that most of the original information is retained after dimensionality reduction. Using the dimensionality-reduced principal components (e.g., PC1 as the horizontal axis and PC2 as the vertical axis) as coordinates, each point represents a sample, and different colors are used to mark sample groups. If the points of the same group of samples are clustered together, and the points of different groups of samples are clearly separated, it indicates that the feature can effectively distinguish the stress state.

[0138] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0139] Based on the same inventive concept, this application also provides a stress analysis system for high-content algal images to implement the stress analysis method for high-content algal images described above. The solution provided by this system is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the stress analysis system for high-content algal images provided below can be found in the limitations of the stress analysis method for high-content algal images described above, and will not be repeated here.

[0140] In one exemplary embodiment, such as Figure 4 As shown, a stress analysis system for high-content algal images is provided, comprising:

[0141] Data module 401 is used to acquire a set of high-content algae images and to perform quality screening and control processing on the set of high-content algae images to obtain high-content algae images that have passed quality control.

[0142] Image segmentation module 402 is used to input high-content algae images that have passed quality control into a pre-trained algae-specific segmentation model based on cellpose-sam to obtain an algae cell image segmentation mask.

[0143] The feature extraction module 403 is used to extract features from algal high-content images based on algal cell image segmentation masks to obtain feature vectors oriented towards algae and their internal structures.

[0144] The stress analysis module 404 is used to input feature vectors into a pre-trained binary classification stress analysis model to obtain stress analysis results.

[0145] In one embodiment, the data module 401 is further configured to:

[0146] Obtain the file metadata and pixel data of each original high-content algae image in the high-content algae image set;

[0147] Based on the preset expected image size range, file integrity is verified according to file metadata, and the original high-content algae image with a complete file is determined as the expected image;

[0148] Pixel statistics are performed on the pixel data of the desired image, and the desired image that meets the preset zero pixel ratio and saturated pixel ratio is identified as a high-content algae image that has passed quality control.

[0149] In one embodiment, the image segmentation module 402 is further configured to:

[0150] Standardized high-content algal images that have passed quality control are preprocessed to obtain standardized high-content algal images;

[0151] Probable algal cell regions are identified from standardized high-content algal images to obtain candidate algal cell regions.

[0152] Based on preset algae-specific parameters, the complete outline of algae cells is identified within the candidate region bounding box of algae cells to obtain the cell region mask; the algae-specific parameters include flow threshold and cell probability threshold;

[0153] The region of adherent cells within the cell region mask is determined based on the average area threshold of algal cells.

[0154] The watershed separation algorithm is used to classify cells in the region of adherent cells, and the separated cells are subjected to morphological closing operation to obtain an algal cell image segmentation mask.

[0155] In one embodiment, the feature extraction module 403 is further configured to:

[0156] Based on a preset filename regularization expression, the corresponding feature extraction channels are determined according to the filename type of the high-content algae image; the feature extraction channels include the bright field channel, chloroplast channel, DAPI channel, and cell mask channel;

[0157] The high-content algae images were sequentially processed by grayscale conversion, intensity standardization, target feature enhancement, and Gaussian filtering for noise reduction to obtain enhanced high-content algae images.

[0158] Based on algal cell image segmentation masks, a multi-level recognition strategy is used to identify cell nuclei, cell bodies, and chloroplasts in enhanced high-content algal images; the multi-level recognition strategy is determined by feature extraction channels.

[0159] Feature measurements were performed on each cell nucleus, cell body, and chloroplast in the feature extraction channel to obtain feature data; the feature data included size and shape features, intensity features, texture features, colocalization features, neighbor relationship features, overlap features, and image quality features;

[0160] The feature data is cleaned, and the cell nucleus-cell body association and chloroplast-cell body association are established based on the cleaned feature data to obtain feature vectors.

[0161] In one embodiment, the feature extraction module 403 is further configured to:

[0162] Based on the pixel spatial distribution of the algal cell image segmentation mask, the size and shape of each cell nucleus, cell body, and chloroplast are calculated to obtain size and shape features. The size and shape features include area, perimeter, eccentricity, compactness, extensibility, axial length, Euler number, compactness, and center coordinates.

[0163] The intensity features of each cell nucleus, cell body, and chloroplast are calculated based on the pixel gray values ​​corresponding to the enhanced high-content images of algae using feature extraction channels; the intensity features include average intensity, integral intensity, intensity standard deviation, maximum intensity, and minimum intensity;

[0164] The gray-level co-occurrence matrix is ​​obtained based on the gray-level pixel spatial distribution of the algal cell image segmentation mask. The texture features of the cell nucleus, cell body, and chloroplast are calculated based on the gray-level co-occurrence matrix. The texture features include the second moment of the angle, contrast, correlation, entropy, and inverse difference moment.

[0165] The feature extraction channels are paired to form channel pairs, and the pixel spatial overlap of each channel pair is calculated to obtain co-localization features. Co-localization features include linear correlation coefficient, rank correlation coefficient, overlap coefficient, and gray-level contribution overlap coefficient.

[0166] The spatial relationships of similar targets in the cell nucleus, cell body, and chloroplast are calculated separately to obtain the neighbor relationship characteristics. The neighbor relationship characteristics include nearest neighbor distance, number of neighbors, and average neighbor distance.

[0167] The overlap relationship between the cell nucleus, cell body and chloroplast is calculated based on the overlap ratio of the algal cell image segmentation mask, and the overlap feature is obtained.

[0168] Image quality features are obtained by performing full-image pixel statistics and calculating quantitative quality indicators for high-content algae images.

[0169] In one embodiment, a model training module is also included, for:

[0170] A training dataset was constructed using high-content images of unstressed algae and high-content images of algae under known stress. The training dataset includes feature vectors and corresponding stress labels.

[0171] Supervised learning methods were used to train a binary classification model on the training dataset to obtain the hyperparameters, random seeds, and training version information of the binary stress analysis model.

[0172] The discriminative ability of the binary stress analysis model is evaluated using assessment indicators, and the models whose discriminative ability meets the expectations are serialized and saved to obtain the binary stress analysis model.

[0173] In one embodiment, an interpretable module is also included for:

[0174] Transform the results of the coercion analysis into textual conclusions;

[0175] When the stress analysis result indicates stress damage, an algal stress visualization map is generated based on the feature vectors. The algal stress visualization map is at least one of the following visualization forms: feature correlation network map, feature distribution heatmap, and principal component analysis (PCA) lithotripsy map.

[0176] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above method embodiments.

[0177] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0178] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0179] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A stress analysis method for high-content images of algae, characterized in that, The method includes: A high-content image set of algae is obtained, and the high-content image set of algae is subjected to quality screening and control processing to obtain high-content images of algae that have passed quality control. The high-content algae images that have passed quality control are input into a pre-trained algae-specific segmentation model based on cellpose-sam to obtain an algae cell image segmentation mask. Based on a preset filename regularization expression, the corresponding feature extraction channels are determined according to the filename type corresponding to the high-content algae image; the feature extraction channels include bright field channel, chloroplast channel, DAPI channel and cell mask channel; The high-content algae image is sequentially subjected to grayscale conversion, intensity standardization, target feature enhancement, and Gaussian filtering for noise reduction to obtain an enhanced high-content algae image. Based on the algal cell image segmentation mask, a multi-level recognition strategy is used to identify the cell nucleus, cell body, and chloroplast in the enhanced algal high-content image; the multi-level recognition strategy is determined by the feature extraction channel. Feature measurements are performed on each of the cell nuclei, cell bodies, and chloroplasts in the feature extraction channel to obtain feature data; the feature data includes size and shape features, intensity features, texture features, colocalization features, neighbor relationship features, overlap features, and image quality features; The feature data is cleaned, and the cell nucleus-cell body association and chloroplast-cell body association are established based on the cleaned feature data to obtain feature vectors; The feature vectors are input into a pre-trained binary classification stress analysis model to obtain stress analysis results; The algae-specific segmentation model based on cellpose-sam obtains the algae cell image segmentation mask through the following steps: The high-content algae images are subjected to standardization preprocessing to obtain standardized high-content algae images; The standardized high-content algae image is subjected to possible algal cell region identification to obtain candidate algal cell region boxes; Based on preset algae-specific parameters, the complete outline of algae cells is identified within the candidate region box of algae cells to obtain a cell region mask; the algae-specific parameters include a flow threshold and a cell probability threshold; The region of adherent cells within the cell region mask is determined based on the average area threshold of algal cells. The watershed separation algorithm is used to classify the cells in the adherent cell region, and morphological closing operation is performed on the separated cells to obtain the algal cell image segmentation mask.

2. The method according to claim 1, characterized in that, The process of performing quality screening and control on the high-content algae image set to obtain high-content algae images that pass quality control includes: Obtain the file metadata and pixel data of each original high-content algae image in the high-content algae image set; Based on a preset expected image size range, file integrity is verified according to the file metadata, and the original algae high-content image whose verification result is that the file is complete is determined as the expected image; Pixel statistics are performed on the pixel data of the desired image, and the desired image whose pixel statistics result meets the preset zero pixel ratio and saturated pixel ratio is determined as the algae high-content image that has passed quality control.

3. The method according to claim 1, characterized in that, The step of performing feature measurements on each of the cell nuclei, cell bodies, and chloroplasts in the feature extraction channel to obtain feature data includes: Based on the pixel spatial distribution of the algal cell image segmentation mask, the size and shape of each cell nucleus, cell body, and chloroplast are calculated to obtain the size and shape features; the size and shape features include area, perimeter, eccentricity, compactness, extensibility, axial length, Euler number, compactness, and center coordinates; The intensity features of each cell nucleus, cell body, and chloroplast are calculated based on the pixel grayscale values ​​corresponding to the enhanced high-content image of algae obtained from the feature extraction channels; the intensity features include average intensity, integral intensity, intensity standard deviation, maximum intensity, and minimum intensity; A gray-level co-occurrence matrix is ​​obtained based on the gray-level pixel spatial distribution of the algal cell image segmentation mask. The texture features of the cell nucleus, cell body, and chloroplast are calculated based on the gray-level co-occurrence matrix. The texture features include angular second moment, contrast, correlation, entropy, and inverse difference moment. The feature extraction channels are paired to form channel pairs, and the pixel spatial overlap of each channel pair is calculated to obtain the co-localization feature; the co-localization feature includes linear correlation coefficient, rank correlation coefficient, overlap coefficient and grayscale contribution overlap coefficient. The spatial positional relationships of similar targets in the cell nucleus, cell body, and chloroplast are calculated respectively to obtain the neighbor relationship features; the neighbor relationship features include nearest neighbor distance, number of neighbors, and average neighbor distance. The overlap relationship between the cell nucleus, the cell body, and the chloroplast is calculated based on the overlap ratio of the algal cell image segmentation mask, and the overlap feature is obtained. The high-content algae image is subjected to full-image pixel statistics and quantitative quality index is calculated to obtain the image quality characteristics.

4. The method according to claim 1, characterized in that, The binary stress analysis model was obtained through the following method: A training dataset was constructed using high-content images of unstressed algae and high-content images of algae under known stress; the training dataset includes feature vectors and corresponding stress labels. A supervised learning method is used to train a binary classification model on the training dataset to obtain the hyperparameters, random seeds, and training version information of the binary stress analysis model. The discriminative ability of the binary stress analysis model is evaluated using evaluation metrics, and models whose discriminative ability meets expectations are serialized and saved to obtain the binary stress analysis model.

5. The method according to claim 1, characterized in that, The method further includes: The stress analysis results are then converted into textual conclusions. When the stress analysis result indicates stress damage, an algal stress visualization map is generated based on the feature vector; the algal stress visualization map is at least one of the following visualization forms: feature correlation network diagram, feature distribution heatmap, and principal component analysis (PCA) lithotripsy diagram.

6. A stress analysis system for high-content images of algae, characterized in that, The system includes: The data module is used to acquire a set of high-content algae images and to perform quality screening and control processing on the set of high-content algae images to obtain high-content algae images that have passed quality control. The image segmentation module is used to input the high-content algae image that has passed quality control into a pre-trained algae-specific segmentation model based on cellpose-sam to obtain an algae cell image segmentation mask. The feature extraction module is used to determine the corresponding feature extraction channel based on the file name type corresponding to the high-content algae image, according to a preset file name regularization expression; the feature extraction channel includes a bright field channel, a chloroplast channel, a DAPI channel, and a cell mask channel; The feature extraction module is also used to sequentially perform grayscale conversion, intensity standardization, target feature enhancement, and Gaussian filtering noise reduction on the algae high-content image to obtain an enhanced algae high-content image; The feature extraction module is further configured to identify the cell nucleus, cell body, and chloroplast in the enhanced algal high-content image based on the algal cell image segmentation mask and employ a multi-level recognition strategy; the multi-level recognition strategy is determined by the feature extraction channel. The feature extraction module is also used to perform feature measurements on each of the cell nuclei, cell bodies, and chloroplasts on the feature extraction channel to obtain feature data; the feature data includes size and shape features, intensity features, texture features, colocalization features, neighbor relationship features, overlap features, and image quality features; The feature extraction module is also used to clean the feature data and establish cell nucleus-cell body association and chloroplast-cell body association based on the cleaned feature data to obtain feature vectors; The stress analysis module is used to input the feature vector into a pre-trained binary classification stress analysis model to obtain stress analysis results; The algae-specific segmentation model based on cellpose-sam obtains the algae cell image segmentation mask through the following steps: The high-content algae images are subjected to standardization preprocessing to obtain standardized high-content algae images; The standardized high-content algae image is subjected to possible algal cell region identification to obtain candidate algal cell region boxes; Based on preset algae-specific parameters, the complete outline of algae cells is identified within the candidate region box of algae cells to obtain a cell region mask; the algae-specific parameters include a flow threshold and a cell probability threshold; The region of adherent cells within the cell region mask is determined based on the average area threshold of algal cells. The watershed separation algorithm is used to classify the cells in the adherent cell region, and morphological closing operation is performed on the separated cells to obtain the algal cell image segmentation mask.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.