Virus particle identification method based on negative staining electron microscope image

Through deep learning and image segmentation technology, CNN is used to automatically identify and classify negatively stained electron microscope images, solving the problem of distinguishing hollow and solid AAV particles, improving the analysis quality and speed, and promoting the development of gene therapy.

CN120220143APending Publication Date: 2025-06-27LOTUSLAKE BIOMEDICAL TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311809273.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to accurately distinguish between hollow and solid AAV particles in negative stained electron microscope images, resulting in challenges in removing and controlling hollow capsids during AAV synthesis.

Method used

Deep learning and advanced image segmentation technology are adopted to automatically identify and classify negatively stained electron microscope images through convolutional neural network (CNN), so as to achieve accurate identification of hollow and solid virus particles.

Benefits of technology

It improves the quality and speed of AAV particle analysis, helps optimize the production quality control of gene therapy vectors, and shortens the development time of gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004631122170000101
    Figure BDA0004631122170000101
  • Figure BDA0004631122170000111
    Figure BDA0004631122170000111
  • Figure HDA0004631122180000011
    Figure HDA0004631122180000011
Patent Text Reader

Abstract

The invention provides a virus particle identification method based on a negative staining electron microscope image. The deep learning and image segmentation technology is adopted to automatically identify virus particles and classify hollow and solid virus particles in the negative staining electron microscope image, so that the real heart rate and the virus particle concentration of the virus are obtained, and the requirement on the purity of a sample is low. Therefore, the production quality control process of the virus vector can be improved, and the development and optimization of gene therapy can be accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biotechnology, and particularly relates to a method for identifying virus particles based on negative staining electron microscope images. Background Art

[0002] Virus-like particles (VLP) and replication-defective viruses such as adeno-associated virus (AAV) cannot proliferate and are widely used as gene delivery vectors in the pharmaceutical industry. In particular, AAV has a wide host cell range, a long in vivo expression time, a fast response time, and a high expression efficiency, and does not cause human diseases, making it the most promising viral gene therapy vector. However, during the synthesis of AAV, natural by-products will appear, namely empty capsids without packaged DNA, and capsids with mis-packaged fragmented ssDNA (referred to as partially complete) or host DNA. It is not yet fully understood how the presence of empty capsids affects the therapeutic properties of rAAV, but removing them can significantly increase transgene expression. Some studies have shown that empty capsids inhibit the transduction of target cells by exacerbating the immune response to rAAV and competing for cell binding sites. Therefore, it is necessary to control the content of empty capsids and partially complete capsids in the product. Currently, several methods have been developed for detecting the empty capsid rate, including the most common anion exchange chromatography (AEX), optical density (OD), size exclusion chromatography multi-angle light scattering (SEC-MALS), enzyme-linked immunosorbent assay (ELISA) combined with quantitative polymerase chain reaction (qPCR), and analytical ultracentrifugation (AUC).

[0003] Negative staining electron microscopy is a commonly used technique for observing AAV particles. It can provide extremely high resolution, enabling researchers to see the detailed structure of virus particles. However, differentiating between hollow (unloaded with genetic material) and solid (loaded with genetic material) AAV particles in these images is a challenge. This differentiation is crucial for understanding the loading efficiency and delivery ability of the virus. Hollow and solid particles may look very similar in electron microscope images, and precise image analysis is required to distinguish them. This involves not only identifying the shape and size of virus particles but also the subtle differences in their internal structures. Traditional manual analysis methods are not only time-consuming but also error-prone.

[0004] Artificial intelligence (AI), especially machine learning and deep learning, has made remarkable progress in the field of image analysis. These technologies can process and analyze large amounts of image data, extract complex features, and thus play a crucial role in many fields such as medical imaging, satellite image processing, and microscopic image analysis. In machine learning, algorithms learn to identify specific patterns and features from a large number of sample data. These patterns are then used for the analysis and prediction of new data. Deep learning, especially convolutional neural networks (CNNs), is particularly effective for image data. CNNs can achieve accurate image recognition and classification by learning hierarchical features in images, from simple edges and textures to more complex shapes and objects. Therefore, the development of efficient and accurate automated technologies, such as AI-based image analysis methods, is crucial for improving the quality and speed of AAV particle analysis. Summary of the Invention

[0005] To overcome the problems existing in differentiating hollow and solid virus particles by conventional negative staining electron microscopy, the present invention provides a method for identifying virus particles based on negative staining electron microscopy images. Through deep learning and advanced image segmentation techniques, it can automatically identify virus particles and classify hollow and solid virus particles in negative staining electron microscopy images, which can improve the production quality control process of virus vectors and contribute to accelerating the development and optimization of gene therapy.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a method for identifying and classifying virus particles based on negative staining electron microscopy images, the method comprising:

[0008] 1) Obtain a crude virus solution;

[0009] 2) Perform negative staining on the virus particles in the crude virus solution, and use a room temperature transmission electron microscope to image the negatively stained virus particle sample to obtain a target image;

[0010] 3) Virus particle detection:

[0011] a. A pre-created convolutional neural network reads the original image of the virus negative staining electron microscope and performs preprocessing;

[0012] b. Based on the pre-created convolutional neural network, perform field-of-view image cutting on the original image, save each cut field of view separately, and record the original position;

[0013] c. Based on the pre-created convolutional neural network, process each field-of-view image separately, perform target detection on the virus particles in each field of view, and output the position of the virus in the current field-of-view image;

[0014] d. Restore the position of the detected virus in the original image and visualize it, and count the total number of viruses.

[0015] 4) Virus particle classification:

[0016] a. According to the original image and the virus detection position, use image processing technology to finely segment each virus particle one by one to generate a virus particle mask; according to the pixel information of the original virus image and the mask, use a pre-created convolutional neural network to perform binary classification on whether the virus is hollow.

[0017] b. According to the classification results and the total number of viruses, calculate the hollow rate and solid rate of virus particles in the negative-stained electron microscopy image.

[0018] The target image is recognized by a pre-created convolutional neural network to obtain the recognition result of a single virus particle, and the recognition result includes one of the virus particle sample being solid, the virus particle sample being hollow, or the virus particle sample being partially solid; the pre-created convolutional neural network is trained based on a negative-stained electron microscopy dataset of virus particle samples.

[0019] The pre-created convolutional neural network includes a classification layer and a segmentation layer. The classification layer consists of a convolutional layer, a pooling layer, and a fully connected layer. The segmentation layer is based on image segmentation technology; this combination enables the neural network to more accurately identify and analyze virus particles in negative-stained electron microscopy images. Among them, the convolutional layer is used to extract features in the image, the pooling layer is used to reduce the spatial size of the features, and the fully connected layer is used for classification decision-making; the segmentation layer specially designed for the CNN architecture of the present invention is a U-Net or a similar structure for image segmentation tasks. The role of image segmentation is to separate virus particles in the image from the background for further analysis, which is crucial for accurately identifying and classifying hollow and solid particles because it provides clearer particle contours and internal structure information; this structure can effectively extract fine spatial features from the image and is an ideal choice for processing high-resolution electron microscopy images.

[0020] Preferably, steps 1) and 2) stain the virus in the crude virus liquid to obtain a target image, including:

[0021] (1) Obtain the crude virus liquid: Transfect cells using an expression vector, where the expression vector carries a nucleic acid molecule encoding a virus capsid protein and a nucleic acid molecule encoding a virus genome, and the density of the cells is sufficient to ensure that the cell density corresponding to the dilution liquid does not exceed 10 6 / mL; culture the transfected cells; lyse the cells to obtain a cell lysate containing the virus; centrifuge at high speed to remove cell debris to obtain the crude virus liquid.

[0022] (2) Gradient dilution of the crude virus solution: Dilute the crude virus solution obtained in step (1) 5-fold with diluent ① to obtain sample A, and then dilute sample A 10-fold with diluent ② to obtain the optimized crude virus solution sample B; Diluent ① is PBS, and diluent ② is PBS containing cell lysate supernatant at a concentration of 1×10 6 cells / mL, that is, resuspend 1×10 6 cells in 1 mL of PBS and freeze-thaw three times repeatedly, then centrifuge at 12,000 rpm for 5 min and take the supernatant; The virus titer of the gradient-diluted crude virus solution is not less than 3×10 10 vp / mL;

[0023] (3) Negative staining of the virus: Suspend the copper grid with a supporting film on the droplet of sample B (with the side with the supporting film facing up) and wait for 1 minute, then invert the copper grid on the uranyl acetate staining solution for 10 seconds, blot the excess staining solution with filter paper, repeat this operation twice, but the third inversion time is 30 seconds, and finally dry the copper grid and use a Thermo fisher Talos120C electron microscope for detection to obtain a negative-stained electron microscope image.

[0024] Preferably, the pre-created convolutional neural network in step (3) is created based on a dataset of negative-stained electron microscope images of virus particle samples. Obtaining a dataset of negative-stained electron microscope (TEM) images of virus particle samples involves the following steps:

[0025] (1) Sample collection and preparation:

[0026] First, collect virus particle samples from laboratories or collaborative research institutions. These samples are from different virus types, including novel viruses under study or known virus strains; The samples need to be purified and concentrated to ensure the visibility of virus particles under the electron microscope;

[0027] (2) Negative staining and microscopy imaging:

[0028] Use heavy metal salts (such as phosphotungstic acid or uranyl acetate) to negatively stain the virus samples to increase the image contrast under the electron microscope; Then, image the stained samples under a transmission electron microscope, and use different magnification and imaging parameters as needed to capture the detailed structure of virus particles;

[0029] (3) Image collection and screening:

[0030] Collect all the images generated during the imaging process, which may include thousands of images depending on the sample size and imaging conditions; perform a preliminary screening of the collected images to eliminate images of poor quality, such as those with inaccurate focus, overexposure or underexposure, and those that do not clearly show virus particles;

[0031] (4) Data annotation and labeling:

[0032] Experts in the fields of virology and electron microscopy imaging annotate and label the screened images; the labeling includes identifying virus particles in the images and marking their positions, types, possible states (such as hollow or solid), etc.;

[0033] (5) Data collation and formatting:

[0034] Collate the annotated and labeled images into a structured dataset; typically includes converting the images to a specific format (such as JPEG or PNG) and resolution, and creating a metadata file containing the annotation information; for the convenience of training and validating machine learning models, the dataset is further divided into a training set, a validation set, and a test set;

[0035] (6) Compliance and ethical considerations:

[0036] When collecting and using virus samples and their image data, relevant legal and ethical regulations must be complied with, especially when dealing with potentially dangerous or restricted virus materials, to ensure that the use of data complies with the requirements of data privacy and intellectual property rights.

[0037] Further preferably, the test dataset and the training dataset are different in statistical characteristics in order to test the generalization ability of the test algorithm under different conditions. To improve the robustness of the algorithm, the dataset should also include various challenging images, such as those with high particle density or poor image quality.

[0038] Preferably, the method for establishing the pre-created convolutional neural network described in step 3) includes the following steps:

[0039] (1) Use the above training dataset for training to obtain an initial convolutional neural network;

[0040] (2) Based on the above test set, adjust the model parameters of the initial convolutional neural network to obtain a convolutional neural network for virus particle recognition and classification;

[0041] Further preferably, the training is performed using the above training dataset. During the training process, by continuously providing images and adjusting the model parameters, the network learns to recognize features related to hollow and solid AAV particles, and the features include the shape, size, edge characteristics, and subtle differences in the internal structure of the particles. To improve the generalization ability of the model, a diverse dataset should be used for training, including images under different conditions, such as different virus concentrations, different staining techniques, etc.

[0042] Further preferably, the training is performed using the above training dataset. The training process involves using these labeled images to adjust the weights of the CNN. Usually, a supervised learning method is adopted, in which the convolutional neural network attempts to predict the correct label for each image and adjusts its parameters by comparing the prediction and the actual label. During the training process, techniques such as cross-validation are used to evaluate the performance of the model and prevent overfitting. In addition, various techniques can be used to optimize the training, such as adjusting the learning rate, using early stopping, or adjusting the batch size. After the training is completed, the performance of the model needs to be evaluated on an independent test set to ensure its good generalization ability for new data.

[0043] Preferably, the preprocessing of the original electron microscopy images of virus negative staining in step 3)a includes the following steps:

[0044] (1) Image size adjustment: Since the CNN requires a fixed-size input, the image usually needs to be scaled to the size required by the network, which usually involves interpolation methods such as bilinear or bicubic interpolation to retain the important features of the image.

[0045] (2) Normalization and standardization: Normalization involves scaling the image pixel values to a specific range (such as 0 to 1 or -1 to 1), which helps the network learn more effectively. Standardization usually refers to subtracting the mean of the dataset from the image pixel values and dividing by the standard deviation, so that the entire dataset has zero mean and unit variance. This can improve the stability and convergence speed of the training.

[0046] (3) Contrast adjustment: In electron microscopy images, contrast adjustment can help highlight the structure of virus particles, especially in negative staining samples, which is achieved through histogram equalization or other contrast enhancement techniques.

[0047] (4) Noise reduction and filtering: Electron microscopy images may contain noise, and noise reduction techniques need to be applied to improve the image quality. Commonly used methods include Gaussian filtering, median filtering, or more advanced noise reduction techniques such as non-local mean filtering.

[0048] (5) Data augmentation: Data augmentation is an important method to improve the generalization ability of the model, including rotation, flipping, scaling, cropping, or color transformation of the image, which helps the network learn to recognize virus particles from different angles and conditions.

[0049] (6) Edge detection and enhancement: Emphasizing the edges in the image so that the outlines of virus particles can be more clearly identified, which is achieved by applying edge detection algorithms (such as Canny edge detection).

[0050] Preferably, the image cutting described in step 3) b refers to dividing the large-size original image into multiple smaller sub-images of fixed size so that the image segments meet the input requirements of the CNN model; the image cutting method includes one or more of the following (1) to (4):

[0051] (1) Fixed-size window sliding: This is a common image cutting technique that involves sliding a fixed-size window across the entire original image and cropping a sub-image at each position. Usually, the size of this window matches the size of the CNN input layer. For example, if the CNN requires an input of 224x224 pixels, the window size should also be set to 224x224 pixels. The window can be moved across the image at a certain stride, which determines the degree of overlap between sub-images. A smaller stride increases the overlapping area and may result in more redundant information, but can sometimes improve detection accuracy.

[0052] (2) Adaptive image cutting: The size and position of the cutting window are dynamically adjusted according to the image content. For example, an image processing algorithm can be used to first identify areas with dense virus particles, and then the optimal position of the cutting window can be determined based on the location and size of these areas.

[0053] (3) Cutting using preprocessing information:

[0054] Before cutting, preprocessing steps (such as contrast enhancement and edge detection) can be applied to more clearly identify the areas that need to be analyzed. This preprocessing information can be used to guide the image cutting process and ensure that the cut sub-image contains valuable fields of view, thereby improving the efficiency and accuracy of subsequent analysis.

[0055] (4) Overlapping and non-overlapping cutting: When cutting, you can choose to create overlapping or non-overlapping sub-images. Overlapping cutting means that there will be a certain degree of overlap between adjacent sub-images, which can ensure that important information in the edge area is not missed; non-overlapping cutting simply divides the image into strictly separated blocks, which is suitable for situations where image features are evenly distributed.

[0056] U-Net is a convolutional neural network (CNN) architecture designed for image segmentation tasks. It was originally developed to address the challenges of medical image segmentation. The core feature of U-Net is its unique structure, which makes it very suitable for processing images with fine structures. The following are some key features of the U-Net structure:

[0057] (1) Symmetrical "U" - shaped structure:

[0058] The architecture of U - Net is in a "U" shape and is divided into two parts: the contracting (downsampling) path and the expanding (upsampling) path. The contracting path is similar to a traditional CNN and includes multiple convolutional layers and pooling layers for capturing context information in the image; the expanding path includes multiple upsampling and convolutional operations for precisely locating and segmenting specific regions of the image;

[0059] (2) Skip Connections: A key feature of U - Net is the skip connections between the contracting and expanding paths. These connections combine the feature maps on the contracting path with the feature maps on the expanding path; skip connections help to recover the detailed information of the image, which is especially important for precise segmentation;

[0060] (3) Fine - grained feature retention: Since U - Net uses skip connections in the expanding path, it can combine high - resolution feature maps with upsampled feature maps, thus retaining more detailed and positional information. This makes U - Net particularly suitable for image segmentation tasks that require precise edge localization;

[0061] (4) Flexibility and adaptability: The U - Net architecture is highly flexible and can be adjusted according to different task requirements. For example, the depth of the network, the number of filters in the convolutional layers, etc. can be changed;

[0062] In addition, although originally designed for medical images, U - Net has also been successfully applied to various other types of image segmentation tasks. Besides U - Net, there are some similar structures such as U - Net ++ or U - Net 3 +, which are improvements and extensions based on the original U - Net, aiming to further improve the segmentation accuracy and network performance. These variants usually include more complex skip connection strategies or more advanced upsampling techniques to improve feature fusion and information flow.

[0063] Preferably, the convolutional neural network needs to carefully adjust the parameters of the segmentation layer and the classification layer to ensure that the final model can balance the requirements of these two tasks and maximize the overall performance.

[0064] Compared with the prior art, the beneficial effects of the present invention are:

[0065] The present invention can directly detect the simply - processed upstream fermentation crude liquid. For the optimization of the upstream process, the detection time is short and the feedback speed is fast, greatly shortening the time of the process optimization process.

[0066] The sampling of the present invention is carried out by dipping a sample with a copper mesh of a tiny area to obtain a sample film attached to the copper mesh. The sample demand is extremely small and multiple samplings can be carried out, so the requirement for sample concentration is relatively low. Also, since the effective particle counting is carried out after image discrimination, the requirement for sample purity is relatively low.

[0067] The solution of the present invention integrates advanced image segmentation technology into the CNN architecture, which can simultaneously process image segmentation and the classification task of virus particles in a single workflow, thereby improving the processing speed and accuracy. This algorithm can more effectively process the complex features in electron microscope images and provide strong quantitative detection technical support for the accurate identification and classification of virus particles. The implementation of this method is of great significance for improving the accuracy and efficiency of the quality control of gene therapy vectors. Brief Description of the Drawings

[0068] Figure 1 The algorithm scores each particle, and the result shown in color coding demonstrates the performance of the algorithm on the actual image;

[0069] Figure 2 The result of the empty shell rate after affinity chromatography purification of virus particles in the crude feed liquid in Example 2 detected by HPLC method. Detailed Embodiments

[0070] The technical solution of the present invention will be further described in detail below in conjunction with specific embodiments. It should be understood that the following embodiments are only for exemplarily illustrating and explaining the present invention and should not be construed as limiting the protection scope of the present invention. All technologies implemented based on the above content of the present invention are covered within the scope of protection intended by the present invention.

[0071] Unless otherwise specified, the raw materials and reagents used in the following embodiments are all commercially available products or can be prepared by known methods.

[0072] Unless otherwise defined or clearly indicated by the context, all technical and scientific terms in this disclosure have the same meaning as commonly understood by those of ordinary skill in the art to which this disclosure belongs.

[0073] Example 1. Detection of Crude Feed Liquid with Known Virus Titer

[0074] 1) AAV virus packaging process: The cells were counted as 5*10^6 / mL, and the AAV three-plasmid system was transfected. After 72 h, the cells were collected, centrifuged at 1000 rpm for 5 min, and then the supernatant was removed. The cell precipitate was added with PBS and repeatedly frozen and thawed 3 times, and then centrifuged at 12000 rpm for 5 min to remove cell debris. The obtained crude virus feed liquid was detected by ELISA, and the number of virus particles was 1.6*10^12 vp / mL.

[0075] 2) Electron microscopy detection: Suspend the copper grid with a supporting film on the sample droplet (with the side having the supporting film facing up) for 1 minute, then invert the copper grid on the uranyl acetate staining solution for 10 seconds, blot the excess staining solution with filter paper, repeat this operation twice, but the inversion time for the third time is 30 seconds, and finally dry the copper grid and detect it using a Thermo fisher Talos120C electron microscope.

[0076] 3) Detection result: There are too many cell impurities to distinguish virus particles.

[0077] Example 2. Optimize the detection method of the crude virus solution

[0078] Gradiently dilute the AAV virus packaged in Example 1, specifically:

[0079] Diluent ①: PBS

[0080] Diluent ②: The cells in Example 1 with a density of 10^6 are resuspended with 1 mL of PBS and repeatedly frozen and thawed three times, and the supernatant is taken after centrifugation at 12000 rpm for 5 min.

[0081] Sample A: Dilute 5-fold with Diluent ① to make the cell density 10^6 / mL and the virus titer 3.2*10^11 vp / mL.

[0082] Sample B: Dilute Sample A 10-fold with Diluent ② to make the cell density 10^6 / mL and the virus titer 3.2*10^10 vp / mL.

[0083] Sample C: Dilute Sample B 2-fold with Diluent ② to make the cell density 10^6 / mL and the virus titer 1.6*10^10 vp / mL.

[0084] Stain the virus particles of Samples A, B, and C by negative staining according to the method of Example 1 and detect them using a Thermo fisher Talos120C electron microscope.

[0085] Detection result: Viruses can be detected in Samples A and B, and virus particles cannot be distinguished in Sample C.

[0086] Example 3 Introduce AI to calculate the virus empty capsid rate and titer in the crude solution

[0087] 1. AI module design:

[0088] The basic process of the module includes virus particle detection and virus particle classification:

[0089] 1.1 Detection module:

[0090] 1.1.1 Read the original AAV virus negative staining electron microscopy image and perform preprocessing;

[0091] 1.1.2 Crop the field-of-view images of the original image to a fixed size and record the original positions.

[0092] 1.1.3 Based on image processing techniques and a pre-created deep learning network, automatically detect virus particles in each field of view and output the corresponding positions of the viruses.

[0093] 1.1.4 Restore and visualize the positions of the detected viruses in the original image and count the total number of viruses.

[0094] 1.2 Identification Module / Classification Module

[0095] 1.2.1 According to the pixel information and masks of the original virus images, use image processing and a pre-created deep learning network to perform binary classification on whether the virus is hollow.

[0096] 1.2.2 Calculate the hollow shell rate of virus particles in the negative-stained electron microscopy images based on the classification results and the total number of viruses.

[0097] The following are the application results of the pre-created convolutional neural network of the present invention on actual electron microscopy images:

[0098] In the experiment, the algorithm was applied to a series of electron microscopy images that contain different types of AAV particles. During this process, the algorithm scores each particle to determine whether it is solid or hollow. The results can be divided into three categories: blue indicates that the algorithm identifies the particle as a solid AAV particle, green indicates that it is identified as a hollow AAV particle, and red indicates that the algorithm cannot determine the particle type (e.g., Figure 1 ). This color-coded result visualization provides an intuitive way to show the performance of the algorithm on actual images. For example, by showing the original electron microscopy image and marking different types of AAV particles with corresponding colors on the image, the classification results of the algorithm can be clearly seen. This method not only helps to evaluate the accuracy of the algorithm but also reveals possible problems of the algorithm, such as difficulties in identifying certain particle types or misclassification cases.

[0099] Quantitative analysis of algorithm performance:

[0100] · The hollow ratio is the ratio of the number of hollow AAV particles identified by the algorithm to the total number of AAV particles detected by the algorithm, Hollow_ratio = hollow_detected / Total_detected.

[0101] · True hollow rate: As mentioned above, it is obtained by manual marking and is defined as the number of AAV particles that are actually hollow divided by the total number of detected AAV particles.

[0102] · The key to performance evaluation lies in comparing these two ratios. Ideally, the hollow rate calculated by the algorithm should be very close to the true hollow rate. This comparison can not only evaluate the accuracy of the algorithm but also reflect its reliability in practical applications. To conduct this comparison, it is necessary to manually mark the AAV particles in a part of the test images to determine whether they are actually hollow or solid, and then compare these marks with the classification results of the algorithm. Currently, the comparison accuracy between the test results and the manual marking results is 100%. The process of validating and testing the algorithm also involves several key steps, aiming to ensure that the performance of the algorithm is not only good on the training data but also highly accurate and reliable for new, unseen data. Cross-validation: To comprehensively evaluate the performance of the algorithm, cross-validation methods such as stratified k-fold cross-validation can be used. In this method, the entire dataset is evenly divided into k parts. Each part takes turns as the test set, while the remaining parts are combined as the training set. This helps ensure that each data point is used for testing, providing a more comprehensive and fair performance evaluation. Independent test set: In addition to cross-validation, the algorithm should also be evaluated on a test set that is completely independent of the training set. This test set should contain electron microscope images that have never been used during the training process to ensure that the test results can accurately reflect the algorithm's ability to process new data.

[0103] 2. Detection of AAV applied after module design

[0104] 2.1 Use sample B in Example 2 for electron microscopy detection and AI calculation.

[0105] 2.2 Detection results

[0106] 2.2.1 Through virus particle detection and AI calculation, the hollow shell rate of sample B is 68%. The virus particle titer is 2.8 * 10^10 vp / mL, and there is no significant difference between the particle titer detection result and the ELISA result.

[0107] Affinity chromatography purification of virus particles in the crude feed liquid in Example 4 and detection of the hollow shell rate by HPLC method

[0108] 1. Affinity chromatography purification of virus particles in the crude feed liquid

[0109] 1.1 Instruments and reagents: Automatic liquid chromatography pilot system, manual chromatography column, 5L disposable storage bag, affinity equilibration solution, affinity rinse solution, affinity elution solution, affinity regeneration solution I, affinity regeneration solution II, affinity preservation solution, 0.5M NaOH. 1.2 Sample: Virus clarified liquid.

[0110] 1.3 Operation steps:

[0111] 1.3.1 Cleaning: When cleaning all pipelines with 0.5M NaOH, set the chromatography column to baypass. Then rinse the flow path of the chromatography system with injection water until the pH is neutral.

[0112] 1.3.2 Equilibration: Rinse the flow path of the system and the AAVX chromatography column with the affinity equilibration solution for 3 column volumes or more until the pH and conductivity baseline are stable and the UV is zeroed.

[0113] 1.3.3 Loading: After evacuating the residual gas in the pipeline through the bubble trap, perform loading.

[0114] 1.3.3 Re - equilibration: After evacuating the residual gas in the pipeline through the bubble trap, rinse the chromatography column with the affinity equilibration solution to the monitoring baseline level.

[0115] 1.3.4 Flushing: After evacuating the residual gas in the pipeline through the bubble trap, rinse the chromatography column with the affinity flushing solution for 3 column volumes or more to the monitoring baseline level.

[0116] 1.3.5 Elution: Add 500 ml - 800 ml of ionic equilibration solution to the collection container in advance; elute the chromatography column with the affinity elution solution and collect the AAV component; 3) Calculate the volume of the virus sample collected in the sterile container and add a 2M MgCl2 solution with a final concentration of 10 mM.

[0117] 1.3.6 After the above operations are completed, the virus affinity purification solution is obtained.

[0118] 2. HPLC detection of the empty - capsid rate of AAV virus particles

[0119] 2.1 Instruments and reagents: High - performance liquid chromatograph (Waters, Arc), chromatographic column (Proteomix POR - Q, Sepax, 405950 - 4605).

[0120] 2.2 Mobile phase A: 20 mM BTP - HCl pH 9.5 (Weigh 11.29 g of BTP, dissolve it in 1800 ml of water, adjust the pH to 9.5 with HCl, make up the volume to 2 L with water, filter with a 0.22 μm filter membrane and sonicate).

[0121] 2.3 Mobile phase B: 20 mM BTP - HCl + 0.5M TMAC pH 9.5 (Weigh 11.29 g of BTP and 109.60 g of TMAC, dissolve them in 1800 ml of water, adjust the pH to 9.5 with HCl, make up the volume to 2 L with water, filter with a 0.22 μm filter membrane and sonicate).

[0122] 2.4 Operating parameters

[0123] 2.5 Integration Parameter

[0124]

[0125] 2.6 Detection Result

[0126] Since the hollow virus does not contain nucleic acid and has a weak electronegativity, as the concentration of the eluent increases, the hollow virus is preferentially eluted with a retention time of 5.453 min. The solid virus contains nucleic acid and has a stronger electronegativity than the hollow virus, so its elution time lags behind that of the hollow virus, and its retention time is 13.201 min. Integrate the elution peaks of the hollow virus and the solid virus respectively, calculate the proportion of the integrated area, and the contents of the hollow and solid viruses. As Figure 2 shown, the content of the hollow virus is 66.52%, and the content of the solid virus is 33.48%.

[0127] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A method for virus particle recognition and classification based on negative staining electron microscope images, characterized in that, The method includes: 1) Obtain a crude virus solution; 2) Negatively stain the virus particles in the crude virus solution, image the negatively stained virus particle sample using a room-temperature transmission electron microscope to obtain a target image; 3) Virus particle detection: a. A pre-created convolutional neural network reads the original electron microscopy image of the negatively stained virus and performs preprocessing; b. Based on the pre-created convolutional neural network, perform field-of-view image cutting on the original image, save each cut field of view separately and record the original position; c. Based on the pre-created convolutional neural network, process each field-of-view image separately, perform target detection on the virus particles in each field of view, and output the position of the virus in the current field-of-view image; d. Restore the position of the detected virus in the original image and visualize it, and count the total number of viruses; 4) Virus particle classification: a. According to the original image and the virus detection position, use image processing technology to finely segment each virus particle one by one to generate a virus particle mask; according to the virus original image pixels and mask information, use the pre-created convolutional neural network to perform binary classification on whether the virus is hollow; b. According to the classification results and the total number of viruses, calculate the hollow rate and solid rate of the virus particles in the negatively stained electron microscopy image; Recognize the target image through the pre-created convolutional neural network to obtain the recognition result of a single virus particle, and the recognition result includes one of the virus particle sample being solid, the virus particle sample being hollow, or the virus particle sample being partially solid; the pre-created convolutional neural network is trained based on a dataset of negatively stained electron microscopy images of virus particle samples; The pre-created convolutional neural network includes a classification layer and a segmentation layer. The classification layer consists of a convolutional layer, a pooling layer, and a fully connected layer. The segmentation layer is based on image segmentation technology; among them, the convolutional layer is used to extract features in the image, the pooling layer is used to reduce the spatial size of the features, and the fully connected layer is used for classification decision-making; the segmentation layer is a U-Net or a similar structure, which is used for image segmentation tasks to separate the virus particles in the image from the background for further analysis.

2. The method according to claim 1, wherein Steps 1) and 2) negatively stain the virus in the crude virus solution to obtain a target image, specifically including: (1) Obtain a crude virus solution: Transfect cells using an expression vector, wherein the expression vector carries a nucleic acid molecule encoding a viral capsid protein and a nucleic acid molecule encoding a viral genome, and the density of the cells is sufficient such that the cell density corresponding to the dilution solution does not exceed 10 6 / mL; Culture the transfected cells; Lyse the cells to obtain a cell lysate containing the virus; Centrifuge at high speed to remove cell debris and obtain a crude virus solution; (2) Gradient dilution of the crude virus stock solution: Dilute the crude virus stock solution obtained in step (1) 5-fold with diluent ① to obtain sample A, and then dilute sample A 10-fold with diluent ② to obtain the optimized crude virus stock solution sample B; diluent ① is PBS, and diluent ② is PBS containing a cell lysate supernatant concentration of 1×10 6 cells / mL, that is, resuspend 1×10 6 cells with 1 mL of PBS and freeze-thaw three times repeatedly, and centrifuge at 12,000 rpm for 5 min to take the supernatant; the virus titer of the gradient-diluted crude virus stock solution is not less than 3×10 10 vp / mL; (3) Negatively stain the virus: Suspend the copper grid with a supporting film on the sample B droplet (with the side with the supporting film facing up) for 1 minute, then invert the copper grid on the uranyl acetate staining solution for 10 seconds, blot the excess staining solution with filter paper, repeat this operation twice, but the third inversion time is 30 seconds, and finally dry the copper grid and use a Thermo fisher Talos120C electron microscope for detection to obtain a negatively stained electron microscopy image.

3. The method according to claim 1, wherein The pre-created convolutional neural network in step 3) is created based on a dataset of negatively stained electron microscopy images of virus particle samples. Obtaining a dataset of negatively stained electron microscopy (TEM) images of virus particle samples involves the following steps: (1) Sample collection and preparation: First, collect virus particle samples from laboratories or collaborative research institutions, including novel viruses or known virus strains under study; the samples are purified and concentrated to ensure the visibility of virus particles under the electron microscope; (2) Negative staining and microscopy imaging: Use heavy metal salts (such as phosphotungstic acid or uranyl acetate) to perform negative staining on the virus samples, and image the stained samples under a transmission electron microscope. As needed, different magnifications and imaging parameters are adopted to capture the detailed structure of virus particles; (3) Image collection and screening: Collect all the images generated during the imaging process, and conduct a preliminary screening of the collected images to eliminate images of poor quality, such as images with inaccurate focus, overexposure or underexposure, and those that do not clearly show virus particles; (4) Data annotation and labeling: Experts in the fields of virology and electron microscopy imaging annotate and label the screened images; the labeling includes identifying virus particles in the images and annotating their positions, types, possible states (such as hollow or solid), etc.; (5) Data organization and formatting: Organize the annotated and labeled images into a structured data set; usually, this includes converting the images to a specific format (such as JPEG or PNG) and resolution, and creating a metadata file containing the annotation information; for the convenience of training and validation of machine learning models, the data set is further divided into a training set, a validation set, and a test set; (6) Compliance and ethical considerations: When collecting and using virus samples and their image data, relevant legal and ethical regulations must be complied with, especially when dealing with potentially dangerous or restricted virus materials, to ensure that the use of data complies with the requirements of data privacy and intellectual property rights.

4. The method according to claim 1, wherein The test data set and the training data set differ in statistical characteristics; the data set should also include various challenging images, such as those with high particle density or poor image quality.

5. The method according to claim 1, characterized in that, The method for establishing the pre-created convolutional neural network described in step 3) includes the following steps: (1) Train using the training data set in claim 3 to obtain an initial convolutional neural network; (2) Adjust the model parameters of the initial convolutional neural network based on the test set in claim 3 to obtain a convolutional neural network for virus particle recognition and classification.

6. The method according to claim 5, wherein Step (1) trains using the training data set in claim 3. During the training process, by continuously providing images and adjusting the model parameters, the network learns to recognize features related to hollow and solid virus particles. The features include the shape, size, edge characteristics of the particles, and subtle differences in the internal structure; use a diverse data set for training, including images under different conditions, such as different virus concentrations, different staining techniques, etc.

7. The method according to claim 5 or 6, characterized in that, Step (1) is to train using the training dataset in claim 3. The training process involves using these labeled images to adjust the weights of the convolutional neural network. Usually, a supervised learning method is adopted, where the convolutional neural network attempts to predict the correct label for each image and adjusts its parameters by comparing the prediction with the actual label. During the training process, techniques such as cross-validation are used to evaluate the performance of the model and prevent overfitting. In addition, various techniques can be used to optimize the training, such as adjusting the learning rate, using early stopping, or adjusting the batch size. After training is completed, it is necessary to evaluate the performance of the model on an independent test set to ensure its good generalization ability for new data.

8. The method according to claim 1, characterized in that, Step 3)a preprocesses the original electron microscopy images of virus negative staining, including the following steps: (1) Image size adjustment: Since the CNN requires a fixed-size input, the image is scaled to the size required by the convolutional neural network. Usually, interpolation methods such as bilinear or bicubic interpolation are involved to preserve the important features of the image. (2) Normalization and standardization: Normalization involves scaling the image pixel values to a specific range (such as 0 to 1 or -1 to 1), which helps the convolutional neural network learn more effectively. Standardization usually means subtracting the mean of the dataset from the image pixel values and dividing by the standard deviation, so that the entire dataset has zero mean and unit variance. (3) Contrast adjustment: In electron microscopy images, contrast adjustment highlights the structure of virus particles. In negative staining samples, it is achieved through histogram equalization or other contrast enhancement techniques. (4) Noise reduction and filtering: Electron microscopy images may contain noise, and noise reduction techniques need to be applied to improve the image quality. Common methods include Gaussian filtering, median filtering, or more advanced noise reduction techniques such as non-local means filtering. (5) Data augmentation: Data augmentation includes rotation, flipping, scaling, cropping, or color transformation of the image. (6) Edge detection and enhancement: It is achieved by applying an edge detection algorithm (such as Canny edge detection).

9. The method according to claim 1, wherein The image cutting described in step 3)b refers to splitting the large-size original image into multiple smaller, fixed-size sub-images to make the image segments adapt to the input requirements of the convolutional neural network. The image cutting methods include one or more of the following (1) to (4): (1) Sliding of a fixed-size window: Slide a fixed-size window over the entire original image and crop out a sub-image at each position. The size of this window matches the size of the input layer of the convolutional neural network. For example, if the convolutional neural network requires an input of 224x224 pixels, the window size should also be set to 224x224 pixels. The window moves on the image at a certain stride, and the stride determines the overlap degree between sub-images. A smaller stride will increase the overlapping area, which may lead to more redundant information, but sometimes it can improve the detection accuracy. (2) Adaptive image cutting: Dynamically adjust the size and position of the cutting window according to the image content. For example, use an image processing algorithm to first identify the areas with dense virus particles, and then determine the best position of the cutting window based on the positions and sizes of these areas. (3) Cutting using preprocessing information: Before cutting, preprocessing steps (such as contrast enhancement and edge detection) are applied to more clearly identify the areas that need to be analyzed, which are used to guide the image cutting process and ensure that the cut sub-image contains valuable fields of view; (4) Overlapping and non-overlapping cutting: When cutting, you can choose to create overlapping or non-overlapping sub-images. Overlapping cutting means that there will be a certain degree of overlap between adjacent sub-images, while non-overlapping cutting simply divides the image into strictly separated blocks.

10. The method according to claim 1, characterized in that The convolutional neural network requires careful tuning of the parameters of the segmentation layer and the classification layer to ensure that the final model can balance the requirements of these two tasks and maximize the overall performance.

Citation Information

Cited By

  • Honeycomb type microorganism culture and mixing control method and device

    CN120843746A

  • A honeycomb type microorganism culture and mixing control method and device

    CN120843746B