Method and device for screening AAV serotypes
By constructing a training dataset and training multiple screening models, and using LSTM and convolutional neural networks to screen AAV serotypes, the problems of high cost and low efficiency in existing technologies are solved, comprehensive screening of multiple traits is achieved, screening efficiency is improved and costs are reduced.
Patent Information
- Application Number
- CN202511432545.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2026-02-27
AI Technical Summary
Existing technologies are costly and inefficient in screening novel AAV capsid variants, and the screening results are not comprehensive enough, especially for screening multiple traits.
By constructing a training dataset, multiple initial screening models are trained. LSTM and convolutional neural networks are used to classify the sequences to be screened. Combined with deduplication, the target screening sequences are obtained, enabling the screening of multiple traits.
Comprehensive screening of multiple traits can be achieved without animal experiments, which improves screening efficiency, reduces costs, and yields more comprehensive screening results.
Smart Images

Figure CN121583330A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of sequence screening, in particular to a screening method and device of AAV serotypes. BACKGROUND
[0002] AAV serotypes have application prospects in the field of gene therapy due to their high safety, low immunogenicity and the ability to mediate long-term stable expression of genes. Among them, new tissue-tropic AAV capsid variants need to be developed to enable AAV serotypes to be widely used in clinical practice.
[0003] In related technologies, new AAV capsids are screened through directed evolution technology. Specifically, a large-capacity capsid variant library is constructed by inserting random peptide segments at specific sites of capsid proteins, and then multiple rounds of screening are performed in live animal models to enrich capsid variants that can specifically target target tissues. However, each round of screening in the above scheme involves complex animal experiments, virus production and high-throughput sequencing, resulting in high cost of the entire screening process and reducing screening efficiency. In addition, the above scheme may only screen for a single trait (e.g., optimizing targeting of the cerebral cortex in a mouse model), resulting in incomplete screening results. SUMMARY
[0004] The present application aims to at least partially solve one of the technical problems in the related art.
[0005] To this end, one object of the present application is to provide a screening method for AAV serotypes, which obtains target classification results of multiple traits of a to-be-screened sequence through a plurality of target screening models, and obtains a target screening sequence based on the target classification results of the to-be-screened sequence, without the need for animal experiments and comprehensive screening of multiple traits, thereby improving screening efficiency, reducing screening cost, and making screening results more comprehensive.
[0006] Another object of the present application is to provide a screening device for AAV serotypes.
[0007] To achieve the above-mentioned objects, one embodiment of the present application provides a screening method for AAV serotypes, comprising: constructing a training data set; training a plurality of initial screening models based on the training data set to obtain a plurality of target screening models, wherein each target screening model corresponds to one trait; determining a to-be-screened sequence and inputting the to-be-screened sequence into each target screening model to obtain target classification results of the to-be-screened sequence for multiple traits; determining an initial candidate sequence set based on the target classification results of the to-be-screened sequence; De-duplicate the initial candidate sequence set to obtain a target screening sequence.
[0008] The screening method of the AAV serotype of the embodiment of the applicationapplicationalso have the following additional technical features: In an embodiment of the application, the training data set is constructed, comprising: obtaining a target length and an amino acid type; encoding the target length and the amino acid type according to a preset rule to obtain a sequence of a variant; determining the sequence of the variant and the quantified values of the corresponding multiple traits as a training data set.
[0009] Further, the training of the initial screening model based on the training data set to obtain multiple target screening models, comprising: inputting the sequence in the training data set into the initial screening model of the corresponding trait to obtain a predicted classification result of the sequence; inputting the predicted classification result of the sequence and the real classification result into a loss function to obtain a loss value corresponding to the sequence; updating the network parameters in the initial screening model using the loss value until the initial screening model converges, or when the iteration number of the network parameters reaches a preset number, obtaining a corresponding target screening model.
[0010] Further, the initial screening model comprises a first LSTM layer, a second LSTM layer and a first full connection layer; the inputting of the sequence in the training data set into the initial screening model of the corresponding trait to obtain the predicted classification result of the sequence, comprising: processing the sequence through the first LSTM layer to obtain a first feature vector; processing the first feature vector through the second LSTM layer to obtain a second feature vector; mapping the second feature vector through the first full connection layer to obtain the predicted classification result of the sequence.
[0011] Further, the initial screening model comprises a one-dimensional convolution layer, a global pooling layer and a second full connection layer; the inputting of the sequence in the training data set into the initial screening model of the corresponding trait to obtain the predicted classification result of the sequence, comprising: extracting features of the sequence through the one-dimensional convolution layer to obtain a third feature vector; pooling the third feature vector through the global pooling layer to obtain a fourth feature vector; Mapping the fourth feature vector through the second fully connected layer obtains a predicted classification result of the sequence.
[0012] Further, the determining the sequence to be screened comprises: a sequence generated by performing target amino acid mutation on a seed sequence is determined as the sequence to be screened; and / or a randomly generated sequence is determined as the sequence to be screened; and / or a first sequence and a second sequence are randomly selected from a seed sequence library; a first sequence is spliced with a second sequence to generate a spliced sequence, wherein m and n are positive integers, and m+n is equal to the total length s of the first sequence or the second sequence; one or more predetermined positions of the spliced sequence are replaced to obtain the sequence to be screened.
[0013] Further, the determining the initial candidate sequence set based on the target classification result of the sequence to be screened comprises: if the target classification result of the sequence to be screened meets a screening condition, the sequence to be screened is added to the initial candidate sequence set, wherein the screening condition is that the target classification result indicates that the sequence to be screened meets all traits.
[0014] Further, the removing duplicate processing of the initial candidate sequence set to obtain the target candidate sequence set comprises: determining the similarity of each sequence in the initial candidate sequence set to the remaining sequences; if the amino acid difference between two sequences is less than or equal to a preset number, the sequence added to the target candidate sequence set is determined from the two sequences according to the target classification result.
[0015] To achieve the above-mentioned purpose, another aspect of the embodiments of the present application proposes a screening device for AAV serotypes, the device comprising: a construction module for constructing a training data set; a training module for training a plurality of initial screening models based on the training data set to obtain a plurality of target screening models, wherein each target screening model corresponds to one trait; a first determination module for determining a sequence to be screened and inputting the sequence to be screened into each target screening model to obtain target classification results of the sequence to be screened for a plurality of traits; a second determination module for determining an initial candidate sequence set based on the target classification results of the sequence to be screened; a third determination module for removing duplicate processing of the initial candidate sequence set to obtain a target screening sequence.
[0016] The application provides a screening method and device for AAV serotypes, and the method comprises the following steps: constructing a training data set; training a plurality of initial screening models based on the training data set to obtain a plurality of target screening models, wherein each target screening model corresponds to a trait; determining a to-be-screened sequence, and inputting the to-be-screened sequence into each target screening model to obtain target classification results of the to-be-screened sequence for a plurality of traits; determining an initial candidate sequence set based on the target classification results of the to-be-screened sequence; and performing deduplication processing on the initial candidate sequence set to obtain a target screening sequence. The target classification results of the to-be-screened sequence for a plurality of traits can be obtained through the target screening models of the plurality of traits, and the target screening sequence can be obtained based on the target classification results of the to-be-screened sequence, without the need for animal experiments, and a plurality of traits can be comprehensively screened, so that the screening efficiency is improved, the screening cost is reduced, and the screening results are more comprehensive.
[0017] Additional aspects and advantages of the application will be set forth in part in the description that follows, and in part will become apparent to those skilled in the art upon examination of the following and / or can be learned by practice of the application. BRIEF DESCRIPTION OF DRAWINGS
[0018] The above and / or additional aspects and advantages of the application will become apparent and be more readily understood through consideration of the following description, taken in conjunction with the accompanying drawings, in which: Figure 1 A flow chart of a screening method for AAV serotypes according to an embodiment of the application; Figure 2 A structural schematic diagram of a screening device for AAV serotypes according to an embodiment of the application. DETAILED DESCRIPTION
[0019] Embodiments of the application are described in detail below with reference to examples illustrated in the accompanying drawings, in which the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the accompanying drawings are exemplary and are intended to explain the application, and cannot be understood as a limitation of the application.
[0020] In an embodiment of the application, in order to find AAV capsids with tissue targeting type, the capsid variants in the capsid library modified by random peptides need to be screened through multiple rounds of selection. The purpose of the screening rounds is to narrow down the sequence exploration to a small number of top candidate vectors. Two or more rounds of in vivo screening may be required to reliably identify the best candidate vector, and each round of screening is performed in a live animal model by injecting a virus library and collecting specific tissue data and enriching quantification, and selecting the superior virus library for the next round of screening, thereby resulting in a high cost of the entire screening process and reducing the screening efficiency.
[0021] The AAV serotype screening method and apparatus according to embodiments of the present invention are described with reference to the accompanying drawings.
[0022] First, the AAV serotype screening method proposed according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0023] Figure 1 This is a flowchart of a method for screening AAV serotypes according to an embodiment of the present invention.
[0024] like Figure 1 As shown, the screening method for this AAV serotype includes the following steps: Step S1: Construct the training dataset.
[0025] In one embodiment of the present invention, each trait is important for the clinical translation of AAV serotypes. Based on this, the present invention considers constructing a training dataset based on multiple traits. Specifically, in one embodiment of the present invention, the aforementioned multiple traits may include: production adaptability, NHP (non-human primate) cortical tissue targeting, mouse cortical tissue targeting, and NHP-low liver targeting.
[0026] In one embodiment of the present invention, the method for constructing the training dataset may include the following steps: Step S11: Obtain the target length and amino acid type.
[0027] In one embodiment of the present invention, the target length can be 7, and the amino acid types and abbreviations may include alanine (A), phenylalanine (F), cysteine (C), selenocysteine (U), aspartic acid (D), asparagine (N), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), leucine (L), isoleucine (I), lysine (K), pyrrolidone (O), methionine (M), proline (P), arginine (R), serine (S), threonine (T), and valine (V).
[0028] Furthermore, in one embodiment of the present invention, the target length and amino acid type can be set as needed, and the present invention does not limit them here.
[0029] Step S12: Encode the target length and amino acid type according to preset rules to obtain the sequence of the variant.
[0030] In one embodiment of the present invention, the aforementioned preset rule can be one-hot encoding, that is, encoding the target length and amino acid type according to one-hot encoding to obtain the variant sequence.
[0031] Specifically, in one embodiment of the present application, the sequence length can be fixed to the target length of amino acid residues, and each amino acid in the sequence can be regarded as an independent categorical variable.
[0032] For example, in one embodiment of the present application, assuming that the target length is 7 and the amino acid types include 20 standard amino acids, a unique binary vector with a length of 20 is assigned to each amino acid type, in which one vector element is “1” (i.e., “hot” state) and the remaining 19 vector elements are “0” (i.e., “cold” state), and the position index corresponds to the specific amino acid type. Based on this, for a sequence with a target length of 7, the encoding process is to replace each amino acid at each position with its predefined 20-dimensional One-Hot vector. Thus, the sequence of variants corresponding to the target length of 7 and the amino acid types including 20 standard amino acids is a structured 2-dimensional numerical matrix of 7 rows and 20 columns, i.e., the 7-amino acid sequence is one-hot encoded to obtain a 2-dimensional matrix of (7, 20).
[0033] In addition, in one embodiment of the present application, the sequence obtained by the above steps is adapted to the corresponding target AAV serotype and target position. For example, in one embodiment of the present application, the above target AAV serotype is AAV9, and the target position is 588-589.
[0034] Step S13, determining the sequence of variants and the quantified values of the corresponding multiple traits as a training data set.
[0035] In one embodiment of the present application, in order to integrate data and avoid experimental systematic error problems in the training process, the sequence of variants needs to be classified.
[0036] In one embodiment of the present application, the production capacity can be ranked from high to low, the Y of the top 20% polypeptide sequences is 1, which is defined as easy to produce, and the Y of the remaining sequences is 0, which is defined as not easy to produce. This index is a quantification of production adaptability; the cortical tissue targeting ability of mouse and monkey sequences is ranked from high to low, respectively, and the Y of the top 20% polypeptide is 1, and the Y of the remaining is 0. This index represents the quantification of the cortical tissue targeting of mice and NHP (non-human primates), respectively; according to the liver tissue targeting from high to low, the Y of the last 20% is 1, and the Y of the remaining is 0. This index is a quantification of NHP low liver targeting.
[0037] In addition, in one embodiment of the present application, the sequence of variants and the quantified values of the corresponding multiple traits are determined as a training data set.
[0038] Step S2, training the plurality of initial screening models based on the training data set to obtain a plurality of target screening models.
[0039] In an embodiment of the present application, after obtaining the training data set through the above steps, the plurality of initial screening models can be trained based on the training data set to obtain a plurality of target screening models, wherein each target screening model corresponds to a trait.
[0040] Specifically, in an embodiment of the present application, the method of training the plurality of initial screening models based on the training data set to obtain the plurality of target screening models can include the following steps: Step S21, inputting the sequence in the training data set into the initial screening model corresponding to the trait state to obtain the predicted classification result of the sequence.
[0041] Step S22, inputting the predicted classification result of the sequence and the true classification result into the loss function to obtain the loss value corresponding to the sequence.
[0042] Step S23, updating the network parameters in the initial screening model using the loss value until the initial screening model converges, or when the number of iterations of the network parameters reaches a preset number, obtaining the corresponding target screening model.
[0043] In an embodiment of the present application, the initial screening model can include a first LSTM layer, a second LSTM layer, and a first full connection layer. In an embodiment of the present application, the method of inputting the sequence in the training data set into the initial screening model to obtain the predicted classification result of the sequence can include the following steps: Step 1, processing the sequence through the first LSTM layer to obtain a first feature vector; In an embodiment of the present application, the number of units of the first LSTM layer can be 64.
[0044] In an embodiment of the present application, B sequences in the training data set can be input into the initial screening model, wherein B is a positive integer greater than or equal to 1.
[0045] For example, in an embodiment of the present application, assuming that the target length is 7, the 2-dimensional matrix of each sequence is input into the first LSTM layer, and the first feature vector obtained after calculation is (B, 7, 64).
[0046] Step 2, processing the first feature vector through the second LSTM layer to obtain a second feature vector; In an embodiment of the present application, the number of units of the second LSTM layer can be 32.
[0047] For example, in an embodiment of the present application, the first feature vector (B, 7, 64) is processed by a second LSTM layer to obtain a second feature vector (B, 7, 32).
[0048] Step 3: mapping the second feature vector through a first fully connected layer to obtain a predicted classification result of the sequence.
[0049] In an embodiment of the present application, the first fully connected layer can include a sigmoid activation function.
[0050] In an embodiment of the present application, the second feature vector is mapped through the first fully connected layer to obtain a classification result of the sequence on the trait (e.g., 1 for meeting and 0 for not meeting), thereby obtaining the predicted classification result of the sequence.
[0051] For example, in an embodiment of the present application, assuming that the initial screening model corresponds to the trait of NHP cortical tissue targeting, when the predicted classification result of the sequence is 1, it indicates that the sequence meets the NHP cortical tissue targeting; when the predicted classification result of the sequence is 0, it indicates that the sequence does not meet the NHP cortical tissue targeting.
[0052] Further, in an embodiment of the present application, during the training process of the initial screening model, a combination of Glorot uniform distribution (forward propagation) and orthogonal matrix (cyclic connection) can be used for parameter initialization, and complete reproducibility can be achieved through random seed control, the loss function uses accuracy, and adam is used for optimization.
[0053] In another embodiment of the present application, the initial screening model can include a one-dimensional convolutional layer, a global pooling layer, and a second fully connected layer. In an embodiment of the present application, the method of inputting the sequence in the training data set into the initial screening model corresponding to the trait to obtain the predicted classification result of the sequence can include the following steps: Step a: feature extraction of the sequence through a one-dimensional convolutional layer to obtain a third feature vector.
[0054] In an embodiment of the present application, the one-dimensional convolutional layer can include 128 filter numbers.
[0055] In an embodiment of the present application, B sequences in the training data set can be input into the initial screening model, where B is a positive integer greater than or equal to 1.
[0056] For example, in an embodiment of the present application, assuming that the target length is 7, the 2-dimensional matrix of each sequence is input into the one-dimensional convolutional layer, and the third feature vector obtained through feature extraction is (B, 7, 128).
[0057] Step b, performing a pooling processing on the third feature vector through a global pooling layer to obtain a fourth feature vector.
[0058] In an embodiment of the present application, the fourth feature vector obtained by performing the pooling processing on the third feature vector (B, 7, 128) through the global pooling layer is (B, 7, 128).
[0059] Step c, performing a mapping on the fourth feature vector through a full connection layer to obtain a predicted classification result of the sequence.
[0060] In an embodiment of the present application, the predicted classification result of the sequence obtained by performing the mapping on the fourth feature vector (B, 7, 128) through the full connection layer is (B, 2).
[0061] In an embodiment of the present application, the initial screening model of each trait in the training data set is trained to obtain the target screening model of the corresponding trait, that is, the target screening model of the production adaptability, the target screening model of the NHP (non-human primate) cortex tissue targeting, the target screening model of the mouse cortex tissue targeting, and the target screening model of the NHP low liver targeting can be obtained through the above steps. It should be noted that in an embodiment of the present application, the initial screening model corresponding to each trait can be the same or different, and the embodiments of the present application do not limit this.
[0062] Step S3, determining a sequence to be screened, and inputting the sequence to be screened into each target screening model to obtain a target classification result of the sequence to be screened for multiple traits.
[0063] In an embodiment of the present application, the method for determining the sequence to be screened includes: a sequence generated by performing a target amino acid mutation on a seed sequence is determined as the sequence to be screened; and / or, a randomly generated sequence is determined as the sequence to be screened; and / or, a first sequence and a second sequence are randomly selected from a seed sequence library; a first sequence is spliced with a second sequence to generate a splicing sequence, wherein m and n are positive integers, and m+n is equal to the total length s of the first sequence or the second sequence; and / or, the amino acid at one or more predetermined positions of the splicing sequence is replaced to obtain the sequence to be screened.
[0064] In an embodiment of the present application, the target amino acid can be set as needed, such as 2-4 amino acids.
[0065] Further, in an embodiment of the present application, the number of times of randomly selecting the first sequence and the second sequence from the seed sequence library can be set as needed, such as 5000 times.
[0066] In an embodiment of the present application, the first sequence S1 can include ACDEFGH, and the second sequence S2 can include OMPRSTV.
[0067] Further, in an embodiment of the present application, m and n can be set as needed, for example, m is 3 and n is 4, that is, the first 3 sequences (ACD) of S1 and the last 4 amino acid sequences (RSTV) of S2 are spliced into a new sequence (ACDRSTV) with a length of 7 amino acids.
[0068] Further, in an embodiment of the present application, the one or more predetermined positions can be set as needed, for example, 2. For example, in an embodiment of the present application, assuming that the spliced sequence is ACDRSTV and the predetermined position is 2, the spliced sequence ACDRSTV is mutated by 1 amino acid at the position 2, and the sequence to be screened is AYDRSTV.
[0069] In an embodiment of the present application, the sequence to be screened is determined by the above steps, which can be systematically and unbiasedly searched in a huge sequence space, so that the subsequent screening results are more comprehensive.
[0070] In an embodiment of the present application, after the sequence to be screened is determined by the above steps, the sequence to be screened can be input into each target screening model to obtain the target classification results of the sequence to be screened for multiple traits, without the need for animal experiments, thereby improving the screening efficiency, reducing the screening cost, and making the screening results more comprehensive. In an embodiment of the present application, the vector composed of the classification results obtained by each target screening model can be determined as the target classification results of the sequence to be screened for multiple traits.
[0071] Step S4: determining an initial candidate sequence set based on the target classification results of the sequence to be screened.
[0072] In an embodiment of the present application, after the target classification results of the sequence to be screened are obtained by the above steps, the initial candidate sequence set can be determined based on the target classification results of the sequence to be screened.
[0073] Specifically, in an embodiment of the present application, the method of determining the initial candidate sequence set based on the target classification results of the sequence to be screened can include: if the target classification results of the sequence to be screened satisfy a screening condition, the sequence to be screened is added to the initial candidate sequence set, wherein the screening condition is that the target classification results indicate that the sequence to be screened satisfies all traits.
[0074] In an embodiment of the present application, the target classification result meeting the above screening condition can be a classification result of 1 for all traits, such as (1, 1, 1, 1), i.e., adding sequences with high production capacity, high CNS targeting in monkeys, high CNS targeting in mice, and negative targeting in monkey liver to the initial candidate sequence set.
[0075] In step S5, the initial candidate sequence set is subjected to a deduplication process to obtain a target screening sequence.
[0076] In an embodiment of the present application, after the initial candidate sequence set is obtained through the above steps, the initial candidate sequence set can be subjected to a deduplication process to obtain a target screening sequence.
[0077] In an embodiment of the present application, the method of deduplicating the initial candidate sequence set to obtain the target screening sequence can include determining the similarity of each sequence in the initial candidate sequence set to the remaining sequences, and if the amino acid difference between two sequences is less than or equal to a preset number, determining the sequence to be added to the target candidate sequence set from the two sequences according to the target classification result.
[0078] In an embodiment of the present application, the preset number can be set as needed, such as 1.
[0079] In an embodiment of the present application, if the amino acid difference between two sequences is less than or equal to 1, the sequence with a classification result of 1 for NHP cortical tissue targeting in the target classification result can be determined to be added to the target candidate sequence set, and the other sequence can be deleted.
[0080] Further, in an embodiment of the present application, after the target screening sequence is determined through the above steps, the target screening sequence can be subjected to a wet test (such as virus synthesis and animal administration) for screening, and the test result data can be statistically analyzed. The target screening sequence can be updated to obtain updated data through the statistical result, and the above steps can be repeated for the next round of screening of the updated data.
[0081] According to an embodiment of the present invention, a method for screening AAV serotypes includes: constructing a training dataset; training multiple initial screening models based on the training dataset to obtain multiple target screening models, wherein each target screening model corresponds to one trait; determining the sequence to be screened and inputting the sequence to be screened into each target screening model to obtain the target classification results of the sequence to be screened for multiple traits; determining an initial candidate sequence set based on the target classification results of the sequence to be screened; and performing deduplication processing on the initial candidate sequence set to obtain the target screening sequence. This invention can obtain the target classification results of the sequence to be screened corresponding to multiple traits through target screening models of multiple traits, and obtain the target screening sequence based on the target classification results of the sequence to be screened, without the need for animal experiments, and comprehensively screens multiple serotypes, thereby improving screening efficiency, reducing screening costs, and making the screening results more comprehensive.
[0082] Next, the AAV serotype screening device according to an embodiment of the present invention is described with reference to the accompanying drawings.
[0083] Figure 2 This is a schematic diagram of an AAV serological screening device according to an embodiment of the present invention.
[0084] like Figure 2 As shown, the AAV serotype screening device 10 includes: a construction module 201, a training module 202, a first determination module 203, a second determination module 204, and a third determination module 205, wherein, Module 201 is used to build the training dataset; Training module 202 is used to train multiple initial screening models based on the training dataset to obtain multiple target screening models, wherein each target screening model corresponds to a trait; The first determining module 203 is used to determine the sequence to be screened and input the sequence to be screened into each target screening model to obtain the target classification results of the sequence to be screened for multiple traits. The second determining module 204 is used to determine an initial candidate sequence set based on the target classification results of the sequences to be screened; The third determining module 205 is used to perform deduplication on the initial candidate sequence set to obtain the target screening sequence.
[0085] In one embodiment of the present invention, the above-mentioned construction module 201 is specifically used for: Obtain the target length and amino acid type; The target length and amino acid type are encoded according to preset rules to obtain the sequence of the variant; The sequences of the variants and the corresponding quantized values of multiple traits are used as the training dataset.
[0086] Further, the training module 202 is specifically configured to: input the sequence in the training data set into the initial screening model corresponding to the trait to obtain a predicted classification result of the sequence; input the predicted classification result of the sequence and the real classification result into a loss function to obtain a loss value corresponding to the sequence; update the network parameters in the initial screening model using the loss value until the initial screening model converges, or when the number of iterations of the network parameters reaches a preset number, obtain a corresponding target screening model.
[0087] Further, the initial screening model includes a first LSTM layer, a second LSTM layer and a first full connection layer; the training module 202 is further configured to: process the sequence through the first LSTM layer to obtain a first feature vector; process the first feature vector through the second LSTM layer to obtain a second feature vector; map the second feature vector through the first full connection layer to obtain the predicted classification result of the sequence.
[0088] Further, the initial screening model includes a one-dimensional convolution layer, a global pooling layer and a second full connection layer; the training module 202 is further configured to: extract features of the sequence through the one-dimensional convolution layer to obtain a third feature vector; perform pooling processing on the third feature vector through the global pooling layer to obtain a fourth feature vector; map the fourth feature vector through the second full connection layer to obtain the predicted classification result of the sequence.
[0089] Further, the first determining module 203 is specifically configured to: determine the sequence generated by the target amino acid mutation of the seed sequence as the screening sequence; and / or determine the randomly generated sequence as the screening sequence; and / or randomly select a first sequence and a second sequence from the seed sequence library; splice the first m amino acid fragments of the first sequence and the last n amino acid fragments of the second sequence to generate a spliced sequence, wherein m and n are positive integers, and m+n is equal to the total length s of the first sequence or the second sequence; replace the amino acids at one or more predetermined positions of the spliced sequence to obtain the screening sequence.
[0090] Further, the second determining module 204 is specifically configured to: If the target classification result of the sequence to be screened meets the screening criteria, the sequence to be screened is added to the initial candidate sequence set. The screening criteria are that the target classification result indicates that the sequence to be screened meets all traits.
[0091] Furthermore, the aforementioned third determining module 204 is specifically used for: Determine the similarity between each sequence in the initial candidate sequence set and the remaining sequences; If the amino acid difference between two sequences is less than or equal to a preset number, then the sequence to be added to the target candidate sequence set is determined from the two sequences based on the target classification results.
[0092] According to an embodiment of the present invention, an AAV serotype screening device is proposed. The device constructs a training dataset; trains multiple initial screening models based on the training dataset to obtain multiple target screening models, each corresponding to a trait; determines the sequence to be screened and inputs it into each target screening model to obtain the target classification results of the sequence for multiple traits; determines an initial candidate sequence set based on the target classification results of the sequence; and performs deduplication on the initial candidate sequence set to obtain the target screening sequence. This invention can obtain target classification results of the sequence to be screened corresponding to multiple traits through target screening models of multiple traits, and obtain the target screening sequence based on the target classification results of the sequence, without the need for animal experiments, and comprehensively screens multiple serotypes, thereby improving screening efficiency, reducing screening costs, and making the screening results more comprehensive.
[0093] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0094] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0095] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for screening AAV serotypes, characterized in that, The method includes: Build the training dataset; Multiple initial screening models are trained based on the training dataset to obtain multiple target screening models, wherein each target screening model corresponds to a trait; The sequence to be screened is determined, and the sequence to be screened is input into each of the target screening models to obtain the target classification results of the sequence to be screened for multiple traits; Based on the target classification results of the sequences to be screened, an initial candidate sequence set is determined; The initial candidate sequence set is deduplicated to obtain the target selection sequence.
2. The method according to claim 1, characterized in that, The construction of the training dataset includes: Obtain the target length and amino acid type; The target length and the amino acid type are encoded according to a preset rule to obtain the sequence of the variant; The sequences of the variants and the quantized values of the corresponding traits are used to determine the training dataset.
3. The method according to claim 1, characterized in that, The process of training multiple initial screening models based on the training dataset to obtain multiple target screening models includes: The sequences in the training dataset are input into the initial screening model corresponding to the trait to obtain the predicted classification results of the sequences; The predicted classification result and the true classification result of the sequence are input into the loss function to obtain the loss value corresponding to the sequence; The network parameters in the initial screening model are updated using the loss value until the initial screening model converges, or the corresponding target screening model is obtained when the number of iterations of the network parameters reaches a preset number.
4. The method according to claim 3, characterized in that, The initial screening model includes a first LSTM layer, a second LSTM layer, and a first fully connected layer; the step of inputting sequences from the training dataset into the initial screening model corresponding to the trait to obtain the predicted classification result of the sequences includes: The sequence is processed by the first LSTM layer to obtain a first feature vector; The second feature vector is obtained by processing the first feature vector through the second LSTM layer; The second feature vector is mapped through the first fully connected layer to obtain the predicted classification result of the sequence.
5. The method according to claim 3, characterized in that, The initial screening model includes a one-dimensional convolutional layer, a global pooling layer, and a second fully connected layer; the step of inputting the sequences in the training dataset into the initial screening model corresponding to the trait to obtain the predicted classification result of the sequences includes: The sequence is subjected to feature extraction through the one-dimensional convolutional layer to obtain a third feature vector; The third feature vector is pooled using the global pooling layer to obtain the fourth feature vector; The fourth feature vector is mapped through the second fully connected layer to obtain the predicted classification result of the sequence.
6. The method according to claim 1, characterized in that, The process of determining the sequence to be screened includes: The sequence generated by mutating the seed sequence to the target amino acid is identified as the sequence to be screened; and / or The randomly generated sequence is selected as the sequence to be screened; and / or Randomly select the first and second sequences from the seed sequence library; The first m amino acid fragments of the first sequence are spliced with the last n amino acid fragments of the second sequence to generate a spliced sequence, where m and n are positive integers, and m+n is equal to the total length s of the first or second sequence; The amino acids at one or more predetermined positions in the spliced sequence are replaced to obtain the sequence to be screened.
7. The method according to claim 1, characterized in that, The step of determining an initial candidate sequence set based on the target classification result of the sequence to be screened includes: if the target classification result of the sequence to be screened meets the screening criteria, then the sequence to be screened is added to the initial candidate sequence set, wherein the screening criteria are that the target classification result indicates that the sequence to be screened meets all traits.
8. The method according to claim 1, characterized in that, The process of deduplicating the initial candidate sequence set to obtain the target candidate sequence set includes: Determine the similarity between each sequence in the initial candidate sequence set and the remaining sequences; If the amino acid difference between two sequences is less than or equal to a preset number, then a sequence to be added to the target candidate sequence set is determined from the two sequences based on the target classification result.
9. A screening device for AAV serotypes, characterized in that, The device includes: Build modules are used to construct training datasets; The training module is used to train multiple initial screening models based on the training dataset to obtain multiple target screening models, wherein each target screening model corresponds to a trait. The first determining module is used to determine the sequence to be screened and input the sequence to be screened into each of the target screening models to obtain the target classification results of the sequence to be screened for multiple traits. The second determining module is used to determine an initial candidate sequence set based on the target classification result of the sequence to be screened; The third determining module is used to perform deduplication processing on the initial candidate sequence set to obtain the target screening sequence.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Battery capacity life evaluation and prediction method, system and device and medium
CN119577404A
Machine learning accelerated protein engineering through fitness prediction
US20210403946A1