A method and apparatus for screening candidate drug molecules

By constructing a position probability model and calculating similarity for the splicing regulation database, candidate splicing regulation sequences were screened out, solving the problem of target omission in existing technologies and achieving efficient and accurate screening of candidate drug molecules.

CN122177278APending Publication Date: 2026-06-09XILI TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610212039.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-13
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing small molecule screening methods targeting splicing regulation cannot cover potential targets that are lowly expressed in screening cell lines but highly expressed in disease-related cells, and have failed to accurately establish the association between sequence similarity and drug response consistency, resulting in low accuracy and efficiency in candidate molecule recommendations.

Method used

By constructing a positional probability model based on a splicing regulation database, the similarity between the target donor sequence and the splicing donor sequence in the database is calculated, candidate splicing regulation sequences are screened out, and candidate drug molecules are obtained based on the splicing change and consistency ratio. Combined with multi-order one-heat encoding and feature vector extraction, the organic coupling of sequence similarity and drug response is achieved.

Benefits of technology

It improves the accuracy and efficiency of candidate drug molecule screening, expands the screening scope, enables drug prediction without relying on high expression of target genes, and enhances the accuracy and efficiency of candidate molecule recommendation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122177278A_ABST
    Figure CN122177278A_ABST
Patent Text Reader

Abstract

The application discloses a candidate drug molecule screening method and device, and relates to the technical field of bioengineering, and comprises the following steps: obtaining a target target point donor sequence and a splicing regulation database; the splicing regulation database is obtained by variable splicing analysis of each small molecule in a preset splicing regulation small molecule library based on preset processing conditions; a position probability model is constructed based on the statistical probability distribution of base combinations of each splicing donor sequence in the splicing regulation database at different positions; the similarity between the target target point donor sequence and each splicing donor sequence in the splicing regulation database is calculated based on the position probability model; a plurality of candidate splicing regulation sequences are screened out from the splicing regulation database based on the similarity, and small molecules corresponding to each candidate splicing regulation sequence and splicing change amounts are obtained; the consistency proportion of each small molecule is calculated based on the splicing change amount, and the candidate drug molecule of the target splicing donor sequence is obtained based on the consistency proportion and a preset proportion threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioengineering technology, and in particular to a method and apparatus for screening candidate drug molecules. Background Technology

[0002] Currently, with traditional protein targets becoming increasingly saturated, RNA-targeted drug development has become an important direction for new drug development. RNA plays a core role in key biological processes such as transcriptional regulation, splicing, and stability, making it an emerging target for disease intervention. Compared to nucleic acid drugs such as antisense oligonucleotides (ASO) or siRNA, small molecule RNA-targeted drugs have significant advantages in terms of oral administration, tissue distribution, synthesis costs, and development pathways, combining the potential of small molecule drugs and gene therapy. Among them, the regulation of mRNA splicing has become a hot topic. Small molecule splicing regulators, represented by Risdiplam, have been successfully applied in the clinical treatment of spinal muscular atrophy (SMA), demonstrating the broad prospects of this field.

[0003] Existing small molecule screening methods targeting splicing regulation primarily rely on high-throughput RNA-seq experiments in specific cell lines, such as HEK293T, to identify splicing changes in highly expressed genes. This method cannot cover potential targets that are lowly expressed in the selected cell lines but highly expressed in disease-related cells, leading to target omissions. Some methods, when recommending small molecules that may regulate new targets from existing data, typically use simple sequence alignment or k-mer (k-length subsequence) frequency comparisons to assess sequence similarity. These methods fail to consider the differences in the importance of bases at different positions in the splicing donor sequence, and do not incorporate the dynamic weighting information of the influence of different positions on splicing under small molecule regulatory conditions. Therefore, it is difficult to accurately establish the correlation between sequence similarity and drug response consistency, limiting the accuracy and efficiency of candidate molecule recommendations. Summary of the Invention

[0004] This invention provides a method and apparatus for screening candidate drug molecules to improve the accuracy and efficiency of candidate molecule screening.

[0005] To address the aforementioned technical problems, this invention provides a method for screening candidate drug molecules, comprising: The target donor sequence and splicing regulation database are obtained; the splicing regulation database is constructed by performing variable splicing analysis on each small molecule in a preset splicing regulation small molecule library based on preset processing conditions. A positional probability model is constructed based on the statistical probability distribution of base combinations at different positions of each splice donor sequence in the splice regulation database. The similarity between the target donor sequence and each splice donor sequence in the splice modulation database is calculated based on the location probability model. Based on the similarity, several candidate splicing regulatory sequences are selected from the splicing regulation database, and the small molecules and splicing changes corresponding to each candidate splicing regulatory sequence are obtained. The consistency ratio of each small molecule is calculated based on the splicing change, and candidate drug molecules of the target donor sequence are obtained based on the consistency ratio and a preset ratio threshold.

[0006] This invention acquires target donor sequences and a splicing regulation database. First, it constructs a positional probability model based on the statistical probability distribution of base combinations at different positions in a large number of splice donor sequences within the database. This model objectively reflects the differences in the importance of each position in the splice regulatory function. Then, it uses the positional probability model to calculate the similarity between the target sequence and the database sequences, achieving precise quantification from sequence motif features to functional similarity. Subsequently, it screens several candidate splicing regulatory sequences based on similarity and associates them with their corresponding small molecules and splicing variations, mapping the nearest neighbor relationships in the sequence space to the drug response space. Finally, by calculating the consistency ratio of each small molecule in the candidate sequences and comparing it with a preset threshold, it screens the candidate drug molecules most likely to regulate the splicing of the target target. This allows for drug prediction based solely on sequence information, without requiring high expression of the target gene in the selected cell lines, thus expanding the screening range of splice-regulating drugs. Simultaneously, it organically couples sequence similarity with drug response consistency, improving the accuracy and efficiency of candidate molecule recommendations.

[0007] Furthermore, the construction of the position probability model based on the statistical probability distribution of base combinations at different positions of each splice donor sequence in the splice regulation database includes: Obtain all splicing donor sequences from the splicing control database, and perform deduplication on the splicing donor sequences to obtain a set of donor sequences; Under several preset values, the frequency of occurrence of different base combinations at each position of all donor sequences in the donor sequence set is counted, and a position frequency matrix corresponding to each value is constructed respectively. Normalize the position frequency matrix corresponding to each value to generate the normalized probability matrix corresponding to each value. The normalized probability matrix is ​​expanded and spliced ​​according to a preset order to generate a position probability model.

[0008] This invention ensures the independence and representativeness of the model training data by acquiring all splicing donor sequences from the splicing regulation database and performing deduplication. Under preset numerical values, the frequency of different base combinations at each position is statistically analyzed to construct a multi-scale positional frequency matrix, thereby simultaneously capturing information from different base combinations. The frequency matrix is ​​then normalized into a probability matrix, eliminating the influence of different sequence quantities on probability estimation and enabling the model to objectively reflect the true distribution. Finally, the probability matrices are expanded and concatenated according to a preset order to generate a positional probability model, integrating multidimensional heterogeneous statistical information into a unified, computable one-dimensional vector form. This not only fully preserves the distribution characteristics of each position and multiple local motifs in the sequence but also lays a standardized and reusable data foundation for subsequent efficient feature extraction and similarity calculation, improving the model's ability to represent sequence patterns related to splicing regulation functions.

[0009] Furthermore, the step of calculating the similarity between the target donor sequence and each splice donor sequence in the splicing modulation database based on a preset location probability model includes: The target feature vector of the target donor sequence is extracted based on a preset location probability model; Based on the preset position probability model, the first feature vector of each splice donor sequence in the splice control database is extracted; The similarity between the target donor sequence and each splice donor sequence in the splice modulation database is calculated based on the target feature vector and the first feature vector.

[0010] This invention extracts the target feature vector of the target donor sequence and the first feature vector of each splice donor sequence in the splice regulation database based on the constructed position probability model. This achieves a unified mapping of the original nucleic acid sequence to the feature space defined by the position probability model. Then, it calculates the similarity between the target feature vector and each first feature vector, transforming the differences in base composition at the sequence level into a geometric distance metric in the feature space. This transforms the position probability model from a static statistic into a dynamic feature extractor, ensuring that each sequence can obtain a vector representation that matches its motif features under the same weighting system. This allows the similarity calculation results to fully incorporate the differentiated contributions of each position and base combination in splice regulation.

[0011] Furthermore, the step of extracting the target feature vector of the target donor sequence based on a preset location probability model includes: Based on the numerical values, the target donor sequence is subjected to multi-level one-hot encoding to generate a one-hot feature vector of the target donor sequence. In the one-hot feature vector, the component corresponding to the actual base combination existing at each position of the target splice donor sequence is 1, and the other components are 0. The unique hot feature vector is multiplied element-wise with the position probability model to generate the target feature vector of the target splicing donor sequence.

[0012] This invention generates an one-hot feature vector that precisely corresponds to the local motif distribution of the donor sequence of the target site by performing multi-level one-hot encoding based on preset values. In this vector, the component corresponding to the actual base combination in the sequence is 1, and the rest are 0, thus completely preserving the precise base information of the original sequence. Then, this one-hot feature vector is multiplied element-wise with a position probability model, so that the value of each dimension in the feature vector is no longer just a Boolean value of 0 and 1, but a weighted value obtained by weighting the occurrence probability of that position and base combination in the database. This allows the target feature vector to inherit the lossless expression capability of one-hot encoding for sequence information, while also injecting the prior regulatory knowledge condensed by the position probability model. The resulting target feature vector can simultaneously reflect the precise base composition of the sequence and characterize the degree of conservation or regulatory tendency of the sequence in a known splicing regulatory network, significantly enhancing the predictive ability of the regulatory potential of new target sequences.

[0013] Furthermore, the step of extracting the first feature vector of each splice donor sequence in the splicing modulation database based on the preset location probability model includes: Based on the aforementioned values, multi-level one-hot encoding is performed on each splice donor sequence to generate one-hot feature vectors for each splice donor sequence. In the one-hot feature vectors, the component corresponding to the actual base combination present at each position of the splice donor sequence is 1, and the remaining components are 0. The unique hot feature vectors of each splice donor sequence are multiplied element-wise by the position probability model to generate the first feature vector of each splice donor sequence.

[0014] This invention ensures that the feature vectors of the database sequence and the target feature vector are completely isomorphic in mathematical definition, dimensional space, and weight system by performing multi-order one-hot encoding and element-wise multiplication operations on each database sequence, which are equivalent to those of the target sequence. This ensures the comparability of subsequent similarity calculations and transforms the entire splicing and manipulation database into a standardized feature vector space. This provides an efficient and consistent computational foundation for subsequent nearest neighbor search and similarity ranking, while avoiding systematic biases caused by inconsistent feature extraction methods.

[0015] Furthermore, the step of screening several candidate splicing regulatory sequences in the splicing regulation database based on the similarity, and obtaining the small molecules and splicing changes corresponding to each candidate splicing regulatory sequence, includes: The similarity between the target donor sequence and each splice donor sequence in the splice modulation database is sorted in descending order to obtain the sorting result; Based on a preset number of candidate sequences, a splicing donor sequence with high similarity is selected from the sorting results as a candidate splicing control sequence; Small molecules associated with the candidate splice regulation sequences and their corresponding splice changes are extracted from the splice regulation database.

[0016] This invention obtains a clear similarity priority sequence by sorting the similarity between the target sequence and sequences in various databases in descending order. Then, based on a preset number of candidate sequences, it extracts several splicing donor sequences with the highest similarity from the sorting results as candidate splicing regulation sequences. This achieves a deterministic transformation from continuous similarity values ​​to a discrete candidate set, ensuring the objectivity and repeatability of the screening process. As a result, it can selectively extract small molecules associated with these candidate sequences and their corresponding splicing changes from the splicing regulation database, completing the mapping from sequence similarity to drug association.

[0017] Furthermore, the step of calculating the consistency ratio of each small molecule based on the splicing change amount, and obtaining candidate drug molecules of the target donor sequence based on the consistency ratio and a preset ratio threshold, includes: Based on the splice variation and the preset variation threshold, the target splice control sequence is selected from the candidate splice control sequences; A target regulation record set is constructed based on the target splicing regulation sequence, the small molecules associated with the target splicing regulation sequence, and the corresponding splicing changes; The frequency of each small molecule appearing in the target regulatory record set is counted, and the identity ratio of each small molecule is calculated based on the frequency and the total number of sequences in the target regulatory record set. Small molecules whose consistency ratio reaches or exceeds a preset threshold are identified as candidate drug molecules for the target donor sequence.

[0018] This invention first filters splicing changes based on a preset threshold, retaining only records with significant regulatory effects and eliminating weak or false-positive interference to ensure that subsequent statistics are based on high-quality regulatory events. Then, a target regulatory record set is constructed using these target splicing regulatory sequences, their associated small molecules, and splicing changes as elements. The frequency of each small molecule appearing in this record set is then counted, and this frequency is divided by the total number of candidate sequences to calculate the consistency ratio. This ratio intuitively quantifies the broad-spectrum and stable regulation of each small molecule in the similar sequence space. Finally, small molecules with a consistency ratio reaching or exceeding the preset threshold are identified as candidate drug molecules for the target targets. This achieves the technical effect of accurately identifying high-potential molecules from multi-sequence synergistic responses, significantly improving the confidence and practical value of molecule recommendations.

[0019] Furthermore, before obtaining the target donor sequence and splicing modulation database, the process also includes: A splice-regulating small molecule library is obtained, and variable splicing analysis is performed on the small molecules in the splice-regulating small molecule library based on preset processing conditions. Several splicing events induced by small molecules are identified, and the splicing change amount of the splicing events is determined. Splice donor sequences are extracted from each splicing event based on a preset length, and a splicing regulation database is generated based on small molecules, splice donor sequences, and splicing changes.

[0020] This invention acquires a library of splicing-regulating small molecules containing diverse small molecules and performs variable splicing analysis under preset processing conditions. It systematically identifies small molecule-induced splicing events and quantifies their splicing changes, ensuring the richness and quality of the database source. Furthermore, it extracts splicing donor sequences from each splicing event based on a preset length and stores the small molecule, splicing donor sequence, and splicing change in a standardized triplet data structure. This provides large-scale, high-quality, and structured prior knowledge of splicing regulation, covering the splicing regulatory fingerprints of thousands of small molecules in representative cell lines. This provides solid data support for statistical learning of positional probability models, sequence similarity comparison, and collaborative filtering recommendations, while also giving the method good transferability and continuous iteration capabilities.

[0021] Furthermore, the process of acquiring a splice-regulating small molecule library and performing alternative splicing analysis on the small molecules in the library based on preset processing conditions, identifying several small molecule-induced splicing events, and based on the splicing change amount of the splicing events, includes: Obtain a preset splice-regulating small molecule library, which includes several different types of small molecules; Based on preset processing conditions, at least one cell line was selected from the preset splicing regulation small molecule library to conduct small molecule treatment experiments and obtain processed samples. The processed sample was subjected to transcriptome sequencing to obtain sequencing data. Alternative splicing analysis was performed on the sequencing data to identify several alternative splicing events induced by small molecules, and the splicing changes of the alternative splicing events were calculated.

[0022] This invention ensures the diversity of compound skeletons and enhances the database's coverage of the chemical space by limiting the splicing regulation small molecule library to include several different types of small molecules. Through small molecule treatment experiments in at least one cell line with controls, realistic and comparable samples from the drug-treated and control groups were obtained. Transcriptome sequencing and systematic alternative splicing analysis of the samples accurately identified small molecule-induced alternative splicing events and calculated their splicing changes. By deeply integrating high-throughput wet experiments with bioinformatics analysis, high-confidence, high-resolution data on the regulatory relationships of small molecule splicing events were generated. This allows the database to accurately capture the direct regulatory signals of small molecules on splice donor sites, providing biologically significant underlying features for subsequent position probability model construction and drug response consistency analysis.

[0023] In a second aspect, the present invention provides a candidate drug molecule screening apparatus, including a module for performing the method. Attached Figure Description

[0024] Figure 1 This is a schematic flowchart of a candidate drug molecule screening method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a candidate drug molecule screening device provided in an embodiment of the present invention. Detailed Implementation

[0025] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0026] The terms "first" and "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0027] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0028] Example 1 See Figure 1 , Figure 1 This is a schematic flowchart illustrating a candidate drug molecule screening method provided by an embodiment of the present invention. The embodiment of the present invention provides a candidate drug molecule screening method, including steps 101 to 105, as detailed below: Step 101: Obtain the target donor sequence and splicing regulation database; the splicing regulation database is constructed by performing variable splicing analysis on each small molecule in the preset splicing regulation small molecule library based on preset processing conditions; In this embodiment, before obtaining the target donor sequence and splicing modulation database, the method further includes: A splice-regulating small molecule library is obtained, and variable splicing analysis is performed on the small molecules in the splice-regulating small molecule library based on preset processing conditions. Several splicing events induced by small molecules are identified, and the splicing change amount of the splicing events is determined. Splice donor sequences are extracted from each splicing event based on a preset length, and a splicing regulation database is generated based on small molecules, splice donor sequences, and splicing changes.

[0029] In this embodiment, firstly, a splicing regulation small molecule library is obtained. This library contains a variety of small molecules with different backbone structures, and its quantity is sufficient to cover a wide range of chemical spaces. Subsequently, alternative splicing analysis is performed on each small molecule in the library based on preset processing conditions. These preset processing conditions include: selecting at least one cell line as an experimental model, setting at least one drug concentration and a corresponding solvent control group, and setting a fixed drug incubation time; after processing, cell samples are collected, RNA is extracted, and transcriptome sequencing is performed to obtain sequencing data.

[0030] In this embodiment, for each small molecule and cell line, a drug-treated group and a control group were set up. After incubation for 24 hours, two groups of sample cells were collected, RNA was extracted and RNAseq was performed to obtain sample sequencing data. The sample sequencing data was in FASTQ format.

[0031] In this embodiment, bioinformatics analysis is performed on sequencing data to identify small molecule-induced alternative splicing events using a quantitative alternative splicing algorithm. The splicing change ΔPSI of each event compared to the control group is calculated as a measure of the splicing change in regulatory intensity. Then, for each identified splicing event, the nucleic acid sequence of its splice donor region is extracted according to a preset sequence length, which covers several adjacent bases upstream and downstream of the splice donor site to completely include the core motif that determines the splicing recognition specificity.

[0032] Finally, the extracted splicing donor sequence is associated with its corresponding small molecule and the splicing change ΔPSI caused by the small molecule to the event containing the sequence, and stored to form a splicing regulation database containing a triplet data structure of "small molecule-splicing donor sequence-splicing change".

[0033] In this embodiment, a splicing regulation small molecule library containing diverse small molecules is acquired, and variable splicing analysis is performed under preset processing conditions. This systematically identifies small molecule-induced splicing events and quantifies their splicing changes, ensuring the richness and quality of the database source. Furthermore, splicing donor sequences are extracted from each splicing event based on a preset length, and the small molecule, splicing donor sequence, and splicing change are associated and stored to form a standardized triplet data structure. This provides large-scale, high-quality, and structured prior knowledge of splicing regulation, covering the splicing regulation fingerprints of thousands of small molecules in representative cell lines. This provides solid data support for statistical learning of position probability models, sequence similarity comparison, and collaborative filtering recommendations, while also giving the method good transferability and continuous iteration capabilities.

[0034] In this embodiment, the process of acquiring a splice-regulating small molecule library, performing alternative splicing analysis on the small molecules in the library based on preset processing conditions, identifying several small molecule-induced splicing events, and determining the splicing change based on the splicing events includes: Obtain a preset splice-regulating small molecule library, which includes several different types of small molecules; Based on preset processing conditions, at least one cell line was selected from the preset splicing regulation small molecule library to conduct small molecule treatment experiments and obtain processed samples. The processed sample was subjected to transcriptome sequencing to obtain sequencing data. Alternative splicing analysis was performed on the sequencing data to identify several alternative splicing events induced by small molecules, and the splicing changes of the alternative splicing events were calculated.

[0035] In this embodiment, a splicing regulation small molecule library containing no less than 1000 small molecules was selected, covering a variety of backbone structures, such as pyridines, pyrimidines, indoles, quinolines, and benzo[a]heterocyclic compounds. Using the HEK293T cell line as the experimental model, two drug concentrations of 0.1 μM and 1 μM were set for each small molecule, with an equal volume of DMSO as the solvent control. The treatment time was 24 hours, and each group had three biological replicates.

[0036] In this embodiment, after processing, cells were collected, RNA was extracted and a library was constructed. High-throughput transcriptome sequencing was performed using the Illumina platform to obtain raw data in FASTQ format. A pre-defined alternative splicing analysis workflow was used to identify small molecule-induced neo-exon insertion (pseudo-exon) events in highly expressed genes, excluding other splicing modes such as classical exon skipping. The splicing rate PSI (Percent Spliced ​​In) for each event was calculated, and the splicing change ΔPSI was determined.

[0037] In this embodiment, the highly expressed gene is one with a TPM (Transcripts Per Million) > 5.

[0038] In this embodiment, ΔPSI = - in, The splice index (PSI) of the drug-treated group. The splice rate (PSI) is the control group.

[0039] In this embodiment, for each identified neonatal exon insertion event, a 10-base-long nucleic acid sequence of its 3′ splice donor region is extracted. This sequence covers the 4th (-4) position upstream of the splice donor site to the 6th (+6) position downstream of the splice donor site.

[0040] In this embodiment, to ensure the confidence level of the database, the splicing regulation database is further filtered, retaining only splicing donor sequences with ΔPSI>0.5 under the action of at least 10 small molecules.

[0041] Finally, each preserved splice donor sequence, its corresponding small molecule identifier, and the ΔPSI value generated by the small molecule for the event containing the sequence are stored in the form of a triplet, thus completing the construction of the splice regulation database.

[0042] In this embodiment, by limiting the splicing regulation small molecule library to include several different types of small molecules, the diversity of the compound skeleton was ensured, enhancing the database's coverage of the chemical space. By conducting small molecule treatment experiments in at least one cell line and setting up controls, real and comparable samples from the drug-treated group and the control group were obtained. Through transcriptome sequencing and systematic alternative splicing analysis of the samples, small molecule-induced alternative splicing events were accurately identified, and their splicing changes were calculated. By deeply integrating high-throughput wet experiments with bioinformatics analysis, high-confidence, high-resolution data on the regulatory relationships of small molecule splicing events were generated. This allows the database to accurately capture the direct regulatory signals of small molecules on splice donor sites, providing biologically significant underlying features for subsequent position probability model construction and drug response consistency analysis.

[0043] Step 102: Construct a position probability model based on the statistical probability distribution of base combinations at different positions of each splice donor sequence in the splice regulation database; In this embodiment, the construction of a position probability model based on the statistical probability distribution of base combinations at different positions of each splice donor sequence in the splice regulation database includes: Obtain all splicing donor sequences from the splicing control database, and perform deduplication on the splicing donor sequences to obtain a set of donor sequences; Under several preset values, the frequency of occurrence of different base combinations at each position of all donor sequences in the donor sequence set is counted, and a position frequency matrix corresponding to each value is constructed respectively. Normalize the position frequency matrix corresponding to each value to generate the normalized probability matrix corresponding to each value. The normalized probability matrix is ​​expanded and spliced ​​according to a preset order to generate a position probability model.

[0044] In this embodiment, firstly, all splice donor sequences in the splice regulation database are obtained, and the sequences are deduplicated to remove redundant records caused by multiple small molecules regulating the same splice donor site, forming a donor sequence set composed of non-repetitive splice donor sequences. Subsequently, under several preset values ​​(i.e., k-mer order), the frequency of occurrence of different base combinations at each position of all sequences in the donor sequence set is counted.

[0045] In this embodiment, each position is defined by a sliding window, the window length is equal to a preset value, and the overlap step of adjacent windows is 1 base, thereby completely covering the continuous local motifs of the donor sequence.

[0046] In this embodiment, based on the statistical results of each preset value and its corresponding position, a positional frequency matrix is ​​constructed for each value. The row index of the matrix represents the base combination type, the column index represents the sequence position or window start site, and the matrix elements represent the frequency of occurrence of the corresponding base combination at that position. Then, each positional frequency matrix is ​​normalized column-wise so that the sum of each column element is 1, generating a normalized probability matrix corresponding to each value. The matrix elements reflect the conditional probability of a specific base combination occurring at the current position.

[0047] In this embodiment, finally, according to a preset order (e.g., k values ​​from smallest to largest), each normalized probability matrix is ​​sequentially expanded into a one-dimensional vector, and the beginning and end of each vector are concatenated to generate a fixed-dimensional numerical vector, which is the position probability model. This model condenses the position-specific features of tens of thousands of splice donor sequences at different scales of local motifs in the splicing modulation database in the form of a probability distribution, providing a standardized mathematical expression basis for subsequent sequence feature extraction and similarity comparison.

[0048] As a specific example of an embodiment of the present invention, the splice donor sequence has a uniform length of 10 bases, covering the 4th (-4) position upstream and the 6th (+6) position downstream of the splice donor site. All splice donor sequences are extracted from the splice regulation database, and after deduplication, 2300 non-repetitive sequences are obtained, constituting a donor sequence set.

[0049] In this embodiment, the preset values ​​(k-mer order) are set to k=1, k=2, and k=3, which correspond to the three local motif scales of single base, double base, and triple base, respectively.

[0050] When k=1, count the frequency of the four bases A, G, C, and T in all sequences in the set from the 1st to the 10th position (a total of 10 positions) and construct a 4×10 position frequency matrix; When k=2, using a sliding window with a step size of 1, the frequency of occurrence of 16 kinds of binary bases AA, AC, ..., TT in positions 1-2, 2-3, ..., 9-10 (a total of 9 windows) is counted, and a 16×9-dimensional position frequency matrix is ​​constructed. When k=3, the frequency of occurrence of 64 triplet bases, namely AAA, AAC, ..., TTT, is counted in positions 1-3, 2-4, ..., 8-10 (a total of 8 windows), and a 64×8-dimensional position frequency matrix is ​​constructed.

[0051] The frequency matrices at the three positions mentioned above are subjected to column normalization. Taking the k=1 matrix as an example, the frequencies of the four bases in the first column (i.e., the first position of the sequence) are summed and then divided by the sum to obtain the probability of A, G, C, and T appearing in the first position. The remaining columns are processed in the same way to generate a 4×10 dimensional normalized probability matrix. The k=2 and k=3 matrices are processed in the same way. After normalization, for example, in the k=1 probability matrix, the probability of the G base in the 5th column (corresponding to the +1 position) is close to 0.95, and the probability of the T base is close to 0.05, while in the 1st column, corresponding to the -4 position, the probability distribution of each base is relatively uniform. This distribution characteristic is highly consistent with the conservatism of the human 5′ splice donor consensus sequence "CAGGUAAGU".

[0052] Subsequently, following the order k=1, k=2, and k=3, the three normalized probability matrices are sequentially expanded into one-dimensional vectors: the k=1 matrix expands into a 40-dimensional vector, the k=2 matrix expands into a 144-dimensional vector, and the k=3 matrix expands into a 512-dimensional vector. These three vectors are then concatenated to generate a positional probability model with a total dimension of 40 + 144 + 512 = 696 dimensions. This positional probability model, in the form of a single numerical vector, fully encodes the base distribution probability of the splicing donor sequence at each specific position or window across the three levels of single-base, double-base, and triple-base sequences. It serves as the core benchmark for subsequently mapping the original sequence to the feature space and calculating sequence similarity in this invention.

[0053] In this embodiment, by acquiring all splicing donor sequences from the splicing regulation database and performing deduplication, the independence and representativeness of the model training data are ensured. The frequency of different base combinations at each position is statistically analyzed under several preset values ​​to construct a multi-scale positional frequency matrix, thereby simultaneously capturing information from different base combinations. The frequency matrix is ​​then normalized into a probability matrix, eliminating the influence of different sequence quantities on probability estimation and enabling the model to objectively reflect the true distribution. Finally, the probability matrices are expanded and concatenated according to a preset order to generate a positional probability model, integrating multidimensional heterogeneous statistical information into a unified, computable one-dimensional vector form. This not only fully preserves the distribution characteristics of each position and multiple local motifs in the sequence but also lays a standardized and reusable data foundation for subsequent efficient feature extraction and similarity calculation, improving the model's ability to represent sequence patterns related to splicing regulation functions.

[0054] Step 103: Calculate the similarity between the target donor sequence and each splice donor sequence in the splice modulation database based on the location probability model; In this embodiment, calculating the similarity between the target donor sequence and each splice donor sequence in the splicing modulation database based on a preset location probability model includes: The target feature vector of the target donor sequence is extracted based on a preset location probability model; Based on the preset position probability model, the first feature vector of each splice donor sequence in the splice control database is extracted; The similarity between the target donor sequence and each splice donor sequence in the splice modulation database is calculated based on the target feature vector and the first feature vector.

[0055] In this embodiment, the target feature vector of the target donor sequence is extracted based on the completed position probability model. The position probability model is a numerical benchmark model in the form of probability feature vectors. Its dimensions are fixed and the physical meaning of each dimension is clear. It corresponds to the statistical probability of each base combination at different positions or on the sliding window under different preset values ​​(k-mer order).

[0056] In this embodiment, when extracting the target feature vector, the original base information of the target sequence is mathematically coupled with the position probability model to generate a feature vector with the same dimension as the model and the values ​​of each dimension are weighted by probability. This vector not only fully preserves the precise base composition of the target sequence, but also incorporates the prior knowledge of position importance condensed from thousands of known splicing donor sequences in the database.

[0057] In this embodiment, the exact same feature extraction operation is performed on each splice donor sequence in the splice control database to generate their respective first feature vectors; this operation ensures that all sequences in the database are uniformly mapped to a feature space with the same definition and the same dimensions as the target sequence.

[0058] In this embodiment, finally, using the target feature vector as the query benchmark, the similarity metric between it and each of the first feature vectors is calculated. The similarity metric can employ common distance or angle calculation methods in vector space models, transforming the proximity of two sequences in the feature space into a single numerical score. A higher score indicates greater similarity between the two sequences in the splicing regulatory motif features encoded by the position probability model, and a higher probability that they will produce similar splicing response behaviors under the influence of small molecules.

[0059] In this embodiment, by extracting the target feature vector of the target donor sequence and the first feature vector of each splice donor sequence in the splice regulation database based on the constructed position probability model, the original nucleic acid sequence is uniformly mapped to the feature space defined by the position probability model. Then, the similarity between the target feature vector and each first feature vector is calculated, transforming the differences in base composition at the sequence level into a geometric distance metric in the feature space. The position probability model is transformed from a static statistic into a dynamic feature extractor, ensuring that each sequence can obtain a vector representation that matches its motif features under the same weighting system. This allows the similarity calculation results to fully incorporate the differentiated contributions of each position and base combination in splice regulation.

[0060] In this embodiment, the step of extracting the target feature vector of the target donor sequence based on a preset location probability model includes: Based on the numerical values, the target donor sequence is subjected to multi-level one-hot encoding to generate a one-hot feature vector of the target donor sequence. In the one-hot feature vector, the component corresponding to the actual base combination existing at each position of the target splice donor sequence is 1, and the other components are 0. The unique hot feature vector is multiplied element-wise with the position probability model to generate the target feature vector of the target splicing donor sequence.

[0061] In this embodiment, based on several preset values ​​(i.e., k-mer order) used when constructing the position probability model, multi-order one-hot encoding is performed on the target donor sequence. The multi-order one-hot encoding refers to mapping the base combinations of the target sequence at each position or sliding window to a high-dimensional sparse vector at each preset value. The dimension of this vector is equal to the total number of all possible base combinations at that preset value multiplied by the total number of positions (or windows). Only the component corresponding to the base combination actually existing at that position in the target sequence has a value of 1, and the remaining components are 0. The one-hot sub-vectors generated at each preset value are concatenated in the same order as the position probability model to form an one-hot feature vector with the same dimension as the position probability model.

[0062] In this embodiment, the one-hot feature vector records the precise base composition information of the target sequence at all scale local motifs in a complete and distortion-free Boolean format. Subsequently, the one-hot feature vector is multiplied element-wise with the position probability model; the position probability model is a dense numerical vector of the same dimension, where each element represents the normalized probability of the corresponding position and base combination appearing in the splicing regulation database. After element-wise multiplication, the components with a value of 1 in the one-hot feature vector are replaced with the corresponding probability weight value, while the components with a value of 0 remain 0, thereby generating the target feature vector of the target sequence.

[0063] In this embodiment, the non-zero element values ​​in the target feature vector are no longer merely binary existence markers, but rather regulatory potential scores weighted by prior knowledge of the database: if the target sequence has a conserved motif that appears frequently in the database at a certain position, the corresponding dimension feature value is larger; conversely, if it is a rare variant motif, the feature value is smaller. This successfully maps the discrete base information of the original sequence to a continuous probability feature space defined by thousands of known splicing regulatory sequences, providing a vectorized representation that combines sequence fidelity and regulatory semantic information for subsequent similarity calculations.

[0064] In this embodiment, multi-level one-hot encoding is performed on the target donor sequence based on preset values ​​to generate an one-hot feature vector that precisely corresponds to its local motif distribution. In this vector, the position component corresponding to the actual base combination in the sequence is 1, and the rest are 0, completely preserving the precise base information of the original sequence. Then, this one-hot feature vector is multiplied element-wise by a position probability model, so that the value of each dimension in the feature vector is no longer just a Boolean value of 0 and 1, but a weighted value calculated by the probability of occurrence of that position and base combination in the database. This allows the target feature vector to inherit the lossless representation of sequence information from one-hot encoding while also incorporating the prior regulatory knowledge condensed by the position probability model. The resulting target feature vector can simultaneously reflect the precise base composition of the sequence and characterize the degree of conservation or regulatory tendency of the sequence in a known splicing regulatory network, significantly enhancing the predictive ability of the regulatory potential of new target sequences.

[0065] In this embodiment, the step of extracting the first feature vector of each splice donor sequence in the splice modulation database based on the preset location probability model includes: Based on the aforementioned values, multi-level one-hot encoding is performed on each splice donor sequence to generate one-hot feature vectors for each splice donor sequence. In the one-hot feature vectors, the component corresponding to the actual base combination present at each position of the splice donor sequence is 1, and the remaining components are 0. The unique hot feature vectors of each splice donor sequence are multiplied element-wise by the position probability model to generate the first feature vector of each splice donor sequence.

[0066] In this embodiment, firstly, using several preset values ​​(i.e., k-mer order) and encoding rules that are exactly the same as the extracted target feature vector, multi-order one-hot encoding is performed on each splice donor sequence in the splice control database. The multi-order one-hot encoding is executed independently under each preset value. For each preset value, according to the sliding window length and step size determined by the value, all positions or windows of the sequence are traversed to generate sparse one-hot sub-vectors of the corresponding dimension. In this embodiment, in the sparse one-hot subvector, only the component corresponding to the actual base combination present at that position or window of the current sequence has a value of 1, and the remaining components are all 0. Subsequently, the one-hot subvectors generated for the same sequence under various preset values ​​are concatenated in an order completely consistent with the position probability model to form an one-hot feature vector corresponding to the sequence and with the same dimension as the position probability model. This one-hot feature vector records the precise base composition information of the splice donor sequence at all scale local motifs in Boolean form. Then, the one-hot feature vector is multiplied element-wise with the position probability model; the position probability model is a dense numerical vector of the same dimension, where each element represents the normalized probability of the corresponding position and base combination appearing in the splice regulation database. After element-wise multiplication, the components with a value of 1 in the one-hot feature vector are replaced with the corresponding probability weight value, while the components with a value of 0 remain 0, thereby generating the first feature vector of the splice donor sequence. By performing feature representation on each splice donor sequence in the database, a set of first feature vectors corresponding one-to-one with the database sequences is obtained. This set shares the exact same dimension definition, physical meaning, and probability weight system as the target feature vector, ensuring that subsequent similarity calculations are performed fairly and comparablely within the same feature space.

[0067] In this embodiment, by performing multi-order one-hot encoding and element-wise multiplication operations equivalent to the target sequence on each database sequence, it is ensured that the feature vectors of the database sequence and the target feature vector are completely isomorphic in mathematical definition, dimensional space, and weight system. This ensures the comparability of subsequent similarity calculations, and transforms the entire splicing and manipulation database into a standardized feature vector space. This provides an efficient and consistent computational foundation for subsequent nearest neighbor search and similarity ranking, while avoiding systematic deviations caused by inconsistent feature extraction methods.

[0068] As a specific example of an embodiment of the present invention, the frequency of occurrence of different k-mer segments at each position in all splice donor sequences in the splice regulation database is statistically analyzed to form a position-dependent frequency probability matrix.

[0069] Preferably, k = 1, 2, 3 (single base, one base, three bases).

[0070] For k=1 (single base), since the donor sequence length is N=10, the matrix dimension is 4×10 (4 types of bases × positions 1-10). For k=2, the dimension is 16×9 (16 kinds of binary bases × positions 1-9, corresponding to the binary base sequences at positions 1-2, 2-3, ..., 9-10 respectively). For k=3, the dimension is 64×8 (64 triplet bases × positions 1-8, corresponding to triplet sequences at positions 1-3, 2-4, ..., 8-10 respectively).

[0071] Thus, three position frequency matrices grouped by different k values ​​are obtained. Normalizing the frequency of each matrix yields a position probability matrix (PPM), whose elements reflect the probability distribution of a specific k-mer at a specific location. Next, the three PPM matrices are sequentially expanded into one-dimensional vectors and concatenated in the order k=1, k=2, k=3, forming a position probability model with a total length of 40 + 144 + 512 = 696 dimensions.

[0072] Subsequently, based on the constructed 696-dimensional position probability model as the preset position probability model, the similarity between the target donor sequence and each splice donor sequence in the splice modulation database is calculated.

[0073] In this embodiment, the target site is the ARPE5A1 gene donor sequence "AAGAGTATTA", which is 10 bp in length. First, the target site donor sequence "AAGAGTATTA" is encoded using multi-order one-hot encoding with k=1, k=2, and k=3 respectively, to obtain a 696-dimensional one-hot feature vector of the target site donor sequence. Then, this one-hot feature vector is multiplied element-wise with the position probability model to obtain a 696-dimensional target feature vector of the target sequence.

[0074] Secondly, for each splice donor sequence in the splice modulation database, multi-order one-hot encoding (k=1, 2, 3) is performed to generate a 696-dimensional one-hot feature vector. This feature vector is then multiplied element-wise with the probability model at the same location to generate the 696-dimensional first feature vector for that sequence. This process is repeated for all sequences in the database to obtain the set of first feature vectors.

[0075] Finally, the cosine similarity between the target feature vector and each first feature vector is calculated. The result ranges from [-1, 1]. The closer the result is to 1, the smaller the angle between the two vectors in the feature space and the more consistent their directions are, meaning that the two sequences are more similar in terms of splicing regulation motif distribution characteristics.

[0076] Step 104: Based on the similarity, select several candidate splicing regulation sequences from the splicing regulation database, and obtain the small molecules and splicing changes corresponding to each candidate splicing regulation sequence; In this embodiment, the step of selecting several candidate splicing regulatory sequences from the splicing regulation database based on the similarity, and obtaining the small molecules and splicing changes corresponding to each candidate splicing regulatory sequence, includes: The similarity between the target donor sequence and each splice donor sequence in the splice modulation database is sorted in descending order to obtain the sorting result; Based on a preset number of candidate sequences, a splicing donor sequence with high similarity is selected from the sorting results as a candidate splicing control sequence; Small molecules associated with the candidate splice regulation sequences and their corresponding splice changes are extracted from the splice regulation database.

[0077] In this embodiment, several candidate splicing regulatory sequences are screened from the splicing regulation database based on similarity, and the small molecules and splicing changes corresponding to each candidate splicing regulatory sequence are obtained.

[0078] In this embodiment, firstly, the similarity score between the target donor sequence and each splicing donor sequence in the splicing modulation database is used as the sorting key. All database sequences are then sorted in descending order to generate a similarity ranking list from highest to lowest similarity. This ranking list fully represents the gradient of proximity between the target sequence and all sequences in the database in the feature space.

[0079] In this embodiment, based on a preset number of candidate sequences M, the top M similarity-ranked splicing donor sequences are selected sequentially from the top of the sorted list and identified as candidate splicing modulation sequences. The preset number of candidate sequences M is an adjustable parameter that can be flexibly configured according to the database size, computing resources, and recommendation sensitivity requirements to balance recall and computational efficiency.

[0080] In this embodiment, the sequence identifiers or sequences themselves of these M candidate splicing regulatory sequences are used as query keys to perform an association search operation in the splicing regulation database, extracting all small molecule records associated with each candidate sequence and the corresponding splicing change ΔPSI.

[0081] In this embodiment, since the splice regulation database is stored using a triplet structure of [small molecule, splice donor sequence, splice change], the association retrieval can quickly and accurately return one or more small molecules corresponding to each candidate sequence and the regulatory intensity value generated by the splice event to which the sequence is located.

[0082] In this embodiment, a clear similarity priority sequence is obtained by sorting the similarity between the target sequence and the sequences in each database in descending order. Then, based on a preset number of candidate sequences, several splicing donor sequences with the highest similarity are selected from the sorting results as candidate splicing regulation sequences. This achieves a deterministic transformation from continuous similarity values ​​to a discrete candidate set, ensuring the objectivity and repeatability of the screening process. As a result, small molecules associated with these candidate sequences and their corresponding splicing changes are extracted from the splicing regulation database, completing the mapping from sequence similarity to drug association.

[0083] Step 105: Calculate the consistency ratio of each small molecule based on the splicing change, and obtain the candidate drug molecule of the target donor sequence based on the consistency ratio and the preset ratio threshold.

[0084] In this embodiment, the step of calculating the consistency ratio of each small molecule based on the splicing change amount, and obtaining candidate drug molecules of the target donor sequence based on the consistency ratio and a preset ratio threshold, includes: Based on the splice variation and the preset variation threshold, the target splice control sequence is selected from the candidate splice control sequences; A target regulation record set is constructed based on the target splicing regulation sequence, the small molecules associated with the target splicing regulation sequence, and the corresponding splicing changes; The frequency of each small molecule appearing in the target regulatory record set is counted, and the identity ratio of each small molecule is calculated based on the frequency and the total number of sequences in the target regulatory record set. Small molecules whose consistency ratio reaches or exceeds a preset threshold are identified as candidate drug molecules for the target donor sequence.

[0085] In this embodiment, the consistency ratio of each small molecule is calculated based on the splicing change, and candidate drug molecules of the target splice donor sequence are obtained based on the consistency ratio and a preset ratio threshold.

[0086] In this embodiment, based on a preset change threshold, all splicing changes associated with the acquired candidate splicing regulatory sequences are filtered for significance. The preset change threshold is a quantitative boundary used to distinguish effective regulation from background noise. Only when the absolute value of the splicing change of a small molecule on a candidate sequence exceeds the threshold is the small molecule considered to have a significant regulatory effect on the sequence, and the corresponding record is retained for subsequent analysis. Otherwise, it is regarded as a weak or false positive effect and is eliminated.

[0087] In this embodiment, the target splicing regulatory sequences retained after filtering are used as the core. They are associated and recombined with small molecules that have significant regulatory effects on these sequences and their corresponding splicing changes to construct a target regulatory record set. This record set is a structured data set, and each record clearly identifies the small molecule, the target splicing regulatory sequence and its splicing change, ensuring the accuracy and traceability of subsequent statistics.

[0088] In this embodiment, the target regulatory record set is used as the data source. Small molecule identifiers are grouped and counted, and the frequency of each small molecule appearing in the record set is statistically analyzed. This frequency essentially reflects the number of candidate splicing regulatory sequences that the small molecule can significantly regulate. The frequency is divided by the total number of candidate splicing regulatory sequences in the target regulatory record set, i.e., the number of sequences actually included in the statistics after filtering, to obtain the consistency ratio of the small molecule. This ratio ranges from 0 to 1; a higher value indicates a stronger regulatory spectrum and higher response consistency of the small molecule in the similar sequence space. Finally, the consistency ratio is compared with a preset threshold, and small molecules that reach or exceed the threshold are screened out and identified as candidate drug molecules for the target splicing donor sequence.

[0089] In this embodiment, splicing changes are first filtered based on a preset threshold, retaining only records with significant regulatory effects and eliminating weak or false positive interference to ensure that subsequent statistics are based on high-quality regulatory events. Then, a target regulatory record set is constructed using these target splicing regulatory sequences, their associated small molecules, and splicing changes as elements. The frequency of each small molecule appearing in this record set is then counted, and this frequency is divided by the total number of candidate sequences to calculate the consistency ratio. This ratio intuitively quantifies the broad-spectrum and stable regulation of each small molecule in the similar sequence space. Finally, small molecules with a consistency ratio reaching or exceeding a preset threshold are identified as candidate drug molecules for the target target. This achieves the technical effect of accurately identifying high-potential molecules from multi-sequence synergistic responses, significantly improving the confidence and practical value of molecule recommendations.

[0090] Please refer to Figure 2 , Figure 2 This is a schematic diagram of a candidate drug molecule screening device provided in an embodiment of the present invention.

[0091] The present invention provides a candidate drug molecule screening device, including modules for performing the method, which include a data acquisition module 201, a position probability model calculation module 202, a similarity calculation module 203, a matching module 204 and a screening module 205; The data acquisition module 201 is used to acquire the target donor sequence and splicing regulation database; the splicing regulation database is constructed by performing variable splicing analysis on each small molecule in the preset splicing regulation small molecule library based on preset processing conditions. The position probability model calculation module 202 is used to construct a position probability model based on the statistical probability distribution of base combinations at different positions of each splice donor sequence in the splice regulation database. The similarity calculation module 203 is used to calculate the similarity between the target donor sequence and each splice donor sequence in the splice modulation database based on the position probability model. The matching module 204 is used to filter out several candidate splicing regulation sequences in the splicing regulation database based on the similarity, and to obtain the small molecules and splicing changes corresponding to each candidate splicing regulation sequence. The screening module 205 is used to calculate the consistency ratio of each small molecule based on the splicing change amount, and to obtain candidate drug molecules of the target donor sequence based on the consistency ratio and a preset ratio threshold.

[0092] The present invention provides a communication device, including a processor and an interface circuit. The interface circuit is used to receive signals from other communication devices and transmit them to the processor, or to send signals from the processor to other communication devices. The processor is used to implement the method through logic circuits or execution code instructions.

[0093] The present invention provides a computer program product, including a computer program or instructions, which, when executed by a communication device, implement the method described therein.

[0094] In this embodiment of the invention, a processing device is also provided, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the above-described candidate drug molecule screening method.

[0095] In this embodiment of the invention, a computer-readable storage medium is also provided, which includes a stored computer program, wherein the computer program controls the device where the computer-readable storage medium is located to perform the above-described candidate drug molecule screening method when it is running.

[0096] For example, a computer program can be divided into one or more modules, one or more of which are stored in memory and executed by a processor to perform the present invention. The one or more modules can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a processing device.

[0097] The processing device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The processing device may include, but is not limited to, a processor, memory, and a display. Those skilled in the art will understand that the above components are merely examples of the processing device and do not constitute a limitation on the processing device. It may include more or fewer components than the specified components, or a combination of certain components, or different components. For example, the processing device may also include input / output devices, network access devices, buses, etc.

[0098] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the processing device, connecting all parts of the processing device through various interfaces and lines.

[0099] Memory can be used to store computer programs and / or modules. The processor performs various functions of the processing device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function (such as sound playback, text conversion, etc.), etc.; the data storage area can store data created based on the use of the mobile phone (such as audio data, text message data, etc.). In addition, memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0100] In this invention, the module for screening candidate drug molecules, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. Those skilled in the art can understand and implement this invention without any inventive effort.

[0101] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A method for screening candidate drug molecules, characterized in that, include: Obtain the target donor sequence and splicing regulation database; The splicing regulation database is constructed by performing variable splicing analysis on each small molecule in the preset splicing regulation small molecule library based on preset processing conditions. A positional probability model is constructed based on the statistical probability distribution of base combinations at different positions of each splice donor sequence in the splice regulation database. The similarity between the target donor sequence and each splice donor sequence in the splice modulation database is calculated based on the location probability model. Based on the similarity, several candidate splicing regulatory sequences are selected from the splicing regulation database, and the small molecules and splicing changes corresponding to each candidate splicing regulatory sequence are obtained. The consistency ratio of each small molecule is calculated based on the splicing change, and candidate drug molecules of the target donor sequence are obtained based on the consistency ratio and a preset ratio threshold.

2. The method for screening candidate drug molecules as described in claim 1, characterized in that, The construction of the position probability model based on the statistical probability distribution of base combinations at different positions of each splice donor sequence in the splice regulation database includes: Obtain all splicing donor sequences from the splicing control database, and perform deduplication on the splicing donor sequences to obtain a set of donor sequences; Under several preset values, the frequency of occurrence of different base combinations at each position of all donor sequences in the donor sequence set is counted, and a position frequency matrix corresponding to each value is constructed respectively. Normalize the position frequency matrix corresponding to each value to generate the normalized probability matrix corresponding to each value. The normalized probability matrix is ​​expanded and spliced ​​according to a preset order to generate a position probability model.

3. The method for screening candidate drug molecules as described in claim 2, characterized in that, The calculation of the similarity between the target donor sequence and each splicing donor sequence in the splicing modulation database based on a preset location probability model includes: The target feature vector of the target donor sequence is extracted based on a preset location probability model; Based on the preset position probability model, the first feature vector of each splice donor sequence in the splice control database is extracted; The similarity between the target donor sequence and each splice donor sequence in the splice modulation database is calculated based on the target feature vector and the first feature vector.

4. The method for screening candidate drug molecules as described in claim 3, characterized in that, The step of extracting the target feature vector of the target donor sequence based on a preset location probability model includes: Based on the numerical values, the target donor sequence is subjected to multi-level one-hot encoding to generate a one-hot feature vector of the target donor sequence. In the one-hot feature vector, the component corresponding to the actual base combination existing at each position of the target splice donor sequence is 1, and the other components are 0. The unique hot feature vector is multiplied element-wise with the position probability model to generate the target feature vector of the target splicing donor sequence.

5. The method for screening candidate drug molecules as described in claim 3, characterized in that, The step of extracting the first feature vector of each splice donor sequence in the splicing modulation database based on the preset position probability model includes: Based on the aforementioned values, multi-level one-hot encoding is performed on each splice donor sequence to generate one-hot feature vectors for each splice donor sequence. In the one-hot feature vectors, the component corresponding to the actual base combination present at each position of the splice donor sequence is 1, and the remaining components are 0. The unique hot feature vectors of each splice donor sequence are multiplied element-wise by the position probability model to generate the first feature vector of each splice donor sequence.

6. The method for screening candidate drug molecules as described in claim 1, characterized in that, The process involves selecting several candidate splicing regulatory sequences from the splicing regulation database based on the similarity, and obtaining the small molecules and splicing changes corresponding to each candidate splicing regulatory sequence, including: The similarity between the target donor sequence and each splice donor sequence in the splice modulation database is sorted in descending order to obtain the sorting result; Based on a preset number of candidate sequences, a splicing donor sequence with high similarity is selected from the sorting results as a candidate splicing control sequence; Small molecules associated with the candidate splice regulation sequences and their corresponding splice changes are extracted from the splice regulation database.

7. The method for screening candidate drug molecules as described in claim 6, characterized in that, The step of calculating the consistency ratio of each small molecule based on the splicing change amount, and obtaining candidate drug molecules of the target donor sequence based on the consistency ratio and a preset ratio threshold, includes: Based on the splice variation and the preset variation threshold, the target splice control sequence is selected from the candidate splice control sequences; A target regulation record set is constructed based on the target splicing regulation sequence, the small molecules associated with the target splicing regulation sequence, and the corresponding splicing changes; The frequency of each small molecule appearing in the target regulatory record set is counted, and the identity ratio of each small molecule is calculated based on the frequency and the total number of sequences in the target regulatory record set. Small molecules whose consistency ratio reaches or exceeds a preset threshold are identified as candidate drug molecules for the target donor sequence.

8. The method for screening candidate drug molecules as described in claim 1, characterized in that, Before obtaining the target donor sequence and splice modulation database, the method further includes: A splice-regulating small molecule library is obtained, and variable splicing analysis is performed on the small molecules in the splice-regulating small molecule library based on preset processing conditions. Several splicing events induced by small molecules are identified, and the splicing change amount of the splicing events is determined. Splice donor sequences are extracted from each splicing event based on a preset length, and a splicing regulation database is generated based on small molecules, splice donor sequences, and splicing changes.

9. The method for screening candidate drug molecules as described in claim 8, characterized in that, The process involves acquiring a splice-regulating small molecule library, performing alternative splicing analysis on the small molecules in the library based on preset processing conditions, identifying several small molecule-induced splicing events, and determining the splicing changes based on these events, including: Obtain a preset splice-regulating small molecule library, which includes several different types of small molecules; Based on preset processing conditions, at least one cell line was selected from the preset splicing regulation small molecule library to conduct small molecule treatment experiments and obtain processed samples. The processed sample was subjected to transcriptome sequencing to obtain sequencing data. Alternative splicing analysis was performed on the sequencing data to identify several alternative splicing events induced by small molecules, and the splicing changes of the alternative splicing events were calculated.

10. A candidate drug molecule screening device, characterized in that, Includes a module for performing the method as described in any one of claims 1 to 9.