Library design and analysis method for high-throughput analysis of coding capability and function of non-classical open reading frame

By designing overexpression library vectors and fluorescent protein labeling technology to screen ncORFs, combined with RNA sequencing and mass spectrometry, the efficient and accurate evaluation of ncORFs encoding ability and functional evaluation were solved, and high-throughput screening and early disease judgment were achieved.

CN120356524APending Publication Date: 2025-07-22GUANGZHOU MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510250324.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately evaluate the coding capabilities and functions of non-classical open reading frames (ncORFs), especially in high-throughput screening, with low throughput, high cost and off-target effects.

Method used

A library design method is designed to analyze the encoding capabilities and functions of non-classical open reading frames with high-throughput. By obtaining the target ncORFs data, designing overexpression library vectors, and obtaining stable transfected cells by transfecting cells. Translated ncORFs are screened out using fluorescent protein labeling and flow cell sorting technology, and their encoding capabilities and functions are verified in combination with RNA sequencing and mass spectrometry technology.

Benefits of technology

It realizes efficient and accurate screening and evaluation of ncORFs' coding capabilities and functions, improves throughput and accuracy, can judge related diseases early and provide a basis for disease treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356524A_ABST
    Figure CN120356524A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a library design and analysis method for high-throughput analysis of non-classical open reading frame coding capability and function, and belongs to the technical field of data analysis. The invention discloses a library design method for high-throughput analysis of non-classical open reading frame coding capability and function. The library design method comprises the following steps: acquiring target non-classical open reading frame (ncORFs) data; designing an overexpression library vector according to the ncORFs data of the target; and obtaining the overexpression library according to the overexpression library vector. The analysis method for analyzing the coding capability and function of the non-classical open reading frame at high throughput comprises the following steps: obtaining an overexpression library; obtaining stable transfected cells; according to the stably transfected cells, obtaining the translation ability information of the ncORFs; acquiring the influence information of the ncORFs on the survival of the cells; and acquiring information of the regulation effect of the ncORFs on the cells. According to the invention, the coding capability and function of ncORFs can be efficiently screened and evaluated, and novel functional protein products can be found.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis technology, and particularly to a library design and analysis method for high-throughput analysis of the coding ability and function of non-classical open reading frames. Background Art

[0002] The evaluation of the coding ability and function of non-classical open reading frames (ncORFs) usually relies on traditional overexpression vectors and ATG mutant vectors, and is verified by transfecting cells and conducting molecular and cellular experiments. This traditional experimental verification method has low throughput, complex processes and high costs. Currently, the efficient functional screening of ncORFs mainly relies on CRISPR / Cas9 functional genomic screening libraries. However, the CRISPR / Cas9 screening library method cannot directly evaluate the coding ability of ncORFs, and due to the usually short length of ncORFs, there are challenges in designing and constructing screening libraries, which may cause serious off-target effects. Therefore, there are many deficiencies in the technology for evaluating the coding ability and biological function of ncORFs, especially the lack of efficient and accurate prediction tools and screening technologies suitable for the functional verification of microproteins and small molecules. The evaluation of the coding ability and functionality of ncORFs is a major challenge faced by current technologies. The functions of many ncORFs may be manifested as polypeptides or microproteins, and these molecules usually do not have the clear structures or functions of classical proteins. For example:

[0003] Lack of high-throughput coding ability and functionality evaluation experiments: Most traditional coding ability and functionality evaluation experiments are designed for classical long proteins, with low throughput and cumbersome experimental processes, and are difficult to be directly applied to the high-throughput evaluation of the coding ability and functionality of ncORFs.

[0004] Unable to fully simulate its role in vivo: Even if the function of ncORFs can be verified in experiments, its role in a complex biological environment may be different from that in a laboratory environment, making it difficult to accurately predict its biological function. Summary of the Invention

[0005] The main purpose of the embodiments of this application is to provide a library design and analysis method for high-throughput analysis of the coding ability and function of non-classical open reading frames.

[0006] The technical solution adopted by the present invention is:

[0007] On the one hand, the embodiments of the present invention provide a library design method for high-throughput analysis of the coding ability and function of non-classical open reading frames. The library design method for high-throughput analysis of the coding ability and function of non-classical open reading frames includes the following steps:

[0008] Obtain data of target non-classical open reading frames;

[0009] Design an overexpression library vector according to the non-classical open reading frame data of the target;

[0010] Obtain an overexpression library according to the overexpression library vector.

[0011] Furthermore, the obtaining of the non-classical open reading frame data of the target includes the following steps:

[0012] Obtain potential target non-classical open reading frame sequences by screening from different gene databases;

[0013] Obtain non-classical open reading frame sequences associated with pathology by identifying and interpreting the translation process in cells;

[0014] Obtain predicted non-classical open reading frame sequences using open reading frame prediction software based on the characteristics of gene sequences;

[0015] Use the potential target non-classical open reading frame sequences, the non-classical open reading frame sequences associated with pathology, and the predicted non-classical open reading frame sequences as the non-classical open reading frame data of the target.

[0016] Furthermore, the designing of the overexpression library vector according to the non-classical open reading frame data of the target includes the following steps:

[0017] Select a vector; the vector includes plasmid vectors and viral vectors;

[0018] Obtain a first fluorescent protein and a second fluorescent protein; the excitation light wavelength of the first fluorescent protein is different from that of the second fluorescent protein;

[0019] Design a tag protein and a signal peptide;

[0020] According to the non-classical open reading frame data of the target, set a first promoter and a second promoter inside the vector to obtain an overexpression library vector;

[0021] The first promoter is used to drive the expression of a resistance gene and the first fluorescent protein;

[0022] The second promoter is used to drive the expression of the non-classical open reading frame data of the target, the tag protein, the signal peptide, and the second fluorescent protein.

[0023] Furthermore, the obtaining of the overexpression library according to the overexpression library vector includes the following steps:

[0024] According to the non-classical open reading frame data of the target and the overexpression library vector, perform chip synthesis, library construction, and screening and amplification to construct an overexpression library.

[0025] On the other hand, an embodiment of the present invention provides a method for analyzing the coding ability and function of non-canonical open reading frames in a high-throughput manner. The method for analyzing the coding ability and function of non-canonical open reading frames in a high-throughput manner includes the following steps:

[0026] Obtain an overexpression library through the library design method for analyzing the coding ability and function of non-canonical open reading frames described in any one of the foregoing;

[0027] Introduce the overexpression library into target cells by transfection to obtain stably transfected cells;

[0028] Obtain translation ability information of non-canonical open reading frames based on the stably transfected cells;

[0029] Obtain information on the effect of non-canonical open reading frames on cell survival based on the stably transfected cells;

[0030] Obtain information on the regulatory effect of non-canonical open reading frames on cells based on the stably transfected cells;

[0031] Systematically analyze and obtain the coding ability and function information of non-canonical open reading frames based on the translation ability information of non-canonical open reading frames, the information on the effect of non-canonical open reading frames on cell survival, and the information on the regulatory effect of non-canonical open reading frames on cells.

[0032] Further, the step of introducing the overexpression library into target cells by transfection to obtain stably transfected cells includes the following steps:

[0033] Introduce the overexpression library into target cells by transfection and perform preliminary screening with antibiotics to obtain a primary cell population;

[0034] Based on the primary cell population, combine flow cytometry sorting technology and fluorescence signals to obtain stably transfected cells.

[0035] Further, the step of obtaining translation ability information of non-canonical open reading frames based on the stably transfected cells includes the following steps:

[0036] Based on the stably transfected cells, perform RNA sequencing, analyze the expression of different non-canonical open reading frames in cells, verify the translation potential and function of different non-canonical open reading frames, and obtain the translation ability information of non-canonical open reading frames.

[0037] Further, the step of obtaining information on the effect of non-canonical open reading frames on cell survival based on the stably transfected cells includes the following steps:

[0038] Based on the stably transfected cells, compare their survival with that of the cells transfected with the empty vector control to evaluate the effects of non-canonical open reading frames on cell proliferation and apoptosis, and obtain information on the effects of non-canonical open reading frames on cell survival.

[0039] Further, the method for obtaining information on the regulatory effects of non-canonical open reading frames on cells based on the stably transfected cells includes the following steps:

[0040] Based on the stably transfected cells, analyze the potential functions of non-canonical open reading frames through gene expression data and cell biological function analysis, and analyze the regulatory effects of non-canonical open reading frames on cell biological processes to obtain information on the regulatory effects of non-canonical open reading frames on cells;

[0041] The cell biological processes include cell proliferation and migration.

[0042] Further, the analysis method for high-throughput analysis of the coding ability and function of non-canonical open reading frames further includes the following steps:

[0043] Verify whether the identified non-canonical open reading frames can naturally encode proteins through mass spectrometry technology and immunoprecipitation enrichment treatment, and confirm the coding function of non-canonical open reading frames.

[0044] The embodiments of the present application at least include the following beneficial effects: The present application provides a library design and analysis method for high-throughput analysis of the coding ability and function of non-canonical open reading frames. The library design method for high-throughput analysis of the coding ability and function of non-canonical open reading frames of the present invention includes: obtaining data of target non-canonical open reading frames; designing an overexpression library vector according to the data of target non-canonical open reading frames; obtaining an overexpression library according to the overexpression library vector. The analysis method for high-throughput analysis of the coding ability and function of non-canonical open reading frames of the present invention includes: obtaining an overexpression library through the library design method for high-throughput analysis of the coding ability and function of non-canonical open reading frames; obtaining stably transfected cells; obtaining information on the translation ability of non-canonical open reading frames according to the stably transfected cells; obtaining information on the effects of non-canonical open reading frames on cell survival; obtaining information on the regulatory effects of non-canonical open reading frames on cells. The present invention can screen and evaluate the functions of ncORFs, which is beneficial to the early diagnosis of related diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flowchart of the library design and analysis method for high-throughput analysis of the coding ability and function of non-canonical open reading frames provided by the embodiments of the present invention;

[0046] Figure 2 is a schematic diagram of the vector design, construction and amplification of the overexpression library of target non-canonical open reading frames provided by the embodiments of the present invention;

[0047] Figure 3 It is a schematic diagram of the vector design of the target non - classical open reading frame over - expression library provided by the embodiments of the present invention;

[0048] Figure 4 It is a schematic diagram of systematically analyzing the coding ability of ncORFs and their impact on cell survival provided by the embodiments of the present invention;

[0049] Figure 5 It is a schematic diagram of systematically evaluating and identifying the molecular functions of encodable ncORFs provided by the embodiments of the present invention. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0051] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".

[0052] The terms "at least one", "a plurality of", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, a plurality of includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0054] Before elaborating on the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application are first explained, and the nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0055] 1) ncORFs, the plural form of non-coding open reading frames, represents multiple non-coding open reading frames existing in the genome;

[0056] 2) ncORF, the singular form of non-coding open reading frame, is a region in the genome that, although its sequence may be transcribed, is not translated into a protein;

[0057] 3) FunPEP, is a database used for predicting and analyzing functional peptides;

[0058] 4) ncEP (non-coding epigenetic protein), is a database used for annotating proteins encoded by non-coding RNAs;

[0059] 5) TransLnc, is a database that focuses on the potential roles of translatable lncRNAs in transcriptional regulation, gene expression regulation, etc.;

[0060] 6) Ribo-seq (Ribosome profiling), is a technique used to capture the positions of ribosomes on mRNA during translation;

[0061] 7) ORF Finder, a tool used for predicting open reading frames (ORFs) in DNA sequences;

[0062] 8) FastQC, is a high-throughput sequencing data quality control tool that can perform preliminary analysis on sequencing data;

[0063] 9) Trimmomatic, a tool for preprocessing high-throughput sequencing data that can remove low-quality sequences, adapter sequences, and overly short reads;

[0064] 10) Cutadapt, a tool used for removing adapter sequences from high-throughput sequencing data;

[0065] 11) STAR (Spliced Transcripts Alignment to a Reference), is an RNA-seq data alignment tool that can align RNA sequences to a reference genome or transcriptome;

[0066] 12) Hisat2, an RNA-seq data alignment tool;

[0067] 13) HTSeq, a tool for RNA-seq data analysis;

[0068] 14) featureCounts, a tool for gene expression quantification in RNA-seq data;

[0069] 15) TPM (Transcripts Per Million), one of the commonly used methods for normalizing gene expression levels in RNA-seq data;

[0070] 16) FPKM (Fragments Per Kilobase of transcript per Million mapped reads), a method for normalizing gene expression levels;

[0071] 17) t-SNE algorithm (t-Distributed Stochastic Neighbor Embedding), an algorithm for dimensionality reduction and data visualization;

[0072] 18) ComBat tool, a tool for batch effect correction in RNA-seq or other gene expression data;

[0073] 19) UMAP (Uniform Manifold Approximation and Projection), an algorithm for dimensionality reduction and data visualization;

[0074] 20) Louvain, a community detection algorithm, commonly used in single-cell RNA-seq analysis to identify cell subpopulations (e.g., through graph-based clustering methods);

[0075] 21) Seurat, a single-cell RNA-seq data analysis tool;

[0076] 22) FindMarker is an analysis function in the Seurat toolbox for analyzing and finding genes that are significantly differentially expressed between different cell populations;

[0077] 23) DESeq analysis tool, a tool for differential expression analysis of RNA-seq data, which processes count data based on the negative binomial distribution model and is suitable for gene differential expression analysis of small sample size data;

[0078] 24) edgeR analysis tool, a tool for differential expression analysis of RNA-seq data;

[0079] 25) The GO analysis tool (Gene Ontology analysis) is a tool for analyzing the functional enrichment of gene sets, which annotates according to the biological processes, molecular functions, and cellular components of genes;

[0080] 26) The KEGG analysis tool (Kyoto Encyclopedia of Genes and Genomes) is a gene set enrichment analysis tool used to analyze the relationship between genes and metabolic pathways.

[0081] The following further elaborates on the embodiments of the present invention in conjunction with the accompanying drawings.

[0082] On the one hand, the embodiments of the present invention provide a library design method for high-throughput analysis of the coding ability and function of non-classical open reading frames, referring to Figure 1 The library design method for high-throughput analysis of the coding ability and function of non-classical open reading frames includes the following steps:

[0083] S100. Obtain the data of the target non-classical open reading frame;

[0084] S200. Design an overexpression library vector according to the data of the target non-classical open reading frame;

[0085] S300. Obtain an overexpression library based on the overexpression library vector.

[0086] S100 for obtaining the data of the target non-classical open reading frame disclosed in the embodiments of the present invention includes the following steps:

[0087] S110. Obtain potential target non-classical open reading frame sequences by screening from different gene databases;

[0088] S120. Obtain non-classical open reading frame sequences associated with pathology by identifying and interpreting the translation process in cells;

[0089] S130. Obtain predicted non-classical open reading frame sequences by using an open reading frame prediction software based on the characteristics of gene sequences;

[0090] S140. Use the potential target non-classical open reading frame sequences, non-classical open reading frame sequences associated with pathology, and predicted non-classical open reading frame sequences as the data of the target non-classical open reading frame.

[0091] S200 for designing an overexpression library vector according to the data of the target non-classical open reading frame disclosed in the embodiments of the present invention includes the following steps:

[0092] S210. Select a vector; the vector includes a plasmid vector and a viral vector;

[0093] S220. Obtain a first fluorescent protein and a second fluorescent protein; the excitation light wavelength of the first fluorescent protein is different from that of the second fluorescent protein;

[0094] S230. Design a tag protein and a signal peptide;

[0095] S240. According to the data of the target non-classical open reading frame, set a first promoter and a second promoter inside the vector to obtain an overexpression library vector;

[0096] The first promoter is used to drive the expression of a resistance gene and the first fluorescent protein;

[0097] The second promoter is used to drive the expression of the target non-classical open reading frame data, the tag protein, the signal peptide, and the second fluorescent protein.

[0098] S300 disclosed in the embodiments of the present invention to obtain an overexpression library according to the overexpression library vector includes the following steps:

[0099] S310. According to the data of the target non-classical open reading frame and the overexpression library vector, perform chip synthesis, library construction, and screening and amplification to construct an overexpression library.

[0100] On the other hand, the embodiments of the present invention also provide an analysis method for high-throughput analysis of the coding ability and function of non-classical open reading frames. The analysis method for high-throughput analysis of the coding ability and function of non-classical open reading frames includes the following steps:

[0101] S400. Obtain an overexpression library through the library design method for high-throughput analysis of the coding ability and function of non-classical open reading frames in any one of the previous items;

[0102] S500. Introduce the overexpression library into target cells by transfection to obtain stably transfected cells;

[0103] S600. Obtain the translation ability information of the non-classical open reading frame according to the stably transfected cells;

[0104] S700. Obtain the information on the influence of the non-classical open reading frame on the survival of cells according to the stably transfected cells;

[0105] S800. Obtain the information on the regulatory effect of the non-classical open reading frame on cells according to the stably transfected cells;

[0106] S900. Based on the translation ability information of non-classical open reading frames, the information on the impact of non-classical open reading frames on cell survival, and the information on the regulatory effects of non-classical open reading frames on cells, the coding ability and functional information of non-classical open reading frames are systematically analyzed and obtained.

[0107] S500 disclosed in the embodiments of the present invention imports an overexpression library into target cells by transfection to obtain stably transfected cells, including the following steps:

[0108] S510. Import an overexpression library into target cells by transfection, and perform preliminary screening with antibiotics to obtain a primary cell population;

[0109] S520. Based on the primary cell population, combine flow cytometry sorting technology and fluorescence signals to obtain stably transfected cells.

[0110] S600 disclosed in the embodiments of the present invention obtains the translation ability information of non-classical open reading frames based on stably transfected cells, including the following steps:

[0111] S610. Based on stably transfected cells, perform RNA sequencing, analyze the expression of different non-classical open reading frames in cells, verify the translation potential and functions of different non-classical open reading frames, and obtain the translation ability information of non-classical open reading frames.

[0112] As an optional implementation manner, the translation ability information of the non-classical open reading frames in the embodiments of the present invention includes translation efficiency, codon optimization information, translation rate, mRNA secondary structure influence, and information entropy.

[0113] The formulas used in the embodiments of the present invention for calculating translation efficiency include:

[0114]

[0115] Among them, Protein Abundance is the abundance of the protein, mRNA Abundance is the abundance of the mRNA, TE is the translation efficiency, and this ratio reflects the translation efficiency from mRNA transcription to protein synthesis.

[0116] The formulas used in the embodiments of the present invention for calculating translation rate include:

[0117]

[0118] Among them, R is the translation rate, w i is the translation efficiency of codon i, f i is the frequency of codon i in the gene, and N is the number of all possible codon types.

[0119] As an alternative embodiment, when predicting the translation potential of non-classical open reading frames in the embodiments of the present invention, information entropy is used to measure the uncertainty of the mRNA sequence. A lower information entropy indicates a higher stability and translation ability of the mRNA sequence. The calculation formula for information entropy is as follows:

[0120]

[0121] where H(X) is the entropy of the sequence, and p(x i ) is the occurrence efficiency of each base (or codon) in the sequence.

[0122] S700 disclosed in the embodiments of the present invention obtains information on the impact of non-classical open reading frames on cell survival based on stably transfected cells, including the following steps:

[0123] S710. Compare the survival of stably transfected cells with that of cells with empty vector control to evaluate the impact of non-classical open reading frames on cell proliferation and apoptosis, and obtain information on the impact of non-classical open reading frames on cell survival.

[0124] As an alternative embodiment, the embodiments of the present invention provide a mathematical model for cell survival to calculate the change in cell number, and the formulas used include:

[0125]

[0126] where N is the cell number, r is the growth rate, K is the carrying capacity of the environment, a is the apoptosis rate, represents the change in cell number.

[0127] S800 disclosed in the embodiments of the present invention obtains information on the regulatory effect of non-classical open reading frames on cells based on stably transfected cells, including the following steps:

[0128] S810. Based on stably transfected cells, analyze the potential function of non-classical open reading frames through gene expression data and cell biological function analysis, and analyze the regulatory effect of non-classical open reading frames on cell biological processes to obtain information on the regulatory effect of non-classical open reading frames on cells;

[0129] Cell biological processes include cell proliferation and migration.

[0130] The analysis method for high-throughput analysis of the coding ability and function of non-classical open reading frames disclosed in the embodiments of the present invention further includes the following steps:

[0131] S1000. Through mass spectrometry technology and immunoprecipitation enrichment treatment, verify whether the identified non-classical open reading frame can naturally encode proteins and confirm the coding function of the non-classical open reading frame.

[0132] As an optional implementation manner, the embodiment of the present invention includes:

[0133] 1. Design, construction and amplification of target overexpression non-canonical open reading frame library

[0134] Design, construction and amplification of target overexpression non-canonical open reading frame libraries Figure 2 The specific contents are as follows:

[0135] (1) Design of target overexpression vector:

[0136] The present invention provides a vector and a supporting method for high-throughput analysis of the coding capacity and function of non-classical open reading frames. The vector has the following significant characteristics, reference Figure 3 , as follows:

[0137] In terms of vector type, it can be a plasmid vector, which can efficiently shuttle into host cells with its unique structural advantages; it can also exist in the form of a viral vector, which can achieve efficient transfection of host cells with the help of the powerful infection characteristics of the virus, laying a solid foundation for subsequent experimental steps.

[0138] The internal structure of the vector houses two promoters that operate independently of each other: Promoter 1 (the first promoter) is used to drive the expression of antibiotic resistance genes and fluorescent protein 1 (the first fluorescent protein). During the experiment, the antibiotic resistance gene makes the cells that have been successfully transfected with the vector (stable transfected cells) tolerant to antibiotics, so that the stable transfected cells survive in the antibiotic screening environment; and fluorescent protein 1 emits fluorescence of a unique wavelength, allowing researchers to intuitively and conveniently identify transfected positive cells (stable transfected cells) through fluorescent signals. Promoter 2 (the second promoter) is used to drive the expression of the target non-classical open reading frame (ncORF) and its closely associated tag protein, signal peptide and fluorescent protein 2 (the second fluorescent protein), opening up a pathway for studying the function of ncORF.

[0139] To ensure the accuracy and purity of the experiment, all codons (ATG) that can initiate protein translation in the tag protein, signal peptide, and fluorescent protein 2 in the vector were mutated without changing the function of the encoded protein. Therefore, the entire translation process will completely rely on the translation start codon of the target ncORF to start, avoiding other unnecessary interference factors, making the exploration of ncORF coding ability more accurate.

[0140] Design features of fluorescent protein 1 and fluorescent protein 2: They have different excitation wavelengths. When researchers observe fluorescence imaging, they can clearly distinguish different expression situations based on fluorescence signals of different wavelengths. They can then identify cells transfected with translatable ncORFs and cells transfected with untranslatable ncORFs, providing a basis for subsequent differential analysis.

[0141] The role played by tag proteins (such as FLAG tags) in immunofluorescence experiments can assist researchers in verifying the presence and characteristics of related proteins. At the same time, the signal peptides linked between the target ncORF and fluorescent protein 2 are designed, and sequences such as the GS sequence and T2A sequence can be selected. These sequences can effectively block the potential interference of fluorescent protein 2 on the translation process of ncORF and the functions of its translated protein polypeptides, improving the accuracy of experimental results.

[0142] (2) Ways to obtain the target non-classical open reading frame:

[0143] The sequences of target ncORFs can be obtained by the following methods:

[0144] Using public gene databases: Such as databases like FunPEP, ncEP, TransLnc, etc., and then screening gene information. According to specific research needs and screening criteria, potential target ncORF sequences are read out from them.

[0145] With the help of translatome technologies such as Ribo-seq: Deeply analyze the translation dynamics in cells. Through the fine interpretation of the translation process, accurately identify those ncORFs that are closely related to new physiological or pathological phenomena, providing new targets for disease mechanism research and innovative drug development.

[0146] Using open reading frame prediction software: Such as ORF Finder, based on the inherent characteristics and rules of gene sequences, efficiently predict new ncORFs that have not been discovered, expanding the research boundaries.

[0147] (3) Construction and amplification of the target overexpressed non-classical open reading frame library:

[0148] According to the scale requirements of library construction, customize the target ncORFs overexpression library:

[0149] When facing the need for small-scale library construction, use PCR amplification technology. By designing primers with restriction enzyme sites, ensure that the gene sequences of target ncORFs are obtained completely and accurately during the amplification process, providing high-quality gene materials for subsequent experiments.

[0150] If a large-scale library needs to be constructed, use chemical synthesis methods to batch-synthesize a large number of ncORFs and implant them into vectors through cloning operations. This method is suitable for high-throughput screening scenarios to improve research efficiency.

[0151] The constructed library is diverse in composition, including empty vectors, which provide a benchmark for subsequent comparative analysis; it also includes several translatable ncORFs with known functions, which can provide reference for the study of unknown ncORFs and help to interpret the characteristics of newly discovered ncORFs faster and better.

[0152] Finally, the constructed library is introduced into host cells such as Escherichia coli by biological methods such as heat shock or electroporation, enabling the vector to initiate cloning and amplification, and reserving sufficient experimental materials for subsequent experimental studies.

[0153] 2. References Figure 4 , systematically analyze the translatability of ncORFs and their effects on cell survival efficiency

[0154] (1) Cell transfection and infection steps:

[0155] The constructed ncORFs overexpression library is introduced into target cells by transfection or infection. With the help of the fluorescence signals emitted by fluorescent protein 1 and fluorescent protein 2, visually monitor and evaluate the efficiency of the transfection process, and preliminarily evaluate the proportion of translatable ncORFs, and real-time control the uptake of the library by cells.

[0156] (2) Stable transfected cell screening stage:

[0157] Use antibiotics for preliminary screening to enrich the cell population that has successfully and stably taken up the vector. On this basis, with the help of flow cytometry sorting technology, precisely isolate the positively transfected cells labeled with fluorescent protein 1 to ensure the consistency and stability of the cell population used in subsequent experiments.

[0158] (3) Cell sorting process based on fluorescent protein 2:

[0159] Focus on the fluorescent protein 1-labeled positive cells obtained in the previous step, and use flow sorting again to subdivide according to whether fluorescent protein 2 is expressed. Among them, cells expressing fluorescent protein 2 represent that the target ncORF they carry has translation activity and can successfully express functional proteins; while cells not labeled with fluorescent protein 2 indicate that the target ncORF they carry cannot be effectively translated.

[0160] (4) Multiple verifications of experimental reliability:

[0161] We collected some cells labeled with fluorescent protein 2 and some cells that were not labeled, and prepared them into cell slides. We used specific antibodies against the label protein to perform cell immunofluorescence experiments. Through careful observation and analysis of the fluorescence signal, we further confirmed that cells labeled with fluorescent protein 2 do express translatable target ncORFs, and vice versa, they express untranslatable target ncORFs, laying a solid data foundation for subsequent research.

[0162] (5) Construction of RNA sequencing library and sequencing operation:

[0163] In order to significantly improve the sensitivity and accuracy of the detection, cells expressing and not expressing fluorescent protein 2 in the previous step were selected to extract RNA. Specific probes were used to accurately capture the fusion expression RNA containing ncORF-tag protein-signal peptide-fluorescent protein 2. Subsequently, according to the standardized process, the captured target RNA was constructed into a conventional RNA sequencing library, and deep sequencing analysis was carried out to explore the key information hidden at the RNA level.

[0164] (6) Quality control and in-depth analysis of RNA sequencing data:

[0165] Initial assessment of raw data quality: Use FastQC software to conduct a comprehensive quality assessment of the raw data obtained by sequencing, check core indicators such as sequence quality, GC content, and repetition rate, and eliminate the risk of data contamination or low-quality sequences.

[0166] Data cleaning and optimization: Use Trimmomatic or Cutadapt tools to remove low-quality sequences and connector contamination, fully guarantee the purity and accuracy of the data, and create excellent conditions for subsequent precise analysis.

[0167] Sequence alignment and genome mapping: Use STAR or Hisat2 software to rigorously align the cleaned high-quality data with the reference genome, accurately evaluate the alignment rate and the proportion of unaligned sequences, and ensure that the vast majority of sequences can be accurately mapped to the corresponding positions in the genome.

[0168] Expression statistics and standardization: Use HTSeq or featureCounts tools to accurately count the expression levels of ncORFs, and use standardization methods such as TPM and FPKM to deeply evaluate the consistency of expression profiles between samples, so that data from different samples are comparable.

[0169] Batch effect and sample difference analysis: PCA or t-SNE algorithm is used to deeply explore batch effects in data and subtle differences between samples. ComBat tool is used to accurately correct batch effects to ensure the reliability and stability of experimental results.

[0170] (7) Precise identification and experimental verification process of translatable ncORFs:

[0171] Identification of translatable ncORFs: Systematically count the ncORFs expressed in the cell population expressing fluorescent protein 2 and the cell population not expressing fluorescent protein 2 respectively. Through comparative analysis, define translatable and non - translatable ncORFs. Specifically, the ncORFs identified in the cells expressing fluorescent protein 2 are those with translation potential, while the ncORFs found in the cells not expressing fluorescent protein 2 are determined to be non - translatable. For those ncORFs not detected in both types of cells, it is most likely due to the failure of vector transfection, so their translation potential cannot be effectively evaluated.

[0172] Verification of the natural coding ability of translatable ncORFs: For all ncORFs determined to be translatable, directly use proteomic techniques to comprehensively detect the total proteins of tissue cells. By searching the protein index library of the target ncORFs, judge whether they can naturally encode proteins. For the ncORF - encoded proteins not identified in the initial proteomic detection, further design targeted specific antibodies. After performing protein immunoprecipitation enrichment treatment, conduct proteomic detection again and re - search the protein index library of the target ncORFs to significantly improve the sensitivity of mass spectrometry detection and firmly verify whether the target ncORFs have the ability to be naturally translated and stably exist.

[0173] (8) Systematic evaluation of the impact on cell survival:

[0174] Further analyze the targeted RNA - sequencing data to deeply evaluate the impact of translatable ncORFs on cell survival. The specific method is as follows: Compare the expression levels of these ncORFs with the empty - vector control group. The ncORFs with significantly increased expression are determined to promote cell survival, those with significantly decreased expression are determined to inhibit cell survival, and those with no significant difference are considered to have no impact on survival. At the same time, verify the "translatable" ncORFs with known biological effects to ensure the reliability of the experimental results. Next, for the "translatable" ncORFs closely related to cell survival, construct over - expression, translation start - codon mutation, knockdown and corresponding control lentiviruses, and stably transfect the target cell line. Then, use a cell counting kit (such as CCK - 8 or MTT method) to accurately evaluate the proliferation ability and survival rate of each group of cells. At the same time, use flow cytometry to deeply detect the cell cycle distribution and apoptosis situation to comprehensively verify whether the over - expression of different "translatable" ncORFs has a substantial impact on cell survival.

[0175] According to the experimental results, the translatable ncORFs were classified into three categories based on their effects on cell survival: first, ncORFs that promote cell survival, manifested as a significant increase in cell proliferation rate or a significant decrease in apoptosis rate; second, ncORFs that inhibit cell survival, showing a slowdown in cell proliferation or an exacerbation of apoptosis; third, ncORFs with no significant effect, that is, changes in their expression levels did not have an obvious impact on the cell survival situation. Through such a systematic and comprehensive survival analysis, the potential functions of each translatable ncORF in the process of cell survival can be deeply revealed, laying a solid foundation for further exploring its mechanism of action.

[0176] 3. Reference Figure 5 , systematically evaluate and identify the molecular functions of cncORFs

[0177] (1) Cell transfection and infection operation procedures:

[0178] The overexpression library of carefully constructed non-classical open reading frames (ncORFs) was introduced into target cells by transfection or infection. With the fluorescence signals emitted by fluorescent protein 1 and fluorescent protein 2, visually monitor and evaluate the efficiency of the transfection process, and preliminarily evaluate the proportion of translatable ncORFs, and monitor the cell uptake of the library in real time.

[0179] (2) Screening strategy for stably transfected cells:

[0180] With the selection pressure imposed by antibiotics, screen out the cell population that has successfully and stably taken up the vector (the primary selected cell population). On this basis, further use the flow cytometry sorting technology to judge and isolate the positively transfected cells labeled with fluorescent protein 1 to obtain the basis of cell samples.

[0181] (3) Cell sorting based on fluorescent protein 2:

[0182] Focus on the positively labeled cell population with fluorescent protein 1 locked in the previous stage of screening, and re-enable the flow sorting technology to perform a secondary sorting operation to isolate the cells labeled with fluorescent protein 2 and the cells without this fluorescent label, providing clearly classified cell samples for subsequent differential analysis.

[0183] (4) Measures for in-depth verification of experimental reliability:

[0184] We carefully collected some cell samples labeled with fluorescent protein 2 and some unlabeled cells and made them into cell slides. Then, we used specific antibodies against the label protein to perform cell immunofluorescence experiments. Through the observation and interpretation of the fluorescence signal, we further verified that the target ncORF carried by cells labeled with fluorescent protein 2 has efficient translation ability and can successfully express functional protein products; on the contrary, the target ncORF carried by unlabeled cells does not have translation function, which builds a solid defense line for the reliability of experimental data.

[0185] (5) Targeted single-cell RNA sequencing data systematically analyze pathways that encode molecular functions of ncORFs:

[0186] Single-cell suspension preparation and sequencing library construction: Single-cell suspension is prepared from cell samples labeled with fluorescent protein 1 and fluorescent protein 2 using professional technology, and a single-cell sequencing library targeting ncORF-fluorescent protein fusion RNA is constructed following standardized processes, thereby advancing single-cell RNA sequencing work.

[0187] Key links in data quality control: Implement quality control procedures on the data obtained by sequencing to screen out low-quality cells, double cells, and cell samples with excessively high mitochondrial gene expression.

[0188] Standardization of sequencing data: Standardization methods such as TPM and RPKM are used to normalize gene expression data, so that data from different samples and batches are comparable and consistent.

[0189] Accurate batch effect correction strategy: Enable professional batch effect correction tools such as ComBat to eliminate technical deviations caused by different experimental batches and ensure the stability and consistency of data throughout the entire process.

[0190] Dimensionality reduction and clustering multivariate analysis methods: Comprehensive use of cutting-edge dimensionality reduction algorithms such as UMAP and t-SNE, combined with clustering analysis techniques such as Louvain and Seurat, to scientifically classify and visualize cell samples, which helps to intuitively understand the internal structure and differential characteristics of cell populations.

[0191] Application of differential expression deep analysis tools: With the help of analysis tools such as FindMarker, DESeq2 or edgeR, deep mining and revealing of subtle functional differences between cell populations expressing different ncORFs can provide key clues for subsequent functional exploration.

[0192] (6) Functional analysis and exploratory research strategies of encoding ncORFs:

[0193] For cells overexpressing empty vector, known function, and unknown function ncORFs, the following analysis steps were performed:

[0194] Analysis quality control: Evaluate the credibility of experimental analysis results by comparing cells overexpressing empty vectors with cells with known functional ncORFs.

[0195] Differential gene analysis: Compared with overexpression and empty vector, evaluate the effects of ncORFs with different unknown functions on gene expression profiles, and gain insights into the changing patterns at the gene expression level.

[0196] Gene co-expression network analysis: In-depth exploration of the key role of ncORFs in complex gene regulatory networks and revealing the synergistic mechanism between genes.

[0197] Functional enrichment analysis: Use GO, KEGG and other analysis tools to deeply explore and obtain the potential functional information of ncORFs in the entire process of cell biology.

[0198] Special research on biological processes such as cell proliferation, apoptosis, and migration: By comparing different ncORFs and their mutants, we can obtain information on their unique roles in core biological processes such as cell growth, survival, and migration.

[0199] (7) In-depth verification of the functional mechanisms of ncORFs:

[0200] In the early high-throughput screening work, after locking in the translatable ncORFs that are closely related to cell survival and function, comprehensive functional verification is performed. A variety of in vivo and in vitro experimental methods are used to comprehensively evaluate the specific roles of candidate ncORFs in key biological processes such as cell growth, survival, apoptosis, migration, and invasion, and further use sequencing analysis technology to deeply reveal their inherent molecular mechanisms of action. For the selected candidate translatable ncORFs, diversified vectors such as normal overexpression, start codon mutation, and gene knockout targeting ncORFs are constructed, and they are transfected into cells to construct stable cell lines, and the following in vivo and in vitro experiments and mechanism research experiments are carried out in order:

[0201] Core tasks of in vivo and in vitro experiments: Accurately evaluate the actual role of candidate ncORFs in biological processes such as cell growth, survival, apoptosis, and migration, and provide intuitive evidence for functional verification.

[0202] Key measures of mechanism research experiments: Through transcriptome sequencing, protein immunoprecipitation combined with protein spectrum analysis and various cutting-edge omics research methods, in-depth analysis and revealing the intrinsic mechanism of action of ncORFs.

[0203] Effects of the embodiments of the present invention:

[0204] 1. Excellent high-throughput screening efficiency: The present invention can quickly lock in the translatable ncORFs closely related to cell survival in a large-scale library. Compared with traditional methods, its screening efficiency has increased exponentially, and the accuracy has been greatly improved, saving a great deal of scientific research time and resource costs.

[0205] 2. Precise insight into coding ability: By means of the innovative dual-labeling technology of fluorescent proteins and combined with advanced RNA sequencing methods, it is possible to accurately distinguish between translatable and non-translatable ncORFs. This has laid a solid data foundation for subsequent in-depth functional analysis and provided the possibility for precisely interpreting the biological code of ncORFs.

[0206] 3. Excellent application prospects: The present invention breaks through the limitations of a single field and is beneficial to cancer research. Due to its universality, it can be extended to many frontier research fields such as cell survival and development processes, providing new ideas and technical support for scientific research breakthroughs in various fields.

[0207] The embodiments of the present invention innovatively provide a vector and a supporting method for high-throughput analysis of the coding ability and function of ncORFs, realizing a comprehensive screening and precise evaluation of the functions of ncORFs, which is conducive to the early diagnosis of related diseases and the optimization of treatment strategies.

[0208] The high-throughput analysis method introduced by the present invention can efficiently and systematically conduct in-depth evaluations of the coding ability and molecular functions of ncORFs. It can not only accurately screen out the codable ncORFs related to cell survival from a large number of gene sequences, but also, with the help of cutting-edge single-cell sequencing technology, delve into the microscopic world of cells and explore its internal molecular mechanisms. The present invention is beneficial to the research of cancer, development, and other various diseases, provides technologies for revealing the role of ncORFs in the process of complex diseases, and is also helpful for multiple scientific research fields such as new drug screening, disease mechanism exploration, and in-depth analysis of gene functions.

[0209] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, which does not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of rights of the embodiments of the present application.

Claims

1. A library design method for high-throughput analysis of the coding ability and function of non-classical open reading frames, characterized in that, The library design method for high-throughput analysis of the coding ability and function of non-canonical open reading frames includes the following steps: Obtain data of the target non-canonical open reading frames; Design an overexpression library vector according to the data of the target non-canonical open reading frames; Obtain an overexpression library according to the overexpression library vector.

2. The library design method for high-throughput analysis of the coding ability and function of non-classical open reading frames according to claim 1, wherein The obtaining of the data of the target non-canonical open reading frames includes the following steps: Obtain potential target non-canonical open reading frame sequences by screening from different gene databases; Obtain non-canonical open reading frame sequences associated with pathology by identifying and interpreting the translation process in cells; Obtain predicted non-canonical open reading frame sequences by using open reading frame prediction software based on the characteristics of gene sequences; Use the potential target non-canonical open reading frame sequences, the non-canonical open reading frame sequences associated with pathology, and the predicted non-canonical open reading frame sequences as the data of the target non-canonical open reading frames.

3. The library design method for high-throughput analysis of the coding ability and function of non-classical open reading frames according to claim 1, wherein The design of the overexpression library vector according to the data of the target non-canonical open reading frames includes the following steps: Select a vector; the vector includes a plasmid vector and a viral vector; Obtain a first fluorescent protein and a second fluorescent protein; the excitation light wavelength of the first fluorescent protein is different from that of the second fluorescent protein; Design a tag protein and a signal peptide; According to the data of the target non-canonical open reading frames, set a first promoter and a second promoter inside the vector to obtain an overexpression library vector; The first promoter is used to drive the expression of a resistance gene and the first fluorescent protein; The second promoter is used to drive the expression of the data of the target non-canonical open reading frames, the tag protein, the signal peptide, and the second fluorescent protein.

4. The library design method for high-throughput analysis of the coding ability and function of non-classical open reading frames according to claim 1, characterized in that The obtaining of the overexpression library according to the overexpression library vector includes the following steps: According to the data of the target non-canonical open reading frames and the overexpression library vector, perform chip synthesis, library construction, and screening and amplification to construct an overexpression library.

5. A method for high-throughput analysis of the coding ability and function of non-classical open reading frames, characterized in that The analysis method for high-throughput analysis of the coding ability and function of non-canonical open reading frames includes the following steps: Obtain an overexpression library through the library design method for high-throughput analysis of the coding ability and function of non-canonical open reading frames according to any one of claims 1 to 4; Introduce the overexpression library into target cells by transfection to obtain stably transfected cells; Obtain information on the translation ability of non-canonical open reading frames according to the stably transfected cells; Obtain information on the impact of non-canonical open reading frames on cell survival according to the stably transfected cells; Obtain information on the regulatory effect of non-canonical open reading frames on cells according to the stably transfected cells; Systematically analyze to obtain the coding ability and function information of non-canonical open reading frames according to the translation ability information of non-canonical open reading frames, the impact information of non-canonical open reading frames on cell survival, and the regulatory effect information of non-canonical open reading frames on cells.

6. The analysis method for high-throughput parsing of the coding ability and function of non-classical open reading frames according to claim 5, characterized in that, The introduction of the overexpression library into target cells by transfection to obtain stably transfected cells includes the following steps: The overexpression library is introduced into target cells by transfection, and preliminary screening is carried out using antibiotics to obtain a primary selection cell population; Based on the primary selection cell population, combined with flow cytometry and fluorescence signals, stable transfected cells are obtained.

7. The analysis method for high-throughput parsing of the coding ability and function of non-canonical open reading frames according to claim 5, characterized in that The obtaining of the translation ability information of non-classical open reading frames based on the stable transfected cells includes the following steps: Based on the stable transfected cells, RNA sequencing is performed to analyze the expression of different non-classical open reading frames in cells, verify the translation potential and functions of different non-classical open reading frames, and obtain the translation ability information of non-classical open reading frames.

8. The analysis method for high-throughput parsing of the coding ability and function of non-classical open reading frames according to claim 5, characterized in that The obtaining of the information on the impact of non-classical open reading frames on cell survival based on the stable transfected cells includes the following steps: Based on the stable transfected cells, the survival conditions are compared with those of cells with empty vector control to evaluate the impact of non-classical open reading frames on cell proliferation and apoptosis, and obtain the information on the impact of non-classical open reading frames on cell survival.

9. The analysis method for high-throughput parsing of the coding ability and function of non-classical open reading frames according to claim 5, wherein The obtaining of the information on the regulatory role of non-classical open reading frames on cells based on the stable transfected cells includes the following steps: Based on the stable transfected cells, through gene expression data and cell biological function analysis, the potential functions of non-classical open reading frames are analyzed, and the regulatory role of non-classical open reading frames on cell biological processes is analyzed to obtain the information on the regulatory role of non-classical open reading frames on cells; The cell biological processes include cell proliferation and migration.

10. The analysis method for high-throughput parsing of the coding ability and function of non-classical open reading frames according to claim 5, wherein The high-throughput analysis method for analyzing the coding ability and functions of non-classical open reading frames further includes the following steps: Through mass spectrometry technology and immunoprecipitation enrichment treatment, it is verified whether the identified non-classical open reading frames can naturally encode proteins, and the coding functions of non-classical open reading frames are confirmed.