Gene insertion site analysis system and analysis method for stem cell therapeutic agent inserted with specific gene
Patent Information
- Application Number
- CN202180093173.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-08
- Filing Date
- 2021-08-27
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2041-08-27
AI Technical Summary
但是,现有的分析方法不仅具有有关灵敏度、重现性、准确性及存在有害基因的问题,而且,因无法积极应用最新信息通信(ICT,Information&Communications Technology)技术而难以有效分析大量碱基序列分析数据
Smart Images

Figure CN116802739B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a gene insertion site analysis system and method for stem cell therapeutic agents containing specific genes. Background Technology
[0002] To treat rare and intractable diseases that cannot be treated with synthetic new drugs, gene therapy, cell therapy, and gene-cell therapy are being developed as next-generation biological drugs.
[0003] Gene therapy is limited by its potential carcinogenicity and technological constraints. Cell therapy, utilizing living cells, initially focused on skin regeneration and cartilage repair using skin or chondrocytes. However, with active research, second-generation cell therapy (inserting functionally enhancing genes) has become increasingly important, surpassing simpler cell therapies targeting tumors and degenerative diseases (first-generation cell therapy). Cell therapy primarily uses adult stem cells (hematopoietic stem cells or mesenchymal stem cells). Hematopoietic stem cells are known to continuously generate blood through proliferation and differentiation. Mesenchymal stem cells, contributing to hematopoietic stem cell proliferation, exist in approximately one million cells in the bone marrow and play a role in differentiating from the bone marrow to aid in the regeneration of various organs. Cell therapy based on mesenchymal stem cells can effectively promote the regeneration of damaged tissue through differentiation, rather than replacing it. Since mesenchymal stem cells alone lack regenerative therapeutic effectiveness, it is necessary to insert functionally enhancing genes into mesenchymal stem cells for passage culture to develop cell therapy.
[0004] The in vitro manipulation processes described above (in vitro manipulation methods - gene insertion for functional enhancement, passage culture, culture conditions, etc.) and the differentiation capacity of the cells themselves exhibit genetic instability. Genetic instability is caused by mechanisms such as mutation, mismatch repair deficiency, and chromosomal instability. Mismatch repair deficiency refers to the hypermutability state of various genes, and is frequently found in length mutations and microsatellites with short, repetitive base sequences. Chromosomal instability refers to the impact on the number or overall structure of chromosomes; due to differences between cells, chromosomal abnormalities can vary from cell to cell.
[0005] When structural defects as described above occur, they can be passed on to the next generation of cells, and with repeated proliferation, chromosomal abnormalities can occur during replication. Considering cell or gene durability, off-target effects (the problem of editing useless genes), etc., the U.S. Food and Drug Administration (FDA) has determined that premarket clinical trials alone cannot confirm all theoretical risks associated with gene therapy or gene-cell therapy. Therefore, assuming that postmarket clinical studies can promptly address these theoretical risks, long-term follow-up observation is considered particularly important. Stem cell therapies using chromosomal insertion viruses (ex-vivotherapy) carry the risk of genotoxicity. For example, there have been cases where leukemia was induced six years after inserting genes such as LMO2, BMI1, CND2, and EVI1 into 3kb–10kb hematopoietic stem cells (HSCs). Unlike hematopoietic stem cells, mesenchymal stem cells (MSCs), as stromal cells, do not exhibit in vivo colonization (the phenomenon where a bacterial species grows more vigorously in the vicinity of other bacterial species). Because they do not leave any residue in vivo, there is no risk of genotoxicity. However, to verify safety, gene insertion site correlation analysis is required for each passaged culture and chromosome of MSCs infected with inserted viruses. Typically, gene base sequence analysis is performed as follows: MSCs are collected from bone marrow and infected with a transgenic virus (retrovirus, adenovirus, lentivirus, etc.). DNA is extracted from each passaged culture and amplified using linear amplification-mediated polymerase chain reaction (LAM-PCR). The DNA is then analyzed using various next-generation sequencing (NGS) platforms or through genome walking methods. However, existing analytical methods not only have problems with sensitivity, reproducibility, accuracy, and the presence of harmful genes, but also cannot effectively analyze large amounts of base sequence analysis data due to the inability to actively apply the latest information and communication technologies (ICT).
[0006] The above background technology is content that the inventors mastered or learned in the process of deriving the disclosure of this application, and therefore should not be regarded as known technology that was publicly disclosed before this application. Summary of the Invention
[0007] Technical issues
[0008] The purpose of this invention is to provide a gene insertion site analysis system and method for stem cell therapeutic agents containing specific genes. After collecting mesenchymal stem cells from bone marrow and performing the first passaging culture for the purpose of manufacturing gene cell therapeutic agents, DNA is extracted from the cell core through in vitro tissue (infected by a virus with a specific gene insertion) and passaging culture. Next-generation sequencing (NGS) technology is used to ensure base sequence analysis data for each passaging culture and chromosome. This allows for the extraction and storage of gene insertion site (integration site) information (start and end positions), quantity, and biotype. Therefore, this system not only provides the ability to generate various morphological analysis reports but also utilizes the latest cloud computing (SaaS) technology, enabling domestic and international organizations (enterprises, public institutions, universities, etc.) developing gene cell therapeutic agents to operate as an independent system.
[0009] The technical objectives to be solved by the embodiments of the present invention are not limited to the above-mentioned objectives. Those skilled in the art to which this invention pertains can clearly understand other technical objectives not mentioned through the following description.
[0010] Technical solution
[0011] The following describes the gene insertion site analysis system and method for stem cell therapeutic agents containing specific genes according to embodiments of the present invention.
[0012] The gene insertion site analysis system for stem cell therapeutic agents containing specific genes includes: a basic information management unit for managing code information, equipment information, staff information, work information, customer information, partner company information, parameter information, reference gene information, oncogene information, and base sequence conversion table information required to operate the gene insertion site analysis system for stem cell therapeutic agents containing specific genes; a gene insertion site analysis unit for managing order information, contract information, project information, gene insertion site analysis information, and project execution result information required to operate the above-mentioned gene insertion site analysis system, including customer orders, work instructions, and material order information; and a database ( The DB management unit is used to manage the basic information database, order information database, contract information database, project information database, gene insertion site information database, base sequence analysis information database, and project progress information database generated or referenced by the aforementioned basic information management unit and gene insertion site analysis unit. The aforementioned basic information database includes company information database, user information database, reference genome information database, catalog of somatic mutations in cancer (COSMIC) information database, equipment information database, human resources information database, material information database, standard working information database, base sequence conversion table database, code information database, and parameter information database.
[0013] According to one embodiment of the present invention, the aforementioned basic information management unit includes a basic information management module, which includes: a company information management module for managing company information; a user information management module for managing a user information database; a reference genome information management module for managing a reference genome information database; a cancer somatic mutation catalog information management module for managing a cancer somatic mutation catalog information database; an equipment information management module for managing an equipment information database; a human resources information management module for managing a human resources information database; a material information management module for managing a material information database; a standard working information management module for managing a standard working information database; a base sequence conversion table management module for managing a base sequence conversion table database; a code information management module for managing a code information database; and a parameter information management module for managing a parameter information database.
[0014] According to one embodiment of the present invention, the gene insertion site analysis unit includes: an order information management module for managing order information; a contract information management module for managing contract information; a project management module for managing project information; a gene insertion site analysis module for managing gene insertion site information; and a project results management module for managing project progress information.
[0015] On the other hand, the gene insertion site analysis method for stem cell therapeutic agents with specific genes includes the following steps: a basic information management step, which records the basic information required to run the gene insertion site analysis system for stem cell therapeutic agents with specific genes in a database; a project start step, which registers the agreement process with the client in the project information database, registers the order information in the order information database, registers the contract terms based on the order information in the contract information database, registers the project information in the project progress database, and executes the project; a sequencing data processing step, which includes data preprocessing, data structuring, data annotation and registration, and data extraction and registration for analysis of the sequencing data generated in the above project execution steps; a gene insertion site analysis step, which includes setting the search conditions for the analysis object, selecting the analysis method, confirming the analysis report for the analysis results, and registering the analysis results in the database; and a project end step, which involves preparing the analysis report to be submitted to the client, registering client feedback, calculating the total input cost, and ending the project after receiving payment.
[0016] According to one embodiment of the present invention, the above-mentioned project execution steps include: a sequencing work information registration step, in which sequencing work information is registered in an order information database and a project progress information database for sequencing work. In the sequencing work information registration step, when the last record is read during the process of sequentially reading sequencing data from beginning to end, the sequencing result information is registered in the gene insertion site information database and the project progress information database. Furthermore, in the sequencing work information registration step, records where the sequencing data is a multiple of 4 + 1 are read, duplicate data are registered as a single intrinsic value, and variable values are separated and registered in additional records. Moreover, in the sequencing work information registration step, records where the sequencing data is a multiple of 4 + 2 are read in units of 4 bytes to convert the corresponding values in the base sequence conversion table database into binary numbers. Wherein, when the value read from the sequencing data in units of 4 bytes contains N or U, or when the corresponding value is not found in the base sequence conversion table database, the value is stored in the base information error table of the gene insertion site information database. Furthermore, in the above-mentioned sequencing work information registration step, the records of sequencing data that are multiples of 4 + 4 are compressed as quality fractions and stored in the base information quality fraction table of the gene insertion site information database.
[0017] According to one embodiment of the present invention, data for the above-mentioned gene insertion site analysis is generated and each passaged culture is analyzed to register its data in a gene insertion site information database, and the correlation between genes in each passaged culture and each chromosome is analyzed.
[0018] According to one embodiment of the present invention, in the above-described gene insertion site analysis step, mesenchymal stem cells collected simultaneously from bone marrow and mesenchymal stem cells with multiple inserted genes can be passaged and cultured under the same conditions, and expression level information can be compared and analyzed using t-value technology. Specifically, for the above-described expression level information, the total expression level of each passage, the total number of genes, the expression level of each biotype, the expression level of each chromosome, the number of genes on each chromosome, the expression level of each chromosome and biotype, the expression level of each gene, the expression level of the inserted gene, and the expression levels of genes adjacent to the inserted gene and the inserted gene itself can be compared and analyzed.
[0019] The effects of the invention
[0020] The embodiments of the present invention have the following effects.
[0021] First, compared to stem cell therapies with multiple inserted genes, there is no theoretical risk that can be predetermined during the Investigational New Drug (IND) stage.
[0022] Second, it can effectively compare the base sequence data of each passaged culture with human base sequence data for stem cells with multiple inserted genes, and can analyze the correlation between genes in each passaged culture and chromosome.
[0023] Third, reliability can be ensured by comprehensively managing the entire process from client requests to analyze gene insertion sites for stem cell therapeutics containing specific genes to the development of reports for clinical trial approval applications.
[0024] Fourth, cloud computing and software as a service (SaaS) technologies, which are the latest computer technologies, can be applied to enable many researchers at home and abroad to use it as an independent system.
[0025] The effects of the gene insertion site analysis system and analysis method for stem cell therapeutic agents with specific genes inserted in the embodiments of the present invention are not limited to the effects mentioned above. Those skilled in the art to which this invention pertains can clearly understand other effects not mentioned through the following description. Attached Figure Description
[0026] Figure 1 This is a schematic diagram illustrating the structure of a gene insertion site analysis system for stem cell therapeutic agents containing specific genes, according to an embodiment of the present invention.
[0027] Figure 2 For brevity Figure 1The diagram shows the business process of the gene insertion site analysis system from the customer's analysis request to the submission of the final gene insertion site analysis report.
[0028] Figure 3a is a structural block diagram illustrating a gene insertion site analysis system for stem cell therapeutic agents containing specific genes according to an embodiment of the present invention.
[0029] Figure 3b is a structural block diagram showing the basic information management module of Figure 3a.
[0030] Figure 3c is a structural block diagram showing the basic information database of Figure 3a.
[0031] Figure 4 The flowchart below briefly illustrates a method for analyzing the gene insertion sites of inserted mesenchymal stem cells using a gene insertion site analysis system for stem cell therapeutic agents containing specific genes, and submitting a gene insertion site analysis report based on a customer request, according to an embodiment of the present invention.
[0032] Figure 5 For further detailed explanation Figure 4 The flowchart shows the basic information management steps.
[0033] Figure 6 For further detailed explanation Figure 4 A flowchart of the project initiation steps.
[0034] Figure 7 For further detailed explanation Figure 6 The flowchart of the project execution steps.
[0035] Figure 8 For further detailed explanation Figure 7 The flowchart shows the steps for registering sequencing work information.
[0036] Figure 9 For illustrative purposes Figure 4 The flowchart shows the sequencing data processing steps.
[0037] Figure 10 For further detailed explanation Figure 9 A flowchart of the data pre-processing steps.
[0038] Figure 11 For a brief explanation Figure 4 The flowchart of the gene insertion site analysis steps.
[0039] Figure 12 For a brief explanation Figure 4 The flowchart for the project completion steps.
[0040] Figure 13a shows the situation in Figure 7A diagram showing the basic information of next-generation sequencing data (NGSData; FastQ) generated during the sequencing job information registration step.
[0041] Figure 13b is a block diagram illustrating the entity-relationship diagram (ERD) used to manage the gene insertion site information database for the next-generation sequencing data (NGS Data; FastQ) in Figure 13a.
[0042] Figure 14 To show in Figure 10 The analysis object is a diagram of the entity group of the entity group in the database of gene insertion site information generated in the data registration step.
[0043] Figure 15a shows a base sequence conversion table database used to convert the base sequence data of Figure 13b into binary numbers without any blanks.
[0044] Figure 15b shows the base sequence conversion table database used to convert the base sequence data of Figure 13b into binary numbers in the presence of blanks. Detailed Implementation
[0045] The following describes several embodiments in detail with reference to the accompanying drawings. However, it should be understood that various modifications can be made to the embodiments; therefore, the scope of protection of this application is not limited to or confined to the embodiments. All modifications, equivalent technical solutions, and alternative technical solutions of the embodiments are within the scope of protection.
[0046] The terminology used in the embodiments is for illustrative purposes only and should not be construed as limiting. Unless the context clearly indicates otherwise, singular expressions include plural expressions. In this specification, terms such as "comprising" or "having" are used only to specify the presence of features, numbers, steps, operations, structural elements, components, or combinations thereof described in this specification, and do not preclude the presence or additional possibilities of one or more other features, numbers, steps, operations, structural elements, components, or combinations thereof.
[0047] Unless otherwise defined, all terms used herein, including technical and scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms as defined in commonly used dictionaries should be interpreted as having the same meaning as they have in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense unless explicitly defined in this specification.
[0048] Furthermore, in the description with reference to the accompanying drawings, the same structural elements are given the same reference numerals regardless of the reference numerals, and repeated descriptions thereof are omitted. In describing embodiments, detailed descriptions of known technologies are omitted when it is determined that such detailed descriptions may unnecessarily obscure the spirit of the invention.
[0049] Furthermore, when describing the structural elements of the present invention, terms such as first, second, A, B, (a), and (b) may be used. These terms are used only to distinguish one structural element from other structural elements, and the nature, order, or sequence of the corresponding structural elements are not limited by the aforementioned terms. When a structural element is "connected," "combined," or "linked" with another structural element, it not only indicates that the structural element can be directly connected or linked to another structural element, but it can also be understood as the various structural elements being "connected," "combined," or "linked" through other structural elements.
[0050] The structural elements included in one embodiment of the present invention, and structural elements having common functions, may be described using the same names in other embodiments. Unless there is a contrary description, the description described in one embodiment of the present invention is also applicable to other embodiments, and specific descriptions within the scope of repetition will be omitted.
[0051] The following is for reference Figures 1 to 1 5b describes a gene insertion site analysis system (hereinafter referred to as the "gene insertion site analysis system") 10 and method for stem cell therapeutic agents containing specific genes according to an embodiment of the present invention. For reference, Figure 1 This is a schematic diagram illustrating the operational concept of a cloud computing-based gene insertion site analysis system 10 for stem cell therapeutic agents containing specific genes, according to an embodiment of the present invention. Figure 2 For brevity Figure 1 The gene insertion site analysis system 10 is a flowchart illustrating a series of business processes for performing gene insertion site analysis, including managing contracts with clients, receiving gene-inserted and passaged stem cells from clients and performing DNA extraction, library construction, sequencing data generation, sequencing data quality management and analysis, gene insertion site analysis, gene insertion site analysis report preparation, and project management. Figure 3a is a structural block diagram illustrating an embodiment of the gene insertion site analysis system 10 of the present invention, Figure 3b is a structural block diagram showing the basic information management module 21 of Figure 3a, and Figure 3c is a structural block diagram showing the basic information database 41 of Figure 3a.
[0052] Reference Figures 1 to 2The Gene Integration Site Analysis System 10 utilizes cloud computing and Software as a Service (SaaS) technologies, enabling numerous researchers both domestically and internationally to use it as a standalone system. Furthermore, the Gene Integration Site Analysis System 10 can perform DNA extraction, library construction, sequencing data processing and analysis, gene integration site analysis, and report preparation based on client requests. This provides comprehensive management of the entire process, including report preparation for Investigational New Drug (IND) applications, specifically for the subculturing of stem cells containing a client-requested gene insertion.
[0053] Referring to Figures 3a to 3c, the gene insertion site analysis system 10 includes a basic information management unit 20, a gene insertion site analysis unit 30, and a database management unit 40.
[0054] The basic information management unit 20 includes: a basic information management module 21, which manages the basic information required to run the gene insertion site analysis system 10. The basic information management module 21 includes: a company information management module 211 for managing the company information database 411; a user information management module 212 for managing the user information database 212; a reference genome information management module 213 for managing the reference genome (human genome (GRCh / hg38)) information database 413; a cancer somatic mutation catalog information management module 214 for managing the cancer somatic mutation catalog (COSMIC, Catalogue of Somatic Mutations in Cancer) information database 414; an equipment information management module 215 for managing the work execution equipment information database 415; a human resources information management module 216 for managing the human resources information database 416; a material information management module 217 for managing the material information database 417; and a standard work information management module 218 for managing various tasks performed to obtain base sequence data of stem cells with inserted specific genes (DNA extraction, library construction, LAM-PCR amplification, next-generation sequencing (NGS)). The database 418 is a standard working information database required for (e.g., Sequencing); the base sequence conversion table management module 219 is used to manage the base sequence conversion table database 419 required to convert base sequence data into binary numbers; the code information management module 21a is used to manage the code information database 41a; and the parameter information management module 21b is used to manage the parameter information database 41b.
[0055] The gene insertion site analysis unit 30 generates an order based on a customer request. After signing the contract, it registers the contract information and begins project management. It receives cells with inserted genes from the customer and performs gene insertion site analysis through DNA extraction, library construction, LAM-PCR amplification, and next-generation sequencing analysis. A final report is then generated and submitted to the customer for project completion. The gene insertion site analysis unit 30 includes: an order information management module 31 for managing various order information (customer orders, DNA extraction work orders, library construction orders, sequencing orders); a contract information management module 32 for managing contract information with customers; a project management module 33 for managing project information required to fulfill customer contracts; a gene insertion site analysis module 34 for analyzing gene insertion sites using sequencing data as the result of sequencing orders; and a project outcome management module 35 for submitting analysis results to the customer and terminating the project.
[0056] The order information management module 31 registers order information, work orders (including self-manufacturing and outsourced processes such as DNA extraction, library construction, LAM-PCR amplification, and next-generation sequencing), and purchase orders for necessary materials in the order information database 42 according to the contract terms agreed upon with the client. The contract information management module 32 registers contract information agreed upon with the client in the contract information database 43. The project management module 33 registers client order information and project information related to the contract information in the project information database 44. The gene insertion site analysis module 34 extracts next-generation sequencing data and data used for gene insertion site analysis and registers them in the gene insertion site information database 45. The project results management module 35 registers project progress information and gene insertion site analysis results in the project progress information database 46.
[0057] The database management unit 40 includes a basic information database 41, an order information database 42, a contract information database 43, a project information database 44, a gene insertion site information database 45, and a project progress information database 46.
[0058] The basic information database 41 includes a company information database 411, a user information database 412, a reference genome information database 413, a cancer somatic mutation catalog information database 414, an equipment information database 415, a human resources information database 416, a material information database 417, a standard operating procedure information database 418, a code information database 41a, and a parameter information database 41b, enabling the database management system (DBMS) to manage the data generated or referenced by each module of the basic information management unit 20 and the gene insertion site analysis unit 30.
[0059] The company information database 411 includes the business registration number or fixed number, serial number, organization name, representative telephone number, fax number, address, and representative name of the operating company and its clients (enterprises, public institutions, universities, etc.). The user information database 412 includes name, password, contact information (mobile phone, office, fax, etc.), email address, job code, access code, email receipt status, and SMS message receipt status. The reference genome information database 413 includes information on the types and functions of human genes identified through the Human Genome Project (HGP), which can be referenced from the database managed by the National Center for Biotechnology Information (NCBI), or by copying and storing it on a local computer. The cancer somatic mutation catalog database 414 includes gene name, Entrez database identification number, gene locus (start and end sites of the chromosome insertion site), and information on its role in cancer. The equipment information database 415 includes equipment number, equipment name, equipment specifications, equipment manufacturer, purchase price, and hourly usage price. The Human Resources Information Database 416 includes information such as the required skills of the staff performing the work and the hourly usage price. The Materials Information Database 417 includes material numbers, material names, input units, unit prices, and supplier codes corresponding to the material information required by the Standard Work Information Database 418. The Standard Work Information Database 418 includes various work information required for managing and analyzing gene insertion site information, such as work numbers, work names, work hours, equipment information for performing the work, required material information, required skills of the staff performing the work, and preceding and following work information. The base sequence conversion table database 419 converts 4 bytes (32 bits) into variable-bit (1 to 8 bits) bytes to efficiently store the base sequence of the genome (a nucleobase (A (adenine), T (thymine), G (guanine), C (cytosine)) of nucleotides, the basic unit of DNA, as a conversion table. The code information database 41a includes standard work information numbers, standard equipment usage time, standard material codes, standard material consumption, standard labor cost codes, standard working hours, and standard work unit price information. The code information database includes various types of code information.The parameter information database 41b includes information such as noise baseline value, total interval of the same gene, sequencing analysis volume, cell type, infecting virus type, inserted gene type, mapping error (noise data) range before and after gene insertion site, specified range value when the transcription site of the same gene is within the specified range, and sequencing analysis volume value (1G, 3G, 10G, etc.) when pooled samples are treated as objects.
[0060] The following is for reference Figures 4 to 1 5b describes a method for analyzing the gene insertion site analysis system 10 of this invention for each passaged culture and each chromosome of a virus-infected stem cell therapeutic agent in which a specific gene has been inserted. For reference, Figure 4 The flowchart below briefly illustrates the method of using the gene insertion site analysis system 10 shown in Figure 3a to perform gene insertion site analysis on stem cells with specific genes requested by a customer and submit the results to the customer. Figure 5 For illustrative purposes Figure 4 The flowchart of the basic information management steps 100. Figure 6 For illustrative purposes Figure 4 The flowchart for the project start step 200. Figure 7 For illustrative purposes Figure 6 The flowchart of project execution steps 250 Figure 8 For illustrative purposes Figure 7 The flowchart for step 254 of the sequencing work information registration. Figure 9 For illustrative purposes Figure 4 The flowchart of sequencing data processing steps 300, Figure 10 For illustrative purposes Figure 9 The flowchart for data pre-processing step 310. Figure 11 For illustrative purposes Figure 4 The flowchart of step 400 for gene insertion site analysis. Figure 12 For illustrative purposes Figure 4 The flowchart for the project completion step 500. Furthermore, Figure 13a shows the process... Figure 7 Figure 13b is a diagram showing the basic information of the next-generation sequencing data (FastQ) generated in the sequencing job information registration step 254. It is a block diagram of the entity-relationship diagram (ERD) used to manage the gene insertion site information database of the next-generation sequencing data (FastQ) in Figure 13a. Figure 14 To show in Figure 10The analysis object is a block diagram of the entity group model of the entity group of the gene insertion site information database generated in step 314. Figures 15a and 15b are base sequence conversion table databases used to convert the base sequence data of Figure 13b into binary numbers. Figure 15a is the database without blanks, and Figure 15b is the database with blanks.
[0061] First, in the basic information management step 100, the basic information is recorded in various databases.
[0062] Specifically, refer to Figure 5 In the basic information management step 100, when the step is divided into a preparation step, the company information of the company running the gene insertion site analysis system 10 is registered in the company information database 411 through the company information management module 211 (111), the user information belonging to the operating company is registered in the user information database 412 through the user information management module 212 (112), the reference genome information is registered in the reference genome information database 413 through the reference genome information management module 213 (113), the cancer somatic mutation catalog information including the cancer somatic mutation catalog is registered in the cancer somatic mutation catalog information database 414 through the cancer somatic mutation catalog information management module 214 (114), and the equipment information for performing various tasks to obtain the base sequence data of stem cells with specific inserted genes is registered in the equipment information database 415 through the equipment information management module 215. (115) Register the required human resources information (required capacity, hourly unit price, etc.) for performing the corresponding work in the human resources information database 416 through the human resources information management module 216 (116). Register the required material information for performing the corresponding work in the material information database 417 through the material information management module 217 (117). Register the standard work information for the corresponding work in the standard work information database 418 through the standard work information management module 218 (118). Register the base sequence conversion table used to compress base sequence information in the base sequence conversion table database 419 through the base sequence conversion table management module 219 (119). Register the code information required for running the system in the code information database 41a through the code information management module 21a (11a). Register the various parameters required for running the system in the parameter information database 41b through the parameter information management module 21b (11b).
[0063] Furthermore, when the steps are divided into operational steps, if new customers are discovered, customer information is registered in the company information database 411 through the company information management module 211 (11c), user information belonging to the customer company is registered in the user information database 412 through the user information management module 212 (11d), and parameter information that meets the customer's requirements is registered in the parameter information database 41b through the parameter information management module 21b (11e).
[0064] Refer again Figure 4 After the basic information management step 100, proceed to the project start step 200.
[0065] Reference Figure 6 In the project initiation step 200, the agreement process with the client is registered in the project information database 44 using the estimation management 210. When the agreement is signed and the result order is received, the order information is registered in the order information database 42 (220), the contract terms based on the order information are registered in the contract information database 43 (230), and the project information is registered in the project progress database 44 for the client's successor training management (240), thereby starting the project execution step 250.
[0066] Moreover, refer to Figure 7 In project execution step 250, the cells and related information of the customer's passage culture are received and registered in the project progress information database 46 (251). For DNA extraction work, DNA extraction work order information is registered in the order information database 42 and the project progress information database 46 (252). For library construction work, library construction work order information is registered in the order information database 42 and the project progress information database 46 (253). For sequencing work, sequencing work information is registered in the order information database 42 and the project progress information database 46 (254).
[0067] Furthermore, referring to Figure 8 In the sequencing work information registration step 254, during the process of the next-generation sequencing staff reading the sequencing data (FastQ) submitted together with the work results from beginning to end (2541), if the last recorded sequencing data is read, the sequencing work result information is registered in the gene insertion site information database 45 and the project progress information database 46 (2548).
[0068] During the sequential reading of sequencing data (2541), if the sequence data record read is a multiple of 4 + 1 and belongs to the first record (n=0), it is registered in the header table (FastQHeader) 1302, corresponding to the leader of the sequencing data (the first row of Figure 13b). Therefore, the inherent value (fixed information) of the header information is extracted and registered in the fixed information table (FastQ_Line#1-Overview) 1303 (2542) of the gene insertion site information database 45. The variable value (variable information) of the header information is registered in the variable information table (FastQ_Line#1-Detail) 1304 (2543) of the gene insertion site information database 45, and then the next record is read (2541).
[0069] Furthermore, during the sequential reading of sequencing data (2541), if the sequence data record read is a multiple of 4 + 2, it belongs to base sequence data. Therefore, the corresponding value (a variable binary number, 1 to 8 bits) can be read from the Reference Sequence 1301 of the base sequence conversion table database 419 in 4 bytes (4 bytes) to confirm the value. For example, referring to Figure 15a, the binary number is 0 for the base sequence AAAA, 1 for the base sequence AAAAC, 1111110 for the base sequence TTTG, and 111111111 for the base sequence TTTT. After the confirmed value is registered in the base information table (FastQ_Line#2_Detail-Char) 1305 of the gene insertion site information database 45, 4-byte processing is performed. When processed in the above manner, as shown in Figure 13b, after reading 151 bytes, 148 bytes (1184 bits) are converted into site information (84 bits) and conversion information (261 bits), achieving a compression rate of 71.9%. Furthermore, if there are 1 to 3 blanks in the data read in 4-byte increments, the corresponding value (a variable binary number, 1 to 8 bits) is identified in the reference sequence table 1301 of the base sequence conversion table database 419, and assigned a corresponding binary number with reference to the base sequence conversion table in Figure 15b. For example, in the case of AAAb, the binary number is 0; in the case of AACb, the binary number is 1; in the case of TTTb, the binary number is 111111; in the case of AAbb, the binary number is 1000000; and in the case of Tbbb, the binary number is 1010011. As shown in Figure 13b, the remaining 3 bytes (1 blank byte, 4 bytes 32 bits), site information (6 bits), and conversion information (6 bits) are managed in 12 bits, thus achieving a compression rate of 62.5%. After the value confirmed in the above manner is registered in the base information table (FastQ_Line#2_Detail-Blank) 1306 of the gene insertion site information database 45, 4 bytes of processing are performed. Furthermore, if the read data includes N or U, it is considered an error message; therefore, it is registered in the base information-error table (FastQ_Line#2_Detail-Not Match) 2545 of the gene insertion site information database 45, and then the next record (2541) is read. In the case shown in Figure 13b, no error data can be registered.
[0070] Furthermore, during the sequential reading of sequencing data (2541), if the sequence data record read is a multiple of 4 + 3, it indicates a simple ligation (value "+"), and then the next record is read (2546).
[0071] Furthermore, during the sequential reading of sequencing data (2541), if the sequence data record read is a multiple of 4 + 4, it represents a quality score value. This value is then compressed using data compression technology (Huffman Coding Method) and registered in the base information-quality score table (FastQ_Line#4_Quality Score) 1308 (2547) of the gene insertion site information database 45, before the next record (2541) is read. For example, as shown in Figure 13b, when converted using the Huffman Coding Method, 151 bytes (1208 bits) are compressed into 39 bytes and 55 bits (7 bytes), achieving a compression rate of 74.2%.
[0072] Refer again Figure 4 After step 200 at the start of the project, step 300, sequencing data processing, is performed.
[0073] Reference Figure 9 The sequencing data processing step 300 includes a data pre-processing step 310, a data structuring step 320, a data annotation and registration step 330, and a data extraction and registration step 340 for analysis.
[0074] Reference Figure 10 The data preprocessing step 310 includes a useless data deletion step 311, a reference data mapping step 312, a duplicate data deletion step 313, and an analysis object data registration step 314.
[0075] In this embodiment, sequencing data processing step 300 refers to a standard genomic base sequence analysis process, used to analyze the degree to which genome expression is affected by specific environmental factors and to analyze genes that cause diseases. However, in order to analyze stem cells with specific genes inserted for functional enhancement, it is necessary to perform insertion site correlation analysis on each passaged culture and chromosome gene. Therefore, the useless data deletion step 311 and the duplicate data deletion step 313 in data preprocessing step 310 can be omitted.
[0076] Furthermore, in the reference data mapping step 312, while performing the sequencing job information registration step 254 in reverse order, the binary portion of the data (4n+2 lines) is converted into raw data, and the converted raw data is mapped into the reference genome information database 413 to confirm the gene information (312). Moreover, duplicate data is deleted from the confirmed gene information (313), or, if not deleted, the analysis object is registered in the analysis object data table of the gene insertion site information database 45 (314).
[0077] In the data structuring step 320, when the data is distributed before and after the specific gene insertion site (Integration Site) of a specific gene, the merging step (320) is performed.
[0078] In the data annotation and registration step 330, tasks such as association with neighboring genes or gene ontology analysis, genome feature association analysis, and association of peak values with gene expression data are performed.
[0079] In the data extraction and registration step 340 for analysis, data for gene insertion site analysis is generated and registered in the analysis object data table of the gene insertion site information database 45.
[0080] Refer again Figure 4 After sequencing data processing step 300, gene insertion site analysis step 400 is performed.
[0081] Reference Figure 11 The gene insertion site analysis step 400 includes a search condition setting step 410, an analysis method selection step 420, an analysis report confirmation step 430, and an analysis result registration step 440.
[0082] In the search criteria setting step 410, condition values are set for searching the gene insertion site information database 45 generated in the sequencing data processing step 300. That is, search criteria are set in order to retrieve customer information, project information, and passage culture information according to the previous passage number or a specific passage number.
[0083] Next, after searching the data using the set search criteria, the analysis method is determined (420). The analysis results are then registered in the project progress information database 46 by executing the determined analysis method and confirming the search execution report (430).
[0084] The process of analyzing gene insertion sites for each passaged culture is as follows: mesenchymal stem cells collected simultaneously from bone marrow and mesenchymal stem cells with multiple inserted genes are passaged under the same conditions, and expression information is compared and analyzed using t-value technology. The expression information includes the total expression level of each passage, the total number of genes, the expression level of each biotype, the expression level of each chromosome, the number of genes on each chromosome, the expression level of each chromosome and biotype, the expression level of each gene, the expression level of the inserted gene, the expression level of genes adjacent to the inserted gene, and the expression level of the inserted gene. The average values between the two sample groups are compared using t-value technology, as shown in the following formula (1):
[0085] Equation (1)
[0086]
[0087] Where t is a statistical indicator of the sample mean difference. The mean difference between the two sample groups. This represents the uncertainty of the average difference between the two sample groups.
[0088] Uncertainty can be represented by the following equation (2):
[0089] Equation (2)
[0090]
[0091] Where s1 and s2 are the standard deviations of each sample, and n1 and n2 are the number of each sample.
[0092] Therefore, based on equations (1) and (2), the following equation (3) can be expressed:
[0093] Equation (3)
[0094]
[0095] On the other hand, if we assume that n1 and n2 of the two sample groups are the same and have the same variance, then equation (3) can be expressed by the following equation (4):
[0096] Equation (4)
[0097]
[0098] Among them, s p The pooled standard deviation is... express.
[0099] Furthermore, if we assume that n1 and n2 of the two sample groups are different and have the same variance, then equation (3) can be expressed by equation (5):
[0100] Equation (5)
[0101]
[0102] in, .
[0103] Refer again Figure 4 After the gene insertion site analysis step 400, proceed to the project end step 500.
[0104] Reference Figure 12 The project completion step 500 includes step 510 of submitting the analysis report, step 520 of registering customer feedback, step 430 of calculating total input costs and collecting payments, and step 540 of project completion processing.
[0105] In step 510 of the analysis report submission, a gene insertion site analysis report for stem cell therapeutic agents containing specific genes is prepared and submitted to the client by referring to the project progress information database 46 registered in step 400 of the gene insertion site analysis and in accordance with the PDF file.
[0106] Next, customer feedback is recorded in the project progress information database 46 upon receiving feedback on the analysis report (520). Furthermore, if customer feedback is received, the total input cost is calculated by referring to the work order database 42 (order information 42) to confirm working hours, required materials, manpower, and time information, and then recorded in the project progress information database 46 to request payment from the customer (430). Additionally, a loss is calculated by comparing standard cost information, the total input cost calculation, and the total input cost calculated in the payment collection step, and then recorded in the project progress information database 46, and the project processing is terminated (540).
[0107] While embodiments have been described above with reference to the limiting drawings, those skilled in the art can make various technical modifications and variations based on the above description. For example, even if the described techniques are performed in a different order than the described methods, and / or the described systems, structures, devices, circuits, and other structural elements are combined or integrated in a different manner than the described methods, or even if they are replaced or substituted by other structural elements or equivalent technical solutions, appropriate results can still be achieved.
[0108] Therefore, other implementation methods, other embodiments, and contents equivalent to the scope of protection claimed in the invention also fall within the protection scope of the present invention.
Claims
1. A gene insertion site analysis system for a stem cell therapeutic agent into which a specific gene is inserted, characterized by, include: The basic information management unit is used to manage the code information, equipment information, staff information, work information, customer information, cooperative enterprise information, parameter information, reference gene information, oncogene information, and base sequence conversion table information required for running the gene insertion site analysis system for stem cell therapeutic agents containing specific genes; The gene insertion site analysis unit is used to manage the order information, contract information, project information, gene insertion site analysis information and project execution result information required by the above gene insertion site analysis system. The order information includes customer orders, work instructions and material order information. as well as The database management unit manages the basic information database, order information database, contract information database, project information database, gene insertion site information database, base sequence analysis information database, and project progress information database generated or referenced by the aforementioned basic information management unit and gene insertion site analysis unit. The aforementioned basic information databases include company information database, user information database, reference genome information database, cancer somatic mutation catalog information database, equipment information database, human resources information database, material information database, standard working table information database, base sequence conversion table database, code information database, and parameter information database. The gene insertion site analysis unit includes: The order information management module registers the agreement process with the customer in the project information database, registers the order information in the order information database, registers the contract terms based on the order information in the contract information database, and registers the sequencing work information used to register the project information in the project progress database. as well as The gene insertion site analysis module is used to manage gene insertion site information. The gene insertion site analysis unit is configured as follows: NGS sequencing data and data for gene insertion site analysis are generated. The correlation between genes in each passaged culture and on each chromosome is compared and analyzed using t-value technology, and the data is registered in the gene insertion site information database. The order information management module is configured as follows: The sequencing data is read in multiples of 4 plus 1 records. Repeated data are recorded as a single intrinsic value, and variable values are separated and recorded in additional records, thereby performing the registration of sequencing job information. If the value read from the above sequencing data in units of 4 bytes contains N or U, or if the corresponding value is not found in the above base sequence conversion table database, the value will be stored in the base information error table of the gene insertion site information database.
2. The gene insertion site analysis system for stem cell therapeutic agents with specific genes inserted according to claim 1, characterized in that, The aforementioned basic information management unit includes a basic information management module, which includes: The company information management module is used to manage company information; The user information management module is used to manage the user information database; The reference genome information management module is used to manage the reference genome information database; The module for managing information on somatic mutations in cancer is used to manage the database of information on somatic mutations in cancer. The equipment information management module is used to manage the equipment information database; The Human Resources Information Management module is used to manage the human resources information database; The material information management module is used to manage the material information database; The standard work information management module is used to manage the standard work information database; The base sequence conversion table management module is used to manage the base sequence conversion table database; The code information management module is used to manage the code information database; and The parameter information management module is used to manage the parameter information database.
3. The gene insertion site analysis system for a stem cell therapeutic agent into which a specific gene is inserted according to claim 1, characterized by, The above-mentioned gene insertion site analysis unit also includes: The order information management module is used to manage order information; The contract information management module is used to manage contract information; The project management module is used to manage project information; and The project deliverables management module is used to manage project progress information.
4. A method for analyzing a gene insertion site of a stem cell therapeutic agent into which a specific gene is inserted, characterized by, Includes the following steps: The basic information management steps involve recording the basic information required to run a gene insertion site analysis system for stem cell therapeutic agents containing specific genes in a database; The sequencing job information registration step and the project start step are as follows: In the sequencing job information registration step, the agreement process with the client is registered in the project information database, the order information is registered in the order information database, the contract terms based on the order information are registered in the contract information database, and the project information is registered in the project progress database; In the project start step, the project is executed. The sequencing data processing steps include data preprocessing, data structuring, data annotation and registration, and data extraction and registration for analysis performed on the sequencing data generated in the above project execution steps. The gene insertion site analysis steps include steps for setting search criteria for the analysis target, selecting the analysis method, confirming the analysis report, and registering the analysis results in the database; and... The project closure steps include preparing an analysis report to be submitted to the client, recording client feedback, calculating the total cost of investment, and collecting payment to close the project. In the sequencing job information registration step, Read the sequencing data at multiples of 4 plus 1, register the duplicate data as an intrinsic value, and separate the variable values and register them in additional records. In the gene insertion site analysis step, data for the gene insertion site analysis is generated, and each passaged culture is analyzed to register its data in the gene insertion site information database. The t-value technique is then used to compare and analyze the correlations between genes in each passaged culture and on each chromosome. If the value read from the above sequencing data in units of 4 bytes contains N or U, or if the corresponding value is not found in the above base sequence conversion table database, the value will be stored in the base information error table of the gene insertion site information database.
5. The method for analyzing the gene insertion site of a stem cell therapeutic agent with a specific gene inserted according to claim 4, characterized in that, In the above sequencing work information registration steps, when the last record is read during the process of reading sequencing data from beginning to end, the sequencing results information is registered in the gene insertion site information database and the project progress information database.
6. The method for analyzing the gene insertion site of a stem cell therapeutic agent with a specific gene inserted according to claim 5, characterized in that, In the above sequencing work information registration step, the records of the above sequencing data that are multiples of 4 + 2 are read in units of 4 bytes to convert the corresponding values in the base sequence conversion table database into binary numbers.
7. The method for analyzing the gene insertion site of a stem cell therapeutic agent with a specific gene inserted according to claim 5, characterized in that, In the above sequencing work information registration step, the sequencing data that is a multiple of 4 + 4 is compressed as a quality score and stored in the base information quality score table of the gene insertion site information database.
8. The method for analyzing the gene insertion site of a stem cell therapeutic agent with a specific gene inserted according to claim 4, characterized in that, In the gene insertion site analysis step described above, mesenchymal stem cells collected simultaneously from bone marrow and mesenchymal stem cells with multiple inserted genes were passaged and cultured under the same conditions, and the expression level information was compared and analyzed using the t-value technique.
9. The method for analyzing the gene insertion site of a stem cell therapeutic agent with a specific gene inserted according to claim 8, characterized in that, For the above expression level information, we compared and analyzed the total expression level of each generation, the total number of genes, the expression level of each biotype, the expression level of each chromosome, the number of genes on each chromosome, the expression level of each chromosome and biotype, the expression level of each gene, the expression level of the inserted gene, the expression level of genes adjacent to the inserted gene, and the expression level of the inserted gene.
Citation Information
Patent Citations
KR20200098189A