Method, device and electronic equipment for processing genetic data
Patent Information
- Application Number
- CN202510702278.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2045-05-28
AI Technical Summary
[0005]本申请提供一种基因数据的处理方法、装置以及电子设备,以解决相关技术中对样本对象的基因数据进行统计分析的效率较低的问题
[0016]通过本申请,采用以下步骤:获取目标种类下的M个样本对象的基因数据集合,得到M个基因数据集合,其中,每个基因数据集合中包含N个基因的基因数据,M、N为正整数;构建M个样本对象和N个基因之间的关系矩阵,并根据基因统计指标中的各个子指标判断是否需要对关系矩阵进行分割,其中,基因统计指标中包括T个子指标,每个子指标用于指示对关系矩阵中的样本数据进行统计的分组要求;对于需要对关系矩阵进行分割的子指标,根据子指标对关系矩阵进行分割,得到P个子矩阵,并根据子指标对各个子矩阵中的基因数据进行并行处理,得到P个数据统计结果,并将P个数据统计结果进行组合,得到目标种类和子指标下的基因数据统计结果,其中,P为正整数;对于不需要对关系矩阵进行分割的子指标,根据子指标对关系矩阵进行处理,得到目标种类和子指标下的基因数据统计结果;将各个子指标下的基因数据统计结果进行组合,得到目标种类下的基因数据统计总结果。解决了相关技术中对样本对象的基因数据进行统计分析的效率较低的问题。通过构建样本对象和基因之间的关系矩阵,并根据基因统计指标中的子指标确定是否需要对关系矩阵进行分割,从而在样本对象和基因的数据量较多的情况下,可以通过将关系矩阵进行分割,并分别对各个子矩阵进行处理的方式,将处理后的子矩阵的统计结果组合为关系矩阵的统计结果,进而达到了提高基因数据的处理效率的技术效果。
Smart Images

Figure CN120673855B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data, and more specifically, to a method, apparatus, and electronic device for processing gene data. Background Technology
[0002] In the field of bioinformatics research, pangenome analysis is a crucial task. It involves comparing the genomes of samples from multiple species or strains to identify shared core genes and unique pangeny, which is of great significance for understanding genome diversity, evolutionary relationships, and species adaptability.
[0003] Currently, when comparing genetic data from sample subjects, the processing of large-scale genetic data mostly adopts a direct, holistic approach, analyzing all the genetic data to be compared at once. However, with the rapid development of gene sequencing technology, since each sample subject represents an independent genome, and the amount of genetic data is growing exponentially, the sheer number of genes and the complex permutations and combinations between sample subjects make direct processing an extremely time-consuming task. This approach is not only computationally inefficient but also often limited by the hardware performance of a single computer, making it difficult to complete pan-genome set analysis of tens of thousands of sample combinations and sample data from multiple samples in a short period of time. This seriously affects the work efficiency of researchers and the timeliness of research progress.
[0004] There is currently no effective solution to the problem of low efficiency in statistical analysis of genetic data of sample subjects in related technologies. Summary of the Invention
[0005] This application provides a method, apparatus, and electronic device for processing gene data to solve the problem of low efficiency in statistical analysis of gene data of sample objects in related technologies.
[0006] According to one aspect of this application, a method for processing genetic data is provided. The method includes: acquiring gene datasets of M sample objects under the target category, resulting in M gene datasets, where each gene dataset contains gene data of N genes, and M and N are positive integers; constructing a relationship matrix between the M sample objects and the N genes, and determining whether the relationship matrix needs to be segmented based on each sub-indicator in the gene statistical indicators, where the gene statistical indicators include T sub-indicators, each sub-indicator indicating the grouping requirements for statistical analysis of the sample data in the relationship matrix, and T is a positive integer; for sub-indicators that require segmentation of the relationship matrix, segmenting the relationship matrix according to the sub-indicators to obtain P sub-matrices, and performing parallel processing on the gene data in each sub-matrices according to the sub-indicators to obtain P data statistical results, and combining the P data statistical results to obtain the gene data statistical results under the target category and sub-indicators, where P is a positive integer; for sub-indicators that do not require segmentation of the relationship matrix, processing the relationship matrix according to the sub-indicators to obtain the gene data statistical results under the target category and sub-indicators; and combining the gene data statistical results under each sub-indicator to obtain the overall gene data statistical result under the target category.
[0007] Optionally, obtaining the gene data set of M sample objects under the target category includes: obtaining the category information of each initial object from the object database, and filtering the category information to the target category of the initial object to obtain M sample objects, wherein the object database contains multiple test objects; detecting the occurrence frequency of each gene carried in each sample object, and determining the occurrence frequency as the gene data of each gene to obtain the gene data set of each sample object.
[0008] Optionally, determining whether the relation matrix needs to be segmented based on each sub-indicator in the gene statistical indicators includes: for any sub-indicator, determining the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; determining the number of groups of sample objects under the grouping requirements based on the target number and the total number of sample objects, and determining a preset threshold; calculating the quotient between the number of groups and the preset threshold, and rounding the quotient up to obtain the target value; determining whether the target value is greater than the preset value, and if the target value is greater than the preset value, determining that the relation matrix needs to be segmented; if the target value is less than or equal to the preset value, determining that the relation matrix does not need to be segmented.
[0009] Optionally, the relationship matrix is segmented according to the sub-indicators to obtain P sub-matrices, including: determining the target value as the number of sub-matrices after segmenting the relationship matrix, and obtaining the value of P; segmenting the relationship matrix along the gene dimension according to the target value to obtain P sub-matrices, wherein the number of sample objects in each sub-matrice is the same, and the genes contained in each sub-matrice are different.
[0010] Optionally, the gene data in each submatrix is processed in parallel according to the sub-indicators to obtain P statistical results, including: for any submatrix, determining the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicators; traversing and grouping the sample objects in the submatrix according to the target number to obtain H groups of sample objects, where H is a positive integer; obtaining the gene data under each group of sample objects from the submatrix to obtain H matrices to be processed, and obtaining the first number of core genes and the second number of pan-genes in each matrix to be processed to obtain the first number of H groups and the second number of H groups; calculating the first statistical result of the first number of H groups and the second statistical result of the second number of H groups, wherein the first statistical result includes at least one of the following: the maximum value, minimum value and average value of the first number of H groups, and the second statistical result includes at least one of the following: the maximum value, minimum value and average value of the second number of H groups; and determining the first statistical result and the second statistical result as a single statistical result.
[0011] Optionally, combining the P statistical results to obtain the gene data statistical results under the target category and sub-indicators includes: obtaining the maximum value of the first statistical result in the statistical results of each sub-matrix, obtaining multiple first maximum values, and determining the maximum value among the multiple first maximum values as the target first maximum value; obtaining the maximum value of the second statistical result in the statistical results of each sub-matrix, obtaining multiple second maximum values, and determining the maximum value among the multiple second maximum values as the target second maximum value; obtaining the minimum value of the first statistical result in the statistical results of each sub-matrix, obtaining multiple first minimum values, and determining the minimum value among the multiple first minimum values as the target first minimum value; obtaining the data of each sub-matrix... Based on the minimum value of the second statistical result in the statistical results, multiple second minimum values are obtained, and the minimum value among these multiple second minimum values is determined as the target second minimum value; the average value of the first statistical result in the statistical results of each submatrix is obtained, multiple first average values are obtained, and the average of these multiple first average values is calculated to obtain the target first average value; the average value of the second statistical result in the statistical results of each submatrix is obtained, multiple second average values are obtained, and the average of these multiple second average values is calculated to obtain the target second average value; the target first maximum value, target second maximum value, target first minimum value, target second minimum value, target first average value, and target second average value are determined as the gene data statistical results.
[0012] Optionally, combining the gene data statistical results under each sub-indicator to obtain the total gene data statistical result for the target category includes: adding the gene data statistical results under each sub-indicator to a preset schematic diagram according to the grouping requirements in the sub-indicator to obtain a gene data statistical graph; drawing a curve based on the average value of the gene data statistical results under each sub-indicator and calculating the slope of the curve at each point to obtain multiple slopes; determining the validity result of the gene data statistical graph based on the slope values of the multiple slopes, and determining the gene data statistical graph and the validity result as the total gene data statistical result.
[0013] According to another aspect of this application, a gene data processing apparatus is provided. The apparatus includes: an acquisition unit, configured to acquire gene data sets of M sample objects under a target category, resulting in M gene data sets, wherein each gene data set contains gene data of N genes, and M and N are positive integers; a judgment unit, configured to construct a relationship matrix between the M sample objects and the N genes, and determine whether the relationship matrix needs to be segmented based on various sub-indicators in the gene statistical indicators, wherein the gene statistical indicators include T sub-indicators, each sub-indicator indicating the grouping requirements for statistical analysis of the sample data in the relationship matrix; and a first processing unit, configured to process the sub-indicators that require segmentation of the relationship matrix. The relation matrix is divided into P sub-matrices based on sub-indicators. The gene data in each sub-matrice is then processed in parallel based on the sub-indicators to obtain P statistical results. These P statistical results are then combined to obtain the gene data statistical results for the target category and sub-indicators, where P is a positive integer. The second processing unit processes the relation matrix based on the sub-indicators for those sub-indicators that do not require partitioning, obtaining the gene data statistical results for the target category and sub-indicators. The combination unit combines the gene data statistical results for each sub-indicator to obtain the total gene data statistical result for the target category.
[0014] According to another aspect of the present invention, a computer program product is also provided, comprising a computer program that, when executed by a processor, implements a gene data processing method provided in the foregoing embodiments of the present application.
[0015] According to another aspect of the present invention, an electronic device is also provided, comprising one or more processors and a memory; the memory stores computer-readable instructions, and the processor is configured to execute the computer-readable instructions, wherein the computer-readable instructions, when executed, perform a gene data processing method provided in the foregoing embodiments.
[0016] This application employs the following steps: First, obtain a set of gene data for M sample objects under the target category, resulting in M gene data sets, each containing gene data for N genes, where M and N are positive integers. Second, construct a relationship matrix between the M sample objects and the N genes, and determine whether the relationship matrix needs to be segmented based on each sub-indicator in the gene statistical indicators. The gene statistical indicators include T sub-indicators, each indicating the grouping requirements for statistical analysis of the sample data in the relationship matrix. Third, for sub-indicators requiring segmentation, segment the relationship matrix according to the sub-indicators to obtain P sub-matrices, and process the gene data in each sub-matrice in parallel according to the sub-indicators to obtain P statistical results. Combine these P statistical results to obtain the gene data statistical results for the target category and its sub-indicators, where P is a positive integer. Fourth, for sub-indicators not requiring segmentation, process the relationship matrix according to the sub-indicators to obtain the gene data statistical results for the target category and its sub-indicators. Finally, combine the gene data statistical results for each sub-indicator to obtain the overall gene data statistical result for the target category. This invention addresses the low efficiency of statistical analysis of gene data from sample objects in related technologies. By constructing a relationship matrix between sample objects and genes, and determining whether to segment the matrix based on sub-indicators in gene statistical indicators, the invention improves the processing efficiency of gene data, especially when dealing with large amounts of sample object and gene data. This is achieved by segmenting the relationship matrix, processing each sub-matrix separately, and then combining the statistical results of the processed sub-matrixes into the statistical result of the relationship matrix. Attached Figure Description
[0017] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0018] Figure 1 This is a flowchart of a gene data processing method provided in the embodiments of this application;
[0019] Figure 2 This is a schematic diagram illustrating the optional statistical summary results provided according to an embodiment of this application;
[0020] Figure 3 This is a schematic diagram of a gene data processing apparatus provided according to an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0022] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present application.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] It should be noted that the gene data processing methods, devices, and electronic devices defined in this disclosure can be used in the field of big data, or in any field other than big data. The application fields of the gene data processing methods, devices, and electronic devices defined in this disclosure are not limited.
[0026] It should be noted that all information, user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data) used in this application are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with the relevant laws, regulations, and standards of the relevant regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse use. If the user chooses to refuse, the process proceeds to the expert decision-making process. For example, this system has interfaces with relevant users or institutions. Before obtaining relevant information, a request to obtain the information needs to be sent to the aforementioned user or institution through the interface, and the relevant information is obtained only after receiving consent from the aforementioned user or institution.
[0027] The embodiments or examples disclosed herein are not exhaustive, but merely illustrative of some embodiments or examples, and are not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment or example can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment or example can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment or example can be arbitrarily interchanged. Furthermore, optional methods or examples in a particular embodiment or example can be arbitrarily combined; moreover, embodiments or examples can be arbitrarily combined. For example, some or all steps of different embodiments or examples can be arbitrarily combined, and a particular embodiment or example can be arbitrarily combined with optional methods or examples of other embodiments or examples.
[0028] For ease of description, the following explains some of the nouns or terms used in the embodiments of this application:
[0029] Core genes: These are genes present in all samples and represent conserved genetic units necessary for a species to maintain basic survival and reproduction.
[0030] Pangenome: refers to the collection of genes of all individuals within a species, which, in addition to the core genes, also includes genes that are unique to all individuals or shared by some individuals. This part of the genes constitutes the pangenome of the species.
[0031] Gene family: A gene family is a group of genes that have similar sequences, structures or functions. These genes usually evolved from a common ancestral gene through gene duplication events.
[0032] According to embodiments of this application, a method for processing gene data is provided.
[0033] Figure 1 This is a flowchart of a gene data processing method provided according to an embodiment of this application. For example... Figure 1 As shown, the method includes the following steps:
[0034] Step S101: Obtain the gene data set of M sample objects under the target category, resulting in M gene data sets, where each gene data set contains gene data of N genes, and M and N are positive integers.
[0035] It should be noted that the execution entity in this embodiment can be a data processing system, the target type refers to the specific biological population to be analyzed, the sample object represents the sample individuals selected from that population, each individual has a genome, and the gene represents the genes (or gene families) contained in the genome of each individual. The gene dataset covers the gene information of each gene of each sample object, including gene sequences, gene annotations, etc.
[0036] Specifically, when processing genetic data, the first step is to obtain all the basic data required for processing. This means first identifying the target category for genetic data analysis and obtaining M representative samples within that category. Then, the genetic data for each sample—that is, information on all N genes for each sample—can be collected. This genetic data can be obtained from various sources, such as public databases or laboratory-generated sequencing data. After data collection, quality control and preprocessing can be performed to ensure data consistency and accuracy, laying the foundation for subsequent analysis.
[0037] Step S102: Construct a relationship matrix between M sample objects and N genes, and determine whether the relationship matrix needs to be segmented based on each sub-indicator in the gene statistical indicators. The gene statistical indicators include T sub-indicators, each of which indicates the grouping requirements for statistical analysis of the sample data in the relationship matrix.
[0038] It should be noted that the relation matrix is a data structure that displays the association between M samples and N genes in a two-dimensional table. The gene statistical index includes multiple sub-indicators. Each sub-indicator is used to indicate the number of sample objects in each group when combining sample data. For example, if the gene statistical index is (2, 29), it indicates that there are 28 sub-indicators, indicating that the sample objects need to be grouped in pairs, then in groups of three, and finally in groups of 29, thus obtaining the number of groups from C(29, 2) to C(29, 29).
[0039] Specifically, after obtaining the basic data, a relationship matrix between sample objects and genes can be constructed. This relationship matrix is an M×N matrix, where each row corresponds to a sample object and each column corresponds to a gene. Thus, the relationship matrix can be used to intuitively determine the genes contained in each sample object and intuitively show the distribution of genes among different samples.
[0040] After obtaining the relation matrix, it is necessary to determine the number of sub-indicators in the gene statistical indicators and calculate whether the relation matrix needs to be segmented under the grouping requirements of each sub-indicator. For example, when there are 2 sub-indicators and M is 29, C(29,2) = 406 groups. The number of groups in this group is relatively small, so the relation matrix does not need to be segmented, and the statistical data of the relation matrix under this sub-indicator can be calculated directly. When there are 5 sub-indicators and M is 29, C(29,5) = 118755 groups. The number of groups in this sub-indicator is relatively large. If the relation matrix is calculated directly, the computational workload will be large. Therefore, the relation matrix can be segmented to obtain multiple sub-matrices, and then the 118755 sub-matrices can be processed separately to obtain the processing results. The calculation results of multiple sub-matrices are combined to form the statistical data of the relation matrix under this sub-indicator, thereby reducing the complexity of each data calculation and ensuring the efficiency of relation matrix analysis through parallel processing.
[0041] It should be noted that each element in the matrix can represent the presence status of a specific gene in a certain sample, usually in binary form, 0 or 1, where 0 indicates that the gene does not exist and 1 indicates that it exists, or 0 or other numbers, representing the number of times (or copy number) the gene appears in the sample object.
[0042] Table 1 shows an optional relational matrix, in which each row corresponds to a sample object, each column corresponds to a gene, and each data point represents the number of times the gene in the sample object appears.
[0043] Table 1
[0044] Gene 1 28 22 23 24 Gene 2 18 17 18 17 Gene 3 4 16 18 14 Gene 4 13 18 15 10 Gene 5 0 13 14 0
[0045] Step S103: For the sub-indicators that need to be segmented in the relation matrix, the relation matrix is segmented according to the sub-indicators to obtain P sub-matrices. The gene data in each sub-matrices is processed in parallel according to the sub-indicators to obtain P data statistical results. The P data statistical results are then combined to obtain the gene data statistical results under the target category and sub-indicators, where P is a positive integer.
[0046] Specifically, when segmenting the relation matrix, the number of sub-matrices can be determined based on the number of groups indicated by the sub-indicators. The larger the number of groups, the more sub-matrices there are, thus ensuring that the total number of calculations in a single calculation remains within a suitable range.
[0047] For example, if there are 10,000 groups, the relation matrix can be divided into 10 groups.
[0048] It should be noted that when cutting the relation matrix, it needs to be cut horizontally, that is, divided according to genes. For example, if the relation matrix contains 1000 rows of genes, when the relation matrix is divided into 10 sub-matrices, the number of sample objects in each sub-matrix remains unchanged, but the number of genes is reduced from 1000 to 100. Thus, when the number of sub-matrices is too large, the computational load for each sub-matrix is kept at a reasonable level by reducing the number of genes.
[0049] Furthermore, after dividing the relation matrix, a distributed server can be used to perform data statistical operations on each sub-matrix in parallel to obtain P data statistical results. These P data statistical results are then combined to obtain the gene data statistical results of the relation matrix, thereby completing accurate statistical analysis of the gene data statistical results of sub-indicators with a large number of groups.
[0050] Step S104: For sub-indicators that do not require segmentation of the relation matrix, process the relation matrix according to the sub-indicators to obtain the statistical results of gene data under the target category and sub-indicators.
[0051] Specifically, when the number of groups indicated by the sub-indicators is small, since the amount of data to be calculated is not large, there is no need to segment the relation matrix; the statistical results of the gene data can be calculated directly.
[0052] Step S105: Combine the statistical results of gene data under each sub-indicator to obtain the total statistical results of gene data under the target category.
[0053] Specifically, after obtaining the statistical results of gene data under each sub-indicator, the statistical results of gene data under each sub-indicator can be combined to obtain the total statistical results of gene data distributed according to the grouping requirements indicated by the sub-indicators. Figure 2 This is a schematic diagram illustrating the optional statistical summary result provided in the embodiments of this application, such as... Figure 2 As shown, the horizontal axis represents the grouping requirements indicated by each sub-indicator in the gene statistics index, that is, the number of sample objects included in each group ranges from 2 to 29. The vertical axis represents the number of gene families. The triangles indicate the statistical distribution of pan-genes, and the circles indicate the statistical distribution of core genes, thus obtaining the total statistical results of gene data under the target category.
[0054] The gene data processing method provided in this application involves obtaining gene data sets of M sample objects under a target category, resulting in M gene data sets, each containing gene data of N genes, where M and N are positive integers. A relationship matrix is constructed between the M sample objects and the N genes. Based on each sub-indicator in the gene statistical indicators, it is determined whether the relationship matrix needs to be segmented. The gene statistical indicators include T sub-indicators, each indicating the grouping requirements for statistical analysis of the sample data in the relationship matrix. For sub-indicators requiring segmentation, the relationship matrix is segmented according to the sub-indicator to obtain P sub-matrices. Gene data in each sub-matrice is processed in parallel according to the sub-indicator to obtain P statistical results. These P statistical results are then combined to obtain the gene data statistical results for the target category and the sub-indicator, where P is a positive integer. For sub-indicators not requiring segmentation, the relationship matrix is processed according to the sub-indicator to obtain the gene data statistical results for the target category and the sub-indicator. Finally, the gene data statistical results for each sub-indicator are combined to obtain the overall gene data statistical result for the target category. This invention addresses the low efficiency of statistical analysis of gene data from sample objects in related technologies. By constructing a relationship matrix between sample objects and genes, and determining whether to segment the matrix based on sub-indicators in gene statistical indicators, the invention improves the processing efficiency of gene data, especially when dealing with large amounts of sample object and gene data. This is achieved by segmenting the relationship matrix, processing each sub-matrix separately, and then combining the statistical results of the processed sub-matrixes into the statistical result of the relationship matrix.
[0055] To ensure the accuracy of gene data statistical analysis, optionally, in the gene data processing method provided in this application embodiment, obtaining a gene data set of M sample objects under the target category, and obtaining the M gene data set includes: obtaining the category information of each initial object from the object database, and filtering the category information as the initial object of the target category to obtain M sample objects, wherein the object database contains a variety of objects to be tested; detecting the occurrence frequency of each gene carried in each sample object, and determining the occurrence frequency as the gene data of each gene, to obtain the gene data set of each sample object.
[0056] It should be noted that the object database is used to store various objects to be tested, such as species and strains.
[0057] Specifically, when obtaining a gene dataset, it is first necessary to obtain sample objects. Since the sample objects required for this statistical analysis are those under the target category, the sample objects can be filtered in the object database based on the category information to obtain M sample objects.
[0058] It should be noted that M sample objects can also be sent to clients who need to perform gene data statistics. This technical solution does not restrict the method of obtaining sample objects.
[0059] After obtaining M sample objects, the genetic data of each sample object can be deeply mined. For example, the genome information of each sample can be traversed through the gene detection module to record the occurrence of each gene, that is, to identify and count the occurrence of each gene in the sample. The statistical results of the occurrence are recorded as genetic data, forming the genetic data set of each sample object.
[0060] For example, assuming 27 Staphylococcus samples are selected as M sample objects, the system will scan the genome sequence of each sample and count the occurrences of each gene family member in the sample. If a gene family appears 10 times in a sample, then the corresponding record for that gene in the sample's gene dataset will be 10.
[0061] This embodiment selects sample objects that meet the requirements of statistical operations and efficiently and accurately obtains the gene data set of M sample objects under the target category, providing high-quality and targeted initial data for subsequent pan-genome statistical analysis and ensuring the accuracy of subsequent gene data statistical analysis.
[0062] To accurately determine whether the relation matrix needs to be segmented, optionally, in the gene data processing method provided in this application embodiment, determining whether the relation matrix needs to be segmented based on each sub-indicator in the gene statistical indicators includes: for any sub-indicator, determining the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; determining the number of groups of sample objects under the grouping requirements based on the target number and the total number of sample objects, and determining a preset threshold; calculating the quotient between the number of groups and the preset threshold, and rounding the quotient up to obtain the target value; determining whether the target value is greater than a preset value, and if the target value is greater than the preset value, determining that the relation matrix needs to be segmented; if the target value is less than or equal to the preset value, determining that the relation matrix does not need to be segmented.
[0063] Specifically, since different sub-indicators indicate different numbers of sample objects per group, it is necessary to first group the sample objects according to the grouping requirements indicated in the sub-indicators to obtain the number of groups. For example, taking 5 samples per group as an example, the target number indicated in the sub-indicator is 5. For 27 sample objects, it is necessary to calculate C(27,5), that is, the number of different combinations of selecting 5 samples from 27 samples. The result is 134,596 groups, which is the number of groups.
[0064] After obtaining the number of groups, the quotient between the number of groups and the preset threshold can be calculated, and the quotient can be rounded up to obtain the target value. For example, assuming the preset threshold is set to 10,000 groups, the quotient between 134,596 groups and 10,000 groups can be calculated. That is, 134,596 divided by 10,000 gives approximately 13.4596, which is then rounded up to obtain 14, which is the target value.
[0065] After obtaining the target value, it can be compared with a preset value, which can be 2. If the target value is greater than the preset value, it is determined that the relation matrix needs to be segmented. If the target value is less than or equal to the preset value, it is determined that the relation matrix does not need to be segmented. This allows the grouping strategy of the relation matrix to be adjusted according to the complexity of the actual task, avoiding the waste of computing resources and low execution efficiency caused by the large scale of a single task.
[0066] For example, if the target value of 14 is greater than the preset value of 2, the system determines that the current task is too heavy and the original data needs to be divided into multiple subtasks for parallel processing. Conversely, if the target value is less than or equal to the preset value, the system considers the size of the original data to be within an acceptable range and no division is necessary.
[0067] This embodiment determines whether the relation matrix needs to be segmented by calculating the target value, ensuring that when processing large-scale pan-genome data, computing resources can be used efficiently while maintaining the accuracy of the analysis.
[0068] Optionally, in the gene data processing method provided in the embodiments of this application, segmenting the relation matrix according to the sub-indicators to obtain P sub-matrices includes: determining the target value as the number of sub-matrices after segmenting the relation matrix, and obtaining the value of P; segmenting the relation matrix along the gene dimension according to the target value to obtain P sub-matrices, wherein the number of sample objects in each sub-matrice is the same, and the genes contained in each sub-matrice are different.
[0069] It should be noted that the number of submatrices is the number of independent data blocks formed after the relation matrix is divided, i.e., the value of P.
[0070] Specifically, after obtaining the target value, the target value can be determined as the number of segments to be divided in the relation matrix. The relation matrix is then divided according to this number of segments. During the segmentation operation, the relation matrix needs to be divided along the gene dimension. At this time, the genes in the matrix need to be evenly distributed to ensure that the number of sample objects in each sub-matrix is the same, while the genes contained in each sub-matrix are different. This avoids the duplication of gene data while ensuring that the difference in the amount of data in each sub-matrix is within an acceptable range.
[0071] For example, if the relation matrix originally contains 1,000 genes and 29 sample objects, the system will attempt to distribute the 1,000 genes evenly into 10 sub-matrices, with each sub-matrix containing 100 genes, while the total number of sample objects, 29, is contained in each sub-matrix.
[0072] This embodiment ensures consistency in the number of sample objects in each submatrix by determining the target value as the number of submatrices and dividing the relation matrix along the gene dimension. At the same time, by distributing the number of genes, the computational complexity of each subtask is significantly reduced, thereby improving the overall efficiency and resource utilization of pan-gene set statistical analysis.
[0073] Optionally, in the gene data processing method provided in this application embodiment, the parallel processing of gene data in each sub-matrix according to sub-indicators to obtain P data statistical results includes: for any sub-matrix, determining the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; traversing and grouping the sample objects in the sub-matrix according to the target number to obtain H groups of sample objects, where H is a positive integer; obtaining gene data under each group of sample objects from the sub-matrix to obtain H matrices to be processed, and obtaining the first number of core genes and the second number of pan-genes in each matrix to be processed to obtain the first number of H groups and the second number of H groups; calculating the first statistical result of the first number of H groups and the second statistical result of the second number of H groups, wherein the first statistical result includes at least one of the following: the maximum value, minimum value and average value of the first number of H groups, and the second statistical result includes at least one of the following: the maximum value, minimum value and average value of the second number of H groups; and determining the first statistical result and the second statistical result as a single data statistical result.
[0074] Specifically, when processing gene data, for any sub-matrix under any sub-index, the number of samples that each group of sample objects should contain can be determined according to the preset sub-index, that is, the target number, and the samples can be grouped according to the target number to obtain H groups of sample objects. The process of traversing the grouping is to calculate the number of groups, that is, H groups of sample objects, according to the formula C(M,K) based on the number of sample objects in each group.
[0075] After obtaining H groups of sample objects, since each object has corresponding gene data, a matrix corresponding to each group of sample objects can be obtained, resulting in H matrices to be processed. Each matrix to be processed contains gene data corresponding to the grouped samples, ensuring the relevance and accuracy of the statistical analysis. The gene data in each matrix corresponds to the segmented gene data.
[0076] After obtaining H matrices to be processed, statistical analysis of core genes and pangenes needs to be performed on each matrix. Core genes are genes that exist in all sample objects, while pangenes are genes that appear in some samples. The first number of core genes and the second number of pangenes in each matrix are obtained, forming a set of H sets of first number (number of core genes) and H sets of second number (number of pangenes). Thus, the number of core genes and pangenes is counted, and the genetic diversity and stability of the sample population are quantified.
[0077] Only after obtaining the first and second quantities of group H, it is necessary to calculate statistical results based on these quantities, such as the maximum, minimum, and average values, to perform subsequent analysis and graphing processes. Finally, the system combines the first statistical result of the calculated core gene quantity with the second statistical result of the pan-gene quantity to form a comprehensive statistical result, ready for further analysis or visualization.
[0078] It should be noted that the data statistics result obtained by this process is the data statistics result of a sub-matrix under a sub-indicator. It is necessary to combine the data statistics results of multiple sub-matrixes to obtain the data statistics result under the sub-indicator, and then combine the data statistics results of multiple sub-indicators to obtain the total data statistics result of this statistical operation.
[0079] This embodiment achieves effective management and efficient analysis of submatrix data by traversing grouping, statistical analysis, and result integration. That is, by setting a target number of sample objects in each group, the large task is broken down into multiple small tasks for parallel computation, which significantly improves the efficiency of pan-genome set statistical analysis. At the same time, through accurate gene data statistics and comprehensive result calculation, the comprehensiveness and accuracy of the analysis are ensured, providing a solid data foundation for subsequent pan-genome set research.
[0080] Optionally, in the gene data processing method provided in this application embodiment, combining P data statistical results to obtain gene data statistical results under the target category and sub-index includes: obtaining the maximum value of the first statistical result in the data statistical results of each sub-matrix, obtaining multiple first maximum values, and determining the maximum value among the multiple first maximum values as the target first maximum value; obtaining the maximum value of the second statistical result in the data statistical results of each sub-matrix, obtaining multiple second maximum values, and determining the maximum value among the multiple second maximum values as the target second maximum value; obtaining the minimum value of the first statistical result in the data statistical results of each sub-matrix, obtaining multiple first minimum values, and determining the minimum value among the multiple first minimum values as the target first minimum value. The process involves: obtaining the minimum value of the second statistical result in the data statistics results of each submatrix, obtaining multiple second minimum values, and determining the minimum value among these multiple second minimum values as the target second minimum value; obtaining the average value of the first statistical result in the data statistics results of each submatrix, obtaining multiple first average values, and calculating the average of these multiple first average values to obtain the target first average value; obtaining the average value of the second statistical result in the data statistics results of each submatrix, obtaining multiple second average values, and calculating the average of these multiple second average values to obtain the target second average value; and determining the target first maximum value, target second maximum value, target first minimum value, target second minimum value, target first average value, and target second average value as the gene data statistics results.
[0081] Specifically, after completing the statistical analysis of each submatrix, the system extracts the first statistical result of the number of core genes in each submatrix and identifies the maximum value, i.e., the highest number of core genes observed in the corresponding submatrix. This process is repeated for all submatrixes, collecting a list containing the maximum number of core genes for each submatrix, thereby obtaining the maximum value of the first statistical result. This aims to identify the peak number of core genes in the entire pangenome set analysis, providing benchmark data for subsequent population genetic diversity and stability analysis.
[0082] Similarly, as with the core gene count, the system also extracts a second statistical result of the pangene count in each submatrix and finds the maximum pangene count in each submatrix. This operation is also performed in all submatrixes, thus collecting a complete list of the maximum pangene count in each submatrix.
[0083] Furthermore, the system also needs to obtain the minimum value of the first statistical result of the number of core genes in each sub-matrix. Through traversal analysis, the system collects and determines the lowest number of core genes in all sub-matrixes, thereby identifying the lower limit of the number of core genes in the pangenome by determining the minimum value of the first statistical result. At the same time, the system will find the minimum value of the second statistical result of the number of pangenes in each sub-matrix, forming a list containing the lowest number of pangenes in all sub-matrixes, and thus determining the target second minimum value.
[0084] Finally, the system extracts the average of the first statistical results regarding the number of core genes in each submatrix, forming a set of averages. It then performs a second average calculation on all the averages in this set to obtain the target first average, which is the global average of the number of core genes in all submatrixes of the pan-gene set. Similarly, it extracts the average of the second statistical results regarding the number of pan-genes in each submatrix, forming another set of averages. Finally, by calculating the average of this set, it obtains the target second average, which is the global average of the number of pan-genes in all submatrixes of the pan-gene set.
[0085] It should be noted that the above three calculation steps can obtain the key data needed for the final statistical operation. By integrating various statistical results, a comprehensive statistical overview reflecting the pan-genome genetic characteristics can be provided for subsequent analysis of gene data, ensuring the accuracy and comprehensiveness of the subsequent data analysis.
[0086] It should be noted that when generating data based on statistical data, such as... Figure 2 When generating a visualization, since the visualization may contain distribution maps of core genes and pan-genes, and the number of core genes and pan-genes is relatively large, gene values can be filtered during visualization generation. The number of core genes and pan-genes under each sub-index can be arranged in descending order, and grouped according to the preset number of groups, such as 30 groups, to obtain 30 groups of gene counts. The average value of the gene count in each group is calculated to obtain 30 average values. The distribution of gene counts under that sub-index is determined by the 30 average values, as well as the maximum and minimum values. In this way, the distribution of points can still be reflected while reducing the number of points in the distribution map of the visualization.
[0087] It should be noted that the curve used to represent the trend of quantity change in the visualization can be determined by the average number of genes under each sub-index, or other mathematical fitting methods can be used to generate the trend curve. This embodiment does not impose any limitations.
[0088] Optionally, in the gene data processing method provided in this application embodiment, combining the gene data statistical results under each sub-index to obtain the total gene data statistical result under the target category includes: adding the gene data statistical results under each sub-index to a preset schematic diagram according to the grouping requirements in the sub-index to obtain a gene data statistical graph; drawing a curve based on the average value of the gene data statistical results under each sub-index, and calculating the slope of the curve at each point to obtain multiple slopes; determining the validity result of the gene data statistical graph based on the slope values of the multiple slopes, and determining the gene data statistical graph and the validity result as the total gene data statistical result.
[0089] Specifically, after obtaining the statistical results of gene data under each sub-indicator, the system adds the statistical results (including maximum, minimum and average values) of the number of core genes and pan-genes under each sub-indicator calculated previously to a pre-designed schematic template based on the predefined sub-indicators. For example, the average number of core genes under different sub-indicators can be plotted as points on the X-axis, and the Y-axis position of each point represents the average value under this indicator.
[0090] Furthermore, the system can analyze the average number of core genes and pan-genes under each sub-indicator. Based on these averages, the system can plot curves on its preset schematic diagrams, reflecting the changing trend of the average number of genes as the sub-indicator changes (such as sample size, gene family type, etc.). By displaying the statistical results of gene data in graphical form, the distribution of gene numbers under different sub-indicators can be intuitively compared, and trend analysis can also be performed.
[0091] It should be noted that after obtaining the curve depicting the trend of the average value, the system can also calculate the slope of the curve at different points. This allows for the evaluation of the validity of the genetic data statistical graph by analyzing multiple slope values. For example, if the slope value shows a clear trend or pattern, such as a stable increase or decrease, it indicates that the number of sample objects used in the statistical graph is insufficient, leading to data non-convergence, and the number of sample objects needs to be increased. Conversely, if the slope value gradually decreases to a stable level, indicating that the current data has converged, the number of sample objects is sufficient, and the obtained statistical data can be used to analyze the genetic information of objects under the target category. Therefore, by evaluating the validity of the genetic data statistical graph, the reliability of the statistical analysis results and the accuracy of the visualization are ensured.
[0092] Finally, the system combines the generated gene data statistical graphs and their validity evaluation results to form a complete gene data statistical summary result. The gene data statistical summary result includes intuitive graphical representations and objective evaluations of the statistical analysis quality, achieving the technical effect of providing both intuitive and rigorous data analysis results, thereby completing an accurate and efficient analysis process for gene data.
[0093] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0094] This application also provides a gene data processing apparatus. It should be noted that the gene data processing apparatus of this application can be used to execute the gene data processing method provided in this application. The gene data processing apparatus provided in this application will be described below.
[0095] Figure 3 This is a schematic diagram of a gene data processing apparatus provided according to an embodiment of this application. Figure 3 As shown, the device includes: an acquisition unit 31, a judgment unit 32, a first processing unit 33, a second processing unit 34, and a combination unit 35.
[0096] The acquisition unit 31 is used to acquire the gene data set of M sample objects under the target category, and obtain M gene data sets, wherein each gene data set contains the gene data of N genes, and M and N are positive integers.
[0097] The judgment unit 32 is used to construct the relationship matrix between M sample objects and N genes, and to determine whether the relationship matrix needs to be segmented based on each sub-indicator in the gene statistical indicators. The gene statistical indicators include T sub-indicators, each of which is used to indicate the grouping requirements for statistical analysis of the sample data in the relationship matrix.
[0098] The first processing unit 33 is used to divide the relation matrix according to the sub-indicators that need to be divided into sub-indicators, obtain P sub-matrices, and perform parallel processing on the gene data in each sub-matrices according to the sub-indicators to obtain P data statistical results. The P data statistical results are then combined to obtain the gene data statistical results under the target category and sub-indicators, where P is a positive integer.
[0099] The second processing unit 34 is used to process the relation matrix according to the sub-indicators for sub-indicators that do not require segmentation of the relation matrix, and obtain the statistical results of gene data under the target category and sub-indicators.
[0100] Combination unit 35 is used to combine the statistical results of gene data under each sub-index to obtain the total statistical result of gene data under the target category.
[0101] The gene data processing apparatus provided in this application embodiment acquires gene data sets of M sample objects under a target category through acquisition unit 31, resulting in M gene data sets, where each gene data set contains gene data of N genes, and M and N are positive integers; judgment unit 32 constructs a relationship matrix between the M sample objects and the N genes, and determines whether the relationship matrix needs to be segmented based on each sub-indicator in the gene statistical indicators, wherein the gene statistical indicators include T sub-indicators, each sub-indicator indicating the grouping requirements for statistical analysis of the sample data in the relationship matrix; the first processing unit 33 processes the relationship matrix as needed. The first processing unit (34) divides the relation matrix into P sub-matrices based on the sub-indicators, and processes the gene data in each sub-matrice in parallel according to the sub-indicators to obtain P statistical results. These P statistical results are then combined to obtain the gene data statistical results for the target category and sub-indicators, where P is a positive integer. The second processing unit (34) processes the relation matrix for sub-indicators that do not require relation matrix division to obtain gene data statistical results for the target category and sub-indicators. The combination unit (35) combines the gene data statistical results for each sub-indicator to obtain the total gene data statistical result for the target category. This solves the problem of low efficiency in statistical analysis of gene data of sample objects in related technologies. By constructing a relation matrix between sample objects and genes, and determining whether the relation matrix needs to be divided based on the sub-indicators in the gene statistical indicators, even with a large amount of sample object and gene data, the relation matrix can be divided, and each sub-matrix can be processed separately. The statistical results of the processed sub-matrices can then be combined to form the statistical results of the relation matrix, thereby improving the technical efficiency of gene data processing.
[0102] Optionally, in the gene data processing apparatus provided in this application embodiment, the acquisition unit 31 includes: a first acquisition module, used to acquire the type information of each initial object from the object database, and filter the initial objects whose type information is the target type to obtain M sample objects, wherein the object database contains a variety of objects to be tested; and a first determination module, used to detect the occurrence frequency of each gene carried in each sample object, and determine the occurrence frequency as the gene data of each gene to obtain the gene data set of each sample object.
[0103] Optionally, in the gene data processing apparatus provided in this application embodiment, the judgment unit 32 includes: a second determination module, used to determine the target number of sample objects in each group of sample objects according to the grouping requirements in any sub-indicator; a third determination module, used to determine the number of groups of sample objects under the grouping requirements according to the target number and the total number of sample objects, and determine a preset threshold; a first calculation module, used to calculate the quotient between the number of groups and the preset threshold, and to round up the quotient to obtain the target value; a judgment module, used to determine whether the target value is greater than a preset value, and if the target value is greater than the preset value, to determine that a relation matrix is needed for segmentation; and a fourth determination module, used to determine that a relation matrix is not needed for segmentation if the target value is less than or equal to the preset value.
[0104] Optionally, in the gene data processing apparatus provided in the embodiments of this application, the first processing unit 33 includes: a fifth determining module, used to determine the target value as the number of sub-matrices after segmenting the relation matrix, to obtain the value of P; and a segmentation module, used to segment the relation matrix along the gene dimension according to the target value, to obtain P sub-matrices, wherein the number of sample objects in each sub-matrice is the same, and the genes contained in each sub-matrice are different.
[0105] Optionally, in the gene data processing apparatus provided in this application embodiment, the first processing unit 33 includes: a sixth determining module, used to determine the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-index for any sub-matrix; a grouping module, used to traverse and group the sample objects in the sub-matrix according to the target number to obtain H groups of sample objects, where H is a positive integer; a second obtaining module, used to obtain gene data under each group of sample objects from the sub-matrix to obtain H matrices to be processed, and obtain the first number of core genes and the second number of pan-genes in each matrix to be processed to obtain the first number of H groups and the second number of H groups; a second calculation module, used to calculate the first statistical result of the first number of H groups and the second statistical result of the second number of H groups, wherein the first statistical result includes at least one of the following: the maximum value, minimum value and average value of the first number of H groups, and the second statistical result includes at least one of the following: the maximum value, minimum value and average value of the second number of H groups; and a seventh determining module, used to determine the first statistical result and the second statistical result as a data statistical result.
[0106] Optionally, in the gene data processing apparatus provided in this application embodiment, the first processing unit 33 includes: a third acquisition module, configured to acquire the maximum value of the first statistical result in the data statistical results of each sub-matrix, obtain multiple first maximum values, and determine the maximum value among the multiple first maximum values as the target first maximum value; a fourth acquisition module, configured to acquire the maximum value of the second statistical result in the data statistical results of each sub-matrix, obtain multiple second maximum values, and determine the maximum value among the multiple second maximum values as the target second maximum value; a fifth acquisition module, configured to acquire the minimum value of the first statistical result in the data statistical results of each sub-matrix, obtain multiple first minimum values, and determine the minimum value among the multiple first minimum values as the target first minimum value; and a sixth acquisition module, configured to acquire the maximum value of the first statistical result in the data statistical results of each sub-matrix; The first module is used to obtain the minimum value of the second statistical result in the data statistics results of each submatrix, obtain multiple second minimum values, and determine the minimum value among the multiple second minimum values as the target second minimum value; the seventh module is used to obtain the average value of the first statistical result in the data statistics results of each submatrix, obtain multiple first average values, and calculate the average of the multiple first average values to obtain the target first average value; the eighth module is used to obtain the average value of the second statistical result in the data statistics results of each submatrix, obtain multiple second average values, and calculate the average of the multiple second average values to obtain the target second average value; the eighth module is used to determine the target first maximum value, target second maximum value, target first minimum value, target second minimum value, target first average value, and target second average value as the gene data statistics results.
[0107] Optionally, in the gene data processing apparatus provided in this application embodiment, the combination unit 35 includes: an adding module, used to add the gene data statistical results under each sub-index to a preset schematic diagram according to the grouping requirements in the sub-index, to obtain a gene data statistical graph; a third calculation module, used to draw a curve based on the average value of the gene data statistical results under each sub-index, and calculate the slope of the curve at each point to obtain multiple slopes; and a ninth determining module, used to determine the validity result of the gene data statistical graph based on the slope values of the multiple slopes, and determine the gene data statistical graph and the validity result as the total gene data statistical result.
[0108] The aforementioned gene data processing device includes a processor and a memory. The aforementioned acquisition unit 31, judgment unit 32, first processing unit 33, second processing unit 34, combination unit 35, etc., are all stored in the memory as program units. The processor executes the aforementioned program units stored in the memory to realize the corresponding functions.
[0109] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured; by adjusting kernel parameters, the low efficiency of statistical analysis of genetic data from sample objects in related technologies can be addressed.
[0110] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0111] This invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements a method for processing the genetic data.
[0112] This invention provides a processor for running a program, wherein the program executes a method for processing the gene data.
[0113] Figure 4 This is a schematic diagram of an electronic device provided according to an embodiment of this application, such as... Figure 4 As shown, this embodiment of the invention provides an electronic device 40, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the above-described gene data processing method. The device described herein can be a server, PC, PAD, mobile phone, etc.
[0114] This application also provides a computer program product that, when executed on a data processing device, is adapted to perform the steps of the processing method for initializing the above-described genetic data.
[0115] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0116] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0117] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0118] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0119] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0120] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0121] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0122] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0123] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for processing gene data, characterized in that, include: Obtain the gene data set of M sample objects under the target category, resulting in M gene data sets, where each gene data set contains the gene data of N genes, and M and N are positive integers; Construct a relationship matrix between the M sample objects and the N genes, and determine whether the relationship matrix needs to be segmented based on each sub-indicator in the gene statistical indicators. The gene statistical indicators include T sub-indicators, each of which is used to indicate the grouping requirements for statistical analysis of the sample data in the relationship matrix, where T is a positive integer. For a sub-index that requires segmentation of the relation matrix, the relation matrix is segmented according to the sub-index to obtain P sub-matrices. The gene data in each sub-matrices are then processed in parallel according to the sub-index to obtain P data statistical results. The P data statistical results are then combined to obtain the gene data statistical results under the target category and the sub-index, where P is a positive integer. For sub-indicators that do not require segmentation of the relation matrix, the relation matrix is processed according to the sub-indicators to obtain the target category and the gene data statistics under the sub-indicators; The statistical results of gene data under each sub-indicator are combined to obtain the total statistical results of gene data under the target category. The gene data in each submatrix is processed in parallel according to the sub-indicators to obtain P statistical results, including: for any submatrix, determining the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicators; traversing and grouping the sample objects in the submatrix according to the target number to obtain H groups of sample objects, where H is a positive integer; obtaining gene data under each group of sample objects from the submatrix to obtain H matrices to be processed, and obtaining the first number of core genes and the second number of pan-genes in each matrix to be processed to obtain the first number of H groups and the second number of H groups; calculating the first statistical result of the first number of H groups and the second statistical result of the second number of H groups, wherein the first statistical result includes at least one of the following: the maximum value, minimum value and average value of the first number of H groups, and the second statistical result includes at least one of the following: the maximum value, minimum value and average value of the second number of H groups; and determining the first statistical result and the second statistical result as a single statistical result.
2. The method according to claim 1, characterized in that, Obtain the gene dataset of M sample objects under the target category. The resulting M gene dataset includes: The object database is used to obtain the type information of each initial object, and the initial objects of the type information are selected as the target type to obtain the M sample objects. The object database contains a variety of objects to be tested. The occurrence frequency of each gene carried in each sample object is detected, and the occurrence frequency is determined as the gene data of each gene, thus obtaining the gene data set of each sample object.
3. The method according to claim 1, characterized in that, Determining whether the relationship matrix needs to be segmented based on each sub-indicator in the gene statistical indicators includes: For any sub-indicator, determine the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; The number of groups of sample objects under the grouping requirement is determined based on the target number and the total number of sample objects, and a preset threshold is determined; Calculate the quotient between the number of groups and the preset threshold, and round the quotient up to obtain the target value; Determine whether the target value is greater than a preset value, and if the target value is greater than the preset value, determine that the relationship matrix needs to be segmented; If the target value is less than or equal to the preset value, it is determined that the relationship matrix does not need to be segmented.
4. The method according to claim 3, characterized in that, The relation matrix is segmented based on the sub-indices to obtain P sub-matrices, including: The target value is determined as the number of sub-matrices after dividing the relation matrix, and the value of P is obtained. The relationship matrix is divided along the gene dimension according to the target value to obtain P sub-matrices, wherein the number of sample objects in each sub-matrice is the same, and the genes contained in each sub-matrice are different.
5. The method according to claim 1, characterized in that, Combining the P statistical results yields the following gene data statistical results for the target category and the sub-indicator: Obtain the maximum value of the first statistical result in the data statistical results of each submatrix, obtain multiple first maximum values, and determine the maximum value among the multiple first maximum values as the target first maximum value; The maximum value of the second statistical result in the data statistical results of each submatrix is obtained, resulting in multiple second maximum values, and the maximum value among the multiple second maximum values is determined as the target second maximum value; Obtain the minimum value of the first statistical result in the data statistical results of each submatrix, obtain multiple first minimum values, and determine the minimum value among the multiple first minimum values as the target first minimum value; Obtain the minimum value of the second statistical result in the data statistical results of each submatrix, obtain multiple second minimum values, and determine the minimum value among the multiple second minimum values as the target second minimum value; Obtain the average value of the first statistical result in the data statistical results of each submatrix, obtain multiple first average values, and calculate the average of the multiple first average values to obtain the target first average value; Obtain the average of the second statistical results in the data statistics of each submatrix to get multiple second averages, and calculate the average of the multiple second averages to obtain the target second average; The first maximum value, the second maximum value, the first minimum value, the second minimum value, the first average value, and the second average value are determined as the statistical results of the gene data.
6. The method according to claim 1, characterized in that, The statistical results of gene data under each sub-indicator are combined to obtain the total statistical results of gene data under the target category, including: According to the grouping requirements in the sub-indicators, the gene data statistics results under each sub-indicator are added to the preset schematic diagram to obtain the gene data statistics diagram; Based on the average value of the gene data statistics results under each sub-index, a curve is plotted, and the slope of the curve at each point is calculated to obtain multiple slopes. The validity result of the gene data statistical graph is determined based on the slope values of the multiple slopes, and the gene data statistical graph and the validity result are determined as the total gene data statistical result.
7. A gene data processing device, characterized in that, include: The acquisition unit is used to acquire the gene data set of M sample objects under the target category, resulting in M gene data sets, where each gene data set contains the gene data of N genes, and M and N are positive integers; The judgment unit is used to construct the relationship matrix between the M sample objects and the N genes, and to determine whether the relationship matrix needs to be segmented according to each sub-indicator in the gene statistical index. The gene statistical index includes T sub-indicators, each of which is used to indicate the grouping requirements for statistical analysis of the sample data in the relationship matrix, where T is a positive integer. The first processing unit is configured to, for the sub-indicators that need to be segmented in the relation matrix, segment the relation matrix according to the sub-indicators to obtain P sub-matrices, and perform parallel processing on the gene data in each sub-matrices according to the sub-indicators to obtain P data statistical results, and combine the P data statistical results to obtain the gene data statistical results under the target category and the sub-indicators, where P is a positive integer; The second processing unit is used to process the relationship matrix according to the sub-indicators that do not require segmentation of the relationship matrix, and obtain the target category and the gene data statistics results under the sub-indicators. The combination unit is used to combine the statistical results of gene data under each sub-index to obtain the total statistical result of gene data under the target category. The first processing unit includes: a sixth determining module, used to determine the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-index for any sub-matrix; a grouping module, used to traverse and group the sample objects in the sub-matrix according to the target number to obtain H groups of sample objects, where H is a positive integer; a second obtaining module, used to obtain gene data under each group of sample objects from the sub-matrix to obtain H matrices to be processed, and obtain the first number of core genes and the second number of pan-genes in each matrix to be processed to obtain the first number of H groups and the second number of H groups; a second calculation module, used to calculate the first statistical result of the first number of H groups and the second statistical result of the second number of H groups, wherein the first statistical result includes at least one of the following: the maximum value, minimum value and average value of the first number of H groups, and the second statistical result includes at least one of the following: the maximum value, minimum value and average value of the second number of H groups; and a seventh determining module, used to determine the first statistical result and the second statistical result as a data statistical result.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the gene data processing method according to any one of claims 1 to 6.
9. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method for processing genetic data according to any one of claims 1 to 6.
Citation Information
Patent Citations
Construction method of generic genome, terminal equipment and storage medium
CN117037912A
Gene regulation relation prediction method based on matrix enhancement and feature fusion
CN117789831A