Gene data processing method and device and electronic equipment
By constructing and dividing the relationship matrix of genetic data for parallel processing, the problem of low efficiency in statistical analysis of genetic data is solved, and efficient genetic data processing is achieved.
Patent Information
- Application Number
- CN202510702278.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The existing technology for statistical analysis of genetic data of sample subjects is inefficient, especially when processing large-scale genetic data. The computational efficiency is low and is limited by the hardware performance of a single computer, which affects the work efficiency and research progress of scientific researchers.
By constructing the relationship matrix between sample objects and genes, it is determined whether the relationship matrix needs to be segmented according to the sub-indicators in the gene statistical indicators. After segmentation, the matrix is divided into multiple sub-matrices for parallel processing, and finally the statistical results are combined.
It improves the processing efficiency of genetic data, can efficiently complete statistical analysis in the case of large-scale data, reduce computational complexity and improve resource utilization.
Smart Images

Figure CN120673855A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data, and more specifically, to a method, device, and electronic device for processing genetic data. Background Art
[0002] In the field of bioinformatics research, pan-genome analysis is a crucial task. It involves comparing the genomes of sample objects from multiple species or strains to identify the core genes shared between them and the unique pan-genes, which is of great significance for understanding genome diversity, evolutionary relationships and species adaptability.
[0003] Currently, when comparing genetic data from sample subjects, large-scale genetic data is often processed directly as a whole, analyzing all the genetic data that needs to be compared at once. However, with the rapid development of gene sequencing technology, each sample subject represents an independent genome, and the amount of genetic data is growing exponentially. The sheer number of genes and the complex permutations and combinations between sample subjects make direct processing an extremely time-consuming task. This processing method is not only computationally inefficient but is often limited by the hardware performance of a single computer. It is difficult to complete pan-gene set analysis of thousands of sample combinations and sample data from multiple samples in a short period of time, seriously affecting the work efficiency of researchers and the timeliness of research progress.
[0004] Currently, no effective solution has been proposed to the problem of low efficiency in statistical analysis of genetic data of sample objects in related technologies. Summary of the Invention
[0005] The present application provides a method, apparatus, and electronic device for processing genetic data to solve the problem of low efficiency in statistical analysis of genetic data of sample objects in related technologies.
[0006] According to one aspect of the present application, a method for processing genetic data is provided. The method includes: acquiring gene data sets of M sample objects under a target category to obtain M gene data sets, wherein each gene data set contains gene data of N genes, where M and N are positive integers; constructing a relationship matrix between the M sample objects and the N genes, and judging whether it is necessary to segment the relationship matrix according to each sub-indicator in a gene statistical index, wherein the gene statistical index includes T sub-indicators, each sub-indicator is used to indicate a grouping requirement for statistically analyzing sample data in the relationship matrix, and T is a positive integer; for sub-indicators that require segmentation of the relationship matrix, segmenting the relationship matrix according to the sub-indicators to obtain P sub-matrices, performing parallel processing on the gene data in each sub-matrix according to the sub-indicators to obtain P data statistical results, and combining the P data statistical results to obtain a gene data statistical result under the target category and the sub-indicator, wherein P is a positive integer; for sub-indicators that do not require segmentation of the relationship matrix, processing the relationship matrix according to the sub-indicators to obtain a gene data statistical result under the target category and the sub-indicator; and combining the gene data statistical results under each sub-indicator to obtain a total gene data statistical result under the target category.
[0007] Optionally, obtaining a genetic data set of M sample objects under a target category, obtaining the M genetic data sets includes: obtaining category information of each initial object from an object database, and screening the initial objects whose category information is the target category, to obtain M sample objects, wherein the object database contains a variety of objects to be tested; detecting the number of occurrences of each gene carried in each sample object, and determining the number of occurrences as genetic data of each gene, to obtain a genetic data set for each sample object.
[0008] Optionally, judging whether the relationship matrix needs to be segmented according to each sub-indicator in the gene statistical indicator includes: for any sub-indicator, determining the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; determining the number of groups of sample objects under the grouping requirements according to the target number and the total number of sample objects, and determining a preset threshold; calculating the quotient between the number of groups and the preset threshold, and rounding up the quotient to obtain a target value; judging whether the target value is greater than a preset value, and if the target value is greater than the preset value, determining that the relationship matrix needs to be segmented; if the target value is less than or equal to the preset value, determining that the relationship matrix does not need to be segmented.
[0009] Optionally, segmenting the relationship matrix according to the sub-indicators to obtain P sub-matrices includes: determining the target value as the number of sub-matrices after segmenting the relationship matrix to obtain the value of P; segmenting the relationship matrix along the gene dimension according to the target value to obtain P sub-matrices, wherein the number of sample objects in each sub-matrix is the same and the genes contained in each sub-matrix are different.
[0010] Optionally, the gene data in each submatrix is processed in parallel according to the sub-indicator to obtain P data statistical results, including: for any submatrix, determining the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; traversing and grouping the sample objects in the submatrix according to the target number to obtain H groups of sample objects, where H is a positive integer; obtaining the gene data under each group of sample objects from the submatrix to obtain H matrices to be processed, and obtaining the first number of core genes and the second number of pan-genes in each matrix to be processed to obtain H groups of first numbers and H groups of second numbers; calculating a first statistical result of the H groups of first numbers and a second statistical result of the H groups of second numbers, wherein the first statistical result includes at least one of the following: the maximum value, the minimum value and the average value of the H groups of first numbers, and the second statistical result includes at least one of the following: the maximum value, the minimum value and the average value of the H groups of second numbers; and determining the first statistical result and the second statistical result as one data statistical result.
[0011] Optionally, combining P data statistical results to obtain the genetic data statistical results under the target category and sub-indicator includes: obtaining the maximum value of the first statistical result in the data statistical results of each submatrix, obtaining multiple first maximum values, and determining the maximum value of the multiple first maximum values as the target first maximum value; obtaining the maximum value of the second statistical result in the data statistical results of each submatrix, obtaining multiple second maximum values, and determining the maximum value of the multiple second maximum values as the target second maximum value; obtaining the minimum value of the first statistical result in the data statistical results of each submatrix, obtaining multiple first minimum values, and determining the minimum value of the multiple first minimum values as the target first minimum value; obtaining the data statistical results of each submatrix According to the minimum value of the second statistical result in the statistical result, multiple second minimum values are obtained, and the minimum value among the multiple second minimum values is determined as the target second minimum value; the average value of the first statistical result in the data statistical result of each submatrix is obtained to obtain multiple first average values, and the average value of the multiple first average values is calculated to obtain the target first average value; the average value of the second statistical result in the data statistical result of each submatrix is obtained to obtain multiple second average values, and the average value of the multiple second average values is calculated to obtain the target second average value; the target first maximum value, the target second maximum value, the target first minimum value, the target second minimum value, the target first average value, and the target second average value are determined as the genetic data statistical results.
[0012] Optionally, combining the genetic data statistical results under each sub-indicator to obtain the total genetic data statistical results under the target category includes: adding the genetic data statistical results under each sub-indicator to a preset schematic diagram according to the grouping requirements in the sub-indicators to obtain a genetic data statistical graph; drawing a curve according to the average value of the genetic data statistical results under each sub-indicator, and calculating the slope of the curve at each point to obtain multiple slopes; determining the validity result of the genetic data statistical graph according to the slope values of the multiple slopes, and determining the genetic data statistical graph and the validity result as the total genetic data statistical result.
[0013] According to another aspect of the present application, a device for processing genetic data is provided. The device includes: an acquisition unit, which is used to acquire genetic data sets of M sample objects under a target type to obtain M genetic data sets, wherein each genetic data set contains genetic data of N genes, and M and N are positive integers; a judgment unit, which is used to construct a relationship matrix between the M sample objects and the N genes, and judge whether it is necessary to segment the relationship matrix according to each sub-indicator in the genetic statistical index, wherein the genetic statistical index includes T sub-indicators, each sub-indicator is used to indicate the grouping requirements for statistical analysis of the sample data in the relationship matrix; a first processing unit, which is used to determine the sub-indicators for which the relationship matrix needs to be segmented. The relationship matrix is divided according to the sub-indicators to obtain P sub-matrices, and the genetic data in each sub-matrix is processed in parallel according to the sub-indicators to obtain P data statistical results, and the P data statistical results are combined to obtain the genetic data statistical results under the target category and sub-indicators, wherein P is a positive integer; the second processing unit is used to process the relationship matrix according to the sub-indicators for sub-indicators that do not require segmentation of the relationship matrix, and obtain the genetic data statistical results under the target category and sub-indicators; the combining unit is used to combine the genetic data statistical results under each sub-indicator to obtain the total genetic data statistical result under the target category.
[0014] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements a method for processing genetic data provided in the aforementioned embodiment of the present application.
[0015] According to another aspect of the present invention, an electronic device is provided, comprising one or more processors and a memory; the memory stores computer-readable instructions, and the processor is used to execute the computer-readable instructions, wherein when the computer-readable instructions are executed, a method for processing genetic data provided in the aforementioned embodiment is executed.
[0016] The present application adopts the following steps: obtaining gene data sets of M sample objects under a target category to obtain M gene data sets, wherein each gene data set contains gene data of N genes, where M and N are positive integers; constructing a relationship matrix between the M sample objects and the N genes, and judging whether it is necessary to segment the relationship matrix according to each sub-indicator in the gene statistical index, wherein the gene statistical index includes T sub-indicators, each sub-indicator being used to indicate a grouping requirement for statistically analyzing the sample data in the relationship matrix; for sub-indicators that require segmentation of the relationship matrix, segmenting the relationship matrix according to the sub-indicators to obtain P sub-matrices, and performing parallel processing on the gene data in each sub-matrix according to the sub-indicators to obtain P data statistical results, and combining the P data statistical results to obtain the gene data statistical results under the target category and the sub-indicator, wherein P is a positive integer; for sub-indicators that do not require segmentation of the relationship matrix, processing the relationship matrix according to the sub-indicators to obtain the gene data statistical results under the target category and the sub-indicator; and combining the gene data statistical results under each sub-indicator to obtain the total gene data statistical result under the target category. This solves the problem of low efficiency in statistical analysis of genetic data from sample objects in related technologies. By constructing a relationship matrix between sample objects and genes, and determining whether to segment the relationship matrix based on sub-indicators within the genetic statistical indicators, the method can be used to segment the relationship matrix and process each sub-matrix separately, combining the statistical results of the processed sub-matrices into the statistical results of the relationship matrix, thereby achieving the technical effect of improving the processing efficiency of genetic data. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0018] Figure 1 is a flow chart of a method for processing genetic data provided in an embodiment of the present application;
[0019] Figure 2 is a schematic diagram of an optional statistical summary result provided according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of a genetic data processing device provided in accordance with an embodiment of the present application;
[0021] Figure 4 This is a schematic diagram of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0023] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] It should be noted that the genetic data processing methods, devices, and electronic devices determined in the present disclosure can be used in the field of big data, and can also be used in any field other than the field of big data. The application fields of the genetic data processing methods, devices, and electronic devices determined in the present disclosure are not limited.
[0026] It should be noted that the collected information, user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) used in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize use or refuse use. If the user chooses to refuse, the expert decision-making process will be entered. For example, an interface is set up between this system and relevant users or institutions. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or institution through the interface, and obtain relevant information after receiving the consent information fed back by the aforementioned user or institution.
[0027] The embodiments or examples of the present disclosure are not exhaustive, but are merely illustrations of some embodiments or examples, and are not intended to be specific limitations on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment or example can be implemented as an independent example, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment or example can also be implemented as an independent example, and the order of the steps in a certain embodiment or example can be arbitrarily exchanged. In addition, the optional methods or optional examples in a certain embodiment or example can be arbitrarily combined; in addition, the various embodiments or examples can be arbitrarily combined. For example, some or all steps of different embodiments or examples can be arbitrarily combined, and a certain embodiment or example can be arbitrarily combined with the optional methods or optional examples of other embodiments or examples.
[0028] For ease of description, some nouns or terms involved in the embodiments of the present application are explained below:
[0029] Core genes: genes that exist in all samples and represent the conserved genetic units required for species to maintain basic survival and reproduction.
[0030] Pan-genome: refers to the gene set of all individuals within a species, which, in addition to the core genes, also includes genes that are unique to all individuals or shared by some individuals. This part of the genes constitutes the pan-genome of the species.
[0031] Gene family: A gene family is a group of genes with similar sequence, structure, or function that have evolved from a common ancestral gene, usually through gene duplication events.
[0032] According to an embodiment of the present application, a method for processing genetic data is provided.
[0033] Figure 1 This is a flow chart of a method for processing genetic data according to an embodiment of the present application. Figure 1 As shown, the method includes the following steps:
[0034] Step S101 : Acquire gene data sets of M sample objects under a target category to obtain M gene data sets, wherein each gene data set contains gene data of N genes, where M and N are positive integers.
[0035] It should be noted that the execution entity of this embodiment can be a data processing system. The target species refers to the specific biological population to be analyzed. The sample objects represent individual samples selected from this population. Each individual has a genome, and the genes represent the genes (or gene families) contained in each individual's genome. The genetic data set covers the genetic information of each gene in each sample object, including gene sequence, gene annotation, etc.
[0036] Specifically, when processing genetic data, all the underlying data involved must first be acquired. This means first determining the target category for genetic data analysis and obtaining M representative samples of that category. Next, the genetic data for each sample can be collected, encompassing all N genes within each sample. This genetic data can be obtained from various sources, such as public databases and in-house sequencing data. After data collection is complete, quality control and preprocessing can be performed to ensure consistency and accuracy, laying the foundation for subsequent analysis.
[0037] Step S102: construct a relationship matrix between M sample objects and N genes, and determine whether the relationship matrix needs to be segmented based on each sub-indicator in the gene statistical index. The gene statistical index includes T sub-indicators, and each sub-indicator is used to indicate the grouping requirements for statistical analysis of the sample data in the relationship matrix.
[0038] It should be noted that the relationship matrix is a data structure that displays the association between M samples and N genes in a two-dimensional table format. The gene statistical index includes multiple sub-indicators, each of which is used to indicate the number of sample objects contained in each group when the sample data is combined. For example, if the gene statistical index is (2, 29), it represents the existence of 28 sub-indicators, indicating that the sample objects need to be grouped in pairs, groups of three, and finally groups of 29, thereby obtaining the number of groups from C(29, 2) to C(29, 29).
[0039] Specifically, after obtaining the basic data, we can first construct a relationship matrix between sample objects and genes. The relationship matrix is an M×N relationship matrix. Each row of the matrix corresponds to a sample object, and each column corresponds to a gene. Therefore, the relationship matrix can be used to intuitively determine the genes contained in each sample object, and intuitively display the distribution of genes among different samples.
[0040] After obtaining the relationship matrix, it is necessary to determine the number of sub-indicators in the gene statistical index and calculate whether the relationship matrix needs to be split under the grouping requirements of each sub-indicator. For example, when the sub-indicator is 2 and M is 29, C(29, 2) = 406 groups. The number of groups in this group is small, so the relationship matrix can be omitted and the statistical data of the relationship matrix under this sub-indicator can be directly calculated. When the sub-indicator is 5 and M is 29, C(29, 5) = 118755 groups. The number of groups in this sub-indicator is large. If the relationship matrix is calculated directly, the amount of calculation is large. Therefore, the relationship matrix can be split to obtain multiple sub-matrices, and then the 118755 groups of sub-matrices are processed separately to obtain the processing results. The calculation results of the multiple sub-matrices are combined into the statistical data of the relationship matrix under this sub-indicator, thereby reducing the complexity of each data calculation and ensuring the efficiency of analyzing and processing the relationship matrix through parallel processing.
[0041] It should be noted that each element in the matrix can represent the presence status of a specific gene in a certain sample, usually a binary 0 or 1, where 0 indicates that the gene does not exist and 1 indicates that it exists, or 0 or other numbers representing the number of occurrences (or copy number) of the gene in the sample object.
[0042] Table 1 is an optional relationship matrix, in which each row corresponds to a sample object, each column corresponds to a gene, and each data represents the number of occurrences of the gene contained in the sample object.
[0043] Table 1
[0044] Sample object 1 Sample object 2 Sample object 3 Sample object 4 Gene 1 28 22 23 24 Gene 2 18 17 18 17 Gene 3 4 16 18 14 Gene 4 13 18 15 10 Gene 5 0 13 14 0
[0045] Step S103: For the sub-indicators that require segmentation of the relationship matrix, the relationship matrix is segmented according to the sub-indicators to obtain P sub-matrices, and the gene data in each sub-matrix is processed in parallel according to the sub-indicators to obtain P data statistical results, and the P data statistical results are combined to obtain the gene data statistical results under the target category and sub-indicator, where P is a positive integer.
[0046] Specifically, when the relationship matrix is segmented, the number of segmented sub-matrices can be determined according to the number of grouping groups indicated by the sub-indicator. The larger the number of grouping groups, the more sub-matrices there are, thereby ensuring that the total number of calculations for a single calculation remains within an appropriate number range.
[0047] For example, when the number of grouping groups is 10,000, the relationship matrix can be divided into 10 groups.
[0048] It should be noted that when cutting the relationship matrix, it is necessary to cut it horizontally, that is, to divide it according to genes. For example, the relationship matrix includes 1000 rows of genes. When the relationship matrix is divided into 10 sub-matrices, the number of sample objects in each sub-matrix remains unchanged, and the number of genes changes from 1000 to 100. Therefore, when the number of sub-matrices is too large, the number of genes can be reduced to ensure that the calculation amount of each sub-matrix is maintained at a reasonable level.
[0049] Furthermore, after the relationship matrix is divided, a distributed server can be used to perform data statistical operations on each sub-matrix in parallel to obtain P data statistical results, and the P data statistical results can be combined to obtain the genetic data statistical results of the relationship matrix, thereby completing accurate statistical analysis of the genetic data statistical results of sub-indicators with a large number of groups.
[0050] Step S104: For sub-indicators that do not require segmentation of the relationship matrix, the relationship matrix is processed according to the sub-indicators to obtain the statistical results of the gene data under the target category and the sub-indicators.
[0051] Specifically, when the number of groups indicated by the sub-indicators is small, since the amount of data to be calculated is not large, there is no need to split the relationship matrix, and the statistical results of the genetic data can be directly calculated.
[0052] Step S105 : combining the genetic data statistical results under each sub-indicator to obtain the overall genetic data statistical results under the target category.
[0053] Specifically, after obtaining the genetic data statistical results under each sub-indicator, the genetic data statistical results under each sub-indicator can be combined to obtain the total genetic data statistical results distributed according to the grouping requirements indicated by the sub-indicators. Figure 2 is a schematic diagram of an optional statistical total result provided according to an embodiment of the present application, such as Figure 2 As shown, the horizontal axis is the grouping requirements indicated by each sub-indicator in the gene statistical indicator, that is, the number of sample objects included in each group ranges from 2 to 29, the vertical axis is the number of gene families, the triangle mark is the statistical quantity distribution of pan-genes, and the circle mark is the statistical quantity distribution of core genes, thereby obtaining the total statistical results of gene data under the target category.
[0054] The genetic data processing method provided in an embodiment of the present application obtains M genetic data sets by acquiring genetic data sets of M sample objects under a target category, wherein each genetic data set contains genetic data of N genes, where M and N are positive integers; constructs a relationship matrix between the M sample objects and the N genes, and determines whether the relationship matrix needs to be segmented based on each sub-indicator in the genetic statistical index, wherein the genetic statistical index includes T sub-indicators, each sub-indicator being used to indicate a grouping requirement for statistically analyzing the sample data in the relationship matrix; for sub-indicators that require segmentation of the relationship matrix, the relationship matrix is segmented based on the sub-indicators to obtain P sub-matrices, and the genetic data in each sub-matrix is processed in parallel based on the sub-indicators to obtain P data statistical results, and the P data statistical results are combined to obtain genetic data statistical results under the target category and sub-indicators, wherein P is a positive integer; for sub-indicators that do not require segmentation of the relationship matrix, the relationship matrix is processed based on the sub-indicators to obtain genetic data statistical results under the target category and sub-indicators; and the genetic data statistical results under each sub-indicator are combined to obtain a total genetic data statistical result under the target category. This solves the problem of low efficiency in statistical analysis of genetic data from sample objects in related technologies. By constructing a relationship matrix between sample objects and genes, and determining whether to segment the relationship matrix based on sub-indicators within the genetic statistical indicators, the method can be used to segment the relationship matrix and process each sub-matrix separately, combining the statistical results of the processed sub-matrices into the statistical results of the relationship matrix, thereby achieving the technical effect of improving the processing efficiency of genetic data.
[0055] In order to ensure the accuracy of the statistical analysis of genetic data, optionally, in the genetic data processing method provided in the embodiment of the present application, a genetic data set of M sample objects under the target category is obtained, and obtaining the M genetic data set includes: obtaining the category information of each initial object from the object database, and screening the initial objects whose category information is the target category to obtain M sample objects, wherein the object database contains a variety of objects to be tested; detecting the number of occurrences of each gene carried in each sample object, and determining the number of occurrences as the genetic data of each gene, to obtain the genetic data set of each sample object.
[0056] It should be noted that the object database is used to store various objects to be tested, such as species, strains, etc.
[0057] Specifically, when obtaining a genetic data set, you first need to obtain sample objects. Since the sample objects required for this statistics are sample objects under the target category, you can screen the sample objects in the object database according to the category information to obtain M sample objects.
[0058] It should be noted that the M sample objects may also be sent by customers who need to perform genetic data statistics, and this technical solution does not limit the method for obtaining the sample objects.
[0059] After obtaining M sample objects, the genetic data of each sample object can be deeply mined. For example, the genome information of each sample can be traversed through the gene detection module to record the number of occurrences of each gene, that is, to identify and count the number of occurrences of each gene in the sample, and the statistical results of the number of occurrences can be recorded as genetic data, forming a genetic data set for each sample object.
[0060] For example, suppose 27 Staphylococcus samples are selected as M sample objects. The system will scan the genome sequence of each sample and count the number of occurrences of each gene family member in the sample. If a gene family appears 10 times in a sample, the gene data set for that sample will have 10 records corresponding to this gene.
[0061] This embodiment selects sample objects that meet the statistical operation requirements and efficiently and accurately obtains the gene data set of M sample objects under the target category, providing high-quality and targeted initial data for subsequent pan-gene set statistical analysis, thereby ensuring the accuracy of subsequent genetic data statistical analysis.
[0062] In order to accurately determine whether the relationship matrix needs to be segmented, optionally, in the genetic data processing method provided in the embodiment of the present application, judging whether the relationship matrix needs to be segmented according to each sub-indicator in the genetic statistical indicator includes: for any sub-indicator, determining the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; determining the number of groups of sample objects under the grouping requirements according to the target number and the total number of sample objects, and determining a preset threshold; calculating the quotient between the number of groups and the preset threshold, and rounding up the quotient to obtain a target value; judging whether the target value is greater than a preset value, and if the target value is greater than the preset value, determining that the relationship matrix needs to be segmented; if the target value is less than or equal to the preset value, determining that the relationship matrix does not need to be segmented.
[0063] Specifically, since different sub-indicators indicate that each group of samples should contain different numbers of sample objects, it is necessary to first group the sample objects according to the grouping requirements indicated in the sub-indicators to obtain the number of groups. For example, taking 5 samples in each group as an example, the target number indicated in the sub-indicator is 5. For 27 sample objects, it is necessary to calculate C(27, 5), that is, the number of different combinations of 5 samples selected from the 27 samples. The result is 134,596 groups, which is the number of groups.
[0064] After obtaining the number of groups, the quotient between the number of groups and the preset threshold can be calculated, and the quotient can be rounded up to obtain the target value. For example, assuming that the preset threshold is set to 10,000 groups, the quotient between 134,596 groups and 10,000 groups is calculated, that is, 134,596 divided by 10,000 is approximately 13.4596, and then rounded up to 14, which is the target value.
[0065] After obtaining the target value, the target value can be compared with a preset value, where the preset value can be 2. When the target value is greater than the preset value, it is determined that the relationship matrix needs to be split. When the target value is less than or equal to the preset value, it is determined that the relationship matrix does not need to be split. In this way, the grouping processing strategy of the relationship matrix is adjusted according to the complexity of the actual task, avoiding the waste of computing resources and low execution efficiency caused by the excessive scale of a single task.
[0066] For example, if the target value of 14 is greater than the preset value of 2, the system determines that the current task is too heavy and needs to split the original data into multiple subtasks for parallel processing. Conversely, if the target value is less than or equal to the preset value, the system considers the original data size to be within an acceptable range and no splitting is required.
[0067] This embodiment determines whether the relationship matrix needs to be segmented by calculating the target value, ensuring that when processing large-scale pan-gene set data, computing resources can be efficiently utilized while maintaining analysis accuracy.
[0068] Optionally, in the genetic data processing method provided in an embodiment of the present application, the relationship matrix is segmented according to sub-indicators to obtain P sub-matrices, including: determining the target value as the number of sub-matrices after the relationship matrix is segmented, to obtain the value of P; segmenting the relationship matrix along the gene dimension according to the target value to obtain P sub-matrices, wherein the number of sample objects in each sub-matrix is the same, and the genes contained in each sub-matrix are different.
[0069] It should be noted that the number of sub-matrices is the number of independent data blocks formed after the relationship matrix is divided, that is, the value of P.
[0070] Specifically, after obtaining the target value, the target value can be determined as the number of splits required to split the relationship matrix, and the relationship matrix can be split according to the number of cuts. When performing the splitting operation, the relationship matrix needs to be split along the gene dimension. At this time, the genes in the matrix need to be evenly distributed to ensure that the number of sample objects in each sub-matrix is the same, while the genes contained in each sub-matrix are different, thereby avoiding repeated calculation of gene data while ensuring that the difference in data volume of each sub-matrix is within an acceptable range.
[0071] For example, if the relationship matrix originally contains 1000 genes and 29 sample objects, the system will try to evenly distribute these 1000 genes into 10 sub-matrices, each containing 100 genes, and the total number of sample objects 29 is contained in each sub-matrix.
[0072] This embodiment ensures the consistency of the number of sample objects in each submatrix by determining the target value as the number of submatrices and partitioning the relationship matrix along the gene dimension. At the same time, by dispersing the number of genes, the computational complexity of each subtask is significantly reduced, thereby improving the overall efficiency and resource utilization of pan-gene set statistical analysis.
[0073] Optionally, in the genetic data processing method provided in the embodiment of the present application, the genetic data in each sub-matrix is processed in parallel according to the sub-indicator to obtain P data statistical results, including: for any sub-matrix, determining the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; traversing and grouping the sample objects in the sub-matrix according to the target number to obtain H groups of sample objects, where H is a positive integer; obtaining the genetic data under each group of sample objects from the sub-matrix to obtain H matrices to be processed, and obtaining the first number of core genes and the second number of pan-genes in each matrix to be processed to obtain H groups of first numbers and H groups of second numbers; calculating the first statistical result of the H groups of first numbers and the second statistical result of the H groups of second numbers, wherein the first statistical result includes at least one of the following: the maximum value, minimum value and average value of the H groups of first numbers, and the second statistical result includes at least one of the following: the maximum value, minimum value and average value of the H groups of second numbers; and determining the first statistical result and the second statistical result as one data statistical result.
[0074] Specifically, when processing genetic data, for any submatrix under any sub-indicator, the number of samples that each group of sample objects should contain, that is, the target number, can be determined based on the preset sub-indicator, and grouped according to the target number to obtain H groups of sample objects. The process of traversing the groups is to calculate according to the formula C(M, K) according to the number of sample objects in each group to obtain the number of groups, that is, H groups of sample objects.
[0075] After obtaining H groups of sample objects, since each object has corresponding genetic data, the matrix corresponding to each group of sample objects can be obtained, and H processing matrices can be obtained. Each processing matrix contains the genetic data corresponding to the grouped samples, ensuring the pertinence and accuracy of the statistical analysis. Among them, the genetic data in each matrix corresponds to the genetic data after segmentation.
[0076] After obtaining H matrices to be processed, it is necessary to perform statistical analysis of core genes and pan-genes on each matrix to be processed. Core genes refer to genes that exist in all sample objects, while pan-genes are genes that appear in some samples. The first number of core genes and the second number of pan-genes in each matrix to be processed are obtained to form a set of H groups of first numbers (number of core genes) and H groups of second numbers (number of pan-genes), thereby counting the number of core genes and pan-genes and quantifying the genetic diversity and stability of the sample population.
[0077] The only requirement is that, after obtaining the first and second counts for the H groups, statistical results, such as the maximum, minimum, and average values, need to be calculated based on these first and second counts, so that subsequent analysis and graphing can be performed based on these statistical results. Ultimately, the system combines the first statistical results for the number of core genes and the second statistical results for the number of pan-genes to form a comprehensive data statistical result, ready for further analysis or visualization.
[0078] It should be noted that the data statistical results obtained by this process are the data statistical results of a sub-matrix under a sub-indicator. It is also necessary to combine the data statistical results of multiple sub-matrices to obtain the data statistical results under the sub-indicator, and then combine the data statistical results under multiple sub-indicators into the total data statistical results of this statistical operation.
[0079] This embodiment achieves effective management and efficient analysis of submatrix data by traversing the steps from grouping to statistical analysis and then to result integration. That is, by setting a target number of sample objects for each group, the large task is split into multiple small tasks for parallel calculation, which significantly improves the efficiency of pan-gene set statistical analysis. At the same time, through precise gene data statistics and comprehensive result calculation, the comprehensiveness and accuracy of the analysis are ensured, providing a solid data foundation for subsequent pan-gene set research.
[0080] Optionally, in the genetic data processing method provided in the embodiment of the present application, P data statistical results are combined to obtain genetic data statistical results under the target category and sub-indicator, including: obtaining the maximum value of the first statistical result in the data statistical results of each submatrix, obtaining multiple first maxima, and determining the maximum value of the multiple first maxima as the target first maximum; obtaining the maximum value of the second statistical result in the data statistical results of each submatrix, obtaining multiple second maxima, and determining the maximum value of the multiple second maxima as the target second maximum; obtaining the minimum value of the first statistical result in the data statistical results of each submatrix, obtaining multiple first minimum values, and determining the minimum value of the multiple first minimum values as the target first minimum value ; Obtain the minimum value of the second statistical result in the data statistical result of each submatrix, obtain multiple second minimum values, and determine the minimum value of the multiple second minimum values as the target second minimum value; obtain the average value of the first statistical result in the data statistical result of each submatrix, obtain multiple first average values, and calculate the average value of the multiple first average values to obtain the target first average value; obtain the average value of the second statistical result in the data statistical result of each submatrix, obtain multiple second average values, and calculate the average value of the multiple second average values to obtain the target second average value; determine the target first maximum value, target second maximum value, target first minimum value, target second minimum value, target first average value, and target second average value as the genetic data statistical results.
[0081] Specifically, after completing the statistical analysis of each submatrix, the system extracts the first statistical result of the number of core genes in each submatrix and finds the maximum value, which is the highest number of core genes observed in the corresponding submatrix. This process is repeated for all submatrices, compiling a list containing the maximum number of core genes in each submatrix. This maximum value of the first statistical result is obtained, aiming to identify the peak number of core genes in the entire pan-gene set analysis, providing benchmark data for subsequent population genetic diversity and stability analysis.
[0082] Similarly, as with the core gene count, the system also extracts a secondary statistic for the number of pan-genes in each submatrix and finds the maximum number of pan-genes in each submatrix. This operation is also performed on all submatrices, resulting in a complete list of the maximum number of pan-genes in each submatrix.
[0083] Furthermore, the system also needs to obtain the minimum value of the first statistical result of the number of core genes in each submatrix. Through traversal analysis, the lowest number of core genes in all submatrices is collected and determined. By determining the minimum value of the first statistical result, the lower limit of the number of core genes in the pan-genome is identified. At the same time, the system searches for the minimum value of the second statistical result of the number of pan-genes in each submatrix, forming a list containing the lowest pan-gene counts in all submatrices, and then determines the target second minimum value.
[0084] Finally, the system will extract the average value of the first statistical results on the number of core genes in each submatrix to form a set of average values, and perform a secondary average calculation on all the average values in this set to obtain the target first average value, that is, the global average of the number of core genes in all submatrices of the pan-gene set. It will also extract the average value of the second statistical results on the number of pan-genes in each submatrix to form another set of average values, and then calculate the average value of this set to obtain the target second average value, that is, the global average of the number of pan-genes in all submatrices of the pan-gene set.
[0085] It should be noted that through the above three calculation steps, the key data needed for the final statistical operation can be obtained. Then, by integrating various statistical results, a statistical overview that comprehensively reflects the pan-genome genetic characteristics is provided for subsequent analysis of genetic data, ensuring the accuracy and comprehensiveness of subsequent data analysis.
[0086] It should be noted that when generating Figure 2 When the visual graph is shown, since there may be distribution point maps of core genes and pan-genes in the visual graph, and the number of core genes and pan-genes is large, the gene values can be screened when generating the visual graph, and the number of core genes and pan-genes under each sub-indicator can be arranged in order from large to small, and according to a preset number of groups, for example, 30 groups, they are grouped in order according to the size of the number to obtain 30 groups of gene numbers, and the average value of the number of genes in each group is calculated to obtain 30 average values, and the 30 average values and the maximum and minimum values are determined as the gene number distribution under the sub-indicator, so that the point distribution can still be reflected while reducing the number of points in the distribution point map in the visual graph.
[0087] It should be noted that the curve used to characterize the quantity change trend in the visual graph can be determined by the average value of the gene quantity under each sub-indicator, and other mathematical fitting methods can also be used to generate the trend curve, which is not limited in this embodiment.
[0088] Optionally, in the genetic data processing method provided in the embodiment of the present application, the genetic data statistical results under each sub-indicator are combined to obtain the total genetic data statistical results under the target category, including: adding the genetic data statistical results under each sub-indicator to a preset schematic diagram according to the grouping requirements in the sub-indicators to obtain a genetic data statistical graph; drawing a curve according to the average value of the genetic data statistical results under each sub-indicator, and calculating the slope of the curve at each point to obtain multiple slopes; determining the validity result of the genetic data statistical graph according to the slope values of the multiple slopes, and determining the genetic data statistical graph and the validity result as the total genetic data statistical result.
[0089] Specifically, after obtaining the statistical results of genetic data under each sub-indicator, the system adds the statistical results (including maximum, minimum and average) of the number of core genes and pan-genes under each sub-indicator calculated previously to a pre-designed schematic template based on the predefined sub-indicators. For example, the average number of core genes under different sub-indicators can be plotted as points on the X-axis, and the Y-axis position of each point represents the average value under this indicator.
[0090] Furthermore, the system can analyze the average number of core genes and pan-genes for each sub-indicator. Based on these averages, the system can draw curves on its preset schematic diagram, which reflect the changing trend of the average number of genes as sub-indicators (such as sample group size, gene family type, etc.) change. By displaying the statistical results of genetic data in graphical form, it is possible to intuitively compare the distribution of gene numbers under different sub-indicators and conduct trend analysis.
[0091] It should be noted that after obtaining the curve depicting the trend of changes in the average value, the system can also calculate the slope of the curve at different points, thereby evaluating the validity of the genetic data statistical graph by analyzing multiple slope values. For example, if the slope value shows an obvious trend or pattern, such as a steady increase or decrease in the slope, this indicates that the number of sample objects used in the statistical graph is insufficient, resulting in non-convergence of the data, and the number of sample objects needs to be increased. Conversely, if the slope value gradually decreases to a stable state, indicating that the current data has converged, then the number of sample objects is sufficient, and the statistical data obtained this time can be used to analyze the genetic information of objects under the target category, thereby ensuring the reliability of the statistical analysis results and the accuracy of the visual presentation by evaluating the validity of the genetic data statistical graph.
[0092] Finally, the system combines the above-generated genetic data statistical graphs and their validity evaluation results to form a complete genetic data statistical total result. The genetic data statistical total result includes intuitive graphical representation and objective evaluation of the quality of statistical analysis, achieving the technical effect of providing both intuitive and rigorous data analysis results, thereby completing the accurate and efficient analysis process of genetic data.
[0093] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0094] The present application also provides a genetic data processing device. It should be noted that the genetic data processing device of the present application can be used to execute the genetic data processing method provided in the present application. The genetic data processing device provided in the present application is described below.
[0095] Figure 3 Schematic diagram of a gene data processing device according to an embodiment of the present application. Figure 3 As shown, the device includes: an acquisition unit 31, a judgment unit 32, a first processing unit 33, a second processing unit 34, and a combination unit 35.
[0096] The acquisition unit 31 is configured to acquire gene data sets of M sample objects under a target category, to obtain M gene data sets, wherein each gene data set contains gene data of N genes, where M and N are positive integers.
[0097] The judgment unit 32 is used to construct a relationship matrix between M sample objects and N genes, and to judge whether the relationship matrix needs to be segmented based on each sub-indicator in the gene statistical index, wherein the gene statistical index includes T sub-indicators, each sub-indicator is used to indicate the grouping requirements for statistical analysis of the sample data in the relationship matrix.
[0098] The first processing unit 33 is used to split the relationship matrix according to the sub-indicators for which the relationship matrix needs to be split, obtain P sub-matrices, and parallelly process the genetic data in each sub-matrix according to the sub-indicators to obtain P data statistical results, and combine the P data statistical results to obtain the genetic data statistical results under the target category and sub-indicators, where P is a positive integer.
[0099] The second processing unit 34 is configured to process the relationship matrix according to the sub-indicators that do not require segmentation of the relationship matrix, and obtain the statistical results of the gene data under the target category and the sub-indicators.
[0100] The combining unit 35 is used to combine the genetic data statistical results under various sub-indicators to obtain the overall genetic data statistical results under the target category.
[0101] The genetic data processing device provided in the embodiment of the present application obtains the genetic data sets of M sample objects under the target category through the acquisition unit 31, and obtains M genetic data sets, wherein each genetic data set contains genetic data of N genes, and M and N are positive integers; the judgment unit 32 constructs a relationship matrix between the M sample objects and the N genes, and judges whether it is necessary to segment the relationship matrix according to each sub-indicator in the genetic statistical index, wherein the genetic statistical index includes T sub-indicators, each sub-indicator is used to indicate the grouping requirements for statistical analysis of the sample data in the relationship matrix; the first processing unit 33 is for the relationship matrix that needs to be segmented. The segmented sub-indicators are segmented according to the sub-indicators to obtain P sub-matrices, and the gene data in each sub-matrix is processed in parallel according to the sub-indicators to obtain P data statistical results. The P data statistical results are then combined to obtain the gene data statistical results under the target category and sub-indicator, where P is a positive integer. The second processing unit 34 processes the relationship matrix according to the sub-indicators for sub-indicators that do not require segmentation of the relationship matrix to obtain the gene data statistical results under the target category and sub-indicator. The combining unit 35 combines the gene data statistical results under each sub-indicator to obtain the total gene data statistical results under the target category. This solves the problem of low efficiency in statistical analysis of gene data of sample objects in related technologies. By constructing a relationship matrix between sample objects and genes and determining whether to segment the relationship matrix based on the sub-indicators in the gene statistical indicators, when the amount of sample object and gene data is large, the relationship matrix can be segmented and each sub-matrix is processed separately, and the statistical results of the processed sub-matrices are combined into the statistical results of the relationship matrix, thereby achieving the technical effect of improving the processing efficiency of gene data.
[0102] Optionally, in the genetic data processing device provided in the embodiment of the present application, the acquisition unit 31 includes: a first acquisition module, used to obtain the category information of each initial object from the object database, and screen the initial objects whose category information is the target category, to obtain M sample objects, wherein the object database contains multiple objects to be tested; a first determination module, used to detect the number of occurrences of each gene carried in each sample object, and determine the number of occurrences as the genetic data of each gene, to obtain a genetic data set for each sample object.
[0103] Optionally, in the genetic data processing device provided in the embodiment of the present application, the judgment unit 32 includes: a second determination module, used to determine, for any sub-indicator, the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; a third determination module, used to determine the number of groups of sample objects under the grouping requirements based on the target number and the total number of sample objects, and determine a preset threshold; a first calculation module, used to calculate the quotient between the number of groups and the preset threshold, and round up the quotient to obtain a target value; a judgment module, used to determine whether the target value is greater than a preset value, and if the target value is greater than the preset value, determine that the relationship matrix needs to be segmented; a fourth determination module, used to determine that the relationship matrix does not need to be segmented if the target value is less than or equal to the preset value.
[0104] Optionally, in the genetic data processing device provided in the embodiment of the present application, the first processing unit 33 includes: a fifth determination module, used to determine the target value as the number of sub-matrices after dividing the relationship matrix, and obtain a value of P; a segmentation module, used to segment the relationship matrix along the gene dimension according to the target value, and obtain P sub-matrices, wherein the number of sample objects in each sub-matrix is the same, and the genes contained in each sub-matrix are different.
[0105] Optionally, in the genetic data processing device provided in the embodiment of the present application, the first processing unit 33 includes: a sixth determination module, which is used to determine, for any sub-matrix, the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; a grouping module, which is used to traverse and group the sample objects in the sub-matrix according to the target number to obtain H groups of sample objects, where H is a positive integer; a second acquisition module, which is used to obtain genetic data under each group of sample objects from the sub-matrix to obtain H matrices to be processed, and obtain the first number of core genes and the second number of pan-genes in each matrix to be processed to obtain H groups of first numbers and H groups of second numbers; a second calculation module, which is used to calculate a first statistical result of the H groups of first numbers and a second statistical result of the H groups of second numbers, wherein the first statistical result includes at least one of the following: the maximum value, minimum value and average value of the H groups of first numbers, and the second statistical result includes at least one of the following: the maximum value, minimum value and average value of the H groups of second numbers; a seventh determination module, which is used to determine the first statistical result and the second statistical result as one data statistical result.
[0106] Optionally, in the genetic data processing device provided in the embodiment of the present application, the first processing unit 33 includes: a third acquisition module, used to obtain the maximum value of the first statistical result in the data statistical results of each submatrix, obtain multiple first maximum values, and determine the maximum value of the multiple first maximum values as the target first maximum value; a fourth acquisition module, used to obtain the maximum value of the second statistical result in the data statistical results of each submatrix, obtain multiple second maximum values, and determine the maximum value of the multiple second maximum values as the target second maximum value; a fifth acquisition module, used to obtain the minimum value of the first statistical result in the data statistical results of each submatrix, obtain multiple first minimum values, and determine the minimum value of the multiple first minimum values as the target first minimum value; a sixth acquisition module, used to obtain the maximum value of the second statistical result in the data statistical results of each submatrix, obtain multiple first minimum values, and determine the minimum value of the multiple first minimum values as the target first minimum value. the minimum value of the second statistical result in the data statistical result of each submatrix, obtain multiple second minimum values, and determine the minimum value of the multiple second minimum values as the target second minimum value; a seventh acquisition module, used to obtain the average value of the first statistical result in the data statistical result of each submatrix, obtain multiple first average values, and calculate the average value of the multiple first average values to obtain the target first average value; an eighth acquisition module, used to obtain the average value of the second statistical result in the data statistical result of each submatrix, obtain multiple second average values, and calculate the average value of the multiple second average values to obtain the target second average value; an eighth determination module, used to determine the target first maximum value, the target second maximum value, the target first minimum value, the target second minimum value, the target first average value, and the target second average value as the genetic data statistical results.
[0107] Optionally, in the genetic data processing device provided in the embodiment of the present application, the combination unit 35 includes: an adding module, used to add the genetic data statistical results under each sub-indicator to a preset schematic diagram according to the grouping requirements in the sub-indicators to obtain a genetic data statistical graph; a third calculation module, used to draw a curve based on the average value of the genetic data statistical results under each sub-indicator, and calculate the slope of the curve at each point to obtain multiple slopes; a ninth determination module, used to determine the validity result of the genetic data statistical graph based on the slope values of the multiple slopes, and determine the genetic data statistical graph and the validity result as the total genetic data statistical result.
[0108] The above-mentioned genetic data processing device includes a processor and a memory. The above-mentioned acquisition unit 31, judgment unit 32, first processing unit 33, second processing unit 34, combination unit 35, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize the corresponding functions.
[0109] The processor includes a core, which retrieves the corresponding program unit from the memory. One or more cores can be set, and by adjusting the core parameters, the problem of low efficiency in statistical analysis of genetic data of sample objects in related technologies is solved.
[0110] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0111] An embodiment of the present invention provides a computer-readable storage medium having a program stored thereon, which implements the method for processing genetic data when executed by a processor.
[0112] An embodiment of the present invention provides a processor, which is used to run a program, wherein the method for processing genetic data is executed when the program is run.
[0113] Figure 4 is a schematic diagram of an electronic device provided according to an embodiment of the present application, such as Figure 4 As shown, an embodiment of the present invention provides an electronic device. The electronic device 40 includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps of the above-described method for processing genetic data are implemented. The device herein may be a server, a PC, a PAD, a mobile phone, or the like.
[0114] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program initialized with the steps of the above-mentioned genetic data processing method.
[0115] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0116] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.
[0117] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0119] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0120] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0121] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0122] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0123] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for processing genetic data, characterized in that: include: Obtaining gene data sets of M sample objects under the target category to obtain M gene data sets, wherein each gene data set contains gene data of N genes, where M and N are positive integers; Constructing a relationship matrix between the M sample objects and the N genes, and determining whether the relationship matrix needs to be segmented according to each sub-indicator in the gene statistical index, wherein the gene statistical index includes T sub-indicators, each sub-indicator is used to indicate a grouping requirement for statistical analysis of the sample data in the relationship matrix, and T is a positive integer; For sub-indicators that require segmentation of the relationship matrix, segment the relationship matrix according to the sub-indicators to obtain P sub-matrices, and perform parallel processing on the gene data in each sub-matrix according to the sub-indicators to obtain P data statistical results, and combine the P data statistical results to obtain the gene data statistical results under the target category and the sub-indicator, where P is a positive integer; For sub-indicators that do not require segmentation of the relationship matrix, the relationship matrix is processed according to the sub-indicators to obtain the target category and the genetic data statistical results under the sub-indicators; The genetic data statistical results under each sub-indicator are combined to obtain the overall genetic data statistical results under the target category.
2. The method according to claim 1, characterized in that Obtain a gene data set of M sample objects under the target category, and obtain the M gene data sets including: Acquire category information of each initial object from an object database, and filter the initial objects whose category information is the target category to obtain the M sample objects, wherein the object database contains multiple objects to be tested; The number of occurrences of each gene carried in each sample object is detected, and the number of occurrences is determined as the gene data of each gene to obtain a gene data set for each sample object.
3. The method according to claim 1, characterized in that Judging whether the relationship matrix needs to be segmented according to each sub-indicator in the gene statistical indicator includes: For any sub-indicator, determine the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicator; Determining the number of groups of sample objects under the grouping requirement according to the target number and the total number of sample objects, and determining a preset threshold; Calculating a quotient between the number of groups and the preset threshold, and rounding up the quotient to obtain a target value; determining whether the target value is greater than a preset value, and determining that the relationship matrix needs to be segmented if the target value is greater than the preset value; When the target value is less than or equal to the preset value, it is determined that the relationship matrix does not need to be segmented.
4. The method according to claim 3, characterized in that The relationship matrix is divided according to the sub-indicators to obtain P sub-matrices including: Determine the target value as the number of submatrices after dividing the relationship matrix, and obtain the value of P; The relationship matrix is divided along the gene dimension according to the target value to obtain the P sub-matrices, wherein the number of sample objects in each sub-matrix is the same and the genes contained in each sub-matrix are different.
5. The method according to claim 1, wherein The gene data in each sub-matrix is processed in parallel according to the sub-indicators to obtain P data statistical results including: For any sub-matrix, determining the target number of sample objects in each group of sample objects according to the grouping requirements in the sub-indicators; Traversing and grouping the sample objects in the submatrix according to the target number to obtain H groups of sample objects, where H is a positive integer; Obtaining gene data under each group of sample objects from the submatrix to obtain H matrices to be processed, and obtaining a first number of core genes and a second number of pan-genes in each matrix to be processed to obtain H groups of first numbers and H groups of second numbers; Calculating a first statistical result of the H groups of first quantities and a second statistical result of the H groups of second quantities, wherein the first statistical result includes at least one of the following: a maximum value, a minimum value, and an average value of the H groups of first quantities, and the second statistical result includes at least one of the following: a maximum value, a minimum value, and an average value of the H groups of second quantities; The first statistical result and the second statistical result are determined as one data statistical result.
6. The method according to claim 5, characterized in that Combining the P data statistical results to obtain the gene data statistical results under the target category and the sub-indicator includes: Obtaining a maximum value of the first statistical result in the data statistical results of each submatrix to obtain multiple first maximum values, and determining the maximum value of the multiple first maximum values as a target first maximum value; Obtaining a maximum value of the second statistical result in the data statistical results of each submatrix to obtain multiple second maximum values, and determining the maximum value of the multiple second maximum values as a target second maximum value; Obtaining a minimum value of the first statistical result in the data statistical results of each submatrix to obtain multiple first minimum values, and determining the minimum value among the multiple first minimum values as a target first minimum value; Obtaining a minimum value of the second statistical result in the data statistical results of each submatrix to obtain multiple second minimum values, and determining the minimum value among the multiple second minimum values as a target second minimum value; Obtaining an average value of the first statistical results in the data statistical results of each submatrix to obtain multiple first average values, and calculating an average value of the multiple first average values to obtain a target first average value; Obtaining an average value of the second statistical results in the data statistical results of each submatrix to obtain multiple second average values, and calculating an average value of the multiple second average values to obtain a target second average value; The target first maximum value, the target second maximum value, the target first minimum value, the target second minimum value, the target first average value, and the target second average value are determined as the genetic data statistical results.
7. The method according to claim 1, characterized in that The genetic data statistical results under each sub-indicator are combined to obtain the overall genetic data statistical results under the target category, including: Adding the genetic data statistical results under each sub-indicator to a preset schematic diagram according to the grouping requirements in the sub-indicators to obtain a genetic data statistical diagram; Draw a curve based on the average value of the genetic data statistical results under each sub-indicator, and calculate the slope of the curve at each point to obtain multiple slopes; The validity result of the gene data statistical graph is determined according to the slope values of the multiple slopes, and the gene data statistical graph and the validity result are determined as the gene data statistical total result.
8. A genetic data processing device, characterized in that: include: an acquisition unit, configured to acquire gene data sets of M sample objects under a target category, to obtain M gene data sets, wherein each gene data set contains gene data of N genes, where M and N are positive integers; a judgment unit, configured to construct a relationship matrix between the M sample objects and the N genes, and to judge whether the relationship matrix needs to be segmented according to each sub-indicator in the gene statistical index, wherein the gene statistical index includes T sub-indicators, each sub-indicator is used to indicate a grouping requirement for statistical analysis of the sample data in the relationship matrix, and T is a positive integer; The first processing unit is configured to segment the relationship matrix according to the sub-indicators for which the relationship matrix needs to be segmented, thereby obtaining P sub-matrices, and to perform parallel processing on the gene data in each sub-matrix according to the sub-indicators to obtain P data statistical results, and to combine the P data statistical results to obtain the gene data statistical results under the target category and the sub-indicators, wherein: P is a positive integer; A second processing unit is configured to process the relationship matrix according to the sub-indicators that do not require segmentation of the relationship matrix, to obtain the target category and the genetic data statistical results under the sub-indicators; The combining unit is used to combine the statistical results of the genetic data under each sub-indicator to obtain the total statistical result of the genetic data under the target category.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the genetic data processing method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: a memory storing an executable program; A processor, configured to run the program, wherein the program, when running, executes the method for processing genetic data according to any one of claims 1 to 7.
Citation Information
Patent Citations
Construction method of generic genome, terminal equipment and storage medium
CN117037912A
Gene regulation relation prediction method based on matrix enhancement and feature fusion
CN117789831A
K-MER based strain typing
US20170364666A1