Typing clustering method and device of whole genome sequence data, terminal equipment and storage medium

By hashing the entire genome sequence data and building a minimum spanning tree network diagram, the problem of high pressure on computing and storage of central databases and data privacy in the existing technology is solved, and efficient data analysis and privacy protection are achieved.

CN119961706AInactive Publication Date: 2025-05-09CHINA NAT CENT FOR FOOD SAFETY RISK ASSESSMENT
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510442652.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When processing genome sequence data in the prior art, the central database bears excessive computing and storage pressure, and at the same time, there are data privacy issues for users to upload complete gene sequences.

Method used

By hashing the genome sequence data, a hash list database is generated, and a relative distance matrix is ​​constructed based on the hash list database, and the minimum spanning tree network map is determined to realize data alignment and clustering.

Benefits of technology

It reduces the server's workload on genome-wide sequence data, realizes decentralized data analysis, reduces the sensitivity of data transmission and sharing, and improves data privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961706A_ABST
    Figure CN119961706A_ABST
Patent Text Reader

Abstract

The invention discloses a typing and clustering method and device for whole genome sequence data, terminal equipment and a storage medium. The typing and clustering method comprises the steps of obtaining genome sequence data; performing Hash operation on the genome sequence data to obtain a Hash list corresponding to the genome sequence data, and generating a Hash list database; according to the Hash list database, Hash lists in the Hash list data are processed, and a relative distance matrix between the Hash lists is obtained; according to the relative distance matrix, a minimum spanning tree network diagram corresponding to the genome sequence data is determined, the minimum spanning tree network diagram is used for comparing and clustering the genome sequence data, data analysis is carried out on a local terminal of a user, support of a central server is not needed, and decentralization is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of gene technology, and in particular relates to a typing and clustering method, apparatus, terminal device and storage medium for whole genome sequence data. Background Art

[0002] cgMLST and wgMLST are commonly used methods for accurate typing of biological whole genomes and genetic tracing and clustering. They are often used for pathogen outbreak tracking, etc. The traditional cg / wgMLST method requires a central database and computing server to record the type and sequence data of gene loci. When users compare sample sequences, they need to upload the sample sequences to the server to complete the sequence annotation and type comparison of each gene locus. This method brings greater storage and computing pressure to the central database, and for users, what is uploaded is the complete gene sequence, which is not conducive to data privacy.

[0003] Currently, there is a distributed cgMLST method based on the md5 hash table conversion method. This method distributes the calculation of genotype to the user end, which reduces the computational pressure of the central database on the one hand, and reduces the sensitivity of data transmission and sharing between laboratories on the other hand, because there is no need to upload the complete genome data to the server, which can reduce users' concerns about data confidentiality. However, this method still needs to be compared with the gene locus reference annotation sequence provided by the central database to determine the gene locus name, so it cannot achieve complete decentralization. How to reduce the workload of the server on the whole genome sequence data is an urgent problem to be solved. Summary of the invention

[0004] The present application aims to provide a typing and clustering method, apparatus, terminal device and storage medium for whole genome sequence data to solve the deficiencies in the prior art. The technical problem to be solved by the present application is achieved through the following technical solutions.

[0005] In a first aspect, an embodiment of the present application provides a typing and clustering method for whole genome sequence data, the method comprising: Obtain genome sequence data; Performing a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generating a hash list database; According to the hash list database, the hash lists in the hash list data are processed to obtain a relative distance matrix between the hash lists; According to the relative distance matrix, a minimum spanning tree network diagram corresponding to the genome sequence data is determined, and the minimum spanning tree network diagram is used to align and cluster the genome sequence data.

[0006] Optionally, performing a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data and generating a hash list database includes: Preprocessing the genome sequence data to obtain a processed predicted coding sequence; Performing hash mapping on each coding sequence in the predicted coding sequence to obtain multiple hash lists corresponding to the genome sequence data; A plurality of hash lists corresponding to the genome sequences are stored in a hash list database, wherein the hash list database includes file names and hash lists corresponding to the file names.

[0007] Optionally, the processing of the hash lists in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists includes: Calculating the intersection length of the hash lists in the hash list database to obtain the number of matching coding sequences; According to the number of the matched coding sequences, a matching number matrix is ​​constructed; The length of the hash list of the genome sequence data in the hash list database is used as the total number of coding sequences of the genome; Determining the absolute distance between the genome sequence data based on the total number of coding sequences in the genome and the number of matching coding sequences; Determining the relative distance of the genome sequence data according to the absolute distance between the genome sequence data and the total coding sequence quantity of the genome; According to the relative distances of the genome sequences, a relative distance matrix is ​​determined.

[0008] Optionally, determining a minimum spanning tree network graph corresponding to the genome sequence data according to the relative distance matrix comprises: Taking a relatively small value of the symmetric position of the relative distance matrix to obtain a genome pair distance list; Generate a minimum spanning tree according to the genome pairwise distance list; According to the minimum spanning tree, a minimum spanning tree network graph is determined.

[0009] Optionally, the method further comprises: Merging the plurality of hash list databases to obtain a merged hash list database; and / or A plurality of the relative distance matrices are merged to obtain a merged distance matrix list.

[0010] Optionally, the method further comprises: Obtain new genome sequence data; The minimum spanning tree network diagram is updated according to the new genome sequence data.

[0011] Optionally, the hash list database includes at least one of JSON data format, text format separated by any characters, table format or SQL database format.

[0012] In a second aspect, an embodiment of the present application provides a typing and clustering device for whole genome sequence data, the device comprising: An acquisition module, used to acquire genome sequence data; A calculation module, used for performing a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generating a hash list database; A processing module, used for processing the hash lists in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists; A generation module is used to determine a minimum spanning tree network diagram corresponding to the genome sequence data based on the relative distance matrix, and the minimum spanning tree network diagram is used to align and cluster the genome sequence data.

[0013] Optionally, the computing module is used to: Preprocessing the genome sequence data to obtain a processed predicted coding sequence; Performing hash mapping on each coding sequence in the predicted coding sequence to obtain multiple hash lists corresponding to the genome sequence data; A plurality of hash lists corresponding to the genome sequences are stored in a hash list database, wherein the hash list database includes file names and hash lists corresponding to the file names.

[0014] Optionally, the processing module is used to: Calculating the intersection length of the hash lists in the hash list database to obtain the number of matching coding sequences; According to the number of the matched coding sequences, a matching number matrix is ​​constructed; The length of the hash list of the genome sequence data in the hash list database is used as the total number of coding sequences of the genome; Determining the absolute distance between the genome sequence data based on the total number of coding sequences in the genome and the number of matching coding sequences; Determining the relative distance of the genome sequence data according to the absolute distance between the genome sequence data and the total coding sequence quantity of the genome; According to the relative distances of the genome sequences, a relative distance matrix is ​​determined.

[0015] Optionally, the generating module is used to: Taking a relatively small value of the symmetric position of the relative distance matrix to obtain a genome pair distance list; Generate a minimum spanning tree according to the genome pairwise distance list; According to the minimum spanning tree, a minimum spanning tree network graph is determined.

[0016] Optionally, the generating module is used to: Merging the plurality of hash list databases to obtain a merged hash list database; and / or A plurality of the relative distance matrices are merged to obtain a merged distance matrix list.

[0017] Optionally, the generating module is used to: Obtain new genome sequence data; The minimum spanning tree network diagram is updated according to the new genome sequence data.

[0018] Optionally, the hash list database includes at least one of JSON data format, text format separated by any characters, table format or SQL database format.

[0019] In a third aspect, an embodiment of the present application provides a terminal device, including: at least one processor and a memory; The memory stores a computer program; the at least one processor executes the computer program stored in the memory to implement the typing and clustering method for whole genome sequence data provided in the first aspect.

[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed, the typing and clustering method of whole genome sequence data provided in the first aspect is implemented.

[0021] The embodiments of the present application include the following advantages: The typing and clustering method, device, terminal device and storage medium of the whole genome sequence data provided in the embodiment of the present application are obtained by obtaining genome sequence data; performing hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generating a hash list database; processing the hash list in the hash list data according to the hash list database to obtain a relative distance matrix between hash lists; determining the minimum spanning tree network diagram corresponding to the genome sequence data according to the relative distance matrix, and the minimum spanning tree network diagram is used to compare and cluster the genome sequence data. In the embodiment of the present application, the terminal device directly predicts the whole genome coding sequence of the sample genome assembly sequence (fasta format), and performs hash calculation on each coding sequence, and does not need to compare the gene position information of each coding sequence and directly forms a hash list as a mapping of the sample genome. The relative distance matrix between samples is calculated by comparing the intersection of the hash value lists between different samples, and the minimum spanning tree network diagram corresponding to the genome sequence data is determined according to the relative distance matrix, and the minimum spanning tree network diagram is used to compare and cluster the genome sequence data, so as to realize data analysis at the user's local terminal without the support of the central server, and realize decentralization. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application or the existing technical solutions, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0023] Figure 1 This is a flow chart of a typing and clustering method for whole genome sequence data in one embodiment of the present application; Figure 2 A flowchart of another method for clustering whole genome sequence data in one embodiment of the present application; Figure 3 This is a schematic diagram of a test comparison in an embodiment of the present application; Figure 4 It is a structural block diagram of an embodiment of a typing and clustering device for whole genome sequence data of the present application; Figure 5 It is a structural diagram of a terminal device of the present application. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0025] An embodiment of the present application provides a typing and clustering method for whole genome sequence data, which is used to compare and cluster biological whole genome data based on hash table encoded sequence typing. The execution subject of this embodiment is a typing and clustering device for whole genome sequence data, which is set on a terminal device, for example, the terminal device at least includes a computer terminal, etc.

[0026] Reference Figure 1 , shows a flow chart of steps of an embodiment of a typing and clustering method for whole genome sequence data of the present application, which method may specifically include the following steps: S101, obtaining genome sequence data; Specifically, genome sequence data is acquired on the terminal device, and the genome sequence data may be genome sequence data of different types, such as a Salmonella genome.

[0027] S102, performing a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generating a hash list database; The terminal device pre-processes the genome sequence data, including excluding sequences with degenerate sites (such as N), deleting low-quality sequences, and deleting sequences that are too short or too long, to obtain processed genome sequence data; A hash operation is performed on each coding sequence in the processed genome sequence data to generate a hash list database. The hash mapping described herein includes but is not limited to MD5 (MD5 Message-Digest Algorithm) mapping.

[0028] S103, processing the hash lists in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists; Specifically, the terminal device calculates the intersection length of each hash list in the hash list database to obtain the number of matching coding sequences and construct a matching number matrix; then obtains the hash list length of the genome sequence data in the hash list database and uses it as the total number of coding sequences of the genome. In this way, the absolute distance between the genome sequence data can be determined based on the total number of coding sequences of the genome and the number of matching coding sequences, and then the relative distance of the genome sequence data can be determined based on the absolute distance between the genome sequence data and the total number of coding sequences of the genome; and the relative distance matrix is ​​determined based on the relative distance of the genome sequences.

[0029] S104. Determine a minimum spanning tree network diagram corresponding to the genome sequence data according to the relative distance matrix. The minimum spanning tree network diagram is used for aligning and clustering the genome sequence data.

[0030] Specifically, the terminal device takes a relatively small value of the symmetric position of the relative distance matrix to obtain a genome pairwise distance list; generates a minimum spanning tree based on the genome pairwise distance list; and determines a minimum spanning tree network diagram based on the minimum spanning tree.

[0031] Since the mapping results of MD5 in any computer are the same, each sample genome sequence theoretically corresponds to a unique coding sequence hash list. Therefore, data analysis can be performed locally without the support of a central server.

[0032] Since the MD5 value cannot be restored to the original text, and because the comparison of the method provided in the embodiment of the present application does not include locus information, the database of the shared coding sequence hash list is used as a method for exchanging information on strain similarity comparisons between laboratories, which can avoid direct exposure of the whole genome data and ensure data privacy to a great extent.

[0033] The typing and clustering method of the whole genome sequence data provided in the embodiment of the present application is obtained by obtaining the genome sequence data; performing a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generating a hash list database; processing the hash list in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists; determining the minimum spanning tree network diagram corresponding to the genome sequence data according to the relative distance matrix, and the minimum spanning tree network diagram is used to compare and cluster the genome sequence data. In the embodiment of the present application, the terminal device directly predicts the whole genome coding sequence of the sample genome assembly sequence (fasta format), and performs a hash operation on each coding sequence, and does not need to compare the gene position information of each coding sequence and directly forms a hash list as a mapping of the sample genome. The relative distance matrix between samples is calculated by comparing the intersection of the hash value lists between different samples, and the minimum spanning tree network diagram corresponding to the genome sequence data is determined according to the relative distance matrix. The minimum spanning tree network diagram is used to compare and cluster the genome sequence data, and data analysis is performed on the user's local terminal without the support of the central server, thereby achieving decentralization.

[0034] Another embodiment of the present application further supplements the typing and clustering method of whole genome sequence data provided in the above embodiment.

[0035] Optionally, a hash operation is performed on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and a hash list database is generated, including: Preprocessing the genome sequence data to obtain a processed predicted coding sequence; Performing hash mapping on each coding sequence in the predicted coding sequence to obtain multiple hash lists corresponding to the genome sequence data; A plurality of hash lists corresponding to genome sequences are stored in a hash list database, wherein the hash list database includes file names and hash lists corresponding to the file names.

[0036] Optionally, the hash list database includes at least one of JSON data format, arbitrary character-delimited text format, table format or SQL database format.

[0037] Specifically, the terminal device performs quality control on the input sequence, including excluding sequences with degenerate sites (such as N) from the input sequence, deleting low-quality sequences, deleting sequences that are too short or too long, etc., and performing hash mapping on each coded sequence in each input file (ignoring the sequence name) to generate a hash list. The hash mapping described here includes but is not limited to MD5 mapping, and the hash list of all input samples is saved as a database containing file names and hash list information. The database described here includes any form such as JSON data format, text format separated by any characters, table format, SQL database format, etc. This database can be used for data storage, data exchange, etc.

[0038] Optionally, according to the hash list database, the hash lists in the hash list data are processed to obtain a relative distance matrix between the hash lists, including: Calculating the intersection length of the hash lists in the hash list database to obtain the number of matching coding sequences; According to the number of matched coding sequences, a matching number matrix is ​​constructed; The length of the hash list of the genome sequence data in the hash list database is used as the total number of coding sequences of the genome; Determine the absolute distance between genomic sequence data based on the total number of coding sequences in the genome and the number of matching coding sequences; Determine the relative distance of the genomic sequence data based on the absolute distance between the genomic sequence data and the total coding sequence amount of the genome; According to the relative distances of the genome sequences, a relative distance matrix is ​​determined.

[0039] Optionally, determining a minimum spanning tree network graph corresponding to the genome sequence data according to the relative distance matrix includes: Take the relatively small value of the symmetric position of the relative distance matrix to obtain a list of genome pair distances; Generate a minimum spanning tree based on the genome pairwise distance list; Based on the minimum spanning tree, determine the minimum spanning tree network graph.

[0040] Optionally, the method further comprises: Merging multiple hash list databases to obtain a merged hash list database; and / or Multiple relative distance matrices are merged to obtain a merged distance matrix list.

[0041] In addition to constructing a minimum spanning tree from scratch, the embodiment of the present application also provides a method for merging databases, including: supporting the merging of multiple constructed sample hash list databases. The database described herein includes any form such as JSON data form, text form separated by any characters, table form, SQL database form, etc. Supporting the merging of multiple distance matrices and generating a new distance matrix table at the same time. The matrix described herein includes any form such as JSON data form, text form separated by any characters, table form, SQL database form, etc.

[0042] Optionally, the method further comprises: Obtain new genome sequence data; The minimum spanning tree network diagram is updated according to the new genome sequence data.

[0043] Specifically, the terminal device compares the newly input coding sequence file with the constructed hash list database file and / or distance matrix, and outputs the sample information in the database closest to the input file and its relative distance.

[0044] Recalculate the distance matrix for the new input coding sequence file and the constructed hash list database file, and obtain the updated minimum spanning tree, including a method of continuously updating a constructed sample relationship tree in this way.

[0045] Figure 2 FIG. 1 is a flow chart of a policy cache in an embodiment of the present application. Figure 2 As shown, the method also includes: 1. Predict the coding sequence for the genome assembly sequence. If the analyzed sequence is a eukaryotic sequence, it also includes exon prediction. The genome assembly file includes FASTA format. The generated CDS sequence includes common gene coding formats such as FASTA format, GBK format, GFF3 format, etc.

[0046] 2. Quality control of input sequences, including excluding sequences with degenerate sites (such as N), deleting low-quality sequences, and deleting sequences that are too short or too long.

[0047] 3. Perform hash mapping on each encoded sequence in each input file (ignore the sequence name) to generate a hash list. The hash mapping described here includes but is not limited to MD5 mapping. Hash algorithm is an algorithm that compresses data of any length into a data summary of a fixed length. The result is called a hash value or hash value. Output fixed length; Determinism: For the same input, the hash function always generates the same hash value, which can be used to verify the introduction of data integrity issues; Irreversibility (one-way): The hash function is irreversible, that is, it is difficult to derive the original input data from the hash value. Security, anti-hash collision: Anti-hash collision means that the probability of different inputs producing the same hash value is very low. Avalanche effect: Even if the input data changes slightly, the hash value will change significantly.

[0048] 4. Save the hash list of all input samples as a database containing the file name and hash list information. The database described here includes any form such as JSON data format, text format separated by any characters, table format, SQL database format, etc. This database can be used for data storage, data exchange, etc.

[0049] 5. By calculating the intersection length of the hash lists of each sample, the number of matching coding sequences between each sample can be obtained.

[0050] 6. Construct a matching number matrix for the number of coding sequences that match each other in all sample sets. The matrix here includes any form such as JSON data format, text format separated by any characters, table format, SQL database format, etc.

[0051] 7. The length of the sample hash list is used as the total number of coded sequences of the sample. The absolute distance between two samples = the total number of CDSs - the number of CDS sequences matching the sample, which can be equivalent to the length of the sample hash list - the length of the intersection of the hash lists. Construct a distance matrix for the absolute distance of the samples. The matrix described here includes any form such as JSON data format, text format separated by any characters, table format, SQL database format, etc.

[0052] 8. The non-matching absolute distance matrix samples are standardized by the following formula, and a relative distance matrix is ​​constructed: relative distance = sample absolute distance / total number of CDS. The matrix described here includes any form such as JSON data format, text format separated by any characters, table format, SQL database format, etc. Hierarchical clustering can be constructed using this distance matrix.

[0053] 9. Take the relatively small value of the symmetric position of the relative distance matrix and convert it into a sample pairwise distance list. The list here includes any form such as JSON data format, text format separated by any characters, table format, SQL database format, etc.

[0054] 10. Calculate the minimum spanning tree using this distance list.

[0055] 11. Convert the minimum spanning tree into a network graph.

[0056] The embodiment of the present application provides a multi-site sequence typing method (hash-CDST) that directly performs hash mapping on the genome coding sequence locally, without the need to annotate the gene position. This method skips the gene position alignment process, directly predicts the whole genome coding sequence for the sample genome assembly sequence (fasta format), and calculates the MD5 hash value for each coding sequence. The gene position information of each coding sequence is not compared, but a hash list is directly formed as a mapping of the sample genome. The evolutionary distance between samples is calculated by comparing the intersection of the MD5 hash value lists between different samples.

[0057] The reason why the MD5 values ​​mapped to the coding sequences of all gene bits can be put in a list and compared at the same time is that the probability of random collision of MD5 mapping is extremely low. The output of MD5 mapping is a 128-bit hash value, so there are 2 possible hash values. 128 For two different strings, the probability of their MD5 values ​​​​colliding is 1 / 2 128 Assuming that the sample genome contains 5000 coding sequences, the probability that at least one of the coding sequences has a collision in its MD5 hash value is approximately 7.35×10 −32 . This probability is negligible. Therefore, if the same MD5 value is detected in two lists, it can be considered that the corresponding coding sequences are completely consistent. This is equivalent to detecting the genotype of gene loci in the cgMLST and wgMLST methods. Moreover, even if an md5 collision occurs, it will only make the distance between samples closer than it actually is, and will not cause the clusters that should have been identified to be mistakenly not identified.

[0058] The embodiment of the present application can realize low storage space requirements, and the file size of each bacterial sample in the hash value table after MD5 mapping is about hundreds of KB, which is about 1 / 20 of the genome assembly file of the same sample. Therefore, the storage pressure required for the storage and data exchange of the database of the present invention is greatly reduced.

[0059] Comparison of storage space usage for 100 Salmonella genomes shows that the total size of the original genome assembly sequence file is 483,188 kb. The hash list database file size of the embodiment of the present application is 19,692 kb. The storage space is only 4.1% of the original file size.

[0060] The embodiment of the present application can realize low computing time requirements, eliminate the comparison of gene sites, greatly shorten the computing time, and provide a database merging function that also supports merging the constructed distance matrix at the same time. This can save the process of repeatedly calculating the distance matrix of the merged data of both parties, and only need to fill in the parts that are compared with each other, so the computing time can be further reduced.

[0061] The operation time comparison shows that, based on the single-threaded real operation time calculation, for 100 Salmonella genomes, starting from the genome assembly file, the time required for the predicted coding sequence using the embodiment of the present application is 22 minutes and 4.032 seconds, and a total of 53.125 seconds is used from building the coding sequence database to generating the minimum generation, a total of about 23 minutes. As a control, the complete process of calculating wgMLST using the core genome clustering analysis process ChewBBACA+GrapeTree takes 116 minutes and 33.175 seconds, and calculating cgMLST takes 120 minutes and 21.333 seconds.

[0062] The calculation results of the hash-CDST method in the embodiment of the present application are more consistent with those of the existing genome clustering comparison method and can replace the existing method.

[0063] In order to verify the consistency of the method of the embodiment of the present application and the distance matrix constructed by the cgMLST, wgMLST, and cgSNP methods, a relative distance matrix was constructed for 1961 complete genome sequences of Salmonella in the NCBI database using the embodiment of the present application, and the result was compared with cgMLST, wgMLST, and cgSNP, respectively. A scatter plot was constructed by excluding the pairwise distances between samples from which the comparison itself was made, and their linear regression fitting R-square, Pearson coefficient, and Spearman coefficient were calculated respectively, and the comparison between cgMLST and wgMLST was used as a reference. Among them, wgMLST uses chewBBACA (v3.3.9) to construct Allele, and GrapeTree (v1.5.0) is used to calculate the pairwise distance matrix between samples for the Allele file, and the distance matrix is ​​normalized to the [0,1] interval according to the maximum and minimum method. cgMLST uses chewBBACA (v3.3.9) to screen 95% coverage loci for wgMLST Allele, and uses GrapeTree (v1.5.0) to calculate the pairwise distance matrix between samples. The cgSNP process uses Prokka (v1.14.6) for genome annotation, then uses Roary (v3.13.0, BLASTP similarity threshold = 0.95) to calculate the core genome and construct sequence alignment files, and finally uses snp-dists (v0.8.2) to construct the distance between samples, log-transform the distance (with e as the base), and standardize it to the [0,1] interval according to the maximum and minimum method.

[0064] The consistency verification results show that the sample distances of the present embodiment (hash-CDST) and the typing clustering methods such as cgMLST, wgMLST, cgSNP, and SourMASH are all consistent. Figure 3 The present embodiment has a strong linear correlation with cgMLST and wgMLST, and the R 2 The Pearson coefficient of the method described in this patent is the highest with cgMLST, reaching 0.997, and with wgMLST it is 0.989. Although the linear fit is lower with cgSNP and SourMASH, R 2 The Pearson coefficient and the Spearman rank correlation coefficient were 0.940 and 0.831 respectively, but the Spearman rank correlation coefficient reached 0.967 and 0.954 respectively, which proves that although the embodiment of the present application (hash-CDST) is not linearly correlated with the two, the rank consistency is high. As a reference, the cgMLST and wgMLST constructed by chewBBACA software are compared, and the Pearson coefficient is 0.988, the Spearman coefficient is 0.742, and the R 2 It is 0.977.

[0065] Figure 3 The scatter plots of the distances of paired samples of the present application embodiment (hash-CDST), cgMLST, wgMLST, cgSNP, and SourMASH for 1961 Salmonella genomes are shown. Each point represents the comparison of the distances of the tested samples obtained by the two methods. The linear regression trend line is marked with a red solid line, and the x=y reference line is marked with a black dotted line. The Pearson coefficient, Spearman coefficient, and R-squared are marked in the middle. A, hash-CDST vs. cgMLST, B, hash-CDST vs. wgMLST, C, hash-CDST vs. cgSNP, and D, hash-CDST vs. SourMASH.

[0066] The typing and clustering method of the whole genome sequence data provided in the embodiment of the present application is obtained by obtaining the genome sequence data; performing a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generating a hash list database; processing the hash list in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists; determining the minimum spanning tree network diagram corresponding to the genome sequence data according to the relative distance matrix, and the minimum spanning tree network diagram is used to compare and cluster the genome sequence data. In the embodiment of the present application, the terminal device directly predicts the whole genome coding sequence for the sample genome assembly sequence (fasta format), and performs a hash calculation on each coding sequence, and does not need to compare the gene position information of each coding sequence and directly forms a hash list as a mapping of the sample genome. The relative distance matrix between samples is calculated by comparing the intersection of the hash value lists between different samples, and the minimum spanning tree network diagram corresponding to the genome sequence data is determined according to the relative distance matrix. The minimum spanning tree network diagram is used to compare and cluster the genome sequence data, so as to realize data analysis at the user's local terminal without the support of the central server, and realize decentralization.

[0067] Another embodiment of the present application provides a typing and clustering device for whole genome sequence data, which is used to execute the typing and clustering method for whole genome sequence data provided in the above embodiment.

[0068] Reference Figure 4 , shows a structural block diagram of an embodiment of a typing and clustering device for whole genome sequence data of the present application, the device may specifically include the following modules: an acquisition module 401, a calculation module 402, a processing module 403 and a generation module 404, wherein: The acquisition module 401 is used to acquire genome sequence data; The operation module 402 is used to perform a hash operation on the genome sequence data, obtain a hash list corresponding to the genome sequence data, and generate a hash list database; The processing module 403 is used to process the hash lists in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists; The generation module 404 is used to determine the minimum spanning tree network diagram corresponding to the genome sequence data according to the relative distance matrix, and the minimum spanning tree network diagram is used to align and cluster the genome sequence data.

[0069] The typing and clustering device of the whole genome sequence data provided in the embodiment of the present application obtains the genome sequence data; performs a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generates a hash list database; processes the hash list in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists; determines the minimum spanning tree network diagram corresponding to the genome sequence data according to the relative distance matrix, and the minimum spanning tree network diagram is used to compare and cluster the genome sequence data. In the embodiment of the present application, the terminal device directly predicts the whole genome coding sequence of the sample genome assembly sequence (fasta format), and performs a hash calculation on each coding sequence, and does not need to compare the gene position information of each coding sequence and directly forms a hash list as a mapping of the sample genome. The relative distance matrix between samples is calculated by comparing the intersection of the hash value lists between different samples, and the minimum spanning tree network diagram corresponding to the genome sequence data is determined according to the relative distance matrix. The minimum spanning tree network diagram is used to compare and cluster the genome sequence data, so as to realize data analysis at the user's local terminal without the support of the central server, and realize decentralization.

[0070] Another embodiment of the present application further supplements the description of the typing and clustering device for whole genome sequence data provided in the above embodiment.

[0071] Optionally, the computing module is used to: Preprocessing the genome sequence data to obtain a processed predicted coding sequence; Performing hash mapping on each coding sequence in the predicted coding sequence to obtain multiple hash lists corresponding to the genome sequence data; A plurality of hash lists corresponding to genome sequences are stored in a hash list database, wherein the hash list database includes file names and hash lists corresponding to the file names.

[0072] Optionally, the processing module is used to: Calculating the intersection length of the hash lists in the hash list database to obtain the number of matching coding sequences; According to the number of matched coding sequences, a matching number matrix is ​​constructed; The length of the hash list of the genome sequence data in the hash list database is used as the total number of coding sequences of the genome; Determine the absolute distance between genomic sequence data based on the total number of coding sequences in the genome and the number of matching coding sequences; Determine the relative distance of the genomic sequence data based on the absolute distance between the genomic sequence data and the total coding sequence amount of the genome; According to the relative distances of the genome sequences, a relative distance matrix is ​​determined.

[0073] Optionally, generate a module for: Take the relatively small value of the symmetric position of the relative distance matrix to obtain a list of genome pair distances; Generate a minimum spanning tree based on the genome pairwise distance list; Based on the minimum spanning tree, determine the minimum spanning tree network graph.

[0074] Optionally, generate a module for: Merging multiple hash list databases to obtain a merged hash list database; and / or Multiple relative distance matrices are merged to obtain a merged distance matrix list.

[0075] Optionally, generate a module for: Obtain new genome sequence data; The minimum spanning tree network diagram is updated according to the new genome sequence data.

[0076] Optionally, the hash list database includes at least one of JSON data format, arbitrary character-delimited text format, table format or SQL database format.

[0077] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0078] The typing and clustering device of the whole genome sequence data provided in the embodiment of the present application obtains the genome sequence data; performs a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generates a hash list database; processes the hash list in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists; determines the minimum spanning tree network diagram corresponding to the genome sequence data according to the relative distance matrix, and the minimum spanning tree network diagram is used to compare and cluster the genome sequence data. In the embodiment of the present application, the terminal device directly predicts the whole genome coding sequence of the sample genome assembly sequence (fasta format), and performs a hash calculation on each coding sequence, and does not need to compare the gene position information of each coding sequence and directly forms a hash list as a mapping of the sample genome. The relative distance matrix between samples is calculated by comparing the intersection of the hash value lists between different samples, and the minimum spanning tree network diagram corresponding to the genome sequence data is determined according to the relative distance matrix. The minimum spanning tree network diagram is used to compare and cluster the genome sequence data, so as to realize data analysis at the user's local terminal without the support of the central server, and realize decentralization.

[0079] Yet another embodiment of the present application provides a terminal device for executing the typing and clustering method for whole genome sequence data provided in the above embodiment.

[0080] Figure 5 It is a schematic diagram of the structure of a terminal device of the present application, such as Figure 5 As shown, the terminal device includes: at least one processor 501 and a memory 502; The memory stores a computer program; at least one processor executes the computer program stored in the memory to implement the typing and clustering method for whole genome sequence data provided in the above embodiment.

[0081] The terminal device provided in this embodiment obtains genome sequence data; performs hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generates a hash list database; processes the hash list in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists; determines the minimum spanning tree network diagram corresponding to the genome sequence data according to the relative distance matrix, and the minimum spanning tree network diagram is used to compare and cluster the genome sequence data. In the embodiment of the present application, the terminal device directly predicts the whole genome coding sequence of the sample genome assembly sequence (fasta format), and performs hash calculation on each coding sequence, and does not need to compare the gene position information of each coding sequence, but directly forms a hash list as a mapping of the sample genome. The relative distance matrix between samples is calculated by comparing the intersection of the hash value lists between different samples, and the minimum spanning tree network diagram corresponding to the genome sequence data is determined according to the relative distance matrix. The minimum spanning tree network diagram is used to compare and cluster the genome sequence data, so as to realize data analysis at the user's local terminal without the support of the central server, thereby realizing decentralization.

[0082] Yet another embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed, the typing and clustering method for whole genome sequence data provided in any of the above embodiments is implemented.

[0083] According to the computer-readable storage medium of the present embodiment, by obtaining genome sequence data; performing hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generating a hash list database; according to the hash list database, processing the hash list in the hash list data to obtain a relative distance matrix between the hash lists; according to the relative distance matrix, determining the minimum spanning tree network diagram corresponding to the genome sequence data, the minimum spanning tree network diagram is used to compare and cluster the genome sequence data, in the embodiment of the present application, the terminal device directly predicts the whole genome coding sequence for the sample genome assembly sequence (fasta format), and performs hash mapping calculation on each coding sequence, without comparing the gene position information of each coding sequence and directly forming a hash list as the mapping of the sample genome. By comparing the intersection of the hash value lists between different samples to calculate the relative distance matrix between samples, according to the relative distance matrix, determining the minimum spanning tree network diagram corresponding to the genome sequence data, the minimum spanning tree network diagram is used to compare and cluster the genome sequence data, and realizing data analysis at the user's local terminal without the support of the central server, and realizing decentralization.

[0084] It should be noted that the above detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present application belongs.

[0085] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0086] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.

[0087] In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0088] For ease of description, spatially relative terms, such as "above", "above", "on the upper surface of", "above", etc., may be used herein to describe the spatial positional relationship between a device or feature and other devices or features as shown in the figure. It should be understood that spatially relative terms are intended to include different orientations of the device in use or operation in addition to the orientation described in the figure. For example, if the device in the accompanying drawings is inverted, the device described as "above other devices or structures" or "above other devices or structures" will be positioned as "below other devices or structures" or "below other devices or structures". Thus, the exemplary term "above" may include both "above" and "below". The device may also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and the spatially relative descriptions used herein are interpreted accordingly.

[0089] In the above detailed description, reference is made to the accompanying drawings, which form a part of this document. In the accompanying drawings, similar symbols typically identify similar components unless the context indicates otherwise. The illustrated embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be used, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein.

[0090] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A typing and clustering method for whole genome sequence data, characterized in that: The method comprises: Obtain genome sequence data; Performing a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generating a hash list database; According to the hash list database, the hash lists in the hash list data are processed to obtain a relative distance matrix between the hash lists; According to the relative distance matrix, a minimum spanning tree network diagram corresponding to the genome sequence data is determined, and the minimum spanning tree network diagram is used to align and cluster the genome sequence data.

2. The typing and clustering method for whole genome sequence data according to claim 1, characterized in that: The step of performing a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data and generating a hash list database comprises: Preprocessing the genome sequence data to obtain a processed predicted coding sequence; Performing hash mapping on each coding sequence in the predicted coding sequence to obtain multiple hash lists corresponding to the genome sequence data; A plurality of hash lists corresponding to the genome sequences are stored in a hash list database, wherein the hash list database includes file names and hash lists corresponding to the file names.

3. The typing and clustering method for whole genome sequence data according to claim 1, characterized in that: The step of processing the hash lists in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists includes: Calculating the intersection length of the hash lists in the hash list database to obtain the number of matching coding sequences; According to the number of the matched coding sequences, a matching number matrix is ​​constructed; The length of the hash list of the genome sequence data in the hash list database is used as the total number of coding sequences of the genome; Determining the absolute distance between the genome sequence data based on the total number of coding sequences in the genome and the number of matching coding sequences; Determining the relative distance of the genome sequence data according to the absolute distance between the genome sequence data and the total coding sequence quantity of the genome; According to the relative distances of the genome sequences, a relative distance matrix is ​​determined.

4. The typing and clustering method for whole genome sequence data according to claim 1, characterized in that: Determining a minimum spanning tree network graph corresponding to the genome sequence data according to the relative distance matrix includes: Taking a relatively small value of the symmetric position of the relative distance matrix to obtain a genome pair distance list; Generate a minimum spanning tree according to the genome pairwise distance list; According to the minimum spanning tree, a minimum spanning tree network graph is determined.

5. The typing and clustering method for whole genome sequence data according to claim 3, characterized in that: The method further comprises: Merging the plurality of hash list databases to obtain a merged hash list database; and / or A plurality of the relative distance matrices are merged to obtain a merged distance matrix list.

6. The typing and clustering method for whole genome sequence data according to claim 1, characterized in that: The method further comprises: Obtain new genome sequence data; The minimum spanning tree network diagram is updated according to the new genome sequence data.

7. The typing and clustering method for whole genome sequence data according to claim 1, characterized in that: The hash list database includes at least one of JSON data format, text format separated by any characters, table format or SQL database format.

8. A typing and clustering device for whole genome sequence data, characterized in that: The device comprises: An acquisition module, used to acquire genome sequence data; A calculation module, used for performing a hash operation on the genome sequence data to obtain a hash list corresponding to the genome sequence data, and generating a hash list database; A processing module, used for processing the hash lists in the hash list data according to the hash list database to obtain a relative distance matrix between the hash lists; A generation module is used to determine a minimum spanning tree network diagram corresponding to the genome sequence data based on the relative distance matrix, and the minimum spanning tree network diagram is used to align and cluster the genome sequence data.

9. A terminal device, characterized in that: include: at least one processor and memory; The memory stores a computer program; The at least one processor executes the computer program stored in the memory to implement the typing and clustering method for whole genome sequence data according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed, implements the typing and clustering method for whole genome sequence data according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Pathogen analysis method and device based on high-throughput sequencing and computer equipment

    CN112259167A

  • Method and system for generating summary data of biological gene sequence

    CN113496762A

  • Large-scale biological data clustering method and system based on spanning tree

    CN114420215A

  • DNA sequence clustering method and system based on locality sensitive hash function, electronic equipment and readable storage medium

    CN118629513A

  • An improved method to identify nucleic acid sequences within a set of sequences obtained by a sequencer and a system

    US20250069697A1