Assembly of overlap group detection methods, apparatus and equipment

CN122575491APending Publication Date: 2026-08-14SHENZHEN HUADA GENE INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

目前,尽管现有的组装软件已经在处理短读长数据方面取得了一定的进展,但由于短读长数据本身缺乏短读段之间的长片段关联信息,导致现有组装技术能得到的组装结果仍然不尽人意

Benefits of technology

[0016]上述组装重叠群检测方法、装置、计算机设备、计算机可读存储介质和计算机程序产品,获取宏基因组组装重叠群和双端测序短读长数据,并将宏基因组组装重叠群与双端测序短读长数据进行比对以得到比对文件;在比对文件中确定目标重叠群片段,并确定目标重叠群片段的比对特征;比对特征至少包括比对断点特征、双端短读长比对非同一重叠群特征或反向比对特征;生成比对特征的二维图像,并将比对特征的二维图像输入目标分类模型中,得到目标重叠群片段的检测结果。本申请提供的组装重叠群检测方法,能够通过生成目标重叠群片段的比对特征的二维图像,并通过将比对特征的二维图像输入目标分类模型中,使得目标分类模型能够通过直观的图像特征更好地在分类过程中学习所需特征信息,从而地,得到的目标重叠群片段的检测结果中包括了比对特征的贡献,进而地,能够通过目标重叠群片段的检测结果的准确度以确保宏基因组组装重叠群的检测结果的准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575491A_ABST
    Figure CN122575491A_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, and device for detecting assembly contigs. The method includes: acquiring metagenomic assembly contigs and paired-end sequencing short-read data, and comparing the metagenomic assembly contigs with the paired-end sequencing short-read data to obtain an alignment file; identifying a target contig fragment in the alignment file, and determining the alignment features of the target contig fragment; the alignment features include at least alignment breakpoint features, paired-end short-read alignment features indicating non-identical contigs, or reverse alignment features; generating a two-dimensional image of the alignment features, and inputting the two-dimensional image of the alignment features into a target classification model to obtain the detection result of the target contig fragment. This method can ensure the accuracy of the metagenomic assembly contig detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of high-throughput sequencing technology, and in particular to an assembly contig detection method, apparatus, and device. Background Technology

[0002] With the development of high-throughput sequencing technology, metagenomic assembly technology has emerged. Metagenomic assembly plays an important role in unlocking the secrets of microorganisms. Among them, the key to constructing a reliable metagenomic assembly is how to obtain accurate metagenomic assembly contigs.

[0003] Current metagenomic assembly techniques primarily rely on short-read sequencing platforms to generate large amounts of paired-end short-read data. This data is then assembled using various assembly software to construct contigs. While existing assembly software has made some progress in processing short-read data, the lack of long-segment association information between short reads in the short-read data itself results in unsatisfactory assembly results. Clearly, failing to ensure the accuracy of the metagenomic contig assembly results will hinder the construction of reliable metagenomic genomes. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for detecting assembly contigs that can ensure the accuracy of the detection results of metagenomic assembly contigs, addressing the aforementioned technical problems.

[0005] Firstly, this application provides a method for detecting assembly overlap groups, including:

[0006] Obtain metagenomic assembly contigs and paired-end sequencing short read data, and align the metagenomic assembly contigs with the paired-end sequencing short read data to obtain an alignment file;

[0007] Identify the target contiguous group segment in the comparison file and determine the comparison features of the target contiguous group segment; the comparison features include at least the comparison breakpoint features, the feature of comparing different contiguous groups with short read lengths at both ends, or the reverse comparison features;

[0008] A two-dimensional image of the alignment features is generated and input into the target classification model to obtain the detection results of the target overlapping group segments.

[0009] Secondly, this application also provides an assembly of an overlap group detection device, comprising:

[0010] The alignment module is used to acquire metagenomic assembly contigs and paired-end sequencing short read data, and to align the metagenomic assembly contigs with the paired-end sequencing short read data to obtain an alignment file;

[0011] The first determining module is used to determine the target contiguous group segment in the comparison file and determine the comparison features of the target contiguous group segment; the comparison features include at least the comparison breakpoint features, the feature of comparing non-same contiguous group at both ends short read lengths, or the reverse comparison features;

[0012] The input module is used to generate a two-dimensional image of the alignment features and input the two-dimensional image of the alignment features into the target classification model to obtain the detection results of the target overlapping group segments.

[0013] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement some or all of the steps described in any method of the first aspect of the embodiments of this application.

[0014] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements some or all of the steps described in any method of the first aspect of the embodiments of this application.

[0015] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements some or all of the steps described in any method of the first aspect of the embodiments of this application.

[0016] The aforementioned assembly contig detection method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire metagenomic assembly contigs and paired-end sequencing short-read data, and align the metagenomic assembly contigs with the paired-end sequencing short-read data to obtain an alignment file; identify the target contig fragment in the alignment file, and determine the alignment features of the target contig fragment; the alignment features include at least alignment breakpoint features, paired-end short-read alignment features of non-identical contigs, or reverse alignment features; generate a two-dimensional image of the alignment features, and input the two-dimensional image of the alignment features into a target classification model to obtain the detection result of the target contig fragment. The assembly contig detection method provided in this application can generate a two-dimensional image of the alignment features of the target contig fragment, and by inputting the two-dimensional image of the alignment features into the target classification model, the target classification model can better learn the required feature information during the classification process through intuitive image features. Therefore, the detection result of the target contig fragment includes the contribution of the alignment features, and thus, the accuracy of the metagenomic assembly contig detection result can be ensured by the accuracy of the target contig fragment detection result. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a diagram illustrating the application environment of the assembly overlap group detection method in one embodiment;

[0019] Figure 2 This is a flowchart illustrating the assembly of an overlap group detection method in one embodiment;

[0020] Figure 3A This is a schematic diagram of a diagonal image comparing breakpoint features in one embodiment;

[0021] Figure 3B This is a schematic diagram of a diagonal image comparing features of different overlapping groups with short read lengths at both ends in one embodiment;

[0022] Figure 3C This is a schematic diagram of a diagonal image comparing depth features in one embodiment;

[0023] Figure 3D This is a schematic diagram of the diagonal image of the average length feature of the paired short read length comparison position in one embodiment;

[0024] Figure 3E This is a schematic diagram of a matrix image comparing depth difference features in one embodiment;

[0025] Figure 4A This is a schematic diagram of the precision-recall curve of a target classification model in one embodiment;

[0026] Figure 4B This is a schematic diagram of the receiver operation feature curve of a target classification model in one embodiment;

[0027] Figure 5A This is a schematic diagram comparing the length of contiguous groups in a misassembled metagenomic assembly before and after correction in one embodiment.

[0028] Figure 5B This is a schematic diagram comparing the number of contiguous groups in a misassembled metagenomic assembly before and after correction in one embodiment.

[0029] Figure 5C This is a schematic diagram comparing the number of sequence sites with assembly errors in a metagenomic assembly contig before and after correction in one embodiment.

[0030] Figure 5DThis is a schematic diagram comparing the number of metagenomic assembly contigs with local assembly errors before and after correction in one embodiment.

[0031] Figure 5E This is a schematic diagram comparing the number of metagenomic assembly contigs before and after correction in one embodiment where the assembly error is a shift error;

[0032] Figure 5F This is a schematic diagram comparing the number of metagenomic assembly contigs before and after correction in one embodiment where the assembly error is inverted assembly.

[0033] Figure 5G This is a schematic diagram comparing the number of metagenomic assembly contigs before and after correction in one embodiment where the assembly error is an interspecies translocation error;

[0034] Figure 5H This is a schematic diagram comparing the number of metagenomic assembly contigs before and after correction in one embodiment where the assembly error is an intra-species translocation error;

[0035] Figure 6 This is a structural block diagram of an assembly of an overlap group detection device in one embodiment;

[0036] Figure 7 This is an internal structural diagram of a computer device in one embodiment;

[0037] Figure 8 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0039] The assembly overlap group detection method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0040] In one exemplary embodiment, such as Figure 2 As shown, an assembly contiguous group detection method is provided, which is applied to... Figure 1 Taking the terminal in the example, the explanation includes the following steps 202 to 206. Wherein:

[0041] Step 202: Obtain metagenomic assembly contigs and paired-end sequencing short read data, and align the metagenomic assembly contigs with the paired-end sequencing short read data to obtain an alignment file.

[0042] Metagenome assembly contigs are intermediate products generated during metagenomic assembly. They are continuous DNA sequences obtained by high-throughput sequencing of DNA from mixed microbial communities from complex environmental samples (such as soil, water, and human gut), followed by sequence alignment and splicing using bioinformatics software.

[0043] Optionally, the obtained metagenomic assembly contigs can be real data or synthetic simulated data. Optionally, the metagenomic assembly contigs can be obtained by the terminal from a database or obtained locally by the terminal using a simulated data generation tool.

[0044] Paired-end sequencing short read data refers to high-throughput DNA sequence information generated using paired-end sequencing technology, in which both ends of each DNA fragment are sequenced, thereby producing paired short read sequences.

[0045] Optionally, the obtained paired-end sequencing short read data can be obtained from a high-throughput sequencing platform; for example, the high-throughput sequencing platform can be Illumina NovaSeq, HiSeq, etc.

[0046] An alignment file is a detailed data file that contains the alignment results between metagenomic assembly contigs and paired-end sequencing short read data.

[0047] In an exemplary embodiment, the above-mentioned acquisition of metagenomic assembly contigs and paired-end sequencing short read data, and the comparison of metagenomic assembly contigs with paired-end sequencing short read data to obtain an alignment file, includes: acquiring metagenomic assembly contigs and paired-end sequencing short read data, and comparing the metagenomic assembly contigs with paired-end sequencing short read data based on the BWA (Burrows-Wheeler Aligner) tool to obtain an alignment file.

[0048] Step 204: Determine the target contiguous group segment in the comparison file and determine the comparison features of the target contiguous group segment; the comparison features include at least the comparison breakpoint features, the feature of comparing short read lengths at both ends to the same contiguous group, or the reverse comparison features.

[0049] The target contig fragment is part or all of the alignment file; it is easy to understand that since the alignment file is obtained by aligning metagenomic assembly contigs and paired-end sequencing short read data, the target contig fragment can actually be understood as part or all of the metagenomic assembly contigs.

[0050] Alignment features refer to the feature information extracted from the target contig fragment, including multiple sequence sites in the target contig fragment. They can be used to evaluate the alignment quality of the target contig fragment, identify assembly errors in the target contig fragment, or perform other bioinformatics analyses.

[0051] Specifically, the alignment features include at least one of the following: alignment breakpoint features, short-read-length pairwise alignment features of non-consolidated groups, or reverse alignment features. That is to say, the alignment features include at least one of the following: alignment breakpoint features, short-read-length pairwise alignment features of non-consolidated groups, and reverse alignment features.

[0052] Alignment breakpoints are sequence sites where short reads cannot be continuously aligned with contiguous groups or reference sequences during alignment. The presence of alignment breakpoints can be used to characterize break points in the target contiguous group fragment or erroneous connections during assembly. Alignment breakpoints refer to break points where short reads cannot be continuously aligned with the target contiguous group fragment.

[0053] The feature of paired short read alignment to non-consolidated groups can be used to characterize erroneous connections between contigs or the presence of repetitive sequences in the genome. The number of short reads at each sequence site that are not aligned to a contig, refers to the number of short reads in a specific region of the target contig fragment where one end of the paired short read is aligned to one contig, while the other end is aligned to a different contig.

[0054] Reverse alignment features can be used to characterize inverted repetitive sequences or assembly errors in the genome. Each sequence site is defined by the number of reverse alignments at one end within a paired short read.

[0055] In an exemplary embodiment, the above-described determination of the target contiguous group segment in the comparison file and determination of the comparison features of the target contiguous group segment includes: determining a target length window; determining the target contiguous group segment in the comparison file based on the target length window and determining the comparison features of the target contiguous group segment.

[0056] The length of the target contiguous group fragment is the same as the length of the target length window. For example, when the length of the target length window is 384 bp, the length of the target contiguous group fragment is also 384 bp.

[0057] It is easily understood that the short read coverage of certain regions in the alignment file may be lower than that of other regions. This uneven coverage is more common at the ends of contigs. Furthermore, there may be repetitive sequence regions at the ends of the alignment file, which can easily lead to ambiguous results during alignment, resulting in lower alignment accuracy. Therefore, it is obvious that the alignment results at the beginning and end of the alignment file are usually poor. Thus, in an exemplary embodiment, the above-described determination of the target length window; determination of the target contig fragment in the alignment file based on the target length window; and determination of the alignment characteristics of the target contig fragment include: determining the target length window; determining the target contig fragment in the target region of the alignment file based on the target length window; and comparing the target contig fragment with paired-end sequencing short read data to determine the alignment characteristics of the target contig fragment.

[0058] The target region refers to the remaining portion of the alignment file excluding the regions at both ends. Optionally, the length of both ends can be 50 bp, meaning that the two ends can be a 50 bp portion at the beginning and a 50 bp portion at the end. For example, if the length of the alignment file is 10000 bp and the length of both ends is 50 bp, then the target region is from the 51st to the 9950th base pair of the alignment file, meaning that the target contiguous group fragment is a fragment among the 51st to the 9950th base pairs.

[0059] Step 206: Generate a two-dimensional image of the alignment features and input the two-dimensional image of the alignment features into the target classification model to obtain the detection results of the target overlapping group segments.

[0060] The two-dimensional image includes multiple pixels, which represent the numerical values ​​of sequence sites in the target contiguous group fragment in terms of alignment features. It is easy to understand that the longer the target contiguous group fragment, the more sequence sites there are, and thus, the more pixels representing substantial features in the two-dimensional image. Consequently, the information content of the two-dimensional image is richer; that is, there is a positive correlation between the length of the target contiguous group fragment and the information content of the two-dimensional image.

[0061] Optionally, the arrangement of pixels in the two-dimensional image for feature comparison can be along the image diagonal or along all positions in the image. When the pixel arrangement is along the image diagonal, the pixel arrangement direction is from the upper left corner to the lower right corner along the image diagonal.

[0062] A target classification model is an image classification model that can classify two-dimensional images. Optionally, the target classification model can be a vision transformer (ViT) model, a convolutional neural network model, a deep residual network model, or other models that can be used to classify images.

[0063] The test result is either "assembled correctly" or "assembled incorrectly." "Assembled correctly" means that the obtained target contiguous fragment is accurate, continuous, and complete. That is, during genome assembly, the DNA sequences corresponding to the target contiguous fragment are correctly linked and arranged to form a continuous sequence that accurately reflects the original genome structure. "Assembly incorrectly" means that there is a deviation between the obtained target contiguous fragment and the original DNA sequence. Optionally, assembly errors can include incorrect connections, deletions, duplications, and sequence errors.

[0064] In an exemplary embodiment, the target classification model is a ViT model; the ViT model includes an image embedding layer, an encoding layer, and a classification layer; the image embedding layer is used to convert multiple image patches in a two-dimensional image into multiple numerical vectors; the encoding layer is responsible for learning the feature relationships between the numerical vectors corresponding to each image patch and outputting the image classification result of the two-dimensional image; the classification layer is used to output the detection result of the target overlapping group segment based on the image classification result output by the encoding layer.

[0065] In an exemplary embodiment, the method further includes: determining an initial ViT model; determining the cross-entropy loss function value of the initial ViT model during training; determining the gradient value of the cross-entropy loss function value based on a gradient descent algorithm; iteratively training the model parameters of the initial ViT model based on the gradient value until the newly obtained cross-entropy loss function value satisfies a preset loss condition; then determining the model including the model parameters corresponding to the cross-entropy loss function value satisfying the preset loss condition as the ViT model. It is easy to understand that the ViT model determined at this point is the target classification model.

[0066] Furthermore, when the detection results of the target contiguous group fragments are obtained and the detection results show that the assembly is correct, it can be used to analyze the microbial community. Specifically, it can be used to perform species abundance analysis, functional gene analysis, evolutionary relationship analysis, environmental adaptability analysis, or other information analysis on the microbial community.

[0067] In an exemplary embodiment, after obtaining the detection result of the target contiguous fragment, the method further includes: sliding a target length window in the alignment file to determine the next target contiguous fragment, and aligning the next target contiguous fragment with paired-end sequencing short read data to determine the alignment features of the next target contiguous fragment; the above-mentioned generation of the two-dimensional image of the alignment features includes: generating the two-dimensional image of the alignment features of the next target contiguous fragment, and inputting the two-dimensional image of the alignment features of the next target contiguous fragment into the target classification model to obtain the detection result of the next target contiguous fragment.

[0068] In another exemplary embodiment, the above-described method of sliding a target length window in the alignment file to determine the next target contiguous fragment and aligning the next target contiguous fragment with paired-end sequencing short-read data to determine the alignment characteristics of the next target contiguous fragment includes: sliding a target length window in the alignment file by one step to determine the next target contiguous fragment and aligning the next target contiguous fragment with paired-end sequencing short-read data to determine the alignment characteristics of the next target contiguous fragment.

[0069] Specifically, the alignment file is slid by one step based on the target length window. In other words, the step size is equal to the length of the target length window. For example, assuming the target length window is 384 bp, if the previous target contiguous fragment is from the 51st to the 435th (51+384)th base pair of the metagenomic assembly contiguous fragment, then after sliding the alignment file by one step based on the target length window, the next target contiguous fragment will be from the 436th (51+384+1)th to the 820th (51+384+1+384)th base pair of the metagenomic assembly contiguous fragment.

[0070] In this embodiment, by sliding a target length window across the metagenomic assembly contiguous group with a step size equal to the target length window, the next target contiguous group fragment is continuously determined. Finally, the detection result of each target contiguous group fragment in the metagenomic assembly contiguous group is obtained. Thus, when the detection result of each target contiguous group fragment is correctly assembled, it can be ensured that the detection result of the final metagenomic assembly contiguous group is also correctly assembled. That is to say, this embodiment can ensure the accuracy and reliability of the detection result of the final metagenomic assembly contiguous group by using the detection results of multiple target contiguous group fragments.

[0071] In the aforementioned assembly contiguous group detection method, metagenomic assembly contiguous group and paired-end sequencing short-read data are acquired, and the metagenomic assembly contiguous group and paired-end sequencing short-read data are compared to obtain an alignment file. Target contiguous group fragments are identified in the alignment file, and their alignment features are determined. These alignment features include at least alignment breakpoint features, paired-end short-read alignment features indicating non-identical contiguous groups, or reverse alignment features. A two-dimensional image of the alignment features is generated and input into a target classification model to obtain the detection result of the target contiguous group fragment. The assembly contiguous group detection method provided in this application can generate a two-dimensional image of the alignment features of the target contiguous group fragment and input this image into a target classification model. This allows the target classification model to better learn the required feature information during the classification process through intuitive image features. Therefore, the detection result of the target contiguous group fragment includes the contribution of the alignment features, and the accuracy of the target contiguous group fragment detection result can be used to ensure the accuracy of the metagenomic assembly contiguous group detection result.

[0072] In an exemplary embodiment, the above-mentioned alignment features further include at least alignment depth features, alignment depth difference features, or average length features of the alignment positions of the two short reads.

[0073] Among them, the alignment depth feature can reflect the sequencing coverage of the target contig fragment in a specific region. The alignment depth indicates the number of times a specific region has been sequenced, and the number of times it has been sequenced is positively correlated with the amount of information.

[0074] By comparing depth difference features, specific regions with uneven coverage within target contig fragments can be identified to indicate structural variations, repetitive regions, or assembly errors in the genome.

[0075] The average length feature of paired short read alignment positions refers to the average distance between the alignment positions of the ends of each pair of short reads within multiple paired short reads aligned to a sequence site. This average length feature of paired short read alignment positions can be used to evaluate the pairing quality of paired short reads.

[0076] Specifically, the alignment features also include at least the alignment depth feature, the alignment depth difference feature, or the average length feature of the alignment positions of the paired short reads. That is to say, in addition to including at least one of the alignment breakpoint feature, the paired short read alignment non-overlapping group feature, and the reverse alignment feature, the alignment features also include at least one of the alignment depth feature, the alignment depth difference feature, and the average length feature of the alignment positions of the paired short reads.

[0077] In this embodiment, since the two-dimensional image of the alignment features also includes at least one of the alignment depth features, alignment depth difference features, and the average length of the alignment positions of the paired short reads, after inputting the two-dimensional image of the alignment features containing more feature information into the target classification model, the detection result of the target contiguous fragment includes the contribution of more feature information. Therefore, the accuracy and reliability of the detection result of the target contiguous fragment can be further improved, so as to further ensure the accuracy and reliability of the detection result of metagenomic assembly contiguous groups.

[0078] In one exemplary embodiment, the target contiguous fragment includes multiple sequence sites;

[0079] Determining the alignment features of target contiguous fragments includes at least one of the following steps:

[0080] Determine the number of alignment breakpoints at each sequence site to obtain alignment breakpoint features;

[0081] Determine the number of short reads at each sequence site that are not aligned to a contiguous group to obtain the feature that the short reads at both ends are not aligned to the same contiguous group;

[0082] The number of sequence sites covered by the short reads that are reverse-aligned from one end of the paired short reads is determined to obtain the reverse alignment features.

[0083] In an exemplary embodiment, the determination of alignment features of the target contiguous group fragments includes at least one of the following steps:

[0084] Determine the short read coverage depth of each sequence site to obtain alignment depth features;

[0085] Based on the alignment depth features, the absolute value of the difference in coverage depth between each sequence site and other sequence sites is determined to obtain the alignment depth difference features;

[0086] The average length of the alignment positions of the paired short reads covered by each sequence site is determined to obtain the average length feature of the paired short read alignment positions.

[0087] In this embodiment, by performing detailed analysis on multiple sequence sites included in the target contiguous fragment, the alignment breakpoint features, the features of non-constiguous fragments in paired short read alignments or reverse alignment features, as well as the alignment depth features, alignment depth difference features, or the average length features of paired short read alignment positions are accurately determined. Thus, based on the accurately obtained alignment features, errors in the assembly process can be more accurately identified and corrected through the target classification model, thereby improving the accuracy and reliability of the detection results of the target contiguous fragment.

[0088] In an exemplary embodiment, generating a two-dimensional image of comparison features includes:

[0089] The first image is generated based on the comparison of breakpoint features, the comparison of non-overlapping group features of short read lengths at both ends, or the reverse comparison features.

[0090] A second image is generated based on the alignment depth features, alignment depth difference features, or the average length feature of the alignment positions of the two-end short read lengths.

[0091] The first and second images are stacked to obtain a two-dimensional image.

[0092] The first image contains the feature information necessary for the target classification model to obtain the detection results of the target overlapping group fragments. In other words, the first image is the base image that enables the target classification model to identify and correct errors in the assembly process.

[0093] The feature information included in the second image can further improve the accuracy and reliability of the detection results of target overlapping group segments. In other words, the second image is a supplementary image used to improve the ability of the target classification model to identify and correct errors in the assembly process.

[0094] Optionally, the arrangement of pixels in the first image can be along the diagonal of the image or along all positions of the image; the same applies to the second image, so it will not be repeated here. The arrangement of pixels in the first image and the second image can be different.

[0095] It is easy to understand that since the two-dimensional image in this embodiment is obtained by stacking the first image and the second image, the two-dimensional image in this embodiment is multi-channel.

[0096] In this embodiment, a base image, namely the first image, is generated based on the comparison breakpoint features, the paired short read length comparison non-identical contiguous group features, or the reverse comparison features, enabling the target classification model to identify and correct errors in the assembly process. Simultaneously, a supplementary image, namely the second image, is generated based on the comparison depth features, the comparison depth difference features, or the paired short read length comparison position average length features, which can improve the target classification model's ability to identify and correct errors in the assembly process. Thus, the multi-channel two-dimensional image obtained after stacking the first and second images can determine the detection results of target contiguous group segments with sufficient accuracy and reliability based on a sufficient number of intuitive image features.

[0097] In one exemplary embodiment, the first image described above is a first diagonal image;

[0098] The second image mentioned above includes a second diagonal image generated based on the alignment depth feature or the average length feature of the alignment position of the two-ended short read length, and a matrix image generated based on the alignment depth difference feature.

[0099] In an exemplary embodiment, when the first image is a first diagonal image, the above-mentioned generation of the first image based on the comparison of breakpoint features, the comparison of non-overlapping group features of short read lengths at both ends, or the reverse comparison features includes at least one of the following steps:

[0100] The number of alignment breakpoints for each sequence locus is arranged sequentially along a preset diagonal direction to generate a diagonal image of alignment breakpoint features; the number of short reads for each sequence locus that are not aligned to a contiguous group is arranged sequentially along a preset diagonal direction to generate a diagonal image of short reads aligned to non-constiguous group features; the number of short reads for each sequence locus that are aligned backwards from one end of the short reads is arranged sequentially along a preset diagonal direction to generate a diagonal image of backward alignment features.

[0101] The preset diagonal direction is from the top left corner to the bottom right corner.

[0102] The first diagonal image includes a diagonal image comparing breakpoint features, a diagonal image comparing features of different contiguous groups with short reads at both ends, or a diagonal image comparing features in reverse.

[0103] In an exemplary embodiment, when the second image includes a second diagonal image generated based on alignment depth features or the average length feature of paired short read alignment positions, and a matrix image generated based on alignment depth difference features, generating the second image based on alignment depth features, alignment depth difference features, or the average length feature of paired short read alignment positions includes at least one of the following steps:

[0104] The short read coverage depths of each sequence site are arranged sequentially along a preset diagonal direction to generate a diagonal image of alignment depth features. The absolute values ​​of the coverage depth differences between each sequence site and other sequence sites are arranged sequentially on each pixel to generate a matrix image of alignment depth difference features. The average length of the alignment positions of the short reads covered by each sequence site is arranged sequentially along a preset diagonal direction to generate a diagonal image of the average length of the alignment positions of the short reads.

[0105] In the matrix image comparing depth difference features, each pixel represents the absolute value of the coverage depth difference between each sequence site and the remaining sequence sites.

[0106] For example, let A represent the array of short read coverage depths of each sequence site included in the alignment depth features, and let B represent the matrix image of the alignment depth difference features. Then, each pixel in the matrix image is the absolute value of the difference in coverage depth between the i-th sequence site and the j-th sequence site in A, i.e., B. ij =|A i -A j |. Further exemplarily, the target contiguous fragment includes 5 sequence sites, each with a coverage depth of 3, 5, 2, 5, and 3, respectively. Then A = [3, 5, 2, 5, 3], and B... 11 =|A1–A1|=|3–3|=0, B 12 =|A1–A2|=|3–5|=2……B 54 =|A5–A4|=|3–5|=2, B 55 If |A5–A5| = |3–3| = 0, then the resulting matrix image is represented as:

[0107]

[0108] It is easy to understand that, since the diagonal positions of the matrix image represent the absolute value of the difference in coverage depth between each sequence site and itself, and the difference in coverage depth between each sequence site and itself is 0, the diagonal positions of the matrix image are all the lowest values. Furthermore, the pixels at the diagonal positions of the matrix image have the same gray level.

[0109] Optionally, in the diagonal images of comparison breakpoint features, comparison of features from different short read lengths of non-overlapping groups, comparison of reverse features, comparison of depth features, comparison of depth difference features, and comparison of the average length of positions of comparison of different short read lengths, the numerical value of the comparison feature represented by a pixel can be positively correlated with its grayscale value. That is, the larger the numerical value of the comparison feature represented by a pixel, the larger the grayscale value of that pixel, resulting in higher brightness in the visual effect. For example, in the matrix image B described above, since the values ​​of all pixels along the diagonal position from the upper left to the lower right are 0, all pixels at the diagonal position are the darkest, appearing as a black diagonal line in the visual effect. For example, the grayscale range can be from 0 to 255.

[0110] For example, if the maximum value of the alignment feature represented by a pixel is represented by a grayscale value of 255, and the minimum value is represented by a grayscale value of 0, then the diagonal image of the alignment breakpoint feature is as follows: Figure 3A As shown; diagonal images of features from different contiguous groups compared by comparing the short read lengths at both ends are shown. Figure 3B As shown; the diagonal image comparing depth features is as follows Figure 3C As shown; the diagonal image of the average length feature of the paired short read length comparison position is as follows. Figure 3D As shown; the matrix image comparing depth difference features is as follows Figure 3E As shown.

[0111] In this embodiment, by comparing breakpoint features, paired short read length comparison of non-same contiguous group features, reverse alignment features, alignment depth features, alignment depth difference features, and paired short read length comparison position average length features, diagonal images of alignment breakpoint features, paired short read length comparison of non-same contiguous group features, reverse alignment features, alignment depth features, alignment depth difference features, and paired short read length comparison position average length features are generated respectively. Thus, the stacked multi-channel two-dimensional images can simultaneously consider multiple different feature information of the target contiguous group fragment, and also enable the target classification model to learn and understand the detection results of the target contiguous group fragment from multiple perspectives, thereby ensuring that the detection results of metagenomic assembly contiguous groups have accuracy and reliability.

[0112] In one exemplary embodiment, the method further includes:

[0113] When the detection result of the target overlapping group segment is that the target overlapping group segment is assembled incorrectly, the comparison breakpoint features of the target overlapping group segment are analyzed, and the target breakpoint is determined based on the analysis results.

[0114] The target breakpoint is the site with the most alignment breakpoints or the midpoint of the target contiguous group segment.

[0115] In an exemplary embodiment, the above-mentioned determination of the target breakpoint based on the analysis results includes: determining whether there is an alignment breakpoint in the target contiguous group segment based on the analysis results; if there is an alignment breakpoint, determining the sequence position with the most alignment breakpoints as the target breakpoint; if there is no alignment breakpoint, determining the midpoint of the target contiguous group segment as the target breakpoint.

[0116] In this embodiment, when the detection result of the target contiguous fragment indicates that the target contiguous fragment is assembled incorrectly, the target breakpoint is determined based on the analysis results of the comparison breakpoint features. Thus, by disconnecting the incorrectly assembled target contiguous fragment, the assembly accuracy of metagenomic contiguous fragments is improved.

[0117] In an exemplary embodiment, the comparison quality of the above-mentioned comparison files meets a preset quality condition, and the length meets a preset length condition.

[0118] The preset quality condition can be that the comparison quality is greater than or equal to the preset quality, and the preset quality can be 40 or other values; thus, when the preset quality is 40, the fact that the comparison quality meets the preset quality condition can indicate that the comparison quality of the comparison file is greater than or equal to 40.

[0119] The preset length condition can be that the length of the comparison file is greater than or equal to the preset length, which can be 3000bp or other values. Thus, when the preset length is 3000bp, the length meeting the preset length condition indicates that the length of the comparison file is greater than 3000bp.

[0120] In an exemplary embodiment, the comparison quality of the comparison file meets the preset quality condition and the length meets the preset length condition, including: the comparison quality of the comparison file meets the preset quality condition, the length meets the preset length condition, and the comparison file does not have supplementary comparison, duplicate comparison, or second comparison.

[0121] In this embodiment, the alignment quality of the alignment file meets the preset quality condition, and the length meets the preset length condition. Thus, it can be ensured that the alignment file has high alignment quality and long length. By avoiding low-quality alignment files, the accuracy and reliability of the detection results of the target contig fragments obtained subsequently are ensured. In turn, the accuracy and reliability of the detection results of the final metagenomic assembly contigs are ensured.

[0122] In one exemplary embodiment, the method further includes:

[0123] An initial classification model and multiple microbial communities were constructed, and short read sequences with different sequencing depths were obtained based on the different abundances of each microbial community.

[0124] Short read sequences at different sequencing depths were assembled to obtain multiple training metagenomic assembly contigs;

[0125] Short read sequences at different sequencing depths were aligned with multiple training metagenomic assembly contigs to obtain multiple training alignment files;

[0126] Each training comparison file is evaluated to obtain the training and detection results for each training comparison file;

[0127] Based on multiple training and detection results, the initial classification model is trained to obtain the target classification model.

[0128] The abundance of a microbial community refers to the quantity of a certain type of microorganism in a specific environment. The abundance of multiple microbial communities is sampled according to a log-normal distribution. Optionally, the multiple microbial communities can have four log-normal distributions. For example, the four log-normal distributions can be (mean = 5, standard deviation = 2), (mean = 10, standard deviation = 1), (mean = 10, standard deviation = 2), and (mean = 10, standard deviation = 0.5).

[0129] Optionally, multiple microbial communities may include the complete genomes of 1000 microorganisms. More optionally, the complete genomes of the 1000 microorganisms may be selected from the Genome Taxonomy Database version 207.

[0130] Optionally, after constructing the initial classification model and multiple microbial communities, the MGSIM tool can be used to obtain short read sequences at different sequencing depths based on the different abundances of each microbial community. The MEGAHIT tool can be used to assemble the short read sequences at different sequencing depths to obtain multiple training metagenomic assembly contigs. The BWA tool can be used to align the short read sequences at different sequencing depths with multiple training metagenomic assembly contigs to obtain multiple training alignment files. The QUAST tool can be used to evaluate each training alignment file to obtain the training detection results of each training alignment file.

[0131] In an exemplary embodiment, the above-mentioned alignment of short read sequences at different sequencing depths with multiple training metagenomic assembly contigs to obtain multiple training alignment files includes: aligning short read sequences at different sequencing depths with multiple training metagenomic assembly contigs to obtain multiple initial training alignment files; and cleaning the multiple initial training alignment files to obtain multiple training alignment files.

[0132] In this embodiment, multiple microbial communities are constructed to simulate short read sequences, thereby obtaining training data that can be used to train the target classification model, overcoming the difficulty of training a suitable deep learning model due to the scarcity of real data.

[0133] In an exemplary embodiment, the above-described method of training an initial classification model based on multiple training detection results to obtain a target classification model includes:

[0134] Based on multiple training detection results, multiple training metagenomic assembly contigs are divided into at least one first training contig and at least one second training contig; the training detection result of the first training contig is assembly error, and the training detection result of the second training contig is assembly correct.

[0135] Based on the assembly breakpoints of each first training overlap group, construct a set of positive example two-dimensional images;

[0136] Based on at least one of the following: alignment breakpoint features of each second training overlap group, features of short-read length comparison of non-overlapping groups at both ends, and reverse alignment features, construct a set of negative example two-dimensional images.

[0137] The initial classification model is trained based on the positive and negative two-dimensional image sets to obtain the target classification model.

[0138] The training detection results may include assembly breakpoints of each training metagenomic assembly contiguous group, thereby enabling the construction of a set of positive example two-dimensional images based on the assembly breakpoints of each first training contiguous group in multiple training detection results.

[0139] In an exemplary embodiment, the above-mentioned construction of a set of positive example two-dimensional images based on the assembly breakpoints of each first training overlap group includes: sliding a training length window left and right along the assembly breakpoints of each first training overlap group, and obtaining multiple first training two-dimensional images including the assembly breakpoints of each first training overlap group during the sliding process; and constructing a set of positive example two-dimensional images based on the multiple first training two-dimensional images.

[0140] In order to make the obtained set of positive example two-dimensional images more abundant, the training length window can have different lengths to obtain more first training two-dimensional images that include the assembly breakpoints of each first training overlap group.

[0141] In an exemplary embodiment, the above-described construction of a negative example two-dimensional image set based on at least one of the alignment breakpoint features of each second training contiguous group, the feature of short-read comparison of two-sided contiguous groups, and the reverse alignment feature includes: sliding a training length window along each second training contiguous group, and obtaining multiple second training two-dimensional images that include at least one of the alignment breakpoint features of each second training contiguous group, the feature of short-read comparison of two-sided contiguous groups, and the reverse alignment feature during the sliding process; and constructing a negative example two-dimensional image set based on the multiple second training two-dimensional images.

[0142] In order to make the resulting set of negative example two-dimensional images more abundant, the training length window can have different lengths to obtain more second training two-dimensional images including at least one of the following: alignment breakpoint features of each second training overlap group, short-read alignment features of two ends that are not in the same overlap group, and reverse alignment features.

[0143] In a readily understandable manner, the second training two-dimensional image includes at least one of the following: alignment breakpoint features of each second training contiguous group, features of short-read length comparison between different contiguous groups, and reverse alignment features. That is, the second training two-dimensional image may individually include alignment breakpoint features of each second training contiguous group, features of short-read length comparison between different contiguous groups, and reverse alignment features of each second training contiguous group. Alternatively, it may simultaneously include alignment breakpoint features of each second training contiguous group, features of short-read length comparison between different contiguous groups, and reverse alignment features of each second training contiguous group.

[0144] In this embodiment, a set of positive two-dimensional images constructed based on the assembly breakpoints of each first training contiguous group, and a set of negative two-dimensional images constructed based on at least one of the alignment breakpoint features, the double-ended short read length alignment features of each second training contiguous group, and the reverse alignment features, are used to train the initial classification model to obtain the target classification model. This enables the target classification model to correctly classify different situations of assembly errors and successful assembly through two-dimensional images. Furthermore, the target classification model obtained by the training method of this embodiment can ensure the accuracy of the detection results of the target contiguous group segments.

[0145] In an exemplary embodiment, the above-mentioned training of the initial classification model based on the positive example two-dimensional image set and the negative example two-dimensional image set to obtain the target classification model includes: training the initial classification model based on the positive example two-dimensional image set and the negative example two-dimensional image set to obtain the trained initial classification model; determining the AUPRC value and AUROC value of the trained initial classification model; and determining the trained initial classification model whose AUPRC value satisfies the preset AUPRC condition and whose AUROC value satisfies the preset AUROC value condition as the target classification model.

[0146] AUPRC stands for Area Under the Precision-Recall Curve. During the training of a classification model, the AUPRC value is the area under the precision-recall curve. Precision refers to the proportion of samples correctly identified as positive out of all samples correctly identified as positive, while recall refers to the proportion of samples correctly identified as positive out of all actual positive samples.

[0147] AUROC stands for Area Under the Receiver Operating Characteristic Curve. In the training process of a classification model, AUROC is the area under the receiver operating characteristic curve (ROC curve). The ROC curve is plotted by changing the threshold of the classifier to show the True Positive Rate (TPR) versus the False Positive Rate (FPR).

[0148] Optionally, satisfying the preset AUPRC condition can be that the AUPRC value reaches the preset AUPRC value; alternatively, satisfying the preset AUROC value condition can be that the AUROC value reaches the preset AUROC value.

[0149] For example, such as Figure 4AThe figure shows the precision-recall curve of the target classification model. The horizontal axis represents precision, and the vertical axis represents recall. At this point, the AUPRC value of the target classification model is 0.688. Figure 4B The figure shows the receiver operating characteristic curve of the target classification model. The horizontal axis represents the false positive rate and the vertical axis represents the true positive rate. At this time, the AUROC value of the target classification model is 0.944.

[0150] In an exemplary embodiment, the target classification model provided in this embodiment can be used to correct metagenomic assembly contigs that are detected as having assembly errors, thereby obtaining new detection results for metagenomic assembly contigs.

[0151] For example, when correcting multiple metagenomic assembly contigs with assembly errors based on the target classification model provided in this embodiment, such as... Figure 5A As shown, the length of the contiguous phase of the misassembled metagenomic assembly was 165,480,116 before correction and became 106,794,977 after correction; Figure 5B As shown, the number of metagenomic assembly contigs with assembly errors was 765 before correction and became 692 after correction; Figure 5C As shown, the number of sequence sites with assembly errors in the metagenomic assembly contiguous group was 1245 before correction and became 1012 after correction; Figure 5D As shown, the number of metagenomic assembly contigs with local assembly errors was 754 before correction and decreased to 534 after correction; Figure 5E As shown, the number of metagenomic assembly contigs with shift errors was 165 before correction and became 138 after correction; Figure 5F As shown, the number of metagenomic contigs in cases of inverted assembly errors was 10 before correction and decreased to 8 after correction; Figure 5G As shown, the number of metagenomic assembly contigs with interspecies translocation errors was 788 before correction and became 663 after correction; Figure 5H As shown, the number of metagenomic assembly contigs, which are intraspecific translocation errors, was 282 before correction and became 203 after correction. This demonstrates that the target classification model provided in this embodiment significantly improves performance in cases of assembly errors.

[0152] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0153] Based on the same inventive concept, this application also provides an assembly overlap group detection apparatus for implementing the assembly overlap group detection method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the assembly overlap group detection apparatus provided below can be found in the limitations of the assembly overlap group detection method described above, and will not be repeated here.

[0154] In one exemplary embodiment, such as Figure 6 As shown, an assembly overlap group detection device is provided, comprising: a comparison module 602, a first determination module 604, and an input module 606, wherein:

[0155] The alignment module 602 is used to acquire metagenomic assembly contigs and paired-end sequencing short read data, and to align the metagenomic assembly contigs with the paired-end sequencing short read data to obtain an alignment file.

[0156] The first determining module 604 is used to determine the target contiguous group segment in the comparison file and determine the comparison features of the target contiguous group segment; the comparison features include at least the comparison breakpoint features, the feature of comparing non-same contiguous group at both ends short read lengths, or the reverse comparison features.

[0157] The input module 606 is used to generate a two-dimensional image of the alignment features and input the two-dimensional image of the alignment features into the target classification model to obtain the detection results of the target overlapping group segments.

[0158] In an exemplary embodiment, the target contiguous group fragment includes multiple sequence sites; the first determining module 604 is further configured to perform at least one of the following steps: determining the number of alignment breakpoints for each sequence site to obtain alignment breakpoint features; determining the number of short reads for each sequence site that are not aligned to a contiguous group to obtain paired short read alignment features that are not in the same contiguous group; determining the number of short reads for each sequence site that are covered by short reads that are reverse aligned at one end of the paired short reads to obtain reverse alignment features.

[0159] In an exemplary embodiment, the target contiguous fragment includes multiple sequence sites; the first determining module 604 is further configured to perform at least one of the following steps: determining the short read coverage depth of each sequence site to obtain alignment depth features; based on the alignment depth features, determining the absolute value of the coverage depth difference between each sequence site and other sequence sites to obtain alignment depth difference features; determining the average length of the alignment positions of the paired short reads covered by each sequence site to obtain the average length feature of the paired short read alignment positions.

[0160] In an exemplary embodiment, the input module 606 is further configured to generate a first image based on the comparison breakpoint feature, the feature of the two-end short read length comparison of non-overlapping groups, or the reverse comparison feature; generate a second image based on the comparison depth feature, the comparison depth difference feature, or the feature of the average length of the two-end short read length comparison position; and stack the first image and the second image to obtain a two-dimensional image.

[0161] In an exemplary embodiment, the first image is a first diagonal image; the second image includes a second diagonal image generated based on alignment depth features or the average length feature of the alignment positions of the two-ended short reads, and a matrix image generated based on alignment depth difference features.

[0162] like Figure 6 As shown, in an exemplary embodiment, the above-mentioned device further includes a second determining module 608, which is used to analyze the alignment breakpoint features of the target overlapping group segment when the detection result of the target overlapping group segment is that the target overlapping group segment is assembled incorrectly, and to determine the target breakpoint based on the analysis result.

[0163] In an exemplary embodiment, the comparison quality of the above-mentioned comparison files meets a preset quality condition, and the length meets a preset length condition.

[0164] like Figure 6 As shown, in an exemplary embodiment, the above-described apparatus further includes a construction module 610, which is used to construct an initial classification model and multiple microbial communities, and obtain short read sequences with different sequencing depths based on the different abundances of each microbial community; assemble the short read sequences with different sequencing depths to obtain multiple training metagenomic assembly contigs; align the short read sequences with different sequencing depths and the multiple training metagenomic assembly contigs to obtain multiple training alignment files; evaluate each training alignment file to obtain the training detection results of each training alignment file; and train the initial classification model based on the multiple training detection results to obtain a target classification model.

[0165] In an exemplary embodiment, the construction module 610 is further configured to divide multiple training metagenomic assembly contigs into at least one first training contig and at least one second training contig based on multiple training detection results; the training detection result of the first training contig is an assembly error, and the training detection result of the second training contig is a correct assembly; a set of positive example two-dimensional images is constructed based on the assembly breakpoints of each first training contig; a set of negative example two-dimensional images is constructed based on at least one of the alignment breakpoint features, paired short read alignment features of different contigs, and reverse alignment features of each second training contig; and an initial classification model is trained based on the set of positive example two-dimensional images and the set of negative example two-dimensional images to obtain a target classification model.

[0166] Each module in the aforementioned overlapping group detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0167] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores metagenomic assembly contiguous group data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements an assembly contiguous group detection method.

[0168] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 8As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements an assembly overlap group detection method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0169] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0170] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0171] In one exemplary embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above-described method embodiments.

[0172] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0173] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0174] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0175] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0176] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for detecting overlapping groups in assembly, characterized in that, The method includes: Obtain metagenomic assembly contigs and paired-end sequencing short read data, and align the metagenomic assembly contigs with the paired-end sequencing short read data to obtain an alignment file; In the comparison file, a target contiguous group segment is identified, and the comparison features of the target contiguous group segment are determined; the comparison features include at least the comparison breakpoint feature, the feature of comparing two short read lengths at both ends to be different contiguous groups, or the reverse comparison feature; A two-dimensional image of the alignment features is generated and input into the target classification model to obtain the detection result of the target overlapping group segment.

2. The method according to claim 1, characterized in that, The comparison features also include at least the comparison depth feature, the comparison depth difference feature, or the average length feature of the comparison positions of the two short reads.

3. The method according to claim 1, characterized in that, The target contiguous fragment includes multiple sequence sites; Determining the alignment features of the target contiguous group fragment includes at least one of the following steps: The number of alignment breakpoints at each of the sequence sites is determined to obtain the alignment breakpoint features; The number of short reads that are not aligned to a contiguous group for each sequence site is determined to obtain the feature that the paired short reads are not aligned to the same contiguous group. The number of sequence sites covered by the short reads that are reverse-aligned at one end of the paired short reads is determined to obtain the reverse alignment feature.

4. The method according to claim 2, characterized in that, Determining the alignment features of the target contiguous group fragment includes at least one of the following steps: The short read coverage depth of each sequence site is determined to obtain the alignment depth feature; Based on the alignment depth features, the absolute value of the coverage depth difference between each sequence site and other sequence sites is determined to obtain the alignment depth difference features; The average length of the alignment positions of the paired short reads covered by each sequence site is determined to obtain the average length feature of the paired short read alignment positions.

5. The method according to claim 2, characterized in that, The generation of the two-dimensional image of the comparison features includes: The first image is generated based on the comparison of breakpoint features, the comparison of non-overlapping group features of short read lengths at both ends, or the reverse comparison features. A second image is generated based on the alignment depth features, alignment depth difference features, or the average length feature of the alignment positions of the two-end short read lengths. The first image and the second image are stacked to obtain the two-dimensional image.

6. The method according to claim 5, characterized in that, The first image is a first diagonal image; The second image includes a second diagonal image generated based on the alignment depth feature or the average length feature of the paired short read alignment positions, and a matrix image generated based on the alignment depth difference feature.

7. The method according to claim 1, characterized in that, The method further includes: When the detection result of the target overlapping group segment indicates that the target overlapping group segment is assembled incorrectly, the comparison breakpoint features of the target overlapping group segment are analyzed, and the target breakpoint is determined based on the analysis results.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: An initial classification model and multiple microbial communities were constructed, and short read sequences with different sequencing depths were obtained based on the different abundances of each microbial community. Short read sequences at different sequencing depths were assembled to obtain multiple training metagenomic assembly contigs; Short read sequences at different sequencing depths were aligned with multiple training metagenomic assembly contigs to obtain multiple training alignment files; Each of the training comparison files is evaluated to obtain the training detection results of each of the training comparison files. Based on multiple training and detection results, the initial classification model is trained to obtain the target classification model.

9. An assembly overlap group detection device, characterized in that, The device includes: The alignment module is used to acquire metagenomic assembly contigs and paired-end sequencing short read data, and to align the metagenomic assembly contigs with the paired-end sequencing short read data to obtain an alignment file; The first determining module is used to determine the target contiguous group segment in the comparison file and determine the comparison features of the target contiguous group segment; the comparison features include at least the comparison breakpoint feature, the feature of comparing different contiguous groups with short read lengths at both ends, or the reverse comparison feature; The input module is used to generate a two-dimensional image of the alignment features and input the two-dimensional image of the alignment features into the target classification model to obtain the detection result of the target overlapping group segment.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.