IGH gene fusion identification method and device, storage medium and program product

By preprocessing sequencing data and comparing it with IGH libraries and DNA sequence databases, the resolution and cost issues of IGH gene fusion detection in existing technologies have been resolved, enabling efficient and low-cost IGH fusion gene identification and promoting the diagnosis and personalized treatment of lymphoma.

CN121237218APending Publication Date: 2025-12-30BOE TECHNOLOGY GROUP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410866827.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing fluorescence in situ hybridization (FISH) methods suffer from low resolution, limited diversity, high cost, and time consumption when identifying IGH gene fusions, making it difficult to effectively monitor IGH fusion patterns in lymphoma patients.

Method used

Sequencing data is preprocessed and aligned to the IGH library to obtain candidate sequences. Unknown regions are aligned to the DNA sequence database. The optimal alignment results that meet the similarity criteria are analyzed to identify IGH region fusion genes. A multi-step sequence classification and filtering method is used to reduce the amount of downstream data analysis.

Benefits of technology

It improves the sensitivity and speed of IGH fusion gene detection, reduces computational resources and analysis time, has strong applicability and low cost, and can quickly identify IGH fusion genes and patterns in various sequencing types, which has important clinical value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237218A_ABST
    Figure CN121237218A_ABST
Patent Text Reader

Abstract

The invention discloses an IGH gene fusion identification method and device, a storage medium and a program product, and the method comprises the following steps: obtaining sequencing data, and preprocessing the obtained sequencing data; the preprocessed data is compared to an IGH library, one or more candidate sequences are obtained, the candidate sequences are sequences containing IGH sections and unknown sections, and the unknown sections are sequences with the length within a preset first length range and are not compared to IGH fragments; acquiring the unknown sections, and comparing the unknown sections to a DNA sequence database to obtain a plurality of comparison results; and analyzing the plurality of comparison results, and identifying the gene corresponding to the optimal comparison result meeting a preset similarity condition as the gene fused with the IGH segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to, but is not limited to, the field of biotechnology, and particularly to a method and apparatus for identifying IGH gene fusions, as well as a storage medium and program product. Background Technology

[0002] Lymphoma is a clinically aggressive and heterogeneous disease, exhibiting diversity in clinical outcomes, genomic characteristics, and cells of origin. Immunochemotherapy is the most common treatment for this disease, with the most frequently used regimen being rituximab combined with cyclophosphamide, doxorubicin, vincristine, and prednisone (R-CHOP). Despite the significant efficacy of R-CHOP, a considerable proportion of patients experience treatment refractory and relapsed outcomes. Reports indicate that patients with dozens of fusion patterns, such as BCL6 / IGH6 and MYC / IGH, have poor overall survival after receiving R-CHOP, exhibit high resistance to R-CHOP, and demonstrate treatment refractory and relapsed characteristics. These fusions may originate from conversions during treatment, as seen in reported cases of transformation from diffuse large B-cell lymphoma (DLBCL) to plasmablastic lymphoma with a MYC / IGH fusion pattern. Therefore, identifying the potential emergence of new IGH fusion patterns is crucial in clinical surveillance.

[0003] Related technologies identify potential IGH fusions through fluorescence in situ hybridization (FISH) experiments, but this identification method has drawbacks such as low resolution (usually within the range of a few hundred base pairs), limited diversity (detectable only for specific fusion modes), high cost, and time consumption. Summary of the Invention

[0004] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.

[0005] This disclosure provides a method for identifying IGH gene fusions, including:

[0006] Acquire sequencing data and preprocess the acquired sequencing data;

[0007] The preprocessed data is compared with the IGH library to obtain one or more candidate sequences. The candidate sequence is a sequence containing IGH segments and unknown segments. The unknown segments are sequences whose length is within a preset first length range and which have not been compared with IGH segments.

[0008] The unknown segment is obtained, and the unknown segment is aligned to a DNA sequence database to obtain multiple alignment results;

[0009] Analyze multiple alignment results and identify the gene corresponding to the best alignment result that meets the preset similarity conditions as the gene fused with the IGH segment.

[0010] This disclosure also provides an IGH gene fusion identification device, including a memory; and a processor connected to the memory, the memory being used to store instructions, and the processor being configured to execute the steps of the IGH gene fusion identification method according to any embodiment of this disclosure based on the instructions stored in the memory.

[0011] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the IGH gene fusion identification method described in any embodiment of this disclosure.

[0012] This disclosure also provides a program product including instructions that, when executed by a computer, perform the IGH gene fusion identification method as described in any embodiment of this disclosure.

[0013] This disclosure also provides an IGH gene fusion identification device, including a preprocessing module, a first comparison module, a second comparison module, and a result analysis module, wherein:

[0014] The preprocessing module is configured to acquire sequencing data and preprocess the acquired sequencing data.

[0015] The first comparison module is configured to compare the preprocessed data with the IGH library to obtain one or more candidate sequences. The candidate sequence is a sequence containing IGH segments and unknown segments. The unknown segments are sequences whose length is within a preset first length range and which have not been compared with IGH segments.

[0016] The second alignment module is configured to acquire the unknown segment, align the unknown segment to a DNA sequence database, and obtain multiple alignment results;

[0017] The result analysis module is configured to analyze multiple alignment results and identify the gene corresponding to the best alignment result that meets the preset similarity conditions as a gene fused with the IGH segment.

[0018] The IGH gene fusion identification method, apparatus, storage medium, and program product disclosed herein involve aligning preprocessed data to an IGH library to obtain one or more candidate sequences. Unknown segments within these candidate sequences are then aligned to a non-redundant nucleic acid sequence database to obtain multiple alignment results. Analysis of these results identifies the gene corresponding to the optimal alignment result that meets preset similarity criteria as the gene fused with the IGH segment. This multi-step sequence classification and filtering approach reduces the amount of downstream data analysis, effectively lowers computational resources and analysis time, and improves processing speed. It exhibits high sensitivity for detecting IGH fusion genes and is applicable to various sequencing data types, demonstrating strong applicability and low cost. It can quickly and effectively identify potential IGH fusion genes and fusion patterns. This method has significant clinical value in tumor diagnosis, prognosis, and personalized treatment, helping to predict patient prognosis and treatment efficacy, contributing to a deeper understanding of the role of gene fusion in disease development, and promoting basic and translational medical research.

[0019] Other features and advantages of this disclosure will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the disclosure. Other advantages of this disclosure may be realized and obtained by means of the methods described in the description and the accompanying drawings. Attached Figure Description

[0020] The accompanying drawings are used to provide an understanding of the technical solutions of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0021] Figure 1 A flowchart illustrating an exemplary embodiment of the present disclosure of an IGH gene fusion identification method;

[0022] Figure 2 This is a schematic diagram illustrating the structure of two candidate sequences as an exemplary embodiment of this disclosure;

[0023] Figure 3 This is a schematic diagram illustrating the results of sequence filtering and sequence partitioning as exemplified in this disclosure;

[0024] Figure 4 This is a schematic diagram of the structure of sequence partitioning as an exemplary embodiment of this disclosure;

[0025] Figure 5 A schematic diagram of the structure of an IGH gene fusion identification device provided for an exemplary embodiment of this disclosure;

[0026] Figure 6 A schematic diagram of another IGH gene fusion identification device provided as an exemplary embodiment of this disclosure. Detailed Implementation

[0027] This disclosure describes several embodiments, but these descriptions are exemplary and not limiting, and it will be apparent to those skilled in the art that many more embodiments and implementations are possible within the scope of the embodiments described herein. Although many possible combinations of features are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with, or may replace, any feature or element of any other embodiment.

[0028] This disclosure includes and contemplates combinations of features and elements known to those skilled in the art. The embodiments, features, and elements disclosed in this disclosure may also be combined with any conventional features or elements to form a unique inventive scheme as defined by the claims. Any feature or element of any embodiment may also be combined with features or elements from other inventive schemes to form another unique inventive scheme as defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in this disclosure may be implemented individually or in any suitable combination. Therefore, the embodiments are not limited except by the limitations imposed by the appended claims and their equivalents. Furthermore, various modifications and changes may be made within the scope of the appended claims.

[0029] Furthermore, in describing representative embodiments, the specification may have presented methods and / or processes as a specific sequence of steps. However, the method or process should not be limited to the specific order of steps described herein, to the extent that the method or process does not depend on the specific order of steps described herein. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation of the claims. Moreover, the claims relating to the method and / or process should not be limited to the steps performed in the order written, and those skilled in the art will readily understand that these orders can be varied and still remain within the spirit and scope of the embodiments disclosed herein.

[0030] like Figure 1 As shown in the embodiments of this disclosure, a method for identifying IGH gene fusions is provided, including:

[0031] Step 101: Obtain sequencing data and preprocess the obtained sequencing data;

[0032] Step 102: Align the preprocessed data with the IGH library to obtain one or more candidate sequences. The candidate sequence is a sequence containing IGH segments and unknown segments. The unknown segments are sequences whose length is within a preset first length range and which have not been aligned with IGH segments.

[0033] Step 103: Obtain the unknown segment, align the unknown segment to the DNA sequence database, and obtain multiple alignment results;

[0034] Step 104: Analyze multiple alignment results and identify the gene corresponding to the best alignment result that meets the preset similarity conditions as the gene fused with the IGH segment.

[0035] The IGH gene fusion identification method provided in this disclosure involves aligning preprocessed data to an IGH library to obtain one or more candidate sequences. Unknown regions within these candidate sequences are then aligned to a non-redundant nucleic acid sequence database to obtain multiple alignment results. Analysis of these results identifies the gene corresponding to the optimal alignment that meets preset similarity criteria as the gene fused with the IGH region. This multi-step sequence classification and filtering approach reduces the amount of downstream data analysis, effectively lowers computational resources and analysis time, and improves processing speed. It exhibits high sensitivity for detecting IGH fusion genes and is applicable to various sequencing data types, demonstrating strong applicability and low cost. It can quickly and effectively identify potential IGH fusion genes and fusion patterns. This method has significant clinical value in tumor diagnosis, prognosis, and personalized treatment, helping to predict patient prognosis and treatment efficacy, contributing to a deeper understanding of the role of gene fusion in disease development, and promoting basic and translational medical research.

[0036] In some exemplary embodiments, the acquired sequencing data is preprocessed, including:

[0037] Preliminary quality control of sequencing data is performed to determine whether the amount and quality values ​​of the sequencing data meet the requirements for subsequent data processing.

[0038] The sequencing data were statistically analyzed, and low-quality data were filtered out. Low-quality data included at least one of the following: low average base quality, containing structural sequences, or sequences with excessively high N content.

[0039] In this embodiment of the disclosure, software such as Fastp (a tool for quality control and data preprocessing of high-throughput sequencing data) and trim_galore (a tool for quality control and trimming of sequencing data) can be used for data filtering; however, this disclosure does not limit this.

[0040] For example, taking the data of 7 samples in Table 1 as an example, all 7 samples were diagnosed with non-Hodgkin's lymphoma, and all 7 samples underwent R-CHOP treatment before sequencing. The data of these 7 samples were statistically analyzed to obtain the sequencing data base number, sequence (reads) number, and Q30 quality value as shown in Table 1.

[0041] Sample name Number of bases Number of Reads Q30 Ratio (%) 1 180054600 1200364 97.31 / 95.12 2 240817200 1605448 97.49 / 92.61 3 247156200 1647708 97.21 / 95.7 4 186371250 1242475 97.43 / 94.76 5 116570250 777135 96.81 / 94.32 6 54637200 364248 95.94 / 94.12 7 39758250 265055 95.9 / 94.87

[0042] Table 1 shows the results of deconnecting and filtering low-quality data from the 7 sample data, as shown in Table 2.

[0043]

[0044]

[0045] Table 2

[0046] In this embodiment of the disclosure, prior to step 102, the method further includes: downloading human IGHV, IGHD, and IGHJ gene data from the International Immunogenetics Information System (IMGT) database for use as an IGH library.

[0047] The IMGT database is a global reference database for immunogenetics and immunoinformatics created by the University of Montpellier and the French National Centre for Scientific Research. It consists of sequence databases, genome databases, structural databases, monoclonal antibody databases, web resources, and interactive tools, providing universal access to sequence, genomic, and structural immunogenetics data.

[0048] In this embodiment of the disclosure, the preprocessed data can be aligned to the IGH library using sequence alignment tools (such as igblast, mixcr, etc.) to obtain the distribution and length of IGHV, IGHD, and IGHJ on each sequence. However, this disclosure does not limit this.

[0049] In some exemplary embodiments, in step 102, the preprocessed data is aligned to the IGH library to obtain one or more candidate sequences, including:

[0050] Detect whether the R1 end sequence and / or R2 end sequence includes the first fragment, and whether the first fragment can be matched with the IGHV fragment or the IGHJ fragment;

[0051] Select the R1 end sequence and / or R2 end sequence including the first segment as the IGH sequence;

[0052] The test detects whether the IGH sequence includes a second fragment located at the left end of the R1 sequence position or the right end of the R2 sequence position, and cannot be compared with any of the following fragments: IGHD fragment, IGHV fragment, and IGHJ fragment;

[0053] The IGH sequence including the second segment was selected as the candidate sequence, and the first segment was divided into the IGH segment, while the second segment was divided into the unknown segment.

[0054] For Fastq files of paired-end sequencing, R1 and R2 are used as identifiers. The sequencing results at both ends are stored in two separate files. This can be understood as a DNA or RNA fragment. Sequencing is performed base by base from both ends toward the middle to read the sequence information, which will generate two reads: the R1 end sequence and the R2 end sequence. These two reads are a pair.

[0055] In this embodiment of the disclosure, the IGH library includes multiple IGH fragments, including IGHV fragments, IGHD fragments, and IGHJ fragments. When the R1 end sequence and / or the R2 end sequence includes a first fragment that can be matched with an IGHV fragment or an IGHJ fragment, the R1 end sequence and / or the R2 end sequence can be identified as an IGH sequence. That is, when R1 or R2 contains an IGHV fragment or an IGHJ fragment, the sequence is considered to be an IGH sequence.

[0056] In some exemplary embodiments, the length of the first fragment that is aligned to the IGHV fragment is greater than n1, where n1 is between 40bp and 80bp, and the length of the first fragment that is aligned to the IGHJ fragment is greater than n2, where n2 is between 15bp and 25bp.

[0057] In this embodiment of the disclosure, the interval length of the aligned IGHV fragment should be greater than n1, and the range of n1 can be between 40bp and 80bp. For example, the value of n1 can be set to 50bp. The interval length of the aligned IGHJ fragment should be greater than n2, and the range of n2 can be between 15bp and 25bp. For example, the value of n2 can be set to 20bp.

[0058] In some exemplary embodiments, the similarity between the first fragment and the IGHV or IGHJ fragment is greater than or equal to n3, where n3 ≥ 85%.

[0059] In this embodiment of the disclosure, the similarity (identity) value between the first fragment and the IGHV fragment or the IGHJ fragment should be greater than or equal to n3, where n3 ≥ 85%. For example, the value of n3 can be set to 90%.

[0060] In this embodiment, the preprocessed data is aligned to the IGH library so that the R1 or R2 end sequence contains an IGH sequence fragment. This portion of the obtained sequence eliminates non-specifically amplified sequences, reducing the amount of data required for subsequent processing and shortening processing time.

[0061] In this embodiment of the disclosure, such as Figure 2 As shown, the unknown segment is located at the left end of the R1 sequence position or the right end of the R2 sequence position.

[0062] In some exemplary embodiments, the preset first length range is between 35bp and 60bp.

[0063] In this embodiment, if a sequence at the left end of the R1 sequence position or the right end of the R2 sequence position does not align to the IGHV, IGHD, or IGHJ fragment, it may be a candidate fusion sequence. The length of this unknown segment is m. The longer the length, the easier it is to identify the target gene in subsequent analysis. The length m can be set to 35–60 bp.

[0064] In step 102, the final sequence obtained is one containing the IGH region and also containing an unknown region; this sequence is used as a candidate sequence. This candidate sequence is used for subsequent fusion gene identification. Figure 3 As shown, this step can reduce the number of sequences by over 95%. Since the NT database used in subsequent analysis contains known, non-repeating sequences, the database is very large. This step can reduce the data analysis time by over 99% and reduce the computing resources required for data analysis.

[0065] Taking the data in Table 2 above as an example, after this step, the candidate sequences obtained are shown in Table 3. In this embodiment, the number of candidate sequences obtained does not exceed 2% of the original number of sequences. A large number of sequences do not need to be compared with the NT database. Through sequence partitioning, more than 50% of the sequence fragments do not need to be analyzed by NT, which saves more than 99% of the analysis time.

[0066] Sample name sequence number Number of clone sequences Percentage of cloned sequences (%) Number of candidate sequences 1 1190437 1059006 88.96 8842 2 1539356 1105683 71.83 20137 3 1641795 1525834 92.94 18528 4 1234869 1089043 88.19 9610 5 765407 399267 52.16 559 6 352225 111076 31.54 7002 7 251239 3265 1.3 140

[0067] Table 3

[0068] In this embodiment of the disclosure, only the paired R1 and R2 contain IGH segments, and only the non-IGH segments (unknown segments) are used for subsequent analysis. For example... Figure 4 As shown, by partitioning the candidate sequences (i.e. dividing them into unknown segments and IGH segments) and using the unknown segments for subsequent analysis, the average sequence length used for subsequent analysis is reduced by more than 50%, which in turn reduces the data analysis time for subsequent analysis by more than 50%.

[0069] In this embodiment of the disclosure, in step 103, the DNA sequence database can be the NT database. The NT database is a non-redundant nucleic acid sequence database, the official nucleic acid sequence database of the National Center for Biotechnology Information (NCBI) in the United States. The data originates from GenBank, EMBL, and DDBJ, and is NCBI's default nucleic acid BLAST alignment database. By aligning unknown regions to the non-redundant nucleic acid sequence database, multiple alignment results can be obtained.

[0070] In some exemplary embodiments, step 104 involves analyzing multiple alignment results, including:

[0071] Filter through multiple comparison results;

[0072] Select the optimal comparison result for each unknown segment from the filtered comparison results;

[0073] Check whether the optimal alignment result meets the preset similarity conditions.

[0074] In some exemplary embodiments, when filtering multiple alignment results, the filtering criteria are: alignment length ≥ n4 and alignment coverage ≥ n5, where n4 ≥ 30 bp and n5 ≥ 85%.

[0075] In this embodiment of the disclosure, when analyzing multiple alignment results, the results are first screened. The screening criteria are: the length of the alignment should be ≥30bp, and the alignment coverage should be ≥85%. For example, if the alignment length is 50bp, the length of the alignment should be greater than or equal to 85%, that is, the length of the alignment should be greater than or equal to 50*85%≈42bp. Then, the optimal alignment result corresponding to each unknown segment is selected from the screened alignment results, and it is then checked whether the selected optimal alignment result meets the preset similarity condition. If the preset similarity condition is met, it is determined that there is IGH gene fusion, and the NT library sequence corresponding to the optimal alignment result can be identified as a gene fused with the corresponding IGH segment; if the preset similarity condition is not met, it is considered that there is no IGH gene fusion.

[0076] In some exemplary embodiments, the preset similarity condition is: the similarity between the unknown segment and the DNA sequence in the DNA sequence database is ≥85%.

[0077] In this embodiment of the disclosure, the preset similarity condition can be that the identity value of the comparison segment should be greater than or equal to 85%.

[0078] Taking the data in Table 3 as an example, the unknown segment sequences of the candidate sequences were annotated using the NT library, and the fusion gene identification results were as follows. Among them, :: indicates the fusion type, such as LOC::IGH indicating the fusion of LOC and IGH, _ indicates structural connection, - indicates number, and / indicates that a specific type cannot be identified, such as IGHJ1 / IGHJ4 / IGHJ5 indicating that IGHJ1, IGHJ4, and IGHJ5 are all possible.

[0079] 1. LOC110742091::IGHV3-74_IGHD3-10_IGHJ6; OST4::IGHV4-39_IGHD4-17_IGHJ4; SRRM2::IGHV2-70_IGHD1-14_IGHJ5;

[0080] 2. MYH9::IGHV3-15_IGHD2-2_IGHJ5; APOC1::IGHV3-65_IGHD1-7_IGHJ4;

[0081] 3. NDRG1::IGHV3-21_IGHD5-24_IGHJ4; LAPTM4B::IGHV3-11_IGHV3-53_IGHJ4; PLXNB2::IGHV2-26_IGHD3-10_IGHJ4; PNN::IGHV3-30-3_IGHD6-13_IGHJ4;

[0082] 4. DNM2::IGHV3-23_IGHD3-3_IGHJ4; LOC100975460::IGHV3-30-3_IGHD5-12_IGHJ1;

[0083] 5. ARF1::IGHV3-72_IGHD1-1_IGHJ4; LY6E::IGHV4-5_IGHD6-13_IGHJ2; LCP1::IGHV3-48_IGHD2-15_IGHJ6;

[0084] 6. GCNT2::IGHV3-23_IGHD1-20_IGHJ3;

[0085] 7. WFIKKN1::IGHV2-5_IGHD3-22_IGHJ6; TNS3::IGHV2-5_IGHJ2; ARHGAP39::IGHV2-26_IGHJ2; GN AS::IGHV2-26_IGHJ1 / IGHJ4; Nxf1::IGHV2-5_IGHJ1 / IGHJ4 / IGHJ5; Adamts9::IGHV2-26_IGHJ2.

[0086] The IGH gene fusion identification method of this disclosure can identify IGH fusions that may exist in sequencing data, and can cover multiple types of IGH and fusion genes. Through multi-step sequence classification, it greatly reduces the overall data analysis time (by up to 99%).

[0087] like Figure 5 As shown in the embodiments of this disclosure, an IGH gene fusion identification device is also provided, including a preprocessing module 510, a first alignment module 520, a second alignment module 530, and a result analysis module 540, wherein:

[0088] The preprocessing module 510 is configured to acquire sequencing data and preprocess the acquired sequencing data.

[0089] The first comparison module 520 is configured to compare the preprocessed data with the IGH library to obtain one or more candidate sequences. The candidate sequences are sequences containing IGH segments and unknown segments. The unknown segments are sequences whose length is within a preset first length range and which have not been compared with IGH segments.

[0090] The second alignment module 530 is configured to acquire unknown segments, align the unknown segments to the DNA sequence database, and obtain multiple alignment results.

[0091] The results analysis module 540 is configured to analyze multiple alignment results and identify the gene corresponding to the best alignment result that meets the preset similarity conditions as a gene fused with the IGH segment.

[0092] In some exemplary embodiments, the unknown segment is located at the left end of the R1 end sequence position or the right end of the R2 end sequence position.

[0093] In some exemplary embodiments, the preset first length range is between 35bp and 60bp.

[0094] In some exemplary embodiments, the first comparison module 520 compares the preprocessed data with the IGH library to obtain one or more candidate sequences, including:

[0095] Detect whether the R1 end sequence and / or R2 end sequence includes the first fragment, and whether the first fragment can be matched with the IGHV fragment or the IGHJ fragment;

[0096] The R1 end sequence and / or R2 end sequence, including the first fragment, are identified as IGH sequences;

[0097] The test detects whether the IGH sequence includes a second fragment located at the left end of the R1 sequence position or the right end of the R2 sequence position, and cannot be compared with any of the following fragments: IGHD fragment, IGHV fragment, and IGHJ fragment;

[0098] The IGH sequence containing the second fragment is identified as the candidate sequence, and the first fragment is divided into the IGH segment, while the second fragment is divided into the unknown segment.

[0099] In some exemplary embodiments, the length of the first fragment that is aligned to the IGHV fragment is greater than n1, where n1 is between 40bp and 80bp, and the length of the first fragment that is aligned to the IGHJ fragment is greater than n2, where n2 is between 15bp and 25bp.

[0100] In some exemplary embodiments, the similarity between the first fragment and the IGHV or IGHJ fragment is greater than or equal to n3, where n3 ≥ 85%.

[0101] In some exemplary embodiments, the result analysis module 540 analyzes multiple alignment results, including:

[0102] Filter through multiple comparison results;

[0103] Select the optimal comparison result for each segment from the filtered comparison results;

[0104] Check whether the optimal alignment result meets the preset similarity conditions.

[0105] In some exemplary embodiments, when the result analysis module 540 filters multiple alignment results, the filtering conditions are: alignment length ≥ n4 and alignment coverage ≥ n5, where n4 ≥ 30 bp and n5 ≥ 85%.

[0106] In some exemplary embodiments, the preset similarity condition is: the similarity between the unknown segment and the DNA sequence in the DNA sequence database is ≥85%.

[0107] In some exemplary embodiments, the DNA sequence database is a non-redundant nucleic acid sequence database.

[0108] This disclosure also provides an IGH gene fusion identification device, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the IGH gene fusion identification method as described in any embodiment of this disclosure based on the instructions stored in the memory.

[0109] like Figure 6As shown, in one example, the IGH gene fusion identification device may include: a processor 610, a memory 620, a bus system 630, and a transceiver 640, wherein the processor 610, the memory 620, and the transceiver 640 are connected via the bus system 630, the memory 620 is used to store instructions, and the processor 610 is used to execute the instructions stored in the memory 620 to control the transceiver 640 to transmit and receive signals. Specifically, transceiver 640 can acquire sequencing data under the control of processor 610. Processor 610 preprocesses the acquired sequencing data; the preprocessed data is aligned to an IGH library to obtain one or more candidate sequences. The candidate sequences are sequences containing an IGH region and an unknown region. The unknown region is a sequence whose length is within a preset first length range and has not been aligned to an IGH fragment. The unknown region is acquired and aligned to a DNA sequence database to obtain multiple alignment results. The multiple alignment results are analyzed, and the gene corresponding to the optimal alignment result that meets the preset similarity conditions is identified as a gene fused with the IGH region.

[0110] It should be understood that processor 610 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0111] Memory 620 may include read-only memory and random access memory, and provides instructions and data to processor 610. A portion of memory 620 may also include non-volatile random access memory. For example, memory 620 may also store device type information.

[0112] In addition to a data bus, the bus system 630 may also include a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 6 The general labeled all buses as Bus System 630.

[0113] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the processor 610 or through software instructions. That is, the method steps of this embodiment can be executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in memory 620, and the processor 610 reads information from memory 620 and, in conjunction with its hardware, completes the steps of the aforementioned method. To avoid repetition, further details are omitted here.

[0114] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the IGH gene fusion identification method as described in any embodiment of this disclosure. The method for identifying IGH gene fusion by executing executable instructions is essentially the same as the IGH gene fusion identification method provided in the above embodiments of this disclosure, and will not be described in detail here.

[0115] In some possible implementations, various aspects of the IGH gene fusion identification method provided in this disclosure may also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the IGH gene fusion identification method according to various exemplary embodiments of this disclosure as described above. For example, the computer device may execute the IGH gene fusion identification method described in the embodiments of this disclosure.

[0116] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0117] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0118] It should be noted that the above embodiments or implementation methods are merely exemplary and not restrictive. Therefore, this disclosure is not limited to the content specifically shown and described herein. Various modifications, substitutions, or omissions can be made to the form and details of the implementations without departing from the scope of this disclosure.

Claims

1. A method for IGH gene fusion identification, characterized in that, The method comprises the following steps: obtaining sequencing data, and preprocessing the obtained sequencing data; aligning the preprocessed data to an IGH library to obtain one or more candidate sequences, wherein the candidate sequence is a sequence comprising an IGH segment and an unknown segment, and the unknown segment is a sequence with a length within a preset first length range and not aligned to an IGH segment; obtaining the unknown segment, and aligning the unknown segment to a DNA sequence database to obtain a plurality of alignment results; analyzing the plurality of alignment results, and identifying a gene corresponding to an optimal alignment result satisfying a preset similarity condition as a gene fused with the IGH segment.

2. The method of claim 1, wherein, The unknown segment is located at the left end of the R1 end sequence position or the right end of the R2 end sequence position.

3. The method of claim 1, wherein, The preset first length range is between 35 bp and 60 bp.

4. The method of claim 1, wherein, The step of aligning the preprocessed data to the IGH library to obtain one or more candidate sequences comprises the following steps: detecting whether the R1 end sequence and / or the R2 end sequence comprises a first segment, wherein the first segment can be aligned to an IGHV segment or an IGHJ segment; identifying the R1 end sequence and / or the R2 end sequence comprising the first segment as an IGH sequence; detecting whether the IGH sequence comprises a second segment, wherein the second segment is located at the left end of the R1 end sequence position or the right end of the R2 end sequence position, and cannot be aligned to any of the following segments: an IGHD segment, an IGHV segment, and an IGHJ segment; identifying the IGH sequence comprising the second segment as the candidate sequence, and dividing the first segment into the IGH segment and the second segment into the unknown segment.

5. The method of claim 4, wherein, The length of the first segment aligned to the IGHV segment is greater than n1, and the length of the first segment aligned to the IGHJ segment is greater than n2, wherein n1 is between 40 bp and 80 bp, and n2 is between 15 bp and 25 bp.

6. The method of claim 4, wherein, The similarity of the first segment to the IGHV segment or the IGHJ segment is greater than or equal to n3, and n3 is greater than or equal to 85%.

7. The method of claim 1, wherein, The step of analyzing the plurality of alignment results comprises the following steps: screening the plurality of alignment results; selecting an optimal alignment result corresponding to each segment from the screened alignment results; detecting whether the optimal alignment result satisfies a preset similarity condition.

8. The method of claim 7, wherein, When screening the plurality of alignment results, the screening condition is that the alignment length is greater than or equal to n4, and the alignment coverage is greater than or equal to n5, wherein n4 is greater than or equal to 30 bp, and n5 is greater than or equal to 85%.

9. The method of claim 1, wherein, The preset similarity condition is that the similarity of the unknown segment to a DNA sequence in the DNA sequence database is greater than or equal to 85%.

10. The method of claim 1, wherein, The DNA sequence database is a non-redundant nucleic acid sequence database.

11. An IGH gene fusion identification device, comprising: The method comprises the following steps:

12. A computer-readable storage medium, characterized in that, a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the IGH gene fusion identification method according to any one of claims 1 to 10 based on the instructions stored in the memory. A computer program is stored thereon, and the program is executed by a processor to implement the IGH gene fusion identification method according to any one of claims 1 to 10.

13. A computer program product, characterised in that, The computer program product comprises instructions for performing the method of identifying an IGH gene fusion according to any one of claims 1 to 10 when the computer program product is executed by a computer.

14. An IGH gene fusion identification device, comprising: The computer program product comprises a preprocessing module, a first alignment module, a second alignment module, and a result analysis module, wherein: The preprocessing module is configured to obtain sequencing data and preprocess the obtained sequencing data; The first alignment module is configured to align the preprocessed data to an IGH library to obtain one or more candidate sequences, the candidate sequence being a sequence comprising an IGH segment and an unknown segment, the unknown segment being a sequence with a length within a preset first length range and not aligned to an IGH fragment; The second alignment module is configured to obtain the unknown segment and align the unknown segment to a DNA sequence database to obtain a plurality of alignment results; The result analysis module is configured to analyze the plurality of alignment results and identify the gene corresponding to the optimal alignment result satisfying a preset similarity condition as the gene fused with the IGH segment.