A method and related device for analyzing and screening specificities of b cell immune repertoire antibody sequences

By performing germline alignment and multidimensional feature analysis on antibody sequencing sequences from B-cell immune repertoires, clonal collections and phylogenetic forests were constructed, solving the problem of difficulty in identifying high-affinity BCR sequences in existing technologies and achieving efficient screening and identification.

CN121789769BActive Publication Date: 2026-05-08广州赛业百沐生物科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
广州赛业百沐生物科技有限公司
Filing Date
2026-03-03
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify high-affinity BCR sequences in B-cell immune repertoires, leading to resource waste and missed potential high-quality sequences. Furthermore, bioinformatics analysis is insufficient to uncover multidimensional and comprehensive features.

Method used

By performing germline alignment, gene integrity and structural integrity filtering on antibody sequencing sequences from B cell immune repertoires, a clonal set and phylogenetic forest were constructed. Multidimensional feature analysis was conducted to generate a visualization graph, and high-affinity BCR sequences were screened out using preset filtering conditions.

Benefits of technology

It enables panoramic characterization and highly specific screening of BCR sequences in B cell immune repertoires, effectively identifying high-affinity BCR sequences and overcoming the limitations of single-dimensional analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789769B_ABST
    Figure CN121789769B_ABST
Patent Text Reader

Abstract

The application discloses a B cell immune repertoire antibody sequence feature analysis and specific screening method and related equipment, which can be applied to the technical field of data processing. According to the application, after the antibody sequencing sequence is subjected to germ line comparison and recognition to obtain corresponding first sequencing Fv sequence and germ line Fv sequence, the first sequencing Fv sequence is subjected to gene integrity filtering and structure integrity filtering, clonotype division and root node sequence reconstruction to obtain a second clonotype collection, and then a first lineage forest is constructed, and according to the second sequencing Fv sequence, the second clonotype collection or the first lineage forest, sequence feature analysis in multiple dimensions is carried out, so that the antibody sequence features can be analyzed from multiple dimensions such as sequence quality, sequence abundance, mutation degree, mutation preference and aggregation degree, the filtered BCR sequence has high affinity, and according to the sequence feature analysis results in multiple dimensions or the first lineage forest, a visualized graph is generated, so that relevant personnel can view the sequence features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method and related equipment for analyzing the sequence characteristics and specific screening of antibodies from a B-cell immune repertoire. Background Technology

[0002] In related technologies, the B cell repertoire refers to the collection of B cell receptor sequences (BCRs) or antibody variable region sequences from all B cells in an individual, reflecting the body's ability to recognize various antigens and the diversity of immune responses. With the rapid development of high-throughput sequencing technologies (such as single-cell paired sequencing), it is now possible to sequence B cell repertoires at extremely high resolution and depth, thereby obtaining massive amounts of antibody sequence data. This antibody sequence data contains rich information, including key immunological events such as B cell clonal expansion, maturation processes, and responses to antigens.

[0003] However, facing such massive antibody sequence data, effectively mining and analyzing valuable BCR sequences has become an urgent problem to be solved. For example, in current antibody discovery projects, identifying which BCR sequences have higher affinity for specific antigens remains difficult because directly determining antigen specificity from the sequence is still challenging. Due to the time-consuming and costly nature of these methods, the entire library is not tested (SPR, BLI, ELISA). The common practice is to select high-abundance sequences for validation, which may lead to the loss of some better BCR sequences. Even after investing significant resources to test hundreds of sequences, strong binding may not be found, often resulting in the loss of better BCR sequences. Furthermore, the limited experimental data makes it insufficient to train deep learning or machine learning models. Current bioinformatics analysis of B cell immune repositories focuses only on clonal diversity, V / J gene usage frequency, somatic hypermutation (SHM), and evolutionary analysis, failing to mine multidimensional and comprehensive characteristics of BCR sequence sets, and even less able to identify high-affinity BCR sequences.

[0004] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention

[0005] The main objective of this application is to propose a method and related equipment for analyzing the characteristics and specificity of B cell immune repertoire antibody sequences, which can effectively identify BCR sequences with high affinity.

[0006] To achieve the above objectives, one aspect of this application proposes a method for analyzing the sequence characteristics and specificity screening of antibodies from a B-cell immune repertoire, the method comprising the following steps:

[0007] Obtain antibody sequencing sequences from the B-cell immune repertoire;

[0008] The antibody sequencing sequences are subjected to phylogenetic alignment to identify the first sequencing Fv sequence corresponding to each antibody sequencing sequence, and to generate the phylogenetic Fv sequence corresponding to each antibody sequencing sequence. The first sequencing Fv sequence includes a V gene fragment, a D gene fragment, a J gene fragment, or a C gene fragment.

[0009] Gene integrity filtering and structural integrity filtering are performed on the first sequencing Fv sequence to obtain the second sequencing Fv sequence;

[0010] Based on the germline Fv sequence, all second sequencing Fv sequences are divided into several first clonal type sets, and the root node sequence of each first clonal type set is reconstructed to obtain a second clonal type set.

[0011] A first lineage forest is constructed based on the same type category conversion probability and the second clone set. The first lineage forest includes several first evolutionary trees, and each first evolutionary tree corresponds to a second clone set.

[0012] Based on the second sequencing Fv sequence, the second clonal set, or the first lineage forest, a multi-dimensional sequence feature analysis is performed. The multi-dimensional sequence feature analysis includes sequence quality feature analysis, sequence abundance feature analysis, mutation degree feature analysis, mutation preference feature analysis, and clustering degree feature analysis.

[0013] A visualization map is generated based on the results of sequence feature analysis from multiple dimensions or the first lineage forest.

[0014] Based on the sequence feature analysis results from multiple dimensions and combined with preset filtering conditions, the second clone collection or the second sequencing Fv sequence is filtered to obtain qualified antibody sequences.

[0015] In some embodiments, the sequence quality feature analysis includes:

[0016] Analyze the number and position of cysteine ​​residues in the second sequencing Fv sequence;

[0017] Analyze the sequence integrity of the second sequencing Fv sequence;

[0018] The first pseudo-likelihood value of each site is obtained by bit-by-bit masking prediction of the second sequencing Fv sequence using an antibody language model, and the AI ​​likelihood value of the second sequencing Fv sequence is calculated based on the first pseudo-likelihood value of all sites in the second sequencing Fv sequence.

[0019] In some embodiments, the feature analysis of sequence abundance includes:

[0020] The frequency of occurrence of DNA or protein sequences in the second sequencing Fv sequence is obtained by analyzing the number of occurrences.

[0021] The frequency of complementarity-determining region (CD) sequences in the second sequencing Fv sequence was analyzed to obtain the CD sequence enrichment.

[0022] In some embodiments, the characteristic analysis of the mutation degree includes:

[0023] Analyze the homology between the second sequencing Fv sequence and the homologous germline sequence;

[0024] The number of somatic hypermutations in the second sequencing Fv sequence was analyzed by comparing it with the corresponding germline sequence.

[0025] Analyze the ratio between the total branch length of the subtree corresponding to the target node in the first phylogenetic forest and the total branch length on the path from the target node to the root node.

[0026] In some embodiments, the mutation preference feature analysis includes:

[0027] Analyze the high-frequency somatic hypermutation frequency of the second clonal collection;

[0028] Analyze the distance between each second sequencing Fv sequence and the consensus sequence within the second clone set;

[0029] Analyze the local branch density corresponding to the target node in the first phylogenetic forest.

[0030] In some embodiments, the feature analysis of the degree of aggregation includes:

[0031] Analyze the connectivity of the complementarity-determining region sequences in the second sequencing Fv sequence;

[0032] The connectivity of the frame region sequences in the second sequencing Fv sequence was analyzed.

[0033] In some embodiments, the step of performing gene integrity filtering and structural integrity filtering on the first sequencing Fv sequence to obtain the second sequencing Fv sequence includes:

[0034] If the V gene fragment or J gene fragment in the first sequencing Fv sequence is not matched in the gene library, the corresponding first sequencing Fv sequence is removed and the remaining first sequencing Fv sequence is used as the third sequencing Fv sequence.

[0035] If the third sequencing Fv sequence has a missing gene fragment in the frame region or complementarity-determining region, or if the third sequencing Fv sequence is an incomplete sequence at both ends, then the corresponding third sequencing Fv sequence is removed and the remaining third sequencing Fv sequence is used as the second sequencing Fv sequence.

[0036] In some embodiments, the step of dividing all second sequencing Fv sequences into several first clonal set groups based on the germline Fv sequence, and reconstructing the root node sequence of each first clonal set to obtain a second clonal set, includes:

[0037] The second sequencing Fv sequences that have the same V gene fragment and J gene fragment in the germline Fv sequences corresponding to the heavy chain and light chain and have the same CDR3 sequence length are classified into the same first clonal set;

[0038] Align all the second sequencing Fv sequences and the germline Fv sequences in each of the first clone sets;

[0039] If the base types at any site in all aligned second sequencing Fv sequences and germline Fv sequences are different, then the bases at the corresponding sites in the germline Fv sequences are set to N bases.

[0040] The modified germline Fv sequence is used as the root node sequence of the first clonal set to obtain the second clonal set.

[0041] To achieve the above objectives, another aspect of this application proposes a device for analyzing the antibody sequence characteristics and specificity screening of a B-cell immune repertoire, the device comprising:

[0042] The first module is used to obtain antibody sequencing sequences from the B-cell immune repertoire;

[0043] The second module is used to perform phylogenetic alignment on the antibody sequencing sequences, identify the first sequencing Fv sequence corresponding to each antibody sequencing sequence, and generate the phylogenetic Fv sequence corresponding to each antibody sequencing sequence. The first sequencing Fv sequence includes a V gene fragment, a D gene fragment, a J gene fragment, or a C gene fragment.

[0044] The third module is used to perform gene integrity filtering and structural integrity filtering on the first sequencing Fv sequence to obtain the second sequencing Fv sequence.

[0045] The fourth module is used to divide all the second sequencing Fv sequences into several first clonal sets according to the germline Fv sequence, and to reconstruct the root node sequence of each first clonal set to obtain a second clonal set.

[0046] The fifth module is used to construct a first phylogenetic forest based on the same type category conversion probability and the second clone set. The first phylogenetic forest includes several first evolutionary trees, and each first evolutionary tree corresponds to a second clone set.

[0047] The sixth module is used to perform multi-dimensional sequence feature analysis based on the second sequencing Fv sequence, the second clonal set, or the first lineage forest. The multi-dimensional sequence feature analysis includes sequence quality feature analysis, sequence abundance feature analysis, mutation degree feature analysis, mutation preference feature analysis, and clustering degree feature analysis.

[0048] The seventh module is used to generate a visualization map based on the analysis results of sequence features from multiple dimensions or the first lineage forest.

[0049] The eighth module is used to filter the second clone collection or the second sequencing Fv sequence to obtain qualified antibody sequences based on the sequence feature analysis results of multiple dimensions and preset filtering conditions.

[0050] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0051] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.

[0052] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.

[0053] The embodiments of this application include at least the following beneficial effects: This application provides a method and related equipment for antibody sequence feature analysis and specific screening of B-cell immune repertoire. This method involves obtaining antibody sequencing sequences from the B-cell immune repertoire, performing germline alignment to identify the first sequencing Fv sequence corresponding to each antibody sequencing sequence, and generating a germline Fv sequence corresponding to each antibody sequencing sequence. Then, the first sequencing Fv sequences are filtered for gene integrity and structural integrity to obtain second sequencing Fv sequences. Based on the germline Fv sequences, all second sequencing Fv sequences are classified into several first clonal sets, and the root node sequence of each first clonal set is reconstructed. After reaching the second clonal set, the first lineage forest is constructed based on the isotype conversion probability and the second clonal set. Then, multi-dimensional sequence feature analysis is performed based on the second sequencing Fv sequence, the second clonal set, or the first lineage forest. This allows for analysis of antibody sequence characteristics from multiple dimensions, such as sequence quality, sequence abundance, mutation degree, mutation preference, and aggregation degree. A visualization map is then generated based on the results of the multi-dimensional sequence feature analysis or the first lineage forest, allowing relevant personnel to view the sequence characteristics. Furthermore, the second clonal set or the second sequencing Fv sequence is filtered based on the results of the multi-dimensional sequence feature analysis and preset filtering conditions, thereby effectively identifying high-affinity BCR sequences. Attached Figure Description

[0054] Figure 1 This is a schematic diagram of the B cell development process in existing technology;

[0055] Figure 2 This is a flowchart of a method for analyzing the sequence characteristics and specific screening of antibodies from a B-cell immune repertoire provided in an embodiment of this application;

[0056] Figure 3 This is a schematic diagram of the VDJC gene fragment structure of the antibody sequencing sequence provided in the embodiments of this application;

[0057] Figure 4 This is a schematic diagram of the antibody sequencing sequence provided in the embodiments of this application;

[0058] Figure 5 This is a schematic diagram of the heavy chain alignment results corresponding to the antibody sequencing sequence provided in the embodiments of this application;

[0059] Figure 6 This is a schematic diagram of the first sequencing Fv sequence alignment of the missing J gene fragment provided in the embodiments of this application;

[0060] Figure 7 This is a schematic diagram of the third sequencing Fv sequence with FWR1 missing, provided in an embodiment of this application;

[0061] Figure 8This is a schematic diagram of clone type division provided in the embodiments of this application;

[0062] Figure 9 This is a schematic diagram of residue inconsistencies provided in the embodiments of this application;

[0063] Figure 10 This is a partial example schematic diagram of the mutation frequency table provided in the embodiments of this application;

[0064] Figure 11 This is a schematic diagram of the feature analysis results and their corresponding preset filtering conditions provided in the embodiments of this application;

[0065] Figure 12 This is a representation of the data processing results of all antibody sequencing sequences in the B-cell immune repertoire provided in the embodiments of this application;

[0066] Figure 13 This is a schematic representation of the V gene family provided in the embodiments of this application;

[0067] Figure 14 This is a visualization diagram of VDJ rearrangement diversity provided in an embodiment of this application;

[0068] Figure 15 This is a distribution map of the number (protein) of somatic hypermutations provided in the embodiments of this application;

[0069] Figure 16 This is a distribution map of the number (DNA) of somatic hypermutations provided in the embodiments of this application;

[0070] Figure 17 This is an AI likelihood value distribution map provided in the embodiments of this application;

[0071] Figure 18 This is a homology distribution diagram provided in the embodiments of this application;

[0072] Figure 19 This is a schematic diagram of a tree in a first-lineage forest provided in an embodiment of this application;

[0073] Figure 20 This is a sequence aggregation connection network diagram provided in the embodiments of this application. Detailed Implementation

[0074] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0075] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0076] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0077] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0078] Before providing a detailed description of the embodiments of this application, some of the nouns and terms used in the embodiments of this application will be explained first. The nouns and terms used in the embodiments of this application shall be interpreted as follows:

[0079] BCR (B-cell receptor) is a molecule located on the surface of B cells that is responsible for specifically recognizing and binding antigens. It is essentially a membrane surface immunoglobulin.

[0080] B-cell clonal lineage forest is a data structure and visualization method that uses B-cell receptor (BCR) sequences as "genetic barcodes" to integrate, present, and infer the differentiation relationships of B-cell clones from the same or different times and tissues within a phylogenetic framework. It treats individual clones as "lineage trees" and organizes multiple clone trees into a "forest" based on phylogenetic and functional associations. This characterizes dynamic processes such as clonal expansion, high-frequency somatic mutations (shm), class-switch recombination (csr), and inter-tissue migration, and connects to antibody specificity and epitope recognition. This framework relies on the structural and diversity mechanisms of BCR loci (V(D)J rearrangements, n-region insertions, allele rejection, and allotype rejection), as well as the differentiation pathways and biomarker systems of B cells in bone marrow development and peripheral germinal center responses.

[0081] The Fv sequence (frequently mutated v-region) in B cells refers to the set of v-region sequences formed after somatic high-frequency mutations (shm) in the B cell receptor (BCR) / antibody variable region gene, exhibiting high-frequency nucleotide substitutions compared to the germline v-region sequence. It is primarily located in the v-regions (vh / vl) of the heavy and light chains, especially the specificity-determining CDR1 / CDR2 / CDR3. The Fv sequence is the direct molecular basis for the production of high-affinity antibodies during affinity maturation and also represents the genetic imprint of B cells after antigen selection in the germinal center.

[0082] The heavy chain (H chain) in the gene sequence has a larger molecular weight, approximately 55-75 kDa, containing 450-550 amino acid residues, and each chain contains 4-5 cyclic peptides formed by intrachain disulfide bonds. The light chain (L chain) in the gene sequence has a smaller molecular weight, approximately 24 kDa, containing 214 amino acid residues, and each chain contains 2 cyclic peptides formed by intrachain disulfide bonds.

[0083] CDRs are the core components of B-cell antigen receptors. Each antibody consists of two heavy chains and two light chains, and each chain contains three CDR regions (CDR1, CDR2, and CDR3). These regions determine the antibody's specificity and affinity through the diversity of their amino acid sequences and are the key sites for antibody-antigen binding.

[0084] CDR3 is an important component of the B-cell antigen receptor (BCR), and its gene sequence plays a central role in the immune response. VDJ cells together form the variable region of the BCR, and CDR3 is a part of the variable region.

[0085] Homologous sequences refer to gene or protein sequences among different species that share the same or similar ancestor. Their similarity is usually determined through sequence alignment analysis. Corresponding lineages in gene sequences refer to groups of species that share a common ancestor or evolutionary relationship, whose gene sequences have retained similarity during evolution.

[0086] A locus is a fixed location of a gene on a chromosome. It is used to identify the coordinates of a specific gene or DNA segment and to describe the spatial distribution and variation characteristics of genes on chromosomes.

[0087] In related technologies, such as Figure 1 As shown, B cells undergo developmental stages in the bone marrow, including hematopoietic stem cells → lymphoid progenitor cells → pre-B cells → immature B cells. Immature B cells differentiate into mature (naive) B cells through a peripheral (e.g., spleen) transition period. Mature B cells recognize antigens via membrane-bound B cell receptors (BCRs) and are activated via T cell-dependent or T cell-independent pathways after antigen binding. T-dependent responses often occur in germinal centers, accompanied by somatic hypermutation and class switching, ultimately leading to the differentiation of plasma cells that secrete large numbers of antibodies and memory B cells carrying BCRs. Plasma cells are responsible for secreting soluble antibodies to mediate neutralization, opsonization, and complement activation effects; memory B cells respond rapidly and produce high-affinity antibodies upon re-exposure to the same antigen.

[0088] B cell repertoire refers to the collection of B cell receptor (BCR) sequences or antibody variable region sequences from all B cells in an individual, reflecting the body's ability to recognize various antigens and the diversity of immune responses. With the rapid development of high-throughput sequencing technologies (such as single-cell paired sequencing), it is now possible to sequence B cell repertoires at extremely high resolution and depth, thereby obtaining massive amounts of antibody sequence data. This antibody sequence data contains rich information, including key immunological events such as B cell clonal expansion, maturation processes, and responses to antigens.

[0089] However, facing such massive antibody sequence data, effectively mining and analyzing valuable BCR sequences has become an urgent problem to be solved. For example, in current antibody discovery projects, identifying which BCR sequences have higher affinity for specific antigens remains difficult because directly determining antigen specificity from the sequence is still challenging. Due to the time-consuming and costly nature of these methods, the entire library is not tested (SPR, BLI, ELISA). The common practice is to select high-abundance sequences for validation, which may lead to the loss of some better BCR sequences. Even after investing significant resources to test hundreds of sequences, strong binding may not be found, often resulting in the loss of better BCR sequences. Furthermore, the limited experimental data makes it insufficient to train deep learning or machine learning models. Current bioinformatics analysis of B cell immune repositories focuses only on clonal diversity, V / J gene usage frequency, somatic hypermutation (SHM), and evolutionary analysis, failing to mine multidimensional and comprehensive characteristics of BCR sequence sets, and even less able to identify high-affinity BCR sequences.

[0090] In view of this, this application provides a method and related equipment for analyzing the characteristics and specificity of B cell immune repertoire antibody sequences. This method integrates five dimensions of characteristics, including sequence quality, mutation degree, sequence abundance, aggregation degree and mutation preference, to achieve panoramic characterization and high-specificity screening of antibody sequences. This can overcome the limitations of single-dimensional analysis and effectively identify BCR sequences with high affinity.

[0091] This application provides a method for analyzing the sequence characteristics and specificity screening of B-cell immune repertoire antibodies, relating to the field of data processing technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the B-cell immune repertoire antibody sequence characteristic analysis and specificity screening method, but is not limited to the above forms.

[0092] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0093] The embodiments of this application will be described in detail below with reference to the accompanying drawings:

[0094] Figure 2 This is an optional flowchart of a method for analyzing the sequence characteristics and specificity screening of B-cell immune repertoire antibodies provided in an embodiment of this application. Figure 2 The method may include, but is not limited to, steps S210 to S280:

[0095] Step S210: Obtain antibody sequencing sequences from the B-cell immune repertoire;

[0096] Step S220: Perform phylogenetic alignment on the antibody sequencing sequences to identify the first sequencing Fv sequence corresponding to each antibody sequencing sequence, and generate the phylogenetic Fv sequence corresponding to each antibody sequencing sequence. The first sequencing Fv sequence includes V gene fragment, D gene fragment, J gene fragment or C gene fragment.

[0097] Step S230: Perform gene integrity filtering and structural integrity filtering on the first sequencing Fv sequence to obtain the second sequencing Fv sequence;

[0098] Step S240: Based on the germline Fv sequence, all second sequencing Fv sequences are divided into several first clonal sets, and the root node sequence of each first clonal set is reconstructed to obtain the second clonal set.

[0099] Step S250: Construct a first lineage forest based on the same type category conversion probability and the second clone set. The first lineage forest includes several first evolutionary trees, and each first evolutionary tree corresponds to a second clone set.

[0100] Step S260: Perform multi-dimensional sequence feature analysis based on the second sequencing Fv sequence, the second clonal set, or the first lineage forest. The multi-dimensional sequence feature analysis includes sequence quality feature analysis, sequence abundance feature analysis, mutation degree feature analysis, mutation preference feature analysis, and clustering degree feature analysis.

[0101] Step S270: Generate a visualization map based on the results of sequence feature analysis in multiple dimensions or the first lineage forest;

[0102] Step S280: Based on the sequence feature analysis results from multiple dimensions and combined with preset filtering conditions, filter the second clone collection or the second sequencing Fv sequence to obtain a qualified antibody sequence.

[0103] It is understandable that in this embodiment, after obtaining the antibody sequencing sequence from the B-cell immune repertoire,

[0104] Germplasm comparison is performed on all antibody sequencing sequences to identify the sequencing Fv sequence corresponding to each receptor sequencing sequence as the first sequencing Fv sequence, and a germline Fv sequence corresponding to each antibody sequencing sequence is generated. Specifically, in this embodiment, during germline comparison, multiple antibody sequencing sequences can be compared with candidate Fv sequences in a reference gene library to identify the first sequencing Fv sequence corresponding to each receptor sequencing sequence. Figure 3 As shown, the first sequencing Fv sequence in this embodiment may include a V (Variable) gene fragment, a D (Diversity) gene fragment, a J (Joining) gene fragment, or a C (Constant) gene fragment. When comparing this embodiment with sequences in a gene library, professional immune sequence analysis tools (such as IgBlast, MiXCR, or Partis) can be used to align the input antibody sequencing sequence with the VDJC gene segment. This allows for rapid retrieval of the most matching homologous sequence from a rich gene library, in-depth analysis of VDJ gene rearrangement information, ensuring the precise location of the gene origin of each antibody sequencing sequence, and providing effective data support for subsequent analysis.

[0105] Understandably, VDJC gene alignment is a crucial bioinformatics analysis process. This technology uses specialized immune sequence analysis tools to trace the origins of gene fragments in the sequencing sequences of B cell or T cell receptors. Specifically, this method utilizes dedicated software such as IgBlast and MiXCR to systematically compare each sequencing sequence with a database of original immune gene fragments stored in the organism's genome. This allows for the precise identification of the origins of four key gene fragments in the sequence: the variable region V gene fragment, which determines antigen-binding specificity; the D gene fragment, mainly located in the heavy chain and increasing sequence diversity; the J gene fragment, connecting the V region and the constant region; and the constant region C gene fragment, which determines the immunoglobulin type. By analyzing the origins of these key gene fragments, the genetic composition of each immune receptor sequence can be fully resolved, providing a reliable gene annotation basis for subsequent clone classification, somatic high-frequency mutation analysis, and lymphocyte lineage tracing.

[0106] For example, Figure 4 When the nucleotide sequence of the antibody heavy chain shown is compared with the sequence in the gene library as the antibody sequencing sequence, the following can be obtained: Figure 5 The comparison results are shown. Among them, Figure 5 In the first line, query_1 is the first Fv sequence (containing a fragment of the VDJC gene region) after alignment with the input antibody sequencing sequence. The line below query_1 is the candidate Fv sequence from the gene library that best matches the antibody sequencing sequence. Figure 5 The sequence alignment shown indicates that the candidate Fv sequence consists of the gene sequences IGHV1-3*01, IGHD3-9*01, IGHJ6*03, and IGHHG2A*01. Since this candidate Fv sequence best matches the antibody sequencing sequence, it can be used as the germline Fv sequence for the antibody sequencing sequence.

[0107] It is understood that in this embodiment, after obtaining the first sequencing Fv sequence of the antibody sequencing sequence, gene integrity filtering and structural integrity filtering are performed on the first sequencing Fv sequence to obtain the second sequencing Fv sequence. Specifically, the integrity filtering process in this embodiment can employ two types of filtering methods to ensure that the sequences input to downstream processes are all complete Fv sequences. The first type of filtering method in this embodiment determines that if the V gene fragment or J gene fragment is not matched in the gene library, the corresponding first sequencing Fv sequence is removed, and the remaining first sequencing Fv sequence is used as the third sequencing Fv sequence. That is, the first sequencing Fv sequence is subjected to V / J gene integrity filtering to remove the first sequencing Fv sequences corresponding to sequences that cannot be successfully matched with the V gene fragment or J gene fragment in the gene library, ensuring that the obtained third sequencing Fv sequences are all complete and traceable BCR sequences. Figure 6The first sequencing Fv sequence (Query-1) shown is from... Figure 6 The sequence alignment shown indicates that the first sequencing Fv sequence lacks the J gene fragment; therefore, it needs to be removed. Figure 6 The first sequencing Fv sequence is shown.

[0108] It is understood that this embodiment, after completing gene integrity filtering, will also perform structural integrity filtering on the obtained third sequencing Fv sequence. Specifically, if this embodiment determines that the third sequencing Fv sequence has gene fragment deletions in the frame region (FWR) or complementarity-determining region (CDR), or that the third sequencing Fv sequence is incomplete at both ends, the corresponding third sequencing Fv sequence is removed, and the remaining third sequencing Fv sequence is used as the second sequencing Fv sequence. This ensures the structural integrity and functional relevance of the obtained second sequencing Fv sequence, effectively removing low-quality or abnormal sequences and improving the accuracy and reliability of subsequent analysis. Here, fragment deletion refers to any one of the three CDRs and four FWR regions being empty, and incomplete at both ends refers to the deletion of two or more consecutive residues at the 3' end or 5' end. For example, as... Figure 7 As shown, since the third sequencing Fv sequence lacks FWR1, it is necessary to remove the third sequencing Fv sequence.

[0109] It is understood that, in this embodiment, after completing the structural integrity filtering of the sequencing Fv sequences, the clones of all second sequencing Fv sequences are divided into several first clonal type sets based on the germline Fv sequences. Then, the root node sequence of each first clonal type set is reconstructed to obtain a second clonal type set, so that all sequences within each second clonal type set can be considered antibody sequencing sequences originating from the same B cell. Specifically, this embodiment can classify second sequencing Fv sequences in all second sequencing Fv sequences where the V gene fragment and J gene fragment corresponding to the heavy and light chains in the germline Fv sequences are the same and have the same CDR3 sequence length into the same first clonal type set; align all second sequencing Fv sequences and germline Fv sequences in each first clonal type set; if the base types at any site in all aligned second sequencing Fv sequences and germline Fv sequences are different, then set the base at the corresponding site in the germline Fv sequence to N bases; use the base-modified germline Fv sequence as the root node sequence of the first clonal type set to obtain the second clonal type set. For example, as shown... Figure 8 As shown, if the V gene fragment and J gene fragment are the same in the germline Fv sequences of the heavy chain and light chain of antibody A and antibody B, and the length of the CDR3 sequence is the same, then antibody A and antibody B are classified into the same clonal set.

[0110] It is understandable that in this embodiment, after completing sequence numbering and alignment—that is, numbering all second sequencing Fv sequences in the first clonal set and aligning them with the germline Fv sequences—the ancestral sequences (root node sequences in the phylogenetic tree) within each first clonal set are reconstructed: first, all sequences within the same clonal set are aligned (since the V and J gene fragments are identical and the CDR3 length is the same, direct alignment is sufficient); then, each site is processed according to the site processing rules to construct the ancestral sequence. For example, Figure 9 As shown, the site processing rule is that for any site, if different sequences have the same base type at that site, then the site corresponding to the ancestral sequence is set to that base; if they do not have the same base type, then the site corresponding to the ancestral sequence is set to an N base.

[0111] It is understood that, after obtaining the second clonal set, this embodiment constructs a first phylogenetic tree for each second clonal set based on the homotypic class transition probability corresponding to each set, so as to form a first phylogenetic forest from all the first phylogenetic trees. Specifically, this embodiment can construct the first phylogenetic tree for each set based on the maximum parsimony method. Taking the maximum parsimony method as an example, this embodiment can infer the ancestral node sequence corresponding to each offspring in each set based on the assumption of minimizing the number of mutation events, thereby forming a tree with the minimum mutation. In the process of inferring the ancestral node sequence corresponding to each offspring node, this embodiment can infer the type of the offspring ancestral node by calculating the class transition probability between each sequence within the same phylogenetic lineage. Among them, the types of offspring ancestral nodes include immunoglobulin M (IGHM), which is common in the initial immune response; immunoglobulin G (IGHG), which is important in adaptive immunity; immunoglobulin A (IGHA); and immunoglobulin E (IGHE).

[0112] It is understood that, in this embodiment, after obtaining the second sequencing Fv sequence, the second clonal set, or the first lineage forest, multi-dimensional sequence feature analysis is performed on the second sequencing Fv sequence, the second clonal set, or the first lineage forest. This multi-dimensional sequence feature analysis may include, but is not limited to, feature analysis of sequence quality, sequence abundance, mutation degree, mutation preference, and clustering degree.

[0113] Specifically, the sequence quality feature analysis process is a key indicator for measuring the reliability of sequence data. The analysis process in this embodiment includes, but is not limited to, analyzing the number and position of cysteine ​​residues in the second sequencing Fv sequence, analyzing the sequence integrity of the second sequencing Fv sequence, and using an antibody language model to perform bit-by-bit masking prediction of the second sequencing Fv sequence to obtain the first pseudo-likelihood value for each site, and calculating the AI ​​likelihood value of the second sequencing Fv sequence based on the first pseudo-likelihood values ​​of all sites in the second sequencing Fv sequence.

[0114] The number of cysteine ​​residues (Cys) in a single chain is typically two, located at positions 23 and 104 of the IMGT numbering, respectively. Some heavy chain variable regions (VH / VHH) with longer CDR3 values ​​may contain four Cys residues. In this embodiment, the processing involves determining whether the number of Cys in the VH count is two or four, whether the VL count is two, and then checking whether the Cys residues are located at positions 23 and 104 of the IMGT numbering. If these conditions are met, the sequence quality characteristic of the corresponding second sequencing Fv sequence is marked as True; otherwise, it is marked as False.

[0115] For the sequence integrity analysis process, this embodiment can assess the amino acid deletion status at the 5' and 3' ends by judging the integrity of the heavy chain variable region / light chain variable region (VH / VL) sequence. The number of missing amino acids should be less than 2. If the above conditions are met, the integrity feature of the corresponding second sequencing Fv sequence is marked as True; otherwise, it is marked as False.

[0116] In this embodiment, the AI ​​likelihood value can be obtained by averaging the first pseudo-likelihood value of the site predicted bit-by-bit by the AntiBERTy antibody language model, and then using this average as the AI ​​likelihood value for the entire second sequencing Fv sequence. The AI ​​likelihood value reflects the "reasonableness" or "naturalness" of the second sequencing Fv sequence in the antibody language model, i.e., the probability that the second sequencing Fv sequence is predicted as a true antibody sequence in the antibody language model. A high AI likelihood value indicates that the second sequencing Fv sequence is closer to the known antibody sequence pattern and may have a more stable structure and function; a low AI likelihood value may indicate that the second sequencing Fv sequence has abnormalities or too many mutations, affecting its function. Specifically, the first pseudo-likelihood value of the i-th site of the second sequencing Fv sequence S... The calculation process is as follows:

[0117] ;

[0118] In the formula, This means "masking the sequence after the i-th bit" (i.e., retaining all positional information except for the i-th bit); This represents the probability distribution of the LLM model output, i.e., the mask position predicted by the antibody language model is the original character. The probability of.

[0119] The arithmetic mean of the first pseudo-likelihood values ​​for all L sites is taken to obtain the AI ​​likelihood value for the entire second sequencing Fv sequence:

[0120] ;

[0121] In the formula, represents the AI ​​likelihood value of the second sequencing Fv sequence S.

[0122] In this embodiment, after analyzing the number and position of cysteine ​​residues, sequence integrity, and AI likelihood value in the second sequencing Fv sequence, all second sequencing Fv sequences can be filtered sequentially based on the characteristic analysis results of these sequence quality.

[0123] Specifically, the sequence abundance feature analysis process in this embodiment can reflect clonal amplification and response intensity, measuring which clones are selected and amplified in the immune response. The analysis process in this embodiment may include, but is not limited to, analyzing the occurrence frequency of DNA or protein sequences in the second sequencing Fv sequence to obtain the sequence occurrence frequency, and analyzing the occurrence frequency of complementarity-determining region (CDG) sequences in the second sequencing Fv sequence to obtain the CDG sequence enrichment.

[0124] Sequence frequency refers to the number of times a DNA or protein sequence appears in the second sequencing Fv sequence. High-frequency sequences may represent dominant clones. These clones may be actively responding to an immune response to a specific antigen, thus amplifying extensively in the sample. High-frequency sequences can also be due to deviations in experimental procedures, such as PCR amplification bias or uneven sequencing depth. Low-frequency sequences may represent rare clones, which may be in the early stages of the immune response or responding weakly to a specific antigen. When a sequence occurs at an extremely low frequency, it may be noise due to sequencing errors or low-level non-specific amplification.

[0125] Complementarity-determining region (CDR) enrichment refers to the frequency of occurrence of CDR3 / CDR sequences in DNA / protein sequences. High CDR enrichment typically indicates the presence of a dominant clone. CDR enrichment may reflect the strength of the immune response. If a particular CDR sequence occurs frequently in a sample, it may mean that the immune system is responding very strongly to that antigen. Low CDR enrichment may indicate an early stage of the immune response or may be related to background noise during sequencing.

[0126] In this embodiment, the second sequencing Fv sequence can be filtered sequentially based on the sequence occurrence frequency and the complementarity determination region sequence enrichment to obtain a sequencing sequence that meets the requirements.

[0127] Specifically, the mutation degree feature analysis process in this embodiment can reflect somatic hypermutation and affinity maturation, and is an important signal for determining whether germinal center response and selection pressure have occurred. The analysis process in this embodiment can analyze the homology between the second sequencing Fv sequence and the homologous germline sequence, compare the second sequencing Fv sequence with the corresponding germline sequence to analyze the number of somatic hypermutations in the second sequencing Fv sequence, and analyze the ratio between the total branch length of the subtree corresponding to the target node in the first phylogenetic forest and the total branch length on the path from the target node to the root node.

[0128] Homology (Identity) refers to the degree of similarity between a B cell receptor (BCR) sequence and a homologous germline sequence, reflecting the extent to which the second-sequencing Fv sequence has retained its original germline characteristics during evolution. A high Identity value indicates that the second-sequencing Fv sequence is highly similar to the germline sequence, possibly indicating an early developmental stage or the absence of significant somatic mutations; a low Identity value may suggest that the second-sequencing Fv sequence has undergone more mutations, thus achieving higher antigen specificity in the immune response. The calculation formula is as follows:

[0129] ;

[0130] Somatic hypermutation number (SHM) refers to the number of mutations introduced into the variable region (V region) sequence of B cells during maturation through somatic mutation mechanisms. These mutations help the B cell receptor (BCR) optimize antibody affinity, thereby binding antigens more effectively. High mutation counts typically indicate a strong antigen-driven and affinity-based maturation process. SHM is calculated by aligning the second sequencing Fv sequence to be analyzed with the corresponding germline sequence and counting the number of mutations within the aligned region.

[0131] The Branch Length Ratio (BLR) measures the ratio of the total branch length below a node to the total branch length above that node in a first-lineage forest. Specifically, BLR measures the relative position of a node by comparing the total branch length of its subtrees to the total branch length on the path from that node to the root node. The formula is as follows:

[0132] ;

[0133] In the formula, v is a node in the target evolutionary tree; descendants(v) are all descendant nodes of node v; ancestors(v) are all ancestor nodes on the path from node v to the root node; and branch_length(u) is the branch length of the target node u.

[0134] This embodiment can filter the second sequencing Fv sequences sequentially based on homology and the number of somatic hypermutations, and filter out the second clone collection that does not meet the requirements based on the ratio index.

[0135] Specifically, the mutation preference feature analysis process in this embodiment can reveal AID target sites, replacement patterns, and repair preferences, which can be used to infer mutation mechanisms and selection pressures (e.g., whether there is affinity selection targeting epitopes). This embodiment can analyze the high-frequency somatic hypermutation frequency of the second clonal set, the distance between each second sequencing Fv sequence and the consensus sequence within the second clonal set, and the local branch density corresponding to the target node in the first lineage forest.

[0136] High-frequency somatic hypermutation (SHM) refers to the most frequent mutations occurring within the same clone in a variable region sequence. These mutations are introduced through the SHM mechanism, retained during immune selection, and may play an important role in the immune response, reflecting the adaptive evolution of the immune system. Specifically, the calculation of high-frequency somatic hypermutation can be performed by counting the number of mutations at each site relative to the root node sequence at each site of all second sequencing Fv sequences using that second clone set, for each set that meets the criteria, and then calculating a "mutation frequency table". For each specific second sequencing Fv sequence, the mutation sites relative to the root node sequence are identified, and based on different regions (non-CDR3 and CDR3), according to... Figure 10 The mutation frequency table (column names are residue types, row names are site numbers) shows the average mutation frequency of each second sequencing Fv sequence, which is used as the mutation confidence level. A high mutation confidence level indicates that the second sequencing Fv sequence mainly contains high-frequency SHM sites. The high mutation confidence level in the CDR region may be a result of antigen selection pressure.

[0137] The consensus sequence within the second clone set is a "typical" or "average" sequence representing the second clone set. Based on the majority rule, the most frequent sequence at each position is taken as the consensus for that position, and these consensus sequences are combined to form a consensus sequence. Then, the Hamming distance / Jukes-Cantor evolutionary distance / edit distance between the second sequencing Fv sequence and the consensus sequence within the second clone set is calculated as the distance between the sequence and the consensus sequence (CDist).

[0138] For the j-th locus, iterate through all characters in the second sequencing Fv sequence at that locus, and select the character with the highest frequency as the consensus character Cj, using the following formula:

[0139] ;

[0140] In the formula, ∑ is the character set, and argmax means "take the character C with the largest value in freq(c,j)".

[0141] The frequency is calculated as follows:

[0142] ;

[0143] In the formula, This is an indicator function; it returns 1 if the condition is true, and 0 otherwise.

[0144] The consensus characters at each point are combined sequentially to obtain the consensus sequence. :

[0145] ;

[0146] Branch Length Index (BLI) is a local metric for measuring evolution, used to measure the length of local branches at a given node in a phylogenetic tree. Specifically, BLI measures the local branch density of a node by calculating the total branch length within a certain radius surrounding that node. The calculation formula is as follows:

[0147] BLI(v)=∑u∈local_area(v)branch_length(u);

[0148] In the formula, v is a node in the evolutionary tree; local_area(v) is all nodes in the local area centered on node v; branch_length(u) is the branch length of node u.

[0149] This embodiment can filter the second clonal set sequentially by high-frequency somatic hypermutation frequency, distance from consensus sequence, and local branch density, thereby ensuring that the second sequencing Fv sequence in the filtered clonal set meets the current requirements.

[0150] Specifically, the clustering characteristic analysis process in this embodiment can reflect clonal structure, clonal diversity, and intraclonal variation / diffusion processes, which helps identify amplified clones and series of common origin. This embodiment can analyze the connectivity of complementarity-determining region sequences in the second sequencing Fv sequence, as well as the connectivity of frame region sequences in the second sequencing Fv sequence.

[0151] Complementarity-determining region (CDR) connectivity represents the similarity between CDR sequences and is calculated by comparing the amino acid differences between them. CDRs are key segments in B-cell receptors or antibodies responsible for recognizing and binding antigens; therefore, their sequence diversity directly reflects the breadth of antigen recognition by the immune system. High CDR connectivity means that multiple CDRs are highly similar in amino acid sequence, potentially possessing similar antigen-binding properties, suggesting a convergent immune response to the same or similar antigens. Low CDR connectivity indicates significant CDR sequence differences, possibly targeting different antigens or occurring in an early stage with less somatic mutation and selection, and a lower degree of clonal expansion.

[0152] Frame region (FWR) sequence connectivity represents the similarity between FWR sequences, measured by calculating amino acid differences. FWRs are primarily responsible for maintaining antibody structural stability and also influence antibody affinity maturation to some extent. FWR connectivity reflects the similarity between different FWR sequences, revealing their potential structural and functional connections. High FWR connectivity indicates that multiple FWR sequences are highly similar in amino acid sequence and may have similar structural stability. This usually means that these sequences may originate from the same clonal type or have similar developmental origins. For example, during B cell maturation, some clonal types may retain high FWR sequence similarity, thus exhibiting high FWR connectivity values. Low FWR connectivity indicates significant differences between FWR sequences, possibly originating from different clonal types or having different developmental origins. This is more common in the early stages of the immune response, when B cell clones have not yet undergone sufficient somatic mutation and selection, resulting in high FWR sequence diversity but low amplification of each clone.

[0153] Specifically, for any unique value u in the second sequencing Fv sequence C i ∈U, the connectivity (u) of its complementarity determination region sequence i The calculation formula for ) is as follows:

[0154] ;

[0155] In the formula, Ⅱ(.) is an indicator function; if the condition within the parentheses is true, Ⅱ(.) = 1; otherwise, Ⅱ(.) = 0. Summation and statistics of u j In and u i The number of items with a distance ≤ 1.

[0156] In this embodiment, the second sequencing Fv sequence can be filtered sequentially by the connectivity of the complementary determinant region sequence and the framework region sequence to obtain a qualified antibody sequencing sequence.

[0157] It is understood that, in this embodiment, after performing feature analysis on the second sequencing Fv sequence, the second clonal set, or the first lineage forest in five dimensions—sequence quality, sequence abundance, mutation degree, mutation bias, and aggregation degree—qualified antibody sequences can be obtained by filtering the second clonal set or the second sequencing Fv sequence using preset filtering conditions. The feature analysis results and their corresponding preset filtering conditions are as follows: Figure 11 As shown.

[0158] It is understood that, after completing the feature analysis and lineage forest construction of all antibody sequencing sequences in the B-cell immune repertoire, this embodiment can generate a corresponding visualization diagram to facilitate information viewing by relevant personnel. For example, as shown... Figure 12 The table shown below illustrates the data processing results for all antibody sequencing sequences in the B-cell immune repertoire library. Figure 13 The V gene family statistics table shown is as follows: Figure 14 The VDJ rearrangement diversity visualization shown is as follows: Figure 15 The somatic cell hypermutation number (protein) distribution map shown is as follows: Figure 16 The somatic cell hypermutation number (DNA) distribution map shown is as follows: Figure 17 The AI ​​likelihood distribution plot shown is as follows: Figure 18 The homology distribution diagram shown is as follows: Figure 19 The diagram shown is of a tree in the first lineage forest. Figure 20 The diagram shows a sequence-aggregated connection network. Figure 19 In this system, each clone set corresponds to a tree, and each tree has an initial ancestor. The nodes in the tree marked with natural ordinal numbers are the inferred ancestors (such as 1, 2, 3, etc.), while other nodes named with ID as a prefix are the actual sequences obtained from sequencing. The darker the color, the lower the AI ​​likelihood value. The color of the node outline describes the isotype; the node size describes the abundance of the sequence; and the numerical value of the branches describes the distance between the node sequences.

[0159] It is understood that when the method of this embodiment is applied to the selection of antibody molecules from B-cell immune repertoires simultaneously with existing similar methods, the method of this embodiment has a better positive rate in terms of expression rate, binding rate, and proportion of high affinity.

[0160] As described above, the method in this embodiment constructs a comprehensive analysis approach based on "five-dimensional fusion," integrating five core dimensions: sequence quality, mutation degree, sequence abundance, aggregation degree, and mutation preference. This enables a leap from basic feature annotation to in-depth biological significance analysis. The method not only encompasses the core functions of existing tools such as sequence alignment, clonal typing, and evolutionary analysis, but also achieves a correlation path between "sequence features - immune mechanisms - functional specificity" through the synergistic verification and cross-interpretation of multi-dimensional features. This allows for the quantification of clonal amplification intensity and response level, precise differentiation between antigen-selective directed mutations and random mutations, and analysis of the molecular mechanisms of affinity maturation. Simultaneously, it can characterize clonal diversity and intraclonal variation diffusion patterns, enabling precise screening of functional antibody sequences with high affinity and strong specificity. This provides crucial technical support for upgrading immune repertoire research from "descriptive analysis" to "functional mining."

[0161] This application embodiment also provides a device for analyzing the sequence characteristics and specificity screening of B-cell immune repertoire antibodies, the device comprising:

[0162] The first module is used to obtain antibody sequencing sequences from the B-cell immune repertoire;

[0163] The second module is used to perform phylogenetic alignment of antibody sequencing sequences, identify the first sequencing Fv sequence corresponding to each antibody sequencing sequence, and generate the phylogenetic Fv sequence corresponding to each antibody sequencing sequence. The first sequencing Fv sequence includes V gene fragment, D gene fragment, J gene fragment or C gene fragment.

[0164] The third module is used to perform gene integrity filtering and structural integrity filtering on the first sequencing Fv sequence to obtain the second sequencing Fv sequence.

[0165] The fourth module is used to divide all the second sequencing Fv sequences into several first clonal sets based on the germline Fv sequence, and to reconstruct the root node sequence of each first clonal set to obtain a second clonal set.

[0166] The fifth module is used to construct a first lineage forest based on the same type category conversion probability and the second clone set. The first lineage forest includes several first evolutionary trees, and each first evolutionary tree corresponds to a second clone set.

[0167] The sixth module is used to perform multi-dimensional sequence feature analysis based on the second sequencing Fv sequence, the second clonal set, or the first lineage forest. The multi-dimensional sequence feature analysis includes sequence quality feature analysis, sequence abundance feature analysis, mutation degree feature analysis, mutation preference feature analysis, and clustering degree feature analysis.

[0168] The seventh module is used to generate visualizations based on the results of sequence feature analysis from multiple dimensions or the first lineage forest.

[0169] The eighth module is used to filter the second clone collection or the second sequencing Fv sequence to obtain qualified antibody sequences based on the sequence feature analysis results of multiple dimensions and preset filtering conditions.

[0170] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0171] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0172] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0173] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0174] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0175] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0176] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0177] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0178] This application provides a method and related equipment for analyzing the sequence characteristics and specific screening of B-cell immune repertoire antibodies. Through synergistic analysis and cross-validation of dimensional features, it achieves a comprehensive assessment of the functional value of antibody sequences. Specifically, this application can effectively improve the sequence quality dimension to ensure the reliability of analytical data, laying the foundation for subsequent feature interpretation; by jointly revealing the occurrence pattern of somatic hypermutation through the mutation degree and mutation preference dimensions, it can accurately distinguish between antigen-driven selective mutations and random mutations, clarifying the rationality of antibody affinity maturation levels and mutation mechanisms; by directly linking sequence abundance and aggregation degree dimensions to clonal amplification efficiency, it reflects the strength of B cell response to specific antigens, while simultaneously characterizing the evolutionary diffusion characteristics of clonal populations; through the deep integration of these five dimensions, it further achieves precise quantification of clonal diversity and systematic analysis of intraclonal variation patterns, ultimately constructing a complete technical chain of "feature characterization - mechanism interpretation - specific screening," efficiently screening antibody sequences with high affinity, strong specificity, and clear functional orientation, providing technical support for mechanism research of immune-related diseases, development of diagnostic biomarkers, and research and development of targeted antibody drugs.

[0179] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0180] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0181] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0182] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0183] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0184] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0185] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0186] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0187] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0188] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0189] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for analyzing the sequence characteristics and specific screening of antibodies from a B-cell immune repertoire, characterized in that, The method includes the following steps: Obtain antibody sequencing sequences from the B-cell immune repertoire; The antibody sequencing sequences are subjected to phylogenetic alignment to identify the first sequencing Fv sequence corresponding to each antibody sequencing sequence, and to generate the phylogenetic Fv sequence corresponding to each antibody sequencing sequence. The first sequencing Fv sequence includes a V gene fragment, a D gene fragment, a J gene fragment, or a C gene fragment. Gene integrity filtering and structural integrity filtering are performed on the first sequencing Fv sequence to obtain the second sequencing Fv sequence; Based on the germline Fv sequence, all second sequencing Fv sequences are divided into several first clonal type sets, and the root node sequence of each first clonal type set is reconstructed to obtain a second clonal type set. A first lineage forest is constructed based on the same type category conversion probability and the second clone set. The first lineage forest includes several first evolutionary trees, and each first evolutionary tree corresponds to a second clone set. Based on the second sequencing Fv sequence, the second clonal set, or the first lineage forest, a multi-dimensional sequence feature analysis is performed. The multi-dimensional sequence feature analysis includes sequence quality feature analysis, sequence abundance feature analysis, mutation degree feature analysis, mutation preference feature analysis, and clustering degree feature analysis. A visualization map is generated based on the results of sequence feature analysis from multiple dimensions or the first lineage forest. Based on the sequence feature analysis results from multiple dimensions and combined with preset filtering conditions, the second clone collection or the second sequencing Fv sequence is filtered to obtain qualified antibody sequences. The sequence quality feature analysis includes: Analyze the number and position of cysteine ​​residues in the second sequencing Fv sequence; Analyze the sequence integrity of the second sequencing Fv sequence; The first pseudo-likelihood value of each site is obtained by bit-by-bit masking prediction of the second sequencing Fv sequence using an antibody language model, and the AI ​​likelihood value of the second sequencing Fv sequence is calculated based on the first pseudo-likelihood value of all sites in the second sequencing Fv sequence.

2. The method according to claim 1, characterized in that, The feature analysis of the sequence abundance includes: The frequency of occurrence of DNA or protein sequences in the second sequencing Fv sequence is obtained by analyzing the number of occurrences. The frequency of complementarity-determining region (CD) sequences in the second sequencing Fv sequence was analyzed to obtain the CD sequence enrichment.

3. The method according to claim 2, characterized in that, The characteristic analysis of the mutation degree includes: Analyze the homology between the second sequencing Fv sequence and the homologous germline sequence; The number of somatic hypermutations in the second sequencing Fv sequence was analyzed by comparing it with the corresponding germline sequence. Analyze the ratio between the total branch length of the subtree corresponding to the target node in the first phylogenetic forest and the total branch length on the path from the target node to the root node.

4. The method according to claim 3, characterized in that, The mutation preference feature analysis includes: Analyze the high-frequency somatic hypermutation frequency of the second clonal collection; Analyze the distance between each second sequencing Fv sequence and the consensus sequence within the second clone set; Analyze the local branch density corresponding to the target node in the first phylogenetic forest.

5. The method according to claim 4, characterized in that, The feature analysis of the degree of aggregation includes: Analyze the connectivity of the complementarity-determining region sequences in the second sequencing Fv sequence; The connectivity of the frame region sequences in the second sequencing Fv sequence was analyzed.

6. The method according to claim 1, characterized in that, The process of performing gene integrity filtering and structural integrity filtering on the first sequencing Fv sequence to obtain the second sequencing Fv sequence includes: If the V gene fragment or J gene fragment in the first sequencing Fv sequence is not matched in the gene bank, the corresponding first sequencing Fv sequence is removed and the remaining first sequencing Fv sequence is used as the third sequencing Fv sequence. If the third sequencing Fv sequence has a missing gene fragment in the frame region or complementarity-determining region, or if the third sequencing Fv sequence is an incomplete sequence at both ends, then the corresponding third sequencing Fv sequence is removed and the remaining third sequencing Fv sequence is used as the second sequencing Fv sequence.

7. The method according to claim 1, characterized in that, The process involves classifying all second sequencing Fv sequences into several first clonal type sets based on the germline Fv sequence, and reconstructing the root node sequence for each first clonal type set to obtain a second clonal type set, including: The second sequencing Fv sequences that have the same V gene fragment and J gene fragment in the germline Fv sequences corresponding to the heavy chain and light chain and have the same CDR3 sequence length are classified into the same first clonal set; Align all the second sequencing Fv sequences and the germline Fv sequences in each of the first clone sets; If the base types at any site in all aligned second sequencing Fv sequences and germline Fv sequences are different, then the bases at the corresponding sites in the germline Fv sequences are set to N bases. The modified germline Fv sequence is used as the root node sequence of the first clonal set to obtain the second clonal set.

8. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of claims 1 to 7.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • End-to-end B cell clone pedigree forest construction method and related equipment

    CN121438931A

  • Identification of targets for immunotherapy in melanoma using splicing-derived neoantigens

    WO2025050009A2