A TCR repertoire framework for multiplex disease diagnosis

JP2024528441A5Pending Publication Date: 2025-06-25BOARD OF RGT THE UNIV OF TEXAS SYST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023578712
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-22
Filing Date
2022-06-17
Publication Date
2025-06-25

AI Technical Summary

Technical Problem

Current TCR clustering methods are unsuitable for analyzing large cohorts of TCR-seq samples due to high computational cost and inability to scale up, as they require quadratic equation calculations and fail to capture structural variations in TCR sequences with shared antigen specificity, leading to fragmented clusters.

Method used

A novel framework, GIANA, transforms CDR3 sequences into high-dimensional Euclidean space using isometric embedding and nearest neighbor search, enabling efficient clustering of large TCR datasets by encoding sequences into numerical vectors and applying machine learning techniques for antigen-specific TCR identification.

Benefits of technology

GIANA significantly reduces computational time and improves accuracy, allowing for the analysis of thousands of TCR repertoire samples, identifying disease-associated TCRs, and providing a non-invasive multi-disease diagnosis platform with high sensitivity and specificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A novel method of Geometric Isometry-Based Antigen-Specific TCR Alignment (GIANA) is described herein. GIANA is an antigen-specific TCR clustering method that can efficiently handle tens of millions of sequences. GIANA achieves higher sensitivity and accuracy than all existing methods and can search for TCRs specific to known antigens with high accuracy. Ultra-large-scale TCR clustering and fast querying of novel samples also enables a novel reference-based repertoire classification framework. GIANA can also analyze single-cell RNA-seq data with elucidated TCR regions, and can query unknown data against a large database of TCR repertoire samples in the public domain, providing new insights into shared antigen specificities. GIANA can also be applied to cluster or query large B cell receptor sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority under 35 USC Section 119(e) of U.S. Provisional Application No. 63 / 202,716, filed June 22, 2021. The disclosure of the prior application is deemed to be part of the disclosure of this application in its entirety and is incorporated herein by reference.

[0002] The present disclosure relates generally to immune repertoire-based disease diagnostic techniques, and more particularly to novel systems and methods for efficiently grouping similar T cell receptor (TCR) sequences, diagnosing patients with disease, and determining a patient's disease status by peripheral blood TCR repertoire. [Background technology]

[0003] The adaptive immune repertoire is a key regulator of diverse human diseases, and over 10,000 TCR repertoire sequencing (TCR-seq) samples have been generated in recent years. However, interpretation of TCR data has been hindered by the scarcity of known antigen specificities. Recent studies have demonstrated that similarities in the hypervariable complementarity determining region 3 (CDR3) of TCRs are responsible for structural similarity for antigen recognition. Thus, clustering of similar CDR3s has become an important method to identify antigen-specific receptors. Summary of the Invention [Means for solving the problem]

[0004] Methods, systems, and devices for improving the computational efficiency for the comparison of T cell receptors (TCRs) are described herein. In one example, the sequence of the complementarity determining region 3 (CDR3) can be identified from a reference TCR sequence (TCR-seq) dataset. The reference TCR-seq dataset can consist of TCRs specific to only one epitope. Each of the CDR3 sequences from the reference TCR-seq dataset can be encoded into a numerical vector, where the numerical vector corresponds to the sequence of amino acids in each of the CDR3 sequences. The numerical vector can be transformed into coordinates in a high-dimensional Euclidean space. A neural network can be used to generate a predictive model. The neural network can be trained to generate a tree data structure of the numerical vectors based on the relative distance of the coordinates, and group the coordinates into pre-clusters based on the relative distance. The CDR3 sequences in the pre-clusters can be filtered using one or more criteria to reduce noise. Antigen-specific CDR3 clusters can be identified from the filtered pre-clusters.

[0005] In another embodiment, unknown TCR-seq samples can be queried against existing reference data, and patients with disease can be diagnosed and their disease status determined by peripheral blood TCR repertoire. The above steps of identifying, encoding, converting, generating, and filtering can also be performed on a query TCR-seq dataset. The query TCR-seq dataset may not have known antigen-specific TCR information. The filtered pre-clusters from the query TCR-seq dataset can be compared with the antigen-specific CDR3 clusters. It can be determined that the filtered pre-clusters from the query TCR-seq dataset match the antigen-specific CDR3 clusters.

[0006] In another example, a large TCR database can be searched and grouped into TCR clusters of common antigen specificity. A nearest neighbor search using one or more TCR dissimilarity metrics can be performed to find pairs of TCRs with common antigen specificity. The one or more TCR dissimilarity metrics can include one or more of Smith-Waterman distance and embedding in a high-dimensional Euclidean space, or any other distance or dissimilarity metric.

[0007] The drawings described below are for illustration purposes only and are not intended to limit the scope of the present disclosure. [Brief description of the drawings]

[0008] [Figure 1] 1 is a diagram of a system according to some embodiments of the present disclosure. [Diagram 2] 1 is a block diagram illustrating components for implementing the methods described herein, according to some embodiments of the present disclosure. [Diagram 3] 1 is a flow chart illustrating GIANA analysis of reference TCR-seq data, according to some embodiments of the present disclosure. [Figure 4] 11 is a chart illustrating the performance of isometric embedding based on multidimensional scaling (MDS), according to some embodiments of the present disclosure. [Figure 5A] FIG. 1 shows a chart illustrating a comparison of equal length distances encoded by G6 for CDR3 strings with Smith-Waterman alignment scores according to some embodiments of the present disclosure. [Figure 5B] FIG. 1 shows a chart illustrating a comparison of equal length distances encoded by G6 for CDR3 strings with Smith-Waterman alignment scores according to some embodiments of the present disclosure. [Figure 5C]FIG. 1 shows a chart illustrating a comparison of equal length distances encoded by G6 for CDR3 strings with Smith-Waterman alignment scores according to some embodiments of the present disclosure. [Figure 5D] FIG. 1 shows a chart illustrating a comparison of equal length distances encoded by G6 for CDR3 strings with Smith-Waterman alignment scores according to some embodiments of the present disclosure. [Figure 5E] FIG. 1 shows a chart illustrating a comparison of equal length distances encoded by G6 for CDR3 strings with Smith-Waterman alignment scores according to some embodiments of the present disclosure. [Figure 5F] FIG. 1 shows a chart illustrating a comparison of equal length distances encoded by G6 for CDR3 strings with Smith-Waterman alignment scores according to some embodiments of the present disclosure. [Figure 6] FIG. 1 illustrates an overview of the Geometric Isometry-Based Antigen-Specific TCR Alignment (GIANA) workflow, according to some embodiments of the present disclosure. [Figure 7] FIG. 1 illustrates an overview of the GIANA workflow using stacked MDS vectors (GIANAsv), according to some embodiments of the present disclosure. [Figure 8] 1 is a chart showing a comparison of time complexity for various TCR clustering algorithms, according to some embodiments of the present disclosure. [Figure 9] 1 is a chart illustrating memory usage of various TCR clustering algorithms when evaluating time complexity, according to some embodiments of the present disclosure. [Figure 10] 1 is a chart illustrating clustering accuracy according to some embodiments of the present disclosure. [Figure 11] 1 is a chart illustrating the sensitivity of clustering according to some embodiments of the present disclosure. [Figure 12] 1 is a chart illustrating a comparison of normalized mutual information (NMI) between four methods of TCR clustering. [Figure 13] 1 is a chart illustrating precision-recall curves that measure the performance of GIANA over a range of parameter settings. [Figure 14] 1 is a chart illustrating precision-recall curves measuring the performance of GIANA with different permutation matrices. [Figure 15A] 1 is a chart illustrating the sensitivity and specificity of GIANA when applied to large and noisy TCR sequence (TCR-seq) samples, according to some embodiments of the present disclosure. [Figure 15B] 1 is a chart illustrating the sensitivity and specificity of GIANA when applied to large and noisy TCR sequence (TCR-seq) samples, according to some embodiments of the present disclosure. [Figure 15C] 1 is a chart illustrating the sensitivity and specificity of GIANA when applied to large and noisy TCR sequence (TCR-seq) samples, according to some embodiments of the present disclosure. [Figure 15D] 1 is a chart illustrating the sensitivity and specificity of GIANA when applied to large and noisy TCR sequence (TCR-seq) samples, according to some embodiments of the present disclosure. [Figure 15E] 1 is a chart illustrating the sensitivity and specificity of GIANA when applied to large and noisy TCR sequence (TCR-seq) samples, according to some embodiments of the present disclosure. [Figure 15F] 1 is a chart illustrating the sensitivity and specificity of GIANA when applied to large and noisy TCR sequence (TCR-seq) samples, according to some embodiments of the present disclosure. [Figure 16A] FIG. 1 shows estimates of sensitivity and specificity for GLIPH2, according to some embodiments of the present disclosure. [Figure 16B] FIG. 1 shows estimates of sensitivity and specificity for GLIPH2, according to some embodiments of the present disclosure. [Figure 16C] FIG. 1 shows an estimation of positive predictive value (PPV) for GLIPH2 and GIANA according to some embodiments of the present disclosure. [Figure 17]1 is a diagram illustrating a fast GIANA query based on isometric transformation according to some embodiments of the present disclosure. [Figure 18] 1 is a chart illustrating the time complexity evaluation of the GIANA query module using reference / query data with various numbers of TCRs according to some embodiments of the present disclosure. [Figure 19] 1 is a chart illustrating the extent to which query COVID-19 patients are separated from healthy controls by clustering against a reference dataset, according to some embodiments of the present disclosure. [Figure 20A] 1 is a chart illustrating a receiver operating characteristic (ROC) curve using COVID-19 fraction as a single predictor, according to some embodiments of the present disclosure. [Figure 20B] 1 is a chart illustrating a receiver operating characteristic (ROC) curve using COVID-19 fraction as a single predictor, according to some embodiments of the present disclosure. [Figure 20C] 1 is a chart illustrating a receiver operating characteristic (ROC) curve using COVID-19 fraction as a single predictor, according to some embodiments of the present disclosure. [Figure 21A] 1 is a chart illustrating the coefficient of variation of COVID-19 fraction with different numbers of reference TCRs, according to some embodiments of the present disclosure. [Figure 21B] 1 is a chart illustrating the coefficient of variation of COVID-19 fraction with different numbers of reference TCRs, according to some embodiments of the present disclosure. [Figure 21C] 1 is a chart illustrating the coefficient of variation of COVID-19 fraction with different numbers of reference TCRs, according to some embodiments of the present disclosure. [Figure 21D] 1 is a chart illustrating the coefficient of variation of COVID-19 fraction with different numbers of reference TCRs, according to some embodiments of the present disclosure. [Figure 22A] 1 is a graphical representation of TCR-seq sample similarity based on TCR co-clustering, according to some embodiments of the present disclosure. [Figure 22B]1 is a graphical representation of TCR-seq sample similarity based on TCR co-clustering, according to some embodiments of the present disclosure. [Figure 22C] 1 is a graphical representation of TCR-seq sample similarity based on TCR co-clustering, according to some embodiments of the present disclosure. [Figure 22D] 1 is a graphical representation of TCR-seq sample similarity based on TCR co-clustering, according to some embodiments of the present disclosure. [Figure 23A] 1 is a Bee Swarm plot illustrating the distribution of TCR clone frequencies in various categories, according to some embodiments of the present disclosure. [Figure 23B] 1 is a Bee Swarm plot illustrating the distribution of TCR clone frequencies in various categories, according to some embodiments of the present disclosure. [Figure 24A] 1 is a graph illustrating the dynamic changes in TCR clonal frequencies during the course of severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) infection, according to some embodiments of the present disclosure. [Figure 24B] 1 is a graph illustrating the dynamic changes in TCR clonal frequencies during the course of severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) infection, according to some embodiments of the present disclosure. [Figure 25A] 1 is a chart illustrating a ROC curve using a leave-one-out validation approach for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 25B] 1 is a chart illustrating a ROC curve using a leave-one-out validation approach for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 25C] 1 is a chart illustrating a ROC curve using a leave-one-out validation approach for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 25D]1 is a chart illustrating a ROC curve using a leave-one-out validation approach for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 25E] 1 is a chart illustrating a ROC curve using a leave-one-out validation approach for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 25F] 1 is a chart illustrating a ROC curve using a leave-one-out validation approach for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 26A] 1 is a chart illustrating a ROC curve using a more stringent method for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 26B] 1 is a chart illustrating a ROC curve using a more stringent method for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 26C] 1 is a chart illustrating a ROC curve using a more stringent method for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 26D] 1 is a chart illustrating a ROC curve using a more stringent method for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 26E] 1 is a chart illustrating a ROC curve using a more stringent method for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 26F] 1 is a chart illustrating a ROC curve using a more stringent method for disease fraction calculated from co-clustered TCRs, according to some embodiments of the present disclosure. [Figure 27] 1 is a chart illustrating cross-cohort similarity of reference TCR-seq samples, according to some embodiments of the present disclosure. [Figure 28A]1 is a violin plot illustrating the distribution of class fractions of cancer, COVID-19, multiple sclerosis (MS) patients, and healthy controls (HC), according to some embodiments of the present disclosure. [Figure 28B] 1 is a violin plot illustrating the distribution of class fractions of cancer, COVID-19, multiple sclerosis (MS) patients, and healthy controls (HC), according to some embodiments of the present disclosure. [Figure 28C] 1 is a violin plot illustrating the distribution of class fractions of cancer, COVID-19, multiple sclerosis (MS) patients, and healthy controls (HC), according to some embodiments of the present disclosure. [Figure 28D] 1 is a violin plot illustrating the distribution of class fractions of cancer, COVID-19, multiple sclerosis (MS) patients, and healthy controls (HC), according to some embodiments of the present disclosure. [Figure 29A] 1 is a chart illustrating a ROC curve using disease class fraction as a single predictor for pairwise separation of four disease classes, according to some embodiments of the present disclosure. [Figure 29B] 1 is a chart illustrating a ROC curve using disease class fraction as a single predictor for pairwise separation of four disease classes, according to some embodiments of the present disclosure. [Figure 29C] 1 is a chart illustrating a ROC curve using disease class fraction as a single predictor for pairwise separation of four disease classes, according to some embodiments of the present disclosure. [Figure 29D] 1 is a chart illustrating a ROC curve using disease class fraction as a single predictor for pairwise separation of four disease classes, according to some embodiments of the present disclosure. [Figure 29E] 1 is a chart illustrating a ROC curve using disease class fraction as a single predictor for pairwise separation of four disease classes, according to some embodiments of the present disclosure. [Figure 29F] 1 is a chart illustrating a ROC curve using disease class fraction as a single predictor for pairwise separation of four disease classes, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Many previous studies have applied TCR clustering to examine antigen-specific T cell responses during disease progression or immunotherapy treatment. It is speculated that integrating a large number of TCR-seq samples from many studies will provide more insight into immune-disease interactions and create new opportunities for prognosis and diagnosis. Nevertheless, high clustering specificity requires pairwise Smith-Waterman alignment of both CDR3 sequences and TCR variable gene (TRBV) alleles, which has quadratic computational complexity that usually cannot be scaled up to the scale of TCR repertoire samples (100K sequences or more). Motif-based clustering can achieve high speed but has very low specificity. Thus, none of the current TCR clustering methods are suitable for analyzing large cohorts of TCR-seq samples.

[0010] Unsupervised TCR clustering is a fundamental analysis of immune repertoire data. In an ideal scenario, all TCRs specific for the same epitope should be included in the same cluster. However, this is not feasible for sequence similarity or motif-based clustering approaches due to the expected diversity in TCR sequences with shared specificity. Such diversity is caused by the unique docking strategies of T cell receptors. For example, TCRs specific for influenza GIL epitopes usually contain classical RSS / RSA motifs in the CDR3 region, but related studies have reported that LGGW motifs also induce strong binding to GIL from different orientations. Such structural variations cannot be captured by simple Smith-Waterman alignment or motif grouping. As a result, CDR3s with dissimilar motifs end up being fragmented into small clusters despite sharing their specificity, which is a general limitation for current methods.

[0011] To address this challenge, we developed a novel framework that transforms CDR3 sequences and transforms the sequence alignment and clustering problem into a nearest neighbor search in a high-dimensional Euclidean space. This transformation significantly improves the computational efficiency of pairwise comparison of TCRs, achieving a 10 6 ~10 7 The novel method described herein can be scaled up to arrays.By pooling thousands of TCR repertoire samples that cannot be pooled by conventional systems and methods, the novel method described herein can identify novel disease-related TCRs.As described further herein, this may open new avenues for new multiplex disease diagnostic platforms.

[0012] The present disclosure will now be described in more detail hereinafter with reference to the attached drawings, which form a part hereof and which show, by way of non-limiting illustration, certain examples. However, it is intended that the subject matter may be described in a variety of different forms, and thus the subject matter encompassed or claimed should not be construed as being limited to any of the examples described herein. Among other things, the subject matter may be described as a method, device, component, or system. Thus, the examples may take the form of hardware, software, firmware, or any combination thereof (other than software per se). Thus, the following detailed description is not intended to be taken in a limiting sense.

[0013] Generally, terms may be understood, at least in part, from the usage in the context. For example, terms used herein, such as "and", "or", or "and / or", may include various meanings, depending at least in part on the context in which they are used. Typically, "or", when used in conjunction with a list such as A, B, or C, is intended to mean A, B, and C used in an inclusive sense, and A, B, or C used in an exclusive sense. Furthermore, the term "one or more", used at least in part depending on the context, may be used to describe any feature, structure, or characteristic in the singular sense, or to describe a combination of features, structures, or characteristics in the plural sense. Similarly, terms such as "a" or "the" may be understood to convey the singular or plural sense, depending at least in part on the context. Furthermore, it is understood that the term "based on" is not necessarily intended to convey an exclusive set of elements, and instead, the presence of additional elements not necessarily explicitly described may be acknowledged, also at least in part depending on the context.

[0014] The present disclosure is described below with reference to block diagrams and operational instructions of methods and devices. It is understood that each block of the block diagrams or operational instructions, and combinations of blocks in the block diagrams or operational instructions, can be implemented by analog or digital hardware and computer program instructions. These computer program instructions are provided to a processor of a general purpose computer, a special purpose computer, an ASIC, or other programmable data processing device to change its functions as detailed herein, so that the instructions executed via the processor of the computer or other programmable data processing device can perform the functions / operations specified in the block diagrams or operational blocks. In some alternative implementations, the functions / operations noted in the blocks may occur in a different order than that noted in the operational instructions. For example, two blocks shown in succession may be executed substantially simultaneously, or sometimes the blocks may be executed in the reverse order, depending on the functions / operations involved.

[0015] For purposes of this disclosure, a non-transitory computer-readable medium (or computer-readable storage medium) stores computer data, which may include computer program code (or computer-executable instructions) in machine-readable form that can be executed by a computer. By way of example, and not by way of limitation, a computer-readable medium may include a computer-readable storage medium for tangible or fixed storage of data, or a communication medium for the transitory interpretation of a signal containing the code. As used herein, computer-readable storage medium means physical or tangible storage (as opposed to a signal), and includes, without limitation, volatile and non-volatile, removable and non-removable media implemented in any method or technology for the tangible storage of information, such as computer-readable instructions, data structures, program modules, or other data. A computer readable storage medium includes, but is not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid-state memory technology, CD-ROM, DVD or other optical storage, cloud storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other physical or material medium that can be used to tangibly store desired information or data or instructions and that can be accessed by a computer or processor.

[0016] For purposes of this disclosure, the term "server" should be understood to mean a service point that provides processing, database, and communication facilities. By way of example, and not by way of limitation, the term "server" may mean a single physical processor with associated communication and data storage and database facilities, or it may mean a networked or clustered complex of processors and associated networks and storage devices, as well as operating software and one or more database systems and application software that facilitate the services provided by the server. A cloud server is an example.

[0017] For purposes of this disclosure, a "network" should be understood to mean a network that couples devices such that communications may be exchanged, such as between a server and a client device or other type of device, including between wireless devices coupled via a wireless network. A network may also include mass storage devices, such as network attached storage (NAS), storage area networks (SAN), content delivery networks (CDN), or other forms of computer or machine readable media. A network may include the Internet, one or more local area networks (LANs), one or more wide area networks (WANs), wireline type connections, wireless type connections, cellular, or any combination thereof. Similarly, sub-networks employing different architectures or conforming to or compatible with different protocols may interoperate within a larger network.

[0018] For purposes of this disclosure, a "wireless network" should be understood to couple client devices to a network. A wireless network may employ a standalone ad-hoc network, a mesh network, a wireless LAN (WLAN) network, a cellular network, and the like. A wireless network may further employ multiple network access technologies, including Wi-Fi, Long Term Evolution (LTE), WLAN, Wireless Router (WR) mesh, or second, third, fourth, or fifth generation (2G, 3G, 4G, or 5G) cellular technologies, Bluetooth, 802.11b / g / n, and the like. A network access technology may enable wide-area coverage for devices, such as client devices, with various degrees of mobility.

[0019] Briefly, a wireless network may include substantially any type of wireless communication mechanism by which signals are communicated between devices, such as client devices or computing devices, between or within a network, etc.

[0020] A computing device may be capable of transmitting or receiving signals, e.g., over a wired or wireless network, or may process or store signals, e.g., in a memory as a physical memory state, and thus may operate as a server, i.e., devices capable of operating as a server may include, e.g., dedicated rack-mounted servers, desktop computers, laptop computers, set-top boxes, integrated devices combining various features, such as features of two or more of the above devices, and the like.

[0021] Referring now to Figure 1, a system 100 is shown. Figure 1 illustrates components of a general environment in which the systems and methods discussed herein may be implemented. Not all components are required to practice the present disclosure, and variations in the arrangement and type of components may be made without departing from the spirit and scope of the present disclosure.

[0022] The system 100 of FIG. 1 includes a network 104, which, as discussed above, may include, but is not limited to, a wireless network, a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof.

[0023] Network 104 may couple with, for example, one or more client devices 102, application servers 106, content servers 108, and databases 107 and components thereof, along with other networks or devices. Network 104 may be configured as various wireless sub-networks, etc., further layered on a standalone ad-hoc network to provide infrastructure oriented coupling for one or more client devices 102, application servers 106, content servers 108, and databases 107. Network 104 may be configured to employ any form of computer-readable medium or network to communicate information from one electronic device to another.

[0024] The one or more client devices 102 may include, for example, a desktop computer or a portable device, such as a mobile phone, a smartphone, a display pager, a radio frequency (RF) device, an infrared (IR) device, a near field communication (NFC) device, a personal digital assistant (PDA), a handheld computer, a tablet computer, a phablet, a laptop computer, a set-top box, a wearable computer, a smart watch, an integrated or distributed device that combines various features, such as those of the above devices, and the like.

[0025] The one or more client devices 102 may include at least one client application configured to receive content from another computing device. The one or more client devices 102 may communicate with other devices or servers over the network 104, where such communication may include sending and / or receiving messages, generating and providing TCR data, searching, rendering, and / or sharing TCR data, or any of a variety of other forms of communication. The one or more client devices 102 may process or store signals in memory, for example as physical memory states, and thus may act as a server.

[0026] The application server 106 and the content server 108 may include one or more devices configured to provide and / or generate any type or form of content to another device over a network. Devices that may operate as the application server 106 and / or the content server 108 may include personal computers, desktop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, servers, and the like. The application server 106 and the content server 108 may store various types of data related to the content and services provided by the respective devices in a database 107.

[0027] Users (e.g., patients, doctors, technicians, etc.) can access services provided by application servers 106 and content servers 108, which may include, for example, application servers, authentication servers, search servers, and exchange servers over network 104 using one or more client devices 102. That is, for example, application servers 106 can store various types of applications as well as information related to the applications, including application data and user profile information.

[0028] 1 illustrates application server 106 and content server 108 as each being a single computing device, the disclosure is not so limited. For example, one or more functions of application server 106 and content server 108 may be distributed across one or more distinct computing devices. In another example, application server 106 and content server 108 may be integrated into a single computing device without departing from the scope of the disclosure.

[0029] Referring now to Figure 2, a block diagram illustrating components for implementing the methods described herein is shown. Figure 2 includes a TCR engine 200, a network 104, and a database 107. The TCR engine 200 may be a special purpose machine or processor and may be included in one or more of an application server 106, a content server 108, a web server, a third party server, a user's computing device, etc.

[0030] In one example, the TCR engine 200 may be a conventional personal computer, and the method described below may be implemented using a single thread on the CPU. In another example, when clustering reference data of 10 million sequences, the TCR engine 200 may be a high performance computing (HPC) supercluster (e.g., memory allocation of 128G, CPU node 8).

[0031] The TCR engine 200 may be a standalone application running on a device (e.g., a user device or a server / device connected to a system / web). In another example, the TCR engine 200 may function as an application installed on a device and / or a web-based application accessed by a device over a network. The TCR engine 200 may be installed as an augmented script, program, or application (e.g., a plug-in or extension) to another application, such as a health management application that aggregates and shares data related to a patient.

[0032] Database 107 may be any type of database or memory and may be associated with a server on a network (e.g., application server 106 and content server 108) or with a user's device (e.g., one or more client devices 102). Database 107 may contain data and metadata datasets associated with local and / or network information related to users, services, applications, content, and the like. Such information may be stored and indexed in database 107 independently and / or as linked or associated datasets. As discussed herein, it should be understood that the data (and metadata) in database 107 may be any type of information and type, known or hereafter known, without departing from the scope of this disclosure.

[0033] The database 107 can store data for users (e.g., user data), which can include, for example, information associated with the reference TCR-seq data, the patient's cancer diagnosis, the patient's chromosomal information, the patient's DNA information, the patient's blood information, the patient's demographic information, the patient's historical information, and the like, or some combination thereof.

[0034] The data (and metadata) in database 107 may be any type of information related to TCR-seq data, patients, physicians, content, devices, applications, service providers, and content providers, whether known or to become known, without departing from the scope of this disclosure.

[0035] Data stored in database 107 may be encrypted, for example using 256-bit encryption, so that the data is private and managed in accordance with the Health Insurance Portability and Accountability Act of 1996 (HIPPA).

[0036] The database 107 may store and index information as linked sets of data and metadata, and the relationships between data and metadata may be stored as n-dimensional vectors. Such storage may be accomplished through known or known vector or array storage, including but not limited to hash trees, queues, stacks, VLists, or any other type of known or known dynamic memory allocation method or technique. It should be understood that any known or known computational analysis method or algorithm, including but not limited to cluster analysis, data mining, Bayesian network analysis, hidden Markov models, artificial neural network analysis, logic models, and / or tree analysis, and others, may be applied to determine, derive, or otherwise identify vector information about the patient and / or healthcare provider.

[0037] As discussed above with reference to Figure 1, the network 104 may be any type of network, such as, but not limited to, a wireless network, a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. The network 104 may facilitate a connection between the TCR engine 200 and a database 107 of stored resources. Indeed, as shown in Figure 2, the TCR engine 200 and the database 107 may be directly coupled by any method known or to become known that couples and / or enables communication between such devices and resources.

[0038] A combination of a major processor, server, or device including hardware programmed according to the special purpose functions herein may be conveniently referred to as a TCR engine 200. The TCR engine 200 may include a sample module 202, an AI module 204, an encoding module 206, a filtering module 208, an identification (ID) module 210, and a conversion module 212. The engines and modules discussed herein are non-exhaustive, as additional or fewer engines and / or modules (or sub-modules) may be applicable to the example systems and methods discussed. The operation, configuration, and functionality of each module, as well as their role in the example of the present disclosure, are discussed below.

[0039] The principles described herein can be embodied in many different forms. Antigen-reactive T cells are central mediators of immunity to various diseases and important targets of immunotherapy, but experimental detection of cancer-associated T cells remains challenging because most cancer antigens are unknown. Recent developments in deep immune repertoire sequencing (TCR-seq) technology have further emphasized the identification of such T cells, as this may open new opportunities for non-invasive clinical diagnosis, prognosis and long-term immune monitoring of cancer patients. However, the human immune repertoire includes public T cells, naive T cells and memory / effector T cells specific for diverse antigens, and this complexity adds challenges (e.g., identifying cancer-associated T cells in TCR-seq data) that cannot be solved by conventional systems.

[0040] Previous studies on TCR repertoires in cancer patients have reported that simple statistics such as diversity and clonality are associated with clinical outcomes under certain conditions, demonstrating the utility of repertoire data as potential prognostic factors. However, with the rapid advances in immunotherapy and the rapid accumulation of TCR-seq data, more computational tools are needed to bridge the gap between basic immunogenomic research and clinical applications to benefit cancer patients.

[0041] The disclosed systems and methods provide these necessary tools through a novel framework that implements ensemble machine learning software (termed TCRBoost) that uses three-chain TCR-seq data to provide de novo predictions of cancer-associated immune repertoires.

[0042] Grouping of similar TCR sequences is related to shared antigen specificity and can be used to discover novel therapeutic targets. Conventional methods suffer from high computational cost and inability to scale up to the size of immune repertoire data sets. Geometric Isometry-Based Antigen-Specific TCR Alignment (GIANA) described herein can be used to close the gap between speed and prediction accuracy, at approximately 600 times the speed of conventional methods (e.g., TCRdist) and with better accuracy and sensitivity than conventional methods. GIANA also enables ultra-fast queries of large reference cohorts and can process over 100 billion sequence comparisons in less than 3 minutes. In one example, GIANA can perform 10 7 10 for the reference sequence 4 The TCRs of different TCRs can be compared. By applying GIANA to cluster large-scale TCR datasets, novel insights of disease-specific receptors can be revealed and new solutions for repertoire classification tasks can be provided. High accuracy is achieved by querying unknown TCR-seq samples against existing references with GIANA and can be used to identify cancer, infectious diseases, and autoimmune disorders. GIANA can be used as a non-invasive multiplex disease diagnostic platform based on TCRs.

[0043] Referring to Figure 3, a flow chart illustrating GIANA analysis of reference TCR-seq data is shown. Note that the steps shown in Figure 3 may be performed by the TCR engine 200 described above with reference to Figure 2.

[0044] At step 302, the sample module 202 can identify CDR3 sequences from a TCR-seq dataset. The sample module 202 can receive a TCR-seq dataset, for example, from the database 107. In one example, the TCR-seq dataset can include a reference TCR-seq dataset consisting of TCRs specific for only one epitope. At step 304, the encoding module 206 can encode each of the CDR3 sequences from the TCR-seq dataset into a numeric vector. The numeric vector can correspond to the sequence of amino acids in each of the CDR3 sequences.

[0045] In step 306, the conversion module 212 can convert the numeric vectors into coordinates in a high-dimensional Euclidean space. In step 308, the AI ​​module 204 can generate a predictive model using a neural network. The neural network can learn to generate a tree data structure of the numeric vectors based on the relative distance of the coordinates, and then group the coordinates into pre-clusters based on the relative distance. In step 310, the filtering module 208 can filter the CDR3 sequences in the pre-clusters. In step 312, the ID module 210 can identify antigen-specific CDR3 clusters from the filtered pre-clusters. The GIANA process is described in more detail below.

[0046] Now referring to FIG. 4, a chart illustrating the performance of isometric embedding based on multidimensional scaling (MDS) is shown. GIANA starts with an approximate solution to the isometric embedding of the BLOSUM62 matrix using MDS, which can generate a vector for each of the 20 amino acids that make up a protein. Each amino acid can be represented by a numerical vector. In one example, all 20 vectors may be calculated using a non-metric multidimensional scaling algorithm available in Python. The Euclidean distance between each pair of amino acids is calculated, resulting in a total of 190 pairs. All 190 distances (squared) may be compared visually with the corresponding scores in the BLOSUM62 matrix. The squared distances may be compared with the corresponding transformed BLOSUM62 dissimilarity scores (4-BLOSUM62 score, with diagonal values ​​taken as 0). For example, amino acids W and F may have a BLOSUM62 score of 1. Their distance from the calculation of isometric embedding vectors may be approximately 2.3. Thus, the point (2.3,1) may be displayed in the scatter plot shown in FIG. 4.

[0047] Spearman correlation can be calculated to evaluate the similarity between two measurements.The CDR3 strings can then be modeled as continuous, non-commutative linear transformations on the MDS vectors and expressed as coordinates in high-dimensional space.The unitary transformation matrix can be an element of any sufficiently large order cyclic group related to the typical length of CDR3 sequences, for example, the cyclic group of order 6 (G6), which can produce a nearly perfect linear correlation between the Euclidean distance of pairs of strings and their alignment scores.

[0048] Referring now to Figures 5A-5F, charts illustrating a comparison of the equal length distances encoded by G6 for CDR3 strings with Smith-Waterman alignment scores are shown. With a default equal length cutoff (-t) of 10, all TCR pairs with high Smith-Waterman alignment scores are included in the downstream cluster. Figures 5A-5F show an analysis for CDR3s with lengths of 12-17 amino acids, respectively. In each chart, the x-axis represents the equal length distance (e.g., squared Euclidean distance) between pairs of CDR3s, and the y-axis can represent the corresponding Smith-Waterman alignment scores using BLOSUM62 as the substitution matrix. The equal length distance may be defined as the squared Euclidean distance between pairs of numerical vectors after G6 encoding. Figures 5A-5F show an analysis of 10,000 CDR3 sequences divided into different length categories (i.e., 12-17 amino acids). Numerical vector representations for all sequences in each length category are obtained and pairwise distances are calculated. The sequence similarity of each pair of CDR3 sequences can be evaluated using the classical Smith-Waterman alignment algorithm, which is based on amino acid substitution matrices (e.g., BLOSUM62). High alignment scores indicate high sequence similarity. For two identical sequences, the maximum score is reached at 4*length. High alignment scores are associated with high similarity, which corresponds to a short distance, so Spearman's correlation value is shown as negative. This is different from the dissimilarity score used in Figure 4.

[0049] To identify CDR3 pre-clusters (i.e., TCRs with high similarity and assumed common antigen specificity) with high computational efficiency, a fast index-based nearest neighbor search and recursive centroid grouping may be performed on the coordinates. Nearest neighbor search methods may include one or more conventional methods, such as Facebook AI Similarity Search (FAISS), Navigable Small World (NSW), Hierarchical Navigable Small World (HNSW), PyNNDescent, and Annoy. The TCR dissimilarity measures used for nearest neighbor search may include one or more of Smith-Waterman distance, embedding in high-dimensional Euclidean space, or any other distance or dissimilarity metric used to estimate the common antigen specificity of two TCRs. The CDR3 pre-clusters may then be filtered for matched TRBV alleles and high Smith-Waterman alignment scores using a k-mer guided search table to produce the final TCR clusters as output.

[0050] Now referring to FIG. 6, a graphical description of the GIANA workflow is shown. In step 602, the GIANA process can begin by encoding short CDR3 peptide sequences into a numerical vector by a series of unitary transformations. As described in more detail below, the transformations can include elements of a cyclic group of order 6. In step 604, each encoded CDR3 sequence can be projected into a high-dimensional Euclidean space. In step 606, a fast nearest neighbor search can be performed. In step 608, an iterative centroid clustering can be performed. In step 610, a filtering step can be performed to match TRBV gene alleles and remove pairs with low alignment scores. In step 612, the final TCR clusters can be output.

[0051] Additionally or alternatively to the G6 conversion performed in step 602, a similar method (GIANAsv) may be used that uses stacked MDS vectors as the coordinates of the input CDR3 strings.

[0052] Now referring to FIG. 7, a graphical description of the GIANAsv workflow is shown. In step 702, the input sequence may be encoded as a concatenated vector of all amino acids in the string. Apart from encoding the input sequence, other processing steps may be similar to GIANA. For example, in step 704, each encoded CDR3 sequence may be projected into a high-dimensional Euclidean space and a fast nearest neighbor search may be performed. In step 706, an iterative centroid clustering may be performed. In step 708, a filtering step may be performed to match TRBV gene alleles and remove pairs with low alignment scores. In step 710, the final TCR clusters may be output. The GIANA and GIANAsv processes are described in further detail herein.

[0053] In one example, the GIANA process is used to identify and classify antigen-specific CDR3 sequences from reference TCR-seq data. TCR repertoire sequencing samples can be accessed from one or more databases, such as the immuneACCESS database provided by Adaptive Biotechnology, which is currently the largest database of TCR-seq samples, all profiled using the immunoSEQ platform. Antigen-specific TCRs and matched antigens can be pooled from the VDJdb, the Immune Epitope Database and Analysis Resource (IEDB), and previous literature. TCRs specific for more than one epitope can be excluded from the reference TCR-seq data to avoid mismatches.

[0054] Using a mathematical framework for isometric embedding of CDR3 sequences, we find a numerical representation (which is also a coordinate in a high-dimensional space) x of any short peptide sequence s, so that s i and s jFor two coordinates x i and x j The Euclidean distance between ||x i -x j may be made to correlate perfectly with sequence similarity scores as measured by a hypothetical evolutionary substitution matrix.

[0055] This problem may be referred to as "isometric embedding of short sequences". This concept is introduced to solve the problem of numerical encoding of CDR3 sequences, which typically have lengths in the range of 12-17 amino acids. A mathematical transformation of a given CDR3 sequence can be found that approximately satisfies isometricity. First, an approximately isometric embedding can be found for the BLOSUM62 matrix, as described below.

[0056] Amino Acid A i In the real space

[0057]

number

[0058] A solution to this problem exists if and only if the EDM is flat, and if the embedding space has dimension no greater than n, where n is the dimension of the EDM. Unfortunately, the BLOSUM62 matrix obeys the triangular rule: For i, j, k, M ik +M kj ≧M ij (Formula 2) It does not meet these criteria, so it is not EDM.

[0059] Therefore, an exact isometric embedding of BLOSUM62 may not exist. However, MDS can provide an approximate solution that applies when M is not EDM. MDS is the embedding vector β i In the case of classical MDS, the maximum dimension for the embedding space is 13. To explore dimensionality higher than 13, a non-metric MDS calculation using the skran package in Python may be applied. To maximize embedding isometry, 2,300 training TCRs of length 14 may be selected from the TCGA dataset and pairwise SW alignment scores may be calculated. MDS may be applied to obtain isometric embedding vectors of various dimensions ranging from 13 to 19. For each length, the Euclidean coordinates of the CDR3 sequences may be calculated as described in the GIANA method. The pairwise distances may be compared to the SW scores. The maximum score was observed for dimension 16, which would be the optimal dimension in the isometric representation. With this representation, the BLOSUM matrix: ||β i -β j ||≒M ij (Formula 3) A similarity of approximately 87% can be achieved.

[0060] A numerical encoding scheme may then be introduced, whereby each amino acid in the CDR3 sequence may be considered as an "operator", which is likened to a concept in quantum physics. In general, the operator A may be a mathematical transformation on the existing wave function Φ. The manipulation may be applied to the wave function denoted by the Dirac bracket A|Φ〉. One example is the angular moment operator L x , L y , L z The operator for amino acid i is A i This is defined as follows:

[0061]

number

[0062] where Ω is the matrix that needs to be determined. The operators are incompatible (if i ≠ j then A i A j ≠A j A i ), this definition emphasizes the ordering of the letters in the sequence. The CDR3 sequence may then be viewed as one or more successive linear operations on some initial vector β0. To simplify the calculations, we take β0=0. Thus, after the operation on the rightmost amino acid, the coordinates are

[0063]

number

[0064] Below we provide some examples that illustrate the desirable properties of Ω.

[0065] In the first example, the two amino acid sequences may be offset by one amino acid (e.g., a single mismatch). For example, the sequence s1=A k A i , and array s2=A k A j These numeric encoding vectors can be calculated as follows: s1=A k A i (Formula 6) s2=A k A j (Formula 7) x1=Ω×(Ω×β i +β k )(Formula 8) x2=Ω×(Ω×β j +β k )(Formula 9)

[0066] where x1 and x2 are the encoding vectors of s1 and s2. The Euclidean distance between s1 and s2 is ||x1-x2||=δ Tδ, where δ = x1-x2 (Eq. 10) =(Ω 2 (β i -β j )) T (Ω 2 (β i -β j ))(Equation 11) =(β i -β j ) T Ω T Ω T ΩΩ(β i -β j )(Equation 12) It may be calculated by:

[0067] Since this is the only mismatch, the top value is amino acid A. i and amino acid A j It may be equal to the distance between (β i -β j ) T Ω T Ω T ΩΩ(β i -β j )=(β i -β j ) T (β i -β j )≒M ij (Formula 13)

[0068] Without loss of generality, Ω T Ω=I. In other words, Ω can be a unitary matrix. It is easy to show that longer sequences with one mismatch follow the same pattern as above.

[0069] In a second example, two amino acid sequences may be off by two consecutive amino acids. The variable s1 is s1=A i A j and s2=A t A k The distance between the embedding vectors x1 and x2 is ||x1-x2||=(Ω 2 (β j -βk )+Ω(β i -β t )) T (Ω 2 (β j -β k )+Ω(β i -β t ))(Equation 14) =(β j -β k ) T (β j -β k )+(β i -β t ) T (β i -β t )-2(β i -β t ) T Ω(β j -β k )(Formula 15) ≒M jk +M it -2(β i -β t ) T Ω(β j -β k )(Equation 16) It is shown that.

[0070] Preferably, the third term may be zero for ∀i,j,t,k. One solution is to transform Ω into

[0071]

number

[0072]

number

[0073] where I may be an r-dimensional identity matrix, 0may be a zero matrix of dimension r. In fact, Ω defined in this way may be a representation of a cyclic group G2 of order 2. G2 may have only two elements, e and g, where g 2 = e. This notation can be useful in cases where there are many consecutive mismatches. Thus, β i teeth

[0074]

number

[0075]

number

[0076] where 0 can be a vector of zeros with dimension r. The new vector is

[0077]

number

[0078] In the third example, there may be multiple consecutive mismatches.

[0079]

number

[0080]

number

[0081] Another array,

[0082]

number

[0083]

number

[0084] x i and x j The distance is

[0085]

number

[0086] In an ideal scenario, all terms in the double Σ can be zero, and s i ands j The distance between

[0087]

number

[0088]

number

[0089] ∀u,v,k = 0. Or in general

[0090]

number

[0091]

number

[0092]

number

[0093] Both G2 and G3 may be normal subgroups of G6;

[0094]

number

[0095]

number

[0096] where

[0097]

number

[0098]

number

[0099] In this expression, the term in the double Σ can be 0 if uv≦6. If uv>6 (i.e., the two strings have more than 6 consecutive mismatches), the application of Ω6 as a transformation matrix can introduce undesirable variations in the final distance. Depending on the vectors on each side of the matrix, the addition can be positive or negative. However, when comparing CDR3 sequences with more than 6 mismatches, it is usually not important what the exact distance between them is. This is because only sequences with the greatest similarity are selected as antigen-specific TCR clusters, and the number of mismatches between two CDR3 sequences at the desired cutoff of the alignment score is usually less than 3.

[0100] In the fourth example, there can be many non-consecutive mismatches. Given two sequences of interest, i.e., s i and s j are different in the first and last positions,

[0101]

number

[0102]

number

[0103]

number

[0104] Those distances are

[0105]

number

[0106] this is,

[0107]

number

[0108] If we choose Ω as Ω6 as in the third example, the cross term can be zero as long as NIC is observed. However, if NIC is disturbed (i.e., if the two mismatches are exactly 6 amino acids apart), the cross term can be nonzero. This term can affect the final result. First, if the cross term remains negative (which has a probability of 1 / 2), the estimated equilength distance can be smaller than the exact value. This may not affect the result, since a strict Smith-Waterman alignment can be applied to ensure high sequence similarity. For a CDR3 with length 16 and the first 3 and last 2 amino acids truncated, assuming there are two mismatches, the probability of having two mismatches exactly 6 amino acids apart is:

[0109]

number

[0110] Thus, NIC perturbations may affect at most 0.091 / 2=4.6% of comparisons with two mismatches. When this occurs, some similar sequences may have large distances and be excluded from downstream clustering. To mitigate this effect, a relatively large default equal long distance cutoff (-t 10) may be applied to be inclusive. The current choice of parameter settings is a balance between clustering accuracy and computational speed.

[0111] Approximate isometric embedding of CDR3 sequences allows efficient search of their nearest neighbors (NN) in Euclidean space for fast clustering. One or more machine learning based classification techniques may be used to perform the NN search.

[0112] As will be appreciated by those skilled in the art, machine learning based classification techniques may vary depending on the desired implementation without departing from the disclosed technology. For example, machine learning classification schemes may utilize one or more of the following: hidden Markov models, recurrent neural networks, convolutional neural networks, Bayesian symbolic methods, general adversarial networks, support vector machines, image registration methods, applicable rule-based systems, alone or in combination. When regression algorithms are used, these may include, but are not limited to, stochastic gradient descent regressors, and / or passive-aggressive regressors, etc.

[0113] The machine learning classification model may be based on a clustering algorithm (e.g., a mini-batch K-means clustering algorithm), a recommendation algorithm (e.g., a mini-wise hashing algorithm or a Euclidean LSH algorithm), and / or anomaly detection algorithm, such as a local outlier factor. Additionally, the machine learning model may employ one or more dimensionality reduction approaches, such as a mini-batch dictionary learning algorithm, an incremental principal component analysis (PCA) algorithm, a latent Dirichlet allocation algorithm, and / or a mini-batch K-means algorithm, etc.

[0114] In one example, a Python package such as FAISS may be used to perform a fast indexed NN search.

[0115]

number

[0116] The coordinates (x) of the CDR3 may be divided into neighboring clusters. Prior to clustering, identical CDR3s may be grouped together. First, each unique sequence x i、 For i = 1, 2, . . . , N, its nearest neighbor x j , j = 1, 2, ..., N; j ≠ i can be located. i and x j Two points are considered to be centroids if the distance between them is within a user-defined cutoff (-t option, thr).

[0117]

number

[0118] K-mer guided rapid Smith-Waterman alignment with TCR variable gene matching can be performed on CDR3 pre-clusters. Although the CDR3s from pre-clusters can be highly similar, they do not necessarily qualify as antigen-specific groups. This is because 1) the sequences are not similar enough due to incomplete isometric embedding, and / or 2) TRBV gene information is not taken into account. Therefore, a filtering step can be performed to select antigen-specific CDR3 clusters based on Smith-Waterman alignment and TRBV gene matching.

[0119] The size of the precluster (m) is large, and traditional direct pairwise comparisons require quadratic time complexity O(m 2 ) to reduce the cluster size. TRBV information may be applied. Specifically, a precomputed matrix of alignment scores between pairs of TRBV alleles may be used. For each pair of CDR3 sequences in the precluster, its TRBV alleles may be compared. If the comparison score exceeds a user-defined score (-G option, thr_v), an edge may be added between the two sequences. A depth-first search (DFS) may be performed on the final graph to generate isolated subgraphs, where each subgraph becomes a new precluster. This step may split the original precluster into several smaller preclusters.

[0120] Next, a Smith-Waterman alignment may be performed using a k-mer approach. Each CDR3 sequence may be split into consecutive 5-mers. To store all sequences, a k-mer dictionary (e.g., in database 107) may be constructed with keys being unique 5-mers and values ​​being CDR3s containing a given 5-mer. One mismatch among the 5-mers may be tolerated when constructing the dictionary. For example, the sequence CASSGVTEAFF is indexed under both SSGVT and SSVAT. In this way, CDR3 sequences can be linked into a graph via shared k-mers. For each edge in this graph, a Smith-Waterman alignment may be performed with the BLOSUM62 substitution matrix and an alignment score may be calculated. If the score is below a user-defined cutoff (-S option, thr_s), the edge may be removed. The actual computational complexity of this step ranges from O(m) to O(m 2 ) The worst case scenario can be reached when every pair of CDR3s in the pre-clusters share a similar k-mer motif. DFS can be performed on the final graph to generate the final CDR3 clusters and report these as the final output.

[0121] In one example, new TCR-seq samples may be searched against the final CDR3 clusters of existing reference TCR-seq data. After generating the TCR clusters of the input data set, GIANA can perform a search (query) of further TCRs against this data (reference). In the query mode, GIANA can analyze one or more of the query file, the reference data, and the clustered reference data.

[0122] First, the reference and query TCR-seq data may be converted to isometric coordinates as described above. A fast nearest neighbor search (e.g., by FAISS) may then be performed, but limited to the query TCR-seq data. TCRs with distances shorter than a user-defined cutoff (-t option, thr) may be exported to a separate file (tmp_query.txt). This file may contain all TCRs that can possibly be clustered with the query sequence. GIANA clustering may be performed on this file to generate TCR clusters that satisfy the stringent cutoff for Smith-Waterman alignment. The query TCR clusters may then be merged with the reference clusters as follows: For each query cluster, if any sequence is derived from an existing cluster in the reference data, the two clusters may be merged. This step is to ensure the inclusion of all nearby TCRs in the reference data. There may be two or more conditions under which the query cluster does not contain any sequences in the reference cluster: 1) all TCRs in the query cluster were similar, but limited to the query sample, and / or 2) the query TCR was similar to some very rare reference TCRs that did not cluster with any other reference sample in the original clustering. According to either condition, the query cluster may be included in the final output.

[0123] In one example, the time cost of the query mode may be evaluated by generating reference data including 200K, 1M, 2M, 6M, and 10M TCRs. Different sizes of query data may be scanned including 10K, 20K, 30K, 40K, and 50K TCRs. Each query file may be clustered against each of the reference data, for example, using a general-purpose computer. The elapsed time may be estimated using Python's time module.

[0124] In the GIANAsv process, after MDS padding of 20 amino acids, the isometric representation of the CDR3 string: s = A1A2···A k, the easiest way to obtain k≧5 is to construct a “stack vector” (i.e., embedding vector β i , i=1, 2, . . . , k in the same order). The stacked vector representation can be

[0125]

number

[0126] The GIANA and GIANAsv processes described above provide several improvements over conventional TCR clustering methods (e.g., iSMART, GLIPH2, and TCRdist). For example, the GIANA and GIANAsv processes can handle large TCR data sets and provide more accurate data, as well as reducing the amount of computer resources required to generate these results. To demonstrate the improvements of the GIANA and GIANAsv methods described herein, a comparison using TCR repertoire sequencing data from healthy donors can be used. In the comparison, TCR clones can be ordered based on their abundance, and the top 10K, 20K,..., 100K sequences can be selected. All five methods can be applied to each of the subsamples. GIANA, GIANAsv, iSMART, and GLIPH2 can be run using default parameters. TCRdist does not provide clustering, and only pairwise distances can be calculated.

[0127] Referring now to Figures 8 and 9, charts illustrating the performance of GIANA and GIANAsv over conventional methods are shown. Figure 8 shows a comparison of time complexity for various TCR clustering algorithms. The chart shows the total number of TCR sequences analyzed (in increments of 10k) on the x-axis and the total computation time (seconds) on the y-axis. Line 802 shows the performance of TCRdist, line 804 shows the performance of iSMART, line 806 shows the performance of GLIPH2, line 808 shows the performance of GIANAsv, and line 810 shows the performance of GIANA. The speedup can be calculated based on the time cost for 100K TCR samples.

[0128] As shown in Figure 8, GIANA (line 810) has the lowest time cost throughout the benchmark, taking 23.9 seconds to process 100K sequences, while TCRdist (line 802) took 14,338 seconds. GIANAsv (line 808) is 2.2 times slower than GIANA (line 810). This is expected, since stacked vector encoding results in a high-dimensional isometric embedding space, increasing the time cost during nearest neighbor search. Notably, GLIPH2 (line 806) is the fastest algorithm after GIANA (line 810) and GIANAsv (line 808). This is because GLIPH2 avoids pairwise alignment through motif-guided search. Table 1 below shows a comparison of computation time and memory consumption for GIANA, GIANAsv, iSMART, TCRdist, and GLIPH2. In one example, the calculations may be performed on a system running macOS Catalina v10.15.2 with a 3.5GHz Dual-Core Intel Core i7 processor and 16GB of 2133MHz LPDDR3 memory.

[0129] [Table 1]

[0130] Figure 9 shows the memory usage of various TCR clustering algorithms when evaluating time complexity. The chart shows the total number of TCR sequences analyzed (in 10k increments) on the x-axis and the peak memory usage (in megabytes) on the y-axis. Line 902 shows the performance of TCRdist, line 904 shows the performance of iSMART, line 906 shows the performance of GLIPH2, line 908 shows the performance of GIANAsv, and line 910 shows the performance of GIANA.

[0131] GIANA and GIANAsv can also achieve higher accuracy than traditional methods in predicting antigen-specific TCRs. Antigen specificity may be the most desirable feature of TCR clustering. Analysis was performed using 61,366 non-redundant known TCR / antigen pairs from the public domain, covering over 900 different epitopes from diverse pathogens. Each method was performed, and the output cluster from each method was described as a "pure cluster" if all TCRs in the output cluster were specific to only one epitope. The purity of a cluster may be defined as the percentage of TCRs specific to the most common epitope in a given cluster. A "pure cluster" is defined as having a purity equal to 1.

[0132] Referring now to Figures 10 and 11, charts illustrating a comparison of clustering accuracy and sensitivity of GIANA compared to conventional methods are shown. Figure 10 shows clustering accuracy on the y-axis, which may be defined as the percentage of pure clusters in the output. GIANA, iSMART, TCRdist, and GLIPH2 are represented by bars 1002, 1004, 1006, and 1008, respectively. As shown by bar 1002, GIANA has the highest accuracy (93%) among all methods, while GLIPH2 has the lowest accuracy (35%). Figure 11 shows clustering sensitivity on the y-axis, which may be defined as the total number of TCRs in all pure clusters divided by the total number of TCRs tested. GIANA, iSMART, TCRdist, and GLIPH2 are represented by bars 1102, 1104, 1106, and 1108, respectively. As shown by bar 1102, GIANA also has the highest sensitivity (29%).

[0133] Now referring to Figure 12, a comparison of normalized mutual information (NMI) between the four methods is shown. NMI and epitope specificity between TCR clusters were measured using the same training data set. Similar NMI levels were observed across all methods, with GLIPH2 remaining the lowest.

[0134] Table 2 below shows the pure cluster sensitivity and clustering accuracy assessment for GIANA, iSMART, TCRdist, and GLIPH2. A total of 61,366 TCRs with known antigen specificity were used in this analysis. After excluding singleton TCRs (only one sequence per epitope), 60,700 remained.

[0135] [Table 2]

[0136] The fractions are similar for GIANA (96%), iSMART (97%), and TCRdist (97%), but significantly lower for GLIPH2 (36%). Retention of pure clusters may be defined as the total number of TCRs in all pure clusters divided by the number of all TCRs tested. GIANA also has a similar level of retention (27%) to the other methods, except for GLIPH2 (19%). For the three methods that rely on Smith-Waterman alignment (GIANA, iSMART, and TCRdist), the effect of a range of alignment score cutoffs (-S option in GIANA) was explored.

[0137] Referring now to FIG. 13, a chart illustrating the precision-recall curve measuring the performance of GIANA over a range of parameter settings is shown. The y-axis illustrates precision, which is defined as the fraction of truly positive calls among the total number of calls. The x-axis illustrates recall, which is the number of truly positive calls divided by the number of positive examples in the population. This analysis was applicable to GIANA, TCRdist, and iSMART since all three are based on SW alignments.

[0138] At precision above 0.95, all three methods share similar curves. The "elbow" shape of the curves is due to the use of pure cluster TCRs in the recall calculation. Lowering the cutoff may allow TCRs from different antigens to cluster, thereby reducing the fraction of pure clusters. A cutoff of 3.6 is preferred over 3.5, as it only slightly reduces recall (from 0.268 to 0.267) but increases precision by almost 3% (from 0.932 to 0.961). Therefore, the default parameter of the -S option in GIANA may be set to 3.6.

[0139] Referring now to FIG. 14, there is shown a chart illustrating precision-recall curves comparing the performance of GIANA, iSMART, TCRdist, with BLOSUM62 as the substitution matrix for mismatches, and GIANA with BLOSUM50 as the substitution matrix for mismatches. As in FIG. 13, the y-axis illustrates precision, which is defined as the fraction of truly positive calls among the total number of calls. The x-axis illustrates recall, which is the number of truly positive calls divided by the number of positive examples in the population.

[0140] The curve labeled "GIANA50" is the curve for GIANA with the BLOSUM50 matrix. This curve is very similar to the curve for the original version of GIANA with the BLOSUM62 matrix labeled "GIANA". This may be in part due to the similarity in their off-diagonal values ​​between the BLOSUM50 and BLOSUM62 matrices. The conversion to a distance matrix in GIANA eliminated the differences in diagonal values. Thus, the clustering accuracy of GIANA is relatively robust to the choice of protein substitution criteria, and the choice of whether to use the BLOSUM62 or BLOSUM50 matrix will not substantially affect the precision or recall of the final output.

[0141] It should be noted that TCRs specific for known epitopes were collected from the Immune Epitope Database and Analysis Resources (IEDB) and VDJdb online browser, among other sources. Only TCRβ CDR3 sequences, TRBV genes, and their associated antigens were preserved. After removing redundant or incomplete sequences, a total of 61,366 CDR3s were obtained, covering approximately 900 epitopes from diverse pathogens. All methods were applied to the dataset to perform antigen-specific clustering with their default parameters. For TCRdist, R code was created to perform a depth-first search for sequence pairs with a distance smaller than 15. The time-complexity calculation for TCRdist does not include the depth-first search to find TCR clusters. This cutoff of 15 has a balanced sensitivity and specificity comparable to that of iSMART4. Selection of a larger cutoff may increase the total number of clustered TCRs, but at the cost of reduced specificity of each cluster.

[0142] As above, a cluster with all TCRs specific for the same antigen is defined as a "pure cluster". Sensitivity is defined as the total number of TCRs contained in all pure clusters divided by the total number of sequences (i.e., 61,366). Clustering accuracy is defined as the number of pure clusters divided by the total number of clusters. These measures were used to compare the antigen-specific clustering performance of all four methods.

[0143] Moreover, unlike conventional methods, GIANA can search for antigen-specific TCRs from real, large and noisy TCR-seq samples using TCRs with known antigen specificity. From the above benchmark antigen-specific TCRs, we analyzed TCRs specific for three epitopes expected to be missing in healthy individuals, namely, the YAW and YLQ epitopes from the recent pandemic Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) virus, and the FRD epitope from Human Immunodeficiency Virus 1 (HIV-1). 20% of these TCRs were mixed with 100,000 TCRs from healthy donors as test data. The remaining 80% non-overlapping antigen-specific TCRs were used as training data to recover test sequences. Any sequences that cluster with the training data may be identified as "positive". The 20% spiked antigen-specific TCRs are true positives, and the false positives are TCRs from healthy donors.

[0144] Referring now to Figures 15A-15C, charts are shown illustrating the sensitivity and specificity of GIANA when applied to large and noisy TCR-seq samples. Figures 15A and 15B show the specificity and sensitivity of the YAW epitope from SARS-CoV-2. Figures 15C and 15D show the specificity and sensitivity of the YLC epitope from SARS-CoV-2. Figures 15E and 15F show the specificity and sensitivity of the FRD epitope from HIV-1. The violin plots shown in Figures 14A-14C illustrate the distribution of the data. The symmetrical curves on the sides of the "violin" are the actual probability density of the data points. The typical box plots in the center illustrate the mean (midpoint) and interquartile range of the data. The x-axis of each chart is the Smith-Waterman alignment score cutoff.

[0145] The violin plots shown in Figures 14A-14C illustrate the distribution of the data. The symmetric curves on the sides of the "violin" are the actual probability density of the data points. The typical box plot in the center illustrates the mean (midpoint) and interquartile range of the data.

[0146] The y-axis of each violin plot is the specificity or sensitivity. Specificity is defined as the number of true negatives divided by the total number of negatives called by the algorithm. Sensitivity is defined as the number of true positives divided by the total number of positives called by the algorithm. Positives and negatives called by the algorithm are defined using GIANA clusters. A sequence is called positive if it clusters with an added CDR3 with known antigen specificity.

[0147] The x-axis is the Smith-Waterman alignment score cutoff, an important parameter in GIANA, with a maximum of 4.0. The cutoff is an adjustable parameter in GIANA. For example, if the cutoff is set to 3.7, any sequence pairs with a Smith-Waterman alignment score (normalized by sequence length, thus a maximum of 4.0) greater than 3.7 will be clustered together. Sequence pairs with scores less than 3.7 will be separated. A higher cutoff results in higher specificity, but at the expense of reduced sensitivity. For all three epitopes, GIANA achieved a specificity of over 99.99% with a sensitivity of 20%-50%.

[0148] Referring now to Figures 16A-16C, charts illustrating the performance of GLIPH2 with large and noisy TCR-seq samples are shown. Figures 16A and 16B show the sensitivity and specificity estimates for GLIPH2 with the YAW epitope from SARS-CoV-2, the YLC epitope from SARS-CoV-2, and the FRD epitope from HIV-1 on the y-axis. Figure 16C shows the positive predictive value (PPV) estimates for the YAW epitope from SARS-CoV-2, the YLC epitope from SARS-CoV-2, and the FRD epitope from HIV-1 on the y-axis using GLIPH2 and GIANA. PPV may be defined as the total number of unique TCRs correctly predicted divided by the total number of unique TCRs clustered with the training data. Since GLIPH2 places one TCR among many clusters, a unique TCR may be required for this analysis. GLIPH2 reaches high sensitivity, but its specificity is lower than GIANA. More importantly, the PPV of GIANA reached over 60% for all epitopes, whereas the PPV of GLIPH2 was below 20% for two of the three epitopes.

[0149] It should be noted that an in silico mixing experiment was performed to evaluate the performance of GIANA in discovering TCRs specific for known antigens. Three antigens that are unlikely to be exposed to healthy donors were selected, namely, the YAW and YLQ epitopes from SARS-CoV-2, and the FRD epitope from the HIV-1 virus. TCRs specific for each epitope were selected by removing redundancies. For each antigen, 20% of the TCRs (test data) were randomly sampled and mixed with 100K sequences from healthy donors. There was no overlap between the remaining 80% of the antigen-specific TCRs (training data) and the test data. The mixed samples were considered to be pseudo-patients with the corresponding pathogen. The mixed samples were aligned with the training data and GIANA was applied with a Smith-Waterman alignment score cutoff (thr_s) ranging from 3.0 to 4.0 (increment of 0.1). For each epitope and parameter setting, 20 in silico mixings were performed to capture data variability.

[0150] From the obtained data, the predictive performance was evaluated. TCR clusters containing at least one TCR were selected from the training data. All TCRs in these clusters except the training data were positive calls. All TCRs that did not co-cluster with any of the training TCRs were negative calls. True positive calls were defined as the sequences labeled "test data", while true negative calls were sequences derived from the original 100K TCRs of healthy donors. Specificity was defined as the number of true negative calls divided by 100K. Sensitivity was defined as the number of true positive calls divided by the total number of tested TCRs.

[0151] Furthermore, unlike conventional methods, GIANA's high speed and specificity enables the query module to cluster new TCR samples with existing reference datasets, a feature lacking in all current tools.

[0152] Now referring to Figure 17, a diagram illustrating fast GIANA query based on isometric transformation is shown. As described above, the reference and query TCRs are transformed into Euclidean space with linear complexity, searched for nearest neighbors of each query sequence, processed into TCR clusters, and merged with the reference data. The dotted arrows shown in Figure 17 indicate the direction of search.

[0153] In step 1702, the reference TCR-seq data 1701 may be encoded into a numerical vector through a series of unitary transformations, and each encoded CDR3 sequence may be projected into a high-dimensional Euclidean space to form a reference isometric coordinate 1703. In step 1704, the query TCR-seq data 1715 may be encoded into a numerical vector through a series of unitary transformations, and each encoded CDR3 sequence may be projected into a high-dimensional Euclidean space to form a query isometric coordinate 1705. In step 1706, a nearest neighbor search may be performed between the reference isometric coordinate 1703 and the query isometric coordinate 1705. In step 1708, a minimum query cluster 1709 may be formed. In step 1710, a nearest neighbor search may be performed between the minimum query cluster 1709 and the reference cluster 1711. In step 1712, a merged cluster 1713 may be formed.

[0154] 18, a chart illustrating the time complexity evaluation of the GIANA query module using reference / query data with various numbers of TCRs is shown. The x-axis shows the number of query TCRs in 10k increments, and the y-axis shows the logarithmic representation of the computation time in seconds. Lines 1802, 1804, 1806, 1808, and 1810 show 200k reference TCRs, 1M reference TCRs, 2M reference TCRs, 6M reference TCRs, and 10M reference TCRs, respectively. As shown, GIANA is extremely efficient. It is a 10 7 10 for the reference sequence 4 It took almost 176 seconds to search for TCRs, a task with a computational load equivalent to 100 billion pairwise comparisons. Table 3 below shows the computational time consumption of GIANA queries of TCR samples of various sizes. Times are measured in seconds.

[0155] [Table 3]

[0156] This type of repertoire classification is a non-trivial task with rapid applications to disease diagnosis and prognosis. Typically, this task has been addressed by large-scale learning or deep learning. A common limitation of these methods is the large computational cost that prevents them from scaling up to large TCR-seq datasets.

[0157] GIANA queries can be used to classify TCR repertoires. For example, three reference datasets with 20, 100, or 200 TCR-seq samples can be evenly divided into COVID-19 patients and healthy controls (HC). Additionally, 154 COVID-19 and 120 HC samples can be searched against each of the references.

[0158] Referring now to FIG. 19, a chart illustrating the extent to which query COVID-19 patients are separated from healthy controls by clustering against a reference dataset is shown. The number of TCRs in the reference data is shown as labels on the x-axis. The t-statistics on the y-axis were performed using the t-test function to perform a two-sample t-test using the COVID-19 fraction to separate COVID-19 and HC query samples. Bars 1902, 1904, and 1906 represent 200k reference TCRs, 1M reference TCRs, and 2M reference TCRs, respectively. The two-sample t-test was performed using the COVID-19 fraction estimated from the query data to obtain the t-statistics. All p-values ​​are 2.2×10 -16 was significant at the level of . For each query sample, the fraction of TCRs co-clustered with COVID-19 reference patients can be calculated. This fraction is significantly higher for COVID-19 patients in the query sample, and separation from the query HCs is likely to increase with increasing size of the reference data. Using this fraction as a predictor, an increase in the area under the receiver operating curve (AUC) is observed for the reference with large sample sizes.

[0159] 20A-20C, charts illustrating Receiver Operating Characteristic (ROC) curves using COVID-19 fraction as the single predictor are shown. The number of COVID-19 and HC samples is shown above each chart. Each sample contains 10K TCR sequences.

[0160] ROC curves are an unbiased way to visualize the predictive ability of a given method. Here we use the COVID-19 fraction as a continuous predictor. By changing the threshold of this fraction, the specificity (x-axis) and sensitivity (y-axis) change. Figure 20A shows 10 HC samples and 10 COVID-19 samples. Figure 20B shows 50 HC samples and 50 COVID-19 samples. Figure 20C shows 100 HC samples and 100 COVID-19 samples. 95% confidence intervals were estimated from stratified 2,000 bootstraps. Notably, with 2 million reference TCRs, the sensitivity (79%) and specificity (100%) of this approach outperformed several existing tests for COVID-19, suggesting the potential utility of this approach in diagnosing the disease. More importantly, the accuracy of the repertoire classification improved with more reference samples. This is likely due to the sharing of disease-specific TCRs across different patients, which are usually shared at a low frequency, and therefore the large reference data results in high clustering probability, small variance, and good accuracy.

[0161] Referring now to Figures 21A-21D, charts illustrating the original COVID-19 fraction scores and the coefficient of variation of COVID-19 fractions with various numbers of reference TCRs for both COVID-19 patients and healthy donors are shown. Figures 21A-21 show the distribution of TCR fractions co-clustered with COVID-19 reference samples under various reference data configurations. Figure 21A shows 10 HC samples and 10 COVID-19 samples. Figure 21B shows 50 HC samples and 50 COVID-19 samples. Figure 21C shows 100 HC samples and 100 COVID-19 samples.

[0162] In Figure 21D, the x-axis shows the number of reference TCRs and the y-axis shows the coefficient of variation. The coefficient of variation may be defined as the standard deviation divided by the average COVID-19 fraction of COVID-19 patients in the query sample. Bars 2102, 2104, and 2106 show 200k, 1M, and 2M reference TCRs, respectively. With more reference samples, a decrease in the coefficient of variation of the COVID-19 fraction is observed.

[0163] The above features can also be demonstrated in larger reference datasets. For example, we used a dataset containing 10 million TCRs from 1,213 samples from cancer, COVID-19, multiple sclerosis (MS) patients, and HCs. We performed antigen-specific clustering of all 10M TCRs and applied GIANA to examine the similarity of different repertoire samples as measured by the level of shared TCR clusters. Table 4 shows the TCR-seq sample cohort used as reference data.

[0164] [Table 4]

[0165] For some cohorts, not all available samples were used in creating the reference data. For each sample, the top 10,000 most abundant TCRs were selected, and if the data contained fewer than 10,000 sequences, all were used. Unique samples indicated the number of independent patients who participated in the study. Sample size recorded the number of all TCR-seq samples in the cohort used in the reference. The Emerson 2017 cohort included 666 healthy donors in batch 1, from which 100 samples were randomly selected. The COVID-19 cohort included over 1,400 patients assembled from multiple international COVID-19 studies. Two cohorts were collected by Adaptive Biotechnology (Adaptive, n=154) and the Institute for System Biology (ISB, n=157), respectively. In one example, GIANA took 19.5 hours to cluster the reference data on a high-performance computing cluster with 8 CPUs and 128G memory.

[0166] Referring now to Figures 22A-22D, we show a graphical representation of the similarity of TCR-seq samples based on TCR co-clustering. In Figures 22A-22D, physical proximity represents similarity, and thus TCR-seq samples that are close to each other (shown as dots) are more similar. To generate Figures 22A-22D, a count sharing matrix per sample was calculated from the original TCR clustering results of the 1,213 reference samples described above. A Spearman correlation matrix was also calculated based on the counts of co-clustered TCRs, setting pairs with a correlation value of 0.4 or less to zero. The resulting sparse matrix was used to create a graph. Sample groups were visualized by excluding nodes with less than 2 connections.

[0167] Figure 22A shows an overview of the first cluster of cancer patients (detailed in Figure 22B), the second cluster of HC and MS patients (detailed in Figure 22C), and the third cluster of lung cancer and COVID-19 patients (detailed in Figure 22D). Figure 22A shows that most cancer patients in the first cluster are clearly separated from HC and MS patients in the second cluster. Interestingly, lung cancer and COVID-19 patients formed a separate third cluster. It is known that local inflammatory conditions such as viral infection or cancer release tissue-resident T cells into the circulation, possibly inducing sharing of TCR repertoires. These findings further suggest that the scale of T cell migration in lung tissue is large enough to transcend disease types.

[0168] Note that to test the feasibility of repertoire classification using TCR clustering, 10, 50, and 100 COVID-19 samples were combined with 10, 50, and 100 healthy controls to generate three reference datasets containing 20, 100, and 200 samples. Each sample contained 10K TCRs selected by ranking the abundance of clones. The query samples included 154 COVID-19 patients and 120 HCs. There was no overlap between the query and reference samples. TCR clusters were generated for each query sample using GIANA. For each sample, TCR clusters with more than 100 samples were excluded because these TCRs were generated from small-world connections and may not be informative about disease specificity. For the remaining clusters, the fraction of reference TCRs with COVID-19 patient contributions was calculated and used as a predictor.

[0169] In the task of multiple disease classification, 712 cancer, 311 COVID-19, 25 MS patient samples, and 100 HC samples were collected together to generate 10M TCR reference data. Another 62 cancer, 193 COVID-19, 12 MS, and 153 HC samples were collected and queries were created assuming the disease labels were unknown. Similar analyses were performed on each query cluster file to estimate the fraction of each disease category, including HC. These fractions were used to predict diseases and perform ROC analysis. Specifically, HC fractions were used for all comparisons with HC samples. As a preliminary approach, the difference between the two disease fractions was used for pairwise separation of the three diseases. For example, when predicting cancer from MS patients, the cancer fraction-MS fraction was used as a predictor.

[0170] Moreover, unlike conventional methods, GIANA can be used as a novel multi-disease detection platform with ultra-large-scale TCR clustering and discovery requirements. Ultra-large-scale clustering by GIANA allows for the examination of disease-specific and tissue-specific TCRs. In one example, TCR clusters from patients with lung cancer and COVID-19 were divided into three categories: i) COVID-19 specific, ii) lung cancer specific, and iii) shared by the two diseases.

[0171] Now referring to Figures 23A-23B, we show BeeSwarm plots illustrating the distribution of TCR clone frequencies for various categories. BeeSwarm plots are box plots showing the actual data points that are TCR clone frequencies (y-axis). TCR clone frequency is defined as the percentage of sequencing reads for a given TCR divided by the total number of reads in the sample. The x-axis of both charts is a side-by-side comparison of the two sample categories.

[0172] Figure 23A shows a comparison between TCR 2302, found only in COVID-19 patients, and TCR 2304, found in both COVID-19 and lung cancer patients. Figure 23B shows a comparison between TCR 2306, found only in lung cancer patients, and TCR 2304, found in both COVID-19 and lung cancer patients. For TCR 2304, found in both COVID-19 and lung cancer patients, the clonal frequency was selected to match the cohort of disease-specific TCRs. The p-value in Figure 23A (p<2.2e-16) is a measure of the type I error of the statistical test. The p-value in Figure 23B (ns) means not significant, which is a p-value greater than 0.05.

[0173] As shown in Figures 23A-23B, the clonal frequency of category i) was significantly higher than that of category iii) for COVID-19 patients, whereas there was no difference between categories ii) and iii) for lung cancer patients. TCR frequencies were matched within the same cohort, avoiding batch effects. Thus, the high abundance of COVID-19-specific TCRs is likely driven by the immune response to SARS-CoV-2. Indeed, only COVID-19-specific TCRs underwent dynamic regulation following viral infection, which peaked within the first 2 weeks after exposure and declined thereafter.

[0174] Referring now to Figures 24A-24B, graphs illustrating the dynamic changes in TCR clonal frequency during the course of SARS-CoV-2 infection are shown. In both charts, the x-axis shows the number of days from diagnosis to sample collection, and the y-axis shows the logarithmic representation of TCR frequency. The central solid line 2402 is the smoothed average of the data. At each time point (x-axis), there are multiple observations of TCR clonal frequency, and the solid line is the average value. Similarly, the upper dotted line 2404 is the upper limit of the 95% confidence interval for the observed data. The lower dotted line 2406 is the lower limit of the confidence interval. The "p-value" is the type I error of the Spearman correlation test performed for this analysis. The Spearman correlation value (rho) is displayed as the first line in each figure.

[0175] The clonal abundance of shared TCRs was not affected by the timeline following SARS-CoV-2 infection. These figures show that clustering on large TCR repertoire samples reveals a high abundance of shared disease-specific TCRs, which may provide a finer resolution for repertoire classification.

[0176] Clustered TCRs may be used as markers to assign repertoire samples to multiple diseases, for example by performing a leave-one-out validation approach. Specifically, for a given sample, the fraction of TCRs co-clustered with cancer, COVID-19, MS patients, or healthy controls in a reference cohort may be calculated, excluding the sample itself. This method results in four class fractions that sum to 1 for each sample. The HC fraction may be used to separate patients from healthy donors.

[0177] Referring now to Figures 25A-25F, charts illustrating ROC curves using a leave-one-out validation approach for disease fractions calculated from co-clustered TCRs are shown. AUC values ​​are shown at the bottom right of each chart. Each chart shows specificity on the x-axis and sensitivity on the y-axis. Figure 25A shows a comparison between cancer TCR clusters and HC TCR clusters. Figure 25B shows a comparison between COVID-19 TCR clusters and HC clusters. Figure 25C shows a comparison between MS clusters and HC clusters. Figure 25D shows a comparison between cancer TCR clusters and COVID-19 clusters. Figure 25E shows a comparison between COVID-19 clusters and MS clusters. Figure 25F shows a comparison between cancer clusters and MS clusters. 95% confidence intervals may be calculated using stratified 2,000 bootstraps. Near perfect accuracy was observed for all three diseases. To discriminate pairs of diseases, the difference between the two corresponding fractions may be used as a predictor, which results in high (>93%) AUC values.

[0178] Referring now to Figures 26A-26F, charts illustrating ROC curves using a more stringent approach for disease fractions calculated from co-clustered TCRs are shown. In one example, 40% of the reference samples were randomly selected as training data, and the remaining 60% were left as test data. The training samples are labeled "COVID-19", "Cancer", "MS", or "HC". Each test data was co-clustered with all training samples to calculate the fraction of TCRs clustered with each sample category. The fraction of "HC" is used to distinguish diseased from healthy individuals. The other fractions are used to distinguish the three diseases. This more stringent method achieved a similar level of predictive accuracy as leave-one-out validation.

[0179] AUC values ​​are shown at the bottom right of each chart. Each chart shows specificity on the x-axis and sensitivity on the y-axis. Figure 25A shows a comparison between cancer TCR clusters and HC TCR clusters. Figure 25B shows a comparison between COVID-19 TCR clusters and HC clusters. Figure 25C shows a comparison between MS clusters and HC clusters. Figure 25D shows a comparison between cancer TCR clusters and COVID-19 clusters. Figure 25E shows a comparison between COVID-19 clusters and MS clusters. Figure 25F shows a comparison between cancer clusters and MS clusters.

[0180] Referring now to FIG. 27, a chart illustrating the cross-cohort similarity of reference TCR-seq samples is shown. Using the TCR clustering data of N samples, the percentage of TCRs of each sample that co-cluster with each of the other samples may be calculated. The auto-coclustering percentage may be assigned as zero, making all vectors of length N. A Spearman correlation matrix may be calculated from the N×N co-clustering fraction matrix. The matrix may then be collapsed according to cancer type. The average of the top five largest correlations is displayed in FIG. 27 as a heat map. The same disease correlations (diagonal values) may be calculated similarly, except that the auto-correlation of each sample may be removed prior to calculation. Color coding represents the values ​​on the display. For example, red represents positive correlation.

[0181] The ability to discriminate between lung cancer and COVID-19 was not inconsistent with the apparent grouping of the two diseases, as intradisease similarities were still high, but the majority of diseases were derived from only one study, raising concerns that the predictability may be contributed by unknown batch effects.

[0182] To explore this possibility, we used GIANA to predict disease labels for unknown samples from independent cohorts. We applied GIANA to query 267 new TCR-seq samples of the same disease and 153 HC samples against the reference dataset. All samples were from peripheral blood. Using the same approach, we calculated the fraction of TCRs that co-clustered with reference cancer, COVID-19, MS, or HC sequences. Table 5 below shows the TCR-seq sample cohorts used as query data.

[0183] [Table 5]

[0184] All 120 of the second batch of healthy donors from Emerson's 2017 study were used as controls. To avoid duplication with references, the Hospital Universitario 12 de Octubre (HUniv120, n=193) cohort from Nolan's 2020 study was used for COVID-19 patients. Patients in this cohort were collected from Madrid, Spain. In one example, it took 20.5 hours for GIANA to complete a query of all 420 samples on a MacBook Pro with a 3.5GHz Dual-Core Intel Core i7 processor and 16GB of 2133MHz LPDDR3 memory.

[0185] Referring now to Figures 28A-28D, we show "violin plots" showing the distribution of class fractions for cancer, COVID-19, MS patients, and HCs. Each of the "violin plots" shown in Figures 28A-28D illustrates the distribution of data. The symmetric curves on the sides of the "violins" are the actual probability density of the data points. In the center, a typical box plot illustrates the mean (midpoint) and interquartile range of the data. The y-axis shows the fraction of TCRs for a given disease category (shown as subpanel captions). The x-axis shows the disease category. Figures 28A-28D show that the TCR fraction estimated by GIANA for a given disease is highest in patients with that disease, justifying its use as a predictor for multiple diseases.

[0186] Class fractions (e.g., cancer fractions) may be calculated as the proportion of query TCRs that cluster with reference TCRs from cancer patients. Without any model training, this simple approach is able to distinguish each sample category from the others. The HC fraction is discriminated from all three diseases with over 91% accuracy.

[0187] Referring now to Figures 29A-29F, charts illustrating ROC curves using disease class fraction as a single predictor for pairwise separation of the four disease classes are shown. Fraction can be the percentage of TCRs co-clustered with samples of a given class in the reference dataset. AUC values ​​are shown at the bottom right of each chart. Each chart shows specificity on the x-axis and sensitivity on the y-axis. Figure 29A shows a comparison between cancer TCR clusters and HC TCR clusters. Figure 29B shows a comparison between COVID-19 TCR clusters and HC clusters. Figure 29C shows a comparison between MS clusters and HC clusters. Figure 29D shows a comparison between cancer TCR clusters and COVID-19 clusters. Figure 29E shows a comparison between COVID-19 clusters and MS clusters. Figure 29F shows a comparison between cancer clusters and MS clusters. 95% confidence intervals were calculated using stratified 2,000 bootstraps. Pairwise separations between the diseases all reached an AUC above 87%. Because the query samples were derived from studies not included in the reference data, the large AUCs were not caused by unknown batch or cohort specific effects, but likely reflect the true predictive potential for the three diseases.

[0188] GIANA query performance was compared with traditional repertoire classification methods based on multiple instance learning (MIL) and fitting of cohort-specific parameters (e.g., DeepRC and others). GIANA does not require any parameter fitting, while traditional methods provide suitable reference data (e.g., repertoire samples from true COVID-19 patients and negative controls) with similar attributes to the query sample. Using a cohort containing HCMV+ and HCMV- subjects, 75% of the samples were applied as references (similar to training) and the remaining 25% were applied as test data. Each test sample was searched against the reference data. For each query sample, the fraction of TCRs that co-clustered with the HCMV+ reference subjects was calculated and used as a predictor. This simple approach reached an AUC of 83.06%, the same as DeepRC and better than other methods, as shown in the chart below. Therefore, GIANA would be a competitive method for repertoire classification.

[0189] [Table 6]

[0190] Note that to generate the ROC curves and estimate the AUC values ​​above, the pROC package in the R programming language was used, and 95% confidence intervals were calculated with 2,000 stratified bootstrap replicates, performed using the ci.auc function in the pROC package. Figure 22 was created using the igraph package. The heatmap in Figure 27 with annotated values ​​was created using the heatmap.2 function in the gplots package. For all box plots shown, the center line defines the median value, and the box boundaries indicate the 25% (Q1) and 75% (Q3) quartiles of the data. The lower and upper whiskers correspond to Q1-1.5IQR and Q3+1.5IQR, where IQR is an abbreviation for interquartile range.

[0191] In summary, GIANA is a novel antigen-specific TCR clustering algorithm that can efficiently handle tens of millions of sequences. GIANA achieves higher sensitivity and accuracy than all existing methods and can search for TCRs specific to known antigens with high accuracy. Ultra-large-scale TCR clustering and fast querying of novel samples also enable a novel reference-based repertoire classification framework. GIANA can also analyze single-cell RNA-seq data with elucidated TCR regions, and can query scRNA-seq data for TCRs against large databases of TCR repertoire samples in the public domain, providing new insights into shared antigen specificities. With minimal modifications, GIANA can also be applied to cluster or query large B cell receptor sequencing data. Furthermore, the mathematical framework for performing isometric embedding may provide an alternative solution to the classical short DNA or protein sequence alignment challenges in the future.

[0192] Note that HLA alleles were not considered in GIANA because such data are not available in most current studies. The inclusion of HLA typing is expected to improve the accuracy of TCR clustering and query methods. GIANA does not support gap alignment, which has better sensitivity than other methods with this functionality. This is because allowing gaps may reduce the specificity of clustering and sacrifice prediction accuracy.

[0193] Also, as mentioned above, simple fractional estimates are used to assign disease classes. With more data, this effort could be improved by machine learning models optimizing predictive accuracy. Furthermore, all cancer patients were compared together with other diseases without distinguishing between cancer localizations. However, the ability to separate cancer types with sufficient TCR-seq samples related as references is contemplated. Although the current GIANA method already achieves high accuracy of repertoire classification, the diagnostic value of this platform would be improved by prospectively collected patient samples.

[0194] As we have demonstrated in autoimmune and infectious diseases, low-frequency shared antigen-specific well-known TCRs are potentially important biomarkers that can be detected by comparing large numbers of TCRs from thousands of individuals. Methods have been developed to use immune repertoires to detect cancer, COVID-19, and MS individually, but none of them can simultaneously diagnose and separate the various diseases. In contrast, GIANA can be used as an integrated platform to diagnose infectious diseases, autoimmune disorders, and cancer.

[0195] This offers several improvements over conventional methods. Traditionally, disease diagnosis has been based primarily on symptoms, with each disease requiring a distinct set of signatures derived from a variety of clinical assays, including radioactive imaging, liquid biopsy, invasive endoscopy, surgery, and others. The feasibility of using the immune system as a single biomarker to indicate multiple diseases creates a paradigm shift from symptom-based to immune response-based, which may provide a universal solution to many immune-related disorders.

[0196] Furthermore, differential diagnosis is usually clinically challenging, and adding more diseases to the platform is expected to decrease diagnostic specificity, but the categorical accuracy of GIANA is actually increased by including more TCR-seq samples.

[0197] Furthermore, because immune responses usually precede any measurable symptoms, the GIANA platform has the potential to detect disease at an early stage when most diseases are curable or easily managed. This has already been shown in cancer diagnosis, and the principles of immune control also apply to autoimmune disorders such as MS. Finally, since the platform only requires a small amount of blood to perform targeted V(D)J capture, it can serve as a low-cost, non-invasive test. Together, GIANA can be widely used to find antigen-specific TCR clusters, search for sequences specific to known pathogens such as SARS-CoV-2, and facilitate disease diagnosis from the rapidly growing body of TCR data in cancer, immunology, and clinical research.

[0198] For purposes of this disclosure, a module is a software, hardware, or firmware (or combination thereof) system, process, or functionality, or component thereof, that performs or facilitates the processes, features, and / or functions described herein (with or without human interaction or augmentation). A module may include sub-modules. The software components of a module may be stored in a computer-readable medium for execution by a processor. A module may be integral to one or more servers or may be hosted and executed by one or more servers. One or more modules may be grouped into an engine or application.

[0199] Those skilled in the art will recognize that the disclosed method and system may be implemented in many ways and are therefore not limited to the above examples. In other words, functional elements and individual functions implemented by single or multiple components in various combinations of hardware and software or firmware may be distributed in software applications at the client level or server level or both. In this regard, any number of features of the various embodiments described herein may be combined in a single or multiple embodiments, and alternative embodiments having less than or more than all of the features described herein are possible.

[0200] The functionality may be distributed, in whole or in part, among multiple components in ways now known or that will become known in the future; that is, multiple software / hardware / firmware combinations are possible in achieving the functions, features, interfaces, and preferences described herein. Moreover, the scope of the disclosure includes conventionally known ways of implementing the described features and functions and interfaces, as well as variations and modifications made to the hardware or software or firmware components described herein, now and in the future, as will be understood by those skilled in the art.

[0201] Furthermore, the method embodiments shown and described as flow charts in this disclosure are provided as examples to provide a more complete understanding of the present technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and in which sub-operations described as part of larger operations are performed independently.

[0202] Although various embodiments have been described for purposes of this disclosure, these embodiments should not be construed as limiting the teachings of this disclosure to these embodiments. Various changes and modifications may be made to the elements and operations described above in order to obtain results that remain within the scope of the systems and processes described in this disclosure.

Claims

【Claim 1】 The invention described in the specification.