Blind denoising of sequencing data

The sequence denoising model addresses the challenge of correcting noisy read sequences in sequencing data by generating accurate scaffold sequences using latent space representations and machine learning, enhancing the accuracy of bioinformatics analyses.

JP2025533581APending Publication Date: 2025-10-07GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517888
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-05
Filing Date
2023-09-27
Publication Date
2025-10-07

AI Technical Summary

Technical Problem

Existing sequencing technologies face challenges in accurately resolving long nucleotide sequences and correcting noisy read sequences containing errors such as insertions, deletions, and substitutions, especially when the original scaffold sequence is unavailable.

Method used

A sequence denoising model, comprising a sequence encoder and decoder, is trained to generate denoised scaffold sequences from noisy read sequences without access to the original scaffold sequence, utilizing latent space representations and machine learning architectures like transformers to minimize edit distances and capture latent features.

Benefits of technology

The model effectively denoises noisy read sequences, providing accurate reconstructions of the original scaffold sequences for downstream bioinformatics tasks, outperforming conventional methods like multiple sequence alignment in accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533581000001_ABST
    Figure 2025533581000001_ABST
Patent Text Reader

Abstract

A method for blind denoising of sequencing data may include receiving a plurality of read sequences associated with a scaffold sequence. Each read sequence may be a repeat of the scaffold sequence, at least one of which is a noisy read sequence that does not match the scaffold sequence. A sequence denoising model may be applied to encode the plurality of read sequences and generate a denoised scaffold sequence corresponding to the scaffold sequence based on the encoding of the plurality of read sequences. The denoised scaffold sequence may be generated without the scaffold sequence. Molecules associated with the scaffold sequence may be analysed based on the denoised scaffold sequence. Related systems and computer program products are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 377,337, filed September 27, 2022, entitled "Blind Noise Removal of Sequencing Data," and U.S. Provisional Application No. 63 / 378,460, filed October 5, 2022, entitled "Blind Noise Removal of Sequencing Data," the disclosures of which are incorporated herein by reference in their entireties.

[0002] Technical Field The subject matter described herein relates generally to gene sequencing, and more specifically to techniques for denoising sequencing data. [Background technology]

[0003] Introduction In the context of genetics and biochemistry, the term "sequencing" refers to various techniques for determining the primary structure of unbranched biopolymers. For example, sequencing a nucleic acid molecule (e.g., deoxyribonucleic acid (DNA), ribonucleic acid (RNA), etc.) or a variant or derivative thereof (e.g., single-stranded DNA) can involve determining the sequence of nucleotide bases forming the nucleic acid molecule. In the case of DNA, sequencing a DNA fragment can involve determining the order of guanine (G), adenine (A), cytosine (C), and thymine (T) in the DNA fragment. Meanwhile, sequencing an RNA fragment can involve determining the order of guanine (G), adenine (A), cytosine (C), and uracil (U) in the RNA fragment. Summary of the Invention

[0004] overview Systems, methods, and products are provided, including computer program products, for blind denoising of sequencing data including a plurality of read sequences associated with a scaffold sequence without the scaffold sequence. In one aspect, a system for blind denoising of sequencing data is provided. The system may include at least one processor and at least one memory. The at least one memory may include program code that, when executed by the at least one processor, results in operations. The operations may include receiving a plurality of read sequences associated with the scaffold sequence, wherein each read sequence of the plurality of read sequences includes a repeat of the scaffold sequence, and at least one read sequence of the plurality of read sequences is a noisy read sequence that does not match the scaffold sequence; determining encodings associated with the plurality of read sequences and applying a sequence denoising model trained to generate denoised scaffold sequences corresponding to the scaffold sequence based at least on the encodings associated with the plurality of read sequences; and analyzing molecules associated with the scaffold sequence based at least on the denoised scaffold sequence.

[0005] In another aspect, a method for blind denoising of sequencing data is provided. The method may include receiving a plurality of read sequences associated with a scaffold sequence, wherein each read sequence of the plurality of read sequences includes a repeat of the scaffold sequence, and at least one read sequence of the plurality of read sequences is a noisy read sequence that does not match the scaffold sequence; determining encodings associated with the plurality of read sequences and applying a sequence denoising model trained to generate denoised scaffold sequences corresponding to the scaffold sequences based at least on the encodings associated with the plurality of read sequences; and analyzing molecules associated with the scaffold sequences based at least on the denoised scaffold sequences.

[0006] In another aspect, a computer program product for blind denoising of sequencing data is provided. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations to occur. The operations may include receiving a plurality of read sequences associated with a scaffold sequence, where each read sequence of the plurality of read sequences includes a repeat of the scaffold sequence, and at least one read sequence of the plurality of read sequences is a noisy read sequence that does not match the scaffold sequence; determining encodings associated with the plurality of read sequences and applying a sequence denoising model that is trained to generate denoised scaffold sequences corresponding to the scaffold sequence based at least on the encodings associated with the plurality of read sequences; and analyzing molecules associated with the scaffold sequence based at least on the denoised scaffold sequence.

[0007] In some variations of the methods, systems, and non-transitory computer-readable media, one or more of the following features may optionally be included in any feasible combination.

[0008] In some variations, the sequence denoising model may generate denoised scaffold sequences without the scaffold sequences.

[0009] In some variations, the sequence denoising model may generate a denoised scaffold sequence by at least encoding a plurality of read sequences to generate a plurality of embeddings, determining an aggregate embedding corresponding to an aggregation of the plurality of embeddings, and decoding the aggregate embedding to generate a denoised scaffold sequence.

[0010] In some variations, the sequence denoising model may include a sequence encoder trained to generate multiple embeddings by at least encoding the multiple read sequences. The sequence encoder may be trained to increase similarity between the multiple embeddings when encoding the multiple read sequences by at least decreasing an edit distance between the multiple embeddings when encoding the multiple read sequences.

[0011] In some variations, the sequence encoder may be trained to generate, for each lead sequence of the multiple lead sequences, a corresponding embedding in the latent space.

[0012] In some variations, the latent space may be a topological space occupied by a reduced dimensional representation of multiple lead sequences.

[0013] In some variations, the constellation encoder may be a transformer.

[0014] In some variations, the sequence denoising model may include a sequence decoder that is trained to generate the denoised scaffold sequences by at least decoding the aggregate embeddings to generate the denoised scaffold sequences.

[0015] In some variations, the sequence decoder may include a transformer.

[0016] In some variations, the sequence denoising model may include a set encoder that is trained to determine the aggregate embedding.

[0017] In some variations, the set encoder may be trained to reduce a first distance between the aggregate embedding and the plurality of embeddings in the latent space when determining the aggregate embedding. The set encoder may be further trained to reduce a second distance between the denoised scaffold sequence and the plurality of lead sequences in the sequence space when determining the aggregate embedding.

[0018] In some variations, the set encoder may include a set transformer.

[0019] In some variations, the noisy read sequence may contain at least one insertion, deletion, or substitution of a nucleobase type contained in the scaffold sequence.

[0020] In some variations, the identity of the molecule associated with the scaffold sequence may be determined based at least on the denoised scaffold sequence.

[0021] In some variations, the binding specificity of an antibody may be determined based at least on the identity of the molecule that is conjugated to the antibody.

[0022] In some variations, the gene expressing the T cell receptor (TCR) may be identified based at least on the identity of the molecule that is conjugated to the T cell receptor (TCR).

[0023] In some variations, the scaffold sequence may be associated with a nucleic acid sequence tag that is conjugated to (i) the heavy or light chain of a B cell receptor, or (ii) the alpha and / or beta chain of a T cell receptor.

[0024] In some variations, the scaffold sequence may be a deoxyribonucleic acid (DNA) sequence or a ribonucleic acid (RNA) sequence.

[0025] In some variations, a sequence denoising model may be trained based at least on a training set to generate one or more denoised scaffold sequences while one or more corresponding scaffold sequences are unknown. Performance of the sequence denoising model may be determined while one or more corresponding scaffold sequences are unknown. Once the performance of the sequence denoising model is determined to satisfy one or more thresholds, the sequence denoising model may be applied to determine encodings associated with the plurality of read sequences and generate denoised scaffold sequences corresponding to the scaffold sequences based at least on the encodings associated with the plurality of read sequences.

[0026] In some variations, a leave-one-out edit distance of the sequence denoising model may be determined by at least determining an edit distance between one lead sequence from the training set and a denoised scaffold sequence generated by the sequence denoising model based on one or more remaining lead sequences in the training set. Performance of the sequence denoising model may be determined based at least on the leave-one-out edit distance.

[0027] Implementations of the present subject matter may include, but are not limited to, methods consistent with the descriptions provided herein, as well as articles comprising tangibly embodied machine-readable media operable to cause one or more machines (e.g., computers, etc.) to perform operations that implement one or more of the described features. Similarly, computer systems, which may include one or more processors and one or more memories coupled to the one or more processors, are also described. The memory, which may include a non-transitory computer-readable or machine-readable storage medium, may include, encode, or store one or more programs that cause the one or more processors to perform one or more of the operations described herein. Computer-implemented methods consistent with one or more implementations of the present subject matter may be implemented by one or more data processors present in a single computing system or in multiple computing systems. Such multiple computing systems may be connected, e.g., to exchange data and / or commands or other instructions, etc., via one or more connections, including, for example, connections via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), direct connections between one or more of the multiple computing systems, etc.

[0028] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will become apparent from the specification and drawings, and from the claims. While certain features of the presently disclosed subject matter are described for illustrative purposes in connection with denoising sequencing data associated with deoxyribonucleic acid (DNA) and ribonucleic acid (RNA), it should be readily understood that such features are not intended to be limiting. The claims following this disclosure are intended to define the scope of the protected subject matter. [Brief explanation of the drawings]

[0029] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, serve to explain some of the principles associated with the disclosed embodiments.

[0030] [Figure 1A] 1 depicts a system diagram showing an example of a sequencing system, according to some exemplary embodiments.

[0031] [Figure 1B] 1 depicts a flowchart illustrating an example of a process for denoising multiple read sequences associated with a scaffold sequence without the scaffold sequence, according to some exemplary embodiments.

[0032] [Figure 2] FIG. 1 depicts a schematic diagram showing an example of blind denoising in which multiple lead sequences associated with a scaffold sequence are denoised without the scaffold sequence, according to some exemplary embodiments.

[0033] [Figure 3A] 1 depicts a schematic diagram illustrating an example of a latent space, according to some exemplary embodiments;

[0034] [Figure 3B] FIG. 1 depicts a schematic diagram illustrating an example of a sequence denoising model that generates a latent space representation of each noisy read sequence, according to some exemplary embodiments.

[0035] [Figure 3C] FIG. 1 depicts a schematic diagram illustrating an example of a sequence denoising model that generates a set encoding of a latent space representation of multiple noisy read sequences, according to some exemplary embodiments.

[0036] [Figure 3D] 1 depicts a schematic diagram illustrating an example of latent space regularization, according to some exemplary embodiments;

[0037] [Figure 3E] 1 depicts a schematic diagram illustrating an example of a leave-one-out (LOO) edit distance, according to some exemplary embodiments;

[0038] [Figure 4] 1 depicts a schematic diagram illustrating losses associated with an example of a sequence denoising model, according to some exemplary embodiments;

[0039] [Figure 5] 1 depicts an exemplary use case in which a lead sequence associated with the light chain of an antibody is denoised using a sequence denoising model and a conventional multi-sequence alignment-based denoising technique, according to some exemplary embodiments.

[0040] [Figure 6A] 1 depicts a graph showing a comparison of the respective performance of a sequence denoising model and a conventional multi-sequence alignment based denoising technique, according to some exemplary embodiments;

[0041] [Figure 6B] 10 depicts a graph showing another comparison of the respective performance of a sequence denoising model and a conventional multi-sequence alignment based denoising technique, according to some exemplary embodiments;

[0042] [Figure 6C] 10 depicts a graph showing another comparison of the respective performance of a sequence denoising model and a conventional multi-sequence alignment based denoising technique, according to some exemplary embodiments;

[0043] [Figure 7] 1 depicts a block diagram of an example computing system, according to some illustrative embodiments.

[0044] Wherever practical, like reference numerals refer to like structures, features, or elements. DETAILED DESCRIPTION OF THE INVENTION

[0045] Detailed Description Sequencing of nucleic acid molecules, such as deoxyribonucleic acid (DNA) and ribonucleic acid (RNA), plays a prominent role in various biotechnological pursuits. As an exemplary use case in the context of large molecule drug discovery, antibody binding nonspecificity (or polyspecificity), an undesirable biophysical property in which an antibody can bind to epitopes on many different antigens, can be assessed by sequencing nucleic acid sequence tags (e.g., plasmids, oligonucleotides, etc.) conjugated to each antibody expressed by yeast cells engineered to display individual antibodies from a library of different antibodies. For example, once a labeled antibody is exposed to one or more target and non-target antigens (e.g., polyspecific reagents), fluorescence-activated cell sorting (FACS) can be performed to separate bound from unbound antibodies before sequencing the nucleic acid sequence tags of bound and unbound antibodies. Nevertheless, existing sequencing techniques for sequencing nucleic acid sequence tags are associated with significant limitations. For example, some next-generation sequencing technologies cannot resolve long sequences of nucleotide bases (e.g., sequences longer than 550 base pairs), while antibody heavy and light chains contain approximately 700 base pairs. Third-generation sequencing technologies (e.g., nanopore sequencing) may be able to resolve longer sequences, but the long-read sequences output by applying these third-generation sequencing technologies tend to be noisy read sequences containing errors such as insertions, deletions, and substitutions of one or more nucleotide bases. Therefore, in some cases, noisy read sequences may be denoised to correct discrepancies between the noisy read sequence and the original scaffold sequence. In particular, the present disclosure describes various blind denoising techniques that can denoise noisy read sequences in the absence of the original scaffold sequence. That is, when the original scaffold sequence is unavailable (as is often the case), the noisy read sequence can still be denoised by applying the blind denoising techniques described herein.

[0046] In some exemplary embodiments, the sequencing controller may perform blind denoising of multiple read sequences associated with an original scaffold sequence. For example, each read sequence may be a repeat of the original scaffold sequence. Furthermore, at least one of the read sequences may be a noisy read sequence (e.g., a noisy repeat) that contains an insertion, deletion, or substitution of at least one nucleotide base present in the original scaffold sequence. The sequencing controller may be trained to denoise read sequences to generate denoised scaffold sequences without access to and / or knowledge of the original scaffold sequence. That is, the original scaffold sequence may not exist such that the sequence of nucleotide bases in the original scaffold sequence may remain unknown (as in the "blind" case) through denoising of the corresponding read sequence. The sequencing controller may generate a denoised scaffold sequence based on at least one noisy read sequence (e.g., having an insertion, deletion, or substitution of at least one nucleotide base present in the scaffold sequence), such that the sequence of nucleotide bases in the denoised scaffold sequence corresponds to the sequence of nucleotide bases in the scaffold sequence.

[0047] In some exemplary embodiments, the sequencing controller may apply a sequence denoising model to generate denoised scaffold sequences. The sequence denoising model may generate the denoised scaffold sequences by at least encoding each lead sequence to generate a corresponding embedding within a latent space learned by the sequence denoising model during training of the sequence denoising model. Furthermore, the sequence denoising model may generate an aggregate embedding corresponding to an aggregation of embeddings associated with each lead sequence before decoding the aggregate embedding to generate the denoised scaffold sequence. In this regard, the latent space may be a theoretical topological space (e.g., a manifold) occupied by embeddings, each of which is a reduced-dimensional representation of a corresponding lead sequence. For example, the sequence denoising model may learn various latent features present in the lead sequences during training. These latent features may not be directly observable in each lead sequence but may capture relationships that exist between different lead sequences. Furthermore, lead sequences may exhibit multiple features in the sequence space. Each lead sequence may represent fewer latent features. Thus, the embedding of a lead sequence is said to be a reduced-dimensional representation of the lead sequence, but the latent space occupied by the embedding of a lead sequence may have lower dimensionality (e.g., fewer dimensions) than the sequence space occupied by the lead sequence.

[0048] In some exemplary embodiments, the sequence denoising model may include a sequence encoder that is trained to generate, for each read sequence associated with an original scaffold sequence, a corresponding embedding in a latent space. For example, the latent space learned during training of the sequence denoising model may be a topological space (e.g., a manifold) occupied by embeddings, each of which is a reduced-dimensional representation of the corresponding read sequence. In some cases, the sequence encoder may be implemented with a machine learning architecture that can correlate different portions of each noisy read sequence. For example, in some cases, the sequence encoder may be implemented with a deep learning architecture with multiple parallel attention mechanisms, such as a transformer. Furthermore, in some cases, the sequence encoder may be trained to maximize or increase the similarity between individual embeddings for each read sequence, for example, by minimizing or decreasing the edit distance (e.g., kernelized maximum mean discrepancy (MMD)) between embeddings when encoding each of the read sequences.

[0049] In some exemplary embodiments, the sequence denoising model may include a sequence decoder configured to decode an aggregate embedding corresponding to an aggregation of embeddings associated with each lead sequence. In some cases, the sequence decoder may be implemented as a transformer. Further, in some cases, the sequence decoder may be configured to decode the aggregate embedding by at least transforming the aggregate embedding from its representation in latent space to a denoised scaffold sequence in sequence space that is similarly occupied by the lead sequences associated with the original scaffold sequence.

[0050] In some exemplary embodiments, the sequence denoising model may include a set encoder configured to generate an aggregate embedding corresponding to an aggregation of embeddings associated with each lead sequence. For example, the set encoder may be implemented as a set transformer. In some cases, the aggregate embedding may provide a solution or an optimal solution to a loss function associated with the sequence denoising model that captures the loss (or error) present in the output of the sequence denoising model in the sequence space occupied by the various sequences as well as in the latent space occupied by the corresponding embeddings. Thus, the set encoder may be trained to minimize or reduce a first distance between the aggregate embedding in the latent space and the embedding of the lead sequence when determining the aggregate embedding. Furthermore, the set encoder may be trained to minimize or reduce a second edit distance between the denoised scaffold sequence and the lead sequence in the sequence space when determining the aggregate embedding.

[0051] In some exemplary embodiments, the original scaffold sequence may be associated with various nucleic acid molecules (e.g., DNA, RNA, etc.) or variants or derivatives thereof (e.g., single-stranded DNA). Thus, the nucleic acid molecules may be analyzed as part of various downstream bioinformatics tasks, e.g., based at least on the denoised scaffold sequence generated by the sequencing controller to correspond to the original scaffold sequence. For example, in some cases, the scaffold sequence may be associated with at least a portion of a molecule such as an antigen or antibody. Alternatively, in some cases, the scaffold sequence may be associated with the heavy and / or light chains of a B cell receptor (BCR), or the alpha and / or beta chains of a T cell receptor (TCR). Thus, in some cases, the denoised scaffold sequence may serve as an identifier (or label) in a protein molecule to which the corresponding nucleic acid sequence is conjugated (e.g., linked by a chemically induced covalent bond between the two molecules). In the aforementioned use case for assessing the binding nonspecificity (or polyspecificity) of an antibody, the sequencing controller may perform blind denoising to denoise long read sequences associated with nucleic acid sequence labels (e.g., oligonucleotides, plasmids, etc.) conjugated to antibodies that may, for example, bind to the target antigen, not bind to the target antigen, not bind to a non-target antigen, or not bind to a non-target antigen. The resulting denoised scaffold sequences may then be used to identify the antibodies and, in some cases, determine the set of antibodies that bound (or did not bind) one or more target antigens and non-target antigens. Another exemplary use case is blind denoising of nucleic acid sequence labels conjugated to T cell receptors (TCRs). The resulting denoised scaffold sequences may be used to identify the corresponding T cell receptors and map them to expressed genes (e.g., as surface proteins).

[0052] FIG. 1A depicts a system diagram illustrating an example of a sequencing system 100, according to some exemplary embodiments. Referring to FIG. 1A, the sequencing system 100 may include a sequencing controller 110, a sequencing platform 120, an analysis engine 130, and a client 140. As shown in FIG. 1A, the sequencing controller 110, the sequencing platform 120, the analysis engine 130, and the client 140 may be communicatively coupled via a network 150. The client 140 may be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc. The network 150 may be a wired and / or wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc.

[0053] In some exemplary embodiments, the sequencing controller 110 may perform blind denoising of sequencing data, such as sequencing data 125 received from a sequencing platform 125. In some cases, the sequencing data 125 may include multiple noisy read sequences associated with an original scaffold sequence, but denoising of the noisy read sequences may be performed without access to the original scaffold sequence. Further, in some cases, the sequencing controller 110 may apply a sequence denoising model 115 to perform blind denoising of the sequence data 125. For example, in some cases, the sequence denoising model 115 may: TIFF2025533581000003.tif14170 In some cases, the sequence denoising model 115 can encode and decode sequences and interpolate between embedding sets. TIFF2025533581000004.tif6170 may undergo self-supervised set learning. By learning such an embedding space, the sequence denoising model 115 For example, in some cases, the sequence denoising model 115 may generate and decode an approximation of the true embedding of TIFF2025533581000005.tif4170 as part of the sequencing data 125. TIFF2025533581000006.tif5170 Corresponding denoised As explained in more detail below, each read sequence is To generate TIFF2025533581000008.tif8170 TIFF2025533581000009.tif7170Then the set of embeddings, in some cases, TIFF2025533581000010.tif8170 to generate a single aggregate embedding for the entire set TIFF2025533581000011.tif15170

[0054] To further illustrate, FIG. 1B depicts a flowchart showing an example of a process 160 for denoising multiple read sequences associated with a scaffold sequence without the scaffold sequence, according to some exemplary embodiments. With reference to FIGS. 1A-1B, process 160 may be performed by sequencing controller 110, for example, to denoise noisy read sequences included in sequencing data 125. The denoising of the noisy read sequences may be blind, meaning that the sequence of nucleotide bases forming the corresponding scaffold sequence remains unknown during the denoising of the noisy read sequence. As described in more detail below, in some cases, sequence controller 110 may apply a sequence denoising model 115 to denoise the noisy read sequences and generate a denoised scaffold sequence corresponding to the original scaffold sequence.

[0055] At 162, a plurality of read sequences associated with the scaffold sequence may be received. For example, in some exemplary embodiments, the sequence controller 110 may receive sequencing data 125 from the sequencing platform 120. As shown in FIG. 2, the sequencing data 125 may include a plurality of read sequences 250 associated with the original scaffold sequence 200. In some cases, the sequencing platform 120 may apply various sequencing technologies (e.g., next-generation sequencing technologies, third-generation sequencing technologies, etc.) to generate each of the plurality of read sequences 250. Each read sequence of the plurality of read sequences 250 may be a repeat of the original scaffold sequence 200. That is, each read sequence of the plurality of read sequences 250 may constitute a separate inference of the sequence of nucleic acid bases forming the original scaffold sequence 200. Furthermore, at least one read sequence of the plurality of read sequences 250 may be a noisy read sequence having an insertion, deletion, or substitution of at least one nucleotide base present in the original scaffold sequence 200. Thus, at least one lead sequence of the plurality of lead sequences 250 does not match the original scaffold sequence 200. Without proper noise removal, the plurality of lead sequences 250 may not provide a sufficiently accurate reconstruction of the original scaffold sequence 200 for downstream bioinformatics tasks performed in the analysis engine 130, such as assessing antigen binding non-specificity (or poly-specificity).

[0056] At 164, a sequence denoising model may be applied to determine an encoding associated with the plurality of read sequences and to generate a denoised scaffold sequence based at least on the encoding associated with the plurality of read sequences. In some exemplary embodiments, the sequencing controller 110 may apply the sequence denoising model 115 to perform blind denoising of the plurality of read sequences 250. As shown in FIG. 2 , the sequencing controller 110 may perform blind denoising of the plurality of read sequences 215 included in the sequencing data 125 to generate a denoised scaffold sequence 250 corresponding to the original scaffold sequence 200. That is, the sequencing controller 110 may denoise the read sequence 250 without or without access to the original scaffold sequence 200, such that the sequence of nucleotide bases in the original scaffold sequence remains unknown through the denoising of the plurality of read sequences 250. In some cases, the sequence controller 110 may apply a sequence denoising model 115 that may denoise the plurality of read sequences 250 and generate a denoised scaffold sequence 215 by at least generating an aggregate embedding of the plurality of read sequences 250. For example, in some cases, the sequence denoising model 115 may include a sequence encoder 310 configured to encode each read sequence of the plurality of read sequences 250 to generate a corresponding embedding. Further, the sequence denoising model 115 may include a set encoder 330 configured to generate an aggregate embedding corresponding to an aggregation of embeddings associated with the plurality of read sequences 250 before a sequence decoder 320 of the sequence denoising model 115 decodes the aggregate embedding to generate the denoised scaffold sequence 215. In some cases, the sequence encoder 310, the sequence decoder 320, and the set encoder 330 may be implemented using machine learning models.For example, in some cases, the sequence encoder 310 may be implemented as a first machine learning model (e.g., a first Transformer), the sequence decoder 320 may be implemented as a second machine learning model (e.g., a second Transformer), and the set encoder 330 may be implemented as a third machine learning model (e.g., a set Transformer).

[0057] In some exemplary embodiments, embeddings associated with the plurality of read sequences 250 may occupy a latent space learned by the sequence denoising model 115 during training. Accordingly, encoding each read sequence of the plurality of read sequences 250 may include converting the read sequence from its representation in sequence space to a corresponding embedding in latent space. That is, in some cases, the sequence encoder 310 may generate, for each read sequence of the plurality of read sequences 250, an embedding corresponding to a reduced-dimensional representation of the read sequence. In this regard, the latent space may be a data distribution (e.g., a topological space such as a manifold) occupied by embeddings of different sequences of nucleotide bases, including embeddings of the plurality of read sequences 250 and, as described in more detail below, aggregate embeddings associated with the plurality of read sequences 250. Further illustration is provided in exemplary FIGS. 3A-3C illustrating denoising of read sequences.

[0058] At 166, molecules associated with the scaffold sequence may be analyzed based at least on the denoised scaffold sequence. As previously mentioned, the original scaffold sequence 200 may be associated with at least a portion of a molecule such as an antigen or antibody (e.g., heavy and / or light chains of a B cell receptor (BCR), alpha and / or beta chains of a T cell receptor (TCR), etc.). Thus, by denoising a plurality of lead sequences 250 associated with the original scaffold sequence 200 without the original scaffold sequence 200 to generate a denoised scaffold sequence 215, various downstream bioinformatics may be performed to analyze molecules having the original scaffold sequence 200 based on the denoised scaffold sequence 215. For example, in one exemplary use case, the analysis engine 130 may evaluate various antibodies for binding non-specificity (or poly-specificity) based on the denoised scaffold sequences of the nucleic acid sequence labels (or tags) conjugated to the individual antibodies to distinguish between antibodies that bind (or do not bind) to one or more target antigens and / or non-target antigens. Another exemplary use case includes denoising nucleic acid sequence labels (or tags) conjugated to T cell receptors (TCRs) so that different T cell receptors can be mapped to genes that expressed the T cell receptor (as a surface protein) based on the corresponding denoised scaffold sequences.

[0059] Denoising the plurality of lead sequences 250 in the manner described above to generate the denoised scaffold sequence 215 may enable the denoised scaffold sequence 215 to be a sufficiently accurate reconstruction of the original scaffold sequence 200 for downstream bioinformatics tasks. For example, in some cases, the original scaffold sequence 200 may be associated with at least a portion of an antigen or antibody. Alternatively, in some cases, the original scaffold sequence 200 may be associated with a heavy and / or light chain of a B cell receptor (BCR), or an alpha and / or beta chain of a T cell receptor (TCR). Thus, in some cases, the denoised scaffold sequence 215 may serve as a molecular identifier for the corresponding protein molecule and / or cells expressing the protein molecule. For example, when using sequencing data 125 to assess the binding non-specificity (or polyspecificity) of an antibody, the use of denoised scaffold sequences 215 may allow for accurate identification of the binding (or non-binding) antibody, whereas the use of any of multiple lead sequences 250 may result in misidentification of the binding (or non-binding) antibody. In other cases where sequencing data 125 is used to map T cell receptors (TCRs) to corresponding genes, the use of denoised scaffold sequences 215 may allow for accurate disambiguation between genes expressing different T cell receptors, whereas the use of multiple lead sequences 250 may prevent such mapping from being performed with reasonable accuracy.

[0060] Denoising multiple read sequences, such as multiple read sequences 250, in the embodiments disclosed herein may achieve better denoising performance than conventional denoising techniques, such as multiple sequence alignment (MSA), in which a consensus sequence corresponding to the original scaffold sequence 200 is determined by aligning the multiple read sequences 250 and identifying the most common nucleotide base at each position. In particular, conventional denoising techniques, such as multiple sequence alignment (MSA), may not provide adequate denoising performance when there are only two read sequences associated with the original scaffold sequence 200. For example, multiple sequence alignment (MSA) requires at least three read sequences. Furthermore, conventional denoising techniques do not take into account biological prior information (e.g., absence of stops or frameshifts) and cannot exploit the overall structure within the sequencing data 125. Further illustration is provided in FIG. 5, which depicts an example case in which a read sequence associated with an antibody light chain is denoised.

[0061] 3A-3C depict schematic diagrams showing blind denoising of multiple read sequences 250 performed by sequence denoising model 115. For example, FIG. 3A illustrates a sequence denoising model 115 performed by sequencing controller 110 during training. TIFF2025533581000012.tif6170 As previously mentioned, in some exemplary embodiments, the sequence encoder 310 generates, for each read sequence of the plurality of read sequences 250, In the example shown in FIGS. 3A to 3C, the sequence encoder 310 detects the first read sequence 250a as follows: TIFF2025533581000014.tif5170 Second read sequence 250b TIFF2025533581000015.tif5170Third read sequence 250c TIFF2025533581000016.tif5170 4th lead sequence 250d TIFF2025533581000017.tif5170 and the fifth read sequence 250e TIFF2025533581000018.tif5170 Furthermore, as shown in FIGS. 3A to 3C, the set encoder 330 The decoder 320 of the TIFF2025533581000019.tif10170 array denoising model 115 is By decoding TIFF2025533581000020.tif5170, a denoised scaffold sequence 215 can finally be generated.

[0062] In some exemplary embodiments, TIFF2025533581000021.tif5170 Sequence space occupied by various sequences (e.g., original scaffold sequence 200, multiple read sequence 250, denoised scaffold sequence 215, etc.) and corresponding embeddings TIFF2025533581000022.tif13170 may provide an optimal solution to the loss function associated with the sequence denoising model 115 that captures the losses (or errors) present in the output of the sequence denoising model 115. Thus, the set encoder 330 TIFF2025533581000023.tif13170. TIFF2025533581000024.tif5170 can be trained to minimize or reduce the second edit distance between each of the decoded denoised scaffold sequences 215 and the read sequences 250. Thus, the set encoder 330 can be trained to minimize or reduce the second edit distance between each of the decoded denoised scaffold sequences 215 and the read sequences 250. TIFF2025533581000025.tif19170

[0063] Possibly includes edit distances between two or more embeddings of different lengths The distance between two or more embeddings of different lengths is calculated based on the kernelized maximum mean discrepancy (MMD) as follows: TIFF2025533581000027.tif5170Calculated based on Eq. (1) This corresponds to the latent space loss function mentioned above. In equation (1), TIFF2025533581000029.tif43170

[0064] As mentioned above, the decoder 320 of the sequence denoising model 115 TIFF2025533581000030.tif5170 may generate a denoised scaffold sequence 215. In some exemplary embodiments, due to the latent space regularisation imposed during training of the sequence denoising model 115, the decoder 320 TIFF2025533581000031.tif10170 may be adjusted to decode. TIFF2025533581000032.tif6170 (or adjustment) is not populated by embedding individual repeats TIFF2025533581000033.tif6170 You can increase the density so that it is still significant. For example, Decoding an embedding that occupies the midpoint between two or more repeat read sequences may not necessarily produce a sequence that makes sense in sequence space. TIFF2025533581000035.tif6170 Increases the likelihood that the sequence decoded from the embedding will result in a sensible sequence within the sequence space.

[0065] It should be understood that regularization of the latent space can be achieved by applying various regularization techniques. For example, as shown in Figure 3D, regularization of the latent space can be achieved by: TIFF2025533581000036.tif10170, thus distributing the probability mass of each embedding across the latent space. Alternatively and / or additionally, L2 regularization, also known as weight decay or ridge regression, may be applied by introducing a regularization term into the loss function associated with the sequence denoising model 115. The embedding parameters may also be adjusted, for example, The sequence denoising model 115 may be regularized to approximate the origin of the latent space. In some cases, the read sequences captured by the sequence denoising model 115 during training may be masked (e.g., masking one or more base pairs in the read sequences) to force the model to use context information. Augmenting the training data in this manner may prevent the sequence denoising model 115 from memorizing repeat read sequences in the training data. Masking one or more random base pairs in the read sequences captured by the sequence denoising model 115 during training may increase the robustness of the sequence denoising model 115. The trained sequence denoising model 115 may make inferences based on the context base pairs around the masked base pairs and may therefore be more tolerant to small omissions in the input read sequences.

[0066] In some exemplary embodiments, the sequence denoising model 115 may be trained in a self-supervised manner to reduce or minimize a loss function of the sequence denoising model 115. In some cases, the loss function of the sequence denoising model 115 may capture errors associated with the sequence encoder 310, the sequence decoder 320, and the set encoder 330. For example, FIG. 4 depicts a schematic diagram showing losses associated with each of the sequence encoder 310, the sequence decoder 320, and the set encoder 330. As shown in FIG. 4, the overall loss function associated with the sequence denoising model 115 may include an auto-encoding loss associated with the encoding performed by the encoder 310 and the decoding performed by the decoder 320, a latent space distance loss associated with the aggregate embedding generated by the set encoder 330, and a sequence space distance loss resulting when the aggregate embedding is decoded into its corresponding sequence space representation. The loss function associated with the sequence denoising model 115 may be further expressed as the following equation: TIFF2025533581000038.tif43170

[0067] As mentioned above, the sequence controller 110 may apply the sequence denoising model 115 to denoise the plurality of lead sequences 250, while the original scaffold sequences 200 associated with the plurality of lead sequences 250 remain unknown. Furthermore, in some exemplary embodiments, the sequence denoising model 115 may be trained in a self-supervised manner, for example, by undergoing self-supervised ensemble learning (SSSL), which means that the sequence denoising model 115 may be trained without ground truth scaffold sequences.

[0068] Therefore, in some exemplary embodiments, the performance of the sequence denoising model 115 cannot be evaluated based on the edit distance between the denoised scaffold sequences generated by the sequence denoising model 115 and the corresponding ground truth scaffold sequences. Instead, the performance of the sequence denoising model 115 may be evaluated based on the leave-one-out (LOO) edit distance. For example, if the training set used to train the sequence denoising model 115 includes multiple lead sequences 250, the sequence controller 110 may apply the sequence denoising model 115 to encode all but one of the multiple lead sequences 250. In the example shown in FIG. 3E, the sequence controller 110 may omit the fifth lead sequence 250e. The edit distance between the denoised scaffold sequence 215 generated by the sequence denoising model 115 by decoding TIFF2025533581000039.tif10170 and the fifth read sequence 250e can serve as a proxy indicator of the performance of the sequence denoising model 115. The edit distance between the denoised scaffold sequence 215 generated by the sequence denoising model 115 by decoding TIFF2025533581000040.tif5170 and the fifth read sequence 250e may correspond to the upper limit of the edit distance between the denoised scaffold sequence 215 and the true original scaffold sequence 210.

[0069] 5 depicts an exemplary use case in which a lead sequence 600 associated with the light chain of an antibody is denoised using the sequence denoising model 115 and a conventional multiple sequence alignment (MSA)-based denoising technique. The respective performance of each denoising technique may be quantified based on the edit distance relative to the original scaffold sequence 610 of the light chain. For example, as shown in FIG. 5, the consensus sequence 620 generated by applying the conventional multiple sequence alignment (MSA)-based technique is associated with an edit distance of 29, while the denoised scaffold sequence 630 generated by applying the sequence denoising model 115 is associated with a much smaller edit distance of 1. These results indicate that the sequence denoising model 115 described herein can achieve superior denoising performance compared to conventional denoising techniques such as multiple sequence alignment (MSA).

[0070] 6A-6C depict graphs illustrating a further comparison of the respective performances of the sequence denoising model 115 and conventional multi-sequence alignment-based denoising techniques, according to some exemplary embodiments. For example, FIG. 6A depicts a graph illustrating a comparison of the respective performances of the sequence denoising model 115 and conventional multi-sequence alignment-based denoising techniques when denoising a simulated V / J sequence that is 100 base pairs in length. Repeat sequences may be generated for each simulated V / J sequence to form a total of 10,000 sequences. As shown in FIG. 6A, the sequence denoising model 115 achieved better denoising performance, as quantified by the edit distance to the original scaffold sequence, than the conventional multi-sequence alignment-based denoising technique. In particular, FIG. 6A illustrates that the sequence denoising model 115 can achieve better denoising performance, as indicated by the lower edit distance to the original scaffold sequence, even when fewer read sequences are available to reconstruct the original scaffold sequence.

[0071] FIG. 6B depicts a graph comparing the performance of the sequence denoising model 115 and a conventional multi-sequence alignment-based denoising technique when denoising an antibody light chain sequenced by the Oxford Nanopore Technology sequencing platform. In the example shown in FIG. 6B, the sequence denoising model 115 and a conventional multi-sequence alignment-based denoising technique were used to denoise 100,000 scaffold sequences with an average length of 327 base pairs per scaffold sequence and an average of 8 repeat read sequences. FIG. 6B shows that the sequence denoising model 115 outperforms the conventional multi-sequence alignment-based denoising technique, even though the sequencing data is filtered to be successful with the conventional multi-sequence alignment-based denoising technique. As shown in FIG. 6B, the sequence denoising model 115 outperforms the conventional multi-sequence alignment-based denoising technique, achieving lower one-base pair edit distances on average across different numbers of available repeat read sequences.

[0072] 6C depicts a graph showing a comparison of the respective performance of the sequence denoising model 115 and conventional multi-sequence alignment-based denoising techniques when denoising antibody heavy chains sequenced by

[0023] In the example shown in Figure 6C, the sequence denoising model 115 and conventional multi-sequence alignment-based denoising techniques were used to denoise 100,000 scaffold sequences with an average length of 362 base pairs and an average of 8 repeat read sequences per scaffold sequence. Despite the antibody heavy chain exhibiting more mutations and therefore being more difficult to denoise, the sequence denoising model 115 still provided superior performance, especially when more repeat sequences per scaffold sequence were available (e.g., lower edit distance relative to the original scaffold sequence).

[0073] In view of the foregoing embodiments of the subject matter, the present application discloses the following list of examples, which are further examples that may be included in the disclosure of this application by combining one feature of a single example or two or more features of said examples, optionally in combination with one or more features of one or more additional examples.

[0074] Item 1: A computer-implemented method comprising: receiving a plurality of read sequences associated with a scaffold sequence, wherein each read sequence of the plurality of read sequences comprises a repeat of the scaffold sequence, and at least one read sequence of the plurality of read sequences is a noisy read sequence that does not match the scaffold sequence; determining encodings associated with the plurality of read sequences and applying a sequence denoising model trained to generate denoised scaffold sequences corresponding to the scaffold sequences based at least on the encodings associated with the plurality of read sequences; and analyzing molecules associated with the scaffold sequences based at least on the denoised scaffold sequences.

[0075] Term 2: The sequence denoising model generates denoised scaffold sequences without scaffold sequences, the method of Term 1.

[0076] Term 3: The sequence denoising model generates a denoised scaffold sequence by at least encoding multiple read sequences to generate multiple embeddings, determining an aggregate embedding corresponding to an aggregation of the multiple embeddings, and decoding the aggregate embedding to generate a denoised scaffold sequence.

[0077] Term 4: The method of Term 3, wherein the sequence denoising model includes a sequence encoder that is trained to generate multiple embeddings by at least encoding multiple read sequences.

[0078] Term 5: The method of Term 4, in which the sequence encoder is trained to increase the similarity between multiple embeddings when encoding multiple read sequences.

[0079] Term 6: The method of Term 5, wherein the sequence encoder increases the similarity between the multiple embeddings by at least decreasing the edit distance between the multiple embeddings when encoding the multiple read sequences.

[0080] Term 7: The edit distance involves the kernelized maximum mean discrepancy (MMD),method of Term 6.

[0081] Term 8: The method of any of Terms 4 to 7, wherein the sequence encoder is trained to generate, for each lead sequence of the multiple lead sequences, a corresponding embedding in the latent space.

[0082] Term 9: The method of Term 8, in which the latent space is occupied by a reduced-dimensional representation of multiple lead sequences.

[0083] Term 10: Any of the methods of terms 8 to 9, where the latent space includes a topological space occupied by multiple embeddings.

[0084] Item 11: Any of the methods of items 8 to 10, where the latent space contains a manifold.

[0085] Item 12: Any of the methods of items 4 to 11, wherein the sequence encoder is a transformer.

[0086] Term 13: Any of the methods of terms 3 to 12, wherein the sequence denoising model includes a sequence decoder that is trained to decode the aggregate embeddings to generate denoised scaffold sequences.

[0087] Term 14: The method of Term 13, wherein decoding the aggregate embedding includes converting the aggregate embedding from its latent space representation to a denoised scaffold array in array space.

[0088] Clause 15: Any of the methods of clauses 13 to 14, wherein the array decoder includes a transformer.

[0089] Item 16: Any of the methods of items 3 to 15, wherein the sequence denoising model includes a set encoder that is trained to determine an aggregate embedding.

[0090] Item 17: The method of Item 16, in which the aggregate embedding determined by the set encoder is invariant to the ordering of multiple read sequences.

[0091] Item 18: Any of the methods of Items 16 to 17, wherein the set encoder is trained to reduce a first distance between the aggregate embedding and the multiple embeddings in the latent space when determining the midpoint embedding.

[0092] Item 19: The method of item 18, wherein the set encoder is further trained to reduce a second distance between the denoised scaffold sequence and the plurality of lead sequences in the sequence space when determining the aggregate embedding.

[0093] Clause 20: Any of the methods of clauses 16 to 19, wherein the set encoder includes a set transformer.

[0094] Clause 21: The method of any of clauses 1 to 20, wherein the noisy read sequence comprises at least one insertion, deletion, or substitution of a nucleobase type contained in the scaffold sequence.

[0095] Paragraph 22: The method of any of paragraphs 1 to 21, wherein the scaffold sequence is a nucleic acid sequence label that is conjugated to at least a portion of an antigen or antibody.

[0096] Clause 23: The method of any of clauses 1 to 22, wherein the scaffold sequence is a nucleic acid sequence tag conjugated to a heavy or light chain of a B cell receptor.

[0097] Paragraph 24: The method of any of paragraphs 1 to 23, wherein the scaffold sequence is a nucleic acid sequence tag conjugated to the alpha and / or beta chain of a T cell receptor.

[0098] Clause 25: The method of any of clauses 1 to 24, wherein the scaffold sequence is a deoxyribonucleic acid (DNA) sequence or a ribonucleic acid (RNA) sequence.

[0099] Clause 26: The method of any of clauses 1 to 25, further comprising: training a sequence denoising model to generate one or more denoised scaffold sequences while one or more corresponding scaffold sequences are unknown based on at least the training set; determining performance of the sequence denoising model while one or more corresponding scaffold sequences are unknown; and, upon determining that the performance of the sequence denoising model satisfies one or more thresholds, applying the sequence denoising model to determine encodings associated with the plurality of read sequences and generating denoised scaffold sequences corresponding to the scaffold sequences based at least on the encodings associated with the plurality of read sequences.

[0100] Term 27: The method of Term 26, wherein the performance of the sequence denoising model is determined based at least on a leave-one-out edit distance of the sequence denoising model.

[0101] Clause 28: The method of Clause 27, further comprising determining a leave-one-out edit distance associated with the sequence denoising model by at least determining an edit distance between one lead sequence from the training set and a denoised scaffold sequence generated by the sequence denoising model based on one or more remaining lead sequences in the training set.

[0102] Clause 29: A system comprising at least one data processor and at least one memory storing instructions that, when executed by the at least one data processor, result in operations including the method of any of clauses 1 to 28.

[0103] Clause 30: A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, result in operations including the method of any of clauses 1 to 28.

[0104] 7 depicts a block diagram of an example computing system 700, according to some exemplary embodiments. Referring to FIGS. 1A and 7, computing system 700 may be used to implement sequencing controller 110, sequencing platform 120, analysis engine 130, client 140, and / or any components therein.

[0105] 7, computing system 700 may include a processor 710, a memory 720, a storage device 830, and an input / output device 740. The processor 710, the memory 720, the storage device 830, and the input / output device 740 may be interconnected via a system bus 750. The processor 710 may process instructions for execution within the computing system 700. Such executed instructions may implement one or more components, such as, for example, the EV profile analysis engine 110, the sequencing platform 120, the analysis engine 130, the client 140, etc. In some exemplary embodiments, the processor 710 may be a single-threaded processor. Alternatively, the processor 710 may be a multi-threaded processor. The processor 710 may process instructions stored in the memory 720 and / or the storage device 830 to display graphical information for a user interface provided via the input / output device 740.

[0106] Memory 720 is a computer-readable medium, such as a volatile or non-volatile medium, that stores information within computing system 700. Memory 720 may store, for example, data structures representing a configuration object database. Storage device(s) 830 may provide persistent storage for computing system 700. Storage device(s) 830 may be a floppy disk drive, a hard disk drive, an optical disk drive, or a tape drive, or other suitable persistent storage means. Input / output device(s) 740 provide input / output operations to computing system 700. In some exemplary embodiments, input / output device(s) 740 include a keyboard and / or a pointing device. In various embodiments, input / output device(s) 740 include a display unit for displaying a graphical user interface.

[0107] According to some exemplary embodiments, input / output devices 740 may provide input / output operations to network devices. For example, input / output devices 740 may include an Ethernet port or other networking port for communicating with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0108] In some exemplary embodiments, computing system 700 can be used to execute various interactive computer software applications that can be used to organize, analyze, and / or store data in various formats. Alternatively, computing system 700 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., creating, managing, editing spreadsheet documents, word processing documents, and / or any other objects), computing functions, communication functions, etc. Applications can include various add-in functions or can be standalone computing products and / or functions. When activated within an application, functions can be used to generate a user interface that is provided via input / output devices 740. The user interface can be generated by computing system 700 and presented to a user (e.g., on a computer screen monitor, etc.).

[0109] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs, field programmable gate array (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0110] These computer programs, which may also be referred to as programs, software, software applications, applications, components, or code, contain machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs), and includes a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium may non-transitory store such machine instructions, such as, for example, a non-transitory solid-state memory or a magnetic hard drive or any equivalent storage medium. Alternatively or additionally, a machine-readable medium may temporarily store such machine instructions, such as, for example, a processor cache or other random access memory associated with one or more physical processor cores.

[0111] To provide for user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having, for example, a display device, such as a cathode ray tube (CRT) or liquid crystal display (LCD) or light-emitting diode (LED) monitor, for displaying information to a user, and a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, voice, or tactile input. Other possible input devices include touchscreens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software.

[0112] In the above description and in the claims, phrases such as "at least one of" or "one or more of" may appear followed by a list of conjunctive elements or features. The term "and / or" may also appear with a list of two or more elements or features. Unless implicitly or explicitly contradicted by the context in which it is used, such phrases are intended to refer to any of the listed elements or features individually, or any of the listed elements or features in combination with any of the other listed elements or features. For example, the phrases "at least one of A and B;," "one or more of A and B;," and "A and / or B" are intended to mean "A alone, B alone, or A and B together," respectively. A similar interpretation is intended for lists containing more than two items. For example, the phrases "at least one of A, B, and C;," "one or more of A, B, and C;," and "A, B, and / or C" are intended to mean "A alone, B alone, C alone, A and B, A and C, B and C, or A, B, and C," respectively. Use of the term "based on" above and in the claims is intended to mean "based at least in part on," allowing for unrecited features or elements.

[0113] The subject matter described herein may be implemented in systems, devices, methods, and / or articles, depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, these implementations are merely some examples consistent with aspects related to the described subject matter. While several variations have been described in detail above, other modifications or additions are possible. In particular, additional features and / or variations may be provided in addition to those described herein. For example, the implementations described above may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several additional features described above. Furthermore, the logic flow illustrated in the accompanying drawings and / or described herein does not necessarily require the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the following claims.

Claims

1. receiving a plurality of read sequences associated with a scaffold sequence, wherein each read sequence of the plurality of read sequences includes a repeat of the scaffold sequence, and at least one read sequence of the plurality of read sequences is a noisy read sequence that does not match the scaffold sequence; determining encodings associated with the plurality of read sequences and applying a sequence denoising model trained to generate denoised scaffold sequences corresponding to the scaffold sequences based at least on the encodings associated with the plurality of read sequences; analyzing molecules associated with the scaffold sequences based at least on the denoised scaffold sequences; 10. A computer-implemented method comprising:

2. 10. The method of claim 1, wherein the sequence denoising model generates the denoised scaffold sequence without the scaffold sequence.

3. The sequence denoising model includes at least: encoding the plurality of read sequences to generate a plurality of embeddings; determining an aggregate embedding corresponding to an aggregation of the plurality of embeddings; decoding the aggregate embedding to generate the denoised scaffold sequence; 3. The method of claim 1 or 2, wherein the denoised scaffold sequences are generated by:

4. 4. The method of claim 3, wherein the sequence denoising model comprises a sequence encoder trained to generate the plurality of embeddings by at least encoding the plurality of read sequences, wherein the sequence encoder is trained to, when encoding the plurality of read sequences, increase similarity between the plurality of embeddings by at least decreasing an edit distance between the plurality of embeddings when encoding the plurality of read sequences.

5. 5. The method of claim 4, wherein the sequence encoder is trained to generate, for each lead sequence of the plurality of lead sequences, a corresponding embedding in a latent space.

6. The method of claim 5 , wherein the latent space is a topological space occupied by a reduced-dimensional representation of the plurality of lead sequences.

7. The method of claim 4 , wherein the sequence encoder is a transformer.

8. 8. The method of claim 3, wherein the sequence denoising model comprises a sequence decoder trained to generate the denoised scaffold sequences by at least decoding the aggregate embeddings to generate the denoised scaffold sequences.

9. The method of claim 8 , wherein the sequence decoder comprises a transformer.

10. The method of claim 3 , wherein the sequence denoising model comprises a set encoder that is trained to determine the aggregate embedding.

11. 11. The method of claim 10, wherein the set encoder is trained to decrease a first distance between the aggregate embedding and the plurality of embeddings in a latent space when determining the aggregate embedding, and wherein the set encoder is further trained to decrease a second distance between the denoised scaffold sequence and the plurality of lead sequences in a sequence space when determining the aggregate embedding.

12. The method of claim 10 or 11, wherein the set encoder comprises a set transformer.

13. 13. The method of any one of claims 1 to 12, wherein the noisy read sequence comprises at least one insertion, deletion, or substitution of a nucleobase type contained in the scaffold sequence.

14. determining the identity of the molecule associated with the scaffold sequence based at least on the denoised scaffold sequence; 14. The method of any one of claims 1 to 13, further comprising:

15. determining the binding specificity of the antibody based at least on the identity of the molecule conjugated to the antibody; The method of claim 14 further comprising:

16. Identifying a gene that expresses a T cell receptor (TCR) based at least on the identity of a molecule that is conjugated to the T cell receptor (TCR); The method of claim 14 further comprising:

17. 17. The method of any one of claims 1 to 16, wherein the scaffold sequence is associated with a nucleic acid sequence tag conjugated to (i) a heavy or light chain of a B cell receptor, or (ii) the alpha and / or beta chain of a T cell receptor.

18. 18. The method of any one of claims 1 to 17, wherein the scaffold sequence is a deoxyribonucleic acid (DNA) sequence or a ribonucleic acid (RNA) sequence.

19. training the sequence denoising model to generate one or more denoised scaffold sequences based on at least a training set, while one or more corresponding scaffold sequences are unknown; determining the performance of the sequence denoising model while the one or more corresponding scaffold sequences are unknown; and applying the sequence denoising model to determine the encodings associated with the plurality of read sequences and generating the denoised scaffold sequences corresponding to the scaffold sequences based at least on the encodings associated with the plurality of read sequences when the performance of the sequence denoising model is determined to meet one or more thresholds; 19. The method of any one of claims 1 to 18, further comprising:

20. determining a leave-one-out edit distance of the sequence denoising model by determining at least an edit distance between one lead sequence from the training set and a denoised scaffold sequence generated by the sequence denoising model based on one or more remaining lead sequences in the training set; determining the performance of the sequence denoising model based at least on the leave-one-out edit distance; 20. The method of claim 19 further comprising:

21. at least one data processor; at least one memory storing instructions that, when executed by said at least one data processor, result in operations comprising the method of any one of claims 1 to 20; A system comprising:

22. A non-transitory computer readable medium storing instructions that, when executed by at least one data processor, result in operations comprising the method of any one of claims 1 to 20.