Systems and methods for training machine learning algorithms to process biologically relevant data, microscopes, and trained machine learning algorithms

By using language recognition and visual recognition machine learning algorithms to generate high-dimensional representations in biological applications, the problem of time-consuming and cost-effective processing of biological data is solved, and efficient semantic classification of data is achieved.

CN114450751BActive Publication Date: 2025-05-02LEICA MICROSYSTEMS CMS GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980099039.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-07
Publication Date
2025-05-02
Estimated Expiration
2039-06-07

AI Technical Summary

Technical Problem

In biological applications, manual analysis of large amounts of biological data generated is time-consuming and costly, and it is difficult for the prior art to effectively process these data.

Method used

Design a system, including processors and storage devices, through language recognition machine learning algorithms and visual recognition machine learning algorithms, generate high-dimensional representations of biologically relevant data, and adjust the algorithm according to comparisons to improve the possibility of semantically correct classification of data.

Benefits of technology

By generating high-dimensional representations, the system can significantly improve the possibility of semantically correct classification of biological related data, reduce the time and cost of manual analysis, and achieve efficient processing of large amounts of biological data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114450751B_ABST
    Figure CN114450751B_ABST
Patent Text Reader

Abstract

The system (100) includes one or more processors (110) and one or more storage devices (120), wherein the system (100) is configured to generate a first high-dimensional representation of input training data (102) based on a biologically relevant language by a language recognition machine learning algorithm executed by the one or more processors (110). In addition, the system (100) is configured to generate output training data based on a biologically relevant language based on the first high-dimensional representation by the language recognition machine learning algorithm, and adjust the language recognition machine learning algorithm based on a comparison of the input training data (102) based on the biologically relevant language and the output training data based on the biologically relevant language. In addition, the system (100) is configured to generate a second high-dimensional representation of input training data (104) based on a biologically relevant image by a visual recognition machine learning algorithm executed by the one or more processors (110), and adjust the visual recognition machine learning algorithm based on a comparison of the first high-dimensional representation and the second high-dimensional representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Examples involve processing of biologically relevant data. Background Art

[0002] In many biological applications, large amounts of data are generated. For example, images are collected from a large number of biological structures and stored in a database. Manual analysis of biological data is time-consuming and costly. Summary of the invention

[0003] Therefore, there is a need for improved concepts for processing biologically relevant data.

[0004] This need may be met by the claimed subject matter.

[0005] Some embodiments relate to a system comprising one or more processors and one or more storage devices. The system is configured to receive input training data based on a biologically relevant language, and generate a first high-dimensional representation of the input training data based on the biologically relevant language through a language recognition machine learning algorithm executed by the one or more processors. The first high-dimensional representation includes at least 3 entries, each entry having a different value. In addition, the system is configured to generate output training data based on the biologically relevant language according to the first high-dimensional representation through a language recognition machine learning algorithm executed by the one or more processors, and adjust the language recognition machine learning algorithm based on a comparison of the input training data based on the biologically relevant language and the output training data based on the biologically relevant language. In addition, the system is configured to receive input training data based on a biologically relevant image associated with the input training data based on the biologically relevant language, and generate a second high-dimensional representation of the input training data based on the biologically relevant image through a visual recognition machine learning algorithm executed by the one or more processors. The second high-dimensional representation includes at least three entries, each entry having a different value. In addition, the system is configured to adjust the visual recognition machine learning algorithm based on a comparison of the first high-dimensional representation and the second high-dimensional representation.

[0006] By using a language recognition machine learning algorithm, textual biological input can be mapped to a high-dimensional representation. By making the high-dimensional representation have entries with various different values ​​(as opposed to a one-hot encoding representation), semantically similar biological input can be mapped to similar high-dimensional representations. By training a visual recognition machine learning algorithm to map images to high-dimensional representations trained by a language recognition machine learning algorithm, images with similar biological content can also be mapped to similar high-dimensional representations. Therefore, the possibility of semantically correct classification or at least semantically close classification of images by the corresponding trained visual recognition machine learning algorithm can be significantly improved. In addition, the corresponding trained visual recognition machine learning algorithm can more accurately map untrained images to high-dimensional representations close to high-dimensional representations with similar meanings or semantically matching high-dimensional representations. Through the proposed concept, a trained language recognition machine learning algorithm and / or a trained visual recognition machine learning algorithm can be obtained, which can provide semantically correct or very accurate classification based on biologically relevant language and / or image-based input data. The trained language recognition machine learning algorithm and / or the trained visual recognition machine learning algorithm can be used to retrieve biologically relevant images from multiple biological images based on language-based retrieval input or image-based retrieval input, labeling of biologically relevant images, finding or generating typical images, and / or similar applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Some examples of apparatus and / or methods will be described below by way of example only and with reference to the accompanying drawings, in which:

[0008] Figure 1 is a schematic diagram of a system for training machine learning algorithms to process biologically relevant data;

[0009] Figure 2 This is a schematic diagram of the training of a language recognition machine learning algorithm;

[0010] Figure 3 This is a schematic diagram of the training of a visual recognition machine learning algorithm;

[0011] Figure 4 This is a computational graph of a portion of a visual recognition neural network based on the ResNet architecture.

[0012] Figure 5 is a computational graph of a portion of a visual recognition neural network based on the ResNet architecture with a modified CBAM block;

[0013] Figure 6 A computational graph of a portion of a visual recognition neural network based on a densely connected convolutional network (DenseNet) architecture;

[0014] Figure 7Computational graph of a portion of a visual recognition neural network based on the DenseNet architecture with an attention mechanism;

[0015] Figure 8 is a schematic diagram of a system for training a machine learning algorithm to process biologically relevant data; and

[0016] Fig. 9 is a flow chart of a method for training a machine learning algorithm to process biologically relevant data. DETAILED DESCRIPTION

[0017] Various examples will now be described more fully with reference to the accompanying drawings in which some examples are shown. In the drawings, the thickness of lines, layers and / or regions may be exaggerated for clarity.

[0018] Accordingly, although other examples can have various modifications and alternative forms, some specific examples thereof are shown in the drawings and will be described in detail later. However, this detailed description does not limit other examples to the specific forms described. Other examples may encompass all modifications, equivalents, and substitutes falling within the scope of the present disclosure. Throughout the description of the drawings, identical or similar reference numerals refer to identical or similar elements, which may be implemented in identical or modified forms when compared to each other, while providing identical or similar functions.

[0019] It will be understood that when an element is referred to as being "connected" or "coupled" to another element, these elements may be directly connected or coupled, or connected or coupled via one or more intermediate elements. Unless explicitly or implicitly defined otherwise, if two elements A and B are combined using "or", this will be understood to disclose all possible combinations, i.e., only A, only B, and A and B. Alternative terms for the same combination are "at least one of A and B" or "A and / or B". This also applies, mutatis mutandis, to combinations of more than two elements.

[0020] The terms used herein to describe specific examples are not intended to limit additional examples. Whenever singular forms such as "a", "an", and "the" are used and the use of only a single element is not explicitly or implicitly defined as mandatory, additional examples may also use plural elements to implement the same function. Similarly, when a function is later described as being implemented using multiple elements, additional examples may use a single element or processing entity to implement the same function. It will be further understood that when used, the terms "comprise", "comprising", "includes", and / or "including" indicate the presence of the features, integers, steps, operations, processes, actions, elements, and / or components described, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, actions, elements, components, and / or any groups thereof.

[0021] Unless otherwise defined, all terms (including technical and scientific terms) are used herein according to the ordinary meaning of the art to which their examples belong.

[0022] Figure 1A schematic diagram of a system 100 for training a machine learning algorithm to process biologically relevant data according to an embodiment is shown. The system 100 includes one or more processors 110 and one or more storage devices 120. The system 100 is configured to receive input training data 102 based on a biologically relevant language. In addition, the system 100 is configured to generate a first high-dimensional representation of the input training data 102 based on the biologically relevant language through a language recognition machine learning algorithm executed by the one or more processors 110. The first high-dimensional representation includes at least 3 entries, each entry having a different value, (or including at least 20 entries, at least 50 entries, or at least 100 entries having different values ​​from each other). Further, the system 100 is configured to generate output training data based on the biologically relevant language according to the first high-dimensional representation through a language recognition machine learning algorithm executed by the one or more processors 110. In addition, the system 100 is configured to adjust the language recognition machine learning algorithm based on a comparison of the input training data 102 based on the biologically relevant language and the output training data based on the biologically relevant language. In addition, the system 100 is configured to receive biologically relevant image-based input training data 104 associated with the biologically relevant language-based input training data 102. In addition, the system 100 is configured to generate a second high-dimensional representation of the biologically relevant image-based input training data 104 by a visual recognition machine learning algorithm executed by one or more processors 110. The second high-dimensional representation includes at least three entries, each entry having a different value, (or includes at least 20 entries, at least 50 entries, or at least 100 entries having values ​​that are different from each other). Further, the system 100 is configured to adjust the visual recognition machine learning algorithm based on a comparison of the first high-dimensional representation and the second high-dimensional representation.

[0023] The input training data 102 based on biology-related language can be textual input related to biological structure, biological function, biological behavior or biological activity. For example, the input training data 102 based on biology-related language can be a nucleotide sequence, a protein sequence, a description of a biological molecule or biological structure, a description of the behavior of a biological molecule or biological structure, and / or a description of a biological function or biological activity. The textual input can be a natural language that describes the behavior of a biological molecule (e.g., a polysaccharide, a poly / oligonucleotide, a protein or a lipid) or a biological molecule in the context of an experiment or a data set. The textual input can also be a text such as a nucleotide sequence, a protein sequence or a controlled query language. For example, the input training data 102 based on biology-related language can be a nucleotide sequence or a protein sequence, because a large number of different sequences are known and can be obtained from a database and / or the biological functions and / or biological activities of these sequences are known. The input training data 102 based on biology-related language can include a length of more than 20 characters (or more than 40 characters, more than 60 characters or more than 80 characters). For example, since three base pairs are encoded into one amino acid, nucleotide sequences (DNA / RNA) are generally about three times longer than polypeptide sequences (e.g., peptides, proteins). For example, if the input training data based on the biology-related language is a protein sequence or an amino acid, the input training data 102 based on the biology-related language may include a length of more than 20 characters. If the input training data based on the biology-related language is a nucleotide sequence or a descriptive text in a natural language, the input training data 102 based on the biology-related language may include a length of more than 60 characters. For example, the input training data 102 based on the biology-related language may include at least one non-numeric character (e.g., an alphabetic character). The input training data 102 based on the biology-related language may also be referred to as a token or an input token. The input training data 102 based on the biology-related language may be received from one or more storage devices 120, a database stored by a storage device, or may be input by a user. The input training data based on the biology-related language may be the first input training data set based on the biology-related language of the training group (e.g., a sequence of input characters (e.g., a nucleotide sequence or a protein sequence)). The training group may include multiple input training data sets based on the biology-related language.

[0024] The output training data based on the biology-related language may be of the same type as the input training data 102 based on the biology-related language that optionally includes a prediction of the next element. For example, the input training data 102 based on the biology-related language may be a biological sequence (e.g., a nucleotide sequence or a protein sequence), and the output training data based on the biology-related language may also be a biological sequence (e.g., a nucleotide sequence or a protein sequence). The language recognition machine learning algorithm may be trained so that the output training data based on the biology-related language is the same as the input training data 102 based on the biology-related language that optionally includes a prediction of the next element of the biological sequence. In another example, the input training data 102 based on the biology-related language may be a biological category of a coarse-grained search term, and the output training data based on the biology-related language may also be a biological category of a coarse-grained search term.

[0025] Alternatively, the output training data based on the biology-related language is of a different type from the input training data 102 based on the biology-related language. For example, the input training data 102 based on the biology-related language is a biological sequence (e.g., a nucleotide sequence or a protein sequence), and the output training data based on the biology-related language is a biological category of a coarse-grained search term. In this example, each biological sequence used as the input training data 102 may belong to a coarse-grained search term of a set of biological terms, and the language recognition machine learning algorithm may be trained to classify each biological sequence used as the input training data into a corresponding coarse-grained search term of the set of biological terms.

[0026] A set of biological terms can include a plurality of coarse-grained search terms (or alternatively referred to as molecular biology subject terms) belonging to the same biological subject. A set of biological terms can be catalytic activity (e.g., as a certain reaction formula using words for educts and products), pathway (e.g., which pathway is involved, e.g., glycolysis), site and / or region (e.g., binding site, active site, nucleotide binding site), GO gene ontology (e.g., molecular function, such as nicotinamide adenine dinucleotide NAD binding, microtubule binding), GO biological function (e.g., apoptosis, gluconeogenesis), enzyme and / or pathway database (e.g., unique identifier for sic function such as in BRENDA / EC numbering or UniPathways), subcellular localization (e.g., cytoplasm, nucleus, cytoskeleton), family and / or domain (e.g., binding site, motif such as for posttranslational modification), open reading frame, single nucleotide polymorphism, restriction site (e.g., oligonucleotide recognized by restriction enzyme) and / or biosynthetic pathway (e.g., biosynthesis of lipids, polysaccharides, nucleotides or proteins). For example, the set of biological terms may be the subcellular localization set, and the coarse-grained search terms may be cytoplasm, nucleus, and cytoskeleton.

[0027] Output training data based on a biologically relevant language can be generated by a decoder of a language identification machine learning algorithm. For example, output training data based on a biologically relevant language can be generated by applying a language identification machine learning algorithm with a current set of parameters (e.g., neural network weights) to generate a first high-dimensional representation. The current set of parameters of the language identification machine learning algorithm can be updated during tuning of the language identification machine learning algorithm.

[0028] The input training data 104 based on biologically relevant images can be image training data (e.g., pixel data of training images) of images of: biological structures including nucleotides or nucleotide sequences; biological structures including proteins or protein sequences; biological molecules; biological tissues; biological structures with specific behaviors; and / or biological structures with specific biological functions or specific biological activities. The biological structure can be a molecule, a viroid or a virus, an artificial or natural membrane-encapsulated vesicle, a subcellular structure (such as an organelle), a cell, a spheroid, an organoid, a three-dimensional cell culture, a biological tissue, an organ slice, or a portion of an organ in vivo or in vitro. For example, the image of the biological structure can be an image of the location of a protein within a cell or tissue, or can be an image of a cell or tissue with endogenous nucleotides (e.g., DNA) to which a labeled nucleotide probe binds (e.g., in situ hybridization). The image training data can include pixel values ​​for each pixel of the image for each color dimension of the image (e.g., three color dimensions for RGB representation). For example, depending on the imaging modality, other channels can be adapted to be associated with excitation or emission wavelength, fluorescence lifetime, light polarization, stage position in three spatial dimensions, and different imaging angles. The input training data 104 based on biologically relevant images can be an XY pixel map, volume data (XYZ), time series data (XY+T), or a combination thereof (XYZT). In addition, additional dimensions depending on the type of image source may be included, such as channels (e.g., spectral emission bands), excitation wavelengths, stage positions, logical positions such as in multi-well plates or multi-position experiments, and / or reflector and / or objective lens positions such as in light sheet imaging. For example, a user may input an image as a pixel map or a higher dimensional picture, or a database may provide an image as a pixel map or a higher dimensional picture. The visual recognition machine learning algorithm may convert this image into a semantic embedding (e.g., a second high dimensional representation). For example, the input training data 104 based on biologically relevant images corresponds to the input training data 102 based on biologically relevant languages. For example, the input training data based on biologically relevant images represents a biological structure described by the input training data 102 based on biologically relevant languages, so that the input training data 104 based on biologically relevant images is associated with the input training data 102 based on biologically relevant languages. The biologically relevant image based input training data 104 may be received from one or more storage devices and a database stored by the storage device, or may be input by a user. The biologically relevant image based input training data 104 may be the first biologically relevant image based input training data set of a training set. The training set may include multiple biologically relevant image based input training data sets.

[0029] The high-dimensional representation (e.g., the first high-dimensional representation and the second high-dimensional representation) can be a hidden representation, a latent vector, an embedding, a semantic embedding and / or a token embedding, and / or can also be referred to as a hidden representation, a latent vector, an embedding, a semantic embedding and / or a token embedding.

[0030] The first high-dimensional representation and / or the second high-dimensional representation can be a digital representation (for example, only including numerical values). The first high-dimensional representation and / or the second high-dimensional representation can only include positive values ​​or entries with positive values ​​and entries with negative values. In contrast, input training data based on biology-related languages ​​can only include alphabetic characters or other non-numeric characters, or include a mixture of alphabetic characters, other non-numeric characters and / or numeric characters. The first high-dimensional representation and / or the second high-dimensional representation can include more than 100 dimensions (or more than 300 dimensions or more than 500 dimensions) and / or less than 10000 dimensions (or less than 3000 dimensions or less than 1000 dimensions). Each entry of the high-dimensional representation can be a dimension of the high-dimensional representation (for example, the high-dimensional representation with 100 dimensions includes 100 entries). For example, the use of a high-dimensional representation with more than 300 dimensions and less than 1000 dimensions can achieve a suitable representation of biology-related data with semantic relevance. The first high-dimensional representation can be a first vector and the second high-dimensional representation can be a second vector. If vector representations are used for the entries of the first high-dimensional representation and the entries of the second high-dimensional representation, efficient comparisons and / or other calculations (e.g., normalization) can be implemented, but other representations (e.g., representations as matrices) are also feasible. For example, the first high-dimensional representation and / or the second high-dimensional representation can be normalized vectors. The first high-dimensional representation and the second high-dimensional representation can be normalized to the same value (e.g., 1). For example, the last layer of a model (e.g., a model of a language recognition machine learning algorithm and / or a visual recognition machine learning algorithm) can represent a non-linear operation that can additionally perform normalization. For example, if the first model (language model) is trained using a cross-entropy loss function, a so-called SoftMax operation can be used:

[0031]

[0032] Among them, y i is the prediction of the model for the corresponding input value, and K is the number of all input values.

[0033] For example, compared to a one hot encoded representation, the first high dimensional representation and / or the second high dimensional representation may include various entries (at least three entries) whose values ​​are not equal to 0. By using a high dimensional representation that allows various entries with values ​​not equal to 0, information about the semantic relationship between the high dimensional representations can be reproduced. For example, the values ​​of more than 50% (or more than 70% or more than 90%) of the entries of the first high dimensional representation and / or the values ​​of more than 50% (or more than 70% or more than 90%) of the entries of the second high dimensional representation may not be equal to 0. Sometimes, the one hot encoded representation also has more than one entry that is not equal to 0, but only one entry has a high value, while the values ​​of all other entries are at a noise level (e.g., less than 10% of the one high value). On the contrary, for example, the values ​​of more than 5 entries (or more than 20 entries or more than 50 entries) of the first high dimensional representation may be 10% (or more than 20% or more than 30%) greater than the maximum absolute value of the entries of the first high dimensional representation. Further, for example, the values ​​of more than 5 entries (or more than 20 entries or more than 50 entries) of the second high-dimensional representation may be 10% (or 20% or 30%) greater than the maximum absolute value of the entries of the second high-dimensional representation. For example, each entry of the first high-dimensional representation and / or the second high-dimensional representation may include a value between -1 and 1.

[0034] The first high-dimensional representation can be generated by an encoder of a language recognition machine learning algorithm. For example, the first high-dimensional representation is generated by applying a language recognition machine learning algorithm with a current parameter set to input training data 102 based on a biologically relevant language. The current parameter set of the language recognition machine learning algorithm can be updated during adjustment of the language recognition machine learning algorithm. For example, the adjustment of the language recognition machine learning algorithm includes adjusting multiple language recognition neural network weights, and the final language recognition neural network weight set can be stored by one or more storage devices 120. Further, a second high-dimensional representation can be generated by applying a visual recognition machine learning algorithm with a current parameter set to input training data based on biologically relevant images. The current parameter set of the visual recognition machine learning algorithm can be updated during adjustment of the visual recognition machine learning algorithm. For example, the adjustment of the visual recognition machine learning algorithm includes adjusting multiple visual recognition neural network weights, and the final visual neural network weight set can be stored by one or more storage devices 120.

[0035] The value of one or more entries of the first high-dimensional representation and / or the value of one or more entries of the second high-dimensional representation can be proportional to the possibility of the presence of a specific biological function or a specific biological activity. By using a mapping that generates a high-dimensional representation that retains the semantic similarity of the input data set, semantically similar high-dimensional representations can have a closer distance to each other than high-dimensional representations that are less semantically similar. Further, if two high-dimensional representations represent input data sets with the same or similar specific biological functions or specific biological activities, one or more entries in the two high-dimensional representations can have the same or similar values. Due to the retention of semantics, one or more entries of the high-dimensional representation can indicate the presence or existence of a specific biological function or a specific biological activity. For example, the higher the value of one or more entries of the high-dimensional representation, the higher the possibility of the presence of a biological function or biological activity associated with these one or more entries.

[0036] The system 100 can repeatedly generate a first high-dimensional representation for each of the multiple biologically-relevant language-based input training data sets of the training group. In addition, the system 100 can generate output training data based on the biologically-relevant language for each generated first high-dimensional representation. The system 100 can adjust the language recognition machine learning algorithm according to each comparison of the input training data based on the biologically-relevant language of the multiple biologically-relevant language-based input training data sets of the training group with the corresponding output training data based on the biologically-relevant language. In other words, the system 100 can be configured to repeatedly generate a first high-dimensional representation, generate output training data based on the biologically-relevant language, and adjust the language recognition machine learning algorithm for each biologically-relevant language-based input training data set of the training group based on the biologically-relevant language input training data set. The training group can include enough biologically-relevant language-based input training data sets so that the training goal (e.g., the output change of the loss function is below a threshold) can be achieved.

[0037] The multiple first high-dimensional representations generated during the training of a language identification machine learning algorithm can be called a latent space or a semantic space.

[0038] The system 100 may repeatedly generate a second high-dimensional representation for each biologically-relevant image-based input training data set in a training group of multiple biologically-relevant image-based input training data sets. Further, the system 100 may adjust the visual recognition machine learning algorithm based on each comparison of the first high-dimensional representation with the corresponding second high-dimensional representation. In other words, the system 100 may repeatedly generate a second high-dimensional representation and adjust the visual recognition machine learning algorithm for each biologically-relevant image-based input training data set in a training group of biologically-relevant image-based input training data sets. The training group may include enough biologically-relevant image-based input training data sets so that the training goal (e.g., the output change of the loss function is below a threshold) can be achieved.

[0039] The training set of the biologically relevant language based input training data set may include more entries than the training set of the biologically relevant image based input training data set. For example, if the biologically relevant language based input training data set is different nucleotide sequences or protein sequences, a database with more different nucleotide sequences or protein sequences than images of biological structures including corresponding nucleotides or corresponding proteins may be used for training. Further, if the number of trained first high-dimensional representations is greater than the number of trained second high-dimensional representations, zero-shot learning of untrained biologically relevant image based input data is possible. The trained visual recognition machine learning algorithm can map unseen biologically relevant image based input data to a second high-dimensional representation that is close to one or more semantically similar first high-dimensional representations of the biologically relevant language based input data. Alternatively, for example, if the biology-related language based input training dataset is a description of different behaviors of biological molecules or biological structures or a description of biological functions or biological activities, the training group of the biology-related language based input training dataset may include fewer entries than the training group of the biology-related image based input training dataset, because the number of different input datasets for these kinds of input data may be limited (e.g., less than 500 or less than 100 or less than 50 different biology-related language based input training datasets).

[0040] For example, system 100 uses a combination of a language recognition machine learning algorithm and a visual recognition machine learning algorithm (e.g., also referred to as a visual semantic model). The language recognition machine learning algorithm and / or the visual recognition machine learning algorithm may be a deep learning algorithm and / or an artificial intelligence algorithm.

[0041] The language identification machine learning algorithm may also be referred to as a textual model, a language model or a linguistic model. The language identification machine learning algorithm may be or may include a language identification neural network. The language identification neural network may include more than 30 layers (or more than 50 layers or more than 80 layers) and / or less than 500 layers (or less than 300 layers or less than 200 layers). The language identification neural network may be a recursive neural network, such as a long short-term memory network. Using a recursive neural network, such as a long short-term memory network, a high-precision language identification machine learning algorithm may be provided for input data based on biologically relevant languages. However, other language identification algorithms may also be applied. For example, the language identification machine learning algorithm may be an algorithm (such as a Transformer-XL algorithm) that can process input data of variable length. For example, the length of the first input training data based on the biologically relevant language of the training set of the input training data set based on the biologically relevant language is different from the length of the second input training data based on the biologically relevant language of the training set of the input training data set based on the biologically relevant language. By using the algorithm as the Transformer-XL algorithm, the model is able to detect the structure of longer and variable-length sequences. The unique properties of Transformer-XL that distinguish it from other language model architectures using neural networks can be attributed to the ability to learn semantic dependencies over variable lengths because the hidden state of each segment being analyzed is reused to obtain the hidden state of the next segment. This state accumulation can allow for the establishment of cyclic semantic connections between consecutive segments. Therefore, long-term dependencies that encode biological functions can be captured. For example, in nucleotide sequences, during gene transcription, long stretches of DNA are excised (e.g., grafted), effectively connecting nucleotide sequences that were previously far apart. Using the Transformer-XL architecture can allow for the capture of those long-term dependencies. In addition, in protein sequences, consecutive secondary polypeptide structures (e.g., alpha helix or beta sheet) often form so-called "folds" (e.g., three-dimensional arrangements of secondary structures in space). These stacks can be part of protein subdomains, each of which has a unique biological function. Therefore, long-term semantic dependencies are important for correctly capturing the biological functions to be encoded in the semantic embedding. Other methods can only learn fixed-length dependencies, which limits the model's ability to learn correct semantics. For example, protein sequences are often tens to hundreds of amino acids long (where an amino acid is represented as a letter in a protein sequence). The "semantics" (e.g. biological functions of substrings in a sequence (called peptides, motifs, or domains in biology) can vary in length. Therefore, architectures that can accommodate variable length dependencies can be used, such as Transformer-XL.

[0042] The language recognition machine learning algorithm can be trained by adjusting the parameters of the language recognition machine learning algorithm based on a comparison of the input training data 102 based on the biologically relevant language and the output training data based on the biologically relevant language. For example, the network weights of the language recognition neural network can be adjusted based on the comparison. The adjustment of the parameters (e.g., network weights) of the language recognition machine learning algorithm can be performed with consideration of a loss function (e.g., a cross-entropy loss function). The loss function can generate an actual value that is the degree of equivalence between the prediction and the existing annotation. Training can change the internal degrees of freedom (e.g., the weights of the neural network) until the loss function is minimized. For example, the comparison of the input training data 102 based on the biologically relevant language and the output training data based on the biologically relevant language in order to adjust the language recognition machine learning algorithm can be based on the cross-entropy loss function. For example, if M>2 (e.g., multi-class classification), a separate loss for each class label in each observation can be calculated and the results summed:

[0043]

[0044] where M is the number of classes (e.g., nucleus, cytoplasm, plasma membrane, and mitochondria (in the case of organelles)), log is the natural logarithm, y is a binary indicator (0 or 1), and p is the predicted probability that observation o belongs to class c if class label c is the correct classification for observation o.

[0045] The training may converge quickly, and / or by training the language recognition machine learning algorithm using a cross entropy loss function (although other loss functions may also be used), the training may provide a trained algorithm for biologically relevant data.

[0046] The visual recognition machine learning algorithm may also be referred to as an image recognition model, a visual model, or an image classifier. The visual recognition machine learning algorithm may be or may include a visual recognition neural network. The visual recognition neural network may include more than 20 layers (or more than 40 layers or more than 80 layers) and / or less than 400 layers (or less than 200 layers or less than 150 layers). The visual recognition neural network may be a convolutional neural network or a capsule network. Using a convolutional neural network or a capsule network, a visual recognition machine learning algorithm with high accuracy can be provided for input data based on biologically relevant images. However, other visual recognition algorithms may also be used. For example, a visual recognition neural network may include multiple convolutional layers and multiple pooling layers. However, for example, if a capsule network is used and / or stride=2 is used instead of stride=1 for convolution, the pooling layer may be avoided. The visual recognition neural network may use a rectified linear unit activation function. Using a rectified linear unit activation function, a highly accurate trained visual recognition machine learning algorithm may be provided for input data based on biologically relevant images, but other activation functions (e.g., hard tanh activation function, sigmoid activation function, or tanh activation function) may also be used.

[0047] For example, the visual recognition neural network may include a convolutional neural network architecture and / or may be a ResNet or DenseNet whose depth depends on the size of the input image. For example, for image pixel sizes up to 384×384 pixels, a Res-Net architecture with a depth of up to 50 layers may provide good results. For image pixel sizes from ~512×512 pixels to 800×800 pixels, a ResNet with a depth of 101 layers may be used. For sizes larger than the above image sizes, deeper architectures such as ResNet151 or DenseNet121 or DenseNet169 may be used.

[0048] The visual recognition machine learning algorithm can be trained by adjusting the parameters of the visual recognition machine learning algorithm based on a comparison of a high dimensional representation generated by the language recognition machine learning algorithm with a high dimensional representation generated by the visual recognition machine learning algorithm of the corresponding input training data. For example, the network weights of the visual recognition neural network can be adjusted based on the comparison. The adjustment of the parameters (e.g., network weights) of the visual recognition machine learning algorithm can be completed with consideration of a loss function. For example, the comparison of the first high dimensional representation with the second high dimensional representation in order to adjust the visual recognition machine learning algorithm can be based on a cosine similarity loss function. The training can converge quickly, and / or the visual recognition machine learning algorithm can be trained using a cosine similarity loss function (although other loss functions can also be used), which can provide a trained algorithm for biologically relevant data.

[0049] For example, a vision model can learn how to represent images in a semantic embedding space (e.g., as vectors). Therefore, a measure of the distance between two vectors can be used, which can represent the prediction A (the second high-dimensional representation) and the ground truth B (the first high-dimensional representation). For example, the measure is the cosine similarity as defined below:

[0050]

[0051] Here, the dot product of the prediction A and the ground truth B is divided by the dot product of their corresponding magnitudes (e.g., as in the L2 norm or Euclidean norm).

[0052] Figure 2 An example of training of a language identification machine learning algorithm 220 is shown (e.g., a lookup of a token embedding is shown). The textuality model 220 can be trained on biological sequences or natural languages ​​210 (e.g., nucleotide sequences (e.g., GATTACA)) from a database 200 or an imaging device (e.g., a microscope) in a running experiment. For example, a natural language processing (NLP) task is to predict the next word (dependent variable) in a sentence (independent variable) or to predict the next character of a short fragment of a given text 250 (e.g., the next nucleotide in a nucleotide sequence (e.g., the C after GATTACA)). Other NLP tasks may involve predicting sentiment in text or translation. In the context of biological sequences, the independent variable may be a protein sequence or a nucleotide sequence or a short fragment thereof. The dependent variable may be the next element in the sequence or any coarse-grained search term mentioned or a combination thereof. During training, data may be passed down to the encoder path 230 to learn a hidden representation 260 (a first high-dimensional representation) and passed up through the decoder path 240 to make useful predictions 250 (e.g., based on output training data of a biologically relevant language). A quantitative metric, such as a loss function, can measure the accuracy of the prediction relative to real data. The gradient of the loss function with respect to the trainable parameters of the model can be used to adjust these trainable parameters. This training can be iterated until a preset threshold of the loss function is met. The result of finding token embeddings during training can be a mapping from each token to its corresponding embedding, such as a latent vector 260 (first high-dimensional representation). The latent space can represent a semantic space. For example, meaning can be assigned to each token (e.g., a word or peptide or polynucleotide) through this embedding.

[0053] The prediction 250 can be represented by the output training data y based on the biologically relevant language. For example, y=W*X, where X is the input training data (e.g., biological sequence) based on the biologically relevant language, and W is the trained parameters of the model. In addition, a bias term can be included.

[0054] Optionally, after training the language recognition machine learning algorithm, the images can be mapped to token embeddings. In other words, one can choose to display images of biological structures corresponding to the input training data based on the biologically relevant language. For example, the input training data based on the biologically relevant language can be nucleotide sequences (e.g. Figure 2 A plurality of images corresponding to a plurality of biologically relevant language-based input training data sets may be selected as training sets for training a visual recognition machine learning algorithm. If such a database of training images is already available, the selection of training images may be avoided.

[0055] The visual model can be responsible for computer vision tasks such as predicting the category of an image, such as which subcellular compartment is shown in the image. In other applications, the visual model uses one-hot encoded labels as dependent variables. For example, the system 100 maps image categories to corresponding token embeddings learned by the textual model as described above. For example, an image classifier that learns to predict the categories "p53", "Histone H1", and "GAPDH" will learn to predict the token embeddings of the corresponding protein sequences of the three proteins (e.g., the same can be applied to token embeddings learned from nucleotide sequences or text descriptions in scientific publications). The mapping in the real data itself can be a lookup table of pictures to show the molecules of interest and their corresponding semantic embeddings of biological sequences or natural languages ​​used for training.

[0056] Only the high-dimensional representation 260 may be of interest, which can be obtained by a forward pass of the language identification machine learning algorithm on the input text. For training, a language classification problem may be defined. For example, a soft max layer may follow the determination of the high-dimensional representation 260 and a cross entropy loss function may be used for training. Figure 2 In , an additional decoder path 240 is shown, which again generates text that represents what the model would have been like if it had output text. For example, if the first few words were input, a prediction for the second half of the sentence could be made. For example, for biologically related applications, one could input the first part of a sequence and predict the second half of the sequence or just the next character with a certain probability. Since only the high-dimensional representation 260 is of interest, this prediction 250 may not be of interest, but it may improve training. Then, Figure 3 The visual model can predict the high-dimensional representation 260 as the fact 330. For this application, the cosine distance function can be used as a loss function instead of the cross entropy loss function. The vectors 260, 330 may not be normalized to 0 or 1. Since Batch Normalization can be used to keep the numbers manageable, the values ​​of the vectors may not be much greater than 1.

[0057] Figure 3 An example of training a visual recognition machine learning algorithm 320 is shown. Training of the visual model 320 may be performed to predict token embeddings. Figure 3 As shown, the visual model 320 can be trained on images 310 from a data repository 300 (e.g., a public image database or a private image database) or from a microscope in a running experiment. The dependent variable can be the corresponding token embedding 330 (a second high-dimensional representation) learned by the textual model and optionally mapped to the image categories as described above. The visual model can learn to predict representations of image categories, where these image categories contain the semantics of biological functions learned by the textual model in the previous training stage.

[0058] Figure 4 An example of a portion 400 of a visual recognition neural network based on a ResNet architecture (e.g., a ResNet block) is shown. For example, the visual recognition neural network can be described by the following parameters (e.g., similar to ResNet). The dimensions of a tensor (e.g., data passed through a deep neural network) can be:

[0059] shape = bs × ch × height × width

[0060] Where bs is the batch size (e.g., the number of images loaded into a mini-batch stochastic gradient descent optimization), ch is the number of filters (e.g., equivalent to the number of "channels" of the input image, e.g., ch=3 for RGB images), the height is the number of rows in the image, and the width is the number of columns in the image. For example, a microscope can produce more dimensions (e.g., an axial dimension (z), a spectral emission dimension, a lifetime dimension, a spectral excitation dimension, and / or a stage dimension), which can also be processed by a visual recognition neural network. However, the following examples may only involve the case with channels, height, and width (e.g., examples with ch>3 may also be implemented).

[0061] Visual recognition neural networks can be represented as computational graphs, and operations can be summarized as "layers" that represent specific operations on input data (such as tensors). The following annotations can be used:

[0062] ch_0 The number of channels of the input tensor before the operation.

[0063] XX can be an n-dimensional tensor of the shape defined above.

[0064] conv(n in ,n #t,k,s)(x) an n-dimensional convolution operation 430 (e.g., a 2D convolution in the case shown here) with n_in input channels (e.g., a spatial filter), n_out output channels, a kernel size of k×k (e.g., 3×3), and a stride of s×s (e.g., 1×1) applied to the tensor X.

[0065] As shown in the figure, the rectified linear unit is a non-linearity performed after the convolution. In the figure, this operation is described as "Relu" 420.

[0066] Batch Normalization normalizes the tensor X to its corresponding batch’s mean μ and standard deviation σ. In the figure, this operation is depicted as “Batch Normalization” 410.

[0067] fc(x)=Wx+b The fully connected layer is a linear operator, where W is the weight and b is the bias term (for example, b is not shown in the figure). Where n_in and n_out are the input and output channel dimensions of the current activation.

[0068] rn(x) Figure 4 A ResNet block 400 is shown in FIG, where a bottleneck configuration is applied to a tensor X of shape (1, 64, 256, 256) starting from the activations of the previous layer.

[0069] Some bottleneck blocks can downsample the spatial dimension by a factor of 2 while upsampling the number of channels (e.g., spatial filters) by a factor of 4. ResNet blocks can be combined into groups to produce an overall architecture of 18 to 152 layers. For example, using 50, 101, or 152 layers and bottleneck ResNet blocks and / or ResNet blocks with pre-activation can be used for the proposed visual recognition neural network.

[0070] For example, the visual recognition neural network may include at least a first batch of normalization operations 410, followed by a first ReLu operation 420, followed by a first convolution operation 430 (e.g., 1×1), followed by a second batch of normalization operations 410, followed by a second ReLu operation 420, followed by a second convolution operation 430 (e.g., 3×3), and finally an addition operation 440 (e.g., adding the output of the second convolution operation and the input of the first batch of normalization operations). One or more additional operations may be performed before the first batch of normalization operations 410, after the addition operation 440, and / or in between.

[0071] Figure 5An example of a portion 500 of a visual recognition neural network 400 based on a ResNet architecture, such as a modified ResNet-Convolutional Block Attention Module (CBAM block), is shown. For example, the ResNet-CBAM block 500 may use a so-called channel attention block in a ResNet block combined with spatial attention.

[0072] In addition to combining Figure 4 In addition to the annotations used, the following annotations can also be used:

[0073] Global average pooling collapses a tensor X of dimensions (bs×ch×h×w) to dimensions (bs×ch×1×1) by averaging the height and width dimensions. In the figure, this operation is depicted as “global average pooling” 510.

[0074] gmp(x)=max i=1,…,h max j=1,…,w x(i,j) Global max pooling collapses the tensor X of dimensions (bs×ch×h×w) to dimensions (bs×ch×1×1) by selecting the maximum of the height and width dimensions. In the figure, this operation is described as “global max pooling” 520.

[0075] For channel attention, a cascade 530 of global average pooling 510 and global max pooling 520 can be used instead of using global average pooling 510 alone. In this way, the model can learn both at the same time, and the "soft" global average pooling makes the model more resilient to outliers while maintaining maximum activation. Therefore, the model is able to decide which item to highlight. For example, the output of the previous operation can be provided as an input to the global average pooling operation 510 and the global max pooling operation 520, and the output of the global average pooling operation 510 and the output of the global max pooling operation 520 can be provided as an input to a subsequent identical operation (e.g., a cascade).

[0076] In addition, a 1×1 kernel size can be used instead of a mini MLP (Multi-Layer Perceptron), which can save some redundant flattening and decompression operations in the channel attention module.

[0077] Both the channel attention module and the spatial attention module can use the sigmoid nonlinear function 540 as the final activation function. In this way, more favorable feature scaling can be obtained than using ReLU activation.

[0078] Optionally, between channel attention and spatial attention, batch normalization 410 can be performed immediately after the scaling of channel attention occurs to avoid gradients becoming too large.

[0079] Add the outputs of the previous ResNet bottleneck block and CBAM block, such as Figure 5 The CBAM block starts with “global average pooling” 510 and “global maximum pooling” 520 and ends with the last “Mul” (multiplication) 550

[0080] From these Rn_CBAM(x) building blocks, Figure 5 Rn_CBAM(x) is replaced by The ResNet architecture is constructed by adding bottleneck blocks. For example, for the proposed concept, deeper architectures with 50, 101, and 152 layers can be used, but other depths can also be used.

[0081] Mean operation 560 and maximum operation 570 can work together by generating an arithmetic mean for dimension ch (e.g., 1×1×256×256 from 1×64×256×256) and generating a maximum projection along dimension ch by means of maximum operation 570. The following concatenation operation 530 connects the results of the two projections.

[0082] For example, the visual recognition neural network may include at least a first batch of normalization operations 410, followed by a first ReLu operation 420, followed by a first convolution operation 430 (e.g., kernel size 1×1), followed by a second batch of normalization operations 410, followed by a second ReLu operation 420, followed by a second convolution operation 430 (e.g., kernel size 3×3), followed by a global average pooling operation 510 and a global maximum pooling operation 520, followed by a first cascade operation 530, followed by a third convolution operation 430 (e.g., 1×1), followed by a third ReLu operation 420, followed by a fourth convolution operation 430 (e.g., kernel size 1×1), followed by a first sigmoi d operation 540, followed by a first multiplication (Mul) operation 550 (e.g., multiplying the output of the first sigmoid operation with the output of the second convolution operation), followed by a third batch normalization operation 410, followed by a mean operation 560 and a maximization operation 570, followed by a second cascade operation 530, followed by a fifth convolution operation 430 (e.g., kernel size 7×7), followed by a second sigmoid operation 540, followed by a second multiplication (Mul) operation 550 (e.g., multiplying the output of the second sigmoid operation with the output of the third batch normalization operation), and finally an addition operation 440 (e.g., adding the output of the second multiplication operation to the input of the previous block). The operation between the second convolution operation and the third batch normalization operation can be called a channel attention module, and the operation between the first multiplication operation and the second multiplication operation can be called a spatial attention module. The operation from the first batch normalization operation to the second convolution operation can be called a ResNet bottleneck block, and the operation between the second convolution operation and the second multiplication operation can be called a CBAM block. The CBAM block can be used to scale the second convolution so that the model focuses on the correct features. One or more additional operations can be performed before the first normalization operation 410, after the addition operation 440, and / or in between.

[0083] Figure 6 An example of a portion 600 of a visual recognition neural network based on a DenseNet architecture (e.g., a dense layer with a bottleneck configuration) is shown. An alternative architecture to ResNet is called DenseNet, which relies on a continuous activation map of connections (e.g., instead of additions as in ResNet) to make the activations of the upstream layer directly available to the downstream layer. For the proposed concept, a DenseNet architecture with an attention mechanism added at the level of a single dense layer Hl_B(x) can be used. The channel attention mechanism can be combined with sparse (sparsified) DenseNets.

[0084] For the proposed concept, both spatial attention and channel attention can be combined with dense layers. Optionally, batch normalization between the channel attention module and the spatial attention module can be used, such as by ResNet architecture (e.g., combined with Figure 4 and Figure 5 ). Instead of adding the output of the attention path to the output of the dense layer, the attention mechanism can be applied only to the k activations newly generated by the dense layer, and the rescaled output of the attention path can be connected to the input of the dense layer at the end. For example, for all layers except the first dense layer, the activations have passed through the previous dense layer to which the attention mechanism is attached. Continuous rescaling may not further improve the results. On the contrary, such rescaling may even prevent the network from learning new attention rescaling in more downstream layers as needed. In addition, focusing only on the newly created k layers can reduce the computational complexity and can omit the need for the reduction ratio r as a patch to limit (cap) the computational complexity. For dense layers and DenseNet blocks, full configuration can be used instead of sparse configuration.

[0085] In addition to combining Figure 4 and 5 In addition to the symbols used, the following symbols can also be used:

[0086] Hl_B(x) Figure 6 A dense layer 600 with a bottleneck configuration is shown in FIG.

[0087] The input tensor X of shape (bs,ch,h,w) passes through two consecutive convolutions, each with pre-activation (bn+relu). The first convolution has a 1×1 kernel and outputs ch activations. The second convolution has a 3×3 kernel and outputs only k activations. In this example, k=16. Finally, the 16 new activations are concatenated with the input of the dense layer. In this example, ch=64, so the output has ch+k=80 activations.

[0088] and Figure 4 Compared to the portion of the visual recognition neural network shown in , the addition operation 440 is replaced by a concatenation operation 530 (e.g., the output of the second convolution operation and the input of the first normalization operation). Figure 4 describe.

[0089] Figure 7 An example of a portion 700 of a visual recognition neural network based on a DenseNet architecture (e.g., a dense layer with an attention mechanism) is shown.

[0090] In addition to combining Figure 4 , 5In addition to the symbols used in and 6, the following symbols can also be used:

[0091] Hl_A Dense layer 700 with attention mechanism.

[0092] This building block of DenseNet can be used for the proposed idea. Similar to the attention mechanism described above for ResNet, two continuous attention modules are introduced with channel attention and spatial attention respectively. The output of the attention path is cascaded with the output of the dense layer.

[0093] Based on these Hl_A(x) building blocks, we can replace elements to get DenseNet.

[0094] and Figure 5 Compared to the portion of the visual recognition neural network shown, the addition operation 440 is replaced by a concatenation operation 530 (e.g., the output of the second multiplication operation and the input of the first normalization operation). Figure 5 describe.

[0095] System 100 may be configured to use a variety of Figures 4 to 7 One of the parts of a visual recognition neural network is shown.

[0096] The system 100 may include or may be a computer device (e.g., a personal computer, a laptop computer, a tablet computer, or a mobile phone) in which one or more processors 110 and one or more storage devices 120 are located, or the system 100 may be a distributed computing system (e.g., a cloud computing system having one or more processors 110 and one or more storage devices 120 distributed in various locations (e.g., a local client and one or more remote service farms and / or data centers)). The system 100 may include a data processing system that includes a system bus for connecting the various components of the system 100. The system bus may provide communication links between the various components of the system 100 and may be implemented as a single bus, a combination of buses, or in any other suitable manner. Electronic components may be connected to the system bus. The electronic components may include any circuit or combination of circuits. In one embodiment, the electronic component includes a processor that may be of any type. As used herein, a processor may refer to any type of computing circuit, such as, but not limited to, a microprocessor, a microcontroller, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a graphics processor, a digital signal processor (DSP), a multi-core processor, a field programmable gate array (FPGA) of a microscope or microscope component (e.g., a camera), or any other type of processor or processing circuit. Other types of circuits that may be included in an electronic assembly may be custom circuits, application specific integrated circuits (ASICs), etc., such as, for example, one or more circuits (such as communication circuits) used in wireless devices such as mobile phones, tablet computers, laptop computers, two-way radios, and similar electronic systems. The system 100 includes one or more storage devices 120, which in turn may include one or more storage elements suitable for a particular application (such as, a main memory in the form of a random access memory (RAM)), one or more hard disk drives, and / or one or more drivers for processing removable media (such as compact disks (CDs), flash memory cards, digital video disks (DVDs), etc.). The system 100 may also include a display device, one or more speakers, and a keyboard and / or controller, which may include a mouse, a trackball, a touch screen, a voice recognition device, or any other device that allows a system user to input information to the system and receive information from the system 100.

[0097] Additionally, the system 100 may include a microscope connected to a computer device or a distributed computing system. The microscope may be configured to generate biologically relevant image-based input training data 104 by taking images of biological samples.

[0098] The microscope can be an optical microscope (e.g., a diffraction-limited microscope or a sub-diffraction-limited microscope, such as, for example, a super-resolution microscope or a nanoscope). The microscope can be a stand-alone microscope or a microscope system with attached components (e.g., a confocal scanner, an additional camera, a laser, a climate chamber, an automatic loading mechanism, a liquid handling system, attached optical components (such as additional multiphoton optical paths, light sheet imaging components, optical tweezers, etc.)). Other image sources can also be used as long as they can capture images of objects related to biological sequences (e.g., proteins, nucleic acids, lipids). For example, a microscope according to the embodiments described above or below can implement a deep discovery microscope.

[0099] Further details and aspects of the system 100 may be combined with the proposed concepts and / or one or more examples described above or below (e.g., Figure 8 and Fig. 9 The system 100 may include one or more additional optional features corresponding to one or more aspects of the proposed concept and / or one or more aspects of one or more examples described above or below.

[0100] Some embodiments relate to a microscope comprising: Figures 1 to 7 Alternatively, the microscope may be a system such as a combination of Figures 1 to 7 A portion of a system described in one or more diagrams. Figure 8 A schematic diagram of a system 800 for training a machine learning algorithm is shown. A microscope 810 configured to take images of biological samples is connected to a computer device 820 (e.g., a personal computer, laptop computer, tablet computer, or mobile phone) configured to train a machine learning algorithm. The microscope 810 and the computer device 820 may be combined as shown in FIG. Figures 1 to 7 The invention may be implemented as described in one or more of the figures.

[0101] Fig. 9A flow chart of a method for training a machine learning algorithm to process biologically relevant data is shown. The method 900 includes receiving 910 input training data based on a biologically relevant language, and generating 920 a first high-dimensional representation of the input training data based on the biologically relevant language by a language recognition machine learning algorithm. The first high-dimensional representation includes at least 3 entries, each entry having a different value. Further, the method 900 includes generating 930 output training data based on the biologically relevant language by the language recognition machine learning algorithm according to the first high-dimensional representation, and includes adjusting 940 the language recognition machine learning algorithm according to a comparison of the input training data based on the biologically relevant language and the output training data based on the biologically relevant language. In addition, the method 900 includes receiving 950 input training data based on a biologically relevant image associated with the input training data based on the biologically relevant language, and includes generating 960 a second high-dimensional representation of the input training data based on the biologically relevant image by a visual recognition machine learning algorithm. The second high-dimensional representation includes at least three entries, each entry having a different value. In addition, the method 900 includes adjusting 970 the visual recognition machine learning algorithm according to a comparison of the first high-dimensional representation and the second high-dimensional representation.

[0102] By using a language recognition machine learning algorithm, textual biological input can be mapped to a high-dimensional representation. By making the high-dimensional representation have entries with various different values ​​(compared to the one-hot encoding representation), semantically similar biological input can be mapped to similar high-dimensional representations. By training a visual recognition machine learning algorithm to map images to high-dimensional representations trained by a language recognition machine learning algorithm, images with similar biological content can also be mapped to similar high-dimensional representations. Therefore, the possibility of semantically correct classification or at least semantically close classification of images by a correspondingly trained visual recognition machine learning algorithm can be significantly improved. Further, the correspondingly trained visual recognition machine learning algorithm can more accurately map untrained images to high-dimensional representations close to high-dimensional representations with similar meanings or to semantically matching high-dimensional representations. Through the proposed concept, a trained language recognition machine learning algorithm and / or a trained visual recognition machine learning algorithm can be obtained, which can provide semantically correct or very accurate classification of input data based on biologically relevant language and / or input data based on biologically relevant images. The trained language recognition machine learning algorithm and / or the trained visual recognition machine learning algorithm can be used to retrieve biologically relevant images from multiple biological images based on language-based retrieval input or image-based retrieval input, label biologically relevant images, find or generate typical images, and / or similar applications.

[0103] Further details and aspects of method 900 may be combined with the proposed concepts and / or one or more examples described above or below (e.g., Figures 1 to 8The method 900 may include one or more additional optional features corresponding to one or more aspects of the proposed concept and / or one or more aspects of one or more examples described above or below.

[0104] Some embodiments relate to a trained machine learning algorithm, the trained machine learning algorithm being trained by the following steps: receiving input training data based on a biologically relevant language; and generating a first high-dimensional representation of the input training data based on the biologically relevant language by a language recognition machine learning algorithm. The first high-dimensional representation includes at least 3 entries, each entry having a different value. Further, the trained machine learning algorithm is trained by the following steps: generating output training data based on the biologically relevant language from the first high-dimensional representation by the language recognition machine learning algorithm; and adjusting the language recognition machine learning algorithm based on a comparison of the input training data based on the biologically relevant language with the output training data based on the biologically relevant language. In addition, the trained machine learning algorithm is trained by the following steps: receiving input training data based on a biologically relevant image associated with the input training data based on the biologically relevant language; and generating a second high-dimensional representation of the input training data based on the biologically relevant image by a visual recognition machine learning algorithm, wherein the second high-dimensional representation includes at least 3 entries, each entry having a different value. Further, the trained machine learning algorithm is trained by adjusting the visual recognition machine learning algorithm based on the comparison of the first high-dimensional representation and the second high-dimensional representation.

[0105] The trained machine learning algorithm may be a trained visual recognition machine learning algorithm (e.g., an adjusted visual recognition machine learning algorithm) and / or a trained language recognition machine learning algorithm (e.g., an adjusted language recognition machine learning algorithm). At least a portion of the trained machine learning algorithm may be learning parameters (e.g., neural network weights) stored by a storage device.

[0106] Further details and aspects of the trained machine learning algorithm may be combined with the proposed concepts and / or one or more examples described above or below (e.g., Figures 1 to 9 ) are described. The trained machine learning algorithm may include one or more additional optional features corresponding to one or more aspects of the proposed concept and / or one or more aspects of one or more examples described above or below.

[0107] In the following, one or more of the above embodiments (for example, in combination with Figures 1 to 9 Some examples of application and / or implementation details of the embodiments) described in one or more figures.

[0108] For example, biology in general and microscopy in particular generates a large amount of data which are often poorly annotated or not annotated at all. Often, it is only clear in retrospect which annotations might have been useful or which new biological findings were not known at the time of the experiment. Based on the proposed concept, such data can be made accessible by allowing semantic retrieval and tagging of large amounts of image data stored in databases or as part of experiments run in microscopes. Such experiments can be single one-off experiments or parts of long-term experiments such as screening campaigns.

[0109] In the context of running experiments, the proposed concept can help automatically retrieve biological structures that are part of a sample (such as proteins displayed in single cells, organoids or tissues) as well as more general structures (such as organs or developmental situations). In this way, the time-consuming step of finding relevant parts within a sample can be automated. Otherwise, this step may require repetitive manual work by human experts under time pressure (for example because expensive research equipment is booked for a period of time) in uncomfortable environments (such as noisy dark rooms). The proposed concept can also make this step more objective by avoiding personal biases.

[0110] The proposed concept can achieve zero-shot learning, which means classifying or annotating images of a type that has never been seen before. Because the image model part of the proposed concept can predict semantic embeddings (e.g., high-dimensional representations) instead of one-hot encoded categories, the proposed concept is able to find the closest match for unknown images in the semantic space (e.g., multiple high-dimensional representations). For example, new discoveries may be made by discovering previously unknown biological functions in microstructures. For example, if no matching information is found in the database, the proposed concept can infer the missing information based on the image or available information. This can enable the retrieval of a large amount of existing data that is unannotated or poorly annotated.

[0111] The proposed concept can use a deep learning approach that combines semantic text embedding with image models (e.g., convolutional neural networks (CNNs)) to make unannotated or poorly annotated biological images, image stacks, time intervals, or combinations thereof, such as from optical or electron microscopy, retrievable or extract biological information from them. According to one aspect, a combination of textual models and visual models (e.g., language recognition algorithms and visual recognition algorithms) can be used in microscopy.

[0112] The proposed visual semantic model (e.g., a combination of a language recognition machine learning algorithm and a visual recognition machine learning algorithm) can be based on a two-stage process. In stage 1, a textual model (e.g., a language recognition algorithm) can be trained for biological sequences to solve text cognition tasks. The semantic embedding discovered by the model in stage 1 can then be used as a target value to be predicted by a visual model (e.g., a visual recognition algorithm) in stage 2. This combination and optionally application in a microscope during a running experiment can allow for a variety of applications.

[0113] For example, since the one-hot encoding category vectors of other visual models trained for classification tasks treat each category as completely unrelated, it is impossible to capture any semantics of the category. In contrast, the textual model of stage 1 can capture semantics as token embeddings (e.g., also called latent vectors, semantic embeddings, or high-dimensional representations). Tokens can be characters, words, or in the context of biomolecules, secondary structures, binding motifs, catalytic sites, promoter sequences, etc. Then, because the visual model can be trained for these semantic embeddings, and therefore the visual model can not only predict the same categories that have been trained, but also predict new categories that are not included in the training set. Therefore, the semantic embedding space can be used as a proxy for biological functions. Molecules with similar functions imaged by the proposed imaging system (e.g., microscope) can appear adjacently in the embedding space. In contrast, using other classifiers to predict information about the one-hot encoding category vectors of biological functions cannot be achieved. Therefore, other classifiers cannot predict categories that have not been seen before ("zero-sample learning"), and if the classification is wrong, the predicted category is usually completely unrelated to the actual category.

[0114] By combining a text model (e.g., a language model) trained on text and learning semantic embeddings as hidden representations of text, the proposed concept can train prediction models just like in deep neural networks. Biological sequences (e.g., protein sequences or nucleotide sequences) can be used as text. Other embodiments can use natural language (e.g., text used in scientific publications) to describe the functions of biomolecules. Visual models (e.g., convolutional neural networks (CNNs)) can be trained to predict their corresponding embeddings (e.g., different from the one-hot encoded feature vectors used by others).

[0115] For example, one aspect of the proposed concept describes systems and embodiments built on a combination of a language model (or text model) and a visual model.

[0116] The language model can be implemented as a deep recurrent neural network (RNN), such as a long short-term memory (LSTM) model. The visual model can be implemented as a deep convolutional neural network (CNN). Other embodiments can use different types of deep learning models or machine learning models. For example, the visual model can be implemented as a capsule network.

[0117] The combination of text and visual information across different knowledge domains can allow a visual model to learn the true semantic representation of the images for which the visual model is trained. For example, in the field of image classification, a CNN can be trained to predict different categories that describe the content of an image with one word. The word is represented as a one-hot encoding vector. In one-hot encoding, the encodings of "Lilium sp. pollen grain" and "endosome" are as close or different as those of "endosome" and "lysosome", even though these two organelles are more similar than organelles and pollen grains. Therefore, a visual model trained to predict one-hot encoding vectors can be completely correct or completely wrong. However, if the model is trained to predict the semantic embedding of a category (e.g., learned through a language model), its predictions can be closer to semantically related objects in this embedding space.

[0118] For example, according to the proposed concept, a language model is trained on text and learns semantic embeddings as hidden representations of the text. For example, a language model trained to predict the next word in a sentence can represent words in a 500-dimensional latent vector. Other dimensions are also possible. Latent vectors between 50 and 1000 dimensions can be used for natural language processing. The proposed concept can use biological sequences such as protein sequences or nucleotide sequences as text and train visual models to predict their corresponding embeddings. Biological sequences can encode biological functions and can thus be understood as a form of "biological language". In addition, natural language can also be used to represent images, because there are a large number of scientific publications describing the functional roles of biological entities such as protein sequences or nucleotide sequences as well as subcellular localization or developmental and / or metabolic states, which makes this information useful for characterizing microscopic images.

[0119] For example, the steps to obtain a trained model may be:

[0120] - Finding token embeddings: A first language / linguistic model (e.g. RNN, LSTM) is trained on a representation of a biomolecule in the form of a nucleotide / protein sequence or a textual description / explanation of the corresponding biomolecule (e.g. nucleotide, protein) in a scientific publication. The generated token embeddings may be derived, for example, during model training. The final result of this first training phase itself (e.g. prediction of the next element in a sequence) may not be of interest. However, the definition of a prediction target may improve the accuracy and / or speed of training.

[0121] - Mapping images (e.g. images of corresponding biomolecules) to corresponding token embeddings. In other words, images can be selected from biological structures representing textual biological inputs for training of language / linguistic models. These images can be used for training in the second stage. Such image mapping may not be needed if a database of images with corresponding textual biological descriptions is used.

[0122] - Perform a second stage of training an image recognition model (e.g. CNN, capsule network) to predict the corresponding token embeddings found by the first model. The input is an image of the corresponding biomolecule. The image can be mapped to the semantics contained in the token embeddings generated by the first model.

[0123] For example, one can construct Figure 2 The textual model shown is used to find token embeddings. The biological sequence 210 can be passed from the repository 200 to the textual model 220 as an independent variable. The textual model can be responsible for tasks in language processing, such as predicting the next character from a short fragment of the sequence (such as an amino acid in a protein sequence or a base in a nucleotide sequence). Other language processing tasks can find suitable but different types of embeddings. Such tasks can involve homology prediction, predicting the next word in a sentence, etc. The data can be passed down to the encoder path 230 to learn the hidden representation and pass through the decoder path so that useful predictions 250 can be made based on it. The hidden representation can be regarded as an embedding (such as a high-dimensional vector) in a latent space. In the trained model, the token embedding can represent the mapping of each token to its corresponding latent vector 260. In the text model responsible for natural language processing tasks, tokens can be equivalent to words, and token embeddings can be word embeddings.

[0124] For example, training a vision model to predict token vectors such as Figure 3 As shown. Images 310 from a data repository 300 or from a microscope during a running experiment can be passed as independent variables to the input of a vision model 320. As dependent variables, token embeddings 330 that have been mapped to the desired image category can be displayed to the model at the output. The vision model can learn to predict the token embedding for each input.

[0125] Embodiments may be based on the use of machine learning models or machine learning algorithms. Machine learning may refer to algorithms and statistical models that a computer system can use to perform specific tasks without explicit instructions but relying on models and reasoning. For example, in machine learning, data transformations inferred from analysis of historical and / or training data may be used instead of rule-based data transformations. For example, a machine learning model may be used or a machine learning algorithm may be used to analyze the content of an image. In order for a machine learning model to analyze image content, a machine learning model may be trained using training images as input and training content information as output. By training a machine learning model with a large number of training images and / or training sequences (e.g., words or sentences) and associated training content information (e.g., labels or annotations), the machine learning model "learns" to recognize the content of the image, so the machine learning model may be used to recognize the content of the image not included in the training data. The same principle may also be used for other kinds of sensor data: by training a machine learning model using training sensor data and desired outputs, the machine learning model "learns" the conversion between sensor data and output, which may be used to provide output based on non-training sensor data provided to the machine learning model.

[0126] The machine learning model can be trained using training input data. The examples described in detail above use a training method called "supervised learning". In supervised learning, a machine learning model is trained using multiple training samples, each of which may include multiple input data values ​​and multiple expected output values, i.e., each training sample is associated with an expected output value. By specifying the training samples and the expected output values, the machine learning model "learns" which output value is provided based on input samples similar to the samples provided during training. In addition to supervised learning, semi-supervised learning can also be used. In semi-supervised learning, some training samples lack corresponding expected output values. Supervised learning can be based on supervised learning algorithms, such as classification algorithms, regression algorithms, or similarity learning algorithms. When the output is limited to a limited set of values, a classification algorithm can be used, i.e., the input is classified into one of a limited set of values. When the output can have any numerical value (within a certain range), a regression algorithm can be used. A similarity learning algorithm can be similar to a classification algorithm and a regression algorithm, but is based on learning from examples using a similarity function that measures the similarity or correlation between two objects. In addition to supervised or semi-supervised learning, unsupervised learning can also be used to train machine learning models. In unsupervised learning, one may (only) provide input data, and an unsupervised learning algorithm may be used to find structure in the input data, e.g., find commonalities in the data by grouping or clustering the input data. Clustering is the assignment of input data comprising multiple input values ​​into subsets (clusters) such that input values ​​within the same cluster are similar according to one or more (predefined) similarity criteria, but are dissimilar to input values ​​contained in other clusters.

[0127] Reinforcement learning is the third group of machine learning algorithms. In other words, reinforcement learning can be used to train machine learning models. In reinforcement learning, one or more software actors (called "software agents") are trained to take actions in an environment. Based on the actions taken, rewards are calculated. Reinforcement learning is based on training one or more software agents to select actions in order to increase the cumulative reward, so that the software agent becomes better at a given task (as evidenced by increasing rewards).

[0128] In addition, some techniques can be applied to some machine learning algorithms. For example, feature learning can be used. In other words, a machine learning model can be trained at least in part using feature learning, and / or a machine learning algorithm can include a feature learning component. Feature learning algorithms (also called representation learning algorithms) can retain information in their inputs, but can also transform them in a way that makes this information useful, usually as a preprocessing step before performing classification or prediction. For example, feature learning can be based on principal component analysis or cluster analysis.

[0129] In some examples, anomaly detection (i.e., outlier detection) can be used, the purpose of which is to provide identification for input values ​​that are significantly different from the majority of inputs or training data and thus raise suspicion. In other words, the machine learning model can be trained at least in part using anomaly detection, and / or the machine learning algorithm can include an anomaly detection component.

[0130] In some examples, a machine learning algorithm may use a decision tree as a prediction model. In other words, a machine learning model may be based on a decision tree. In a decision tree, observations about an item (e.g., a set of input values) may be represented by branches of the decision tree, and output values ​​corresponding to the item may be represented by leaves of the decision tree. Decision trees may support both discrete and continuous values ​​as output values. If discrete values ​​are used, the decision tree may be indicated as a classification tree, and if continuous values ​​are used, the decision tree may be indicated as a regression tree.

[0131] Association rules are another technique that can be used in machine learning algorithms. In other words, a machine learning model can be based on one or more association rules. Association rules are created by identifying relationships between variables in a large amount of data. A machine learning algorithm can identify and / or utilize one or more relationship rules that represent knowledge derived from the data. The above rules can be used, for example, to store, manipulate, or apply this knowledge.

[0132] Machine learning algorithms are typically based on machine learning models. In other words, the term "machine learning algorithm" may indicate a set of instructions that can be used to create, train, or use a machine learning model. The term "machine learning model" may indicate, for example, a data structure and / or a set of rules that represent learned knowledge based on training performed by a machine learning algorithm. In an embodiment, the use of a machine learning algorithm may mean the use of an underlying machine learning model (or multiple underlying machine learning models). The use of a machine learning model may mean that the machine learning model and / or the data structure / rule set that is the machine learning model is trained by a machine learning algorithm.

[0133] For example, a machine learning model can be an artificial neural network (ANN). ANN is a system inspired by biological neural networks (such as those found in the retina or brain). ANN includes multiple interconnected nodes and multiple connections between nodes, so-called edges. There are usually three types of nodes, namely: input nodes that receive input values, hidden nodes that are (only) connected to other nodes, and output nodes that provide output values. Each node can represent an artificial neuron. Each edge can transfer information from one node to another. The output of a node can be defined as a (nonlinear) function of the sum of its inputs. The input of a node can be used in the above function based on the "weight" of the edge or the "weight" of the node providing the input. The weights of nodes and / or edges can be adjusted during the learning process. In other words, the training of an artificial neural network can include adjusting the weights of the nodes and / or edges of the artificial neural network, i.e., achieving the desired output for a given input.

[0134] Alternatively, the machine learning model can be a support vector machine, a random forest model or a gradient boosting model. A support vector machine (i.e., a support vector network) is a supervised learning model with an associated learning algorithm, which can be used to analyze data, for example, in classification or regression analysis. The support vector machine can be trained by providing a plurality of training input values ​​belonging to one of the two categories for input. The support vector machine can be trained to assign new input values ​​to one of the two categories. Alternatively, the machine learning model can be a Bayesian network, which is a probabilistic directed acyclic graphical model. A Bayesian network can use a directed acyclic graphical model to represent a set of random variables and their conditional dependencies. Alternatively, the machine learning model can be based on a genetic algorithm, which is a retrieval algorithm and heuristic technique that mimics the process of natural selection.

[0135] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ".

[0136] Although some aspects have been described in the context of an apparatus, it is apparent that these aspects also represent descriptions of corresponding methods, where blocks or devices correspond to method steps or features of method steps. Similarly, some aspects described in the context of method steps also represent descriptions of corresponding blocks or items or features of corresponding apparatuses. Some or all of the method steps may be performed by (or using) hardware devices (e.g., processors, microprocessors, programmable computers, or electronic circuits). In some embodiments, one or more of the most important method steps may be performed by such an apparatus.

[0137] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software. The above implementation may be performed using a non-transitory storage medium such as a digital storage medium (e.g., a floppy disk, DVD, Blu-Ray, CD, ROM, PROM and EPROM, EEPROM or FLASH memory) on which electronically readable control signals are stored, which cooperate (or can cooperate) with a programmable computer system to perform the corresponding method. Therefore, the digital storage medium may be computer readable.

[0138] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0139] Generally, embodiments of the present invention may be implemented as a computer program product having a program code, which, when the computer program product is run on a computer, is operable to perform one of the methods. For example, the program code may be stored on a machine-readable carrier. For example, the computer program may be stored on a non-transitory storage medium. Some embodiments relate to a non-transitory storage medium containing machine-readable instructions, which, when executed, implement a method according to the proposed concept or one or more of the examples described above.

[0140] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0141] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0142] Therefore, another embodiment of the present invention is a storage medium (or a data carrier, or a computer-readable medium) comprising a computer program stored thereon, for performing one of the methods described herein when the computer program is executed by a processor. The data carrier, the digital storage medium or the recorded medium is typically tangible and / or non-transitory. Another embodiment of the present invention is an apparatus as described herein, comprising a processor and a storage medium.

[0143] Therefore, another embodiment of the present invention is a data stream or a sequence of signals representing a computer program for executing one of the methods described herein. The data stream or the sequence of signals may be configured, for example, to be transmitted via a data communication connection (e.g., via the Internet).

[0144] Another embodiment comprises a processing device, for example a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0145] A further embodiment comprises a computer on which the computer program for performing one of the methods described herein is installed.

[0146] Another embodiment according to the invention comprises an apparatus or system configured to transmit (e.g., electronically or optically) to a receiver a computer program for performing one of the methods described herein. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system may, for example, comprise a file server for transmitting the computer program to the receiver.

[0147] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions in the methods described herein. In some embodiments, a field programmable gate array can collaborate with a microprocessor to perform one of the methods described herein. Typically, these methods are preferably performed by any hardware device.

[0148] Reference numerals list

[0149] 100 Systems for training machine learning algorithms to process biologically relevant data

[0150] 102 Input training data based on biologically relevant languages

[0151] 104 Input training data based on biologically relevant images

[0152] 110 One or more processors

[0153] 120 One or more storage devices

[0154] 200 Database; Repository

[0155] 210 Input training data based on biologically relevant languages; biological sequences

[0156] 220 Language Identification Machine Learning Algorithms; Textual Models

[0157] 230 Encoder Path for Language Identification Machine Learning Algorithms

[0158] 240 Decoder Path for Language Identification Machine Learning Algorithms

[0159] 250 Output training data based on biologically relevant languages; prediction

[0160] 260 First High Dimensional Representation; Hidden Representation; Latent Vector; Token Embedding

[0161] 300 Repositories

[0162] 310 Input training data based on biologically relevant images; Image

[0163] 320 Visual Recognition Machine Learning Algorithms; Visual Models

[0164] 330 Second Highest Dimensional Representation; Hidden Representation; Latent Vector; Token Embedding

[0165] 400 Partial Visual Recognition Neural Network; ResNet Block

[0166] 410 Batch Normalization Operation

[0167] 420 ReLu operation

[0168] 430 Convolution Operation

[0169] 440 Addition

[0170] 500 Partial Visual Recognition Neural Network; ResNet-CBAM Block

[0171] 510 Global Average Pooling Operation

[0172] 520 Global Max Pooling Operation

[0173] 530 Cascade Operation

[0174] 540 sigmoid operation

[0175] 550 Multiplication

[0176] 560 Mean Operation

[0177] 570 Maximum value operation

[0178] 600 Partial Visual Recognition Neural Network; Dense Layers with Bottleneck Configuration

[0179] 700 Partial Visual Recognition Neural Network; Dense Layers with Attention Mechanism

[0180] 800 Systems for training machine learning algorithms

[0181] 810 Microscope

[0182] 820 Computer equipment

[0183] 900 Methods for training machine learning algorithms to process biologically relevant data

[0184] 910 Receiving input training data based on a biologically relevant language

[0185] 920 Generate the first high-dimensional representation

[0186] 930 Generate output training data based on biologically relevant language

[0187] 940 Adjusting the language recognition machine learning algorithm

[0188] 950 receiving input training data based on biologically relevant images

[0189] 960 Generate the second high-dimensional representation

[0190] 970 Tuning visual recognition machine learning algorithms.

Claims

1. A system (100), comprising one or more processors (110) and one or more storage devices (120), wherein: The system (100) is configured as follows: Receiving input training data based on a biology-related language (102, 210), wherein the input training data based on a biology-related language (102, 210) is at least one of: a nucleotide sequence or a protein sequence; generating, by a language identification machine learning algorithm (220) executed by the one or more processors (110), a first high-dimensional representation (260) of the biologically relevant language-based input training data (102, 210), wherein the first high-dimensional representation (260) includes at least 3 entries, each entry having a different value; generating, by the language identification machine learning algorithm (220) executed by the one or more processors (110), biologically relevant language based output training data (250) from the first high dimensional representation (260), wherein the biologically relevant language based output training data (250) comprises a prediction of a next element in the biological sequence; adjusting the language recognition machine learning algorithm (220) based on a comparison of the biologically relevant language based input training data (102, 210) and the biologically relevant language based output training data (250); receiving biologically relevant image-based input training data (104, 310) associated with the biologically relevant language-based input training data (102, 210), wherein the biologically relevant image-based input training data (104, 310) is image training data of an image of at least one of: a biological structure including nucleotides or nucleotide sequences; a biological structure including proteins or protein sequences; generating a second high-dimensional representation (330) of the biologically relevant image-based input training data (104, 310) by a visual recognition machine learning algorithm (320) executed by the one or more processors (110), wherein the second high-dimensional representation (330) includes at least three entries, each entry having a different value; and The visual recognition machine learning algorithm (320) is adjusted based on a comparison of the first high-dimensional representation (260) and the second high-dimensional representation (330).

2. The system according to claim 1, wherein: The input training data (102, 210) based on biology-related language is further at least one of the following: a description of a biological molecule or a biological structure; a description of a behavior of a biological molecule or a biological structure; or a description of a biological function or biological activity.

3. The system according to claim 1 or 2, wherein: The biologically relevant image-based input training data (104, 310) is further image training data of at least one of the following images: biological molecules; biological tissues; biological structures with specific behaviors; or biological structures with specific biological functions or specific biological activities.

4. The system according to claim 1 or 2, wherein: The value of one or more entries of the first high-dimensional representation (260) is proportional to the likelihood of the presence of a particular biological function or a particular biological activity.

5. The system according to claim 1 or 2, wherein: The value of one or more entries of the second high-dimensional representation (330) is proportional to the likelihood of the presence of a specific biological function or a specific biological activity.

6. The system according to claim 1 or 2, wherein: The first high-dimensional representation (260) and the second high-dimensional representation (330) are digital representations.

7. The system according to claim 1 or 2, wherein: The first high-dimensional representation (260) and the second high-dimensional representation (330) both include more than 100 dimensions.

8. The system according to claim 1 or 2, wherein: The first high-dimensional representation (260) is a first vector and the second high-dimensional representation (330) is a second vector.

9. The system according to claim 1 or 2, wherein: More than 50% of the values ​​of the entries of the first high-dimensional representation (260) and more than 50% of the values ​​of the entries of the second high-dimensional representation (330) are not equal to zero.

10. The system according to claim 1 or 2, wherein: The values ​​of more than five entries of the first high-dimensional representation (260) are 10% greater than the maximum absolute value of the entries of the first high-dimensional representation (260), and the values ​​of more than five entries of the second high-dimensional representation (330) are 10% greater than the maximum absolute value of the entries of the second high-dimensional representation (330).

11. The system according to claim 1 or 2, wherein: The comparison of the biologically relevant language based input training data (102, 210) and the biologically relevant language based output training data (250) for tuning the language identification machine learning algorithm (220) is performed based on a cross entropy loss function.

12. The system according to claim 1 or 2, wherein: The comparison of the first high-dimensional representation (260) and the second high-dimensional representation (330) for tuning the visual recognition machine learning algorithm (320) is performed based on a cosine similarity loss function.

13. The system according to claim 1 or 2, wherein: The biologically relevant language based input training data (102, 210) includes a length of more than 20 characters.

14. The system according to claim 1 or 2, wherein: The adjustment of the language identification machine learning algorithm (220) includes adjustment of a plurality of language identification neural network weights, wherein a final set of language identification neural network weights is stored by the one or more storage devices (120).

15. The system according to claim 1 or 2, wherein: The adjustments to the visual recognition machine learning algorithm (320) include adjustments to a plurality of visual recognition neural network weights, wherein a final set of visual neural network weights is stored by the one or more storage devices (120).

16. The system according to claim 1 or 2, wherein: The language identification machine learning algorithm (220) includes a language identification neural network.

17. The system of claim 16, wherein: The language recognition neural network includes more than 30 layers.

18. The system of claim 16, wherein: The language recognition neural network is a recursive neural network.

19. The system of claim 16, wherein: The language recognition neural network is a long short-term memory network.

20. The system according to claim 1 or 2, wherein: The visual recognition machine learning algorithm (320) includes a visual recognition neural network.

21. The system of claim 20, wherein: The visual recognition neural network includes more than 30 layers.

22. The system of claim 20, wherein: The visual recognition neural network is a convolutional neural network or a capsule network.

23. The system of claim 20, wherein: The visual recognition neural network includes multiple convolutional layers and multiple pooling layers.

24. The system of claim 20, wherein: The visual recognition neural network uses a rectified linear unit activation function.

25. The system according to claim 1 or 2, wherein: The system is configured to repeatedly generate a first high-dimensional representation (260), generate biologically relevant language-based output training data (250), and adjust the language identification machine learning algorithm (220) for each biologically relevant language-based input training data (102, 210) of a training set of biologically relevant language-based input training data sets.

26. The system of claim 25, wherein: The length of the first biologically relevant language based input training data (102, 210) of the training set of the biologically relevant language based input training data set is different from the length of the second biologically relevant language based input training data (102, 210) of the training set of the biologically relevant language based input training data set.

27. The system according to claim 1 or 2, wherein: The system is configured to repeatedly generate a second high-dimensional representation (330) and adjust the visual recognition machine learning algorithm (320) for each biologically relevant image based input training data (104, 310) of a training set of biologically relevant image based input training data sets.

28. The system of claim 27, wherein: The training set based on the input training dataset of biologically relevant language includes more entries than the training set based on the input training dataset of biologically relevant images.

29. A microscope comprising a system according to any preceding claim.

30. A method (900) for training a machine learning algorithm to process biologically relevant data, the method comprising: Receiving (910) input training data based on a biology-related language, wherein the input training data based on a biology-related language is at least one of: a nucleotide sequence or a protein sequence; generating (920) a first high-dimensional representation of the biologically relevant language-based input training data by a language identification machine learning algorithm, wherein the first high-dimensional representation includes at least three entries, each entry having a different value; generating (930) biologically relevant language based output training data from the first high-dimensional representation by the language recognition machine learning algorithm, wherein the biologically relevant language based output training data comprises a prediction of a next element in the biological sequence; adjusting (940) the language recognition machine learning algorithm based on the comparison of the biologically relevant language based input training data and the biologically relevant language based output training data; Receiving (950) biologically relevant image-based input training data associated with the biologically relevant language-based input training data, wherein the biologically relevant image-based input training data is image training data of an image of at least one of: a biological structure including nucleotides or nucleotide sequences; biological structures including proteins or protein sequences; generating (960) a second high-dimensional representation of the biologically relevant image-based input training data by a visual recognition machine learning algorithm, wherein the second high-dimensional representation includes at least three entries, each entry having a different value; and The visual recognition machine learning algorithm is adjusted (970) based on the comparison of the first high-dimensional representation and the second high-dimensional representation.