Spectral state space models

WO2025125364A1PCT designated stage expired Publication Date: 2025-06-19DEEPMIND TECH LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/EP2024/085753
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-12-11
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing machine learning models, particularly state space models, face challenges in efficiently processing long sequences and handling long-range dependencies, often resulting in unstable training and lower accuracy in predicting time-series data.

Method used

The proposed neural network model incorporates a spectral state space model with a spectral transform layer that applies multiple spectral filters to item embeddings, generating feature vectors and modifying them using weight matrices, allowing for efficient parallel processing and reduced memory requirements.

Benefits of technology

This approach enables the model to predict time-series data with high accuracy, stabilize the learning process, and operate effectively even with negative eigenvalues in the system matrix, outperforming traditional state space models in terms of learning rate and prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024085753_19062025_PF_FP_ABST
    Figure EP2024085753_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems for processing sequences using spectral state space models One of the methods includes, for successive time steps: processing an initial item embedding using an analysis network comprising processing layers arranged in a sequence, a first processing layer being configured to receive the initial item embedding, and to output a modified item embedding, and each other processing layer being configured to receive the item embedding output by the preceding layer and output a modified item embedding; wherein at least one of the processing layers is a spectral transform layer which, for each time step: generates a plurality of feature vectors by processing a sequence embedding using a plurality of spectral filters; multiplies the feature vectors by weight matrices, to form respective weighted feature vectors; and generates the modified item embedding for the time step including a term based on the weighted feature vectors.
Need to check novelty before this filing date? Find Prior Art

Description

SPECTRAL STATE SPACE MODELSCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U.S. Provisional Application No. 63 / 608,804, filed on December 11, 2023. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.BACKGROUND

[0002] This specification relates to processing data using machine learning models.

[0003] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model.

[0004] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.SUMMARY

[0005] This specification describes method and a system, implemented as computer programs on one or more computers in one or more locations, for processing a sequence of data items, for example to predict further data items in the sequence and / or to characterize the sequence (e.g. to recognize that content in the sequence of data items is in one of a set of categories). The method is particularly suitable for implementation in a parallel-processing computer system.

[0006] In general terms, a neural network model is proposed, i.e., an artificial neural network, which may be implemented on a computing system in software, or may be implemented in one or more integrated circuits as hardware. The network input to the neural network model is an input sequence of data items. Item embeddings based on the respective data items are successively processed by the processing layers of an analysis network, which successively modify each item embedding. The function performed by each processing layer is defined by a respective set of numerical parameters. At least one of the processing layers is a “spectral transform layer” (or “spectral analysis layer”) which applies (e.g., in parallel) a plurality of spectral filters (also called here simply “filters”) to a layerinput, to generate respective feature vectors. The feature vectors are combined with a weighting based on corresponding weight matrices, to form at least a term of the modified item embedding for the processing layer.

[0007] For at least one of the spectral transform layers, the input to the filters at a given time (time step) is a sequence embedding comprising (or consisting of) item embeddings received by the spectral transform layer and corresponding to one or more previously received data items (if any) of the network input, and optionally a data item received in a current time step.

[0008] A first processing layer of the analysis network receives “initial” item embeddings. The initial item embeddings may be obtained from respective ones of the data items of the network input as the output of an “embedding network” of the neural network model, upon receiving the data item as an input. Alternatively, in some embodiments the item embeddings may be equal to the corresponding data items (i.e., no embedding network is present).

[0009] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.

[0010] First, the proposed neural network model can be implemented efficiently using a parallel-processing computer system. This makes it more suitable for implementation in a parallel-processing system than some existing state space models.

[0011] Also, the neural network model may be capable of operating with a lower effective memory requirement than some previous neural network models which operate as state space models. This is particularly true if a matrix A, defined below, and which is part of a state space model, is symmetric. Furthermore, the present neural network model can operate stably in the case that the matrix^ includes negative eigenvalues.

[0012] Also, it has been experimentally demonstrated that the present models are capable of predicting time-series data with high accuracy and / or of generating, with high accuracy, data characterizing an input sequence. The learning operation is more stable than for some known state space models, and the learning rate is higher, e.g., time-series data can be predicted with greater accuracy for a given number of example data items.

[0013] A first specific expression of a concept disclosed by the present disclosure is a computer-implemented method for processing a sequence of data items corresponding to a plurality of time steps, the method employing a neural network model and comprising for successive time steps:processing an initial item embedding based on the data item for the time step using an analysis network of the neural network model comprising a plurality of processing layers arranged in a sequence, each processing layer performing a function defined by a corresponding set of trained numerical parameters, a first processing layer of the sequence being configured to receive the initial item embedding, and to output a corresponding modified item embedding, and each other processing layer of the sequence being configured to receive the item embedding output by the preceding layer of the sequence and output a corresponding modified item embedding; wherein at least one of the processing layers is a spectral transform layer which, for each current time step: generates a plurality of feature vectors by processing a sequence embedding using a respective plurality of spectral filters, the sequence embedding comprising as components the item embeddings received by the processing layer at one or more time steps preceding the current time step; multiplies the feature vectors by respective weight matrices defined by at least some of the corresponding set of trained numerical parameters, to form respective weighted feature vectors; and generates the modified item embedding for the current time step including a term based on the weighted feature vectors.

[0014] A second specific expression of a concept proposed by the present disclosure system for processing a sequence of data items corresponding to a plurality of time steps, the system comprising a neural network model which comprises: an analysis network comprising a plurality of processing layers arranged in a sequence, each of the processing layers performing a function defined by a respective set of numerical parameters, a first processing layer of the sequence being configured to receive the initial item embeddings, and output corresponding modified item embeddings, and each other processing layer of the sequence being configured to receive the item embeddings output by the preceding layer of the sequence and output corresponding modified item embeddings; at least one of the processing layers being a spectral transform layer configured, for each current time step of the sequence, to: generate a plurality of feature vectors by processing a sequence embedding using a respective plurality of spectral filters, the sequence embedding comprising as components the item embeddings received by the spectral transform layer for one or more time steps preceding the current time step;to multiply the feature vectors by respective weight matrices to form respective weighted feature vectors, the weight matrices being defined by at least some of the corresponding set of numerical parameters; and to generate the modified item embedding for the current time step including a term based on a sum of the weighted feature vectors.

[0015] The disclosure further proposes a method for training a system of this kind.

[0016] The proposed concepts may be expressed as a method, a system, or a computer program product (e.g. recorded on a tangible recording medium or downloadable over a communications system).

[0017] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Fig. 1 depicts an example neural network model incorporating a spectral transform layer.

[0019] Fig. 2 depicts the operation of an example spectral transform layer.

[0020] Fig. 3 depicts an example system for training a neural network model.

[0021] Fig. 4 is a flow diagram of a method carried out by a spectral transform layer.

[0022] Fig. 5 is a flow diagram of a method carried out by a neural network model.

[0023] Fig. 6 compares learning curves for a spectral transfer layer and a linear recurrent unit (LRU).

[0024] Fig. 7 shows an error is a prediction obtained a single spectral transfer layer as a function of a model parameter

[0025] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0026] Examples of the present disclosure will now be described. The examples are each a neural network model, implemented as computer programs on one or more computers in one or more locations, for processing a sequence of data items, for example to predict further data items in the sequence and / or to characterize the sequence.

[0027] Sequence to sequence (Seq2seq) models are a family of machine learning approaches use for processing an input sequence of data items, to generate an output sequence of data items. Example applications, as widely described in the literature, include sequence prediction, language translation, image captioning, conversational models and text summarization. Some Seq3seq models include a neural network encoder which maps an input sequence to a (real-numerical) vector, and a neural network decoder which maps the real-numerical vector to an output sequence, but the term is used in this document to embrace any model which processes an input sequence of data items, to generate an output sequence of data items. For example, transformer-based Seq2seq models discussed below often do not include an encoder-decoder architecture.

[0028] Handling long-range dependences efficiently remains a core problem in sequence prediction and modelling. Recurrent neural networks (RNN) are notoriously hard to trains, and often suffer from vanishing or exploding gradients. They are also hard to scale given the inherently sequential nature of their computation.

[0029] Transformer-based Seq2seq models which implement attention mechanisms have also been proposed. However, attention layers have memory and computation requirements which scale quadratically with context length (that is, the number of data items of the input sequence which are processed by the attention mechanism to generate a given data item of the output sequence).

[0030] A state space model is a type of sequence-to-sequence (seq2seq) model which has recently been proposed as an alternative to other seq2seq models, and it may be employed in the same wide range of applications as other seq2seq models.

[0031] In general terms, a state space model maps an input sequence of data items ut(e.g. a discrete form of a continuous time-dependent input signal u(t)), for corresponding time steps denoted t, where utis a vector with one or more components, to a corresponding sequence of hidden state (latent state) items xt, where xtis a vector with a dimensionality d (i.e., xtG IRd) which may be different from the dimensionality of ut. The hidden state items are projected to form an output sequence of output data items yt(a discrete form of a timedependent output signal y(t)). ytis a vector which may, or may not, have the same dimensionality as u(t) (e.g., both may have a single, time-dependent component). The mapping is as follows: xt+1= Axt+ Butyt+i= Cxt+ Dut)

[0032] The matrices A, B, C, D govern the evolution of the system and are called system matrices. Note that the term “time step” (which may alternatively be called simply a “step”) is not used to imply that the data items are associated with particular times (e.g., in the real world). Although this is one possibility, the sequence of data items may alternatively be a non-temporal sequence (e.g., a sequence of data items describing successive pixels of a still image, e.g., as a raster pattern across the image, or across each image frame of a moving image).

[0033] State space models have a wide range of known applications, and the present neural network model may be employed in any of the known applications of a state space model. Examples are given below.

[0034] An example of the present disclosure is a neural network model (a “spectral state space model”) which implements a state space model, and obtains a network output based on that data.

[0035] The neural network model receives an input sequence of data items. In the neural network model, item embeddings based on the respective data items are successively processed by a portion of the neural network model referred to as an “analysis” network and comprising a sequence of one or more processing layers. The processing layers successively modify each item embedding.

[0036] At least one of the processing layers is a “spectral transform layer” as discussed below, performing a function defined by a respective set of numerical parameters. The analysis network may further comprise one or more non-linear processing layers. For example, there may be a plurality of such layers, interleaved with the spectral transform layer(s). For example, the non-linear processing layers may alternate with the spectral transform layers. These non-linear processing layers may for example include multi-layer perceptron (MLP) and / or a gate linear unit (GLU). Each of the non-linear processing layer(s) also performs a function defined by a set of trained numerical parameters. The numerical parameters may optionally be trained, during the training process described above, jointly with the training of the sets of numerical parameters of the spectral transform layers (i.e., with substantially simultaneous updates to the numerical parameters of the nonlinear processing layer(s) and the spectral transform layer(s); or with updates to the numerical parameters of the non-linear processing layer(s) being interleaved with updates to the numerical parameters of the spectral transform layer(s)).

[0037] An example neural network model is depicted in Fig. 1. An input sequence 101 to the neural network model is a sequence of L data items, each of which has d components.

[0038] Each data item is encoded by an (optional) embedding layer 101 which the same for all data items and applied individually to each data item (“fixed across T”), to generate an initial item embedding 103 for the corresponding data item. Thus, there may be L initial item embeddings 103, each of which has a plurality of components. The L initial item embeddings 103 may optionally be generated sequentially in respective time steps. The embedding layer 101 is an example of an embedding network which, in other implementations, may include multiple layers and optionally additional inputs.

[0039] The initial item embeddings 103 are processed by an analysis network 104, including at least one spectral transform unit (STU) which is a spectral transform layer 105. As depicted in Fig. 1, the analysis network 104 is made up of a sequence of X processing units (where X is an integer which is at least one), and each processing unit comprises a spectral transform layer 105, followed by a second “non-linear” processing layer 106. The non-linear processing layer 106 may, for example, process the output of the spectral transform unit 105 of the processing unit by performing the operation of a multi-layer perceptron (MPO) and followed by a non-linear operation. The output from each processing unit (except the last) is the input to the next processing unit of the sequence. Thus, the analysis network 104 comprises (or consists of) 2X processing layers, divided into X processing units, where each processing unit is a spectral transform layer 105 followed by a non-linear processing layer 106.

[0040] The first spectral analysis layer 105 of the analysis network 104 receives at each time step a corresponding initial item embedding 103. It processes the initial item embeddings 103, to generate modified item embeddings. Each other of the processing layers of the analysis network 104 (e.g. the processing layer 106 of each processing unit, and the spectral transform layer 150 of each processing unit but the first) receives, at each time step, an item embedding output by the respective preceding processing layer, and processes it to generate and output a corresponding modified item embedding.

[0041] Each spectral transform layer 105 operates in a number of time steps. In each time step (e.g. denoted by an integer index t in the range 1 to E) it processes together the first t of item embeddings 103 it has received (i.e. the items embeddings 103 for time steps 1, . . 7), to generate a modified item embedding for the time step.

[0042] For example, the spectral transform layer 105 of the first of the X processing units, receives, one by one in respective time steps Z, the corresponding initial item embeddings 103 generated by the embedding layer 102, and processes together the item embedding it receives in each time step t and the item embeddings it received in the previous t-1 time steps, to generate a modified item embedding for the current time step.

[0043] Each of the other spectral transform layers 105 (i.e. the respective spectral transform layer 105 of the each of the X-l processing units other than the first processing unit), receives, one by one in respective time steps Z, the item embeddings generated by the second processing layer 106 of the preceding processing unit, and processes together the item embedding it receives in each time step Z and the item embeddings it received in the previous t-1 time steps, to generate a modified item embedding for the current time step.

[0044] As discussed below, each spectral transform layer 105 may also, to generate the modified item embedding for the current time step, process modified item embedding(s) it has generated in previous time steps (i.e. it may be auto-regressive).

[0045] Each second processing layer 106 applies a different corresponding function defined by a respective set of numerical parameters. Each second processing layer 106 performs the corresponding function (e.g. a MLP operation followed by a non-linear operation) individually to each item embedding it receives from the preceding respective spectral transform layer 105, to generate a modified item embedding. Thus, the second processing layer 106 is also “fixed across T”. The second processing layer 106 of the last of the X processing units generates embeddings 107. There are L such embeddings, corresponding to respective ones of the data items 101, and respective ones of the initial item embeddings 103. Each of the embeddings 107 has multiple components.

[0046] In some applications, the item embeddings output by the last processing layer 106 of the analysis network may constitute a network output of the neural network model. This possibility is particularly appropriate if the neural network model is to function as a seq2seq model.

[0047] Alternatively, the L embeddings 107 may be processed by an output network 108. The output network 108 is configured to generate a network output 111 based on one or more of the item embeddings generated by the last processing layer 106 of the analysis network 104. The output network 108 and network output 111 may take many forms depending upon the application.

[0048] In the example of Fig. 1, the output network 108 performs a time pool operation to combine the / . embeddings 107 into a single multi-component embedding 109.The multi-component embedding 109 is then processed by a dense layer 110 (again defined by numerical parameters) to generate a network output 111.

[0049] The network output 111 may, for example, specify a classification for the input sequence of data items 101, i.e. an indication of one of a number of classes. For example, the network output 111 be a vector having a number of components corresponding to a number of pre-defined classification classes, and the network output 111 may be a vector (e.g. a “one-hot”) vector which specifies one of the classification classes. For example, if the input sequence of data items 101 represents a sequence of images (a video, e.g. of the real world captured by one or more cameras), the network output 111 may specify which of a number of categories of activity is depicted in the video.

[0050] For example, if the sequence of data items 101 defines a video segment (i.e. a sequence of images, in which each image is at least one value for each pixel of an (at least) two-dimensional array of pixels), the network output 111 may be a video captioning output. The video captioning output can include a natural language output sequence, e.g., a sequence of words, that is descriptive of the video segment.

[0051] As another example, the output 111 for each video segment is an action recognition output. An action recognition output for a video segment can recognize an action that spans multiple video frames in the video segment. The action can for example be an action, activity, and / or other temporally varying occurrence which involves a human actor and / or a non-human actor, such as an animal, a robot, an inanimate object, or portions thereof.

[0052] When the action recognition output of the video is expressed as text, the action recognition output of the video may include at least one verb that is descriptive of the action, activity, and / or other temporally varying occurrence. Action recognition outputs may be used to facilitate video retrieval, video captioning, and / or visual question-and-answer, among other tasks.

[0053] As another example, the output 111 for each video segment may be an action localization output. The action localization output can identify an action spatially, e.g., by defining the coordinates of bounding boxes that enclose respective actions depicted in the video frames. Additionally or alternatively, the action localization output can identify an action temporally, e.g., by identifying one or more time step within the video segment during which the action is depicted in the corresponding video frames.

[0054] As another example, the output 111 for each video segment is an object detection output. The object detection output can identify regions within each respectivevideo frame in the video segment that are predicted to include objects. For example, the object detection output can include data defining a plurality of bounding boxes in a video frame; and, optionally, for each of the plurality of bounding boxes, a respective confidence score that represents a likelihood that an object belonging to an object category from a predetermined set of one or more object categories is present in the region of the video frame shown in the bounding box.

[0055] As noted, the function performed by each processing layer 105, 106 is defined by a respective set of numerical parameters, and at least one of the processing layers is a spectral transform layer 105. As discussed below with reference to Fig. 2, at each time step each spectral transform layer 105 applies (e.g., in parallel) a plurality of filters to a sequence embedding comprising (or consisting of) item embeddings corresponding to one or more earlier data items it has received (if any), to generate respective feature vectors. The feature vectors may also include the item embedding the spectral transform layer 105 received in current time step. The feature vectors are combined with a weighting based on corresponding weight matrices, to form at least a term of the modified item embedding for the processing layer. The numerical parameters of the spectral analysis layer may define the weight matrices.

[0056] The embedding layer 102 (if any) and output network 108 (if any) may each be defined by respective sets of additional numerical parameters. As discussed below with reference to Fig. 3, the training process may comprise iteratively modifying the additional numerical parameters to reduce a loss function. This iterative modification may be performed jointly with the training of the processing layers of the analysis network (i.e., with simultaneous updates to all the numerical parameters, or interleaved updates to different corresponding (proper) subsets of the numerical parameters).

[0057] The weight matrices of each spectral transform layer 105 (which may be different for each spectral transform layer) may be defined by some or all of the set of numerical parameters for the spectral transform layer 105.

[0058] Note that the modified item embeddings produced by a given processing layer 105, 106 (e.g., a spectral transform layer 105) may have a dimensionality, e.g., denoted dout, which is not the same as the dimensionality, e.g., denoted din, of the item embeddings received by the given processing layer 105 106. The values of dinand doutmay be different for different ones of the processing layers.

[0059] A description is now presented, with reference to Fig. 2, of the possible operation of a single spectral transform layer 105. This description is used explain a possible expression (Eqn. (7) below) for the output of the spectral transform layer 105 based on the item embedding input to the single spectral transform layer. In this discussion, the notation ut and y, (or yt) is used, which is not to be confused with the use of the same two variables in other parts of this document to describe respectively the inputs and outputs of the neural network model 100 as a whole.

[0060] In this implementation of the present concept, current time is denoted t which is assumed to be in the range 1 to / .. The spectral transform layer 105 includes a plurality of spectral filters 202, more simply also called “filters”. The number of filters is denoted K which is greater than one, and the filters may be denotedwhere k is an integer index in the range 1 to an integer K. Note that the set of filters may optionally be different for different ones of the spectral transform layers 105 (if there is more than one spectral transform layer 105 in the neural network model), in which case (pkwould have a further index labelling one of the spectral transform layers. However, this omitted for simplicity. The filters may alternatively be the same for each of the spectral transform layers 105.

[0061] In a given time step, each filter can process up to L of the item embeddings. If the number of preceding time steps t-1 is less than L-l, the first t components of each filter are applied respectively to the item embedding received in the current time step t and the item embeddings received in the preceding t-1 time steps. Thus, other words, at each time step / , each filter is applied to up to L item embeddings received by the spectral transform layer (i.e., in the item embedding received in the current time step and up to L-l item embeddings received in the t-1 preceding time steps). (Note that, in alternative implementations of the present concept, L may be greater than the number (say J) of item embeddings a given filter can process in a certain time step, and in this case, if the number of preceding time steps t-1 is at least J-l, then each filter is applied to the item embeddings received in the current time step and the J-l preceding time steps. However, this possibility is not considered further here).

[0062] At a current time t, the item embedding the spectral transform layer receives is denoted utE IRdin, and the item embeddings the spectral transform layer has previously received are {u ...These t data items are shown in Fig. 2 as the input signal 201.

[0063] Some or all of the item embeddings 201 are formed into a sequence embedding. For example, for a current time step t, the sequence embedding include all {i^ , ... ut-, ut}. In alternative implementations, the feature embedding may not include the item embedding utreceived for the current time step and / or the preceding time step, e.g., the sequence embedding may be {u^ ... ut-2], i.e., not including the item embedding utfor the current time step, or the item embedding ut-for the preceding time step.

[0064] From the sequence embedding, the spectral transform layer 105 generates an output ytE IRd°ut. This is a modified item embedding corresponding to the received item embedding ut. Thus, the output of the spectral transform layer in the current time step and the t-1 preceding time steps is a sequence of corresponding modified item embeddings {y1(... yt] 206 where each ytE Ikdout. Note that in some cases in this document, notation ytis used in place of yt, to denote an actual output of the neural network model 100.

[0065] Each filter <pk202 may be defined by a respective set of L filter components, (pkE IRL. That is, the number of filter components may be the same as the maximum number of item embeddings in the sequence embedding. Multiplying the feature vector (e.g. the input signal 201 if all received item embeddings in the input signal are used to form the sequence embedding) component- wise with each filter 202 gives a respective feature vector 203. Each feature vector 203 may be a sum, over components of the sequence embedding, of a product of the corresponding received item embedding and a corresponding component of the respective spectral filter. The feature vectors 203 are multiplied by corresponding weight matrices 204, and the result summed by an addition unit 205, to produce the respective output yt206. In one possibility, there is single weight matrix denoted } for each respective filter (and this possibility is assumed in Fig. 2 where the filters 204 are indicated asHowever, other possibilities exist. For example, as described below, there may for each filter k be two weight matrices M^+

[0066] The K filters {(pk} 202 may be chosen to be the eigenvectors of a predetermined matrix, such as one having entirely positive eigenvalues. For example, the matrix may be the Hankel matrix. This is real-valued matrix with L X L components, given by:

[0068] The corresponding eigenvalues of this matrix, denoted {o / , are all real values, and <J > <JffL- The K eigenvectors corresponding to the largest Teigenvalues are used as the K filters. Note that K may be less than L so that not all eigenvectors of the Hankel matrix are used as filters.

[0069] Some (or all) of the feature vectors may be “first feature vectors”, formed as a sum, over components of the sequence embedding, of a product of the corresponding received item embedding and a corresponding component of the respective spectral filter.

[0070] For example, the feature vectors may include a plurality of first feature vectors 203 defined as:

[0071] Note that, as the value of t rises, correspondingly more components of each filter are used to work out the corresponding first and second feature vector for the sequence embedding. In a variant of the current implementation, in which L is greater than the dimensionality of the filter (i.e., the current value for t is greater than the number of components J of the filters), then, as noted above, the sequence embedding may be generated based a number of the most recently received item embeddings which is equal to the number of dimensions of the filter. However, for simplicity, this possibility is not considered in the following discussion.

[0072] A term of the modified item embedding yt206 for the current time step t, is based on a sum produced by addition unit 205, over the filters, of a product of a first feature vector for the filter and a respective weight matrix 204 for the filter. The components of the sum are optionally further weighted by a function of the respective eigenvalue ok. For example, the term of the modified embedding yt206 for the current time step t may be defined as

[0074] Here(a doutx dinmatrix of real values) is the weight matrix 204 for the &-th first feature vector, i.e., the first feature vector generated using the &-th filter). The components of M^+are a (proper) subset of the set of numerical parameters defining the spectral transform layer. These numerical parameters iteratively trained during the training procedure.

[0075] Note that this term of the modified item embedding is based on first feature vectors {Af+_2 k} generated using a sequence embedding which includes the item embeddings received up to two time steps before the current time step t (i.e., not including the item embedding received in the current time-step or the preceding time step). In the casethat t is less than 3, this term of the modified item embedding may be set to a default vector, e.g., a vector of zeros (e.g., a vector of dout zeros).

[0076] As noted, this expression for this term of the modified item embedding includes an optional value o-,1 / 4. If such a value is included, it may alternatively take another value which varies inversely with the corresponding eigenvalue of the matrix having the filters as eigenvectors (e.g., the Hankal matrix).

[0077] Alternatively or additionally, some or all of the feature vectors 203 may be “second feature vectors”, formed as a sum, over components of the sequence embedding, of a product of (i) the corresponding item embedding, (ii) a corresponding component of the respective spectral filter, and (iii) a value which alternates in sign for successive components of the sequence embedding. For example, the feature vectors 203 may include a plurality of second feature vectors:

[0079] A term of the modified item embedding ytmay then be defined as

[0081] This provides a negative part of the spectral component for the k-th filter.Here Mk~(a doutX dinmatrix of real values) is the weight matrix 204 for the &-th second feature vector, i.e., the second feature vector generated using the A th filter. The components of Mk~ are a (proper) subset of the set of numerical parameters defining the spectral transform layer. These numerical parameters are iteratively trained during the training procedure.

[0082] Note that this term defined by Eqn. (6) of the modified item embedding ytis based on second feature vectors {AfL2 k} generated using a sequence embedding which includes the item embeddings received up to two time steps before the current time step t (i.e., not including the item embedding received in the current time-step or the preceding time step). In the case that t is less than 3, the feature vector XfL2 kmay be replaced with a vector of zeros.

[0083] Again, this expression includes an optional value■ Again, if such a value is included, it may alternatively take another value which varies inversely with the corresponding eigenvalue of the matrix having the filters as eigenvectors (e.g., the Hankal matrix).

[0084] The modified item embedding ytmay further include one or more autoregressive terms which together provided the modified item embedding ytwith an autoregressive component that allows for stable learning of the spectral component as the memory grows.

[0085] A first of these may be the modified item embedding generated by the spectral transform layer at a preceding time step, such as the modified item embedding two steps earlier yt-2.

[0086] Another auto-regressive term may be based on the item embedding utfor the current time step received by the spectral transform layer. This may be weighted by a second weight matrix, which may be denoted M . The components of M are a (proper) subset of the set of numerical parameters defining the spectral transform layer. These numerical parameters are iteratively trained during the training procedure.

[0087] Another auto-regressive term may be based on an item embedding received by the spectral transform layer for one of the recent preceding time step(s), e.g., the item embedding ut-for the immediately preceding time step, and / or the item embedding ut-2for the time step which is two time steps earlier. This / these may be weighted by a corresponding weight matrix, which may be denoted M2and M respectively. The components of M2and M (when they are present) are each a (proper) subset of the set of numerical parameters defining the spectral transform layer. These numerical parameters are iteratively trained during the training procedure.

[0088] One possibility for the modified item embedding yt, which combines several of the possibilities discussed above is:

[0089] Here yt-2isanauto-regressive component, andS iMt+ (Jl / 'xt-2,k + S iMt~ °k ^ t-2,k is a spectral component.

[0090] This system of equations is particularly suited for implementation in a parallel processing system, that is with multiple processors operating in parallel. This is because different components of the input item embedding ut_i (input dimensions) may initially be processed separately in parallel, followed by an addition operation. Furthermore, the effects of the different spectral filters may be processed separately in parallel, followed by an addition operation. The multiple processors may be a multi-core processors of an integrated circuit, and / or multiple integrated circuits provided within a single computer apparatus (e.g., sharing a clock signal and / or a power supply and / or on a single printedcircuit board), or multiple items of computer apparatus (e.g., with different respective clock signals, power supplies, and / or printed circuit boards). The exact way in which the parallel implementation is implemented by multiple processors operating simultaneously (i.e., the selection of which processor performs which operations) may be selected based on the different capacities of the multiple processors.

[0091] Turning to Fig. 3, a training system 300 is shown for training an example neural network model including a spectral transform layer. The neural network model 305 may be the neural network 100 explained above with reference to Fig. 1.

[0092] The training system 300 employs a training database 301 of training examples which are sample data item sequences and corresponding sample target vectors. The target vectors may be possible network outputs which it would be desirable for the neural network model 305 (e.g. the neural network model 100) to produce as a network output (i.e., the output of the analysis network 104 or the output network 108 if any), upon receiving the corresponding sample data item sequences as a network input (e.g. as the input sequence 101).

[0093] The training operation may comprise iteratively modifying the weight matrices (i.e., the numerical parameters defining them) and / or the sets of numerical parameters defining the processing layers 106 (and optionally numerical parameters defining the embedding layer 102 (if any) and / or numerical parameters defining the output network 108 (if any)), to reduce a loss function characterizing a discrepancy between (i) the network output (i.e., values based on item embeddings output by the last layer of the analysis network upon the embedding network receiving the sample data item sequences), and (ii) the corresponding sample target vectors. The discrepancy may be defined, for example, as a Euclidean distance measure (or any other distance measure, such as a Manhattan distance measure). The loss function may be defined for each time step, denoted t, for example as a measure of the discrepancy between the network output for that time step (which may be denoted yt) and the target vector yt.

[0094] The training system 300 includes a data item generation model 302 which extracts from the training database 301 a sample data item sequence and corresponding sample target vector. The data item generation model 302 transmits the sample data item sequence to be the network input of the neural network model 305, and transmits the corresponding sample target vector to a loss calculation module 303. The loss calculation module 303 also receives the network output of the neutral network model 305, and calculates the discrepancy between them. This process may be repeated for a batch oftraining examples, i.e. each of multiple time steps for single sample data item sequence and corresponding sample target vector, and / or plural sample data item sequences and corresponding sample target vectors. The loss calculation module 303 sums the respective discrepancies over the batch, and generates a loss function which is transmitted to a numerical parameter update module 304.

[0095] The numerical parameter update module 304 updates the numerical parameters defining the neural network module 305. In this case that the neural network module 305 is the neural network module 100 of Fig. 1, the numerical parameter update module 304 updates the numerical parameters defining the weight matrices of the spectral transform layer(s) 105 (but not the filters, which may not be changed in the training procedure) and / or the numerical parameters defining the processing layer(s) 106. It may optionally additionally or alternatively update numerical parameters defining the embedding layer 102 and / or the output network 108.

[0096] This process is then repeated iteratively, to successively reduce the loss function. The iterative process may be performed until a termination criterion is met. The termination criterion may, for example, be that a certain number of the most recent iterations have reduced the loss function by less than a threshold amount. Alternatively, the termination criterion may be that the iterative process has used a threshold number of computing operations or taken a threshold amount of computing time by the computer system which implements it.

[0097] Turning to Fig. 4, a method 400 is illustrated which is performed by a spectral transform layer. For example, the spectral transform layer 105 shown in Fig. 1 and Fig. 2 may perform the method 400. Method 400 is an example of a method performed by one or more computer systems in one or more locations. For example, the method 400 may include activity performed in parallel by multiple processors (e.g. multi-core processors of a single integrated circuit, and / or multiple integrated circuits provided within a single computer apparatus, and / or processors of multiple computer apparatuses). The method 400 is performed at each of a series of time steps, e.g. labelled by corresponding values of the integer variable t.

[0098] In a first step 401, the spectral transform layer receives an item embedding for the current time step. This may be from an embedding network (e.g. the embedding layer 101) or from a preceding layer (if any) of an analysis network of which the spectral transform layer is part.

[0099] In step 402, the spectral transform layer processes a sequence embedding. The sequence embedding comprises, as components, item embeddings received by the processing layer one or more time steps preceding the current time step, and optionally the item embedding received at the current time step (e.g. ut). Thus, if the spectral analysis layer is the first spectral transform layer 105 of the analysis network 104, the sequence embedding comprises initial item embeddings 103 produced by the embedding layer 102. Alternatively, if the spectral analysis layer is another spectral transform layer 105 of the analysis network 104, the sequence embedding comprises item embeddings 103 output by the immediately preceding one of the processing layers 106.[000100] In step 402, the spectral transform layer processes the sequence embedding using a plurality of spectral filters, to generate a plurality of feature vectors. For example, as shown in Fig. 2, the input signal 201 to the spectral transform layer (feature embedding) is processed using the K filters 202, to generate a plurality of feature vectors 203.[000101] In step 403, the feature vectors are multiplied by respective weight matrices (e.g. the weight matrices 204 of Fig. 2) defined by at least some of the corresponding set of trained numerical parameters, to form respective weighted feature vectors.[000102] In step 404, a modified item embedding (e.g. yj for the current time step is generated including a term based on the weighted feature vectors. Eqn. (7) is an example of how this case be done. Two other examples are given below (Eqns. (8) and (9)).[000103] Turning to Fig. 5, a method 500 is illustrated which is performed by a neural network model. For example, the neural network model 100 shown in Fig. 1 may perform the method 500. Method 500 is an example of a method performed by one or more computer systems in one or more locations. The method 500 is performed at each of a series of time steps, e.g. labelled by corresponding values of the integer variable t.[000104] In a first step 501, the neural network model receives a data item for the current time. For example, for the neural network model 100 of Fig. 1, this would be one of the data items in the input sequence 101 (i.e. the data item of the input sequence 101 corresponding to time f).[000105] In step 502, an initial item embedding is generated based on the data item by an embedding network. For example, in the case of the neural network model 100 of Fig. 1, step 502 is performed by the embedding layer 102, to generate the initial data embedding 103.[000106] In step 503, the initial item embedding is processed using an analysis network comprising a sequence of processing layers, which generate respective modifieditem embeddings. For example, for the neural network model 100 of Fig. 1, the analysis network 104 processes the initial item embedding 103, and comprises a sequence of processing layers 105, 106. The operation of each spectral transform layer 105 of the analysis network for the current time step may be the method 400, explained above with reference to Fig. 4.[000107] In step 504, a network output is generated based on a modified item embedding generated by the last processing layer of the analysis network. For example, for the neural network model 100 of Fig. 1, the modified item embedding 107 generated by the last non-linear processing layer 106 of the analysis network 104 is used to generate a network output 111.[000108] Fig. 6 shows experimental results of the learning curve for a neural network model as shown in Fig. 1 with a single spectral transform layer “STU” 105 (as explained above with reference to Fig. 2). This may be the neural network model 100 of Fig. 1 with the value of X equal to one. The quality of learning is measured by a reconstruction loss denoted “L2 loss”. The experiment was performed with a low-dimensional linear system as defined by Eqn. (1), with the components of the matrices B 6 Ik4x3and C 6 Ik3x4being random variables which are chosen independently from identical Gaussian distributions. Matrix D was a diagonal matrix with idd (independent and identically distributed) Gaussian entries and A was a diagonal matrix with i4^~0.9999 * Z where Z was a random sign. The training system was as shown in Fig. 3, using a variable number of training examples in the training database 301. Each training example was generated with a random input sequence utand with ytbeing the result of applying Eqn (1) with the random matrices. Mini-batch (batch size 1) training was performed using an L2 loss.[000109] For comparison, Fig. 6 also shows training loss results for a single linear recurrent unit (LRU) (as proposed by A. Orvieto, et al., “Resurrecting recurrent neural networks for long sequences”, 2023) which directly parametrizes the linear system.[000110] The experiments were carried using the Adam optimizer (D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization”, 2014). The learning rate of the Adam algorithm was turned for both models.[000111] In both the LRU and the STU, the initialization / normalization techniques proposed by A. Orvieto et al. (2023) were used, including stable exponential parameterization, / -normalisation, and ring-initialization.[000112] The STU training used as a hyper-parameter the learning rate of the Adam algorithm, which was selected from the set [2xl0‘2, 10’1, 5xl0-1, 1, 5, 10],[000113] It was observed that for the STU the L2 loss decreased rapidly for a number of samples under 100. By contrast, the LRU shows a plateauing of learning for a number of samples up to 250, and needed a much larger number of samples to achieve the same L2 loss as the STU.[000114] Fig. 7 shows the error, in log-scale, obtained by the neural network model with the single STU as a function of the model parameter K indicating the number of filters. An exponential drop in the L2 loss was observed with respect to K.[000115] A number of variations are possible to the spectral transform layer explained above with reference to Fig. 2. For example, one extension is to parametrize the dependence of yton yt-2with a matrix Mydefined by a further set of numerical parameters, which may be iteratively varied during the training. The modified item embedding in this case is ytgiven by:The numerical parameters defining Mymay be iteratively varied jointly with the other numerical parameters, e.g. during training by the training system explained above with reference to Fig. 3.[000116] In a further variation, the auto-regression may be extended to depend on multiple previous yt, as opposed to just yt-2as in Eqns. (7) and (8). It can be shown that adding sufficiently long auto-regression is powerful enough to capture any linear dynamical system (LDS).[000117] Based on this, a further implementation of the spectral transform layer may be performed called auto-regressive STU (AR-STU) where the output of the STU is given byThat is, the auto-regressive component depends on a parameter ky. The output of the STU for a current time t is a function which includes an auto-regressive term including a respective component Myyt_ifor each of the immediately preceding kytime steps (labelled by integer index z), and the component is based on a respective matrix Mydefined by a respective set of numerical parameters which are trained during the training process.[000118] As noted, Figs. 6 and 7 show the case of a single STU layer. Experiments were also carried out for the “stacked” case, e.g. a model as shown in Fig. 1 in which X is greater than one. The experiments again followed the approach of A. Orvieto, et al. (2023), replacing the LRU layers described there by STU layers. In an experiment using a model which is an example of Fig. 1, the input sequence is first embedded via a time-invariant embedding function, followed by multiple repetitions of alternating STU layers and nonlinear layers (GLUs were used). Finally, the resulting output is time-pooled followed by a final readout layer based on the task the model is being trained to perform. This composite model is trained in a standard fashion via back-propagation and other commonly used deeplearning optimization techniques. The experiments were performed based on the Long Range Area (LRA) benchmark of Yi Tay, et al., “Long range area: A benchmark for efficient transformers”, in International Conference on Learning Representations, 2021. This is benchmark uses six tasks, called ListOps, Text, Retrieval, CIFAR, Pathfinder and PathX.[000119] Both “vanilla” STU layers (e.g. as described above with reference to Eqn. (7)) and AR-STU layers were tried out. In the case of AR-STU layers, the experiments were carried out for the case ky= 2 (i.e. the case of Eqn. (8)) and the case ky= 32.[000120] For both the case of using vanilla STU layers, and the case of using AR-STU layers, multiple values of two hyper-parameters were tried out. First, the learning rate of the Adam algorithm, which was selected from the set [2xl0‘2, 10’1, 5xl0-1, 1, 5, 10], Other optimization hyperparameters of the Adam algorithm were set to their default values. Second, a weight decay parameter (used as in A. Orvieto, et al. (2023)), which was selected from the set [ 10"3, 10'2, 10’1, 5xl0-1, 1.0], The number of filters K was fixed at 24.[000121] For the case of using STU layers, superior performance (compared to the LRU of A. Orvieto, et al. (2023)) was observed for the tasks ListOps, Text and Retrieval. [000122] In the case of AR-STU layers, improved performance over STU layers was observed for all of the tasks except Text. For AR-STU layers, it was found that the setting ky= 2 was sufficient to obtain optimal results for non-image tasks, ListOps, Text and Retrieval tasks. However, for image tasks, CIFAR, Pathfinder and PathX, it was found that ky= 32 lead to significant performance gains. In this case, the AR-STU outperformed LRU in the case of CIFAR and Pathfinder also.[000123] In general terms, the neural network model described above may be used for any of the applications, including (but not limited to) those any application for whichseq2seq models have been employed. A discussion is now presented of some of these applications. In many of these applications, it is more computationally efficient (e.g., efficient in terms of memory resources and / or computer operations) than some other known systems, particularly in the case that the input sequence of data items has long-range time dependencies. In this discussion, the values {u ... ut, ... , uL], where t is an integer index in the range 7 to L, are used to denote vector inputs to the neural network model (e.g., the inputs to the embedding network, if any, or the analysis network otherwise), and the values {y1(■■■ yt> ■■■ ’ VL}areused to denote the outputs of the neural network model (e.g., the outputs of the output network, if any, or otherwise the analysis network).[000124] In many of these applications, image embedding outputs of the analysis network, or network outputs of neural network model as a whole, can be fed back recursively as, or to form, new inputs to the analysis network. In this way, the neural network model can be used recursively. For example, outputs of the output network (if any) can be fed back as inputs to the embedding network (if any); or image embedding outputs of the last layer of analysis network can be fed back to be image embeddings which are input to the first layer of the analysis network. The first one or more data items of the sequence of data items may be used to condition the generation of initial outputs of the analysis network, while later data items of the sequence (i.e., data items of the sequence which are after the first one or more data items) may be based on, or may be, image embedding outputs of the last layer of the analysis network.[000125] A first application of the present neural network model is for the simulation (e.g., prediction of future states of) and / or control of an environment. The environment may be a physical system, such as real-world physical system. The physical system may for example, be a linear dynamic system which is described by Eqn. (1). In this case, at least some of the data items ut (e.g., one or more data items which are at the start of the sequence of data items) may comprise “observation data items”, which are data items characterizing a state of the environment (e.g., sensor data captured by a sensor, such as a (still or moving) camera or a microphone or a medical sensor) at corresponding times, and / or data characterizing inputs (e.g., forces or voltages exerted on) to the environment at those corresponding times. For example, if the environment comprises one or more objects, a given observation data items may characterize the spatial positions, relative spatial positions, orientations and / or relative orientations of the objects at the corresponding time. The outputs of the neural network model yt may be predicted observation data items whichcharacterize the state of the environment, e.g., at a current prediction time which is later than the time(s) corresponding to the data item(s) which were used to generate them. [000126] Predicted observation data item(s) y, generated by the neural network model from one or more data items at the start of the sequence of data items (e.g., data items which encode sensor data), may be used as, or to generate, data items later in the sequence of data items (“predicted item embeddings”). The predicted item embeddings may be processed by the neural network model to generate new predicted observation data item(s). The new predicted observation data item(s) may characterize the state of the environment at corresponding time(s) later than the times corresponding to the predicted observation items. In other words, the neural network model may be used recursively, conditioned on the one or more data items at the start of the sequence of data items, to generate an ongoing sequence of predictions progressively further into the future.[000127] The predictions may be used in various ways. For example, they may be used to generate characterization data which is displayed to a user, and which characterizes the predictions. For example, the characterization data may be classification data indicating that the predictions fall into one or more (e.g., predetermined) categories. For example, the classification data may indicate that the predictions are in one or more categories, e.g., classifications associated with undesirable behaviors of the environment. Upon the neural network model generating classification data indicating that the predictions are in one of these categories, a corresponding warning may be issued to the user. In one example, if the environment is a system which predicts the future levels of water in a river system based on present measurements and / or data indicating rainfall levels, the classification data may indicate that the predictions are in a category associated with an increased risk of flooding, and based on this classification data a flood warning may be issued to a user, and / or to people located in the region affected by potential flooding.[000128] Another way of using the predictions is to control physical equipment (e.g., an electromechanical system) based on the predicted data items. For example, in the case of controlling water levels in a river system, the predicted data items may be used to control water control apparatus such as a dam, e.g., to reduce the risk of flooding.[000129] Another example of an environment which can be predicted is the weather in a certain geographical location. Based on a weather prediction, a warning can be broadcast to individuals in the geographical location. Furthermore, in a way similar to the system discussed at https: / / deepmind.google / discover / blog / machine-leaming-can-boost-the-value- of-wind-energy / the neural network model can be trained on weather forecasts and historicalturbine data, to predict wind power output ahead (e.g., 36 hours ahead) of actual generation. The predictions can be used to predict optimal delivery of power to the power grid in advance. This allows the operation of other power generation systems which supply power to the power grid to be controlled. Furthermore, the weather predictions may be used to control other electromechanical apparatus, e.g., to open or close shutter mechanisms.[000130] Another application of the present neural network model is in a control system for a controlled system.[000131] For example, the neural network model can also be used in a linear dynamic controller. This can also be represented by Eqn. (1) in which the output signal y(t) represents one or more control variables of the controlled system, and u(t) represents a state of the controlled system, or observation data presenting a state of the controlled system, and the (initially unknown) matrices A, B, C and D. During the training of the neural network model, the analysis network is trained to predict desirable values y(t) of the control variable(s), for given values of u(t).[000132] In another example, u(t) is an observed control input to a dynamic system, and t) is an observation of a system (e.g., an observed linear transformation of the state). During the training of the neural network model, the analysis network is trained to predict a state y(t) of a controlled system, for given choices of the control input u(t).[000133] In some implementations the controlled system is the real-world environment of a service facility comprising a plurality of items of electronic equipment, such as a server farm or data center, for example a telecommunications data center, or a computer data center for storing or processing data, or any service facility. The service facility may also include ancillary control equipment that controls an operating environment of the items of equipment, for example environmental control equipment such as temperature control, e.g., cooling equipment, or air flow control or air conditioning equipment. The neural network model may be used to control operation of the items of equipment, or to control operation of the ancillary, e.g., environmental, control equipment. This may be done in such a way as to minimize, use of a resource, such as electrical power consumption or water consumption. [000134] More generally, the controlled system may be an agent, and the control system may select actions for the agent to perform. The agent may be an electro-mechanical agent which interacts with an environment (e.g., a robot, which may be capable of changing its configuration and of navigating within the environment). In these implementations, the observations may include, e.g., one or more of: images, object position data, and sensor data to capture observations as the agent interacts with the environment, for example sensor datafrom an image, distance, or position sensor or from an actuator. For example in the case of a robot, the observations may include data characterizing the current state of the robot, e.g., one or more of: joint position, joint velocityjoint force, torque or acceleration, e.g., gravity- compensated torque feedback, and global or relative pose of an item held by the robot. In the case of a robot or other mechanical agent or vehicle the observations may similarly include one or more of the position, linear or angular velocity, force, torque or acceleration, and global or relative pose of one or more parts of the agent. The observations may be defined in 1, 2 or 3 dimensions, and may be absolute and / or relative observations. The observations may also include, for example, sensed electronic signals such as motor current or a temperature signal; and / or image or video data for example from a camera or a LIDAR sensor, e.g., data from sensors of the agent or data from sensors that are located separately from the agent in the environment. Furthermore, the observations may include sound signals (e.g., collected by one or more microphones of the agent, or microphones that are located separately from the agent in the environment), such as voice commands issued by a human user.[000135] In these implementations, the actions may be control signals to control the robot or other mechanical agent, e.g., torques for the joints of the robot or higher-level control commands, or the autonomous or semi-autonomous land, air, sea vehicle, e.g., torques to the control surface or other control elements, e.g., steering control elements of the vehicle, or higher-level control commands. The control signals can include for example, position, velocity, or force / torque / acceleration data for one or more joints of a robot or parts of another mechanical agent. The control signals may also or instead include electronic control data such as motor control data, or more generally data for controlling one or more electronic devices within the environment the control of which has an effect on the observed state of the environment. For example in the case of an autonomous or semi-autonomous land or air or sea vehicle the control signals may define actions to control navigation, e.g., steering, and movement e.g., braking and / or acceleration of the vehicle.[000136] In some implementations the environment is a real-world manufacturing environment for manufacturing a product, such as a chemical, biological, or mechanical product, or a food product. As used herein a “manufacturing” a product also includes refining a starting material to create a product, or treating a starting material, e.g., to remove pollutants, to generate a cleaned or recycled product. The manufacturing plant may comprise a plurality of manufacturing units such as vessels for chemical or biological substances, or machines, e.g., robots, for processing solid or other materials. Themanufacturing units are configured such that an intermediate version or component of the product is moveable between the manufacturing units during manufacture of the product, e.g., via pipes or mechanical conveyance. As used herein manufacture of a product also includes manufacture of a food product by a kitchen robot.[000137] The agent may comprise an electronic agent configured to control a manufacturing unit, or a machine such as a robot, that operates to manufacture the product. That is, the agent may comprise a control system configured to control the manufacture of the chemical, biological, or mechanical product. For example the control system may be configured to control one or more of the manufacturing units or machines or to control movement of an intermediate version or component of the product between the manufacturing units or machines. As one example, the neural network model may be control the manufacturing units or machine to manufacture the product or an intermediate version or component thereof. As another example, the neural network model may be control the agent to control, e.g., minimize, use of a resource such as a task to control electrical power consumption, or water consumption, or the consumption of any material or consumable used in the manufacturing process.[000138] In some implementations the environment is the real-world environment of a power generation facility, e.g., a renewable power generation facility such as a solar farm or wind farm. The neural network model may be used to control power generated by the facility, e.g., to control the delivery of electrical power to a power distribution grid, e.g., to meet demand or to reduce the risk of a mismatch between elements of the grid, or to maximize power generated by the facility. The agent may comprise an electronic agent configured to control the generation of electrical power by the facility or the coupling of generated electrical power into the grid. The actions may comprise actions to control an electrical or mechanical configuration of an electrical power generator such as the electrical or mechanical configuration of one or more renewable power generating elements, e.g., to control a configuration of a wind turbine or of a solar panel or panels or mirror, or the electrical or mechanical configuration of a rotating electrical power generation machine. Mechanical control actions may, for example, comprise actions that control the conversion of an energy input to an electrical energy output, e.g., an efficiency of the conversion or a degree of coupling of the energy input to the electrical energy output. Electrical control actions may, for example, comprise actions that control one or more of a voltage, current, frequency or phase of electrical power generated.[000139] In general observations of a state of the environment may comprise any electronic signals representing the electrical or mechanical functioning of power generation equipment in the power generation facility. For example a representation of the state of the environment may be derived from observations made by any sensors sensing a physical or electrical state of equipment in the power generation facility that is generating electrical power, or the physical environment of such equipment, or a condition of ancillary equipment supporting power generation equipment. Such observations may thus include observations of wind levels or solar irradiance, or of local time, date, or season. Such sensors may include sensors configured to sense electrical conditions of the equipment such as current, voltage, power or energy; temperature or cooling of the physical environment; fluid flow; or a physical configuration of the equipment; and observations of an electrical condition of the grid, e.g., from local or remote sensors. Observations of a state of the environment may also comprise one or more predictions regarding future conditions of operation of the power generation equipment such as predictions of future wind levels or solar irradiance or predictions of a future electrical condition of the grid.[000140] As another example, the environment may be a chemical synthesis or protein folding environment such that each state is a respective state of a protein chain or of one or more intermediates or precursor chemicals and the agent is a computer system for determining how to fold the protein chain or synthesize the chemical. In this example, the actions are possible folding actions for folding the protein chain or actions for assembling precursor chemicals / intermediates and the result to be achieved may include, e.g., folding the protein so that the protein is stable and so that it achieves a particular biological function or providing a valid synthetic route for the chemical. As another example, the agent may be a mechanical agent that performs or controls the protein folding actions or chemical synthesis steps selected by the system automatically without human interaction. The observations may comprise direct or indirect observations of a state of the protein or chemical / intermediates / precursors and / or may be derived from simulation.[000141] In a similar way the environment may be a drug design environment such that each state is a respective state of a potential pharmaceutically active compound pharmaceutically active compound and the agent is a computer system for determining elements of the pharmaceutically active compound and / or a synthetic pathway for the pharmaceutically active compound.[000142] In some applications the agent may be a software agent, i.e., a computer program, configured to perform a task. For example the environment may be a circuit or anintegrated circuit design or routing environment and the agent may be configured to perform a design or routing task for routing interconnection lines of a circuit or of an integrated circuit, e.g., an ASIC. The observations may be, e.g., observations of component positions and interconnections; the actions may comprise component placing actions, e.g., to define a component position or orientation and / or interconnect routing actions, e.g., interconnect selection and / or placement actions. The method may include making the circuit or integrated circuit to the design, or with interconnection lines routed as determined by the method.[000143] In some applications the agent is a software agent and the environment is a real-world computing environment. In one example the agent manages distribution of tasks across computing resources, e.g., on a mobile device and / or in a data center. In these applications, the observations may include observations of computing resources such as compute and / or memory capacity, or Internet-accessible resources; and the actions may include assigning tasks to particular computing resources.[000144] In another example the software agent manages the processing, e.g., by one or more real -world servers, of a queue of continuously arriving jobs. The observations may comprise observations of the times of departures of successive jobs, or the time intervals between the departures of successive jobs, or the time a server takes to process each job, e.g., the start and end of a range of times, or the arrival times, or time intervals between the arrivals, of successive jobs, or data characterizing the type of job(s). The actions may comprise actions that allocate particular jobs to particular computing resources.[000145] As another example the environment may comprise a real-world computer system or network, the observations may comprise any observations characterizing operation of the computer system or network, the actions performed by the software agent may comprise actions to control the operation, e.g., to limit or correct abnormal or undesired operation, e.g., because of the presence of a virus or other security breach.[000146] In some applications, the environment is a real-world computing environment and the software agent manages distribution of tasks / jobs across computing resources, e.g., on a mobile device and / or in a data center. In these implementations, the observations may comprise observations that relate to the operation of the computing resources in processing the tasks / jobs, the actions may include assigning tasks / jobs to particular computing resources.[000147] In some applications the environment is a data packet communications network environment, and the agent is part of a router to route packets of data over thecommunications network. The actions may comprise data packet routing actions and the observations may comprise, e.g., observations of a routing table which includes routing metrics such as a metric of routing path length, bandwidth, load, hop count, path cost, delay, maximum transmission unit (MTU), and reliability.[000148] In some other applications the environment is an Internet or mobile communications environment and the agent is a software agent which manages a personalized recommendation for a user. The observations may comprise previous actions taken by the user, e.g., features characterizing these; the actions may include actions recommending items such as content items to a user.[000149] In some cases, the observations may include textual or spoken instructions provided to the agent by a third-party (e.g., an operator of the agent). For example, the agent may be an autonomous vehicle, and a user of the autonomous vehicle may provide textual or spoken instructions to the agent (e.g., to navigate to a particular location).[000150] As another example the environment may be an electrical, mechanical or electro-mechanical design environment, e.g., an environment in which the design of an electrical, mechanical or electro-mechanical entity is simulated. The simulated environment may be a simulation of a real-world environment in which the entity is intended to work. The task may be to design the entity. The observations may comprise observations that characterize the entity, i.e., observations of a mechanical shape or of an electrical, mechanical, or electro-mechanical configuration of the entity, or observations of parameters or properties of the entity. The actions may comprise actions that modify the entity, e.g., that modify one or more of the observations. The design process may include outputting the design for manufacture, e.g., in the form of computer executable instructions for manufacturing the entity. The process may include making the entity according to the design. Thus the design of an entity may be optimized, e.g., by reinforcement learning, and then the optimized design output for manufacturing the entity, e.g., as computer executable instructions; an entity with the optimized design may then be manufactured.[000151] Another application of the present neural network model is to obtain data characterizing the sequence of data items (i.e., time series analysis). For example, the data items may be sensor data, such as sensor data which is medical data (e.g., electrocardiogram (ECG) measurements at a plurality of respective times), audio data (e.g., captured by a microphone) or image data (e.g., image captured by a camera or medical imaging equipment at respective times).[000152] The characterization data may be such as to indicate that the sequence of data items is in one of a plurality of categories. For example, the neural network model may indicate that a sequence of electro-cardiogram measurements is in an abnormal category associated with an elevated risk of a medical condition. In another example, the neural network may process a set of data items that represent the pixels of a still or moving image (e.g., in a raster pattern; in the case of a moving image, successive image frames may define respective successive portions of the data item sequence), or that represent audio-visual data, to generate a classification output, e.g., a multi-label classification output, that includes a respective score for each category in a set of categories. The categories in the set of categories can be, e.g., object categories, e.g., corresponding to vehicle, pedestrian, bicyclist, etc. For audio-visual data the categories can comprise event categories, where an event is characterized by a combination of sound and vision, e.g., a baby crying, tool use, a cymbal, a dog barking, fireworks, a crowd cheering, wind blowing, and so forth. The score for an object or event category can define a likelihood that the image depicts an object that belongs to the object category or that an event belongs to the event category.[000153] Alternatively or additionally, the characterization data may identify a (proper) subset of the data items as having a characteristic, e.g., a subset of a sequence of medical images which exhibit a certain characteristic (e.g., in ultrasound images of a heart, images showing times at which the heart malfunctions), and / or identifies a portion of the data items as having a characteristic (e.g., in ultrasound images of an unborn child, a respective portion of the images which shows the heart of the child).[000154] Another application of the present neural network model is as a language model, in which the sequence of data items is an input sequence of tokens selected from a vocabulary (e.g., a natural language vocabulary), to generate an item embedding (e.g., the item embedding generated by the last layer of the analysis network) using which an output token from a vocabulary (e.g., the same vocabulary or a different vocabulary) is selected, e.g., by the output network. Repeating this process a plurality of times produces an output sequence of output tokens. The process may be recursive, i.e., the selected output tokens may be used to generate a corresponding predicted item embedding, which is processed using the analysis network, to obtain an item embedding (e.g., output by the last processing layer of the analysis network) using which a further output token is selected. In this way, a sequence of output tokens of arbitrary length can be generated. The neural network model may be a large language model (e.g., a presently known large language model, but using a respective neural network model as described in place of one or more transformer models ofthe known large language model). The large language model may include over a billion trained numerical parameters. Furthermore, the neural network model may be a foundation model, that is trained using a large database of training data (e.g., language training data), such that it can be adapted (i.e., further trained, or used in combination with additional trained layers) for to perform any of a plurality of other “downstream” computational tasks. [000155] One form of large language model, for which the neural network model may be employed, is a multimodal model which processes an input comprising both a media element (comprising an image (e.g. captured from the real world by a camera) and / or audio data (e.g. captured from the real world by a microphone)) and an input sequence of (e.g. text) tokens selected from a vocabulary. The neural network model processes the media element together with the input sequence of tokens, to generate an output sequence which is an (appropriate) response to the media element and the input sequence of tokens. For example, the input sequence of tokens may frame a question to which the output sequence provides an answer which is dependent on the content of the media element.[000156] One form of multimodal model is a vocabulary visual language model (VLM), which processes an input comprising both at least one image (still image(s) or moving image(s)) and an input sequence of (e.g. text) tokens selected from a vocabulary. The neural network model processes pixel-level data in the image(s), together with the input sequence of tokens, to generate an output sequence which is an (appropriate) response to image content depicted in the image(s) (e.g. objects and / or people depicted in the image(s)) and the input sequence of tokens. For example, the input sequence of tokens may frame a question to which the output sequence provides an answer which is dependent on the content depicted in the image(s).[000157] In some implementations the input tokens and the output tokens each represent words, wordpieces or characters in a natural language. A wordpiece may be a subword (part of a word), and may be an individual letter or character. As used here, “characters” includes Chinese and other similar characters, as well as logograms, syllabograms and the like. The tokens may include marker tokens, such as a start of sequence token, an end of sequence token, and a separator token (indicating a separation or break between two distinct parts of a sequence).[000158] Some of these implementations may be used for natural language tasks such as providing a natural language response to a natural language input, e.g., for question answering, or for text completion. In some implementations the input sequence may represent text in a natural language and the output sequence may represent text in the samenatural language, e.g., a longer item of text. For example in some implementations the input sequence may represent text in a natural language and the output sequence may represent the same text with a missing portion of the text added or filled in. For example the output sequence may represent a predicted completion of text represented by the input sequence. Such an application may be used, e.g., to provide an auto-completion function, e.g., for natural language-based search. In some implementations the input sequence may represent a text in a natural language, e.g., posing a question or defining a topic, and the output sequence may represent a text in a natural language which is a response to the question or about the specified topic.[000159] As another example the input sequence may represent a first item of text and the output sequence may represent a second, shorter item of text, e.g., the second item of text may be a summary of a passage that is the first item of text. As another example the input sequence may represent a first item of text and the output sequence may represent a simplification of the first item of text. As another example the input sequence may represent a first item of text and the output sequence may represent an aspect of the first item of text, e.g., it may represent an entailment task, a paraphrase task, a textual similarity task, a sentiment analysis task, a sentence completion task, a grammaticality task, a parsing task, e.g., constituency parsing, and in general any natural language understanding task that operates on a sequence of text in some natural language, e.g., to generate an output that classifies or predicts some property of the text. For example some implementations may be used to identify a natural language of the first item of text, or of spoken words where the input is audio (as described below).[000160] Some implementations may be used to perform neural machine translation. Thus in some implementations the input tokens represent words, wordpieces, or characters in a first natural language and the output tokens represent words, wordpieces or characters in a second, different natural language. That is, the input sequence may represent input text in the first language and the output sequence may represent a translation of the input text into the second language.[000161] Some implementations may be used for automatic code generation. For example the input tokens may represent words, wordpieces or characters in a first natural language and the output tokens may represent instructions in a computer programming or markup language, or instructions for controlling an application program to perform a task, e.g., build a data item such as an image or web page.[000162] Some implementations may be used for speech recognition. In such applications the input sequence may represent spoken words and the output sequence may represent a conversion of the spoken words to a machine-written representation, e.g., text. Then the input tokens may comprise tokens representing an audio data input including the spoken words, e.g., characterizing a waveform of the audio in the time domain or in the time-frequency domain. The output tokens may represent words, wordpieces, characters, or graphemes of a machine-written, e.g., text, representation of the spoken input, that is representing a transcription of the spoken input.[000163] Some implementations may be used for handwriting recognition. In such applications the input sequence may represent handwritten words, syllabograms or characters and the output sequence may represent a conversion of the input sequence to a machine-written representation, e.g., text. Then the input tokens may comprise tokens representing portions of the handwriting and the output tokens may represent words, wordpieces, characters or graphemes of a machine-written, e.g., text, representation of the spoken input.[000164] Some implementations may be used for text-to-speech conversion. In such applications the input sequence may represent text and the output sequence may represent a conversion of the text to spoken words. Then the input tokens may comprise tokens representing words or wordpieces or graphemes of the text and the output tokens may represent portions of audio data for generating speech corresponding to the text, e.g., tokens characterizing a portion of a waveform of the speech in the time domain or in the timefrequency domain, or phonemes.[000165] In some implementations the input sequence and the output sequence represent different modalities of input. For example the input sequence may represent text in a natural language and the output sequence may represent an image or video corresponding to the text; or vice-versa. In general the tokens may represent image or video features and a sequence of such tokens may represent an image or video. There are many ways to represent an image (or video) using tokens. As one example an image (or video) may be represented as a sequence of regions of interest (Rols) in the image, optionally including one or more tokens for global image features. For example an image may be encoded using a neural network to extract Rol features; optionally (but not essentially) a token may also include data, e.g., a position encoding, representing a position of the Rol in the image. As another example, the tokens may encode color or intensity values for pixels of an image. As anotherexample, some image processing neural network systems, e.g., autoregressive systems, naturally represent images as sequences of image features.[000166] Thus in some implementations at least one of the input sequence and the output sequence is a sequence representing an image or video, and the tokens represent the image or video. For example the input sequence may be a sequence of text, the input tokens may represent words, wordpieces, or characters and the output sequence may comprise output tokens representing an image or video, e.g., described by the text, or providing a visual answer to a question posed by the text, or providing a visualization of a topic of the text. In another example the input sequence may comprise a sequence of input tokens representing an image or video, and the output tokens may represent words or wordpieces, or characters representing text, e.g., for a description or characterization of the image or video, or providing an answer to a question posed visually by the image or video, or providing information on a topic of a topic of the image or video.[000167] In some other implementations both the input sequence and the output sequence may represent an image or video, and both the input tokens and the output tokens may represent a respective image or video. In such implementations the method / system may be configured to perform an image or video transformation. For example the input sequence and the output sequence may represent the same image or video in different styles, e.g., one as an image the other as a sketch of the image; or different styles for the same item of clothing.[000168] In some implementations the input sequence represents data to be compressed, e.g., image data, text data, audio data, or any other type of data; and the output sequence a compressed version of the data. The input and output tokens may each comprise any representation of the data to be compressed / compressed data, e.g., symbols or embeddings generated / decoded by a respective neural network.[000169] In some implementations the input sequence represents a sequence of actions to be performed by an agent, e.g., a mechanical agent in a real -world environment implementing the actions to perform a mechanical task. The output sequence may comprise a modified sequence of actions, e.g., one in which an operating parameter, such as a speed of motion or power consumption, has a limited value; or one in which or safety or other boundary is less likely to be crossed. Then both the input tokens and the output tokens may represent the actions to be performed.[000170] In some implementations the input sequence represents a sequence of health data and the output sequence may comprise a sequence of predicted treatment. Then theinput tokens may represent any aspect of the health of a patient, e.g., data from blood and other medical tests on the patient and / or EHR (Electronic Health Record) data; and the output tokens may represent diagnostic information, e.g., relating to a disease status of the patient and / or relating to suggested treatments for the patient, and / or relating to a likelihood of an adverse health event for the patient.[000171] In some implementations the input sequence represents a time series and the output sequence may comprise a continuation of the time series. For example the input sequence may be a sequence representing the output of an electricity generating plant, e.g., a solar or wind electricity generating plant, or a sequence representing electricity consumption, and the output sequence may provide a forecast of the electricity generated or consumed. As another example the input sequence may be a sequence representing a level of traffic on one or more roads and the output sequence may provide a forecast of the future traffic.[000172] In some cases, the operation of the analysis network may be conditioned on a media item (e.g., a still or moving image and / or audio data), so that the neural network model operates a multi-modal language model.[000173] Another application of the present neural network model is as a generative model for generating a media item (e.g., comprising a still or moving image and / or audio data; and optionally additionally comprising tokens selected from a vocabulary) or generating an item of language (e.g., a passage of text made up of tokens) by selecting a series of tokens from a vocabulary. Successive outputs of the neural network model may be used to generate successive elements of the media items and / or select successive tokens from the vocabulary. For example, the elements may be one or more intensity values for each pixel of an image, or a pixel of a frame of a moving image. The successive elements may correspond to the pixels in a raster pattern; in the case of a moving image, successive image frames may define respective successive portions of the output sequence of the neural network model. In another example, the elements may define one or more values (e.g., Fourier components) characterizing a corresponding time portion of the audio data.[000174] As in some other applications above, outputs of the neural network model can be fed back recursively, e.g., outputs of the output network can be fed back to the embedding network as new data items of the sequence (i.e., data items which follow the initial data items in the sequence of data items), or image embedding outputs of the analysis network can be fed back to be image embeddings which are input to the first layer of the analysis network.[000175] One or more data items of the sequence (e.g., one or more initial data items of the sequence) may be used to condition the generation of the media item or item of language, while optionally later data items of the sequence may be based on elements of the media item which have already been generated and / or tokens which have been selected. Thus, the one or more data items (i.e., data items which are not generated based on outputs of the analysis network) determine which media item or item of language is generated. These one or more data items of the sequence (e.g., an initial one or more of the data items in the sequence) may each comprise one or more tokens selected from a vocabulary (e.g., a vocabulary of the types discussed above, such as natural language tokens), e.g., by a user. In this way, the media item is generated conditioned on these tokens. For example, a still or moving image, or audio data, can be generated conditioned on data items in the sequence which are based on selected tokens, e.g., as an image described by the data items, or as sound based on (e.g., a reading out of, or which is described by) the data items.[000176] This specification uses the term “configured” in connection with computing systems and environments, as well as computer program components. A computing system or environment is considered "configured" to perform specific operations or actions when it possesses the necessary software, firmware, hardware, or a combination thereof, enabling it to carry out those operations or actions during operation. For instance, configuring a system might involve installing a software library with specific algorithms, updating firmware with new instructions for handling data, or adding a hardware component for enhanced processing capabilities. Similarly, one or more computer programs are "configured" to perform particular operations or actions when they contain instructions that, upon execution by a computing device or hardware, cause the device to perform those intended operations or actions.[000177] The embodiments and functional operations described in this specification can be implemented in various forms, including digital electronic circuitry, in tangibly- embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or inaddition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. Additionally or alternatively, the program instructions can be encoded on a transmitted signal, such as a machine-generated electrical, optical, or electromagnetic signal, designed to carry information for transmission to a receiving device or system for execution by a computing device or hardware. Furthermore, implementations may leverage emerging technologies like quantum computing or neuromorphic computing for specific applications, and may be deployed in distributed or cloud-based environments where components reside on different machines or within a cloud infrastructure.[000178] The terms “data processing apparatus” or “computing device or hardware” refer to the physical components involved in data processing and encompass all types of devices and machines used for this purpose. Examples include processors or processing units, computers, multiple processors or computers working together, graphics processing units (GPUs), tensor processing units (TPUs), and specialized processing hardware such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). In addition to hardware, a computing device or hardware may also include code that creates an execution environment for computer programs. This code can take the form of processor firmware, a protocol stack, a database management system, an operating system, or a combination of these elements. Embodiments may particularly benefit from utilizing the parallel processing capabilities of GPUs, in a General-Purpose computing on Graphics Processing Units (GPGPU) context, where code specifically designed for GPU execution, often called kernels or shaders, is employed. Similarly, TPUs excel at running optimized tensor operations crucial for many machine learning algorithms. By leveraging these accelerators and their specialized programming models, the system can achieve significant speedups and efficiency gains for tasks involving artificial intelligence and machine learning, particularly in areas such as computer vision, natural language processing, and robotics.[000179] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or otherunit suitable for use in a computing environment. A program may or may not correspond to a single file in a file system and can be stored in various ways. This includes being embedded within a file containing other programs or data (e.g., scripts within a markup language document), residing in a dedicated file, or distributed across multiple coordinated files (e.g., files storing modules, subprograms, or code segments). A computer program can be executed on a single computer or across multiple computers, whether located at a single site or distributed across multiple sites and interconnected through a data communication network. The specific implementation of the computer programs may involve a combination of traditional programming languages and specialized languages or libraries designed for GPGPU programming or TPU utilization, depending on the chosen hardware platform and desired performance characteristics.[000180] In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.[000181] In this specification, the term "engine" broadly refers to a software-based system, subsystem, or process designed to perform one or more specific functions. An engine is typically implemented as one or more software modules or components installed on one or more computers, which can be located at a single site or distributed across multiple locations. In some instances, one or more dedicated computers may be used for a particular engine, while in other cases, multiple engines may operate concurrently on the same one or more computers. Examples of engine functions within the context of Al and machine learning could include data pre-processing and cleaning, feature engineering and extraction, model training and optimization, inference and prediction generation, and postprocessing of results. The specific design and implementation of engines will depend on the overall architecture and the distribution of computational tasks across various hardware components, including CPUs, GPUs, TPUs, and other specialized processors.[000182] The processes and logic flows described in this specification can be executed by one or more programmable computers running one or more computer programs to perform functions by operating on input data and generating output. Additionally, graphics processing units (GPUs) and tensor processing units (TPUs) can be utilized to enable concurrent execution of aspects of these processes and logic flows, significantly accelerating performance. This approach offers significant advantages for computationallyintensive tasks often found in Al and machine learning applications, such as matrix multiplications, convolutions, and other operations that exhibit a high degree of parallelism. By leveraging the parallel processing capabilities of GPUs and TPUs, significant speedups and efficiency gains compared to relying solely on CPUs can be achieved. Alternatively or in combination with programmable computers and specialized processors, these processes and logic flows can also be implemented using specialized processing hardware, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), for even greater performance or energy efficiency in specific use cases.[000183] The processes and logic flows described in this specification can be executed by one or more programmable computers running one or more computer programs to perform functions by operating on input data and generating output. Additionally, graphics processing units (GPUs) and tensor processing units (TPUs) can be utilized to enable concurrent execution of aspects of these processes and logic flows, significantly accelerating performance. This approach offers significant advantages for computationally intensive tasks often found in Al and machine learning applications, such as matrix multiplications, convolutions, and other operations that exhibit a high degree of parallelism. By leveraging the parallel processing capabilities of GPUs and TPUs, significant speedups and efficiency gains compared to relying solely on CPUs can be achieved. Alternatively or in combination with programmable computers and specialized processors, these processes and logic flows can also be implemented using specialized processing hardware, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs), for even greater performance or energy efficiency in specific use cases.[000184] Computer-readable media suitable for storing computer program instructions and data encompass all forms of non-volatile memory, media, and memory devices. Examples include semiconductor memory devices such as read-only memory (ROM), solid- state drives (SSDs), and flash memory devices; hard disk drives (HDDs); optical media; and optical discs such as CDs, DVDs, and Blu-ray discs. The specific type of computer-readable media used will depend on factors such as the size of the data, access speed requirements, cost considerations, and the desired level of portability or permanence.[000185] To facilitate user interaction, embodiments of the subject matter described in this specification can be implemented on a computing device equipped with a display device, such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display, for presenting information to the user. Input can be provided by the user through various means, including a keyboard), touchscreens, voice commands, gesture recognition, or other inputmodalities depending on the specific device and application. Additional input methods can include acoustic, speech, or tactile input, while feedback to the user can take the form of visual, auditory, or tactile feedback. Furthermore, computers can interact with users by exchanging documents with a user's device or application. This can involve sending web content or data in response to requests or sending and receiving text messages or other forms of messages through mobile devices or messaging platforms. The selection of input and output modalities will depend on the specific application and the desired form of user interaction.[000186] Machine learning models can be implemented and deployed using machine learning frameworks, such as TensorFlow or JAX. These frameworks offer comprehensive tools and libraries that facilitate the development, training, and deployment of machine learning models.[000187] Embodiments of the subject matter described in this specification can be implemented within a computing system comprising one or more components, depending on the specific application and requirements. These may include a back-end component, such as a back-end server or cloud-based infrastructure; an optional middleware component, such as a middleware server or application programming interface (API), to facilitate communication and data exchange; and a front-end component, such as a client device with a user interface, a web browser, or an app, through which a user can interact with the implemented subject matter. For instance, the described functionality could be implemented solely on a client device (e.g., for on-device machine learning) or deployed as a combination of front-end and back-end components for more complex applications. These components, when present, can be interconnected using any form or medium of digital data communication, such as a communication network like a local area network (LAN) or a wide area network (WAN) including the Internet. The specific system architecture and choice of components will depend on factors such as the scale of the application, the need for real-time processing, data security requirements, and the desired user experience.[000188] The computing system can include clients and servers that may be geographically separated and interact through a communication network. The specific type of network, such as a local area network (LAN), a wide area network (WAN), or the Internet, will depend on the reach and scale of the application. The client-server relationship is established through computer programs running on the respective computers and designed to communicate with each other using appropriate protocols. These protocols may include HTTP, TCP / IP, or other specialized protocols depending on the nature of the data beingexchanged and the security requirements of the system. In certain embodiments, a server transmits data or instructions to a user's device, such as a computer, smartphone, or tablet, acting as a client. The client device can then process the received information, display results to the user, and potentially send data or feedback back to the server for further processing or storage. This allows for dynamic interactions between the user and the system, enabling a wide range of applications and functionalities.[000189] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.[000190] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.[000191] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

Claims

CLAIMS1. A computer-implemented method for processing a sequence of data items corresponding to a plurality of time steps, the method employing a neural network model and comprising for successive time steps: processing an initial item embedding based on the data item for the time step using an analysis network of the neural network model comprising a plurality of processing layers arranged in a sequence, each processing layer performing a function defined by a corresponding set of trained numerical parameters, a first processing layer of the sequence being configured to receive the initial item embedding, and to output a corresponding modified item embedding, and each other processing layer of the sequence being configured to receive the item embedding output by the preceding layer of the sequence and output a corresponding modified item embedding; wherein at least one of the processing layers is a spectral transform layer which, for each current time step: generates a plurality of feature vectors by processing a sequence embedding using a respective plurality of spectral filters, the sequence embedding comprising as components the item embeddings received by the processing layer at one or more time steps preceding the current time step; multiplies the feature vectors by respective weight matrices defined by at least some of the corresponding set of trained numerical parameters, to form respective weighted feature vectors; and generates the modified item embedding for the current time step including a term based on the weighted feature vectors.

2. The method of claim 1, in which the plurality of feature vectors include a plurality of first feature vectors formed from respective ones of the spectral filters, each feature vector being a sum, over components of the sequence embedding, of a product of the corresponding received item embedding and a corresponding component of the respective spectral filter.

3. The method of claim 1 or claim 2, in which the plurality of feature vectors include a plurality of second feature vectors formed from respective ones of the spectral filters, as a sum, over components of the sequence embedding, of a product of (i) the correspondingitem embedding, (ii) a corresponding component of the respective spectral filter, and (iii) a value which alternates in sign for successive components of the sequence embedding.

4. The method of any of claims 1 to 3, in which the spectral filters are corresponding eigenvectors of a Hankel matrix.

5. The method of any of preceding claim, in which the term based on the weighted feature vectors is a sum of the weighted feature vectors.

6. The method of claim 5 when dependent on claim 4, in which the sum of the weighted feature vectors is weighted by a value which varies inversely with the corresponding eigenvalue of the Hankel matrix.

7. The method of any preceding claim, in which the modified item embedding for the current time step further includes at least one auto-regressive term based on a modified item embedding generated by the processing layer in a preceding time step.

8. The method of claim 7, in which the preceding time step for at least one of the autoregressive terms is two time steps before the current time step.

9. The method of any preceding claim, in which the modified item embedding for the current time step further includes a term which is a product of the item embedding for the time step with a corresponding second weight matrix defined by some of the corresponding set of trained numerical parameters.

10. The method of any preceding claim, in which the modified item embedding for the current time step further includes a term which is a product of the item embedding for a preceding time step with a corresponding third weight matrix defined by some of the corresponding set of trained numerical parameters.

11. The method of any preceding claim, in which the analysis network comprises a plurality of spectral transform layers, each spectral transform layer employing a corresponding plurality of weight matrices defined by at least some of the corresponding set of trained numerical parameters.

12. The method of any preceding claim in which at least one of the processing layers is a non-linear processing layer which applies a corresponding non-linear function to a received item embedding to generate a modified item embedding, the corresponding non-linear function being defined by at least some of the corresponding set of trained numerical parameters.

13. The method of claim 12 when dependent upon claim 11, the plurality of spectral transform layers being interleaved with a plurality of the non-linear processing layers.

14. The method of any preceding claim, further comprising generating the initial item embedding for the time step from the data item for the time step using an embedding network of the neural network model.

15. The method of any preceding claim, in which the data items comprise sensor data collected by a sensor.

16. The method of claim 15 in which the sensor data comprises medical data.

17. The method of claim 15 or claim 16, in which the sensor data comprises audio data.

18. The method of any of claim 15 to 17, in which the sensor data comprises image data.

19. The method of any preceding claim, in which the data items are observation data items characterizing an environment at a set of corresponding times, the method further comprising generating from the item embedding output by the last processing layer of the analysis network, a current predicted observation data item characterizing a state of the environment at a current prediction time later than the corresponding times.

20. The method of claim 19, further comprising, at least once: inputting to the analysis network a predicted item embedding generated from the item embedding output by the last processing layer of the analysis network,processing the predicted item embedding using the analysis network, to obtain a new item embedding output by the last processing layer of the analysis network; and generating from the new item embedding output by the last processing layer, a new predicted observation data item characterizing a state of the environment at a new current prediction time later than previous current prediction time.

21. The method of claim 19 or claim 20, in which the environment is a real -world environment.

22. The method of claim 20, in which the environment comprises one or more objects, the observation data item characterizing the spatial positions, relative spatial positions, orientations and / or relative orientations of the objects at the corresponding time.

23. The method of any of claims 19 to 22, further including generating and outputting to a user, characterization data characterizing the predicted observation data item.

24. The method of any of claims 19 to 23, further including generating control data for an electromechanical system based on the predicted observation data item.

25. The method of any of claims 1 to 18, in which the data items characterize a state of a controlled system, the method further comprising generating a control input for a controlled system based on the output of the item embedding output by the last processing layer.

26. The method of claim 23 in which the controlled system is a real-world apparatus.

27. The method of claim 25, in which the controlled system is an electromechanical system.

28. The method of any of preceding claim, further comprising, from the item embeddings generated by a last processing layer of the analysis network, generating characterization data characterizing the sequence of data items.

29. The method of claim 28, in which the characterization data is data which identifies a subset of the data items as data items having a characteristic and / or which identifies a portion of the one or more of the data items as having a characteristic.

30. The method of claim 23 or claim 28 in which the characterization data is data which identifies one of a plurality of predefined categories.

31. The method of any of claims 1 to 14, in which the data items are input tokens from a first vocabulary, the method further comprising selecting an output token from a second vocabulary based on an item embedding output by the last layer of the analysis network.

32. The method of claim 31, further comprising, at least once: generating a predicted item embedding corresponding to the selected token, processing the predicted item embedding using the analysis network, to obtain an item embedding output by the last processing layer of the analysis network; and selecting a new output token from the second vocabulary based on an item embedding output by the last layer of the analysis network.

33. The method of any of claims 1-14, further comprising, a plurality of times: generating successive elements of a media item, or selecting successive tokens from a vocabulary, from successive item embeddings output by a last layer of the analysis network; and transmitting item embeddings based on item embeddings output by the last layer of the analysis network as inputs to the analysis network.

34. The method of claim 33, in which the media item comprises image data or audio data.

35. The method of claim 34, in which the media item further comprises one or more tokens selected from a vocabulary based on item embeddings output by the last layer of the analysis network.

36. The method of any of claims 33 to 35, in which one or more of the data items of the sequence comprise tokens selected from a vocabulary.

37. A system for processing a sequence of data items corresponding to a plurality of time steps, the system comprising a neural network model which comprises: an analysis network comprising a plurality of processing layers arranged in a sequence, each of the processing layers performing a function defined by a respective set of numerical parameters, a first processing layer of the sequence being configured to receive the initial item embeddings, and output corresponding modified item embeddings, and each other processing layer of the sequence being configured to receive the item embeddings output by the preceding layer of the sequence and output corresponding modified item embeddings; at least one of the processing layers being a spectral transform layer configured, for each current time step of the sequence, to: generate a plurality of feature vectors by processing a sequence embedding using a respective plurality of spectral filters, the sequence embedding comprising as components the item embeddings received by the spectral transform layer for one or more time steps preceding the current time step; to multiply the feature vectors by respective weight matrices to form respective weighted feature vectors, the weight matrices being defined by at least some of the corresponding set of numerical parameters; and to generate the modified item embedding for the current time step including a term based on a sum of the weighted feature vectors.

38. A method of training the system of claim 37, using a training database of sample data item sequences and sample target vectors, the method comprising iteratively modifying the numerical parameters defining the weight matrices to reduce a loss function characterizing a discrepancy between (i) values based on item embeddings output by the last layer of the analysis network upon the embedding network receiving the sample data items sequences, and (ii) the sample target vectors.

39. The method of claim 38 in which the neural network model further comprises (i) an embedding network configured to receive the data items, and to generate from the data items corresponding initial item embeddings for each of the time steps, and / or (ii) an outputnetwork configured to receive an item embedding from the last processing layer of the analysis network, and generate a network output, the embedding network and / or output network being defined by respective sets of additional numerical parameters, and the method comprising iteratively modifying the additional numerical parameters to reduce the loss function.

40. One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the operations of the respective method of any one of claims 1-36, 38 or 39.

41. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-36, 38 or 39.

Citation Information

Cited By

  • RGB-T image multi-modal semantic segmentation method based on state space

    CN121305091A