Computer-based method, computer system and computer program (overcoming maximum token number limitations of large language models)
By dividing the attention matrix into submatrices, encoding them with a GRU, constructing a DAG, and using a GNN for dynamic graph construction, the method addresses the maximum token number limit in large language models, achieving efficient processing of long texts and preventing information loss.
Patent Information
- Application Number
- JP2024199168
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-11-14
- Publication Date
- 2025-06-12
AI Technical Summary
Large language models face limitations due to the maximum token number limit, which results in information loss when processing text exceeding this limit, and existing methods struggle to perform fine-grained reading comprehension and are not immediately applicable to existing models.
The method involves receiving target text, dividing the attention matrix into submatrices, encoding these submatrices using a gated recurrent unit (GRU) neural network, constructing a directed acyclic graph (DAG) with nodes representing encoded vectors, and using a graph neural network (GNN) for dynamic graph construction and node feature transfer to generate summaries from the updated graph.
This approach effectively overcomes the maximum token number limit by reducing computational complexity, capturing local semantic relationships, and dynamically calculating only necessary nodes, thereby enabling the processing of long texts without information loss.
Smart Images

Figure 2025089268000001_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to computer processing, and more specifically, to overcoming the maximum token limit in large language models.
Summary of the Invention
Problems to be Solved by the Invention
[0002] Large language models are becoming increasingly popular due to their ability to understand and generate human-like text. Many companies are actively investing in discovering and leveraging opportunities to utilize large language models for various end uses designed to enhance efficiency and competitiveness. For example, companies can use large language models to automate tasks, gain insights, improve the customer experience, generate content, and many other things. Therefore, large language models with improved flexibility and practicality are highly desirable.
Means for Solving the Problems
[0003] According to one embodiment, a method, a computer system, and a computer program product for overcoming the maximum token number limit in a large language model are provided. The embodiment may include receiving target text. The embodiment may also include dividing an attention matrix associated with the target text into a series of submatrices. The embodiment may further include encoding fixed-length vectors corresponding to the series of submatrices by leveraging a gated recurrent unit (GRU) neural network. The embodiment may also include constructing a directed acyclic graph, in which the encoded fixed-length vectors have nodes, where connections between the nodes are defined based on a target task. The embodiment may further include repeatedly generating an updated graph including a series of most relevant node features and connection relationships by performing dynamic graph construction and node feature transfer by leveraging a graph neural network (GNN). The embodiment may also include generating one or more summaries regarding the received target text by extracting information from the updated graph.
Brief Description of the Drawings
[0004] These and other objects, features, and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments to be read in conjunction with the accompanying drawings. The illustrations are for the purpose of clarity in facilitating the understanding of the invention by those skilled in the art, so various features of the drawings are not to scale. The drawings include the following:
[0005]
Figure 1
[0006]
Figure 2
[0007]
Figure 3
[0008]
Figure 4
[0009]
Figure 5
Best Mode for Carrying Out the Invention
[0010] Detailed embodiments of the structures and methods in the claims are disclosed herein; however, it can be understood that the disclosed embodiments are merely exemplary of the structures and methods in the claims and may be embodied in various forms. The present disclosure, however, may be embodied in many different forms and should not be construed as limited to the exemplary embodiments described herein. In the description, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments.
[0011] Unless the context clearly indicates otherwise, it should be understood that the singular forms "a", "an", and "the" include plural referents. Thus, for example, a reference to "a component surface" includes a reference to one or more of such surfaces, unless the context clearly indicates otherwise.
[0012] Embodiments of the present application generally relate to computer processing, and more specifically, to overcoming the maximum token limit in large language models. The exemplary embodiments described below, among other things, receive target text, divide an attention matrix associated with the target text into a series of submatrices, utilize a gated recurrent unit neural network to encode fixed-length vectors corresponding to the series of submatrices, construct a directed acyclic graph, in which the encoded fixed-length vectors have nodes, where the connections between the nodes are defined based on a target task, utilize a graph neural network to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of most relevant node features and connection relationships, and extract information from the updated graph to generate one or more summaries regarding the received target text, providing a system, method, and program product.
[0013] As previously explained, large language models (LLMs) are becoming increasingly popular due to their ability to understand and generate human-like text. Many companies are actively investing in discovering and leveraging opportunities to utilize LLMs for various end uses designed to enhance efficiency and competitiveness. For example, companies can utilize LLMs to automate tasks, gain insights, improve the customer experience, generate content, and many other things. Thus, there is a great desire for more flexible and practical LLMs.
[0014] However, there are multiple challenges and limitations associated with using large language models. For example, many large language models have undesirable limitations regarding the maximum number of tokens that can be processed by a given large language model. For instance, an exemplary large language model may only be able to process token sequences that are less than or equal to a length of 32,000 tokens. If a given large language model receives text that includes a token sequence that exceeds the maximum number of tokens associated with that given large language model, all text after the token number limit will thereby be discarded, which can result in information loss. Recently proposed methods for dealing with the maximum token number limit typically involve shortening the received "long text" (text that exceeds the given maximum token number limit) by combining retrieval or summarization techniques. However, since these methods do not directly handle the received long text, they are often unable to perform fine-grained reading comprehension. Additionally, the proposed methods often require consideration during the training phase and cannot be immediately applied to existing LLM models. As a result, an improved method for overcoming the maximum token number limit for large language models that avoids these described drawbacks would be advantageous for companies seeking to use LLM with improved model flexibility and practicality.
[0015] Accordingly, a method, a computer system, and a computer program product for overcoming the maximum token number limit in a large language model are provided. The method, system, and computer program product may receive target text. The method, system, and computer program product may identify a defect in a printing operation based on tracked print data. Next, the method, system, and computer program product may divide an attention matrix associated with the target text into a series of submatrices. The method, system, and computer program product may utilize a gated recurrent unit neural network to encode fixed-length vectors corresponding to the series of submatrices. Next, the method, system, and computer program product may construct a directed acyclic graph, in which the encoded fixed-length vectors comprise nodes, where the connections between the nodes are defined based on a target task. Next, the method, system, and computer program product may utilize a graph neural network to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of most relevant node features and connection relationships. Thereafter, the method, system, and computer program product may generate one or more summaries regarding the received target text by extracting information from the updated graph. As a result, the method, system, and computer program product provide an improved method for overcoming the maximum token number limit for large language models. The described embodiments functionally combine the use of a gated recurrent unit neural network as a long-term memory storage with the naive Bayes method to overcome the maximum token number limit imposed by a given large language model.The embodiments described herein utilize a gated recurrent unit neural network to group encode a partial sequence of tokens in the received target text, calculate an attention matrix, and later transform the grouped calculation units into a directed acyclic graph (DAG) using the Naive Bayes algorithm. In the described embodiments, each node in the constructed DAG represents a computational unit, and whether each computational unit needs to be computed is dynamic, which is in contrast to previously proposed methods where all units had to be computed. The DAG decomposes the attention matrix calculation process into multiple parts, and each part corresponds to a node on the DAG. During the construction of the DAG, the computational units are not executed, which means that the calculation of the DAG is essentially "lazy" and each computational unit is computed only when it is needed. In the embodiments described herein, whether a computational node is computed is determined by the probabilistic relationship between previously computed nodes and a given current node, which is also calculated using the Naive Bayes method. As a result, the described embodiments overcome the limitations and problems associated with previously described methods and enable the generation of summaries (and the execution of other tasks) for received "long texts" that include token sequences exceeding a given maximum token count limit for a given large language model.
[0016] The present invention may be an integrated system, method, and / or computer program product integrated at any possible technical detail level. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0017] Various aspects of the present disclosure are described by a description, a flowchart, a block diagram of a computer system, and / or a block diagram of machine logic included in a computer program product (CPP) embodiment. With respect to any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or at least partially in a time-overlapping manner.
[0018] An embodiment of a computer program product (referred to herein as a "CPP embodiment" or "CPP") is a term used in this disclosure to describe any set of one or more storage media (also referred to as "media") collectively included in a set of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. A computer-readable storage medium can be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some well-known types of storage devices that include these media are floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on the major surfaces of discs), or any suitable combination of the foregoing. A computer-readable storage medium is not to be construed as storage in the form of a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, electrical signals transmitted through a wire, and / or other transmission media, as the term is used in this disclosure. As will be understood by those skilled in the art, data is typically moved during normal operation of a storage device at some irregular points in time, such as during access, defragmentation, or garbage collection, but the data is not transient while it is stored, and thus the storage device is not considered to be transient for the purposes of the foregoing.
[0019] Referring now to FIG. 1, computing environment 100 includes an example of an environment for executing at least some of the computer code involved in performing the method of the present invention, such as data processing program / code 150. In addition to data processing code 150, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In this embodiment, computer 101 includes a processor set 110 (including processing circuit 120 and cache 121), communication fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and the data processing code 150 identified above), a set of peripheral devices 114 (including user interface (UI) device set 123, storage 124, and Internet of Things (IoT) sensor set 125), and network module 115. Remote server 104 includes remote database 130. Public cloud 105 includes gateway 140, cloud orchestration module 141, host physical machine set 142, virtual machine set 143, and container set 144.
[0020] Computer 101 can take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch, or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device currently known or developed in the future that is capable of executing programs, accessing a network, or querying a database such as remote database 130. As is well understood in the field of computer technology and depending on the technology, the execution of computer implementation methods can be distributed among multiple computers and / or among multiple locations. On the other hand, in this description of computing environment 100, for the sake of simplicity as much as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not shown within the cloud in FIG. 1, it can be located within the cloud. On the other hand, computer 101 does not need to exist within the cloud, except within any range that can be assertively shown.
[0021] Processor set 110 includes one or more computer processors of any type currently known or developed in the future. Processing circuit 120 can be distributed among multiple packages, for example, multiple tuned integrated circuit chips. Processing circuit 120 can implement multiple processor threads and / or multiple processor cores. Cache 121 is a memory located within the processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuit. Alternatively, some or all of the cache for the processor set can be located "off-chip". In some computing environments, processor set 110 can be designed to operate using qubits and execute quantum computing.
[0022] Computer-readable program instructions cause a set of operation steps to be executed by the processor set 110 of the computer 101, thereby realizing a computer-implemented method, which is normally loaded onto the computer 101. As a result, the instructions thus executed instantiate the method specified in the flowchart and / or description of the computer-implemented method (collectively referred to as "the method of the present invention") included in this document. These computer-readable program instructions are stored in various types of computer-readable storage media, such as the cache 121 and other storage media discussed below. The program instructions and the associated data are accessed by the processor set 110 to control and direct the execution of the method of the present invention. In the computing environment 100, at least some of the instructions for executing the method of the present invention may be stored in the data processing code 150 within the persistent storage 113.
[0023] The communication fabric 111 is a signal conduction path that enables various components of the computer 101 to communicate with each other. Typically, this fabric is created by switches and conductive paths, such as buses, bridges, physical input / output ports, and switches and conductive paths that make up similar components. Other types of signal communication paths, such as optical fiber communication paths and / or wireless communication paths, may be used.
[0024] The volatile memory 112 is any type of volatile memory known currently or developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory is characterized by random access, but this is not essential unless expressly stated. In the computer 101, the volatile memory 112 is located within a single package and exists inside the computer 101. Alternatively or additionally, the volatile memory may be distributed across multiple packages and / or located externally to the computer 101.
[0025] The persistent storage 113 is any form of non-volatile storage for a computer, known currently or developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is directly supplied to the computer 101 and / or to the persistent storage 113. The persistent storage 113 can be read-only memory (ROM), but usually at least a portion of the persistent storage enables writing of data, deletion of data, and re-writing of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 can take a plurality of forms, such as various known proprietary operating systems, or an open-source portable operating system interface type operating system that employs a kernel. The code included within the data processing program 150 typically includes at least some of the computer code involved in executing the method of the present invention.
[0026] The peripheral device set 114 includes a set of peripheral devices of the computer 101. The data communication connections between the peripheral devices of the computer 101 and other components may be implemented in various ways, such as a Bluetooth (registered trademark) connection, a Near Field Communication (NFC) connection, a connection formed by a cable (such as a Universal Serial Bus (USB) type cable), an insertion type connection (for example, a Secure Digital (SD) card), a connection formed through a local area communication network, and even a connection formed through a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, a speaker, a microphone, wearable devices (such as Google glasses and smartwatches), a keyboard, a mouse, a printer, a touchpad, a game controller, and a haptic device. The storage 124 is external storage such as an external hard drive or removable storage such as an SD card. The storage 124 may be persistent and / or volatile. In some embodiments, the storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 is required to have a large amount of storage (for example, when the computer 101 locally stores and manages a large-scale database), this storage may be provided by a peripheral storage device designed to store a very large amount of data, such as a Storage Area Network (SAN) shared by a plurality of geographically distributed computers. The IoT sensor set 125 is composed of a plurality of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.
[0027] The network module 115 is an aggregation of computer software, hardware, and firmware that enables the computer 101 to communicate with other computers via the WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi (registered trademark) signal transceiver, software for packetizing and / or depacketizing data for communication over a communication network, and / or web browser software for communicating data over the Internet. In some embodiments, the network control function and the network transfer function of the network module 115 are executed on the same physical hardware device. In other embodiments (e.g., embodiments that utilize Software-Defined Networking (SDN)), the control function and the transfer function of the network module 115 are executed on physically separate devices such that the control function manages multiple different network hardware devices. The computer-readable program instructions for executing the method of the present invention can typically be downloaded to the computer 101 from an external computer or an external storage device through a network adapter card or a network interface included in the network module 115.
[0028] The WAN 102 is any wide area network (e.g., the Internet) capable of communicating computer data over a non-local distance by any technology for communicating computer data that is currently known or developed in the future. In some embodiments, the WAN can be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi (registered trademark) network. The WAN and / or the LAN typically includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.
[0029] The end - user device (EUD) 103 is any computer system used and controlled by an end - user (e.g., a customer of the enterprise operating computer 101) and can take any of the forms discussed above in relation to computer 101. The EUD 103 typically receives beneficial and useful data from the operation of computer 101. For example, in a virtual case where computer 101 is designed to provide recommendations to the end - user, this recommendation will typically be communicated from the network module 115 of computer 101, via the WAN 102, to the EUD 103. In this way, the EUD 103 can display or otherwise present the recommendation to the end - user. In some embodiments, the EUD 103 can be a client device such as a thin - client, a thick - client, a mainframe computer, and a desktop computer, etc.
[0030] The remote server 104 is any computer system that provides at least some data and / or functions to computer 101. The remote server 104 can be controlled and used by the same entity that operates computer 101. The remote server 104 represents a machine that collects and stores data that is beneficial and useful for use by other computers such as computer 101. For example, in a virtual case where computer 101 is designed and programmed to provide recommendations based on historical data, this historical data may be provided from the remote database 130 of the remote server 104 to computer 101.
[0031] The public cloud 105 is any computer system that provides on-demand availability of computer system resources and / or other computer functions, particularly data storage (cloud storage) and computing power, for use by multiple entities without direct and active management by the user. Cloud computing typically exploits resource sharing to achieve coherence and economies of scale. The direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments that run on various computers that make up the host physical machine set 142, which is the universe of physical computers within and / or available to the public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from a virtual machine set 143 and / or containers from a container set 144. It is understood that these VCEs can be stored as images and transferred as images or after instantiation of the VCE among and within various physical machine hosts. The cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of the VCE, and manages the active instantiation of the VCE deployment. The gateway 140 is an aggregate of computer software, hardware, and firmware that enables the public cloud 105 to communicate through the WAN 102.
[0032] Here, some further explanation of a virtualized computing environment (VCE) is provided. A VCE can be stored as an "image". A new active instance of a VCE can be instantiated from the image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of the operating system where the kernel enables the existence of multiple isolated instances of user space, called containers. These isolated instances of user space typically behave as actual computers from the perspective of the programs running within them. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices allocated to the container, and this feature is known as containerization.
[0033] The private cloud 106 is similar to the public cloud 105, except that computing resources are available only for use by a single enterprise. The private cloud 106 is shown as being in communication with the WAN 102, but in other embodiments, the private cloud may be completely disconnected from the Internet and accessible only via a local / private network. A hybrid cloud is a composite of multiple different types of clouds (e.g., private cloud, community cloud, or public cloud types), and is often implemented by different vendors. Each of the multiple clouds remains a separate and distinct discrete entity, but the larger hybrid cloud architecture is coupled by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple component clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.
[0034] According to this embodiment, the data processing program 150 may be a program capable of receiving a target text. Next, the data processing program 150 may divide the attention matrix associated with the target text into a series of sub-matrices. Next, the data processing program 150 may utilize a gated recurrent unit neural network to encode a fixed-length vector corresponding to the series of sub-matrices. Next, the data processing program 150 may construct a directed acyclic graph, in which the encoded fixed-length vectors serve as nodes, and the connections between the nodes are defined based on the target task. Next, the data processing program 150 may utilize a graph neural network to perform dynamic graph construction and node feature transfer, and repeatedly generate an updated graph including a series of most relevant node features and connection relationships. Thereafter, the data processing program 150 may generate one or more summaries regarding the received target text by extracting information from the updated graph. As a result, the data processing program 150 provides an improved method for overcoming the maximum token number limit for large language models. The described embodiment functionally combines the use of a gated recurrent unit neural network as a long-term memory storage with the naive Bayes method to overcome the maximum token number limit imposed by a given large language model. The embodiment described herein utilizes a gated recurrent unit neural network to group-encode a partial sequence of tokens in the received target text, enabling the calculation of the attention matrix and the subsequent conversion of the grouped calculation units into a directed acyclic graph (DAG) using the naive Bayes algorithm. In the described embodiment, each node in the constructed DAG represents a calculation unit, and whether each calculation unit needs to be calculated is dynamic, which is in contrast to previously proposed methods where all units must be calculated. The DAG decomposes the attention matrix calculation process into multiple parts, each part corresponding to a node on the DAG.During the construction of the DAG, the computing units are not executed, which means that the calculation of the DAG is essentially "lazy" and each computing unit is calculated only when it is needed. In the embodiments described herein, whether a computing node is calculated is determined by the probabilistic relationship between previously calculated nodes and a given current node, which is also calculated using the Naive Bayes method. As a result, the embodiments described overcome the limitations and problems associated with the previously described methods, enabling the generation of summaries (and the execution of other tasks) for received "long texts" that include token sequences exceeding a given maximum token count limit for a given large language model.
[0035] Referring now to FIG. 2, there is provided an operational flowchart regarding an exemplary process 200 for overcoming the maximum token count limit in a large language model according to at least one embodiment.
[0036] At 202, the data processing program 150 may receive target text. In the context of the present disclosure, the target text may refer to any natural language text, coding or programming language, data or structured text, mathematical equations, scientific and technical text, or any other type of desired target text, which may include any number of desired characters or tokens. In some embodiments, the received target text may be included within any suitable desired format from which the text may be extracted therefrom using known text extraction techniques. The data processing program 150 is configured to process target text ("long text") that includes a number of tokens that may exceed the maximum token count limit for the target large language model. For example, in an embodiment, the data processing program 150 may receive an exemplary target text "T1" that is 40,000 tokens in length, which is intended to be input into an exemplary target large language model "LLM1" that has a maximum token count limit of 32,000.
[0037] In 204, the data processing program 150 may divide the attention matrix associated with the target text into a series of submatrices. FIG. 3 shows an exemplary process of dividing the attention matrix associated with the received target text into a series of submatrices according to at least one embodiment. As shown in the exemplary process 300 of FIG. 3, at this stage, the data processing program 150 may feed the received target text 310 into an exemplary pointer network 320 to perform semantic segmentation. The resulting acquired text segments at 330 are semantically coherent and are of a length shorter than the original received target text at 310. As shown at 340, each of the acquired text segments corresponds to its own context. The data processing program 150 may extract N*N submatrices (where N is the length of the context) from the attention matrix to obtain a series of submatrices. Thus, the series of submatrices are likewise associated with their own unique context. Next, this process may be repeated, and as shown at 350, the original attention matrix is decomposed into a plurality of different submatrices, where each submatrix represents a semantically independent context. Next, the data processing program 150 may further process the series of submatrices obtained from the divided attention matrix.
[0038] Next, at 206, the data processing program 150 may encode fixed-length vectors corresponding to a series of submatrices by utilizing a gated recurrent unit (GRU) neural network. For example, at this stage, the data processing program 150 may utilize the submatrices 410 and 420 shown in FIG. 4 as input features by feeding a series of submatrices into the GRU network. The GRU network may, for example, reshape the submatrices into 784-dimensional embeddings. In an embodiment, any suitable recurrent neural network (RNN) capable of performing the functions described above may be utilized to encode fixed-length vectors corresponding to a series of submatrices. As shown in FIG. 4, the data processing program 150 may, for example, input the first exemplary submatrix 410 and the second exemplary submatrix 420 into their respective RNNs 430, and fixed-length vectors 440 and 450 may be obtained respectively.
[0039] In 208, the data processing program 150 may construct a directed acyclic graph, in which the encoded fixed-length vectors correspond to nodes, and the connections between the nodes are defined based on the target task. FIG. 4 shows an exemplary process 400 for constructing a directed acyclic graph according to at least one embodiment. As shown in FIG. 4 and as described above, each of the partial matrices 410 and 420 is fed into the RNN 430 to obtain fixed-length vectors 440 and 450, respectively. In step 208, the data processing program 150 may be configured to construct a directed acyclic graph and define the nodes therein, where each of the encoded fixed-length vectors (e.g., vectors 440 and 450) functions as nodes 460 and 470, respectively, each representing a unit of computation. Next, the data processing program 150 may establish connections between the nodes based on the relationship between the requirements of the target task and the associated information of interest. For example, in an embodiment, the data processing program 150 may use a known similarity measure such as cosine similarity or correlation coefficient, which can be used to determine the strength of the connection between the nodes, to connect the nodes based on the relevance relationship between the partial matrices, thereby establishing connections between the nodes based on the relevance connection 475. Nodes corresponding to partial matrix encodings with high relevance have strong connections and may exhibit a tight association of information between them. In an embodiment, the data processing program 150 may, at 480, establish further connections between the nodes based on context connections. In the context of the present disclosure, the connections between the nodes may be determined based on the contextual relationships within the partial matrices. For example, if two partial matrices are adjacent or have a logical relationship in the original text, the data processing program 150 may establish a connection between them. In an embodiment, the data processing program 150 may, at 485, establish further connections between the nodes based on importance connections. In the context of the present disclosure, the importance connection or importance relationship between the nodes may be determined based on the importance or level of focus of the partial matrices.For example, if a certain submatrix contains important information or key perspectives, the nodes connected to the encoding nodes of the submatrix may have stronger connections. In an embodiment, when establishing a connection, the data processing program 150 may be configured to calculate an overall score at 490 by considering the three connection types discussed above and determine whether two nodes should be connected. The score may be calculated using any suitable known method. In an embodiment, the overall score may be compared to a predetermined and user-adjustable threshold to control the case where a connection is established between the pair of nodes being considered, as shown at 495.
[0040] Next, at 210, the data processing program 150 may iteratively generate an updated graph including a series of most relevant node features and connection relationships by leveraging a graph neural network (GNN) to perform dynamic graph construction and node feature transfer. At this stage, the data processing program 150 may iteratively facilitate information propagation by leveraging a GNN model and update based on node features and connectivity relationships. In an embodiment, the data processing program 150 may leverage a graph convolutional network (GCN), a GraphSAGE algorithm, a graph attention network (GAT), and any other suitable GNN model or algorithm. An exemplary example of stage 210 is shown in FIG. 5, which includes an exemplary process 500 of leveraging a graph neural network (GNN), specifically a GCN model, to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of most relevant node features and connection relationships according to at least one embodiment. In a first stage, the data processing program 150 may initialize node features such that the encoded vectors of each submatrix are utilized as initial features regarding each node by leveraging a GCN. Next, the data processing program 150 may construct a preliminary directed acyclic graph 510 based on the connectivity relationships between nodes (e.g., based on probabilities). Thus, the preliminary directed acyclic graph 510 describes the strength of the connections or relationships between nodes. Next, in loop 520, the data processing program may perform GNN layer iterations by leveraging a GNN. In each iteration of the GNN layer, exemplary stages may be used to update and propagate node features. For example, the layer iteration executed in loop 520 may include aggregating neighborhood features by using the preliminary directed acyclic graph 510 to aggregate the features of neighboring nodes of each node. This may be realized by a weighted averaging or splicing operation on the neighboring node features. The layer iteration executed in loop 520 may further include updating the node features.For example, the collected neighborhood features may be fused with the features of a given current node to generate new node features. This may be achieved by applying an update function, such as a graph convolution operation, a gated recurrent unit (GRU), a graph attention mechanism (GAT), or any other suitable mechanism or model. In an embodiment, the layer iteration executed in loop 520 may further include transferring the node features. For example, data processing program 150 may transfer the updated node features to the next iteration of the GNN layer associated with the next round of feature updates. Multiple rounds of iteration may be executed through a multi-layer GNN structure. In each round of GNN iteration, the GNN model updates the node features and transfers information based on the node features and connection relationships. Thus, by leveraging relevant probabilistic relationships, it is ensured that only the nodes that require calculation are calculated, while other nodes can be ignored, reducing the amount of calculation.
[0041] In an embodiment, data processing program 150 may be configured to include a stop condition such that the end of the GNN iteration (executed in loop 520) can be determined according to a specific stop condition. For example, in an embodiment, the stop condition may be that after reaching a specific number of iterations, after convergence of the node features, or any other desired and configurable custom condition. At 530, an updated graph 530 may be obtained, and the updated graph 530 includes a series of the most relevant node features and connection relationships.
[0042] After that, at 212, the data processing program 150 may generate one or more summaries regarding the received target text by extracting information from the updated graph. For example, at this stage, the data processing program 150 may generate a summary by extracting information from the updated graph, extracting key sentences based on node features, or classifying node features. In an embodiment, the data processing program 150 may generate other desired outputs such as recommendations or any other output that can be generated based on the information extracted from the updated graph at this stage, and ultimately limit or reduce the amount of tokens associated with the received target text that can be input into a given large language model.
[0043] Therefore, it can be understood that the data processing program 150 provides an improved method for overcoming the maximum token number limit in a large language model, which overcomes the problems observed in conventional and known methods for overcoming the maximum token number limit in a large language model.
[0044] For example, as discussed above, the described embodiments reduce the global computational complexity. In a conventional attention mechanism, as the number of tokens increases, the computational complexity increases significantly because each token needs to calculate an attention score using all other tokens. In the described embodiments, by splitting the attention matrix, the global attention calculation is transformed into calculations between local submatrices, and the computational load is significantly reduced.
[0045] Furthermore, the embodiments described herein provide the benefit of establishing local relationships. By splitting the original attention matrix into multiple submatrices and using a DAG to construct connections, local semantic relationships can be captured. This enables the localization of important information, avoids processing the entire text as a continuous sequence, and as a result, reduces the processing of irrelevant information while maintaining task relevance.
[0046] The embodiments described herein enable dynamic calculation of nodes. In the constructed DAG, unlike the conventional method that requires calculating the entire attention matrix, the calculation of each node is dynamic. Based on the probabilistic relationships between nodes, only the nodes that need to be calculated are evaluated, while others can be ignored. This further reduces the computational load and focuses only on the significant nodes with respect to a given current task and context.
[0047] It can be further understood that the described embodiments uniquely combine the Naive Bayes method with GRU (Gate Recurrent Unit) to address the problem of rapidly expanding token amounts. GRU groups-encodes a partial sequence of tokens, enables the calculation of the attention matrix, and then the grouped calculation units can be transformed into a directed acyclic graph (DAG) using the Naive Bayes algorithm. Each node on the DAG represents a calculation unit, and whether each calculation unit needs to be calculated is dynamic, which is in contrast to the previous method where all units had to be calculated. The DAG decomposes the calculation process of the attention matrix into multiple parts, and each part corresponds to a node on the DAG. During the construction of the DAG, the calculation units are not executed, which means that the calculation of the DAG is "lazy" and each calculation unit is calculated only when it is needed. As a result, whether a calculation node is calculated or not is determined by the probabilistic relationship between the previously calculated nodes and the current node, which is also calculated using the Naive Bayes method.
[0048] Therefore, the described method of splitting the attention matrix and constructing a DAG realizes efficient processing of long texts and information saving by reducing global computational complexity, establishing local relationships, and dynamically calculating nodes. This approach fully utilizes local correlations and task relatedness, enables the model to handle long texts more efficiently, and avoids information loss during the processing of "long" texts that exceed a given maximum token count limit for a target large language model.
[0049] The embodiments described herein may relate to the following items:
[0050] Item 1: A computer-based method for overcoming the maximum token limit in a large language model, the method comprising: receiving a target text; dividing an attention matrix associated with the target text into a series of sub-matrices; encoding a fixed-length vector corresponding to the series of sub-matrices by leveraging a gated recurrent unit (GRU) neural network; constructing a directed acyclic graph in which the encoded fixed-length vectors form nodes, where the connections between the nodes are defined based on a target task; repeatedly generating an updated graph including a series of the most relevant node features and connection relationships by performing dynamic graph construction and node feature transfer using a graph neural network (GNN); and generating one or more summaries regarding the received target text by extracting information from the updated graph. Thereby, the described embodiments are capable of overcoming the maximum token limit imposed by a given large language model by functionally combining the use of a gated recurrent unit neural network as a long-term memory storage with a naive Bayes method. As a result, the described embodiments are capable of leveraging a gated recurrent unit neural network to group-encode a partial sequence of tokens in the received target text, calculate an attention matrix, and later convert the grouped calculation units into a directed acyclic graph (DAG) using a naive Bayes algorithm, improving the generality of the large language model using the described embodiments.
[0051] Item 2: The computer-based method according to item 1, wherein the received target text has a number of tokens exceeding the maximum token limit associated with the target large language model. In such an embodiment, the received target text is not processable by the target large language model until the steps are executed according to the described embodiment. Thus, the received text containing a number of tokens exceeding the maximum token limit further enables the execution of the described method to overcome the maximum token limit while providing additional input that can be processed by the target large language model.
[0052] Item 3: The computer-based method according to any one of the preceding items 1 to 2, wherein each sub-matrix of the series of sub-matrices represents a semantically independent context. This ensures that any subsequently generated representation associated with a portion of the target text still corresponds to a relevant context that can be utilized during subsequent steps to ensure that the meaning and features of the target text are maintained.
[0053] Item 4: The computer-based method according to any one of the preceding items 1 to 3, wherein the connection between the nodes is determined using at least one of a relevance relationship between the sub-matrices based on a similarity measure, a context relationship based on a logical association between the sub-matrices, and an importance relationship based on the focus level of the sub-matrices. In an embodiment, the determined relevance relationship ensures that nodes corresponding to sub-matrix encodings with high relevance have strong connections and exhibit a tight association of information between them. This determination is then utilized to determine whether the nodes should be connected within a directed acyclic graph based on an associated scoring step.
[0054] Item 5: The target task has at least one of classification, summary generation, and recommendation, and is the computer-based method according to any one of Items 1 to 4 above. Thereby, generality is provided to the target large language model using the described embodiments, since the target tasks to be executed each include various useful tasks that are uniquely beneficial, but the described embodiments that overcome the maximum token limit associated with the target large language model tasked with processing the received long text can utilize the same data and features that become available.
[0055] Item 6: The step of using the graph neural network to perform the dynamic graph construction and the node feature transfer to repeatedly generate the updated graph including the series of most relevant node features and the connection relationships further includes: defining a preliminary directed acyclic graph, and for each of a series of nodes in the defined preliminary directed acyclic graph, aggregating the features of neighboring nodes by weighted averaging or splicing operations on the neighboring node features, and is the computer-based method according to any one of Items 1 to 5 above. In such an embodiment, in each round of GNN iteration, the GNN model updates the node features and transfers information based on the node features and connection relationships. As a result, by utilizing relevant probabilistic relationships, it is ensured that only the nodes that need to be calculated are calculated while other nodes can be ignored, reducing the amount of calculation, thereby improving the efficiency and performance of the target large language model using the described embodiments for processing the received text that exceeds a given maximum token limit.
[0056] Item 7: The method further comprises: applying an update function to fuse the collected neighborhood features with a series of current features of the target node to generate updated node features; and transferring the generated updated node features to the next iteration of the GNN layer. The computer-based method according to any one of the preceding items 1 to 6. This also serves to reduce the amount of computation and thereby improve the efficiency and performance of the target large language model using the described embodiments that process received text exceeding a given maximum token count limit.
[0057] Item 8: A computer system, the computer system comprising: one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored in at least one of the one or more computer-readable tangible storage media for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories, the computer system comprising: receiving target text; dividing an attention matrix associated with the target text into a series of submatrices; encoding a fixed-length vector corresponding to the series of submatrices by utilizing a gated recurrent unit (GRU) neural network; constructing a directed acyclic graph, in which the encoded fixed-length vector has nodes, where connections between the nodes are defined based on a target task; repeatedly generating an updated graph including a series of most relevant node features and connection relationships by utilizing a graph neural network (GNN) to perform dynamic graph construction and node feature transfer; and generating one or more summaries regarding the received target text by extracting information from the updated graph. Thus, the described embodiments can functionally combine utilizing the naive Bayes method with a gated recurrent unit neural network as long-term memory storage to overcome the maximum token number limit imposed by a given large language model. As a result, the described embodiments can utilize a gated recurrent unit neural network to group-encode a partial sequence of tokens in the received target text, calculate an attention matrix, and later convert the grouped calculation unit into a directed acyclic graph (DAG) using the naive Bayes algorithm, thereby improving the generality of the large language model using the described embodiments.
[0058] Item 9: The computer system according to item 8, wherein the received target text contains a number of tokens exceeding the maximum token count limit associated with the target large language model. In such an embodiment, the received target text cannot be processed by the target large language model until the steps are executed according to the described embodiment. Thus, the received text containing a number of tokens exceeding the maximum token count limit further enables the execution of the described method to overcome the maximum token count limit while providing additional input that can be processed by the target large language model.
[0059] Item 10: The computer system according to any one of the preceding items 8 to 9, wherein each sub-matrix of the series of sub-matrices represents a semantically independent context. This ensures that any subsequently generated representation associated with a portion of the target text still corresponds to a relevant context, which can be utilized during subsequent steps to ensure that the meaning and features of the target text are maintained.
[0060] Item 11: The computer system according to any one of the preceding items 8 to 10, wherein the connection between the nodes is determined using at least one of a relevance relationship between the sub-matrices based on a similarity measure, a context relationship based on a logical association between the sub-matrices, and an importance relationship based on the focus level of the sub-matrices. In an embodiment, the determined relevance relationship ensures that nodes corresponding to highly relevant sub-matrix encodings have strong connections and exhibit a tight association of information between them. This determination is then utilized to decide whether the nodes should be connected within a directed acyclic graph based on an associated scoring step.
[0061] Item 12: The target task includes at least one of classification, summary generation, and recommendation, and is the computer system according to any one of Items 8 to 11 above. Thereby, generality is provided to the target large language model using the described embodiments, since the target tasks to be executed each include various useful tasks that are uniquely beneficial, but the described embodiments that overcome the maximum token number limit associated with the target large language model tasked with processing the received long text can utilize the same data and features available for use.
[0062] Item 13: The step of using the graph neural network to perform the dynamic graph construction and the node feature transfer to repeatedly generate the updated graph including the series of most relevant node features and the connection relationships further includes: defining a preliminary directed acyclic graph; and for each of a series of nodes in the defined preliminary directed acyclic graph, aggregating the features of neighboring nodes by weighted averaging or splicing operations on the neighboring node features, and is the computer system according to any one of Items 8 to 12 above. In such an embodiment, in each round of GNN iteration, the GNN model updates the node features and transfers information based on the node features and connection relationships. As a result, by utilizing relevant probabilistic relationships, it is ensured that only the nodes that need to be calculated are calculated, while other nodes can be ignored, reducing the amount of calculation, thereby improving the efficiency and performance of the target large language model using the described embodiments in processing the received text that exceeds a given maximum token number limit.
[0063] Item 14: The method to be executed is as follows: applying an update function to fuse the collected neighborhood features with a series of current features of the target node to generate updated node features; and further transferring the generated updated node features to the next iteration of the GNN layer. The computer system according to any one of Items 8 to 13 above. This also reduces the amount of calculation and functions to improve the efficiency and performance of a target large language model using the described embodiments to process received text that exceeds a given maximum token number limit.
[0064] Item 15: A computer program product, the computer program product comprising: one or more computer-readable tangible storage media, and program instructions stored in at least one of the one or more computer-readable tangible storage media, the program instructions being executable by a processor to execute a method, the method comprising: receiving a target text; dividing an attention matrix associated with the target text into a series of submatrices; encoding fixed-length vectors corresponding to the series of submatrices by leveraging a gated recurrent unit (GRU) neural network; constructing a directed acyclic graph in which the encoded fixed-length vectors have nodes, where the connections between the nodes are defined based on a target task; repeatedly generating an updated graph including a series of most relevant node features and connection relationships by leveraging a graph neural network (GNN) to perform dynamic graph construction and node feature transfer; and generating one or more summaries regarding the received target text by extracting information from the updated graph.
[0065] Item 16: The computer program product according to item 15, wherein the received target text contains a number of tokens exceeding the maximum token number limit associated with the target large language model. In such an embodiment, the received target text is not processable by the target large language model until the steps are executed according to the described embodiment. Thus, the received text containing a number of tokens exceeding the maximum token number limit further enables the execution of the described method to overcome the maximum token number limit while providing additional input that can be processed by the target large language model.
[0066] Item 17: The computer program product according to any one of the preceding items 15 to 16, wherein each sub-matrix of the series of sub-matrices represents a semantically independent context. This ensures that any subsequently generated representation associated with a portion of the target text still corresponds to a relevant context, which can be utilized during subsequent stages to ensure that the meaning and features of the target text are maintained.
[0067] Item 18: The computer program product according to any one of the preceding items 15 to 17, wherein the connection between the nodes is determined using at least one of a relevance relationship between the sub-matrices based on a similarity measure, a context relationship based on a logical association between the sub-matrices, and an importance relationship based on the focus level of the sub-matrices. In an embodiment, the determined relevance relationship ensures that nodes corresponding to highly relevant sub-matrix encodings have strong connections and exhibit a tight association of information between them. This determination is then utilized to decide whether the nodes should be connected within a directed acyclic graph based on an associated scoring stage.
[0068] Item 19: The target task includes at least one of classification, summary generation, and recommendation, and is the computer program product according to any one of Items 15 to 18 above. Thereby, generality is provided to the target large language model using the described embodiments, since the target tasks to be executed each include various useful tasks that are uniquely beneficial, but the described embodiments that overcome the maximum token number limit associated with the target large language model tasked with processing the received long text can utilize the same data and features that become available.
[0069] Item 20: The procedure of using the graph neural network to perform the dynamic graph construction and the node feature transfer to repeatedly generate the updated graph including the series of most relevant node features and the connection relationships: the procedure of defining a preliminary directed acyclic graph; and for each of a series of nodes in the defined preliminary directed acyclic graph, further including the procedure of aggregating the features of neighboring nodes by weighted averaging or splicing operations on the neighboring node features, and is the computer program product according to any one of Items 15 to 19 above. In such an embodiment, in each round of GNN iteration, the GNN model updates the node features and transfers information based on the node features and connection relationships. As a result, by utilizing relevant probabilistic relationships, it is ensured that only the nodes that need to be calculated are calculated, while other nodes can be ignored, reducing the amount of calculation, thereby improving the efficiency and performance of the target large language model using the described embodiments in processing the received text that exceeds a given maximum token number limit.
[0070] It can be understood that FIGS. 2 to 5 only provide illustrations of exemplary implementations and do not imply any limitations on how different embodiments can be implemented. Based on design and implementation requirements, many modifications can be made to the shown environment.
[0071] The descriptions of the various embodiments of the present invention are presented for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, the practical application, or a technical improvement over the technologies found in the marketplace, or to enable other skilled artisans to understand the embodiments disclosed herein.
Claims
1. 1. A computer-based method for overcoming a maximum token number limitation in a large language model, the method comprising: receiving a target text; dividing an attention matrix associated with the target text into a set of sub-matrices; utilizing a gated recurrent unit (GRU) neural network to encode fixed length vectors corresponding to the set of sub-matrices; constructing a directed acyclic graph in which the encoded fixed-length vector has nodes, where connections between the nodes are defined based on the target task; Leveraging a graph neural network (GNN) to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph that includes a set of the most relevant node features and connections; and generating one or more summaries for the received target text by extracting information from the updated graph; A computer-based method comprising:
2. The computer-based method of claim 1 , wherein the received target text has a number of tokens that exceeds a maximum token limit associated with a target large language model.
3. The computer-based method of claim 1 , wherein each sub-matrix of the set of sub-matrices represents a semantically independent context.
4. 2. The computer-based method of claim 1, wherein the connections between the nodes are determined using at least one of: relevance relationships between the sub-matrices based on a similarity measure, context relationships based on logical associations between the sub-matrices, and importance relationships based on focus levels of the sub-matrices.
5. The computer-based method of claim 1 , wherein the target task comprises at least one of classification, summary generation, and recommendation.
6. Utilizing the graph neural network to perform the dynamic graph construction and the node feature transfer to iteratively generate the updated graph including the set of most relevant node features and the connectivity relationships, comprising: defining a preliminary directed acyclic graph; and aggregating the features of neighboring nodes by performing a weighted averaging or splicing operation on the features of neighboring nodes for each of a series of nodes in the defined preliminary directed acyclic graph; 6. The computer-based method of claim 1, further comprising:
7. applying an update function to fuse the collected neighborhood features with the set of current features of the target node to generate updated node features; and transferring the generated updated node features to a next iteration of the GNN layer; The computer-based method of claim 6 further comprising:
8. 1. A computer system, comprising: A computer system comprising one or more processors, one or more computer readable memories, one or more computer readable tangible storage media, and program instructions stored on at least one of the one or more computer readable tangible storage media for execution by at least one of the one or more processors via at least one of the one or more computer readable memories, wherein the computer system: receiving a target text; dividing an attention matrix associated with the target text into a set of sub-matrices; utilizing a gated recurrent unit (GRU) neural network to encode fixed length vectors corresponding to the set of sub-matrices; constructing a directed acyclic graph in which the encoded fixed-length vector has nodes, where connections between the nodes are defined based on the target task; Leveraging a graph neural network (GNN) to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph that includes a set of the most relevant node features and connections; and generating one or more summaries for the received target text by extracting information from the updated graph; It is possible to carry out a method comprising the steps of: Computer system.
9. The computer system of claim 8 , wherein the received target text includes a number of tokens that exceeds a maximum token limit associated with a target large language model.
10. The computer system of claim 8 , wherein each submatrix of the set of submatrices represents a semantically independent context.
11. 9. The computer system of claim 8, wherein the connections between the nodes are determined using at least one of: relevance relationships between the sub-matrices based on a similarity measure, context relationships based on logical associations between the sub-matrices, and importance relationships based on focus levels of the sub-matrices.
12. The computer system of claim 8 , wherein the target task includes at least one of classification, summary generation, and recommendation.
13. Utilizing the graph neural network to perform the dynamic graph construction and the node feature transfer to iteratively generate the updated graph including the set of most relevant node features and the connectivity relationships, comprising: defining a preliminary directed acyclic graph; and aggregating the features of neighboring nodes by performing a weighted averaging or splicing operation on the features of neighboring nodes for each of a series of nodes in the defined preliminary directed acyclic graph; 13. The computer system of claim 8, further comprising:
14. applying an update function to fuse the collected neighborhood features with the set of current features of the target node to generate updated node features; and transferring the generated updated node features to a next iteration of the GNN layer; The computer system of claim 13 further comprising:
15. 1. A computer program comprising: The program instructions, executable by a processor, are capable of performing a method, the method comprising: Steps for receiving the target text: dividing an attention matrix associated with the target text into a set of sub-matrices; utilizing a gated recurrent unit (GRU) neural network to encode fixed length vectors corresponding to the set of sub-matrices; constructing a directed acyclic graph in which the encoded fixed-length vector has nodes, where connections between the nodes are defined based on a target task; Leveraging a graph neural network (GNN) to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph containing a set of the most relevant node features and connections; and generating one or more summaries for the received target text by extracting information from the updated graph; having Computer program.
16. 16. The computer program product of claim 15, wherein the received target text includes a number of tokens that exceeds a maximum token limit associated with a target large language model.
17. 16. The computer program product of claim 15, wherein each sub-matrix of the series of sub-matrices represents a semantically independent context.
18. 16. The computer program product of claim 15, wherein the connections between the nodes are determined using at least one of: relevance relationships between the sub-matrices based on a similarity measure, context relationships based on logical associations between the sub-matrices, and importance relationships based on focus levels of the sub-matrices.
19. The computer program product of claim 15 , wherein the target task comprises at least one of classification, summary generation, and recommendation.
20. The steps of performing the dynamic graph construction and the node feature transfer using the graph neural network to iteratively generate the updated graph including the set of most relevant node features and the connectivity relationships include: A procedure for defining a preliminary directed acyclic graph; and aggregating, for each of a set of nodes in the defined preliminary directed acyclic graph, features of neighboring nodes by weighted averaging or splicing operations on the features of neighboring nodes; 20. The computer program of claim 15, further comprising: