Covering maximum lexical constraints for large language models
By splitting the attention matrix and building directed acyclic graphs, using graph neural networks to construct dynamic graph construction and node feature transfer, the problem of maximum word element limitation of large language models is solved, and effective processing of long texts and information saving is achieved.
Patent Information
- Application Number
- CN202411723167.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-03
AI Technical Summary
Due to the maximum word element limitation, large language models cannot effectively process long texts exceeding this limit, resulting in information loss and increased computational complexity.
By receiving the target text, the associated attention matrix is split into a series of submatrices, the gated recurrent unit neural network encodes fixed length vectors, and a directed acyclic graph is constructed, and the graph neural network is used to perform dynamic graph construction and node feature transfer, and an update graph is generated to extract information.
Overcome the problem of the maximum word element limitation of large language models, reduce the complexity of global computing, establish local semantic relationships, and dynamic computing nodes, thereby effectively processing long texts and reducing information loss.
Smart Images

Figure CN120087334A_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] The present application generally relates to computer processing, and more particularly to overcoming the maximum token limit in large language models.
[0002] Large language models have become increasingly popular due to their ability to understand and generate human-like text. Many enterprises are actively investing in discovering and leveraging opportunities to utilize large language models for various end-use applications designed to improve efficiency and competitiveness. For example, enterprises can use large language models to automate tasks, gain insights, improve customer experience, generate content, and so on. Therefore, large language models with increased flexibility and practicality are highly desirable. SUMMARY OF THE INVENTION
[0003] According to one embodiment, there is provided a method, computer system, and computer program product for overcoming the maximum token limit in large language models. The embodiment may include receiving a target text. The embodiment may also include splitting an attention matrix associated with the target text into a series of sub-matrices. The embodiment may further include using a gated recurrent unit (GRU) neural network to encode fixed-length vectors corresponding to the series of sub-matrices. The embodiment may also include constructing a directed acyclic graph, where the encoded fixed-length vectors include nodes, and where the connections between the nodes are defined based on the target task. The embodiment may further include using a graph neural network (GNN) to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of most relevant node features and connection relationships. The embodiment may also include generating one or more summaries for the received target text by extracting information from the updated graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] These and other objects, features, and advantages of the present disclosure will become apparent from the following detailed description of illustrative embodiments of the present disclosure when read in conjunction with the accompanying drawings. The various features of the drawings are not to scale as the illustrations are for the purpose of assisting a person skilled in the art in understanding the invention in conjunction with the detailed description. In the drawings:
[0005] Figure 1 An exemplary networked computer environment according to at least one embodiment is shown;
[0006] Figure 2 An operational flowchart of an exemplary process for overcoming the maximum token limit in large language models according to at least one embodiment is shown; and
[0007] Figure 3 An exemplary process for splitting an attention matrix associated with a received target text into a series of sub-matrices according to at least one embodiment is shown;
[0008] Figure 4illustrates an exemplary process for constructing a directed acyclic graph according to at least one embodiment; and
[0009] Figure 5 depicts an illustrative process for performing dynamic graph construction and node feature transfer using a graph neural network (GNN) to iteratively generate an updated graph including a series of most relevant node features and connection relationships according to at least one embodiment. DETAILED DESCRIPTION
[0010] Detailed embodiments of the claimed structures and methods are disclosed herein; however, it is to be understood that the disclosed embodiments are merely illustrative of the claimed structures and methods that may be implemented in various forms. However, this disclosure may be implemented in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. In the description, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments.
[0011] It should be understood that the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more such surfaces unless the context clearly dictates otherwise.
[0012] Embodiments of the present application generally relate to computer processing, and in particular to overcoming the maximum token limit in large language models. The exemplary embodiments described below provide a system, method, and program product for receiving a target text, splitting an attention matrix associated with the target text into a series of sub-matrices, encoding fixed-length vectors corresponding to the series of sub-matrices using a gated recurrent unit neural network, constructing a directed acyclic graph, performing dynamic graph construction and node feature transfer using a graph neural network to iteratively generate an updated graph including a series of most relevant node features and connection relationships, and generating one or more summaries for the received target text by extracting information from the updated graph, wherein, in the directed acyclic graph, the encoded fixed-length vectors include nodes and the connections between the nodes are defined based on the target task.
[0013] As previously mentioned, large language models (LLMs) have become increasingly prevalent due to their ability to understand and generate human-like text. Many enterprises are actively investing in discovering and exploiting opportunities to utilize LLMs for various end uses designed to improve efficiency and competitiveness. For example, enterprises can use LLMs to automate tasks, gain insights, improve the customer experience, generate content, and so on. Therefore, LLMs with increased flexibility and utility are highly desirable.
[0014] However, there are several challenges and limitations associated with leveraging large language models. For example, many large language models have undesirable limitations related to the maximum language token limit that a given large language model can handle. For example, an exemplary large language model may only be able to process a token sequence that is less than or equal to 32,000 tokens in length. If a given large language model receives text that includes a token sequence that exceeds the maximum token limit associated with the given large language model, this may result in any text after the token limit being discarded, leading to information loss. Recently proposed methods for addressing the maximum token limit typically involve shortening the received "long text" (text that exceeds the given maximum token limit) through combinatorial retrieval or summarization techniques. However, because these methods do not directly process the received long text, they generally cannot perform fine-grained reading comprehension. Additionally, the proposed methods typically need to be considered during the training phase and cannot be easily applied to existing LLM models. Therefore, improved methods for overcoming the maximum token limit of large language models and avoiding these drawbacks are beneficial for enterprises seeking to adopt LLMs with increased model flexibility and utility.
[0015] Therefore, a method, computer system, and computer program product for overcoming the maximum token limit in large language models are provided. The method, system, and computer program product can receive a target text. The method, system, and computer program product can identify defects in a printing operation based on tracked printing data. Then, the method, system, and computer program product can divide the attention matrix associated with the target text into a series of sub-matrices. The method, system, and computer program product can utilize a gated recurrent unit neural network to encode fixed-length vectors corresponding to the series of sub-matrices. Next, the method, system, and computer program product can construct a directed acyclic graph, where the encoded fixed-length vectors include nodes, and the connections between the nodes are defined based on the target task. Then, the method, system, and computer program product can utilize a graph neural network to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of the most relevant node features and connection relationships. Thereafter, the method, system, and computer program product can generate one or more summaries for the received target text by extracting information from the updated graph. Furthermore, the method, system, and computer program product provide an improved method for overcoming the maximum token limit of large language models. The described embodiments functionally combine the naive Bayes method with the use of a gated recurrent unit neural network as long-term memory storage to overcome the maximum token limit imposed by a given large language model. The currently described embodiments utilize a gated recurrent unit neural network to group-encode subsequences of tokens in the received target text, enabling the calculation of the attention matrix, and subsequently using the naive Bayes algorithm to transform the grouped computational units into a directed acyclic graph (DAG). In an embodiment, each node in the constructed DAG represents a computational unit, and whether each computational unit needs to be computed is dynamic, contrary to previously proposed methods that must compute all units. The DAG decomposes the calculation process of the attention matrix into several parts, each corresponding to a node on the DAG. During the construction of the DAG, the computational units are not executed, which means that the calculation of the DAG is essentially "lazy", and each computational unit is only computed when it is needed. In the currently described embodiments, whether a computational node is computed depends on the probabilistic relationship between the previously computed nodes and the given current node, which also uses the naive Bayes method to compute. Thus, the described embodiments overcome the limitations and challenges associated with the previously described methods and allow for summary generation (and the execution of other tasks) for the received "long text", which includes a token sequence that exceeds the given maximum token limit for a given large language model.
[0016] The present invention can be a system, method, and / or computer program product at any possible level of integration of technical details. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0017] Aspects of the present disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). With respect to any flowchart, depending on the technology involved, operations may be performed in an order different from the order shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.
[0018] An embodiment of a computer program product (“CPP embodiment” or “CPP”) is a term used in the present disclosure to describe any collection of one or more storage media (also referred to as “media”) jointly included in a collection of one or more storage devices, the collection of one or more storage devices jointly including machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device that can hold and store instructions used by a computer processor. By way of non-limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded devices such as punched cards or pits / lands formed in the main surface of a disk, or any suitable combination of the foregoing. A computer-readable storage medium, as the term is used in the present disclosure, should not be construed to store in the form of a transient signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses propagating through an optical fiber cable, electrical signals transmitted through a wire, and / or other transmission media. As will be understood by those skilled in the art, data is typically moved at certain incidental points in time during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but this does not render the storage device transient because the data is not transient when it is stored.
[0019] Now refer to Figure 1, The computing environment 100 includes an example of an environment for executing at least some of the computer code involved in performing the methods of the present invention, such as data processing program / code 150. In addition to the data processing code 150, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end-user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including processing circuitry 120 and a cache 121), a communication fabric 111, volatile memory 112, a permanent storage device 113 (including an operating system 122 and data processing code 150, as described above), a peripheral device set 114 (including a user interface (UI), a device set 123, a storage device 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. The remote server 104 includes a remote database 130. The public cloud 105 includes a gateway 140, a cloud coordination module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0020] The computer 101 can take the form of a desktop computer, a laptop computer, a tablet computer, a smart phone, a smart watch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device now known or developed in the future that is capable of running programs, accessing a network, or querying a database such as the remote database 130. As is well known in the computer art, and depending on the technology, the performance of computer-implemented methods can be distributed among multiple computers and / or among multiple locations. On the other hand, in this presentation of the computing environment 100, the discussion focuses in detail on a single computer, specifically the computer 101, to keep the presentation as simple as possible. The computer 101 can be located in the cloud, even if Figure 1 it is not shown in the cloud, on the other hand, the computer 101 does not need to be in the cloud, unless to any extent that can be definitely indicated.
[0021] The processor set 110 includes one or more computer processors of any type now known or developed in the future. The processing circuitry 120 can be distributed across multiple packages, such as multiple cooperative integrated circuit chips. The processing circuitry 120 can implement multiple processor threads and / or multiple processor cores. The cache 121 is a memory located within the processor chip package and is typically used for data or code that should be made available for rapid access by threads or cores running on the processor set 110. The cache memory is typically organized into multiple levels based on its relative proximity to the processing circuitry. Alternatively, some or all of the cache in the processor set can be located "off-chip". In some computing environments, the processor set 110 can be designed to work with qubits and perform quantum computing.
[0022] Computer-readable program instructions are typically loaded onto computer 101 so that a processor set 110 of the computer 101 executes a series of operational steps to implement a computer-implemented method, such that the instructions so executed will instantiate the method specified in the flowchart and / or the narrative description of the computer-implemented method included in this document (collectively referred to as "the method of the present invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the method of the present invention. In computing environment 100, at least some of the instructions for executing the method of the present invention may be stored in data processing code 150 in permanent storage device 113.
[0023] Communication structure 111 is a signal conduction path that allows the various components of computer 101 to communicate with each other. Typically, this structure consists of switches and conductive paths, such as switches and conductive paths that make up a bus, a bridge, a physical input / output port, etc. Other types of signal communication paths can be used, such as fiber optic communication paths and / or wireless communication paths.
[0024] Volatile memory 112 is any type of volatile memory known now or developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory is characterized by random access, but this is not required unless specifically stated. In computer 101, volatile memory 112 is located in a single package and inside the computer 101, but, alternatively or additionally, volatile memory can be distributed in multiple packages and / or located external to the computer 101.
[0025] Permanent storage 113 is any form of non-volatile storage for a computer known now or developed in the future. The non-volatility of this memory means that the stored data is retained regardless of whether power is supplied to computer 101 and / or directly to permanent memory 113. Permanent memory 113 can be read-only memory (ROM), but typically at least a portion of the permanent memory allows for the writing, deletion, and re-writing of data. Some common forms of persistent storage include disk and solid-state storage devices. Operating system 122 can take several forms, such as various known proprietary operating systems or open-source portable operating system interface type operating systems that employ a kernel. The code included in data processing program 150 typically includes at least some of the computer code involved in executing the method of the present invention.
[0026] The peripheral device set 114 includes the peripheral device set of the computer 101. Data communication connections between the peripheral devices and other components of the computer 101 can be implemented in various ways, such as Bluetooth connections, near-field communication (NFC) connections, connections made by cables (such as universal serial bus (USB)-type cables), plug-in connections (e.g., secure digital (SD) cards), connections made through local communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, the UI device set 123 may include components such as display screens, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. The storage device 124 is an external storage device, such as an external hard disk drive, or a pluggable storage device, such as an SD card. The storage device 124 can be permanent and / or volatile. In some embodiments, the storage device 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 needs to have a large amount of storage (e.g., in the case where the computer 101 locally stores and manages a large database), the storage can be provided by a peripheral storage device designed to store a very large amount of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor can be a thermometer, and another sensor can be a motion detector.
[0027] The network module 115 is a collection of computer software, hardware, and firmware that allows the computer 101 to communicate with other computers via the WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi signal transceiver, software for packetizing and / or depacketizing data transmitted over the communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control function and the network forwarding function of the network module 115 are executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networks (SDN)), the control function and the forwarding function of the network module 115 are executed on physically separate devices, such that the control function manages several different network hardware devices. The computer-readable program instructions for performing the methods of the present invention can generally be downloaded to the computer 101 from an external computer or an external storage device via a network adapter card or a network interface included in the network module 115.
[0028] The WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances by any technology known now or developed in the future for transmitting computer data. In some embodiments, the WAN may be replaced and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area, such as a Wi-Fi network. The WAN and / or LAN typically includes computer hardware such as copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers.
[0029] The end user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating the computer 101) and can take any form discussed above in connection with the computer 101. The EUD 103 typically receives useful and valuable data from the operation of the computer 101. For example, in the hypothetical case where the computer 101 is designed to provide recommendations to an end user, the recommendation will typically be transmitted from the network module 115 of the computer 101 to the EUD 103 via the WAN 102. In this way, the EUD 103 can display or otherwise present the recommendation to the end user. In some embodiments, the EUD 103 can be a client device such as a thin client, thick client, mainframe, desktop computer, etc.
[0030] The remote server 104 is any computer system that provides at least some data and / or functionality to the computer 101. The remote server 104 can be controlled and used by the same entity operating the computer 101. The remote server 104 represents a machine that collects and stores useful and valuable data used by other computers such as the computer 101. For example, in the hypothetical case where the computer 101 is designed and programmed to provide recommendations based on historical data, the historical data can be provided to the computer 101 from the remote database 130 of the remote server 104.
[0031] A public cloud 105 is any computer system that can be used by multiple entities and provides on-demand availability of computer system resources and / or other computing capabilities, particularly data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically utilizes the sharing of resources to achieve economies of scale and consistency. The direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments running on various computers that make up the set of host physical machines 142, which is the universe of physical computers within and / or available to the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from a set of virtual machines 143 and / or containers from a container group 144. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after instantiation of the VCE. The cloud orchestration module 141 manages the transfer and storage of the images, deploys new instantiations of the VCE, and manages the active instantiations of the VCE deployment. The gateway 140 is a collection of computer software, hardware, and firmware that allows the public cloud 105 to communicate via the WAN 102.
[0032] Some further explanations of the virtualized computing environment (VCE) will now be provided. A VCE can be stored as an "image". New active instances of the VCE can be instantiated from this image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows for the existence of multiple isolated user space instances, called containers. From the perspective of the programs running within them, these isolated user space instances typically behave as actual computers. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU capabilities, and quantifiable hardware capabilities. However, a program running within a container can only use the contents of the container and the devices allocated to the container, which is a feature known as containerization.
[0033] A private cloud 106 is similar to a public cloud 105, except that computing resources are only available to a single enterprise. Although the private cloud 106 is depicted as communicating with the WAN 102, in other embodiments, the private cloud can be completely disconnected from the Internet and only accessible through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types) that are typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable coordination, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.
[0034] According to this embodiment, the data processing program 150 can be a program capable of receiving a target text. The data processing program 150 can then divide the attention matrix associated with the target text into a series of sub-matrices. Next, the data processing program 150 can use a gated recurrent unit neural network to encode fixed-length vectors corresponding to the series of sub-matrices. The data processing program 150 can then construct a directed acyclic graph, where the encoded fixed-length vectors include nodes, and where the connections between the nodes are defined based on the target task. Next, the data processing program 150 can use a graph neural network to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of the most relevant node features and connection relationships. Thereafter, the data processing program 150 can generate one or more summaries for the received target text by extracting information from the updated graph. Furthermore, the data processing program 150 provides an improved method for overcoming the maximum token limit of large language models. The described embodiment functionally combines the naive Bayes method with the use of a gated recurrent unit neural network as long-term memory storage to overcome the maximum token limit imposed by a given large language model. The currently described embodiment uses a gated recurrent unit neural network to group-encode subsequences of tokens in the received target text, enabling the calculation of the attention matrix, and then uses the naive Bayes algorithm to transform the grouped computational units into a directed acyclic graph (DAG). In an embodiment, each node in the constructed DAG represents a computational unit, and whether each computational unit needs to be computed is dynamic, contrary to previously proposed methods that must compute all units. The DAG decomposes the calculation process of the attention matrix into several parts, each corresponding to a node on the DAG. During the construction of the DAG, the computational units are not executed, which means that the calculation of the DAG is essentially "lazy", and each computational unit is only computed when it is needed. In the currently described embodiment, whether a computational node is computed depends on the probabilistic relationship between the previously computed nodes and the given current node, which also uses the naive Bayes method to compute. Thus, the described embodiment overcomes the limitations and challenges associated with the previously described methods and allows for summary generation (and the execution of other tasks) for the received "long text", where the "long text" includes a token sequence that exceeds the given maximum token limit for a given large language model.
[0035] Now refer to Figure 2 , an operational flowchart of an illustrative process 200 for overcoming the maximum token limit in a large language model is provided according to at least one embodiment.
[0036] At 202, the data processing program 150 can receive the target text. In the context of the present disclosure, the target text can refer to any natural language text, encoded or programming language, data or structured text, mathematical equations, scientific and technical texts, or any other type of desired target text that can include any number of desired characters or tokens. In some embodiments, the received target text can be contained within any suitable desired format, and the text can be extracted from the desired format using known text extraction techniques. The data processing program 150 is configured to process the target text including multiple tokens, and the number of tokens may exceed the maximum token limit of the target large language model (“long text”). For example, in an embodiment, the data processing program 150 can receive an exemplary target text “T1” with a length of 40,000 tokens, which is intended to be input into an exemplary target large language model “LLM1” with a maximum token limit of 32,000.
[0037] At 204, the data processing program 150 can split the attention matrix associated with the target text into a series of sub-matrices. Figure 3 Illustrated is an exemplary process of splitting the attention matrix associated with the received target text into a series of sub-matrices according to at least one embodiment. As Figure 3 shown in the illustrative process 300, at this step, the data processing program 150 can feed the received target text 310 into an exemplary pointer network 320 to perform semantic segmentation. The text segments obtained at 330 are semantically consistent and shorter in length than the original received target text at 310. As shown at 340, the obtained text segments each correspond to their own context. The data processing program 150 can extract one of the N*N sub-matrices from the attention matrix (where N is the length of the context) to obtain a series of sub-matrices. Thus, the series of sub-matrices is also associated with its own unique context. Then, this process can be repeated to decompose the original attention matrix into several different sub-matrices, as shown at 350, where each sub-matrix represents a semantically independent context. The data processing program 150 can then further process the series of sub-matrices obtained from the separated attention matrix.
[0038] Next, at 206, the data processing program 150 can use a gated recurrent unit (GRU) neural network to encode fixed-length vectors corresponding to the series of sub-matrices. For example, at this step, the data processing program 150 can use the Figure 4 sub-matrices 340 shown as input features by feeding the series of sub-matrices into the GRU network. The GRU network can, for example, reshape the sub-matrices into 784-dimensional embeddings. In an embodiment, any suitable recurrent neural network (RNN) capable of performing the above functions can be used to encode fixed-length vectors corresponding to the series of sub-matrices. As Figure 3As shown, data processing program 150 can input, for example, the first exemplary sub-matrix 350 and the second exemplary sub-matrix 360 into the corresponding RNNs 370 to obtain fixed-length vectors 380 and 390 respectively.
[0039] At 208, data processing program 150 can construct a directed acyclic graph, where the encoded fixed-length vectors correspond to nodes, and the connections between the nodes are defined based on the target task. Figure 4 An exemplary process 400 for constructing a directed acyclic graph according to at least one embodiment is shown. As Figure 4 shown and as described above, the corresponding sub-matrices 410 and 420 are fed into the RNN 430 to obtain fixed-length vectors 440 and 450 respectively. At step 208, data processing program 150 can be configured to construct a directed acyclic graph and define nodes therein, where each encoded fixed-length vector (such as vectors 440 and 450) serves as nodes 460 and 470 respectively, and each node represents a computing unit. Then, data processing program 150 can establish connections between the nodes based on the relationship between the requirements of the target task and the relevant information of interest. For example, in an embodiment, data processing program 150 can establish connections between the nodes based on the correlation relationship 475 by using a known similarity metric, such as cosine similarity or correlation coefficient, based on the correlation between the sub-matrices. The known similarity metric can be used to determine the connection strength between the nodes. Nodes corresponding to sub-matrices encoded with high correlation can have strong connections, indicating a close association of the information between them. In an embodiment, at 480, data processing program 150 can establish further connections between the nodes based on context connections. In the context of the present disclosure, the connections between the nodes can be determined based on the context relationship within the sub-matrices. For example, if two sub-matrices are adjacent or have a logical relationship in the original text, data processing program 150 can establish a connection between them. In an embodiment, at 485, data processing program 150 can also establish further connections between the nodes based on importance connections. In the context of the present disclosure, the importance connection or importance relationship between the nodes can be determined based on the importance or focus level of the sub-matrices. For example, if a sub-matrix contains important information or a key viewpoint, the nodes connected to the encoded node of that sub-matrix can have stronger connections. In an embodiment, when establishing connections, data processing program 150 can be configured to calculate a comprehensive score at 490 by considering the above three types of connections to determine whether two nodes should be connected. Any suitable known method can be used to calculate the score. In an embodiment, as shown at 495, the comprehensive score can be compared with a predetermined and user-adjustable threshold to control when a connection is established between the pair of nodes under consideration.
[0040] At 210, the data processing program 150 can then utilize a graph neural network (GNN) to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of the most relevant node features and connection relationships. In this step, the data processing program 150 can utilize a GNN model to iteratively facilitate information propagation and update based on node features and connection relationships. In an embodiment, the data processing program 150 can utilize a graph convolutional network (GCN), a GraphSAGE algorithm, a graph attention network (GAT), and any other suitable GNN model or algorithm. In Figure 5 An illustrative example of step 210 is depicted in Figure 5 , which includes an illustrative process 500 of performing dynamic graph construction and node feature transfer using a graph neural network (GNN), particularly a GCN model, according to at least one embodiment to iteratively generate an updated graph including a series of the most relevant node features and connection relationships. In a first step, the data processing program 150 can utilize a GCN to initialize node features such that the encoding vectors of each sub-matrix are used as the initial features of each node. Next, the data processing program 150 can construct a preliminary directed acyclic graph 510 based on the connection relationships between nodes (e.g., based on probabilities). Thus, the preliminary directed acyclic graph 510 describes the strength of the connections or relationships between nodes. Next, at loop 520, the data processing program can utilize a GNN to perform GNN layer iterations. In each iteration of the GNN layer, exemplary steps can be taken to update and propagate node features. For example, the layer iteration executed at loop 520 can include aggregating neighbor features to aggregate the neighbor node features of each node by using the preliminary directed acyclic graph 510. This can be achieved by a weighted average or concatenation operation on adjacent node features. The layer iteration executed at loop 520 can also include updating node features. For example, the collected neighbor features can be fused with the features of a given current node to generate new node features. This can be achieved by applying an update function, such as a graph convolution operation, a gated recurrent unit (GRU), a graph attention mechanism (GAT), or any other suitable mechanism or model. In an embodiment, the layer iteration executed at loop 520 can also include the transfer of node features. For example, the data processing program 150 can transfer the updated node features to the next iteration of the GNN layer associated with the next round of feature updates. Multiple rounds of iterations can be performed through a multi-layer GNN structure. In each round of GNN iteration, the GNN model updates node features and transmits information based on node features and connection relationships. Thus, using the relevant probability relationships ensures that only the nodes that need to be calculated will be calculated, while other nodes can be ignored, thereby reducing the computational amount.
[0041] In an embodiment, the data processing program 150 can be configured to include a stop condition, so that the end of the GNN iteration (performed at loop 520) can be determined according to the specific stop condition. For example, in an embodiment, the stop condition can be met after reaching a certain number of iterations, after the node features converge, or any other desired and configurable custom conditions. At 530, an updated graph 530 can be obtained, which includes a series of most relevant node features and connection relationships.
[0042] Thereafter, at 212, the data processing program 150 may generate one or more summaries of the received target text by extracting information from the updated graph. For example, at this step, the data processing program 150 may extract information from the updated graph to generate a summary by extracting key sentences based on node features or by classifying node features. In an embodiment, the data processing program 150 may generate other desired outputs at this step, such as recommendations or any other output that may be generated based on information extracted from the updated graph to limit or reduce the amount of tokens associated with the received target text, which may ultimately be input into a given large language model.
[0043] It will be appreciated that the data processing program 150 thus provides an improved method for overcoming the maximum word-gram limitation in large language models that overcomes the challenges observed in conventional and known methods for overcoming the maximum word-gram limitation in large language models.
[0044] For example, as described above, the described embodiments reduce the global computational complexity. In the traditional attention mechanism, each word needs to calculate the attention score together with all other words, resulting in a significant increase in computational complexity as the number of words increases. In the described embodiments, by separating the attention matrix, the global attention calculation is converted into calculations between local sub-matrices, thereby greatly reducing the computational load.
[0045] In addition, the embodiments described so far provide the benefit of establishing local relationships. By dividing the original attention matrix into multiple sub-matrices and using DAG to construct connections, local semantic relationships can be captured. This enables locating important information and avoids processing the entire text as a continuous sequence, thereby reducing the processing of irrelevant information while maintaining task relevance.
[0046] The embodiments described so far also allow for dynamic calculation of nodes. In the constructed DAG, the calculation of each node is dynamic, which is different from the traditional method that requires the calculation of the entire attention matrix. Based on the probabilistic relationship between nodes, only the nodes that need to be calculated are evaluated, while other nodes can be ignored. This further reduces the computational load and focuses only on the nodes that are meaningful given the current task and context.
[0047] It can also be understood that the described embodiments uniquely combine Naive Bayes with GRU (Gated Recurrent Unit) to solve the problem of rapidly expanding the number of tokens. The GRU can group-encode subsequences of representations, enabling the calculation of an attention matrix, and then using the Naive Bayes algorithm to transform the grouped computational units into a directed acyclic graph (DAG). Each node on the DAG represents a computational unit, and whether each computational unit needs to be computed is dynamic, contrary to previous methods that had to compute all units. The DAG decomposes the calculation process of the attention matrix into several parts, with each part corresponding to a node on the DAG. During the construction of the DAG, the computational units are not executed, meaning that the calculation of the DAG is "lazy", and each computational unit is only computed when it is needed. Thus, whether a computational node is computed depends on the probability relationship between previously computed nodes and the current node, which is also computed using the Naive Bayes method.
[0048] Therefore, the described method of splitting the attention matrix and constructing the DAG achieves efficient processing and information saving of long texts by reducing the global computational complexity, establishing local relationships, and dynamically computing nodes. This method makes full use of local and task relevance, enabling the model to process long texts more effectively and avoiding information loss during the processing of "long" texts that exceed the given maximum token limit of the target large language model.
[0049] The presently described embodiments may relate to the following examples:
[0050] Example 1: A computer-based method for overcoming the maximum token limit in a large language model, the method comprising: receiving a target text, splitting an attention matrix associated with the target text into a series of submatrices, using a Gated Recurrent Unit (GRU) neural network to encode fixed-length vectors corresponding to the series of submatrices, constructing a directed acyclic graph, wherein the encoded fixed-length vectors include nodes and wherein the connections between the nodes are defined based on a target task, using a Graph Neural Network (GNN) to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of most relevant node features and connection relationships, and generating one or more summaries for the received target text by extracting information from the updated graph. This allows the described embodiments to functionally combine the Naive Bayes method with the use of a gated recurrent unit neural network as long-term memory storage to overcome the maximum token limit imposed by a given large language model. This improves the generality of the large language model employing the described embodiments, as a result of the described embodiments using a gated recurrent unit neural network to group-encode subsequences of symbols in the received target text, enabling the calculation of an attention matrix, and subsequently using the Naive Bayes algorithm to transform the grouped computational units into a directed acyclic graph (DAG).
[0051] Example 2: The computer-based method according to Example 1, wherein the received target text includes a number of tokens that exceeds the maximum token limit associated with the target large language model. In such an embodiment, the received target text will not be processable by the target large language model until the steps are performed according to the described embodiments. Thus, the received text including a number of tokens that exceeds the maximum token limit provides additional input that can be processed by the target large language model while further enabling the described method to be performed functionally to overcome the maximum token limit.
[0052] Example 3: The computer-based method according to any one of the foregoing Examples 1-2, wherein each sub-matrix in the series of sub-matrices represents a semantically independent context. This ensures that any subsequently generated representation associated with a portion of the target text still corresponds to the relevant context, which can be utilized during subsequent steps to ensure that the meaning and characteristics of the target text are maintained.
[0053] Example 4: The computer-based method according to any one of the foregoing Examples 1-3, wherein the connections between nodes are determined using at least one of a correlation relationship between sub-matrices based on a similarity metric, a context relationship based on a logical association between sub-matrices, and an importance relationship based on the attention level of sub-matrices. In an embodiment, the determined correlation relationship ensures that nodes encoded with sub-matrices having high correlation will have strong connections, indicating a close association of information between them. Then this determination is utilized to determine whether nodes should be connected within the directed acyclic graph based on an associated scoring step.
[0054] Example 5: The computer-based method according to any one of the above Examples 1-4, wherein the target task includes at least one of classification, summary generation, and recommendation. This provides generality to the target large language model that employs the described embodiments, since the target tasks performed can include various useful tasks that are each uniquely valuable, but utilize the same data and features made available by using the described embodiments to overcome the maximum token limit associated with the target large language model that is assigned the task of processing the received long text.
[0055] Example 6: A computer-based method according to any one of the foregoing Examples 1-5, wherein the use of a graph neural network to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of most relevant node features and connection relationships also includes: defining a preliminary directed acyclic graph, and for each of a series of nodes in the defined preliminary directed acyclic graph, aggregating neighbor node features by weighted averaging or concatenating neighbor node features. In such an embodiment, in each round of GNN iteration, the GNN model updates node features and transmits information based on node features and connection relationships. Therefore, the use of relevant probabilistic relationships ensures that only nodes that need to be calculated are calculated, while other nodes can be ignored, thereby reducing the amount of calculation, and thereby improving the efficiency and performance of the target large language model that uses the described embodiment to process received text exceeding a given maximum word limit.
[0056] Example 7: A computer-based method as in any of Examples 1-6 above, wherein the method further comprises: applying an update function to fuse the collected neighbor features with a series of current features of the target node to generate updated node features; transferring the generated updated node features to the next iteration of the GNN layer. This is similarly used to reduce the amount of computation, thereby improving the efficiency and performance of the target large language model that uses the described embodiments to process received text exceeding a given maximum word limit.
[0057] Example 8: A computer system, the computer system comprising: one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored on at least one of the one or more computer-readable tangible storage media for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories, wherein the computer system is capable of performing a method that includes: receiving a target text, splitting an attention matrix associated with the target text into a series of sub-matrices, encoding fixed-length vectors corresponding to the series of sub-matrices using a gated recurrent unit (GRU) neural network, constructing a directed acyclic graph, wherein the encoded fixed-length vectors include nodes and wherein connections between the nodes are defined based on a target task, using a graph neural network (GNN) to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of most relevant node features and connection relationships, and generating one or more summaries for the received target text by extracting information from the updated graph. This allows the described embodiments to functionally combine the naive Bayes method with using a gated recurrent unit neural network as long-term memory storage to overcome the maximum token limit imposed by a given large language model. This improves the generality of the large language model employing the described embodiments, as a result of the described embodiments using a gated recurrent unit neural network to group-encode subsequences of tokens in the received target text such that an attention matrix can be computed and subsequently the grouped computational units are transformed into a directed acyclic graph (DAG) using the naive Bayes algorithm.
[0058] Example 9: The computer system according to Example 8, wherein the received target text includes a number of tokens that exceeds the maximum token limit associated with the target large language model. In such an embodiment, the received target text will not be processable by the target large language model until the steps are performed according to the described embodiments. Thus, the received text including a number of tokens that exceeds the maximum token limit provides additional input that can be processed by the target large language model while further functionally enabling the described method to be performed to overcome the maximum token limit.
[0059] Example 10: The computer system according to any one of the preceding Examples 8-9, wherein each sub-matrix in the series of sub-matrices represents a semantically independent context. This ensures that any subsequently generated representation associated with a portion of the target text still corresponds to the relevant context, which can be utilized during subsequent steps to ensure that the meaning and features of the target text are maintained.
[0060] Example 11: A computer system according to any of the foregoing examples 8-10, wherein the connection between nodes is determined using at least one of a correlation relationship between sub-matrices based on a similarity measure, a contextual relationship based on a logical association between sub-matrices, and an important relationship based on the attention level of the sub-matrices. In an embodiment, the determined correlation relationship ensures that nodes corresponding to sub-matrix encodings with high correlation will have strong connections, indicating a close association of information between them. This determination is then used to determine whether the nodes should be connected within the directed acyclic graph based on the associated scoring step.
[0061] Example 12: A computer system according to any one of the foregoing examples 8 to 11, wherein the target task includes at least one of classification, summary generation, and recommendation. This provides versatility for the target large language model using the described embodiments, because the target tasks performed can include a variety of useful tasks that are each uniquely valuable, but utilize the same data and features that become available using the described embodiments to overcome the maximum word-unit limitation associated with the target large language model that is assigned the task of processing the received long text.
[0062] Example 13: A computer system according to any one of the aforementioned Examples 8-12, wherein the use of a graph neural network to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of most relevant node features and connection relationships also includes: defining a preliminary directed acyclic graph, and for each of a series of nodes in the defined preliminary directed acyclic graph, aggregating neighbor node features by weighted averaging or concatenating neighbor node features. In such an embodiment, in each round of GNN iteration, the GNN model updates node features and transmits information based on node features and connection relationships. Therefore, the use of relevant probabilistic relationships ensures that only nodes that need to be calculated are calculated, while other nodes can be ignored, thereby reducing the amount of calculation, and thereby improving the efficiency and performance of the target large language model that uses the described embodiment to process received text that exceeds a given maximum word limit.
[0063] Example 14: A computer system according to any one of the foregoing examples 8 to 13, wherein the method executed further comprises: applying an update function to fuse the collected neighbor features with a series of current features of the target node to generate updated node features; transferring the generated updated node features to the next iteration of the GNN layer. This is similarly used to reduce the amount of computation, thereby improving the efficiency and performance of the target large language model that uses the described embodiment to process the received text exceeding the given maximum word limit
[0064] Example 15: A computer program product, the computer program product comprising one or more computer-readable tangible storage media and program instructions stored on at least one of the one or more computer-readable tangible storage media, the program instructions being executable by a processor capable of performing a method, the method comprising: receiving a target text, splitting an attention matrix associated with the target text into a series of sub-matrices, encoding a fixed-length vector corresponding to the series of sub-matrices using a gated recurrent unit (GRU) neural network, constructing a directed acyclic graph, wherein the encoded fixed-length vector comprises nodes and wherein connections between the nodes are defined based on a target task, using a graph neural network (GNN) to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph comprising a series of most relevant node features and connection relationships, and generating one or more summaries for the received target text by extracting information from the updated graph.
[0065] Example 16: The computer program product according to Example 15, wherein the received target text comprises a number of tokens that exceeds the maximum token limit associated with the target large language model. In such an embodiment, the received target text will not be processable by the target large language model until the steps are performed according to the described embodiment. Thus, the received text comprising a number of tokens that exceeds the maximum token limit provides additional input that can be processed by the target large language model while further functionally enabling the described method to be performed to overcome the maximum token limit.
[0066] Example 17: The computer program product as in any of the preceding Examples 15 - 16, wherein each sub-matrix in the series of sub-matrices represents a semantically independent context. This ensures that any subsequently generated representation associated with a portion of the target text still corresponds to the relevant context, which can be utilized during subsequent steps to ensure that the meaning and features of the target text are maintained.
[0067] Example 18: The computer program product as in any of the preceding Examples 15 - 17, wherein at least one of a correlation relationship between sub-matrices based on a similarity metric, a context relationship based on a logical association between sub-matrices, and an importance relationship based on the attention level of sub-matrices is used to determine the connections between nodes. In an embodiment, the determined correlation relationship ensures that nodes encoded corresponding to sub-matrices with high correlation will have strong connections, indicating a close association of information between them. The determination is then utilized to determine whether nodes should be connected within the directed acyclic graph based on an associated scoring step.
[0068] Example 19: A computer program product according to any one of the preceding Examples 15 to 18, wherein the target task includes at least one of classification, summary generation, and recommendation. This provides generality for the target large language model adopting the described embodiments, because the target tasks to be performed can include various useful tasks that are each uniquely valuable, but use the same data and features made available by using the described embodiments to overcome the maximum token limit associated with the target large language model, which is assigned the task of processing the received long text.
[0069] Example 20: A computer program product according to any one of the preceding Examples 15 - 19, wherein using a graph neural network to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of most relevant node features and connection relationships further includes: defining a preliminary directed acyclic graph, and for each of a series of nodes in the defined preliminary directed acyclic graph, aggregating neighbor node features by a weighted average or concatenation operation of neighbor node features. In such an embodiment, in each round of GNN iteration, the GNN model updates node features and transmits information based on node features and connection relationships. Therefore, using the relevant probability relationships ensures that only the nodes that need to be calculated are calculated, while other nodes can be ignored, thereby reducing the amount of computation, and thus improving the efficiency and performance of the target large language model that adopts the described embodiments to process the received text exceeding a given maximum token limit.
[0070] It can be understood that Figures 2 - 5 only illustrations of exemplary implementations are provided, without implying any limitation on how to implement different embodiments. Many modifications can be made to the described environment based on design and implementation requirements.
[0071] Descriptions of various embodiments of the present invention have been given for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be obvious to a person of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, practical applications, or improvements to technologies found in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A computer-based method for overcoming a maximum word-gram limitation in a large language model, the method comprising: receiving a target text; Splitting an attention matrix associated with the target text into a series of sub-matrices; Using a gated recurrent unit (GRU) neural network to encode a fixed-length vector corresponding to the series of sub-matrices; constructing a directed acyclic graph in which the encoded fixed-length vector includes nodes and in which connections between the nodes are defined based on a target task; Utilize graph neural network (GNN) to perform dynamic graph construction and node feature transfer to iteratively generate an updated graph including a series of most relevant node features and connection relationships; as well as One or more summaries for the received target text are generated by extracting information from the update graph.
2. The computer-based method of claim 1, wherein the received target text includes a number of tokens that exceeds a maximum token limit associated with a target large language model.
3. The computer-based method of claim 1, wherein each sub-matrix in the series of sub-matrices represents a semantically independent context.
4. A computer-based method according to claim 1, wherein the connection between the nodes is determined using at least one of the following: a correlation relationship between the sub-matrices based on a similarity measure, a contextual relationship based on a logical association between the sub-matrices, and an important relationship based on an attention level of the sub-matrices.
5. The computer-based method of claim 1, wherein the target task comprises at least one of: classification, summary generation, and recommendation.
6. The computer-based method of claim 1, wherein utilizing the graph neural network to perform the dynamic graph construction and the node feature transfer to iteratively generate the updated graph including the series of most relevant node features and the connection relationships further comprises: Define a preliminary directed acyclic graph; as well as For each node in a series of nodes in the defined preliminary directed acyclic graph, features of neighbor nodes are aggregated by weighted averaging or concatenating features of neighbor nodes.
7. The computer-based method of claim 6, further comprising: Applying an update function to fuse the collected neighbor features with a set of current features for the target node to generate updated node features; as well as The generated updated node features are transmitted to the next iteration of the GNN layer.
8. A computer system, comprising: One or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored on at least one computer-readable tangible storage medium of the one or more computer-readable tangible storage media, the program instructions being used to be executed by at least one of the one or more processors via at least one computer-readable memory of the one or more computer-readable memories, wherein the computer system is capable of executing the method according to any one of claims 1 to 7.
9. A computer program product, the computer program product comprising: One or more computer-readable tangible storage media and at least one computer-readable tangible storage medium program instructions stored in the one or more computer-readable tangible storage media, the program instructions being executable by a processor capable of performing the method according to any one of claims 1 to 7.
10. A system comprising modules respectively configured to perform the steps of the method according to any one of claims 1 to 7.