Text information processing method and related equipment

By performing full inference operations on the first processor and storing key information and value information of text information using video memory and CPU memory, the problem of waste of computing power caused by large storage space when the GPU independently processes text information is solved, and more efficient text information processing is achieved.

CN120218223APending Publication Date: 2025-06-27HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311803634.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the process of independently processing text information using the GPU, the key information and value information require a large amount of storage space, which requires multiple GPUs to be deployed, resulting in wasting the GPU's computing power.

Method used

By performing a full inference operation on the first processor, and using the video memory of the first processor and the memory of the CPU to jointly store part of the key information and value information of the text information, the waste of computing power of the first processor is avoided.

Benefits of technology

The computing power of the first processor and the storage space of the CPU are fully utilized, the efficiency of text information processing is improved, and the computing power of the GPU is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218223A_ABST
    Figure CN120218223A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text information processing method which can be used in the field of text processing in the field of artificial intelligence. The first processor transmits first key information and first value information of the first text information to the second processor so as to be stored in a memory of the second processor, second key information and second value information of the first text information are stored in a video memory of the first processor, the first processor is a GPU or an NPU, the second processor is a CPU, and the first processor and the second processor are connected with the first processor. The first key information and the first value information are generated by X first neural network modules in a machine learning model, and the second key information and the second value information are generated by other first neural network modules except the X first neural network modules; the advantage of high computing power of the first processor is fully utilized, and the advantage of large storage space of the CPU is also utilized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and particularly to a method for processing text information and related devices. Background Art

[0002] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0003] Using a machine learning model based on the attention mechanism to process text information is an application method in a scenario of artificial intelligence. With the development of artificial intelligence technology, the number of neural network layers in the machine learning model is increasing, and the number of model parameters is also increasing. To improve the efficiency of the machine learning model in processing text information, a caching technology for key (K) information and value (V) information has emerged.

[0004] Specifically, the process of using a machine learning model to process the input first text information may include a full-scale inference process and an incremental inference process. In the full-scale inference process, the entire first text information is input into the machine learning model, and a first word output by the machine learning model can be obtained, and the key information and value information of the first text information generated in the full-scale inference process are stored. In the incremental inference process, the first word can be used as the input of the machine learning model in the current step. In the calculation process of the current step of the machine learning model, the key information and value information of the first text information are combined to obtain a second word output by the machine learning model in the current step, and the key information and value information of the first word generated in the calculation process of the current step of the machine learning model are stored; the second word is used as the input of the machine learning model in the next step, and in the calculation process of the next step of the machine learning model, the key information and value information of the first text information and the first word are combined to obtain a third word output by the machine learning model in the next step; the above operation is repeatedly executed until all the words output by the machine learning model are obtained.

[0005] However, if all operations for processing text information using a machine learning model are independently completed by a GPU, since the above-mentioned key information and value information require a large storage space, in order to meet the storage requirements of the key information and value information, multiple GPUs need to be deployed, which will cause waste of the computing power of the GPUs. Summary of the Invention

[0006] Embodiments of the present application provide a method for processing text information, a method for processing text information, and related devices. The first processor performs a full-scale inference operation, and the video memory of the first processor and the memory of the CPU jointly store partial key information and value information of the first text information, which not only makes full use of the advantage of the strong computing power of the first processor but also makes use of the advantage of the large storage space of the CPU, avoiding waste of the computing power of the first processor.

[0007] Embodiments of the present application provide the following technical solutions:

[0008] In a first aspect, an embodiment of the present application provides a method for processing text information, which can be used in the field of text processing in the field of artificial intelligence. This method is used in the process of processing first text information through a machine learning model (for convenience of description, hereinafter referred to as the "first machine learning model"). Each of the L first neural network modules in the first machine learning model is a neural network module based on an attention mechanism, and L is an integer greater than or equal to 1.

[0009] In this method, during the process of the first processor performing full-scale inference on the first text information using the machine learning model, the first processor can obtain the first key information and the first value information (i.e., the first information) of the first text information generated by X of the L first neural network modules, and then transmit the first information to the second processor, so that the second processor stores the first information in its memory; where X is an integer greater than or equal to 1 and less than L, the first processor is a graphics processing unit GPU or an embedded neural network processing unit NPU, and the second processor is a central processing unit CPU.

[0010] During the process of the first processor performing full-scale inference on the first text information using the machine learning model, the first processor can also obtain the second key information and the second value information (i.e., the second information) of the first text information, and then store the second information in the video memory of the first processor; where the second information is generated by the other L - X first neural network modules except for the X first neural network modules among the L first neural network modules, that is, the X first neural network modules and the L - X first neural network modules are different.

[0011] In this implementation manner, the first processor is used to perform the computational operation of full-scale inference on the first text information using a machine learning model. However, during the aforementioned full-scale inference process, the key information and value information generated by some neural network modules in the machine learning model are transmitted to the CPU. That is, a part of the key information and value information of the first text information generated during the aforementioned full-scale inference process are transmitted to the CPU. There is a part of the key information and value information of the first text information in the memory of the CPU, and another part of the key information and value information of the first text information are stored in the video memory of the first processor. Thus, the large-capacity memory of the CPU can be utilized to assist in storing a part of the key information and value information of the first text information. That is, not only the advantage of the strong computing power of the first processor is fully utilized, but also the advantage of the large storage space of the CPU is utilized, avoiding the waste of the computing power of the first processor.

[0012] In a possible implementation manner, the above X first neural network modules can be the first X first neural network modules among the L first neural network modules included in the first machine learning model, and the above L-X first neural network modules can be the last L-X first neural network modules among the L first neural network modules included in the first machine learning model. That is, during the process of performing full-scale inference on the first text information through the first machine learning model, the first text information will be processed by the first X first neural network modules first, and then processed by the last L-X first neural network modules.

[0013] In this implementation manner, since the "sending operation of the first key information and the first value information" and the "computational operation corresponding to the first neural network module" do not interfere with each other, while the first processor sends the first key information and the first value information generated by the Xth first neural network module to the second processor, the first processor can start executing the computational operations corresponding to the last L-X first neural network modules, which is beneficial to improving the processing speed of the entire first machine learning model for the first text information. In addition, if the X first neural network modules are selected as the first X first neural network modules among the L first neural network modules, the storage space occupied by the aforementioned first key information and first value information in the video memory can be released in a timely manner. Then, the first processor can use the video memory with a larger storage space during the data processing process through the last L-X first neural network modules, which is beneficial to improving the performance when the first processor performs full-scale inference on the first text information.

[0014] In a possible implementation manner, the process of processing the first text information by the first machine learning model may further include: a process of performing at least one incremental inference using the first machine learning model. The process of performing each incremental inference using the first machine learning model includes: performing each incremental inference using L first neural network modules included in the first machine learning model. Among them, the execution entity of the computing operations corresponding to X first neural network modules among the foregoing L first neural network modules includes a second processor, and the execution entity of the computing operations corresponding to the L-X first neural network modules among the foregoing L first neural network modules is the first processor.

[0015] In this implementation manner, since in the process of incremental inference, the first key information and the first value information of the first text information are used when processing through the foregoing X first neural network modules, and the foregoing first key information and first value information are stored in the memory of the second processor. If the first processor is used to execute all the computing operations generated by the foregoing X first neural network modules, it is necessary to transfer the first key information and the first value information from the memory of the second processor to the first processor. In the case where the data volume of the first key information and the first value information is large, the process of transferring the "first key information and the first value information" from the memory of the second processor to the first processor will cause a large amount of time consumption. In this solution, the execution entity of the computing operations generated by the foregoing X first neural network modules includes a second processor. When the second processor is used to execute all the computing operations generated by the foregoing X first neural network modules, the second processor can directly read the first key information and the first value information from the memory, which helps to avoid the time consumption caused by the process of transferring the "first key information and the first value information" from the memory of the second processor to the first processor, and thus helps to improve the efficiency of the process of processing the first text information using the first machine learning model.

[0016] In a possible implementation manner, the first neural network module is a Transformer module. The first neural network module includes a neural network layer based on the attention mechanism and a feed-forward neural network layer FFN. In the process of performing incremental inference using the first machine learning model based on the second text information, the execution entity of the computing operations corresponding to the neural network layer based on the attention mechanism in the X first neural network modules includes a second processor, and the execution entity of the computing operations corresponding to the FFN in the X first neural network modules is the first processor.

[0017] Exemplarily, each first neural network module may include a neural network layer based on an attention mechanism and a feed-forward neural network layer FFN. The neural network layer based on the attention mechanism may include a linear neural network layer and a matrix multiplication neural network layer; FFN may include a linear neural network layer. Exemplarily, the neural network layer based on the attention mechanism may include Linear1, Linear2, Linear3, matmul 1, matmul 2, and Linear4. Then, the computational operations corresponding to the neural network layer based on the attention mechanism may include the operation operations of the aforementioned Linear1, Linear2, Linear3, matmul 1, matmul 2, and Linear4. FFN may include at least one linear neural network layer, and the computational operations corresponding to FFN may include the operation operations of the aforementioned at least one linear neural network layer.

[0018] In this implementation, during the process of incremental inference, since the key information and value information of the first text information are used in the computational operations corresponding to the neural network layer based on the attention mechanism in the first neural network module, and the key information and value information of the first text information are not required in the computational operations corresponding to FFN in the first neural network module, the execution entity of the computational operations corresponding to the neural network layer based on the attention mechanism in X first neural network modules includes a second processor, and the execution entity of the computational operations corresponding to FFN in X first neural network modules is the first processor. This not only helps avoid the time consumption caused by transferring the "first key information and first value information of the first text information" from the memory of the second processor to the first processor, but also makes full use of the computing power of the first processor, which is beneficial to further improving the efficiency of processing the first text information using the first machine learning model.

[0019] In a possible implementation manner, key information and value information may also be generated during the process of each incremental inference performed by the first machine learning model. The storage location of the key information and value information generated during the process of performing incremental inference by the first machine learning model is determined based on the execution entity that generates the key information and value information. For example, if the key information and value information are generated by the first processor, the key information and value information may be stored in the video memory of the first processor. If the key information and value information are generated by the second processor, the key information and value information may be stored in the memory of the second processor. Alternatively, the key information and value information generated during each incremental inference process may also be stored in the memory of the second processor, etc.

[0020] Optionally, during each incremental reasoning process performed by the first machine learning model, the key information and value information generated by the LX first neural network modules can be stored in the video memory of the first processor. Alternatively, the first processor can also transmit the key information and value information generated by the LX first neural network modules to the second processor, and then store them in the memory of the second processor.

[0021] Optionally, during each incremental reasoning process performed by the first machine learning model, the key information and value information generated by the X first neural network modules may be stored in the memory of the second processor.

[0022] In one possible implementation, the method is applied to a text information processing system, and the processor system of the aforementioned text information may include at least two first processors and a second processor, and the process of the first processor using the first machine learning model to perform full inference on the first text information includes synchronous summation Allreduce between at least two first processors. In the case where there is no separate first interconnection link between at least two first processors, that is, there is no first interconnection link for direct communication between at least two first processors, then the Allreduce between at least two first processors needs to be completed with the help of the second interconnection link between the first processor and the second processor. Exemplarily, each of the at least two first processors sends the data to be synchronously summed to the second processor through the second interconnection link, and the second processor aggregates the data to be synchronously summed sent by all the first processors to obtain aggregated data, and the second processor then sends the aggregated data to each first processor through the second interconnection link, thereby completing the Allreduce between at least two first processors.

[0023] Since the operation of "the first processor sending the first key information and the first value information of the first text information to the second processor" is performed through the second interconnection link, and the operation of "All reduce between at least two first processors" is performed with the help of the second interconnection link, when a conflict occurs between the "operation of sending the first key information and the first value information" and the "operation of All reduce between at least two first processors", the priority of All reduce between at least two first processors is higher than the priority of transmitting the first key information and the first value information to the second processor, that is, "All reduce between at least two first processors" is performed first, and then the "operation of sending the first key information and the first value information" is performed.

[0024] In this implementation manner, when there is no separate interconnection link between multiple first processors, the All reduce between multiple first processors needs to be implemented by means of the interconnection link between the first processor and the second processor. Since the transmission of the first key information and the first value information to the second processor is for use in the subsequent incremental inference process, setting a higher priority for the All reduce between multiple first processors is beneficial to executing the full-scale inference process as soon as possible, and then using the first key information and the first value information stored in the memory of the second processor in the subsequent incremental inference process. This design is beneficial to improving the efficiency of the overall process of processing the first text information using the first machine learning model.

[0025] In a possible implementation manner, the video memory of the first processor is also used to store the weight parameters and intermediate results used by the first processor in the process of performing full-scale inference on the first text information using the machine learning model. The determining factors for the value of L-X include at least one of the following factors: the video memory capacity of the first processor, the first data volume corresponding to the weight parameters, the second data volume corresponding to the intermediate results, or the third data volume corresponding to the first key information and the first value information.

[0026] Exemplarily, at least one first processor includes a total of T first processors, where T is an integer greater than or equal to 1. Since the operation of "performing full-scale inference on the first text information through the first machine learning model" is jointly executed by T first processors, the first data volume can be determined based on the total first data volume occupied by all the weight parameters in the first machine learning model and the value of T. The total first data volume refers to the data volume of all the weight parameters in the first machine learning model. For example, the first data volume can be greater than or equal to the ratio of the aforementioned total first data volume to T.

[0027] The second data volume refers to the data volume occupied by the intermediate results generated in a single first processor in the process of jointly executing the operation of "performing full-scale inference on the first text information through the first machine learning model" by T first processors. The third data volume represents the data volume of the key information and the value information generated by a first neural network module.

[0028] In this implementation, the video memory of the first processor is also used for the weight parameters used by the first processor and the intermediate results generated in the process of the first processor using the first machine learning model to perform full inference on the first text information, thereby avoiding the transmission of the weight parameters and the intermediate results between the memory of the second processor and the first processor, which is beneficial to improving the efficiency of the process of performing full inference through the first processor; in addition, it provides factors including the factors for determining LX, that is, it improves the factors for determining which of the L first neural network modules among the L first neural network modules generate key information and value information reside in the video memory of the first processor, and which of the first neural network modules generate key information and value information transmitted to the memory of the second processor, thereby improving the feasibility of this solution; and the factors for determining LX include the aforementioned weight parameters and the amount of data occupied by the intermediate results generated, thereby ensuring that in the process of processing the first text information using the first machine learning model, the video memory of the first processor is sufficient, which is beneficial to ensuring the smoothness of the process of processing the first text information using the first machine learning model.

[0029] In one possible implementation, the process of using a machine learning model to perform full reasoning on the first text information is parallel to the process of transmitting the first information. In this implementation, since the "process of using the first machine learning model to perform full reasoning on the first text information" and the "process of transmitting the first information" are parallel, that is, after obtaining the key information and value information generated by any one of the X first neural network modules, the transmission can be performed, and while performing the aforementioned transmission operation, other computing operations generated when the first machine learning model is used to perform full reasoning on the first text information can be performed in parallel, so as to minimize the additional time overhead caused by the transmission process of the first information.

[0030] Second aspect, an embodiment of the present application provides a method for processing text information, which can be used in the field of text processing in the field of artificial intelligence. The method is used in the process of processing first text information through a machine learning model. The machine learning model includes L first neural network modules. The first neural network module is a neural network module based on the attention mechanism, and L is an integer greater than or equal to 1. The method includes: A second processor obtains first information sent by a first processor. The first information is obtained during the process of performing full-scale inference on the first text information using the machine learning model. The execution entity of the calculation operations generated during the full-scale inference of the first text information using the machine learning model is the first processor. The first processor is a GPU or an NPU, and the second processor is a CPU; The second processor stores the first information in the memory. Among them, the first information includes the first key information and the first value information of the first text information. The first information is obtained through X first neural network modules among the L first neural network modules, and X is an integer greater than or equal to 1 and less than L; The second information is stored in the video memory of the first processor. The second information is obtained during the process of performing full-scale inference on the first text information using the machine learning model. The second information includes the second key information and the second value information of the first text information. The second information is obtained through L-X first neural network modules among the L first neural network modules. The L-X first neural network modules are the first neural network modules other than the X first neural network modules among the L first neural network modules.

[0031] In the second aspect of the present application, the second processor is further configured to execute the steps executed by the second processor in the first aspect and various possible implementation manners of the first aspect. For the specific implementation manners of the steps, the meanings of the terms, and the beneficial effects brought in the second aspect, reference can be made to the first aspect, which will not be elaborated here.

[0032] In a third aspect, an embodiment of the present application provides a text information processing device, which can be used in the field of text processing in the field of artificial intelligence. The text information processing device is used to process the first text information through a machine learning model. The machine learning model includes L first neural network modules. The first neural network module is a neural network module based on an attention mechanism. L is an integer greater than or equal to 1. The text information processing device is applied to a first processor. The device includes: a transmission module, configured to transmit first information to a second processor during the process of the first processor performing full-scale inference on the first text information using the machine learning model. The first information is stored in the memory of the second processor. The first processor is a graphics processing unit (GPU) or an embedded neural network processor (NPU), and the second processor is a central processing unit (CPU). The execution entity of the calculation operation generated when performing full-scale inference on the first text information using the machine learning model is the first processor; a storage module, configured to store second information in the video memory of the first processor; wherein the first information includes the first key information and the first value information of the first text information, and the first information is obtained through X first neural network modules among the L first neural network modules. X is an integer greater than or equal to 1 and less than L. The second information includes the second key information and the second value information of the first text information, and the second information is obtained through L - X first neural network modules among the L first neural network modules. The L - X first neural network modules are the first neural network modules other than the X first neural network modules among the L first neural network modules.

[0033] In the third aspect of the present application, the text information processing device is further configured to execute the steps performed by the first processor in the first aspect and various possible implementation manners of the first aspect. For the specific implementation manners of the steps, the meanings of the terms, and the beneficial effects brought in the third aspect, reference can be made to the first aspect, and details are not described herein again.

[0034] Fourthly, an embodiment of the present application provides a text information processing device, which can be used in the field of text processing in the field of artificial intelligence. The text information processing device is used for processing the first text information through a machine learning model. The machine learning model includes L first neural network modules, and the first neural network module is a neural network module based on the attention mechanism. L is an integer greater than or equal to 1. The device is applied to a second processor. The text information processing device includes: an acquisition module, configured to acquire the first information sent by the first processor. The first information is obtained during the process of performing a full-scale inference on the first text information using the machine learning model. The execution entity of the computing operation generated during the full-scale inference of the first text information using the machine learning model is the first processor, and the first processor is a GPU or an NPU, and the second processor is a CPU; a storage module, configured to store the first information in the memory. The first information includes the first key information and the first value information of the first text information, and the first information is obtained through X first neural network modules among the L first neural network modules, where X is an integer greater than or equal to 1 and less than L. The second information is stored in the video memory of the first processor and is obtained during the process of performing a full-scale inference on the first text information using the machine learning model. The second information includes the second key information and the second value information of the first text information, and the second information is obtained through L - X first neural network modules among the L first neural network modules, and the L - X first neural network modules are the first neural network modules other than the X first neural network modules among the L first neural network modules.

[0035] In the fourth aspect of the present application, the text information processing device is further configured to execute the steps executed by the second processor in the first aspect and various possible implementation manners of the first aspect. The specific implementation manners of the steps, the meanings of the terms, and the beneficial effects brought in the fourth aspect can all be referred to in the first aspect and will not be elaborated here.

[0036] Fifthly, an embodiment of the present application provides an execution device, including a processor and a memory. The processor is coupled to the memory. The memory is configured to store a program. The processor is configured to execute the program in the memory, so that the execution device executes the text information processing method described in the first aspect or the second aspect above.

[0037] Sixthly, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, it causes the computer to execute the method described in the first aspect or the second aspect above.

[0038] Seventh aspect, an embodiment of the present application provides a computer program product, which includes a program. When the program runs on a computer, it causes the computer to execute the method described in the first aspect or the second aspect above.

[0039] Eighth aspect, the present application provides a chip system, which includes a processor for supporting the implementation of the functions involved in the above aspects. For example, it sends or processes the data and / or information involved in the above method. In a possible design, the chip system further includes a memory for storing the necessary program instructions and data of the terminal device or the communication device. The chip system can be composed of chips or include chips and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic structural diagram of an artificial intelligence entity framework provided by an embodiment of the present application;

[0041] Figure 2 It is a system architecture diagram of a data processing system provided by an embodiment of the present application;

[0042] Figure 3 It is a schematic structural diagram of a first neural network module provided by an embodiment of the present application;

[0043] Figure 4 It is a schematic diagram of using the first neural network module to process the first text information provided by an embodiment of the present application;

[0044] Figure 5 It is a schematic flowchart of a method for processing text information provided by an embodiment of the present application;

[0045] Figure 6 It is a schematic diagram of generating and transmitting the execution order between the first key information and the first value information provided by an embodiment of the present application;

[0046] Figure 7 It is a schematic diagram of the sequence of "All reduce between at least two first processors" and "the first processor sending the first key information and the first value information of the first text information to the second processor" provided by an embodiment of the present application;

[0047] Figure 8 It is a schematic diagram of the data stored in the video memory of the first processor and the memory of the second processor provided by an embodiment of the present application;

[0048] Figure 9 It is a schematic diagram of the arithmetic operations performed by the first processor and the second processor during the incremental inference process provided by an embodiment of the present application;

[0049] Figure 10 A structural schematic diagram of a processing device for text information provided by an embodiment of the present application;

[0050] Figure 11 Another structural schematic diagram of a processing device for text information provided by an embodiment of the present application;

[0051] Figure 12 A structural schematic diagram of an execution device provided by an embodiment of the present application;

[0052] Figure 13 A structural schematic diagram of a chip provided by an embodiment of the present application. Detailed implementation manners

[0053] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will understand that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0054] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0055] First, the overall working process of the artificial intelligence system will be described. Please refer to Figure 1 , Figure 1 which shows a structural schematic diagram of an artificial intelligence main framework. The above artificial intelligence main framework will be elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes a refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.

[0056] (1) Infrastructure

[0057] The infrastructure provides computing power support for the AI system, enables communication with the external world, and is supported by the basic platform. It communicates with the external world through sensors; the computing power is provided by intelligent chips, which can specifically use hardware acceleration chips such as central processing unit (CPU), neural-network processing unit (NPU), graphics processing unit (GPU), application specific integrated circuit (ASIC), or field programmable gate array (FPGA), etc.; the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external world to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.

[0058] (2) Data

[0059] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.

[0060] (3) Data Processing

[0061] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0062] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on the data.

[0063] Reasoning refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formal information to perform machine thinking and solve problems according to the reasoning control strategy. The typical function is search and matching.

[0064] Decision-making refers to the process of making decisions after the intelligent information is reasoned, and usually provides functions such as classification, sorting, prediction, etc.

[0065] (4) General Capabilities

[0066] After the data is processed by the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system. For example, it can be translation, text analysis, computer vision processing, speech recognition, image recognition, and so on.

[0067] (5) Intelligent Products and Industry Applications

[0068] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields. It is the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes its implementation. Its application fields mainly include: intelligent terminals, intelligent manufacturing, intelligent transportation, smart homes, intelligent healthcare, intelligent security, autonomous driving, smart cities, etc.

[0069] The method provided by this application can be applied to various application fields of artificial intelligence technology. Specifically, it can be used to perform natural language processing tasks through a machine learning model (for convenience of description, hereinafter referred to as the "first machine learning model") in various application scenarios; optionally, the aforementioned natural language processing tasks can be performed by a large language model (LLM).

[0070] Before introducing multiple application scenarios of this application, the "natural language processing task" will be introduced first. Exemplarily, natural language processing is the processing of human language. Natural language processing is a process of systematically analyzing, understanding, and extracting information from text using the first machine learning model. Performing natural language processing tasks using machine learning models may exist in application fields such as intelligent terminals, smart homes, and autonomous driving.

[0071] In the above-mentioned various application fields, by using the aforementioned machine learning model, we can manage very large chunks of text information, or perform a large number of automated tasks, and solve various problems, such as automatic summarization, machine translation (MT), named entity recognition (NER), relation extraction (RE), information extraction (IE), sentiment analysis, speech recognition, question answering, and topic segmentation, etc.

[0072] Exemplarily, natural language processing tasks can be classified into the following categories.

[0073] Sequence labeling: For each word in the text, the machine learning model is required to give a classification category based on the context. Such as Chinese word segmentation, part-of-speech tagging, named entity recognition, semantic role labeling, etc.

[0074] Classification task: The machine learning model outputs a classification value for the entire input text. Such as sentiment classification, topic classification, or classification of whether the grammar is used correctly, etc.

[0075] Sentence relationship inference: The input of the machine learning model is two texts, and the machine learning model is used to judge whether these two texts have a certain nominal relationship. Such as question answering systems, semantic rewriting, or natural language inference, etc.

[0076] Generative task: Given a piece of text, another piece of text is generated through the machine learning model. Such as machine translation, automatic summarization, or writing poems and sentences, etc.

[0077] Information extraction task: At least one type of information is obtained from the input text through the machine learning model.

[0078] Before the method provided in this application is described in detail, please first refer to Figure 2 , Figure 2 which is a system architecture diagram of the data processing system provided by the embodiment of this application. In Figure 2 , the data processing system 200 includes a training device 210, a database 220, an execution device 230, a data storage system 240, and a client device 250. The execution device 230 includes a computing module 231.

[0079] Among them, the training data set is stored in the database 220. In the training stage of the first machine learning model 201, the training device 210 generates the first machine learning model 201 that has not yet performed training operations, and iteratively trains the foregoing first machine learning model 201 using the training data set to obtain the trained first machine learning model 201 (which can also be referred to as the "trained first machine learning model 201"). The first machine learning model 201 can be specifically embodied as a neural network or a non-neural network model. In the embodiments of this application, only the case where the first machine learning model 201 is embodied as a neural network is used for illustration.

[0080] The trained first machine learning model 201 obtained by the training device 210 can be deployed to the computing module 231 of the execution device 230. The execution device 230 can call data, code, etc. in the data storage system 240, and can also store data, instructions, etc. in the data storage system 240. The data storage system 240 can be placed in the execution device 230, or the data storage system 240 can be an external memory relative to the execution device 230.

[0081] In the application stage of the trained first machine learning model 201, the first text information can be processed by the trained first machine learning model 201. Exemplarily, in one case, please refer to Figure 2 , the execution device 230 and the client device 250 can be separate devices. The execution device 230 is configured with an input / output (I / O) interface to interact with the client device 250. After determining the first text information, the client device 250 sends the first text information to the execution device 230 through the I / O interface. After generating a processing result corresponding to the first text information through the first machine learning model 201 in the computing module 231, the execution device 230 can send the foregoing processing result to the client device 250 through the I / O interface. Exemplarily, the method for processing text information provided in this application can be applied to the application stage of the first machine learning model 201.

[0082] It should be noted that Figure 2 is only a schematic diagram of the architectures of two data processing systems provided by the embodiments of the present invention. The positional relationships among the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in another case, the execution device 230 and the client device 250 can also be integrated into the same device, and the user can directly interact with the execution device 230. Exemplarily, when the client device 250 is a mobile phone or a tablet, the execution device 230 can be a module in the main processor (Host CPU) of the mobile phone or tablet that uses the first machine learning model for data processing. The execution device 230 can also be a graphics processing unit (GPU) or a neural network processor (NPU) in the mobile phone or tablet. The GPU or NPU is mounted on the main processor as a coprocessor, and tasks are allocated by the main processor.

[0083] To understand this solution more conveniently, the meanings of multiple terms used in this application are explained as follows.

[0084] (1) First machine learning model

[0085] The first machine learning model can include L first neural network modules. Each of the L first neural network modules can be a neural network module based on the attention mechanism; optionally, the first machine learning model can also include other neural network modules, which are not limited in the embodiments of the present application. Optionally, the first machine learning model can be specifically embodied as a large language model.

[0086] Exemplarily, each first neural network module may be embodied as a transformer module, and each first neural network module may include a neural network layer based on an attention mechanism and a feed-forward network (FFN).

[0087] For a more intuitive understanding of the meaning of the "first neural network module", please refer to Figure 3 , Figure 3 which is a schematic structural diagram of the first neural network module provided in an embodiment of the present application. As Figure 3 shown, the first neural network module may include a neural network layer based on an attention mechanism and an FFN; the neural network layer based on an attention mechanism may include a Linear neural network layer and a matmul neural network layer; the FFN may include a Linear neural network layer.

[0088] Exemplarily, when processing input information through the neural network layer based on an attention mechanism, query information of the foregoing input information may be generated through Linear1, key information of the foregoing input information may be generated through Linear2, and value information of the foregoing input information may be generated through Linear3; exemplarily, the query information, key information, and value information may all be embodied as vectors. Matrix multiplication is performed on the query information and key information of the input information through matmul 1 to obtain Information 1, and matrix multiplication is performed on the value information of the input information and the foregoing Information 1 through matmul 2 to obtain Information 2. Linear calculation is performed on Information 2 through Linear4 to obtain a first processing result generated by the neural network layer based on an attention mechanism.

[0089] In the process of processing the foregoing first processing result through the FFN, two linear operations are respectively performed through Linear5 and Linear6 to obtain a second processing result generated by the entire first neural network module. It should be understood that Figure 3 the examples in

[0090] (2) Full-context (Context / Prefill) inference

[0091] The process of processing the first text information input by the user through the first machine learning model may include: the process of performing full-scale inference on the first text information using the first machine learning model. In the process of the aforementioned full-scale inference, the input of the first machine learning model may be all the words included in the first text information, that is, all the words input by the user are input into the first machine learning model as a whole, and the first machine learning model performs feature extraction and feature processing on the entire first text information to obtain the second text information generated by the first machine learning model; exemplarily, the second text information may include a word generated by the first machine learning model.

[0092] To further understand this solution, the following takes the natural language processing task performed by the first machine learning model as a "question and answer system", and the first text information is specifically "Where is XX?" as an example. Then the process of performing full-scale inference on the first text information using the first machine learning model may include feature extraction and feature processing on the whole of "Where is XX?". Here, taking the answer corresponding to the question "Where is XX?" as "ABCD" as an example, the second text information obtained after performing full-scale inference on the first text information using the first machine learning model may be the word "A". It should be understood that the example here is only for facilitating the understanding of this solution and is not used to limit this solution.

[0093] Exemplarily, since the first machine learning model includes L first neural network modules, in the process of performing full-scale inference on the first text information using the first machine learning model, the key information and value information of the first text information generated by each of the L first neural network modules can be obtained.

[0094] To further understand this solution, the following takes the natural language processing task performed by the first machine learning model as a "question and answer system", and the first text information is specifically "Where is XX?" as an example. Each first neural network module will generate the key information and value information of "Where is XX?". It should be understood that the example here is only for facilitating the understanding of this solution and is not used to limit this solution.

[0095] (3) Incremental (Decode) Inference

[0096] The process of processing the first text information input by the user through the first machine learning model may further include: performing at least one incremental inference operation using the first machine learning model. In each incremental inference operation, the output of the first machine learning model includes one word; repeating the incremental inference operation at least once until the word output by the first machine learning model is a specific termination word, or the number of times of performing incremental inference reaches a preset number, thereby terminating the execution of the incremental inference operation. Combining the second text information obtained during full inference and the words generated during each incremental inference, the processing result corresponding to the first text information output by the first machine learning model is obtained.

[0097] Exemplarily, during the first incremental inference process, the input of the first machine learning model includes the above-mentioned second text information. The first machine learning model can generate third text information, and the third text information can be represented as one word; the third text information obtained during the first incremental inference process can be used as the input for the next incremental inference. Repeat the foregoing steps until the output word is a specific termination word, or the number of times of performing incremental inference reaches a preset number.

[0098] To further understand this solution, the following takes the natural language processing task performed by the first machine learning model as a "question answering system", the first text information is specifically expressed as "Where is XX?", and the answer corresponding to the question "Where is XX?" is "ABCD" as an example. The first machine learning model can perform 3 incremental inferences. When the first machine learning model performs the first inference, the input of the first machine learning model includes "A", and the output of the first machine learning model can be "B"; when the first machine learning model performs the second inference, the input of the first machine learning model includes "B", and the output of the first machine learning model can be "C"; when the first machine learning model performs the third inference, the input of the first machine learning model includes "C", and the output of the first machine learning model can be "D"; then the processing result corresponding to the first text information output by the first machine learning model is "ABCD". It should be understood that the example here is only for facilitating the understanding of the relationship between "full inference" and "incremental inference", and is not used to limit this solution.

[0099] Optionally, when the first machine learning model performs the first incremental inference for the first time, the input of the first machine learning model further includes the key information and value information of the first text information generated during the full inference process; that is, during the first incremental inference process of the first machine learning model, after generating the key information and value information of the second text information through the first neural network module, combining the key information and value information of the first text information, the third text information is jointly generated.

[0100] To further understand this solution, the following combinationFigure 3 Describe how to use the "key information and value information of the first text information" during the incremental process. When performing incremental inference for the first time, the query information of the second text information can be generated through Linear1 in the first neural network module; the key information of the second text information can be generated through Linear2 in the first neural network module. Concatenate the key information of the first text information and the key information of the second text information to obtain the key information of (the first text information + the second text information); generate the value information of the second text information through Linear3 in the first neural network module, and concatenate the value information of the first text information and the value information of the second text information to obtain the value information of (the first text information + the second text information).

[0101] Perform matrix multiplication on the query information of the second text information and the key information of (the first text information + the second text information) through matmul 1 to obtain Information 3; perform matrix multiplication on Information 3 and the value information of (the first text information + the second text information) through matmul 2 to obtain Information 4; perform linear calculation on Information 4 through Linear4 to obtain the processing result generated by the neural network layer based on the attention mechanism in the first neural network module. Since the subsequent calculation operations performed through FFN are similar to those performed through FFN during the full-scale inference process, they will not be elaborated in the embodiments of this application.

[0102] When performing incremental inference for the second time through the first machine learning model, the input of the first machine learning model also includes the key information and value information of the first text information generated during the full-scale inference process and the key information and value information of the second text information generated during the first incremental inference; that is, during the second incremental inference process of the first machine learning model, after generating the key information and value information of the third text information through the first neural network module, combine the key information and value information of the first text information, the key information and value information of the second text information to jointly generate a word corresponding to the second incremental inference.

[0103] To further understand this solution, the following is combined with Figure 3Describe how to use the "key information and value information of the first text information and the third text information" in the incremental process. When performing incremental inference for the second time, the query information of the third text information can be generated by Linear1 in the first neural network module; the key information of the third text information can be generated by Linear2 in the first neural network module. Concatenate the key information of the first text information, the key information of the second text information, and the key information of the third text information to obtain the key information of (the first text information + the second text information + the third text information); the value information of the third text information can be generated by Linear3 in the first neural network module. Concatenate the value information of the first text information, the value information of the second text information, and the value information of the third text information to obtain the value information of (the first text information + the second text information + the third text information).

[0104] Perform matrix multiplication on the query information of the third text information and the key information of (the first text information + the second text information + the third text information) through matmul 1 to obtain Information 5; perform matrix multiplication on Information 5 and the value information of (the first text information + the second text information + the third text information) through matmul 2 to obtain Information 6; perform linear calculation on Information 5 through Linear4 to obtain the processing result generated by the neural network layer based on the attention mechanism in the first neural network module. Since the subsequent calculation operations performed by FFN are similar to those performed by FFN in the full-scale inference process, they will not be elaborated in the embodiments of this application.

[0105] And so on. That is, when performing incremental inference for the Nth time through the first machine learning model, the input of the first machine learning model also includes the key information and value information of the first text information generated in the full-scale inference process and the key information and value information generated in the previous N - 1 incremental inference processes, where N is an integer greater than or equal to 1.

[0106] (4) Synchronous summation (All reduce) between at least two first processors

[0107] The execution subject of the method provided in this application is a text information processing system. Exemplarily, Figure 2The execution device 230 therein may be specifically embodied as a processing system for text information. The aforementioned processing system for text information includes at least one first processor and a second processor. The first processor is a GPU or an NPU, and the second processor is a CPU. During the full-scale inference of the first text information through the first machine learning model, the execution entity of the computing operations corresponding to the aforementioned full-scale inference is the at least one first processor; exemplarily, the computing operations corresponding to the aforementioned full-scale inference may include the computing operations generated when processing the first text information using the neural network layer in the first machine learning model. When the at least one first processor includes at least two first processors, the aforementioned full-scale inference process further includes an All reduce between the at least two first processors, and the purpose of the aforementioned All reduce operation is to implement the at least two first processors.

[0108] To more intuitively understand the meanings of "computing operation" and "All reduce", the process of processing the first text information through the first neural network module is described herein to determine which operations are computing operations and which operations are All reduce in the aforementioned process. Please refer to Figure 4 , Figure 4 which is a schematic diagram provided by an embodiment of the present application for processing the first text information using the first neural network module. Figure 4 It can be understood in combination with the above Figure 3 . As shown in Figure 4 , the operations when processing the first text information through Linear1, Linear2, Linear3, matmul 1, matmul 2, Linear4, Linear5, and Linear6 are all arithmetic operations. Since the aforementioned arithmetic operations will execute the arithmetic operations of Linear1, Linear2, Linear3, matmul 1, matmul 2, and Linear4 on at least two first processors, after each of the at least two first processors has executed the arithmetic operations of Linear1, Linear2, Linear3, matmul 1, matmul 2, and Linear4, an All reduce will be performed between the at least two first processors. After executing the aforementioned All reduce, each of the at least two first processors executes the arithmetic operations of Linear5 and Linear6, and then an All reduce is performed again by the at least two first processors. It should be understood that Figure 4 the examples in

[0109] Combined with the above description, the method provided in the embodiments of the present application will be introduced below. The method for processing text information provided by the present application can be applied to the application stage of the first machine learning model. Specifically, please refer to Figure 5 , Figure 5 which is a schematic flowchart of a method for processing text information provided in an embodiment of the present application. The method for processing text information provided in an embodiment of the present application may include:

[0110] 501. During the process of the first processor performing full-scale inference on the first text information using the first machine learning model, the first processor obtains first information, where the first information includes the first key information of the first text information and the first value information of the first text information. The first key information and the first value information of the first text information are generated by X of the L first neural network modules. L is an integer greater than or equal to 1, and X is an integer greater than or equal to 1 and less than L. The first processor is a GPU or an NPU, and the second processor is a CPU.

[0111] In the embodiments of the present application, each of at least one first processor, after obtaining the first text information, may perform full-scale inference on the first text information using the first machine learning model. Among them, the execution subject of the computing operation generated when performing full-scale inference on the first text information using the first machine learning model is the first processor. Exemplarily, since the first machine learning model includes L first neural network modules, during the process of each first processor performing full-scale inference on the first text information through the first machine learning model, a key information and a value information of the first text information can be generated by each of the L first neural network modules. Then, each first processor can obtain the first key information and the first value information of the first text information generated by X of the aforementioned L first neural network modules (that is, the first information is obtained). It should be noted that the first key information of the first text information is also the key information of the first text information, and the first value information of the first text information is also the value information of the first text information. In the present application, the first key information specifically refers to the key information generated by X of the L first neural network modules, and the first value information specifically refers to the value information generated by X of the L first neural network modules.

[0112] Exemplarily, each of the above key information and each value information may be specifically represented as a tensor. The concept of "generating the key information and value information of the first text information through the first neural network module" can be understood in combination with the above description, and will not be further described in detail in combination with the structure of the first neural network module here.

[0113] Both L and X are integers greater than or equal to 1; optionally, the above X first neural network modules are the first X first neural network modules among the L first neural network modules. Exemplarily, if the value of L is greater than X, then in the process of performing full-scale inference on the first text information by the first machine learning model, the first text information will be processed by these X first neural network modules first, and then the first text information will be processed by the remaining L - X first neural network modules.

[0114] 502. The first processor transmits the first information to the second processor.

[0115] In the embodiments of the present application, in the process of each first processor performing full-scale inference on the first text information using the first machine learning model, each first processor can also transmit the first information including the first key information of the first text information and the first value information of the first text information to the second processor, that is, send the first key information and the first value information of the first text information to the second processor; it should be noted that the "transmission process of the first key information and the first value information of the first text information (that is, the transmission process of the first information)" and the "process of performing full-scale inference on the first text information using the first machine learning model" are parallel, that is, they do not interfere with each other.

[0116] Exemplarily, after each first processor obtains the first key information and the first value information of the first text information generated by any one of the X first neural network modules (for the convenience of distinction, any one of the X first neural network modules will be referred to as the "target neural network module" hereinafter), it can execute the operation of "sending the first key information and the first value information of the first text information to the second processor"; since the value of X is an integer greater than or equal to 1, when the value of X is an integer greater than 1, the embodiments of the present application do not limit the execution order between steps 501 and 502.

[0117] To more intuitively understand the relationship between the execution order of "generating the first key information and the first value information of the first text information" and "transmitting the first key information and the first value information of the first text information", please refer to Figure 6 , Figure 6 which is a schematic diagram of the execution order between generating and transmitting the first key information and the first value information provided by the embodiments of the present application. Figure 6Taking the first X neural network modules out of the L first neural network modules as an example, the X first neural network modules include a plurality of first neural network modules, and each first neural network module includes a neural network layer based on an attention mechanism (i.e., Figure 6 the attention in Figure 6 ), and the "transmission of KV information" in Figure 6 represents the "transmission operation of the first key information and the first value information". Among them, the "computing operation corresponding to each first neural network module" and the "transmission operation of the first key information and the first value information" are executed in parallel, that is, the "computing operation corresponding to each first neural network module" and the "transmission operation of the first key information and the first value information" do not interfere with each other. After generating a first key information and a first value information of the first text information through the "neural network layer based on the attention mechanism" in each of the X first neural network modules, the "transmission operation of the first key information and the first value information" can be started. At the same time, the first processor continues to execute the computing operations corresponding to other neural network layers in the first neural network module. It should be understood that Figure 6 the examples in

[0118] are only for facilitating the understanding of this solution and are not used to limit this solution.

[0119] Optionally, the method provided in this application is applied to a text information processing system. If the text information processing system includes at least two first processors and a second processor, then each of the at least two first processors in the process of performing a full-scale inference on the first text information using the first machine learning model not only includes the computing operations generated when processing the first text information using the neural network layer in the first machine learning model, but also includes the All reduce between the at least two first processors. For the concept of "Allreduce between at least two first processors generated during the full-scale inference process", reference can be made to the above description and will not be elaborated here.

[0120] For the specific implementation of "All reduce between at least two first processors", in one case, there is a separate first interconnection link between different first processors among at least two first processors, that is, different first processors can directly communicate through the aforementioned separate first interconnection link; then different first processors can execute "All reduce between at least two first processors" through the aforementioned first interconnection link. Exemplarily, the separate interconnection link between different first processors can be an Nvlink link.

[0121] It should be noted that since there will be a second interconnection link between the first processor and the second processor, exemplarily, the second interconnection link can be a Pci-e link, and the "first key information and first value information of the first text information" are transmitted through the second interconnection link between the first processor and the second processor; if there is a separate first interconnection link between different first processors, then "All reduce between at least two first processors" and "the first processor sends the first key information and first value information of the first text information to the second processor" do not interfere with each other.

[0122] To understand this solution more intuitively, please refer to Figure 7 , Figure 7 which is a schematic diagram of the sequence of "All reduce between at least two first processors" and "the first processor sends the first key information and first value information of the first text information to the second processor" provided by the embodiments of this application. Figure 7 Taking the first X first neural network modules among the L first neural network modules as an example, the first X first neural network modules include a plurality of first neural network modules, and each first neural network module includes a neural network layer based on the attention mechanism (that is, Figure 6 the attention in Figure 7As shown, since the priority of "All reduce between at least two first processors" is higher than the priority of "the first processor sends the first key information and the first value information of the first text information to the second processor", when there is a conflict between "All reduce between at least two first processors" and "the first processor sends the first key information and the first value information of the first text information to the second processor", the execution of "the first processor sends the first key information and the first value information of the first text information to the second processor" will be suspended, and "All reduce between at least two first processors" will be executed in a limited manner. It should be understood that Figure 7 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0123] In another case, if there is no separate first interconnection link between different first processors in at least two first processors, then All reduce between at least two first processors needs to be completed with the help of the second processor, that is, All reduce between at least two first processors needs to be completed with the help of the second interconnection link between the first processor and the second processor; illustratively, each first processor in at least two first processors sends the data to be synchronously summed to the second processor through the second interconnection link, the second processor aggregates the data to be synchronously summed sent by all the first processors to obtain aggregated data, and the second processor then sends the aggregated data to each first processor through the second interconnection link, thereby completing All reduce between at least two first processors.

[0124] If the operation of "the first processor sending the first key information and the first value information of the first text information to the second processor" is to be performed through the second interconnection link, and the operation of "All reduce between at least two first processors" is to be performed with the help of the second interconnection link, a conflict may occur between the "operation of sending the first key information and the first value information" and the "All reduce between at least two first processors"; since the priority of All reduce between at least two first processors is higher than the priority of transmitting the first key information and the first value information to the second processor, when a conflict occurs between the "operation of sending the first key information and the first value information" and the "All reduce between at least two first processors", the "All reduce between at least two first processors" is performed first, and then the "operation of sending the first key information and the first value information" is performed.

[0125] In the embodiments of the present application, when there is no separate interconnection link between multiple first processors, the All reduce between multiple first processors needs to be implemented by means of the interconnection link between the first processor and the second processor. Since the transmission of the first key information and the first value information to the second processor is for use in the subsequent incremental inference process, setting a higher priority for the All reduce between multiple first processors is beneficial to quickly execute the full-scale inference process, and then use the first key information and the first value information stored in the memory of the second processor in the subsequent incremental inference process. This design is beneficial to improving the efficiency of the overall process of processing the first text information using the first machine learning model.

[0126] 503. The second processor stores the first information in the memory.

[0127] In the embodiments of the present application, after the second processor receives the first key information and the first value information (i.e., the first information) of the first text information sent by each first processor, the second processor may store the first key information and the first value information of the first text information in the memory of the second processor.

[0128] 504. The first processor stores the second information in the video memory of the first processor. The second information includes the second key information of the first text information and the second value information of the first text information. The second key information and the second value information are generated by L - X first neural network modules among the L first neural network modules. The L - X first neural network modules are the first neural network modules other than the X first neural network modules among the L first neural network modules.

[0129] In the embodiments of the present application, during the full-scale inference process of using the first machine learning model to process the first text information by each first processor in at least one first processor, the second key information and the second value information of the first text information generated by the remaining L - X first neural network modules among the L first neural network modules can also be obtained. The second processor may store the second key information of the first text information and the second value information of the first text information in the video memory of the first processor.

[0130] Among them, the above-mentioned X first neural network modules and the above-mentioned L-X first neural network modules are different neural network modules among the L first neural network modules. Optionally, the aforementioned L-X first neural network modules are the last L-X first neural network modules among the L first neural network modules. That is, in the process of performing full-scale inference on the first text information by the first machine learning model, the first text information will be processed by the first X first neural network modules first, and then processed by the last L-X first neural network modules.

[0131] In the embodiment of the present application, since the "sending operation of the first key information and the first value information" and the "computing operation corresponding to the first neural network module" do not interfere with each other, when the first processor sends the first key information and the first value information generated by the Xth first neural network module to the second processor, the first processor can start executing the computing operations corresponding to the last L-X first neural network modules, which is beneficial to improving the processing speed of the entire first machine learning model for the first text information. In addition, if the X first neural network modules are selected as the first X first neural network modules among the L first neural network modules, the storage space occupied by the aforementioned first key information and first value information in the video memory can be released in time. Then, when the first processor processes data through the last L-X first neural network modules, it can use the video memory with a larger storage space, which is beneficial to improving the performance when the first processor performs full-scale inference on the first text information.

[0132] It should be understood that the "X first neural network modules" can also be the last X first neural network modules among the L first neural network modules. Correspondingly, the "L-X first neural network modules" are the first L-X first neural network modules among the L first neural network modules. Or, the "X first neural network modules" can also be X non-consecutive first neural network modules among the L first neural network modules. That is, the positions of the X first neural network modules and the L-X first neural network modules in the first machine learning model can be crossed, etc. Specifically, the "X first neural network modules" and the "L-X first neural network modules" can be flexibly determined according to the actual situation and are not limited here.

[0133] Exemplarily, the second key information of the first text information is the key information of the first text information, and the second value information of the first text information is the value information of the first text information. The second key information specifically refers to the key information generated by the above-mentioned L-X first neural network modules, and the second value information specifically refers to the value information generated by the above-mentioned L-X first neural network modules.

[0134] Exemplarily, before each first processor executes step 501, it can also determine which of the L first neural network modules (i.e., X first neural network modules) need to send the generated key information and value information to the second processor, and which first neural network modules (i.e., L-X first neural network modules) store the generated key information and value information in the video memory of the first processor. When the X first neural network modules are the first X first neural network modules among the L first neural network modules, and the L-X first neural network modules are the last L-X first neural network modules among the L first neural network modules, each first processor also needs to obtain the segmentation positions of the L first neural network modules.

[0135] Optionally, the video memory of each first processor can also be used to store the weight parameters and intermediate results generated during the full-scale inference of the first text information using the first machine learning model. For a more intuitive understanding of this solution, please refer to Figure 8 , Figure 8 which is a schematic diagram of the data stored in the video memory of the first processor and the memory of the second processor provided in the embodiments of the present application. Figure 8 Taking the text information processing system including 4 first processors and 1 second processor as an example, as Figure 8 shown, the video memory of each first processor can be used to store the weight parameters used during the full-scale inference of the first text information using the first machine learning model, the intermediate results generated during the full-scale inference of the first text information using the first machine learning model, and the second key information and second value information generated by the L-X first neural network modules among the L first neural network modules included in the first machine learning model; the memory of the second processor can be used to store the first key information and first value information generated by the X first neural network modules among the L first neural network modules included in the first machine learning model. It should be understood that Figure 8 the example in

[0136] Optionally, the determining factors of L-X (i.e., the determining factors of the segmentation positions of "L first neural network modules" and "L-X first neural network modules") can include at least one of the following factors: the video memory capacity of the first processor, the first data volume corresponding to the weight parameters of the first machine learning model, the second data volume corresponding to the above intermediate results, the third data volume corresponding to each first key information and each first value information, or other factors, which are not limited here.

[0137] Exemplarily, the at least one first processor includes a total of T first processors, where T is an integer greater than or equal to 1. Since the operation of "performing full-scale inference on the first text information through the first machine learning model" is jointly executed by the T first processors, the first data volume can be determined based on the first total data volume occupied by all weight parameters in the first machine learning model and the value of T. For example, the first data volume can be greater than or equal to the ratio of the aforementioned first total data volume to T.

[0138] Exemplarily, the second data volume refers to the data volume occupied by the intermediate results generated in a single first processor during the process of jointly executing, by the T first processors, the operation of "performing full-scale inference on the first text information through the first machine learning model".

[0139] Since multiple different text information can be processed through the first machine learning model and the lengths of different text information may vary, the data volume of the intermediate results generated during the process of performing full-scale inference on the first text information using the first machine learning model may also vary, that is, the data volume of the intermediate results corresponding to different lengths of text information may be different. Exemplarily, in one case, the second data volume is preset, and the preset second data volume corresponds to text information of a preset length, and the text information of the preset length can be the longest text information that the first machine learning model can process in a single time. In another case, the second data volume is actually calculated according to the length of the first text information processed by the first machine learning model each time.

[0140] The third data volume represents the data volume of the key information and value information generated by one first neural network module, and the data volumes of the key information and value information corresponding to different lengths of text information are also different. Exemplarily, in one case, the third data volume is preset, and the preset third data volume corresponds to text information of a preset length, that is, the third data volume is the data volume occupied by the key information and value information of the first text information obtained through one first neural network module when the input first text information is of the preset length. In another case, the third data volume can be actually calculated according to the length of the first text information processed by the first machine learning model each time.

[0141] To further understand this solution, an example of the calculation formula for the splitting positions of L first neural network modules is disclosed as follows:

[0142] L - X = (the video memory capacity of the first processor - the first total data volume / T - the second data volume) / the third data volume

[0143] After determining the value of L-X through the above formula, the value of X can be determined according to the value of L. It should be noted that the above formula is only an example for facilitating the understanding of this solution and is not used to limit this solution. In other implementation manners, the value of "L-X" can also be a preset value. For example, "L-X" can take one-tenth, one-fifth, etc. of L; or, the video memory of the first processor may not be used to store the weight parameters and / or intermediate results of the first machine learning model, then the value of "L-X" can be increased, and it can be flexibly determined specifically in combination with the actual application scenario, which is not limited in the embodiments of the present application.

[0144] In the embodiments of the present application, the video memory of the first processor is further used for the weight parameters and the generated intermediate results used by the first processor in the process of performing full-scale inference on the first text information using the first machine learning model, avoiding the transmission of the weight parameters and the intermediate results between the memory of the second processor and the first processor, which is beneficial to improving the efficiency of the process of performing full-scale inference through the first processor; in addition, the factors for determining L-X are provided, that is, it is improved according to which factors to determine which of the L first neural network modules generate the key information and value information that stay in the video memory of the first processor, and which of the first neural network modules generate the key information and value information that are transmitted to the memory of the second processor, improving the feasibility of this solution; and the factors for determining L-X include the data volume occupied by the foregoing weight parameters and the generated intermediate results, thereby ensuring that the video memory of the first processor is sufficient during the process of processing the first text information using the first machine learning model, which is beneficial to ensuring the smoothness of the process of processing the first text information using the first machine learning model.

[0145] Optionally, the embodiments of the present application further provide a specific implementation process of performing at least one incremental inference through the first machine learning model. In one implementation, each incremental inference performed through the first machine learning model can be executed by the second processor. In another implementation, each incremental inference performed through the first machine learning model can be executed by T first processors. In another implementation, the process of each incremental inference can be executed collaboratively by the first processor and the second processor. Exemplarily, key information and value information can also be generated during the process of each incremental inference performed through the first machine learning model. The storage locations of the key information and value information generated during the process of performing incremental inference through the first machine learning model are determined based on the execution entity that generates the key information and value information. For example, if the key information and value information are generated by the first processor, the key information and value information can be stored in the video memory of the first processor. If the key information and value information are generated by the second processor, the key information and value information can be stored in the memory of the second processor. Alternatively, the key information and value information generated during each incremental inference process can also be stored in the memory of the second processor, etc., which can be specifically determined in combination with the actual application scenario and is not limited in the embodiments of the present application.

[0146] For example, during the process of performing each incremental inference using the first machine learning model, the execution entity of the computing operations corresponding to X first neural network modules includes the second processor, and the execution entity of the computing operations corresponding to L - X first neural network modules is the first processor.

[0147] Exemplarily, during the process of performing each incremental inference using the first machine learning model, all the computing operations corresponding to L - X first neural network modules can be executed by T first processors. Optionally, during the process of performing each incremental inference through the first machine learning model, the key information and value information generated by the L - X first neural network modules can be stored in the video memory of the first processor. Alternatively, the first processor can also transmit the key information and value information generated by the L - X first neural network modules to the second processor and then store them in the memory of the second processor.

[0148] Exemplarily, in one case, the first neural network module includes a neural network layer based on an attention mechanism and a feed-forward neural network layer FFN. In the process of performing each incremental inference using the first machine learning model, the execution entity of the computational operations corresponding to the neural network layers based on the attention mechanism in the X first neural network modules is the second processor, and the execution entity of the computational operations corresponding to the FFN in the X first neural network modules is the first processor. Optionally, in the process of performing each incremental inference using the first machine learning model, the key information and value information generated by the X first neural network modules can be stored in the memory of the second processor.

[0149] Further, in one implementation, please refer to Figure 9 , Figure 9 which is a schematic diagram of the arithmetic operations performed by the first processor and the second processor during incremental inference provided by the embodiments of the present application. Figure 9 It needs to be understood in combination with the above Figure 3 . As shown in Figure 9 , when performing the Nth incremental inference, the computational operations corresponding to Linear1, Linear2, and Linear3 in the X first neural network modules can be executed by the first processor. After generating the query information, key information, and value information of the input information for the Nth incremental inference, the first processor can send the query information, key information, and value information of the input information to the second processor. The second processor can execute the computational operations corresponding to matmul 1 and matmul 2 based on the first key information and first value information of the first text information, the key information and value information generated in the previous N - 1 incremental inferences (such as the key information and value information of the second text information, the key information and value information of the third text information, etc.), and the query information, key information, and value information of the aforementioned input information. The computational operations corresponding to Linear4 in the X first neural network modules can be executed by the first processor. It should be understood that Figure 9 the examples in

[0150] Exemplarily, during each incremental inference process, when the second processor performs incremental inference through each of the X first neural network modules, it can obtain intermediate result 1 after completing the computational operations corresponding to Linear1, Linear2, Linear3, matmul 1, and matmul 2, and send intermediate result 1 to the first processor. The first processor then continues to execute the computational operation corresponding to Linear4 and the computational operation corresponding to FFN; thereby implementing all the computational operations of one of the X first neural network modules.

[0151] In another implementation, all the computational operations corresponding to the neural network layers based on the attention mechanism in the X first neural network modules can be executed by the second processor. Exemplarily, during each incremental inference process, when the second processor performs incremental inference through each of the X first neural network modules, the second processor can obtain intermediate result 2 after completing all the computational operations corresponding to the neural network layers based on the attention mechanism, and send intermediate result 2 to the first processor. The first processor then continues to execute the computational operation corresponding to FFN; thereby implementing all the computational operations of one of the X first neural network modules.

[0152] In the embodiments of the present application, since during the incremental inference process, the key information and value information of the first text information are used in the computational operations corresponding to the neural network layers based on the attention mechanism in the first neural network modules, and the key information and value information of the first text information are not required in the computational operations corresponding to the FFN in the first neural network modules, the execution entity of the computational operations corresponding to the neural network layers based on the attention mechanism in the X first neural network modules includes the second processor, and the execution entity of the computational operations corresponding to the FFN in the X first neural network modules is the first processor. This not only helps to avoid the time consumption caused by transferring the "first key information and first value information of the first text information" from the memory of the second processor to the first processor, but also makes full use of the computing power of the first processor, which is beneficial to further improving the efficiency of processing the first text information using the first machine learning model.

[0153] In another case, during the process of performing each incremental inference using the first machine learning model, all the computational operations corresponding to the X first neural network modules can be executed by the second processor; then during each incremental inference process, the second processor can send the processing results generated by the aforementioned X first neural network modules to the first processor, and the first processor is used to execute all the computational operations corresponding to the L - X first neural network modules.

[0154] In the embodiments of the present application, during the incremental inference process, when processing through the foregoing X first neural network modules, the first key information and the first value information of the first text information are used. The data volume of the first key information and the first value information of the first text information is large, and the first key information and the first value information of the first text information are stored in the memory of the second processor. If the execution entity of the computing operation corresponding to the X first neural network modules includes the second processor, it is beneficial to avoid the time consumption brought by the process of transferring the "first key information and the first value information of the first text information" from the memory of the second processor to the first processor, and is beneficial to improving the efficiency of the process of processing the first text information using the first machine learning model.

[0155] The first processor is used to execute the computing operation of performing full-scale inference on the first text information using the machine learning model. However, during the foregoing full-scale inference process, the key information and the value information generated by some neural network modules in the machine learning model are transferred to the CPU, that is, a part of the key information and the value information of the first text information generated during the foregoing full-scale inference process are transferred to the CPU. A part of the key information and the value information of the first text information exist in the memory of the CPU, and another part of the key information and the value information of the first text information are stored in the video memory of the first processor. Thus, the large-capacity memory of the CPU can be used to assist in storing a part of the key information and the value information of the first text information, that is, not only the advantage of the strong computing power of the first processor is fully utilized, but also the advantage of the large storage space of the CPU is utilized, avoiding the waste of the computing power of the first processor.

[0156] To further understand the beneficial effects brought by the method provided in the present application, the following is described in combination with experimental data. In the experiment, a text information processing system including 8 GPUs (i.e., the first processor) and a 52-core CPU (i.e., the second processor) is taken as an example. The video memory capacity of each first processor is 32GB, and the memory capacity of the second processor is 800GB. The original solution can support processing 8 text information with a length of 3.75K each time, while the method provided in the present application can support processing 8 text information with a length of 35K each time, greatly improving the ability of the entire text information processing system.

[0157] In addition, when using the same long text information for testing, the time taken by the original solution during the full-scale inference process is 27602 + 25000 ms, and the time taken for a single incremental inference is 68.1 + 25000 ms; the time taken by the method provided in this application during the full-scale inference process is 25000 ms, and the time taken for a single incremental inference is 70.2 ms; that is, when processing long text information through this application, the processing speed of long text information will also be greatly improved.

[0158] Based on Figures 1 to 9 the corresponding embodiments, in order to better implement the above solutions of the embodiments of this application, the following also provides related devices for implementing the above solutions. Specifically, refer to Figure 10 , Figure 10 FIG.

[0159] is a schematic structural diagram of a text information processing device provided in an embodiment of this application. The text information processing device 1000 is used to process the first text information through a machine learning model. The machine learning model includes L first neural network modules. The first neural network module is a neural network module based on the attention mechanism, and L is an integer greater than or equal to 1. The text information processing device 1000 is applied to the first processor. The text information processing device 1000 includes: a transmission module 1001, configured to transmit the first information to the second processor during the full-scale inference of the first text information by the first processor using the machine learning model. The first information is stored in the memory of the second processor. The first processor is a graphics processing unit GPU or an embedded neural network processor NPU, and the second processor is a central processing unit CPU. The execution subject of the calculation operation generated when performing full-scale inference on the first text information using the machine learning model is the first processor; a storage module 1002, configured to store the second information in the video memory of the first processor;

[0160] Optionally, the X first neural network modules are the first X first neural network modules among the L first neural network modules, and the L - X first neural network modules are the last L - X first neural network modules among the L first neural network modules.

[0161] Optionally, the processing device 1000 for text information further includes: a processing module 1003, configured to perform an incremental inference process using a machine learning model. During the incremental inference process using the machine learning model, the execution entities of the computational operations generated by X first neural network modules in the machine learning model include a second processor, and the execution entities of the computational operations generated by L-X first neural network modules in the machine learning model are the first processor.

[0162] Optionally, the first neural network module is a Transformer module. The first neural network module includes a neural network layer based on an attention mechanism and a feed-forward neural network layer FFN. During the incremental inference process using the machine learning model, the execution entities of the computational operations generated by the neural network layer based on the attention mechanism in X first neural network modules include a second processor, and the execution entities of the computational operations generated by FFN in X first neural network modules are the first processor.

[0163] Optionally, the processing device 1000 for text information is applied to a text information processing system. The system includes at least two first processors and one second processor. The process of the first processor performing a full-scale inference on the first text information using the machine learning model includes an All Reduce for synchronization summation between at least two first processors. In the case where there is no separate interconnection link between at least two first processors, the All Reduce between at least two first processors is completed with the help of the second processor. Among them, the priority of the All Reduce between at least two first processors is higher than the priority of transmitting the first information to the second processor.

[0164] Optionally, the video memory of the first processor is further used to store the weight parameters and intermediate results used by the first processor during the full-scale inference process of the first text information using the machine learning model. The determining factors for the value of L-X include at least one of the following factors: the video memory capacity of the first processor, the data volume of the weight parameters, the data volume of the intermediate results, or the data volume of the first information.

[0165] Optionally, the process of performing a full-scale inference on the first text information using the machine learning model is parallel to the transmission process of the first information.

[0166] It should be noted that the information interaction, execution process, etc. between the modules / units in the processing device 1000 for text information are based on the same concept as the corresponding method embodiments in this application. Figures 3 to 9 For specific content, reference can be made to the descriptions in the method embodiments shown above in this application, and details are not described herein again.

[0167] Please continue to refer to Figure 11 , Figure 11Another structural schematic diagram of the text information processing device provided by the embodiment of the present application. The text information processing device 1100 is used to process the first text information through a machine learning model. The machine learning model includes L first neural network modules. The first neural network module is a neural network module based on the attention mechanism. L is an integer greater than or equal to 1. The device is applied to the second processor. The text information processing device 1100 includes:

[0168] An acquisition module 1101, configured to acquire the first information sent by the first processor. The first information is obtained during the process of performing full-scale inference on the first text information using the machine learning model. The execution entity of the calculation operation generated during the full-scale inference of the first text information using the machine learning model is the first processor. The first processor is a GPU or an NPU, and the second processor is a CPU. A storage module 1102, configured to store the first information in the memory. The first information includes the first key (key) information and the first value (value) information of the first text information. The first information is obtained through X first neural network modules among the L first neural network modules. X is an integer greater than or equal to 1 and less than L.

[0169] The second information is stored in the video memory of the first processor. The second information is obtained during the process of performing full-scale inference on the first text information using the machine learning model. The second information includes the second key information and the second value information of the first text information. The second information is obtained through L - X first neural network modules among the L first neural network modules. The L - X first neural network modules are the first neural network modules other than the X first neural network modules among the L first neural network modules.

[0170] Optionally, the X first neural network modules are the first X first neural network modules among the L first neural network modules, and the L - X first neural network modules are the last L - X first neural network modules among the L first neural network modules.

[0171] It should be noted that the information interaction, execution process, etc. between the modules / units in the text information processing device 1100 are based on the same concept as the corresponding method embodiments in the present application. For specific content, reference can be made to the description in the method embodiments shown above in the present application, and details will not be repeated here. Figures 3 to 9 Next, an execution device provided by the embodiment of the present application will be introduced. Please refer to

[0172] Figure 12 Figure 12 ​​A schematic structural diagram of an execution device provided by an embodiment of the present application. Specifically, the execution device 1200 may include: a receiver 1201, a transmitter 1202, a processor 1203, and a memory 1204 (where the number of processors 1203 in the execution device 1200 may be one or more, Figure 12 and here one processor is taken as an example), where the processor 1203 may include an application processor 12031 and a communication processor 12032. In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203, and the memory 1204 may be connected through a bus or other means.

[0173] The memory 1204 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1203. A part of the memory 1204 may also include a non-volatile random access memory (NVRAM). The memory 1204 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, where the operation instructions may include various operation instructions for implementing various operations.

[0174] The processor 1203 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, where the bus system may include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clarity, all kinds of buses are referred to as the bus system in the figure.

[0175] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 1203. The above-mentioned processor 1203 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1203 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1204, and the processor 1203 reads the information in the memory 1204 and combines its hardware to complete the steps of the above method.

[0176] The receiver 1201 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 1202 can be used to output digital or character information through the first interface; the transmitter 1202 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1202 can also include a display device such as a display screen.

[0177] In the embodiments of the present application, the processor 1203 is used to execute Figures 3 to 9 the text information processing method executed by the first processor in the corresponding embodiment, or the processor 1203 is used to execute Figures 3 to 9 the text information processing method executed by the second processor in the corresponding embodiment. It should be noted that the specific manner in which the processor 1203 executes the foregoing various steps is based on the same concept as the corresponding method embodiments in the present application, and the technical effects brought by it are the same as those of the corresponding method embodiments in the present application. For specific content, reference can be made to the description in the foregoing method embodiments shown in the present application, and details are not described herein again. Figures 3 to 9 The corresponding method embodiments in the present application are the same, and for specific content, reference can be made to the description in the foregoing method embodiments shown in the present application, and details are not described herein again. Figures 3 to 9 The corresponding method embodiments in the present application are the same, and for specific content, reference can be made to the description in the foregoing method embodiments shown in the present application, and details are not described herein again.

[0178] In an embodiment of the present application, a computer-readable storage medium is further provided. A program for signal processing is stored in the computer-readable storage medium. When it runs on a computer, the computer is caused to execute the steps performed by the first processor in the method described in the foregoing Figures 3 to 9 illustrated embodiment, or the computer is caused to execute the steps performed by the second processor in the method described in the foregoing Figures 3 to 9 illustrated embodiment.

[0179] In an embodiment of the present application, a computer program product is further provided. When it runs on a computer, the computer is caused to execute the steps performed by the first processor in the method described in the foregoing Figures 3 to 9 illustrated embodiment, or the computer is caused to execute the steps performed by the second processor in the method described in the foregoing Figures 3 to 9 illustrated embodiment.

[0180] The first processor and the second processor provided in the embodiments of the present application may specifically be chips. The chips include: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, or the like. The processing unit may execute computer execution instructions stored in the storage unit to cause the chip to execute the above Figures 3 to 9 illustrated embodiment of the processing method of text information. Optionally, the storage unit is a storage unit inside the chip, such as a register, a cache, or the like. The storage unit may also be a storage unit outside the chip and inside the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), or the like.

[0181] Specifically, please refer to Figure 13 , Figure 13 which is a schematic structural diagram of the chip provided in the embodiment of the present application. The chip may be embodied as a neural network processor NPU 130. The NPU 130 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 1303. The arithmetic circuit 1303 is controlled by the controller 1304 to extract matrix data from the memory and perform multiplication operations.

[0182] In some implementations, the arithmetic circuit 1303 includes multiple processing units (Process Engine, PE) internally. In some implementations, the arithmetic circuit 1303 is a two-dimensional systolic array. The arithmetic circuit 1303 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1303 is a general matrix processor.

[0183] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1302 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1301 and performs matrix operations with matrix B, and the partial results of the obtained matrix or the processing results corresponding to the first text information are stored in the accumulator 1308.

[0184] The unified memory 1306 is used to store input data and output data. The weight data is directly transferred through the Direct Memory Access Controller (DMAC) 1305 and is transported to the weight memory 1302. The input data is also transported to the unified memory 1306 through the DMAC.

[0185] The BIU is the Bus Interface Unit, that is, the bus interface unit 1310, which is used for the interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 1309.

[0186] The bus interface unit 1310 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch memory 1309 to obtain instructions from the external memory, and is also used for the storage unit access controller 1305 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0187] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1306, or transfer the weight data to the weight memory 1302, or transfer the input data to the input memory 1301.

[0188] The vector calculation unit 1307 includes multiple arithmetic processing units, and in case of need, further processes the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization (batch normalization), pixel-level summation, upsampling of the feature plane, etc.

[0189] In some implementations, the vector computing unit 1307 can store the processed output vectors into the unified memory 1306. For example, the vector computing unit 1307 can apply a linear function and / or a non-linear function to the output of the arithmetic circuit 1303, such as performing linear interpolation on the feature planes extracted by the convolutional layer, or for another example, vectors of accumulated values, to generate activation values. In some implementations, the vector computing unit 1307 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vectors can be used as activation inputs to the arithmetic circuit 1303, such as for use in subsequent layers in the neural network.

[0190] The instruction fetch buffer 1309 connected to the controller 1304 is used to store the instructions used by the controller 1304;

[0191] The unified memory 1306, the input memory 1301, the weight memory 1302, and the instruction fetch memory 1309 are all On-Chip memories. The external memory is private to this NPU hardware architecture.

[0192] Among them, the operations of each layer in the above first machine learning model can be executed by the arithmetic circuit 1303 or the vector computing unit 1307.

[0193] Among them, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the method in the above first aspect.

[0194] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits, etc. However, for this application, in more cases, software program implementation is a better embodiment. Based on such understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, a first processor, or a network device, etc.) to execute the methods described in various embodiments of this application.

[0196] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0197] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a first processor or a data center to another website, a computer, a first processor or a data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a first processor, a data center, etc. that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. A method for processing text information, characterized in that, The method is used in the process of processing first text information by a machine learning model. The machine learning model includes L first neural network modules, and the first neural network modules are neural network modules based on an attention mechanism. L is an integer greater than or equal to 1. The method includes: During the process of the first processor performing full-scale inference on the first text information using the machine learning model, the first processor transmits first information to the second processor, and the first information is stored in the memory of the second processor. The first processor is a graphics processing unit (GPU) or an embedded neural network processor (NPU), and the second processor is a central processing unit (CPU). The execution entity of the computing operations generated when performing full-scale inference on the first text information using the machine learning model is the first processor; The first processor stores second information in the video memory of the first processor; Among them, the first information includes the first key information and the first value information of the first text information, and the first information is obtained through X of the L first neural network modules. X is an integer greater than or equal to 1 and less than L. The second information includes the second key information and the second value information of the first text information, and the second information is obtained through L - X of the L first neural network modules. The L - X first neural network modules are the first neural network modules other than the X first neural network modules among the L first neural network modules.

2. The method according to claim 1, wherein The X first neural network modules are the first X first neural network modules among the L first neural network modules, and the L - X first neural network modules are the last L - X first neural network modules among the L first neural network modules.

3. The method according to claim 1 or 2, characterized in that, The process of processing the first text information by the machine learning model further includes: the process of performing incremental inference using the machine learning model. Among them, during the process of performing incremental inference using the machine learning model, the execution entity of the computing operations generated by the X first neural network modules in the machine learning model includes the second processor, and the execution entity of the computing operations generated by the L - X first neural network modules in the machine learning model is the first processor.

4. The method according to claim 3, wherein The first neural network module is a Transformer module. The first neural network module includes a neural network layer based on an attention mechanism and a feed-forward neural network layer (FFN). During the process of performing incremental inference using the machine learning model, the execution entity of the computing operations generated by the neural network layer based on the attention mechanism in the X first neural network modules includes the second processor, and the execution entity of the computing operations generated by the FFN in the X first neural network modules is the first processor.

5. The method according to claim 1 or 2, characterized in that, The method is applied to a text information processing system, which includes at least two of the first processors and one second processor. The process of the first processors performing full-scale inference on the first text information using the machine learning model includes synchronous summation (All reduce) among the at least two first processors. In the case where there is no separate interconnection link between the at least two first processors, the All reduce among the at least two first processors is completed with the help of the second processor. Among them, the priority of the All reduce among the at least two first processors is higher than the priority of transmitting the first information to the second processor.

6. The method according to claim 1 or 2, characterized in that, The video memory of the first processor is also used to store the weight parameters and intermediate results used by the first processor during the process of performing full-scale inference on the first text information using the machine learning model. The determining factors of the value of L-X include at least one of the following factors: the video memory capacity of the first processor, the data volume of the weight parameters, the data volume of the intermediate results, or the data volume of the first information.

7. The method according to claim 1 or 2, characterized in that, The process of performing full-scale inference on the first text information using the machine learning model is parallel to the transmission process of the first information.

8. A method for processing text information, characterized in that, The method is used in the process of processing the first text information through a machine learning model. The machine learning model includes L first neural network modules, and the first neural network module is a neural network module based on the attention mechanism. L is an integer greater than or equal to 1. The method includes: The second processor obtains the first information sent by the first processor. The first information is obtained during the process of performing full-scale inference on the first text information using the machine learning model. The execution entity of the calculation operations generated when performing full-scale inference on the first text information using the machine learning model is the first processor. The first processor is a GPU or an NPU, and the second processor is a CPU. The second processor stores the first information in the memory. Among them, the first information includes the first key (key) information and the first value (value) information of the first text information. The first information is obtained through X of the L first neural network modules. X is an integer greater than or equal to 1 and less than L. The second information is stored in the video memory of the first processor. The second information is obtained during the process of performing full-scale inference on the first text information using the machine learning model. The second information includes the second key information and the second value information of the first text information. The second information is obtained through L-X of the L first neural network modules. The L-X first neural network modules are the first neural network modules other than the X first neural network modules among the L first neural network modules.

9. The method according to claim 8, characterized in that The X first neural network modules are the first X first neural network modules among the L first neural network modules, and the L-X first neural network modules are the last L-X first neural network modules among the L first neural network modules.

10. A processing device for text information, characterized in that, In the process of the device using a machine learning model to process first text information, the machine learning model includes L first neural network modules. The first neural network module is a neural network module based on an attention mechanism. L is an integer greater than or equal to 1. The device is applied to a first processor. The device includes: A transmission module, configured to transmit first information to a second processor during the process of the first processor performing full-scale inference on the first text information using the machine learning model. The first information is stored in the memory of the second processor. The first processor is a graphics processing unit (GPU) or an embedded neural network processor (NPU), and the second processor is a central processing unit (CPU). The execution entity of the calculation operation generated when performing full-scale inference on the first text information using the machine learning model is the first processor. A storage module, configured to store second information in the video memory of the first processor. Wherein, the first information includes the first key information and the first value information of the first text information. The first information is obtained through X first neural network modules among the L first neural network modules. X is an integer greater than or equal to 1 and less than L. The second information includes the second key information and the second value information of the first text information. The second information is obtained through L-X first neural network modules among the L first neural network modules. The L-X first neural network modules are the first neural network modules among the L first neural network modules except the X first neural network modules.

11. The device according to claim 10, characterized in that, The X first neural network modules are the first X first neural network modules among the L first neural network modules, and the L-X first neural network modules are the last L-X first neural network modules among the L first neural network modules.

12. The device according to claim 10 or 11, characterized in that, The text information processing device further includes: a processing module, configured to perform incremental inference using the machine learning model. During the process of performing incremental inference using the machine learning model, the execution entity of the calculation operation generated by the X first neural network modules in the machine learning model includes the second processor, and the execution entity of the calculation operation generated by the L-X first neural network modules in the machine learning model is the first processor.

13. The device according to claim 12, characterized in that, The first neural network module is a Transformer module, and the first neural network module includes a neural network layer based on an attention mechanism and a feed-forward neural network layer FFN. During the process of incremental inference using the machine learning model, the execution entity of the computational operations generated by the neural network layer based on the attention mechanism in the X first neural network modules includes the second processor, and the execution entity of the computational operations generated by the FFN in the X first neural network modules is the first processor.

14. The device according to claim 10 or 11, characterized in that, The device is applied to a text information processing system, and the system includes at least two first processors and one second processor. The process of the at least two first processors performing full-scale inference on the first text information using the machine learning model includes synchronous summation (All reduce) between the at least two first processors. In the case where there is no separate interconnection link between the at least two first processors, the All reduce between the at least two first processors is completed with the help of the second processor. Among them, the priority of the All reduce between the at least two first processors is higher than the priority of transmitting the first information to the second processor.

15. The device according to claim 10 or 11, characterized in that, The video memory of the first processor is also used to store the weight parameters and intermediate results generated during the process of the first processor performing full-scale inference on the first text information using the machine learning model. The determining factors of the value of L - X include at least one of the following factors: the video memory capacity of the first processor, the data volume of the weight parameters, the data volume of the intermediate results, or the data volume of the first information.

16. The device according to claim 10 or 11, characterized in that, The process of performing full-scale inference on the first text information using the machine learning model is parallel to the transmission process of the first information.

17. A processing device for text information, characterized in that, When the device is used to process the first text information through a machine learning model, the machine learning model includes L first neural network modules. The first neural network module is a neural network module based on an attention mechanism, and L is an integer greater than or equal to 1. The device is applied to a second processor, and the device includes: An acquisition module, configured to acquire the first information sent by the first processor. The first information is obtained during the process of the first processor performing full-scale inference on the first text information using the machine learning model. The execution entity of the computational operations generated when the machine learning model performs full-scale inference on the first text information is the first processor. The first processor is a GPU or an NPU, and the second processor is a CPU. A storage module, configured to store the first information in the memory. Among them, the first information includes the first key information and the first value information of the first text information. The first information is obtained through X first neural network modules among the L first neural network modules, and X is an integer greater than or equal to 1 and less than L. The second information is stored in the video memory of the first processor, and is obtained during the process of performing full-scale inference on the first text information using the machine learning model. The second information includes the second key information and the second value information of the first text information, and the second information is obtained through L-X first neural network modules among the L first neural network modules. The L-X first neural network modules are the first neural network modules other than the X first neural network modules among the L first neural network modules.

18. The device according to claim 17, characterized in that, The X first neural network modules are the first X first neural network modules among the L first neural network modules, and the L-X first neural network modules are the last L-X first neural network modules among the L first neural network modules.

19. An execution device, characterized in that, comprising a processor and a memory, the processor being coupled to the memory, the memory for storing a program; the processor for executing the program in the memory, such that the execution device executes the method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that, A program is stored in the computer-readable storage medium, and when the program runs on a computer, the computer is caused to execute the method according to any one of claims 1 to 9.

21. A computer program product, characterized in that, The computer program product includes a program, and when the program runs on a computer, the computer is caused to execute the method according to any one of claims 1 to 9.