Data processing method and device
By reusing the network layer of the large model for data verification, the problem of low parallelism in the incremental inference stage of large language models is solved, which reduces the computing and memory overhead, and reduces the deployment cost of small models, achieving efficient inference acceleration.
Patent Information
- Application Number
- CN202311524978.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
The existing large language models have low parallelism in the incremental inference stage, resulting in low hardware utilization, and the speculative sampling method requires additional training and deployment of small models, which increases cost and complexity.
Data verification by multiplexing the network layer of the large model reduces additional inference deployment overhead, and memory can reuse the previous inference results during incremental verification, avoiding additional delay and computation overhead.
It improves the parallelism of large models in the incremental inference stage, reduces the overhead of computing and memory, reduces the deployment cost of small models, and ensures a high acceleration ratio of large model inference.
Smart Images

Figure CN120012906A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a data processing method and device thereof. Background Art
[0002] Inference for large language models is primarily divided into two phases. The first is the full prefill inference phase, where all user-entered words are fed into the network for computation, obtaining the output of the final word. This phase allows for parallel inference of multiple words, effectively utilizing hardware computing and memory bandwidth. The second is the incremental decoding inference phase, where the output of the full inference phase serves as the input for the current step and is fed into the network for computation. The output of the current word is then used as the input for the next step, continuing until decoding is complete. This phase is an autoregressive process that requires decoding word by word, resulting in low parallelism and hardware utilization. Repeated memory accesses for weights waste memory bandwidth and do not fully utilize the compute units. However, if a large model can decode multiple words in parallel in a single step, latency will be minimal compared to decoding a single word, improving hardware utilization.
[0003] To improve the parallelism of large model increments, the concept of predictive execution, also known as speculative sampling, is applied to large model inference. This approach ensures the same sampling distribution as the original model. It uses two models: the original target model and a much smaller approximate model. The approximate model performs serial autoregressive decoding, while the large model verifies the results. The large model can verify multiple tokens simultaneously, significantly reducing the number of autoregressive cycles and improving the compute-to-memory ratio of large model inference, thereby enhancing inference performance.
[0004] Speculative reasoning can accelerate the reasoning of large models without loss, but it requires additional training and deployment of speculative small models. Small models also require maintenance and deployment costs. When the matching degree between large and small models is low, speculative reasoning may even lead to performance regression and increased reasoning costs. Therefore, how to reduce the deployment cost of small models while ensuring a high speedup ratio for large model reasoning is a key issue in speculative reasoning. Summary of the Invention
[0005] In a first aspect, the present application provides a data processing method, the method comprising: generating second data through a first network based on first data; generating third data through a second network based on the first data; wherein the first network includes a first network layer, the second network includes a second network layer, and the first network layer is a partial network of the second network layer; and verifying the second data based on the third data.
[0006] This application reuses the network (second network) for the data verification process to generate the data network (part of it, that is, the first network), which can reduce the additional inference deployment overhead, and the second network can reuse the inference results of the first network in the memory during verification, without the need for additional delay, thereby reducing the computational overhead.
[0007] Among them, the embodiments of the present application can be applied to speculative reasoning. Specifically, it can be an incremental reasoning process of speculative reasoning. In the incremental reasoning process, the model (first network) is required to perform autoregressive reasoning, and the model is required to verify the results obtained by the autoregressive reasoning. The autoregressive reasoning can be the above-mentioned step of "obtaining the second data through the first network according to the first data". The result obtained by the autoregressive reasoning is verified for the above-mentioned step of "generating the third data through the second network according to the first data; wherein the first network includes a first network layer, the second network includes a second network layer, and the first network layer is a partial network of the second network layer; and the second data is verified according to the third data". There can be multiple verifications in incremental reasoning, and the data that finally passes the verification can be used as output.
[0008] The first data may be a generated token, that is, a data unit, such as a text unit (e.g., a character unit, a word unit, a punctuation mark, etc.), an image unit (e.g., an image block), or an audio unit. The second data may be a data unit located adjacent to the first data in the autoregressive inference result.
[0009] In a possible implementation, before generating the second data based on the first data through the first network, the method further includes: generating the first data based on fourth data through the first network.
[0010] In one possible implementation, the method further includes: generating fifth data through the second network based on the fourth data; wherein the action of generating the fifth data and the action of generating the third data are performed in parallel; and verifying the first data based on the fifth data.
[0011] In one possible implementation, the method is applied to an incremental reasoning process, and the first network layer and the second network layer both belong to a target neural network, which is a model used in a full reasoning process corresponding to the incremental reasoning process.
[0012] Among them, the full reasoning process is the reasoning process before the incremental reasoning process. When performing speculative reasoning, the input data is first fully reasoned through the large model, and then the incremental reasoning is performed using the results of the full reasoning. The above-mentioned full reasoning and incremental reasoning with a sequential dependency relationship can be considered to have a corresponding relationship.
[0013] Among them, full reasoning (also known as context reasoning, or prefill reasoning) means: inputting the input data as a whole into the network (such as the target neural network in the embodiment of the present application) for calculation, and obtaining one of the data units as the output of the full reasoning.
[0014] Among them, incremental decoding reasoning means: using the output of full reasoning as the input of the current step, inputting it into the network (for example, including the first network, the second network, and the third network in the embodiment of the present application) for calculation; the data unit output of each step can be used as the input of the next step, and looping until the output data unit is a specific termination data unit, or the number of decoding steps reaches the specified maximum value.
[0015] On the one hand, in the embodiment of the present application, the network (second network) performing the incremental verification process during speculative reasoning is reused with the network performing autoregressive reasoning (part of it, that is, the first network), which can reduce additional reasoning deployment overhead, and the second network can reuse the reasoning results of the first network in the memory when performing incremental verification, without the need for additional delay, thereby reducing computational overhead. On the other hand, the first network layer in the first network and the second network layer in the second network can both belong to the large model of full reasoning, so there is no need to maintain and train separate small models, saving costs.
[0016] In a possible implementation, the second network layer includes the first network layer and a third network layer connected to the first network layer, and the third network layer is closer to the output layer of the target neural network than the first network layer.
[0017] In a possible implementation, the first network includes a first network layer and a first decoding layer, and the second network includes a second network layer and a second decoding layer; generating the second data through the first network based on the first data includes: obtaining a first processing result through the first network layer based on the first data, and obtaining the second data through the first decoding layer based on the first processing result; generating the third data through the second network based on the first data includes: obtaining a second processing result through the third network layer based on the first processing result, and generating the third data through the second decoding layer based on the second processing result.
[0018] Since the reasoning branches of the large model are reused, the reasoning results of the sub-model can be directly reused without additional delay. The large model reasoning can cover the overhead of the small model.
[0019] In a possible implementation, the verifying the second data based on the third data includes: when the third data and the second data are consistent, determining that the verification of the second data has passed; when the third data and the second data are inconsistent, determining that the verification of the second data has failed.
[0020] In one possible implementation, the method further includes: when it is determined that the verification of the second data is passed, generating sixth data based on the first data through a third network; wherein the third network includes a fourth network layer, and the second network layer is a partial network of the fourth network layer; and verifying the second data based on the sixth data.
[0021] In one possible implementation, the method further includes: when it is determined that the verification of the second data fails, generating sixth data based on the first data through a third network; wherein the third network includes a fourth network layer, and the second network layer is a partial network of the fourth network layer; and verifying the third data based on the sixth data.
[0022] In a possible implementation, the method further includes: when it is determined that the verification of the second data fails, generating seventh data through the second network according to the third data.
[0023] In one possible implementation, the method is applied to an incremental reasoning process, and when it is determined that the verification of the second data fails based on the fourth data, seventh data is generated based on the third data through the second network, including: when the historical verification pass rate in the incremental reasoning process meets a preset condition (for example, greater than a threshold), and when it is determined that the verification of the second data fails, seventh data is generated based on the third data through the second network.
[0024] In one possible implementation, the method further includes: when it is determined that the verification of the second data fails, generating eighth data based on the third data through the first network; generating seventh data based on the third data through the second network; and verifying the eighth data based on the seventh data.
[0025] In one possible implementation, when it is determined that the verification of the second data fails, eighth data is generated according to the third data through the first network, including: during the incremental reasoning process, when the historical verification pass rate does not meet the preset conditions (for example, less than a threshold value), and when it is determined that the verification of the second data fails, eighth data is generated according to the third data through the first network.
[0026] Among them, the generation of the seventh data through the autoregressive reasoning of the second network can be called Scheme 1, and the generation of the seventh data through the autoregressive reasoning of the first network can be called Scheme 2. The advantage of the above-mentioned Scheme 1 is that the amount of generated data can be controlled, but speculative acceleration is not applied. Too many autoregressive times will slow down the overall speed; the advantage of Scheme 2 is that it can accelerate the generation of multiple autoregressives, but it will cause regression when the matching degree is not good. In actual applications, the two schemes can be combined according to the matching degree.
[0027] In one possible implementation, the method is applied to an incremental reasoning process, which is a response generation process for input data; generating second data based on the first data through the first network includes: determining sub-attention information corresponding to the length of the generated data based on the attention information in the full reasoning process corresponding to the incremental reasoning process and the length of the data generated in the response generation process; and generating second data through the first network based on the sub-attention information and the first data.
[0028] The attention information can be an attention mask. In this embodiment of the present application, a decoded_length parameter is added to the decoded sequence length of each batch during incremental inference. The attention_mask in each batch during each inference is dynamically updated based on the decoded_length and verification length n to ensure the correctness of the attention calculation. The attention_mask diagram is as follows. If the length of n in each batch is not equal, the attention calculation logic needs to be modified to adapt. This solution can reuse the original multi-batch full-scale inference logic, without redundant calculations and unnecessary memory overhead, and can significantly improve the throughput compared to single-batch inference.
[0029] In one possible implementation, the method is executed on a first node, and a full reasoning process corresponding to the incremental reasoning process is executed on a second node; before generating the second data based on the first data through the first network, the method also includes: receiving a processing result obtained through the full reasoning process and sent by the second node; the processing result is used to perform the incremental reasoning process.
[0030] In one possible implementation, the first data is data generated during the incremental reasoning process for a first batch, the method is applied to a first node, and when executing the method, the first node also performs the incremental reasoning process for a second batch in parallel. The method also includes: when the incremental reasoning process for the first batch is completed and the first node has not yet completed the incremental reasoning process for the second batch, clearing the data in the memory space corresponding to the incremental reasoning process for the first batch, and starting the incremental reasoning process for the third batch.
[0031] In a possible implementation, the method further includes: when it is determined that the verification of the second data fails and the number of data that has passed the verification of the second network in the current verification cycle is less than a preset value, performing autoregressive data generation through the second network until the number of data that has passed the verification of the second network in the current verification cycle for the first batch reaches the preset value; or, when it is determined that the verification of the second data fails, performing autoregressive data generation through the first network and verifying the data generated by the autoregression of the first network through the second network until the number of data that has passed the verification of the second network in the current verification cycle reaches the preset value.
[0032] If stage = 1, combined with dynamic batching, the overall throughput can be further improved. Because the data matching degree in each batch varies, the number of decoding steps is also different. If the decoding-completed sentences are exited early, hardware utilization can be improved. In addition, it is necessary to ensure that valid tokens are swapped out during each dynamic batch scheduling, and the KV cache at the corresponding location is cleared and the corresponding decode_len is updated. If stage > 1, multi-level caches must be set to ensure the matching degree of multi-level speculative reasoning. All caches at the corresponding batch location must be cleared during dynamic batch scheduling.
[0033] In a possible implementation, the first data is a token, the second data is a token adjacent to the first data, and the third data is a token adjacent to the second data.
[0034] In a possible implementation, the first data, the second data, and the third data are text data units, image data units, or audio data units.
[0035] In a possible implementation, the first data is a data unit obtained through a full-scale reasoning process or a data unit generated by performing autoregressive reasoning on the data unit obtained through the full-scale reasoning process through the first network.
[0036] In addition, the present application also provides a data processing method, which includes:
[0037] Get input data;
[0038] Performing full inference on the target neural network based on the input data to obtain a full inference result;
[0039] Based on the full inference result, incremental inference is used to obtain a response result corresponding to the input data; the incremental inference is implemented through a first network and a second network, the first network is used for autoregressive inference, and the second network is used for verification inference, the first network includes a first network layer, the second network includes a second network layer, and the first network layer is a partial network of the second network layer.
[0040] In a second aspect, the present application provides a data processing device, comprising:
[0041] A processing module is used to generate second data based on the first data through the first network; generate third data based on the first data through the second network; wherein the first network includes a first network layer, the second network includes a second network layer, and the first network layer is a partial network of the second network layer; and verify the second data based on the third data.
[0042] In a possible implementation, before generating the second data based on the first data through the first network, the processing module is further configured to:
[0043] The first data is generated through the first network according to the fourth data.
[0044] In a possible implementation, the processing module is further configured to:
[0045] generating fifth data through the second network according to the fourth data; wherein the action of generating the fifth data and the action of generating the third data are performed in parallel;
[0046] The first data is verified according to the fifth data.
[0047] In one possible implementation, applied to an incremental reasoning process, the first network layer and the second network layer both belong to a target neural network, and the target neural network is a model used in a full reasoning process corresponding to the incremental reasoning process.
[0048] In a possible implementation, the second network layer includes the first network layer and a third network layer connected to the first network layer, and the third network layer is closer to the output layer of the target neural network than the first network layer.
[0049] In one possible implementation, the first network includes a first network layer and a first decoding layer, and the second network includes a second network layer and a second decoding layer;
[0050] The processing module is specifically used to:
[0051] Obtaining a first processing result through the first network layer according to the first data, and obtaining second data through the first decoding layer according to the first processing result;
[0052] According to the first processing result, a second processing result is obtained through the third network layer, and according to the second processing result, third data is generated through the second decoding layer.
[0053] In a possible implementation, the processing module is specifically configured to:
[0054] When the third data and the second data are consistent, determining that the verification of the second data is passed;
[0055] When the third data and the second data are inconsistent, it is determined that the verification of the second data fails.
[0056] In a possible implementation, the processing module is further configured to:
[0057] When it is determined that the verification of the second data is passed, generating sixth data based on the first data through a third network; wherein the third network includes a fourth network layer, and the second network layer is a partial network of the fourth network layer;
[0058] The second data is verified according to the sixth data.
[0059] In a possible implementation, the processing module is further configured to:
[0060] When it is determined that the verification of the second data fails, generating sixth data based on the first data through a third network; wherein the third network includes a fourth network layer, and the second network layer is a partial network of the fourth network layer;
[0061] The third data is verified according to the sixth data.
[0062] In a possible implementation, the processing module is further configured to:
[0063] When it is determined that the verification of the second data fails, seventh data is generated based on the third data through the second network.
[0064] In one possible implementation, applied to an incremental reasoning process, the processing module is specifically configured to:
[0065] When the historical verification pass rate in the incremental reasoning process meets the preset condition and it is determined that the verification of the second data fails, seventh data is generated according to the third data through the second network.
[0066] In a possible implementation, the processing module is further configured to:
[0067] When it is determined that the verification of the second data fails, generating eighth data based on the third data through the first network;
[0068] generating seventh data through the second network according to the third data;
[0069] The eighth data is verified based on the seventh data.
[0070] In a possible implementation, the processing module is specifically configured to:
[0071] When the historical verification pass rate does not meet the preset conditions during the incremental reasoning process and it is determined that the verification of the second data fails, eighth data is generated based on the third data through the first network.
[0072] In one possible implementation, the method is applied to an incremental reasoning process, wherein the incremental reasoning process is a response generation process for input data; the processing module is specifically configured to:
[0073] Determining sub-attention information corresponding to the length of the generated data according to the attention information in the full inference process corresponding to the incremental inference process and the length of the generated data in the reply generation process;
[0074] Generate second data through the first network based on the sub-attention information and the first data.
[0075] In a possible implementation, the apparatus is executed on a first node, and a full reasoning process corresponding to the incremental reasoning process is executed on a second node;
[0076] Before generating the second data through the first network based on the first data, the processing module is further configured to:
[0077] Receive the processing result obtained through the full reasoning process sent by the second node; the processing result is used to perform the incremental reasoning process.
[0078] In one possible implementation, the first data is data generated during an incremental reasoning process for a first batch, the apparatus is applied to a first node, and when executing the apparatus, the first node also performs an incremental reasoning process for a second batch in parallel, and the processing module is further configured to:
[0079] When the incremental reasoning process for the first batch is completed and the first node has not yet completed the incremental reasoning process for the second batch, the data in the memory space corresponding to the incremental reasoning process of the first batch is cleared, and the incremental reasoning process of the third batch is started.
[0080] In one possible implementation, the processing module is further configured to: when it is determined that the second data fails verification and the number of data that passes verification in the second network in the current verification cycle is less than a preset value, perform autoregressive data generation through the second network until the number of data that passes verification in the second network in the current verification cycle for the first batch reaches the preset value; or
[0081] When it is determined that the verification of the second data fails, autoregressive data generation is performed through the first network, and the data generated by the autoregressive of the first network is verified through the second network until the amount of data that passes the verification of the second network in the current verification cycle reaches the preset value.
[0082] In a possible implementation, the first data is a token, the second data is a token adjacent to the first data, and the third data is a token adjacent to the second data.
[0083] In a possible implementation, the first data, the second data, and the third data are text data units, image data units, or audio data units.
[0084] In a possible implementation, the first data is a data unit obtained through a full-scale reasoning process or a data unit generated by performing autoregressive reasoning on the data unit obtained through the full-scale reasoning process through the first network.
[0085] In a third aspect, an embodiment of the present application provides a data processing device, which may include a memory, a processor, and a bus system, wherein the memory is used to store programs, and the processor is used to execute the programs in the memory to perform the first aspect and any optional method thereof.
[0086] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned first aspect and any optional method thereof.
[0087] In a fifth aspect, an embodiment of the present application provides a computer program, which, when executed on a computer, enables the computer to execute the above-mentioned first aspect and any optional method thereof.
[0088] In a sixth aspect, the present application provides a chip system comprising a processor for supporting an execution device or a training device in implementing the functions described in the aforementioned aspects, such as transmitting or processing data or information described in the aforementioned methods. In one possible design, the chip system further comprises a memory for storing program instructions and data necessary for the execution device or the training device. The chip system may consist of a single chip or may include a chip and other discrete components. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1A A structural diagram of the main framework of artificial intelligence;
[0090] Figure 1B Hezhi Figure 1C This is a schematic diagram of the application system framework of the present invention;
[0091] Figure 1D Schematic diagram of an optional hardware structure of the terminal;
[0092] Figure 2 A schematic diagram of the structure of a server;
[0093] Figures 3 to 5 This is a schematic diagram of the system architecture of this application;
[0094] Figure 6 A process for a cloud service;
[0095] Figure 7 A flowchart of a data processing method provided in an embodiment of the present application;
[0096] Figures 8 to 12 This is a schematic diagram of the model processing process;
[0097] Figure 13 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;
[0098] Figure 14 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application;
[0099] Figure 15A schematic diagram of the structure of a server provided in an embodiment of the present application;
[0100] Figure 16 A schematic diagram of the structure of the chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0101] The following describes the embodiments of the present invention in conjunction with the accompanying drawings. The terms used in the embodiments of the present invention are only used to explain the specific embodiments of the present invention, and are not intended to limit the present invention.
[0102] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0103] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0104] As used herein, the terms "substantially," "about," and similar terms are used as terms of approximation, not as terms of degree, and are intended to take into account the inherent variations in measurements or calculations that one of ordinary skill in the art would recognize. Furthermore, the use of "may" when describing embodiments of the present invention refers to "one or more possible embodiments." As used herein, the terms "use," "using," and "used" may be considered synonymous with the terms "utilize," "utilizing," and "utilized," respectively. Additionally, the term "exemplary" is intended to refer to an example or illustration.
[0105] First, the overall workflow of the artificial intelligence system is described. Figure 1A , Figure 1AThe following diagram illustrates a structural diagram of the AI framework. This framework is explained below from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it encompasses the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed progression from "data-information-knowledge-wisdom." The "IT value chain," encompassing the entire process from the underlying infrastructure of human intelligence, information (provided and processed by technology), to the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.
[0106] (1) Infrastructure
[0107] Infrastructure provides computing power for AI systems, enabling communication with the outside world and supporting this through a foundational platform. External communication occurs through sensors; computing power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, FPGAs, and other hardware accelerators). The foundational platform includes a distributed computing framework and network-related platform guarantees and support, including cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to acquire data, which is then fed into the intelligent chips within the distributed computing system provided by the foundational platform for computation.
[0108] (2) Data
[0109] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0110] (3) Data processing
[0111] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0112] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.
[0113] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.
[0114] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.
[0115] (4) General ability
[0116] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0117] (5) Smart products and industry applications
[0118] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart transportation, smart medical care, autonomous driving, smart cities, etc.
[0119] First, we will introduce the application scenarios of this application. This application can be, but is not limited to, applications with generative artificial intelligence (AIGC) functions (hereinafter referred to as synthetic applications) or cloud services provided by cloud-side servers. The following are introduced respectively:
[0120] 1. Synthetic Applications
[0121] The product form of the embodiment of the present application can be a composite application. The composite application can be run on a terminal device or a cloud-side server.
[0122] In one possible implementation, a synthesis application can implement a data generation task based on input data (such as images, text, audio, video, etc.), wherein the synthesis application can perform the data generation task in response to the input data (such as images, text, audio, video, etc.) to obtain generated data.
[0123] For example, the above data generation tasks can be, but are not limited to:
[0124] Text generation tasks: Various types of text content can be generated, including news reports, blog articles, product descriptions, social media posts, etc. It can generate logical and coherent text based on given topics and requirements.
[0125] Image generation tasks: can generate images, including illustrations, artwork, design drafts, etc. It can generate related image content based on a given description or keyword.
[0126] Audio generation tasks: This can generate speech content, including reading text aloud, voice assistant responses, etc. It can simulate human voice characteristics and intonation to make the generated speech sound more natural.
[0127] Content summarization and summary tasks: can read a large amount of text content and generate a summary or summary. It can extract key information from the text and present it to the user in a concise manner.
[0128] Language translation tasks: It can perform language translation, translating text from one language into another. It can handle multiple language pairs and provide accurate translation results.
[0129] Auto-reply and customer service: It can be used to automatically reply to user questions and provide customer service. It can understand the user's intention and give accurate answers or suggestions.
[0130] In one possible implementation, a user may open a synthesis application installed on a terminal device and input input data (such as images, text, audio, video, etc.). The synthesis application may generate data from the input data using the method provided in an embodiment of the present application and present the generated data to the user (the presentation method may be, but is not limited to, display, saving, uploading to the cloud, etc.).
[0131] In one possible implementation, a user can open a synthesis application installed on a terminal device and enter input data. The synthesis application can send the input data to a server on the cloud side. The server on the cloud side generates data for the input data using the method provided in an embodiment of the present application and transmits the generated data back to the terminal device. The terminal device can present the generated data to the user (the presentation method can be but is not limited to display, saving, uploading to the cloud side, etc.).
[0132] Next, the synthetic application in the embodiment of this application is introduced from the perspective of functional architecture and product architecture that implements the functions.
[0133] Reference Figure 1B , Figure 1B This is a schematic diagram of the functional architecture of the synthesis application in the embodiment of this application:
[0134] In one possible implementation, Figure 1B As shown, a synthesis application 102 can receive input parameters 101 (e.g., including input data) and generate production data 103. The synthesis application 102 can be executed on (for example) at least one computer system and includes computer code that, when executed by one or more computers, causes the computers to execute a natural language model trained by the method provided in the embodiments of the present application.
[0135] Reference Figure 1C , Figure 1C This is a schematic diagram of the physical architecture for running a composite application in an embodiment of the present application:
[0136] See also Figure 1C , Figure 1C A schematic diagram of a system architecture is shown. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1C (The example includes one server, and the server 200 is used for explanation) and can provide synthesis function services for one or more terminals.
[0137] Among them, the terminal 100 can be installed with a synthesis application, or a web page related to the synthesis function can be opened. The above application and web page can provide an interface. The terminal 100 can receive the relevant parameters entered by the user on the synthesis function interface and send the above parameters to the server 200. The server 200 can obtain the processing results based on the received parameters and return the processing results to the terminal 100.
[0138] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the processing result based on the received parameters by itself without the need for the cooperation of the server, and the embodiments of the present application are not limited to this.
[0139] Next describe Figure 1C The product form of the mid-terminal 100;
[0140] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a wearable device, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.
[0141] Figure 1D A schematic diagram of an optional hardware structure of the terminal 100 is shown.
[0142] refer to Figure 1D As shown, the terminal 100 may include a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a processor 170, an external interface 180, a power supply 190 and other components. Those skilled in the art will understand that Figure 1D These are merely examples of terminals or multi-function devices and do not limit the terminal or multi-function device. The terminal or multi-function device may include more or fewer components than shown in the figure, or may combine certain components or different components.
[0143] The input unit 130 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the portable multifunction device. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can detect user touch operations on or near it (for example, operations performed on or near the touch screen using a finger, joint, stylus, or any other suitable object) and drive corresponding connected devices according to pre-set programs. The touch screen can detect user touch actions on the touch screen, convert the touch actions into touch signals and transmit them to the processor 170. It can also receive and execute commands sent by the processor 170; the touch signals include at least touch point coordinate information. The touch screen 131 provides an input interface and an output interface between the terminal 100 and the user. Touch screens can be implemented using various types, including resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch screen 131, the input unit 130 may also include other input devices. Specifically, the other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as a volume control button 132 , a switch button 133 , etc.), a trackball, a mouse, a joystick, and the like.
[0144] Among them, the input device 132 can receive input data and the like.
[0145] The display unit 140 may be used to display information input by the user or provided to the user, various menus of the terminal 100, interactive interfaces, file display, and / or playback of any multimedia file. In an embodiment of the present application, the display unit 140 may be used to display the interface of a synthesis application, generated data, etc.
[0146] Memory 120 can be used to store instructions and data. It primarily includes an instruction storage area and a data storage area. The data storage area can store various data, such as multimedia files and text. The instruction storage area can store software units such as the operating system, applications, and instructions required for at least one function, or subsets or extensions thereof. It may also include non-volatile random access memory (RAM). It provides processor 170 with management functions for the hardware, software, and data resources within the computing and processing device, supporting control software and applications. It is also used to store multimedia files and running programs and applications.
[0147] The processor 170 is the control center of the terminal 100. It connects all components of the terminal 100 using various interfaces and circuits. By executing instructions stored in the memory 120 and accessing data stored therein, it executes various functions of the terminal 100 and processes data, thereby providing overall control of the terminal device. Optionally, the processor 170 may include one or more processing units. Preferably, the processor 170 may integrate an application processor and a modem processor, with the application processor primarily processing the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 170. In some embodiments, the processor and memory may be implemented on a single chip; in other embodiments, they may be implemented on separate chips. The processor 170 may also generate corresponding operational control signals and send them to the corresponding components of the computing and processing device. It may also read and process data in the software, particularly the data and programs in the memory 120, to enable the various functional modules therein to perform their corresponding functions, thereby controlling the corresponding components to operate as instructed.
[0148] Among them, the memory 120 can be used to store software codes related to the data processing method, the processor 170 can execute the steps of the chip's data processing method, and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to achieve corresponding functions.
[0149] The RF unit 110 (optional) can be used to send and receive information or receive and send signals during a call. For example, after receiving downlink information from the base station, it is passed to the processor 170 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF unit 110 can also communicate with network devices and other devices via wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0150] In this embodiment of the present application, the RF unit 110 may send input data to the server 200 and receive generated data sent by the server 200.
[0151] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network port.
[0152] The terminal 100 also includes a power supply 190 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 170 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.
[0153] The terminal 100 further includes an external interface 180 , which may be a standard Micro USB interface or a multi-pin connector, and may be used to connect the terminal 100 to other devices for communication, or to connect a charger to charge the terminal 100 .
[0154] Although not shown, the terminal 100 may also include a flashlight, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which will not be described in detail here. Some or all of the methods described below can be applied to Figure 1D In the terminal 100 shown.
[0155] Next describe Figure 1C The product form of the server 200;
[0156] Figure 2 A structural diagram of a server 200 is provided, such as Figure 2 As shown, the server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other via the bus 201.
[0157] The bus 201 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 2 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0158] The processor 202 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0159] The memory 204 may include volatile memory, such as random access memory (RAM). The memory 204 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard drive (HDD), or solid state drive (SSD).
[0160] The memory 204 may be used to store software codes related to the data processing method, and the processor 202 may execute the steps of the data processing method of the chip, and may also schedule other units to implement corresponding functions.
[0161] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above-mentioned terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.
[0162] It should be understood that the steps related to the model reasoning process in the embodiments of the present application involve AI-related operations. When performing AI operations, the instruction execution architecture of the terminal device and the server is not limited to the processor combined with the memory architecture described above. Figure 5 The system architecture provided in the embodiments of the present application is introduced in detail.
[0163] Figure 5 This is a schematic diagram of the system architecture provided in the embodiment of this application. Figure 5 As shown, the system architecture 500 includes an execution device 510 , a training device 520 , a database 530 , a client device 540 , a data storage system 550 , and a data collection system 560 .
[0164] The execution device 510 includes a calculation module 511, an I / O interface 512, a pre-processing module 513, and a post-processing module 514. The calculation module 511 may include the target model / rule 501, and the pre-processing module 513 and the post-processing module 514 are optional.
[0165] The execution device 510 may be a terminal device or a server that runs the aforementioned composite application.
[0166] The data acquisition device 560 is used to collect training samples. The training samples can be program files (including program codes and program input data), etc. After collecting the training samples, the data acquisition device 560 stores them in the database 530.
[0167] The training device 520 can train the neural network based on the training samples maintained in the database 530 to obtain the target model / rule 501.
[0168] It should be noted that, in actual applications, the training samples maintained in the database 530 may not all be collected by the data acquisition device 560, but may also be received from other devices. It should also be noted that the training device 520 may not train the target model / rule 501 entirely based on the training samples maintained in the database 530, but may also obtain training samples from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.
[0169] The target model / rule 501 obtained by training the training device 520 can be applied to different systems or devices, such as Figure 5 The execution device 510 shown may be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, a vehicle terminal, etc., or a server, etc.
[0170] Specifically, the training device 520 may transfer the trained model to the execution device 510 .
[0171] exist Figure 5 In the embodiment, the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with an external device. The user can input data to the I / O interface 512 through the client device 540 (for example, input data in the embodiment of the present application, etc.).
[0172] Preprocessing module 513 and preprocessing module 514 are used to preprocess the input data received by I / O interface 512. It should be understood that preprocessing module 513 and preprocessing module 514 may be absent or only one preprocessing module may be present. If preprocessing module 513 and preprocessing module 514 are absent, computing module 511 may be used directly to process the input data.
[0173] When the execution device 510 preprocesses the input data, or when the computing module 511 of the execution device 510 performs calculations and other related processing, the execution device 510 can call the data, code, etc. in the data storage system 550 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 550.
[0174] Finally, the I / O interface 512 provides the processing results (eg, generated data, etc.) to the client device 540 , and thus provides them to the user.
[0175] exist Figure 5In the illustrated case, the user can manually input data, and this "manual input data" can be operated through the interface provided by I / O interface 512. In another case, client device 540 can automatically send input data to I / O interface 512. If the automatic transmission of input data by client device 540 requires user authorization, the user can set the corresponding permissions in client device 540. The user can view the results output by execution device 510 on client device 540, and the specific presentation form can be a display, sound, action, etc. Client device 540 can also serve as a data acquisition terminal, collecting input data input into I / O interface 512 and output results from I / O interface 512 as new sample data and storing them in database 530. Of course, collection can also be performed without client device 540, and instead the I / O interface 512 directly stores the input data input into I / O interface 512 and output results from I / O interface 512 as new sample data in database 530.
[0176] It is worth noting that Figure 5 This is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, Figure 5 In the embodiment, the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510. It should be understood that the execution device 510 can be deployed in the client device 540.
[0177] From the inference side of the model:
[0178] In the embodiment of the present application, the computing module 511 of the above-mentioned execution device 510 can obtain the code stored in the data storage system 550 to implement the steps related to the model reasoning process in the embodiment of the present application.
[0179] In an embodiment of the present application, the computing module 511 of the execution device 510 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.
[0180] Specifically, the computing module 511 of the execution device 510 can be a hardware system with an execution instruction function, and the steps related to the model reasoning process provided in the embodiment of the present application can be software codes stored in the memory. The computing module 511 of the execution device 510 can obtain the software code from the memory and execute the obtained software code to implement the steps related to the model reasoning process provided in the embodiment of the present application.
[0181] It should be understood that the computing module 511 of the execution device 510 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to the model reasoning process provided in the embodiment of the present application can also be implemented by the hardware system that does not have the function of executing instructions in the computing module 511 of the execution device 510, which is not limited here.
[0182] From the training side of the model:
[0183] In the embodiment of the present application, the training device 520 can obtain the memory ( Figure 5 Not shown in the figure, the code stored in the training device 520 can be integrated into or deployed separately from the training device 520 to implement the steps related to model training in the embodiments of the present application.
[0184] In an embodiment of the present application, the training device 520 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.
[0185] It should be understood that the training device 520 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to model training provided in the embodiments of the present application can also be implemented by the hardware system in the training device 520 that does not have the function of executing instructions, which is not limited here.
[0186] 2. Synthesis function cloud services provided by the server:
[0187] In a possible implementation, the server may provide the synthesis function service to the terminal side through an application programming interface (API).
[0188] Among them, the terminal device can send relevant parameters (such as input data) to the server through the API provided by the cloud. The server can obtain processing results (such as generated data, etc.) based on the received parameters and return the processing results to the terminal.
[0189] The description of the terminal and the server can be the same as that of the above embodiments, and will not be repeated here.
[0190] like Figure 6 The process of using a synthetic function cloud service provided by a cloud platform is shown.
[0191] 1. Activate and purchase content review services.
[0192] 2. Users can download the software development kit (SDK) corresponding to the content review service. Usually, the cloud platform provides multiple development versions of the SDK for users to choose according to the requirements of the development environment, such as JAVA version SDK, Python version SDK, PHP version SDK, Android version SDK, etc.
[0193] 3. After the user downloads the corresponding version of the SDK to the local computer as needed, the SDK project is imported into the local development environment for configuration and debugging. The local development environment can also be used to develop other functions, forming an application that integrates synthetic functional capabilities.
[0194] 4. When a composite function application is used and needs to perform a composite function, it can trigger an API call for the composite function. When the application triggers the composite function, it initiates an API request to the running instance of the composite function service in the cloud environment. The API request carries input data, which is then processed by the running instance in the cloud environment to obtain a result (e.g., generated data).
[0195] 5. The cloud environment returns the processing results to the application, thus completing a synthetic function service call.
[0196] In addition to applications and cloud services, the implementation form of this application can also be in large-model inference acceleration libraries and large-model application SDKs.
[0197] In order to better understand the solution of the embodiment of the present application, the following takes text generation as an example. Figures 2 to 4 A brief introduction to possible application scenarios of the embodiments of the present application is given.
[0198] Figure 3 A natural language processing system is shown, comprising a user device and a data processing device. The user device includes an intelligent terminal such as a mobile phone, personal computer, or information processing center. The user device is the initiator of natural language data processing, initiating requests such as language questions and answers or inquiries. Typically, users initiate requests through their user devices.
[0199] The aforementioned data processing devices can be devices or servers with data processing capabilities, such as cloud servers, network servers, application servers, and management servers. The data processing devices receive query statements, voice, text, and other information from smart terminals via interactive interfaces. They then use their memory and processors to perform language data processing, including machine learning, deep learning, search, reasoning, and decision-making, and then feed the results back to the user device. The memory in a data processing device is a general term encompassing both local storage and databases storing historical data. The databases can be located on the data processing device or on other network servers.
[0200] exist Figure 3 In the natural language processing system shown, the user device can receive user instructions. For example, the user device can receive a piece of text input by the user, and then initiate a request to the data processing device, so that the data processing device executes a natural language processing application (such as natural language generation, text classification, text reasoning, named entity recognition, translation, etc.) for the piece of text obtained by the user device, thereby obtaining the processing results of the corresponding natural language processing application for the piece of text (such as predicted word results, classification results, reasoning results, named entity recognition results, translation results, etc.).
[0201] In an embodiment of the present application, a user device may receive instructions from a user. For example, the user device may receive a piece of text (such as input data) input by a user, and then initiate a request to a data processing device so that the data processing device executes a natural language processing application (such as text synthesis, etc.) for the piece of text obtained by the user device, thereby obtaining a processing result of the corresponding natural language processing application for the piece of text (such as generated data, etc.).
[0202] Text in Figure 3 In the embodiment of the present invention, the data processing device can process the above text data using the method provided in the embodiment of the present application.
[0203] Figure 4 Another natural language processing system is shown in Figure 4 In the process, the user device directly acts as a data processing device. The user device can directly receive input from the user and process it directly by the hardware of the user device itself. The specific process is the same as Figure 3 Similarly, please refer to the above description and will not be repeated here.
[0204] Figure 4 It is a schematic diagram of a natural language processing related device 300 provided in an embodiment of the present application.
[0205] Figure 3 and Figure 4The processor in the embodiment can perform data training / machine learning / deep learning through a neural network model or other model, and use the model finally trained or learned by the data (such as the natural language model in the embodiment of the present application, etc.) to perform natural language processing applications (such as program synthesis, etc.) on text data (such as the input data text described in the embodiment of the present application), thereby obtaining corresponding processing results.
[0206] Since the embodiments of the present application involve the application of a large number of neural networks, in order to facilitate understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application are first introduced below.
[0207] (1) Neural Network
[0208] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input. The output of the operation unit can be:
[0209]
[0210] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0211] (2) Transformer layer
[0212] The neural network includes an embedding layer and at least one transformer layer, and the at least one transformer layer can be N transformer layers (N is an integer greater than 0), wherein each transformer layer includes an attention layer, an add&norm layer, a feed forward layer, and an add&norm layer that are adjacent in sequence. In the embedding layer, the current input is embedded to obtain multiple embedding vectors; in the attention layer, P input vectors are obtained from the previous layer of the first transformer layer, and with any first input vector among the P input vectors as the center, based on the correlation between each input vector within a preset attention window and the first input vector, the intermediate vector corresponding to the first input vector is obtained, thereby determining the P intermediate vectors corresponding to the P input vectors; in the pooling layer, the P intermediate vectors are merged into Q output vectors, wherein the multiple output vectors obtained by the last transformer layer in the transformer layer are used as feature representations of the current input.
[0213] (3) Attention mechanism
[0214] The attention mechanism mimics the internal process of biological observation behavior, namely, a mechanism that aligns internal experience and external sensations to increase the observation precision of certain areas. It can quickly filter out high-value information from a large amount of information using limited attention resources. The attention mechanism can quickly extract important features from sparse data and is therefore widely used in natural language processing tasks, especially machine translation. The self-attention mechanism is an improvement on the attention mechanism, which reduces dependence on external information and is better at capturing the internal correlation of data or features. The essential idea of the attention mechanism can be rewritten as the following formula:
[0215] Here, Lx = ||Source|| represents the length of the Source. This formula implies that the elements in the Source are imagined to consist of a series of data pairs. Given a Query element in the target, the similarity or correlation between the Query and each Key is calculated to obtain the weight coefficient for each Key's corresponding Value. The weighted sum of the Values is then taken to obtain the final Attention value. Essentially, the Attention mechanism performs a weighted sum of the Values of the Source elements, with the Query and Key used to calculate the weight coefficient for the corresponding Value. Conceptually, Attention can be understood as selectively filtering out a small amount of important information from a large amount of information and focusing on this important information, while ignoring the majority of less important information. This focusing process is reflected in the calculation of the weight coefficients: the larger the weight, the more focus is placed on the corresponding Value. In other words, the weight represents the importance of the information, while the Value represents the corresponding information. The self-attention mechanism can be understood as internal attention. The attention mechanism occurs between the Query element of the Target and all elements of the Source. The self-attention mechanism refers to the attention mechanism that occurs between the internal elements of the Source or the internal elements of the Target. It can also be understood as the attention calculation mechanism in the special case of Target = Source. The specific calculation process is the same, only the calculation object has changed.
[0216] (4) Natural language processing (NLP)
[0217] Natural language refers to human language, and natural language processing (NLP) is the processing of human language. Natural language processing is the process of systematically analyzing, understanding, and extracting information from text data in an intelligent and efficient manner. By using NLP and its components, we can manage very large amounts of text data, perform a large number of automated tasks, and solve a wide variety of problems, such as automatic summarization, machine translation (MT), named entity recognition (NER), relation extraction (RE), information extraction (IE), sentiment analysis, speech recognition, question answering, and topic segmentation.
[0218] (5) Pre-trained language model
[0219] A pretrained language model is a natural language sequence encoder that encodes each word in a natural language sequence into a vector representation for prediction. Its training consists of two phases. In the pre-training phase, the model is trained on a large amount of unsupervised text for language modeling tasks, thereby learning a word representation. In the fine-tuning phase, the model is initialized using the parameters learned in the pre-training phase and trained in a relatively small number of steps on downstream tasks such as text classification and sequence labeling. This allows the semantic information gained from pre-training to be successfully transferred to downstream tasks.
[0220] (6) Autoregressive language model
[0221] An autoregressive language model is a model that can predict the next possible word (such as "good") based on a given context (such as "the phone is very"), which usually predicts the word in the context on the right given the context on the left, but can also predict a word in the middle given the context on both the left and right sides.
[0222] (7) Large language model: A language model with a relatively large number of parameters. Sentences are segmented and then encoded into floating-point vectors. The output is a new word through projection mapping and top-1 calculation.
[0223] (8) Inference / deployment: The forward computation process of a neural network.
[0224] (9) Inference latency: The time required for the forward computation of a neural network.
[0225] (10) KV Cache: The KV vector is the output of the linear mapping layer k and v of the Transformer layer. It concatenates the output of multiple words in the full part into two Batch*Sequence*Dim tensors. When the graph is incremented, it is dynamically appended to the last position of the tensor.
[0226] (11) Token: refers to the smallest unit in a text. Generally, a token can be a word, number, punctuation mark, single letter, or any single element that can be used for text analysis.
[0227] (12) Speculative sampling: This method is used to improve the computational memory access ratio when inferring large language models, ensuring that the sampling distribution is exactly the same as that of the original model. It uses two models: one is the original target model, and the other is an approximate model that is much smaller than the original model. The approximate model is used for serial autoregressive decoding, and the large model is used to verify the results. The large model can verify K tokens in parallel at a time, thus significantly reducing the number of autoregressive times of the large model, thereby improving the inference performance of the large model.
[0228] (13) Speculative small model: An approximate model that is much smaller than the large model used in speculative sampling.
[0229] (14) Separate deployment: Separate the full inference and incremental inference of large models to different machines for deployment. Full inference and incremental inference can use different batch sizes for inference.
[0230] (15) Dynamic batch: that is, continuous batch, which means that during batch inference, the decoded sentences are returned in advance and new sentences are dynamically added for decoding.
[0231] (16) Backpropagation algorithm
[0232] Convolutional neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during training, reducing the reconstruction error loss of the super-resolution model. Specifically, the forward propagation of the input signal to the output generates an error loss. This error loss information is then backpropagated to update the parameters of the initial super-resolution model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by the error loss, aiming to obtain the optimal super-resolution model parameters, such as the weight matrix.
[0233] (17) Loss function
[0234] During the training of a deep neural network, because we want the output of the deep neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, there is usually an initialization process before the first update, which is to pre-configure the parameters for each layer in the deep neural network.) For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the deep neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss as much as possible.
[0235] Inference for large language models is primarily divided into two phases. The first is the full prefill inference phase, where all user-entered words are fed into the network for computation, obtaining the output of the final word. This phase allows for parallel inference of multiple words, effectively utilizing hardware computing and memory bandwidth. The second is the incremental decoding inference phase, where the output of the full inference phase serves as the input for the current step and is fed into the network for computation. The output of the current word is then used as the input for the next step, continuing until decoding is complete. This phase is an autoregressive process that requires decoding word by word, resulting in low parallelism and hardware utilization. Repeated memory accesses for weights waste memory bandwidth and do not fully utilize the compute units. However, if a large model can decode multiple words in parallel in a single step, latency will be minimal compared to decoding a single word, improving hardware utilization.
[0236] To improve the parallelism of large model increments, the concept of predictive execution, also known as speculative sampling, is applied to large model inference. This approach ensures the same sampling distribution as the original model. It uses two models: the original target model and a much smaller approximate model. The approximate model performs serial autoregressive decoding, while the large model verifies the results. The large model can verify multiple tokens simultaneously, significantly reducing the number of autoregressive cycles and improving the compute-to-memory ratio of large model inference, thereby enhancing inference performance.
[0237] Speculative reasoning can accelerate large model reasoning without loss, but it requires additional training and deployment of speculative small models. Small models also require maintenance and deployment costs. When the large and small models are poorly matched, speculative reasoning may even lead to performance regression and increased reasoning costs. Therefore, minimizing the cost of small models while ensuring a high speedup ratio for large model reasoning is a key issue in speculative reasoning.
[0238] In order to solve the above problems, the present invention provides a data processing method. The model training method of the present invention is described in detail below with reference to the accompanying drawings.
[0239] Reference Figure 7 , Figure 7 A data processing method according to an embodiment of the present invention is shown in FIG. Figure 7 As shown, a data processing method provided in an embodiment of the present application may include steps 701 to 703, and these steps are described in detail below.
[0240] 701. Generate second data through a first network based on the first data;
[0241] 702. Generate third data based on the first data through a second network; wherein the first network includes a first network layer, the second network includes a second network layer, and the first network layer is a partial network of the second network layer.
[0242] Inference for large language models is divided into two main phases. The first is the full prefill inference phase, in which all user-entered words are fed into the network computation to obtain the output of the final word. This phase allows for parallel inference of multiple words, effectively utilizing hardware computing and memory bandwidth. The second is the incremental decoding inference phase, in which the output of the full inference is used as the input for the current step, fed into the network computation, and the output of the current word is then used as the input for the next step until decoding is complete. Figure 7 The corresponding embodiment may be a step in the incremental reasoning phase.
[0243] For example, refer to Figure 9 , Figure 9The large-model inference architecture and process involved in the embodiment of the present application are shown. In the overall large-model training and inference system, the embodiment of the present application is mainly applied to the inference deployment node of the large model, and can be widely used to accelerate the inference service of large-model application scenarios such as dialogue, text, AI image or video generation. In the large-model inference process within the deployment node, the embodiment of the present application mainly involves the incremental inference part, which is used to reduce the number of incremental autoregressive steps and improve the incremental inference performance of the large model.
[0244] Among them, the first network can be an approximate model, which is used to perform serial autoregressive decoding, and the second network can be a large model, which is used to verify the results. The large model can verify multiple tokens in parallel at a time, thereby significantly reducing the number of autoregressive times of the large model.
[0245] Step 701 may be an autoregressive generation process performed by the first network. For example, the first data may be a token obtained after full inference, and the first data may be processed by the first network to obtain the next token adjacent to the first data (i.e., the second data). Alternatively, the first data may be a token generated by the first network during an autoregressive process, and the first data may be processed by the first network to obtain the next token adjacent to the first data (i.e., the second data).
[0246] That is, in one possible implementation, before generating the second data using the first network based on the first data, the first data may also be generated using the first network based on the fourth data. This step is a step of generating the first data using the first network through an autoregressive method.
[0247] Step 702 may be a verification process of the second data via the second network.
[0248] In one possible implementation, the first network and the second network can reuse network layers of different depths in the original model. Since the second network is used to verify the data obtained by the first network, the network layer corresponding to the second network is deeper. For example, the first network may include a first network layer (and may also include a first decoding layer), the second network may include a second network layer (and may also include a second decoding layer), and the first network layer is a partial network of the second network layer.
[0249] In one possible implementation, the first network layer and the second network layer both belong to the target neural network (that is, the original model, that is, the target neural network is the model used in the full reasoning process corresponding to the incremental reasoning process), and the second network layer includes the first network layer and a third network layer connected to the first network layer, and the third network layer is closer to the output layer of the target neural network than the first network layer.
[0250] Multiple available sub-models can be formed according to the different depths of the original model. These sub-models can be used as speculative small models, and the original large model reasoning can be divided into multiple levels of speculative reasoning for step-by-step acceleration.
[0251] Reference Figure 8 , Figure 8 A schematic diagram of a network structure is shown in Figure 8 As shown, the original large model includes 20 network layers, which are divided into 4 sub-models according to the depth, where LM Head i corresponds to the decoder Head of the i-th small model. Figure 8 In the inference scheme shown, stage = 4, that is, four levels of "nesting doll" speculative inference. The i-th level of speculative inference uses the i-th model to predict the decoding result of the i+1-th model, accelerating the inference of the i+1-th model. The acceleration is gradually accelerated, ultimately achieving an exponential speedup ratio.
[0252] For example, in Figure 8 In the example, network layers 1-4 and LM head1 may be the first network, network layers 1-4 may be the first network layer, network layers 5-8 and LM head2 may be the second network, and network layers 5-8 may be the second network layer.
[0253] The second network, as a model for verification results, can verify multiple tokens in parallel at one time, thereby significantly reducing the number of autoregressions. Similarly, the second network can also verify the first data in parallel. In one possible implementation, fifth data can be generated based on the fourth data through the second network; wherein the action of generating the fifth data and the action of generating the third data are performed in parallel; and the first data is verified based on the fifth data.
[0254] In a possible implementation, the first network includes a first network layer and a first decoding layer, and the second network includes a second network layer and a second decoding layer; when the first network performs the verification process for the second data, the processing result of the first network layer can be reused.
[0255] Specifically, based on the first data, a first processing result can be obtained through the first network layer, and based on the first processing result, second data can be obtained through the first decoding layer; based on the first processing result, a second processing result can be obtained through the third network layer, and based on the second processing result, third data can be generated through the second decoding layer.
[0256] In an embodiment of the present application, the network (second network) that performs the incremental verification process during speculative reasoning is reused with the network that performs autoregressive reasoning (part of it, that is, the first network), which can reduce additional reasoning deployment overhead, and the second network can reuse the reasoning results of the first network in memory when performing incremental verification, without the need for additional delay, thereby reducing computational overhead.
[0257] Furthermore, the first network layer in the first network and the second network layer in the second network can both belong to the large model for full inference, eliminating the need to maintain and train separate small models, saving costs. The small models do not consume additional deployment resources, and their memory can reuse the model parameters and KV Cache of the large model. By reusing the inference branches of the large model, the inference results of the sub-models can be directly reused without additional latency, allowing the large model inference to cover the overhead of the small model.
[0258] It should be understood that in addition to the second network, more-order verification networks (such as the third network in the embodiments of the present application) can also be set. Similar to the relationship between the first network and the second network described above, the second network and the third network can also be constructed by network layers of different depths in the large model. For example, the third network can include a third network layer, and the second network layer can be part of the third network layer.
[0259] 703. Verify the second data according to the third data.
[0260] In a possible implementation, the verifying the second data based on the third data includes: when the third data and the second data are consistent, determining that the verification of the second data has passed; when the third data and the second data are inconsistent, determining that the verification of the second data has failed.
[0261] In one possible implementation, when it is determined that the second data has passed verification, sixth data may be generated based on the first data via a third network; wherein the third network includes a fourth network layer, and the second network layer is a portion of the fourth network layer; and the second data may be verified based on the sixth data. In other words, the verification process may continue through the next-level verification network until the final level of verification is reached.
[0262] In one possible implementation, when the third data and the second data are different, it can be determined that the second data has failed verification. In this case, the third data can be used as the next token adjacent to the first data, and the third data can be verified data. After that, the next-level verification network (for example, the third network in the embodiment of the present application can verify the third data). Specifically, the sixth data can be generated based on the first data through the third network; and the third data can be verified based on the sixth data.
[0263] In one possible implementation, when the third data and the second data are different, it can be determined that the second data has failed verification. In this case, the third data can be used as the next token adjacent to the first data, and the third data can be verified data. Since the third data is generated by the second network, in order to ensure the matching degree in the next stage, the second network can perform autoregressive inference based on the third data. In another possible implementation, when it is determined that the second data has failed verification, the second network generates seventh data based on the third data.
[0264] For example, refer to Figure 8 The data generated by the first network include a, b, c, d. The data generated by the second network verification are a, b, c and e. Among them, the verification for d fails, so e is used as the currently verified data, and a, b, c and e are used as the objects of verification for the next level (third network). The data generated by the third network verification are a, b, f. Since the verification for c fails, the third network can generate g based on f, that is, the third network generates data g based on data f through autoregressive inference.
[0265] In one possible implementation, when the third data and the second data are different, it can be determined that the second data has failed verification. In this case, the third data can be used as the next token adjacent to the first data, and the third data can be verified data. Since the third data is generated by the second network, in order to ensure the matching degree in the next stage, autoregressive inference can be performed based on the third data through the first network, and corresponding verification can be performed through the second network. In one possible implementation, when it is determined that the second data has failed verification, the eighth data is generated based on the third data through the first network; the seventh data is generated based on the third data through the second network; and then the eighth data can be verified based on the seventh data.
[0266] In one possible implementation, the method is applied to an incremental reasoning process. When the historical verification pass rate in the incremental reasoning process meets a preset condition and it is determined that the verification of the second data fails, seventh data is generated based on the third data through the second network.
[0267] In a possible implementation, when the historical verification pass rate does not meet the preset condition during the incremental reasoning process and it is determined that the verification of the second data fails, eighth data is generated based on the third data through the first network.
[0268] In one possible implementation, when it is determined that the verification of the second data fails and the number of data that has passed the verification of the second network in the current verification cycle is less than a preset value, autoregressive data generation is performed through the second network until the number of data that has passed the verification of the second network in the current verification cycle for the first batch reaches the preset value.
[0269] In a possible implementation, when it is determined that the verification of the second data fails, autoregressive data generation is performed through the first network, and the data autoregressively generated by the first network is verified through the second network until the amount of data that passes the verification of the second network in the current verification cycle reaches the preset value.
[0270] The embodiments of the present application can also be applied to the reasoning process of dynamic batches, especially in the scenario of multi-stage verification, when the number of tokens that pass the verification in one stage cannot fill bs*n (bs is the number of batches, n is the number of data that needs to pass the verification in the verification cycle of each batch, which is the preset value in the above embodiment), if it is directly sent to the next stage for verification, the matching degree will decrease step by step, which directly affects the computational parallelism and the final acceleration ratio. Therefore, the embodiments of the present application propose a multi-level speculative batch reasoning solution based on multi-level cache, in which a cache is set in each level of speculative reasoning, and the cache of the i-th level is Li Cache. When the cache is filled to a threshold, the next level of reasoning is triggered. This triggering strategy can be determined according to the actual matching degree of the model. It is assumed here that the triggering strategy is to fill the cache to bs*n. When the cache is not full, there are two implementation schemes. Solution 1: Regroup the tokens that need to be autoregressed into a new batch, perform unaligned batch autoregression on the corresponding sub-models until the cache is filled, triggering the next stage of reasoning. Solution 2: Based on the number of tokens matched in the current stage, fall back to stage 0, make small model 0 autoregressively predict the next n tokens, and then reason forward to the current stage until the cache is filled to stage*n, and then proceed to the next stage of reasoning. In the actual multi-batch reasoning process, you can dynamically choose between these two filling strategies. You can choose a strategy with lower overhead based on the degree of cache missing. You can refer to the schematic diagram. Figure 10 .
[0271] Assume that the time for inference of the model at level 4 is x. Multiple tokens can be inferred in parallel, and the time for inference of 1 token, 4 tokens, and 4x4 tokens is x.
[0272] Assume that the time required for four autoregressions of small model 0 is 4x, and the cache is not full after small model 1 is verified, so the next level of verification cannot be triggered. In this case, if we choose solution 1, the cost is 2x the time required for one autoregression of small model 1. If we choose solution 2, the cost is 5x the time required for four autoregressions of small model 0 and one verification of small model 1. Therefore, we choose solution 1.
[0273] Small model 2 verification results in an unaligned token cache and cannot trigger the next level of verification. If option 1 is chosen, the overhead is three autoregressions for small model 2, which takes 3*3x = 9x the time. If option 2 is chosen, the overhead is minimal, with one autoregression for small model 0 and one verification for small models 1 and 2, for a total time of 6x. In this case, option 2 is greedily chosen. Option 2 allows for dynamic adjustment of the triggering strategy when filling the cache, for example, ensuring that the number of tokens per batch exceeds the number of missing tokens in the corresponding batch of the next stage.
[0274] In this embodiment, the idle waiting time of computing units during speculative reasoning multi-batch calculations can be reduced. A token cache is added to each level to ensure that the matching length of each level is not reduced due to a low matching degree at the previous level. The advantage of the above-mentioned solution one is that it can control the number of tokens generated, but it does not apply speculative acceleration, and too many autoregressions will slow down the overall speed. The advantage of solution two is that it can accelerate the generation of multiple autoregressions, but poor matching can lead to rollbacks. In practical applications, the two solutions can be combined based on the matching degree.
[0275] In one possible implementation, the method is applied to an incremental reasoning process, which is a response generation process for input data; sub-attention information corresponding to the length of the generated data can be determined based on the attention information in the full reasoning process corresponding to the incremental reasoning process (for example, it can be an attention mask) and the length of the data generated in the response generation process; and second data can be generated through the first network based on the sub-attention information and the first data.
[0276] Speculative inference increases the complexity of inference deployment when using batch inputs, requiring complex engineering implementation to manage the KV cache and the misalignment of incremental inference steps and decode steps. Multi-stage speculative inference also faces the problem of vacuous batch verification, which reduces throughput and mismatch. Therefore, this paper proposes an efficient speculative inference solution for multiple batches. This solution is not only applicable to the proposed "nesting doll" multi-stage speculative small model inference, but can also be extended to standard single-stage speculative small model inference applications, improving the speedup ratio of multi-batch speculative inference. The above examples describe the implementation of multi-stage verification.
[0277] When speculative inference stage = 1, the small model generates bs*n tokens each time through autoregression and submits them to the large model for verification. The large model may have a different number of tokens that pass verification in each batch, resulting in misalignment of the decoding step batches during the next generation and verification of the large and small models. Furthermore, the lengths n in each batch may vary, which is unsupported by the original attention calculation logic. The solution proposed in this paper is to add a decoded_length parameter for the decoded sequence length of each batch during incremental inference. The attention_mask in each batch is dynamically updated during each inference based on the decoded_length and verification length n to ensure the correctness of the attention calculation. The diagram of the attention_mask is shown below. If the lengths n in each batch vary, the attention calculation logic needs to be modified to adapt. This solution reuses the original multi-batch full-scale inference logic without redundant computation or excessive memory overhead, significantly improving throughput compared to single-batch inference.
[0278] For example, refer to Figure 11 , Figure 11 This is an illustration of updating attention_mask based on decoded_length and verification length n, where the blank space indicates masked data.
[0279] In one possible implementation, the method is executed on a first node, and a full reasoning process corresponding to the incremental reasoning process is executed on a second node; the first node can receive the processing results (for example, including output tokens and KV cache) obtained through the full reasoning process sent by the second node; the processing results are used to perform the incremental reasoning process.
[0280] In one possible implementation, the first data is data generated during the incremental reasoning process for a first batch, the method is applied to a first node, and when executing the method, the first node also performs the incremental reasoning process for a second batch in parallel. When the incremental reasoning process for the first batch is completed and the first node has not yet completed the incremental reasoning process for the second batch, the data in the memory space corresponding to the incremental reasoning process for the first batch is cleared, and the incremental reasoning process for the third batch is started.
[0281] Separate deployment and dynamic batching are effective means to improve the reasoning throughput of large models. Separate deployment refers to deploying full reasoning and incremental reasoning separately to different computing devices (for example, the first node and the second node in the embodiment of this application). Optionally, full reasoning can use bs=1 reasoning to ensure low latency, while incremental reasoning can use a larger bs to ensure high throughput. The embodiment of this application can organically combine speculative reasoning with separate deployment and dynamic batching. Based on the multi-batch multi-level reasoning scheme proposed by the present invention, the throughput of speculative reasoning can be further improved. The basic working method of this scheme is shown in the figure below. This scheme is not only applicable to the "nesting doll-style" multi-level speculative reasoning introduced in the above embodiment, but can also be extended to the application scenarios of ordinary speculative reasoning.
[0282] Reference Figure 12 , Figure 12 The figure shows a split deployment example. Full inference is deployed to four nodes with bs = 1. If a separate small model is used for speculative inference, a new small model full inference node is required. If a "nesting doll" small model solution is adopted, the full inference node of the large model can be reused without any additional overhead. Incremental inference is deployed to one node with bs = 2. After full inference is complete, the output token and KV cache are sent to the incremental node. The incremental node then performs multi-stage speculative inference in sequence. The large and small models are deployed to the same incremental node for alternating inference. If stage = 1, combined with dynamic batching, the overall throughput can be further improved. Because the data matching degree in each batch varies, the number of decoding steps is also different. If the decoding-completed sentences are exited early, hardware utilization can be improved. In addition, it is necessary to ensure that valid tokens are swapped out during each dynamic batch scheduling, and the KV cache at the corresponding location is cleared and the corresponding decode_len is updated. If stage > 1, multi-level caches must be set to ensure the matching degree of multi-level speculative reasoning. All caches at the corresponding batch location must be cleared during dynamic batch scheduling.
[0283] Reference Figure 13 , Figure 13 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application is shown in FIG. Figure 13 As shown, an embodiment of the present application provides a data processing device 1300, including:
[0284] Processing module 1301 is used to generate second data based on the first data through the first network; generate third data based on the first data through the second network; wherein the first network includes a first network layer, the second network includes a second network layer, and the first network layer is a partial network of the second network layer; and verify the second data based on the third data.
[0285] For a detailed description of the processing module, please refer to the above embodiment. Figure 7 The similarities between the corresponding embodiments are not repeated here.
[0286] In a possible implementation, before generating the second data based on the first data through the first network, the processing module 1301 is further configured to:
[0287] The first data is generated through the first network according to the fourth data.
[0288] In a possible implementation, the processing module 1301 is further configured to:
[0289] generating fifth data through the second network according to the fourth data; wherein the action of generating the fifth data and the action of generating the third data are performed in parallel;
[0290] The first data is verified according to the fifth data.
[0291] In one possible implementation, applied to an incremental reasoning process, the first network layer and the second network layer both belong to a target neural network, and the target neural network is a model used in a full reasoning process corresponding to the incremental reasoning process.
[0292] In a possible implementation, the second network layer includes the first network layer and a third network layer connected to the first network layer, and the third network layer is closer to the output layer of the target neural network than the first network layer.
[0293] In one possible implementation, the first network includes a first network layer and a first decoding layer, and the second network includes a second network layer and a second decoding layer;
[0294] The processing module 1301 is specifically configured to:
[0295] Obtaining a first processing result through the first network layer according to the first data, and obtaining second data through the first decoding layer according to the first processing result;
[0296] According to the first processing result, a second processing result is obtained through the third network layer, and according to the second processing result, third data is generated through the second decoding layer.
[0297] In a possible implementation, the processing module 1301 is specifically configured to:
[0298] When the third data and the second data are consistent, determining that the verification of the second data is passed;
[0299] When the third data and the second data are inconsistent, it is determined that the verification of the second data fails.
[0300] In a possible implementation, the processing module 1301 is further configured to:
[0301] When it is determined that the verification of the second data is passed, generating sixth data based on the first data through a third network; wherein the third network includes a fourth network layer, and the second network layer is a partial network of the fourth network layer;
[0302] The second data is verified according to the sixth data.
[0303] In a possible implementation, the processing module 1301 is further configured to:
[0304] When it is determined that the verification of the second data fails, generating sixth data based on the first data through a third network; wherein the third network includes a fourth network layer, and the second network layer is a partial network of the fourth network layer;
[0305] The third data is verified according to the sixth data.
[0306] In a possible implementation, the processing module 1301 is further configured to:
[0307] When it is determined that the verification of the second data fails, seventh data is generated based on the third data through the second network.
[0308] In one possible implementation, applied to an incremental reasoning process, the processing module 1301 is specifically configured to:
[0309] When the historical verification pass rate in the incremental reasoning process meets the preset condition and it is determined that the verification of the second data fails, seventh data is generated according to the third data through the second network.
[0310] In a possible implementation, the processing module 1301 is further configured to:
[0311] When it is determined that the verification of the second data fails, generating eighth data based on the third data through the first network;
[0312] generating seventh data through the second network according to the third data;
[0313] The eighth data is verified based on the seventh data.
[0314] In a possible implementation, the processing module 1301 is specifically configured to:
[0315] When the historical verification pass rate does not meet the preset conditions during the incremental reasoning process and it is determined that the verification of the second data fails, eighth data is generated based on the third data through the first network.
[0316] In one possible implementation, the incremental reasoning process is applied to an incremental reasoning process, which is a response generation process for input data. The processing module 1301 is specifically configured to:
[0317] Determining sub-attention information corresponding to the length of the generated data according to the attention information in the full inference process corresponding to the incremental inference process and the length of the generated data in the reply generation process;
[0318] Generate second data through the first network based on the sub-attention information and the first data.
[0319] In a possible implementation, the apparatus is executed on a first node, and a full reasoning process corresponding to the incremental reasoning process is executed on a second node;
[0320] Before generating the second data through the first network based on the first data, the processing module 1301 is further configured to:
[0321] Receive the processing result obtained through the full reasoning process sent by the second node; the processing result is used to perform the incremental reasoning process.
[0322] In one possible implementation, the first data is data generated during an incremental reasoning process for a first batch. The apparatus is applied to a first node, and when executing the apparatus, the first node also performs an incremental reasoning process for a second batch in parallel. The processing module 1301 is further configured to:
[0323] When the incremental reasoning process for the first batch is completed and the first node has not yet completed the incremental reasoning process for the second batch, the data in the memory space corresponding to the incremental reasoning process of the first batch is cleared, and the incremental reasoning process of the third batch is started.
[0324] In one possible implementation, the processing module 1301 is further configured to: when it is determined that the second data fails verification and the number of data that passes verification in the second network in the current verification cycle is less than a preset value, perform autoregressive data generation through the second network until the number of data that passes verification in the second network in the current verification cycle for the first batch reaches the preset value; or
[0325] When it is determined that the verification of the second data fails, autoregressive data generation is performed through the first network, and the data generated by the autoregressive of the first network is verified through the second network until the amount of data that passes the verification of the second network in the current verification cycle reaches the preset value.
[0326] Next, we will introduce an execution device provided by the embodiment of the present application. Figure 14 , Figure 14 This is a structural diagram of an execution device provided in an embodiment of the present application. The execution device 1400 can be specifically manifested as a virtual reality VR device, a mobile phone, a tablet, a laptop, a smart wearable device, a monitoring data processing device or a server, etc., which is not limited here. Specifically, the execution device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403 and a memory 1404 (wherein the number of processors 1403 in the execution device 1400 can be one or more, Figure 14 (taking one processor as an example), the processor 1403 may include an application processor 14031 and a communication processor 14032. In some embodiments of the present application, the receiver 1401, the transmitter 1402, the processor 1403 and the memory 1404 may be connected via a bus or other means.
[0327] Memory 1404 may include read-only memory and random access memory, and provides instructions and data to processor 1403. A portion of memory 1404 may also include non-volatile random access memory (NVRAM). Memory 1404 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0328] Processor 1403 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.
[0329] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1403. Processor 1403 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1403. The above processor 1403 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1403 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium is located in memory 1404. Processor 1403 reads information from memory 1404 and, in conjunction with its hardware, completes the steps involved in the model inference process in the above method.
[0330] Receiver 1401 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1402 can be used to output digital or character information through the first interface. Transmitter 1402 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1402 can also include a display device such as a display screen.
[0331] The present application also provides a server device. Figure 15 , Figure 15This is a structural diagram of a server provided in an embodiment of the present application. Specifically, the server 1500 is implemented by one or more servers. The server 1500 may have relatively large differences due to different configurations or performances. It may include one or more central processing units (CPU) 1515 (for example, one or more processors) and memory 1532, and one or more storage media 1530 (for example, one or more mass storage devices) for storing application programs 1542 or data 1544. Among them, the memory 1532 and the storage medium 1530 can be short-term storage or persistent storage. The program stored in the storage medium 1530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1515 can be configured to communicate with the storage medium 1530 to execute a series of instruction operations in the storage medium 1530 on the server 1500.
[0332] The server 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input and output interfaces 1558; or one or more operating systems 1541, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0333] In the embodiment of the present application, the central processing unit 1515 is used to execute the data processing method in the above embodiment.
[0334] An embodiment of the present application also provides a computer program product, which, when running on a computer, enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0335] A computer-readable storage medium is also provided in an embodiment of the present application, which stores a program for signal processing. When the computer-readable storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0336] The execution device, training device or terminal device provided in the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the data processing method described in the above embodiment, or so that the chip in the training device executes the data processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0337] For details, please refer to Figure 16 , Figure 16 This is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1600. NPU 1600 is mounted on the host CPU as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1603, which is controlled by controller 1604 to extract matrix data from memory and perform multiplication operations.
[0338] In some implementations, arithmetic circuit 1603 includes multiple processing units (PEs). In some implementations, arithmetic circuit 1603 is a two-dimensional systolic array. Arithmetic circuit 1603 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, arithmetic circuit 1603 is a general-purpose matrix processor.
[0339] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1602 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1601 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1608.
[0340] Unified memory 1606 is used to store input and output data. Weight data is directly transferred to weight memory 1602 through the Direct Memory Access Controller (DMAC) 1605. Input data is also transferred to unified memory 1606 through the DMAC.
[0341] BIU stands for Bus Interface Unit 1610 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1609 .
[0342] The bus interface unit 1610 (BIU) is used for the instruction fetch memory 1609 to obtain instructions from the external memory, and is also used for the storage unit access controller 1605 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0343] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1606 or move weight data to the weight memory 1602 or move input data to the input memory 1601.
[0344] The vector calculation unit 1607 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1603, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.
[0345] In some implementations, vector calculation unit 1607 can store the processed output vector to unified memory 1606. For example, vector calculation unit 1607 can apply a linear function or a nonlinear function to the output of operation circuit 1603, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values to generate an activation value. In some implementations, vector calculation unit 1607 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to operation circuit 1603, for example, for use in subsequent layers in a neural network.
[0346] An instruction fetch buffer 1609 connected to the controller 1604 is used to store instructions used by the controller 1604;
[0347] Unified memory 1606, input memory 1601, weight memory 1602, and instruction fetch memory 1609 are all on-chip memories. External memories are private to the NPU hardware architecture.
[0348] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.
[0349] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0350] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0351] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0352] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A data processing method, characterized in that: The method comprises: Generate second data through the first network according to the first data; Generate third data through a second network according to the first data; wherein the first network includes a first network layer, the second network includes a second network layer, and the first network layer is a partial network of the second network layer; The second data is verified according to the third data.
2. The method according to claim 1, characterized in that Before generating the second data through the first network according to the first data, the method further includes: The first data is generated through the first network according to the fourth data.
3. The method according to claim 2, characterized in that The method further comprises: Generate fifth data through the second network according to the fourth data; wherein the action of generating the fifth data and the action of generating the third data are performed in parallel; The first data is verified according to the fifth data.
4. The method according to any one of claims 1 to 3, characterized in that: The method is applied to an incremental reasoning process, wherein the first network layer and the second network layer both belong to a target neural network, and the target neural network is a model used in a full reasoning process corresponding to the incremental reasoning process.
5. The method according to claim 4, characterized in that The second network layer includes the first network layer and a third network layer connected to the first network layer, and the third network layer is closer to the output layer of the target neural network than the first network layer.
6. The method according to claim 5, characterized in that The first network includes a first network layer and a first decoding layer, and the second network includes a second network layer and a second decoding layer; The step of generating second data according to the first data through the first network includes: Obtaining a first processing result through the first network layer according to the first data, and obtaining second data through the first decoding layer according to the first processing result; The step of generating third data through a second network according to the first data includes: According to the first processing result, a second processing result is obtained through the third network layer, and according to the second processing result, third data is generated through the second decoding layer.
7. The method according to any one of claims 1 to 6, characterized in that: The verifying the second data according to the third data includes: When the third data and the second data are consistent, determining that the verification of the second data passes; When the third data and the second data are inconsistent, it is determined that the verification of the second data fails.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: When it is determined that the verification of the second data is passed, sixth data is generated according to the first data through a third network; wherein the third network includes a fourth network layer, and the second network layer is a partial network of the fourth network layer; The second data is verified according to the sixth data.
9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: When it is determined that the verification of the second data fails, generating sixth data through a third network according to the first data; wherein the third network includes a fourth network layer, and the second network layer is a partial network of the fourth network layer; The third data is verified according to the sixth data.
10. The method according to any one of claims 1 to 9, characterized in that: The method further comprises: When it is determined that the verification of the second data fails, seventh data e is generated according to the third data through the second network.
11. The method according to claim 10, characterized in that The method is applied to an incremental reasoning process, and when it is determined according to the fourth data that the verification of the second data fails, generating seventh data according to the third data through the second network, including: When the historical verification pass rate in the incremental reasoning process meets the preset condition and it is determined that the verification of the second data fails, seventh data is generated according to the third data through the second network.
12. The method according to any one of claims 1 to 11, characterized in that: The method further comprises: When it is determined that the verification of the second data fails, generating eighth data through the first network according to the third data; Generate seventh data through the second network according to the third data; The eighth data is verified according to the seventh data.
13. The method according to claim 12, characterized in that When it is determined that the verification of the second data fails, generating eighth data through the first network according to the third data includes: When the historical verification pass rate does not meet the preset condition during the incremental reasoning process and it is determined that the verification of the second data fails, eighth data is generated according to the third data through the first network.
14. The method according to any one of claims 1 to 13, characterized in that: The method is applied to an incremental reasoning process, which is a response generation process for input data; The step of generating the second data through the first network according to the first data includes: Determine sub-attention information corresponding to the length of the generated data according to the attention information in the full reasoning process corresponding to the incremental reasoning process and the length of the generated data in the reply generation process; Generate second data through the first network based on the sub-attention information and the first data.
15. The method according to any one of claims 1 to 14, characterized in that: The method is executed on a first node, and a full reasoning process corresponding to the incremental reasoning process is executed on a second node; Before generating the second data through the first network according to the first data, the method further includes: Receiving a processing result obtained through the full inference process and sent by the second node; The processing result is used to perform the incremental reasoning process.
16. The method according to any one of claims 1 to 15, characterized in that: The first data is data generated during an incremental reasoning process for a first batch, the method is applied to a first node, and when executing the method, the first node also performs an incremental reasoning process for a second batch in parallel, and the method further includes: When the incremental reasoning process for the first batch is completed and the first node has not completed the incremental reasoning process for the second batch, the data in the memory space corresponding to the incremental reasoning process of the first batch is cleared, and the incremental reasoning process of the third batch is started.
17. The method according to any one of claims 1 to 16, characterized in that: The method further comprises: When it is determined that the verification of the second data fails and the number of data that has passed the verification of the second network in the current verification cycle is less than a preset value, the second network is used to generate data through autoregression until the number of data that has passed the verification of the second network in the current verification cycle for the first batch reaches the preset value; or When it is determined that the verification of the second data fails, the first network is used to generate autoregressive data and the second network is used to verify the data generated by the autoregression of the first network until the number of data that passes the verification of the second network in the current verification cycle reaches the preset value.
18. The method according to any one of claims 1 to 17, characterized in that: The first data is a token, the second data is a token adjacent to the first data, and the third data is a token adjacent to the second data.
19. The method according to any one of claims 1 to 18, characterized in that: The first data, the second data and the third data are text data units, image data units or audio data units.
20. The method according to any one of claims 1 to 19, characterized in that: The first data is a data unit obtained in a full-amount reasoning process or a data unit generated by performing autoregressive reasoning on the data unit obtained in the full-amount reasoning process through the first network.
21. A data processing method, characterized in that: The method comprises: Get input data; According to the input data, full inference is performed through the target neural network to obtain a full inference result; According to the full reasoning result, a reply result corresponding to the input data is obtained through incremental reasoning; the incremental reasoning is implemented through a first network and a second network, the first network is used for autoregressive reasoning, the second network is used for verification reasoning, the first network includes a first network layer, the second network includes a second network layer, and the first network layer is a partial network of the second network layer.
22. A data processing device, characterized in that: The device comprises: A processing module, used to generate second data through a first network based on first data; generate third data through a second network based on the first data; wherein the first network includes a first network layer, the second network includes a second network layer, and the first network layer is a partial network of the second network layer; and verify the second data based on the third data.
23. The device according to claim 22, characterized in that Before generating the second data through the first network according to the first data, the processing module is further used to: The first data is generated through the first network according to the fourth data.
24. The device according to claim 23, characterized in that The processing module is further used for: Generate fifth data through the second network according to the fourth data; wherein the action of generating the fifth data and the action of generating the third data are performed in parallel; The first data is verified according to the fifth data.
25. The device according to any one of claims 22 to 24, characterized in that The method is applied to an incremental reasoning process, wherein the first network layer and the second network layer both belong to a target neural network, and the target neural network is a model used in a full reasoning process corresponding to the incremental reasoning process.
26. The device according to any one of claims 22 to 25, characterized in that The first network layer and the second network layer both belong to the target neural network, the second network layer includes the first network layer and a third network layer connected to the first network layer, and the third network layer is closer to the output layer of the target neural network than the first network layer.
27. The device according to any one of claims 22 to 26, characterized in that The processing module is further used for: When it is determined that the verification of the second data is passed, sixth data is generated according to the first data through a third network; wherein the third network includes a fourth network layer, and the second network layer is a partial network of the fourth network layer; The second data is verified according to the sixth data.
28. The device according to any one of claims 22 to 27, characterized in that The processing module is further used for: When it is determined that the verification of the second data fails, generating sixth data through a third network according to the first data; wherein the third network includes a fourth network layer, and the second network layer is a partial network of the fourth network layer; The third data is verified according to the sixth data.
29. The device according to any one of claims 22 to 28, characterized in that The processing module is further used for: When it is determined that the verification of the second data fails, seventh data is generated according to the third data through the second network.
30. The device according to any one of claims 22 to 29, characterized in that The processing module is further used for: When it is determined that the verification of the second data fails, generating eighth data through the first network according to the third data; Generate seventh data through the second network according to the third data; The eighth data is verified according to the seventh data.
31. The device according to claim 30, characterized in that The processing module is specifically used for: When the historical verification pass rate does not meet the preset condition during the incremental reasoning process and it is determined that the verification of the second data fails, eighth data is generated according to the third data through the first network.
32. The device according to any one of claims 22 to 31, characterized in that Applied to the incremental reasoning process, the incremental reasoning process belongs to the response generation process for input data; the processing module is specifically used to: Determine sub-attention information corresponding to the length of the generated data according to the attention information in the full reasoning process corresponding to the incremental reasoning process and the length of the generated data in the reply generation process; Generate second data through the first network based on the sub-attention information and the first data.
33. The device according to any one of claims 22 to 32, characterized in that The device is executed on a first node, and a full reasoning process corresponding to the incremental reasoning process is executed on a second node; Before generating the second data through the first network according to the first data, the processing module is further used to: Receive a processing result obtained through the full reasoning process and sent by the second node; the processing result is used to perform the incremental reasoning process.
34. The device according to any one of claims 22 to 33, characterized in that The first data is data generated during an incremental reasoning process for a first batch, the device is applied to a first node, and when executing the device, the first node also performs an incremental reasoning process for a second batch in parallel, and the processing module is further used to: When the incremental reasoning process for the first batch is completed and the first node has not completed the incremental reasoning process for the second batch, the data in the memory space corresponding to the incremental reasoning process of the first batch is cleared, and the incremental reasoning process of the third batch is started.
35. The device according to any one of claims 22 to 34, characterized in that The processing module is further used for: When it is determined that the verification of the second data fails and the number of data that has passed the verification of the second network in the current verification cycle is less than a preset value, the second network is used to generate data through autoregression until the number of data that has passed the verification of the second network in the current verification cycle for the first batch reaches the preset value; or When it is determined that the verification of the second data fails, the first network is used to generate autoregressive data and the second network is used to verify the data generated by the autoregression of the first network until the number of data that passes the verification of the second network in the current verification cycle reaches the preset value.
36. A computer storage medium, characterized in that The computer storage medium stores one or more instructions which, when executed by one or more computers, cause the one or more computers to perform the operations of the method of any one of claims 1 to 21.
37. A computer program product, characterized in that The method comprises computer-readable instructions, and when the computer-readable instructions are executed on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 21.
38. A system comprising at least one processor and at least one memory; the processor and the memory are connected via a communication bus and communicate with each other; The at least one memory is used to store code; The at least one processor is configured to execute the code to perform the method according to any one of claims 1 to 21.
39. A chip, characterized in that: The method comprises at least one processing unit and an interface circuit, wherein the interface circuit is used to provide program instructions or data to the at least one processing unit, and the at least one processing unit is used to execute the program instructions to implement the method according to any one of claims 1 to 21.
Citation Information
Cited By
Data processing method and apparatus thereof
EP4797164A1
Data processing method and apparatus thereof
WO2025103049A1