A data processing method and related device
Patent Information
- Application Number
- CN202110415349.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-18
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2041-04-18
AI Technical Summary
[0064]本申请实施例提供了一种数据处理方法,所述方法包括:获取M个第一嵌入向量、以及第二嵌入向量;其中,每个第一嵌入向量用于表示目标数据中的一个已知数据单元以及所述一个已知数据单元在所述目标数据中的第一位置,所述第二嵌入向量用于表示所述目标数据中的第一待预测数据单元在所述目标数据中的第二位置;所述M为正整数;通过目标编码器,对所述M个第一嵌入向量进行处理,以得到M个已知数据单元对应的M个第一输出向量,其中,每个所述已知数据单元对应的第一输出向量为根据所述M个第一嵌入向量生成的;通过目标预测网络,对所述M个第一输出向量以及所述第二嵌入向量进行处理,以得到所述第一待预测数据单元。通过上述方式,针对于M个已知数据单元对应的M个第一嵌入向量,目标编码器可以将M个第一嵌入向量作为输入,其中第一嵌入向量包括了位置信息和已知数据单元的数据信息,而不需要再单独设置额外的M个位置信息来作为目标编码器的输入,此外,目标编码器的中间输出的隐变量的数量也和输入的嵌入向量的数量保持一致,减少了目标编码器的计算量和内存消耗。
Smart Images

Figure CN115292439B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more particularly to a data processing method and related equipment. Background Technology
[0002] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0003] A language model is a model that can predict unknown words in a sentence based on a given semantic fragment. For example, given the natural language sequence "Huawei _ _ is very good.", a language model can generate unknown words based on that fragment. In this example, the language model can generate the word "phone" based on the given fragment, thus obtaining the sentence "Huawei phones are very good."
[0004] In existing natural language generation models (refer to) Figure 6b In this paper, an autoencoder model and an autoregressive language model are fused. Compared to the individual autoencoder and autoregressive language models, this model doubles the number of hidden states. The white portion corresponds to the autoencoder model, and the gray portion corresponds to the autoregressive language model. The latent variables related to the autoencoder model represent positional information, while the autoregressive model provides contextual information for the autoencoder model's predictions. This model consumes twice the computational and memory resources of the individual autoencoder and autoregressive models. Therefore, there is a need to provide a language model with lower computational and memory consumption. Summary of the Invention
[0005] In a first aspect, this application provides a data processing method, characterized in that the method includes: Obtain M first embedding vectors and a second embedding vector; wherein each first embedding vector is used to represent a known data unit in the target data and a first position of the known data unit in the target data, and the second embedding vector is used to represent a second position of a first data unit to be predicted in the target data; M is a positive integer; The target data refers to data with missing data. It includes both non-missing data (referred to as known data units in this embodiment) and missing data (referred to as data units to be predicted in this embodiment, such as a first data unit to be predicted and a second data unit to be predicted). Known data units are data units within the non-missing data. For example, if the target data is text data, the known data units in the target data can be known words or characters in the text data, and the data units to be predicted can be words or characters to be predicted in the text data. Similarly, if the target data is speech data, the known data units in the target data can be known audio sequences in the speech data, and the data units to be predicted can be audio sequences to be predicted in the speech data. Likewise, if the target data is image data, the known data units in the target data can be known pixels in the speech data, and the data units to be predicted can be pixels to be predicted in the speech data. It should be understood that the granularity of the known data unit and the data unit to be predicted is related to the type of the target data. The granularity of the known data unit and the data unit to be predicted can be the smallest data unit in the target data or multiple data units composed of the smallest data units. The granularity of the known data unit and the data unit to be predicted is not limited here.
[0006] The M first embedding vectors are processed by the target encoder to obtain M first output vectors corresponding to the M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors. Here, each first output vector is obtained based on M first embedding vectors. This can be understood as each first output vector can use M first embedding vectors as a reference. In other words, when generating each first output vector, each first embedding vector is visible, or each first output vector has a dependency relationship with M first embedding vectors. In one implementation, the target encoder can be a transformation layer, where each first output vector is obtained based on M first embedding vectors. This can be understood as an attention association between any two first embedding vectors among the M first embedding vectors.
[0007] The first output vector and the second embedding vector are processed by the target prediction network to obtain the first data unit to be predicted. In this embodiment of the application, for M first embedding vectors corresponding to M known data units, the target encoder can use the M first embedding vectors as input, where the first embedding vectors include the position information and data information of each known data unit, without needing to set additional M position information as input to the target encoder. In addition, the number of latent variables in the intermediate output of the target encoder is consistent with the number of input embedding vectors, reducing the computational load and memory consumption of the target encoder.
[0008] In one possible implementation, the first position is used to indicate the relative positional relationship between the known data unit and other known data units, as well as between the known data unit and the first data unit to be predicted; the second position is used to indicate the relative positional relationship between the first data unit to be predicted and each known data unit in the target data.
[0009] In one possible implementation, the target encoder is a first transformer layer and the target prediction network is a second transformer layer.
[0010] In one possible implementation, the first transformer layer includes a series of sub-transformer layers; the step of processing the M first embedding vectors through the target encoder to obtain M first output vectors corresponding to the M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M first output vectors.
[0011] In other words, the input of each sub-transformer layer includes M feature vectors corresponding to M known data units, and the output of each sub-transformer layer includes M output vectors corresponding to M known data units. This ensures that the number of latent variables in the intermediate output of the target encoder is consistent with the number of embedding vectors in the input, reducing the computational load and memory consumption of the target encoder.
[0012] In one possible implementation, the target encoder includes an attention head, and the processing of the M first embedding vectors by the target encoder includes: Obtain attention information, which is used to indicate that when the attention head processes the M first embedding vectors, there is an attention association between any two first embedding vectors among the M first embedding vectors; Based on the attention information, the M first embedding vectors are processed by the target encoder.
[0013] In one possible implementation, the method further includes: The target data is processed by an embedding layer to embed M known data units into M known data units, resulting in M third embedding vectors. This embedding layer can be called the input embedding layer. The current input can consist of M known data units. After acquiring the current input, the embedding layer performs embedding processing on each known data unit to obtain the corresponding embedding vector (i.e., the third embedding vector) for each known data unit. Obtain the position vector of each of the M known data units, the position vector being used to indicate the first position; in some embodiments, the position vector of each of the M known data units can also be obtained, the position vector being used to indicate the first position; wherein, the first position is used to represent the position of the known data unit in the target data, specifically, the first position can be used to indicate the relative positional relationship between the known data unit in the target data and other known data units besides itself and the first data unit to be predicted; Each of the M third embedding vectors is fused with its corresponding position vector to obtain the M first embedding vectors. It should be understood that the fusion method may be to perform an addition operation on the third embedding vector and the position vector, or to perform other operations so that the first embedding vector can carry a known data unit in the target data and the information of the first position of the known data unit in the target data. The specific fusion method is not limited here.
[0014] In one possible implementation, the target data further includes a second data unit to be predicted, and the order in which the second data unit to be predicted and the first data unit to be predicted are randomly determined.
[0015] In one possible implementation, if the second data unit to be predicted is predicted after the first data unit to be predicted, the method further includes: Obtain a fourth embedding vector and a fifth embedding vector, wherein the fourth embedding vector is used to represent the first data unit to be predicted and the second position of the first data unit to be predicted in the target data, and the fifth embedding vector is used to represent the third position of the second data unit to be predicted in the target data; The target encoder processes the M first embedding vectors and the fourth embedding vector to obtain M known data units and M+1 second output vectors corresponding to the first data unit to be predicted. The target prediction network processes the M+1 second output vectors and the fifth embedding vector to obtain the second data unit to be predicted.
[0016] In this embodiment, a random order method is used for prediction, which makes full use of the order information of the data units to be predicted and explicitly integrates the order information into the output vector.
[0017] In one possible implementation, the second output vector corresponding to each known data unit is generated based on the M first embedding vectors; the second output vector corresponding to the first data unit to be predicted is generated based on the M first embedding vectors and the fourth embedding vector.
[0018] In one possible implementation, the target data is text data, the known data unit is a known word in the text data, and the first data unit to be predicted is the word to be predicted in the text data; or... The target data is speech data, the known data unit is a known audio sequence in the speech data, and the first data unit to be predicted is the audio sequence to be predicted in the speech data; or... The target data is image data, the known data unit is a known pixel in the image data, and the first data unit to be predicted is a pixel to be predicted in the image data.
[0019] Secondly, this application provides a data processing method, the method comprising: Obtain M first embedding vectors and a second embedding vector; wherein each first embedding vector is used to represent a data unit in the target data and the first position of the data unit in the target data, and the second embedding vector is used to indicate the target processing task; M is a positive integer; The target encoder processes the M first embedding vectors to obtain M output vectors corresponding to the M data units, wherein the output vector corresponding to each data unit is generated based on the M first embedding vectors. The task network is used to process the M output vectors and the second embedded vector according to the target processing task to obtain the task processing result.
[0020] In one possible implementation, the first position is used to indicate the relative positional relationship between the data unit and other data units.
[0021] In one possible implementation, the target encoder is a first transformer layer and the task network is a second transformer layer.
[0022] In one possible implementation, the first transformer layer includes a series of sub-transformer layers; the step of processing the M first embedding vectors through the target encoder to obtain M first output vectors corresponding to the M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M output vectors.
[0023] In one possible implementation, the target encoder includes an attention head, and the processing of the M first embedding vectors by the target encoder includes: Obtain attention information, which is used to indicate that when the attention head processes the M first embedding vectors, there is an attention association between any two first embedding vectors among the M first embedding vectors; Based on the attention information, the M first embedding vectors are processed by the target encoder.
[0024] In one possible implementation, the target data is text data, and the data unit is a word in the text data; or, The target data is speech data, and the known data unit is an audio sequence within the speech data; or... The target data is image data, and the known data unit is a pixel in the image data.
[0025] In one possible implementation, the target processing task includes short text classification, long text classification, natural language inference, text similarity matching, or text sentiment classification.
[0026] Thirdly, this application provides a data processing method, the method comprising: A first encoder, a first prediction network, M first embedding vectors, and a second embedding vector are obtained; wherein each first embedding vector is used to represent a known data unit in the target data and a first position of the known data unit in the target data, and the second embedding vector is used to represent a second position of a first data unit to be predicted in the target data; M is a positive integer; The first encoder processes the M first embedding vectors to obtain M first output vectors corresponding to the M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors. The first prediction network processes the M first output vectors and the second embedding vector to obtain the third prediction data unit. Based on the difference between the third prediction data unit and the first prediction data unit, the first encoder and the first prediction network are updated to obtain the target encoder and the target prediction network.
[0027] In one possible implementation, the first position is used to indicate the relative positional relationship between the known data unit and other known data units, as well as between the known data unit and the first data unit to be predicted; the second position is used to indicate the relative positional relationship between the first data unit to be predicted and each known data unit in the target data.
[0028] In one possible implementation, the first encoder is a first transformer layer and the first prediction network is a second transformer layer.
[0029] In one possible implementation, the first transformer layer includes a series of sub-transformer layers; the step of processing the M first embedding vectors through the first encoder to obtain M first output vectors corresponding to M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M first output vectors.
[0030] In one possible implementation, the target data further includes a second data unit to be predicted, and the order in which the second data unit to be predicted and the first data unit to be predicted are randomly determined.
[0031] In one possible implementation, if the second data unit to be predicted is predicted after the first data unit to be predicted, the method further includes: Obtain a fourth embedding vector and a fifth embedding vector. The fourth embedding vector is used to represent the first data unit to be predicted and the second position of the first data unit to be predicted in the target data. The fifth embedding vector is used to represent the third position of the second data unit to be predicted in the target data. The first encoder processes the M first embedding vectors and the fourth embedding vector to obtain M known data units and M+1 second output vectors corresponding to the first data unit to be predicted. The first prediction network processes the M+1 second output vectors and the fifth embedding vector to obtain the fourth data unit to be predicted. The step of updating the first encoder and the first prediction network based on the difference between the third prediction data unit and the first data unit to be predicted, to obtain the target encoder and the target prediction network, includes: Based on the difference between the third prediction data unit and the first prediction data unit, and the difference between the fourth prediction data unit and the second prediction data unit, the first encoder and the first prediction network are updated to obtain the target encoder and the target prediction network.
[0032] In one possible implementation, the second output vector corresponding to each known data unit is generated based on the M first embedding vectors; the second output vector corresponding to the first data unit to be predicted is generated based on the M first embedding vectors and the fourth embedding vector.
[0033] In one possible implementation, the target data is text data, the known data unit is a known word in the text data, and the first data unit to be predicted is the word to be predicted in the text data; or... The target data is speech data, the known data unit is a known audio sequence in the speech data, and the first data unit to be predicted is the audio sequence to be predicted in the speech data; or... The target data is image data, the known data unit is a known pixel in the image data, and the first data unit to be predicted is a pixel to be predicted in the image data.
[0034] Fourthly, this application provides a data processing apparatus, comprising: The acquisition module is used to acquire M first embedding vectors and second embedding vectors; wherein each first embedding vector is used to represent a known data unit in the target data and a first position of the known data unit in the target data, and the second embedding vector is used to represent a second position of a first data unit to be predicted in the target data; M is a positive integer; The encoding module is used to process the M first embedding vectors through the target encoder to obtain M first output vectors corresponding to the M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors; The prediction module is used to process the M first output vectors and the second embedding vector through the target prediction network to obtain the first data unit to be predicted.
[0035] In one possible implementation, the first position is used to indicate the relative positional relationship between the known data unit and other known data units, as well as between the known data unit and the first data unit to be predicted; the second position is used to indicate the relative positional relationship between the first data unit to be predicted and each known data unit in the target data.
[0036] In one possible implementation, the target encoder is a first transformer layer and the target prediction network is a second transformer layer.
[0037] In one possible implementation, the first transformer layer includes a series of sub-transformer layers; the step of processing the M first embedding vectors through the target encoder to obtain M first output vectors corresponding to the M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M first output vectors.
[0038] In one possible implementation, the target encoder includes an attention head and an encoding module for acquiring attention information, the attention information being used to indicate that when the attention head processes the M first embedding vectors, there is an attention association between any two of the M first embedding vectors; Based on the attention information, the M first embedding vectors are processed by the target encoder.
[0039] In one possible implementation, the device further includes: An embedding module is used to embed M known data units in the target data through an embedding layer to obtain M third embedding vectors; Obtain the position vector of each of the M known data units, the position vector being used to indicate the first position; Each of the M third embedding vectors is fused with its corresponding position vector to obtain the M first embedding vectors.
[0040] In one possible implementation, the target data further includes a second data unit to be predicted, and the order in which the second data unit to be predicted and the first data unit to be predicted are randomly determined.
[0041] In one possible implementation, if the second data unit to be predicted is predicted after the first data unit to be predicted, the method further includes: Obtain a fourth embedding vector and a fifth embedding vector, wherein the fourth embedding vector is used to represent the first data unit to be predicted and the second position of the first data unit to be predicted in the target data, and the fifth embedding vector is used to represent the third position of the second data unit to be predicted in the target data; The target encoder processes the M first embedding vectors and the fourth embedding vector to obtain M known data units and M+1 second output vectors corresponding to the first data unit to be predicted. The target prediction network processes the M+1 second output vectors and the fifth embedding vector to obtain the second data unit to be predicted.
[0042] In one possible implementation, the second output vector corresponding to each known data unit is generated based on the M first embedding vectors; the second output vector corresponding to the first data unit to be predicted is generated based on the M first embedding vectors and the fourth embedding vector.
[0043] In one possible implementation, the target data is text data, the known data unit is a known word in the text data, and the first data unit to be predicted is the word to be predicted in the text data; or... The target data is speech data, the known data unit is a known audio sequence in the speech data, and the first data unit to be predicted is the audio sequence to be predicted in the speech data; or... The target data is image data, the known data unit is a known pixel in the image data, and the first data unit to be predicted is a pixel to be predicted in the image data.
[0044] Fifthly, this application provides a data processing apparatus, comprising: The acquisition module is used to acquire M first embedding vectors and a second embedding vector; wherein each first embedding vector is used to represent a data unit in the target data and a first position of the data unit in the target data, and the second embedding vector is used to indicate the target processing task; M is a positive integer; The encoding module is used to process the M first embedding vectors through the target encoder to obtain M output vectors corresponding to the M data units, wherein the output vector corresponding to each data unit is generated based on the M first embedding vectors; The task processing module is used to process the M output vectors and the second embedded vector through the task network to obtain the task processing result.
[0045] In one possible implementation, the first position is used to indicate the relative positional relationship between the data unit and other data units.
[0046] In one possible implementation, the target encoder is a first transformer layer and the task network is a second transformer layer.
[0047] In one possible implementation, the first transformer layer includes a series of sub-transformer layers; the step of processing the M first embedding vectors through the target encoder to obtain M first output vectors corresponding to the M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M output vectors.
[0048] In one possible implementation, the target encoder includes an attention head and an encoding module for acquiring attention information, the attention information being used to indicate that when the attention head processes the M first embedding vectors, there is an attention association between any two of the M first embedding vectors; Based on the attention information, the M first embedding vectors are processed by the target encoder.
[0049] In one possible implementation, the target data is text data, and the data unit is a word in the text data; or, The target data is speech data, and the known data unit is an audio sequence within the speech data; or... The target data is image data, and the known data unit is a pixel in the image data.
[0050] In one possible implementation, the target processing task includes short text classification, long text classification, natural language inference, text similarity matching, or text sentiment classification.
[0051] Sixthly, this application provides a data processing apparatus, comprising: The acquisition module is used to acquire a first encoder, a first prediction network, M first embedding vectors, and a second embedding vector; wherein each first embedding vector is used to represent a known data unit in the target data and a first position of the known data unit in the target data, and the second embedding vector is used to represent a second position of a first data unit to be predicted in the target data; M is a positive integer; The encoding module is used to process the M first embedding vectors through the first encoder to obtain M first output vectors corresponding to M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors; The prediction module is used to process the M first output vectors and the second embedding vector through the first prediction network to obtain a third prediction data unit. The model training module is used to update the first encoder and the first prediction network based on the difference between the third prediction data unit and the first prediction data unit to obtain the target encoder and the target prediction network.
[0052] In one possible implementation, the first position is used to indicate the relative positional relationship between the known data unit and other known data units, as well as between the known data unit and the first data unit to be predicted; the second position is used to indicate the relative positional relationship between the first data unit to be predicted and each known data unit in the target data.
[0053] In one possible implementation, the first encoder is a first transformer layer and the first prediction network is a second transformer layer.
[0054] In one possible implementation, the first transformer layer includes a series of sub-transformer layers; the step of processing the M first embedding vectors through the first encoder to obtain M first output vectors corresponding to M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M first output vectors.
[0055] In one possible implementation, the target data further includes a second data unit to be predicted, and the order in which the second data unit to be predicted and the first data unit to be predicted are randomly determined.
[0056] In one possible implementation, if the second data unit to be predicted is predicted after the first data unit to be predicted, the method further includes: Obtain a fourth embedding vector and a fifth embedding vector. The fourth embedding vector is used to represent the first data unit to be predicted and the second position of the first data unit to be predicted in the target data. The fifth embedding vector is used to represent the third position of the second data unit to be predicted in the target data. The first encoder processes the M first embedding vectors and the fourth embedding vector to obtain M known data units and M+1 second output vectors corresponding to the first data unit to be predicted. The first prediction network processes the M+1 second output vectors and the fifth embedding vector to obtain the fourth data unit to be predicted. The step of updating the first encoder and the first prediction network based on the difference between the third prediction data unit and the first data unit to be predicted, to obtain the target encoder and the target prediction network, includes: Based on the difference between the third prediction data unit and the first prediction data unit, and the difference between the fourth prediction data unit and the second prediction data unit, the first encoder and the first prediction network are updated to obtain the target encoder and the target prediction network.
[0057] In one possible implementation, the second output vector corresponding to each known data unit is generated based on the M first embedding vectors; the second output vector corresponding to the first data unit to be predicted is generated based on the M first embedding vectors and the fourth embedding vector.
[0058] In one possible implementation, the target data is text data, the known data unit is a known word in the text data, and the first data unit to be predicted is the word to be predicted in the text data; or... The target data is speech data, the known data unit is a known audio sequence in the speech data, and the first data unit to be predicted is the audio sequence to be predicted in the speech data; or... The target data is image data, the known data unit is a known pixel in the image data, and the first data unit to be predicted is a pixel to be predicted in the image data.
[0059] In a seventh aspect, embodiments of this application provide an execution device, which may include a memory, a processor, and a bus system, wherein the memory is used to store a program, and the processor is used to execute the program in the memory to perform the methods described in the first aspect and any optional methods thereof, and the second aspect and any optional methods thereof.
[0060] Eighthly, embodiments of this application provide a training device that may include a memory, a processor, and a bus system, wherein the memory is used to store a program, and the processor is used to execute the program in the memory to perform the methods described in the third aspect above and any of its optional methods.
[0061] Ninthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods described in the first aspect and any optional methods thereunder, the second aspect and any optional methods thereunder, and the third aspect and any optional methods thereunder.
[0062] In a tenth aspect, embodiments of this application provide a computer program that, when run on a computer, causes the computer to perform the methods described in the first aspect and any optional methods thereunder, the second aspect and any optional methods thereunder, and the third aspect and any optional methods thereunder.
[0063] Eleventhly, this application provides a chip system including a processor for supporting an execution device or training device in implementing the functions involved in the foregoing aspects, such as transmitting or processing data involved in the foregoing methods; or, information. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the execution device or training device. This chip system may be composed of chips or may include chips and other discrete devices.
[0064] This application provides a data processing method, comprising: acquiring M first embedding vectors and a second embedding vector; wherein each first embedding vector represents a known data unit in target data and a first position of the known data unit in the target data, and the second embedding vector represents a second position of a first data unit to be predicted in the target data; M is a positive integer; processing the M first embedding vectors through a target encoder to obtain M first output vectors corresponding to the M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors; and processing the M first output vectors and the second embedding vectors through a target prediction network to obtain the first data unit to be predicted. Through the above method, for the M first embedding vectors corresponding to the M known data units, the target encoder can use the M first embedding vectors as input, wherein the first embedding vectors include position information and data information of the known data units, without needing to separately set additional M position information as input to the target encoder. Furthermore, the number of latent variables in the intermediate outputs of the target encoder is consistent with the number of input embedding vectors, reducing the computational load and memory consumption of the target encoder. Attached Figure Description
[0065] Figure 1 A structural diagram illustrating the main framework of artificial intelligence; Figure 2 It is a natural language processing system; Figure 3a For another natural language processing system; Figure 3b This is a schematic diagram of the system structure; Figure 4 A schematic diagram of the natural language processing related devices provided in the embodiments of this application; Figure 5 This is a schematic diagram of a transformer layer architecture; Figure 6a This is an illustration of an embodiment of a data processing method provided in this application. Figure 6b This is an example of a data processing method. Figure 6c This is an illustration of an embodiment of a data processing method provided in this application. Figure 7 This is a schematic diagram of the structure of a neural network model in an embodiment of this application; Figure 8 This is a schematic diagram of a transformer layer structure; Figure 9 A schematic diagram of the operation of an attention head; Figures 10 to 19 A schematic diagram illustrating an embodiment of a data processing method provided in this application; Figure 20 A schematic diagram of the structure of a data processing apparatus provided in an embodiment of this application; Figure 21 A schematic diagram of the structure of a data processing apparatus provided in an embodiment of this application; Figure 22 A schematic diagram of the structure of a data processing apparatus provided in an embodiment of this application; Figure 23 A schematic diagram of the structure of the execution device provided in the embodiments of this application; Figure 24 A schematic diagram of the structure of the training device provided in the embodiments of this application; Figure 25 This is a schematic diagram of a chip structure provided in an embodiment of this application. Detailed Implementation
[0066] The embodiments of the present invention will now be described with reference to the accompanying drawings. The terminology used in the embodiments section is for illustrative purposes only and is not intended to limit the scope of the invention.
[0067] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0068] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0069] First, the overall workflow of the artificial intelligence system is described; please refer to [link / reference]. Figure 1 , Figure 1 The diagram illustrates a structural framework for artificial intelligence (AI). The framework is further elaborated below along two dimensions: the "Intelligent Information Chain" (horizontal axis) and the "IT Value Chain" (vertical axis). The "Intelligent Information Chain" reflects a series of processes from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT Value Chain" reflects the value that AI brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed by technology) to the industrial ecosystem of the system.
[0070] (1) Infrastructure Infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. This communication occurs through sensors; computing power is provided by intelligent chips (hardware acceleration chips such as CPUs, NPUs, GPUs, ASICs, and FPGAs); and the basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.
[0071] (2) Data The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0072] (3) Data processing Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.
[0073] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data, including symbolization and formalization.
[0074] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.
[0075] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.
[0076] (4) General ability After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0077] (5) Smart products and industry applications Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They are the encapsulation of overall artificial intelligence solutions, productizing intelligent information decision-making and realizing practical applications. Their application areas mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.
[0078] This application can be applied to the fields of natural language processing, image processing, and audio-visual processing in the field of artificial intelligence. The following will take natural language processing as an example to introduce several application scenarios that have been implemented in products.
[0079] To better understand the solutions of the embodiments of this application, the following will first combine... Figures 1 to 3a A brief introduction to the possible application scenarios of the embodiments of this application is provided.
[0080] Figure 2 A natural language processing (NLP) system is illustrated, comprising user devices and data processing devices. The user devices include smart terminals such as mobile phones, personal computers, or information processing centers. The user devices are the initiators of natural language data processing, acting as the initiators of requests such as language question answering or queries; typically, users initiate requests through their user devices.
[0081] The aforementioned data processing equipment can be cloud servers, network servers, application servers, management servers, or other devices or servers with data processing capabilities. The data processing equipment receives queries / voice / text from smart terminals via an interactive interface, then performs language data processing through a storage device and a data processing processor, employing methods such as machine learning, deep learning, search, reasoning, and decision-making. The processing results are then fed back to the user device. The storage device in the data processing equipment can be a general term, including local storage and a database storing historical data. The database can be located on the data processing equipment or on other network servers.
[0082] exist Figure 2 In the natural language processing system shown, the user device can receive instructions from the user. For example, the user device can receive a piece of text input by the user and then send a request to the data processing device, so that the data processing device can perform natural language processing applications (such as natural language generation, text classification, text reasoning, named entity recognition, translation, etc.) on the piece of text obtained by the user device, thereby obtaining the processing results of the corresponding natural language processing applications on the piece of text (such as word prediction results, classification results, reasoning results, named entity recognition results, translation results, etc.).
[0083] Taking natural language generation (NLP) as an example, NLP, also known as text prediction or natural language synthesis, refers to the task of generating missing or subsequent text given a given text. NLP is widely used in search engines, input methods, and other scenarios. It can predict a user's next input based on partial input, greatly improving the efficiency of using the product. Furthermore, it can recover missing text. For example, in this embodiment, the user device can receive a piece of text data input by the user (e.g., the target data described in this embodiment). The text data includes known words and words to be predicted. The words to be predicted are not visible; only their position in the text data is known. The user device can then send a request (carrying the text data) to the data processing device, causing the data processing device to predict the words to be predicted in the text data, thereby obtaining the words to be predicted and feeding them back to the user device.
[0084] For example, a user equipment can receive a piece of text data input by a user, and then send a request to a data processing device to enable the data processing device to perform entity classification on the piece of text data, thereby obtaining the entity classification result for the piece of text data, and feeding the entity classification result back to the user equipment. For example, a user device can receive a piece of text data (the text data is Chinese text) input by a user, and then send a request to a data processing device to translate the text data into English, thereby obtaining an English translation of the text data, and then sending the English translation back to the user device.
[0085] exist Figure 2 In this application, the data processing device can process the above-mentioned text data using the data processing method of the embodiments of this application.
[0086] Figure 3a This demonstrates another natural language processing system, in Figure 3a In this context, the user equipment (UE) directly functions as a data processing device. This UE can directly receive input from the user and process it directly through its own hardware. The specific process is similar to... Figure 2 Similar to the description above, it will not be repeated here.
[0087] Figure 4 This is a schematic diagram of the natural language processing related device 300 provided in the embodiments of this application.
[0088] The above Figure 2 and Figure 3a The user equipment in the context can specifically be Figure 4 Local device 301 or local device 302 in the system. Figure 2 The data processing equipment in the middle can specifically be Figure 4 The execution device 310 in the process includes a data storage system 350 that can store the data to be processed by the execution device 310. The data storage system 350 can be integrated into the execution device 310 or set up in the cloud or on other network servers.
[0089] Figure 2 and Figure 3a The processor in the application can perform data training / machine learning / deep learning using neural network models or other models, and use the models finally trained or learned from the data (such as the target encoder, target prediction network, task network, etc. in the embodiments of this application) to perform natural language processing applications (such as natural language generation, text classification, sequence labeling, reading comprehension, text generation, text reasoning, translation, etc.) on text data (such as the target data described in the embodiments of this application) to obtain the corresponding processing results (such as the first data unit to be predicted, the second data unit to be predicted, and the task processing results, etc. in the embodiments of this application).
[0090] It should be understood that the embodiments of this application can also be applied to the fields of image processing and audio-visual processing, in which case the above-mentioned data processing device processes the target data using the data processing method of the embodiments of this application.
[0091] It should be understood that the above-mentioned data processing device may also be referred to as a data processing apparatus, execution device, server, terminal device, etc. in subsequent embodiments.
[0092] The following is combined Figure 3b The system architecture provided in the embodiments of this application will be described in detail. Figure 3b This is a schematic diagram of a system architecture provided for an embodiment of this application. For example... Figure 3b As shown, the system architecture 500 includes an execution device 510, a training device 520, a database 530, a client device 540, a data storage system 550, and a data acquisition device 560.
[0093] The execution device 510 includes a calculation module 511, an I / O interface 512, a preprocessing module 513, and a preprocessing module 514. The calculation module 511 may include a target model / rule 501, while the preprocessing modules 513 and 514 are optional.
[0094] Data acquisition device 560 is used to collect training data. Specifically, in the natural language synthesis task, the training data can be text data with missing text and the corresponding complete text data; in the audio synthesis task, the training data can be speech data with missing audio sequences and the corresponding complete speech data; in the image synthesis (or image reconstruction) task, the training data can be image or video data with missing pixels and the corresponding complete image or video data. After collecting the training data, data acquisition device 560 stores this training data in database 530, and training device 520 trains the target model / rule 501 based on the training data maintained in database 530.
[0095] Taking the target model / rule 501 for implementing a natural language synthesis task as an example, the target model / rule 501 (e.g., the target encoder and target prediction network in the embodiments of this application) can be used to implement a natural language synthesis task. That is, by inputting text data with missing text into the target model / rule 501, the missing text (e.g., the first data unit to be predicted and the second data unit to be predicted in the embodiments of this application) can be obtained.
[0096] Taking the target model / rule 501 as an example of implementing target processing tasks (such as short text classification, long text classification, natural language inference, text similarity matching, text sentiment classification, etc.), the target model / rule 501 (such as the target encoder and task network in the embodiments of this application) can be used to implement target processing tasks. That is, by inputting the target data into the target model / rule 501, the task processing result can be obtained.
[0097] It should be noted that in practical applications, the training data maintained in database 530 may not all come from the data acquisition device 560; it may also be received from other devices. Furthermore, it should be noted that training device 520 may not necessarily train the target model / rule 501 entirely based on the training data maintained in database 530; it may also obtain training data from the cloud or other sources for model training. The above description should not be construed as limiting the embodiments of this application.
[0098] The target model / rule 501 trained using training device 520 can be applied to different systems or devices, such as... Figure 3b The execution device 510 shown can be a terminal, such as a mobile phone terminal, tablet computer, laptop computer, augmented reality (AR) / virtual reality (VR) device, vehicle terminal, etc., or it can be a server or cloud, etc. Figure 3b In the process, the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with external devices. Users can input data (such as target data in the embodiments of this application) into the I / O interface 512 through the client device 540.
[0099] Preprocessing modules 513 and 514 are used to preprocess the input data received from I / O interface 512 (e.g., obtaining the positions of known data units and data units to be predicted in the target data, or generating attention information, etc.). It should be understood that preprocessing modules 513 and 514 may be absent, or only one preprocessing module may be used. When preprocessing modules 513 and 514 are absent, the calculation module 511 can be used directly to process the input data.
[0100] During the preprocessing of input data by the execution device 510, or during the calculation module 511 of the execution device 510 performing calculations and other related processes, the execution device 510 can call data, code, etc. in the data storage system 550 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing into the data storage system 550.
[0101] Finally, the I / O interface 512 presents the processing results, such as the missing text, missing audio sequence, and missing pixels (e.g., the first data unit to be predicted, the second data unit to be predicted, and the task processing result in this embodiment) to the client device 540, thereby providing them to the user.
[0102] exist Figure 3bIn the illustrated scenario, the user can manually provide input data, which can be done through the interface provided by I / O interface 512. Alternatively, the client device 540 can automatically send input data to I / O interface 512. If user authorization is required for the client device 540 to automatically send input data, the user can set the corresponding permissions in the client device 540. The user can view the output results of the execution device 510 on the client device 540, which can be presented in various forms such as display, sound, or animation. The client device 540 can also act as a data acquisition terminal, collecting the input data and output results of the input I / O interface 512 as shown in the figure, and storing them as new sample data in database 530. Alternatively, data can be collected directly from the I / O interface 512 without going through the client device 540, using the input data and output results of the input I / O interface 512 as shown in the figure, and storing them as new sample data in database 530.
[0103] It is worth noting that, Figure 3b This is merely a schematic diagram of a system architecture provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc., shown in the diagram do not constitute any limitation. For example, in Figure 3b In this context, the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 may also be placed within the execution device 510.
[0104] It should be understood that the aforementioned execution device 510 can also be deployed in customer device 540.
[0105] From the perspective of model reasoning, in this embodiment of the application, the data storage system 550 may store code related to implementing the data processing method in this embodiment of the application. The calculation module 511 can obtain the code related to implementing the data processing method in this embodiment of the application from the data storage system 550 to execute the data processing method in this embodiment of the application.
[0106] In this embodiment of the application, the computing module 511 may include hardware circuits (such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), general-purpose processors, digital signal processors (DSPs), microprocessors or microcontrollers, etc.) or combinations of these hardware circuits. For example, the computing module 511 may be a hardware system with instruction execution capabilities, such as a CPU or DSP, or a hardware system without instruction execution capabilities, such as an ASIC or FPGA, or a combination of the aforementioned hardware systems without instruction execution capabilities and hardware systems with instruction execution capabilities.
[0107] Specifically, the computing module 511 can be a hardware system with instruction execution capabilities. The data processing method provided in this application embodiment can be software code stored in the data storage system 550. The computing module 511 can obtain the software code from the data storage system 550 and execute the obtained software code to implement the data processing method provided in this application embodiment.
[0108] It should be understood that the computing module 511 can be a combination of a hardware system without the function of executing instructions and a hardware system with the function of executing instructions. Some steps of the data processing method provided in the embodiments of this application can also be implemented by the hardware system without the function of executing instructions in the computing module 511 or by the preprocessing module 513 or the preprocessing module 514. This is not limited here.
[0109] Since the embodiments of this application involve a large number of neural network applications, for ease of understanding, the relevant terms and concepts such as neural networks involved in the embodiments of this application will be introduced below.
[0110] (1) Neural Network A neural network can be composed of neural units, which can be defined as a computational unit that takes xs (i.e., input data) and an intercept of 1 as input. The output of this computational unit can be: ; Where s = 1, 2, ..., n, where n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer, and the activation function can be the sigmoid function. A neural network is a network formed by connecting multiple of the above-mentioned individual neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, which can be a region composed of several neural units.
[0111] (2) Transformer layer Reference Figure 5 , Figure 5 This is a schematic diagram of a transformer layer architecture, such as Figure 5 As shown, the neural network includes an embedding layer and at least one transformer layer. The at least one transformer layer can be N transformer layers (N being an integer greater than 0). Each transformer layer includes sequentially adjacent attention layers, add and normalize layers, feed-forward layers, and add and normalize layers. In the embedding layer, the current input is embedded to obtain multiple embedding vectors. In the attention layer, P input vectors are obtained from the layer above the first transformer layer. Using any first input vector among the P input vectors as the center, based on the correlation between each input vector within a preset attention window and the first input vector, an intermediate vector corresponding to the first input vector is obtained. This process determines P intermediate vectors corresponding to the P input vectors. In the pooling layer, the P intermediate vectors are merged into Q output vectors, where the multiple output vectors obtained from the last transformer layer are used as feature representations of the current input.
[0112] (3) Attention mechanism Attention mechanisms mimic the internal processes of biological observation—aligning internal experience with external senses to increase the precision of observation in specific areas. They enable the rapid sifting of high-value information from a large volume of data using limited attentional resources. Attention mechanisms can quickly extract important features from sparse data and are therefore widely used in natural language processing tasks, particularly machine translation. Self-attention mechanisms, an improvement on attention mechanisms, reduce reliance on external information and are better at capturing the internal correlations of data or features. The core idea of attention mechanisms can be rewritten as follows: In this formula, Lx = ||Source|| represents the length of the Source. The meaning is that the elements in the Source are imagined as a series of data pairs. Given a Query element in the Target, the similarity or relevance between the Query and each Key is calculated to obtain the weight coefficient of the Value corresponding to each Key. Then, the Values are weighted and summed to obtain the final Attention value. Therefore, the Attention mechanism essentially performs a weighted sum of the Values of the elements in the Source, while the Query and Key are used to calculate the weight coefficients of their corresponding Values. Conceptually, Attention can be understood as selectively filtering a small amount of important information from a large amount of information and focusing on this important information, ignoring most of the unimportant information. The focusing process is reflected in the calculation of the weight coefficients; the larger the weight, the more focused it is on its corresponding Value. That is, the weight represents the importance of the information, and the Value is the corresponding information. Self-attention can be understood as intra attention. The attention mechanism occurs between the elements of the Target (Query) and all elements of the Source. Self-attention refers to the attention mechanism that occurs between elements within the Source or between elements within the Target. It can also be understood as the attention calculation mechanism in the special case where Target = Source. The specific calculation process is the same, only the calculation object changes.
[0113] (4) Natural Language Processing (NLP) Natural language is human language, and Natural Language Processing (NLP) is the processing of human language. NLP is a systematic process of analyzing, understanding, and extracting information from text data in an intelligent and efficient manner. By using NLP and its components, we can manage very large amounts of text data, perform numerous automated tasks, and solve a wide variety of problems, such as automatic summarization, machine translation (MT), named entity recognition (NER), relation extraction (RE), information extraction (IE), sentiment analysis, speech recognition, question answering systems, and topic segmentation, among others.
[0114] (5) Pre-trained language model A pre-trained language model is a natural language sequence encoder that encodes each word in a natural language sequence into a vector representation for prediction tasks. Its training consists of two phases. In the pre-training phase, the model is trained on a large-scale unsupervised text environment to learn word representations. In the fine-tuning phase, the model is initialized using the parameters learned in the pre-training phase and then trained on downstream tasks such as text classification and sequence labeling with fewer steps, successfully transferring the semantic information obtained in pre-training to downstream tasks.
[0115] (6) Autoregressive language model Autoregressive language models are models that can predict the next word that might follow (e.g., "good") based on a given context (e.g., "the phone is good"). Typically, these models predict the word in the context on the right given the preceding text on the left, but they can also predict a word in the middle given the context on both the left and right sides.
[0116] First, the data processing method provided in this application embodiment will be explained using the model inference stage as an example.
[0117] Reference Figure 6a , Figure 6a This is an illustrative embodiment of a data processing method provided in this application. The data processing method provided in this application can be applied to the data processing device and execution device described above. Specifically, the data processing method can be applied to terminal devices such as mobile phones, tablets, laptops, and smart wearable devices, or to cloud-based servers, such as... Figure 6a As shown, an embodiment of this application provides a data processing method, including: 601. Obtain M first embedding vectors and a second embedding vector; wherein each first embedding vector is used to represent a known data unit in the target data and the first position of the known data unit in the target data, and the second embedding vector is used to represent the second position of the first data unit to be predicted in the target data; M is a positive integer.
[0118] The target data refers to data with missing data. It includes both non-missing data (referred to as known data units in this embodiment) and missing data (referred to as data units to be predicted in this embodiment, such as a first data unit to be predicted and a second data unit to be predicted). Known data units are data units within the non-missing data. For example, if the target data is text data, the known data units in the target data can be known words or characters in the text data, and the data units to be predicted can be words or characters to be predicted in the text data. Similarly, if the target data is speech data, the known data units in the target data can be known audio sequences in the speech data, and the data units to be predicted can be audio sequences to be predicted in the speech data. Likewise, if the target data is image data, the known data units in the target data can be known pixels in the speech data, and the data units to be predicted can be pixels to be predicted in the speech data. It should be understood that the granularity of the known data unit and the data unit to be predicted is related to the type of the target data. The granularity of the known data unit and the data unit to be predicted can be the smallest data unit in the target data or multiple data units composed of the smallest data units. The granularity of the known data unit and the data unit to be predicted is not limited here.
[0119] Specifically, in the embodiments of this application, the target data may include M known data units and at least one data unit to be predicted (including a first data unit to be predicted), wherein the data unit to be predicted is invisible data in the target data and needs to be determined by the M known data units.
[0120] Taking text data as the target data as an example, in this embodiment of the application, the text data may include M known words and at least one word to be predicted (including the first word to be predicted). The text data may be Chinese text, English text, or text in other languages, and may be sentences, paragraphs, chapters, etc.
[0121] For example, the target data could be “_ _ sat on the mat”, where “sat”, “on”, “the”, and “mat” are known data units, and “_” and “_” are not visible in the target data and are data units to be predicted. It should be understood that the symbol “_” here means empty, not an underscore.
[0122] In this embodiment of the application, M first embedding vectors can be obtained; wherein each first embedding vector is used to represent a known data unit in the target data and the first position of the known data unit in the target data.
[0123] Next, we will first describe how to generate M first embedding vectors: In one implementation, M known data units in the target data can be embedded through an embedding layer to obtain M third embedding vectors.
[0124] The embedding layer can be called the input embedding layer. The current input can be M known data units. After obtaining the current input, the embedding layer can perform embedding processing on each known data unit in the current input to obtain the embedding vector (that is, the third embedding vector) corresponding to each known data unit.
[0125] In some embodiments, the position vector of each of the M known data units can also be obtained, and the position vector is used to indicate the first position; wherein, the first position is used to indicate the position of the known data unit in the target data, specifically, the first position is used to indicate the relative positional relationship between the known data unit and other known data units and between the known data unit and the first data unit to be predicted.
[0126] In one implementation, the embedding layer may include an input embedding layer and a positional encoding layer. In the input embedding layer, word embedding processing can be performed on each known data unit in the current input to obtain a third embedding vector for each known data unit. In the positional encoding layer, the position of each known data unit in the current input can be obtained, and a position vector can be generated for each known data unit.
[0127] In some examples, the first position of each known data unit in the target data may be the absolute position of each known data unit in the target data. Taking the current input "when should I pay off Huabei" as an example, the position of "when" can be represented as the first position, the position of the auxiliary word can be represented as the second position, ... In some examples, the first position of each known data unit in the target data may be the relative position of each known data unit in the target data. Still taking the current input "when should I pay off Huabei" as an example, the position of "when" can be represented as before the auxiliary word, the position of the auxiliary word can be represented as after "when" and before "should", ... When the third embedding vectors and position vectors of each known data unit in the current input are obtained, the position vectors of each known data unit and the corresponding third embedding vectors can be fused to obtain the first embedding vector of each known data unit, that is, a plurality of first embedding vectors corresponding to the current input are obtained. It should be understood that the fusion method may be an addition operation on the third embedding vector and the position vector, or other operations may enable the first embedding vector to carry information of a known data unit in the target data and information of the first position of the known data unit in the target data, and the specific fusion method is not limited herein. The plurality of first embedding vectors may be represented as an embedding matrix with a preset dimension. If the number of the plurality of first embedding vectors is set to M and the preset dimension is H-dimensional, the plurality of first embedding vectors may be represented as an M×H embedding matrix.
[0128] In the embodiments of the present application, a second embedding vector may be obtained, where the second embedding vector is used to represent a second position of a first to-be-predicted data unit in the target data in the target data, and the second position may be used to indicate a relative positional relationship between the first to-be-predicted data unit and each known data unit in the target data.
[0129] Next, how to generate the second embedding vector is described: In one implementation, an embedding layer may be used to perform embedding processing on the second position of the first to-be-predicted data unit in the target data, so as to obtain a second embedding vector that represents the second position of the first to-be-predicted data unit in the target data, and the second embedding vector can be used as an input of a subsequent target prediction network. Wherein, the second position is used to indicate the relative positional relationship between the first to-be-predicted data unit and each known data unit in the target data. For the description of the second position, reference may be made to the description of the first position in the above embodiments, and similarities will not be repeated herein.
[0130] Further, M first embedding vectors for M known data units and a second embedding vector for the first to-be-predicted data unit can be obtained.
[0131] 602. The M first embedding vectors are processed by the target encoder to obtain M first output vectors corresponding to the M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors.
[0132] In this embodiment of the application, the target encoder can process M first embedding vectors to obtain M first output vectors corresponding to M known data units, that is, it can obtain a first output vector corresponding to each known data unit.
[0133] In existing natural language generation models (refer to...) Figure 6b In this model, autoencoder and autoregressive language models are fused together. Compared with the autoencoder and autoregressive language models, this model doubles the number of hidden states. The white part corresponds to the autoencoder model, and the gray part corresponds to the autoregressive language model. The latent variables related to the autoencoder model are used to represent positional information, and the autoregressive model is used to provide contextual information for the prediction of the autoencoder language model. The computational cost and memory consumption of this model are twice that of the autoencoder and autoregressive models.
[0134] In this embodiment, during the process of the target encoder processing M first embedding vectors, the number of hidden states is consistent with the number of hidden states in the autoencoder and autoregressive language models. Specifically, for the M first embedding vectors corresponding to the M known data units, the target encoder can use the M first embedding vectors as input, where the first embedding vectors include position information and data information of the known data units, without needing to set additional M position information as input to the target encoder. In addition, the number of hidden variables in the intermediate output of the target encoder is also consistent with the number of input embedding vectors, reducing the computational load and memory consumption of the target encoder.
[0135] For details, please refer to Figure 6c The target encoder takes M first embedding vectors as input and outputs M first output vectors.
[0136] In this embodiment of the application, each first output vector is obtained based on M first embedding vectors.
[0137] The statement that each first output vector is obtained based on M first embedding vectors can be understood as each first output vector having the M first embedding vectors as references. In other words, when generating each first output vector, each first embedding vector is visible, or each first output vector has a dependency relationship with the M first embedding vectors.
[0138] In one implementation, the target encoder can be a first transformation transformer layer, and each first output vector is obtained based on M first embedding vectors. This can be understood as an attention association between any two first embedding vectors among the M first embedding vectors.
[0139] Among them, reference Figure 7 The first transformer layer may include multiple sub-transformer layers in sequence. The N sub-transformer layers include adjacent first sub-transformer layers and second sub-transformer layers. That is, the first sub-transformer layer and the second sub-transformer layer can be any two adjacent sub-transformer layers in the first transformer layer.
[0140] Each sub-transformer layer can process the output data of the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and output the M intermediate vectors to the next sub-transformer layer adjacent to it; wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M first output vectors.
[0141] In other words, the input of each sub-transformer layer includes M feature vectors corresponding to M known data units, and the output of each sub-transformer layer includes M output vectors corresponding to M known data units. This ensures that the number of latent variables in the intermediate output of the target encoder is consistent with the number of embedding vectors in the input, reducing the computational load and memory consumption of the target encoder.
[0142] In other words, the input of each sub-transformer layer includes M feature vectors corresponding to M known data units, and the output of each sub-transformer layer includes M output vectors corresponding to M known data units. This ensures that the number of latent variables in the intermediate output of the target encoder is consistent with the number of embedding vectors in the input, reducing the computational load and memory consumption of the target encoder.
[0143] The core feature of the transformer layer lies in its unique attention mechanism. When processing natural language, such as a sentence, the transformer model uses this attention mechanism to assign different attention coefficients to the embedding vectors of each word in the sentence, thus more comprehensively considering the influence of the context on each word. A specific transformer layer may include sequentially adjacent multi-head attention layers, add and normalize layers, feed-forward layers, and add and normalize layers. The attention layer is connected to the embedding layer, obtaining M embedding vectors as input vectors. Based on the correlation between the M embedding vectors, it synthesizes the embedding vectors to obtain M output vectors, which are then output to subsequent transformer layers. The transformer layer takes the output of the previous layer as its input vector and performs similar operations as the previous transformer layer.
[0144] Reference Figure 8 , Figure 8 This is a schematic diagram of a transformer layer structure. Each sub-transformer layer in the embodiments of this application can be referred to. Figure 8 The structure shown in the figure, such as Figure 8 As shown, the transformer layer consists of a multi-head attention layer, an add & normalization layer, a feed forward layer, and another add & normalization layer, which are sequentially adjacent to each other.
[0145] The multi-head attention layer obtains M input vectors X from the layer above it. l This can also be represented as matrix X. Employing a self-attention mechanism, it transforms each vector based on the correlation between them, resulting in M output vectors, which can also be represented as matrix Y. It can be understood that when this multi-head attention layer is directly connected to the embedding layer, for example... Figure 7 In a transformer layer directly connected to the embedding layer, the input vector it receives is the embedding vector output by the embedding layer; when this multi-head attention layer is a multi-head attention layer included in subsequent transformer layers, for example... Figure 7 The transformer layer directly connected to the previous transformer layer includes a multi-head attention layer, whose input vector is the output vector of the previous transformer layer. In the multi-head attention layer, the MHA layer includes multiple attention heads (e.g., ...). Figure 8 The following are Head 1, Head 2, ..., Head N shown in the figure.
[0146] Figure 9 This is a schematic diagram illustrating the operation of an attention head, showing how the attention head transforms an input matrix X into an output matrix Y. For example... Figure 9 As shown, the first transformation matrix Q, the second transformation matrix K, and the third transformation matrix V are respectively applied to the M input vectors.<X1 ,X2 ,… ,XN> The input vectors Xi are transformed to obtain the first intermediate vector (q vector), second intermediate vector (k vector), and third intermediate vector (v vector) corresponding to each input vector. Operationally, the input matrix X composed of N input vectors can be linearly transformed using the first transformation matrix Q, the second transformation matrix K, and the third transformation matrix V, respectively, to obtain the Q matrix, K matrix, and V matrix of the input matrix. Then, the matrices are split to obtain the q vector, k vector, and v vector corresponding to each input vector. For any i-th input vector Xi among the M input vectors, the correlation degree between the i-th input vector Xi and each input vector Xj is determined based on the dot product operation of the first intermediate vector (q vector, qi) corresponding to the i-th input vector and the second intermediate vector (k vector, kj) corresponding to each input vector Xj. Although the dot product result of qi and kj can be directly determined as the correlation degree, a more classic approach is to first divide the dot product result by a constant, then perform a softmax operation, and use the result as the correlation degree between the input vector Xi and Xj, that is: ; Therefore, the correlation degrees αi,j between the i-th input vector Xi and each input vector Xj can be used as weighting factors to perform a weighted combination of the third intermediate vectors (v vector, vj) corresponding to each input vector Xj, resulting in the i-th combined vector Ci corresponding to the i-th input vector Xi: ; Therefore, we can obtain a vector sequence of M combined vectors corresponding to the M input vectors.<C1 ,C2 ,… ,CN> Or matrix C. Based on this combined vector sequence, M output vectors can be obtained. Specifically, in one embodiment, the vector sequence of N combined vectors can be directly used as the M output vectors, i.e., Yi = Ci. In this case, the output matrix Y is the combined vector matrix C, which can also be written as: ; The above describes the processing flow of an attention head. In the MHA architecture, the MHA layer maintains m sets of transformation matrices. Each set of transformation matrices includes the aforementioned first transformation matrix Q, second transformation matrix K, and third transformation matrix V, allowing the above operations to be performed in parallel to obtain m combined vector sequences (i.e., m matrices C). Each vector sequence includes N combined vectors obtained based on a set of transformation matrices. In this case, the MHA layer concatenates the m combined vector sequences to obtain a concatenated matrix; then, it transforms this concatenated matrix using the fourth transformation matrix W to obtain the final output matrix Y. This output matrix Y is then split into M output vectors.<Y1 ,Y2 ,…,YN> Through the above operations, the MHA layer performs transformation operations based on the correlation between N input vectors to obtain M output vectors.
[0147] like Figure 8 As shown, a transformer layer may include a feedforward layer, which comprises an input layer, an intermediate layer, and an output layer. As previously mentioned, a neural network model may contain multiple transformer layers. In one embodiment, these multiple transformer layers may be stacked and connected in a residual network manner.
[0148] In this embodiment, the target encoder includes an attention head. Since the known data units in the target data are mutually visible, when processing the M first embedding vectors, there is an attentional association between any two of the M first embedding vectors. Specifically, attention information can be obtained, which is used to indicate that there is an attentional association between any two of the M first embedding vectors when the attention head processes them. Then, based on the attention information, the target encoder can process the M first embedding vectors, thereby making each output vector dependent on the M first embedding vectors.
[0149] 603. The first output vectors and the second embedding vector are processed by the target prediction network to obtain the first data unit to be predicted.
[0150] In this embodiment, after obtaining M output vectors, the M output vectors can be input into a target prediction network. The target prediction network then processes the M first output vectors and the second embedding vector to obtain the first data unit to be predicted. The target prediction network can be a transformer layer.
[0151] The target prediction network can take M first output vectors and the second embedding vector as input to obtain the vector representation of the first data unit to be predicted. It should be understood that the vector representation of the first data unit to be predicted can be reconstructed by a classifier (e.g., support vector machine, softmax classifier, K-nearest neighbor algorithm, etc.).
[0152] Taking text data as an example, in the data processing of the target prediction network, the first word to be predicted can see its corresponding position vector (second embedding vector) and each known word (first embedding vector). Then, the target prediction network can take M first output vectors and second embedding vectors as input to obtain the word vector representation of the first word to be predicted.
[0153] For example, given that the words at positions 3 to 6 in the target data are "sat on the mat", and the goal is to predict the first two words in the sentence, the target prediction network can first determine the first word to be predicted as "that" based on the four input vectors corresponding to "sat on the mat" and the prediction position 1. Similarly, the target prediction network then predicts the word at position 2 based on "that _ sat on the mat".
[0154] In this embodiment of the application, the target data further includes a second data unit to be predicted. Before processing the M first embedding vectors through the target encoder, the prediction order of the first data unit to be predicted and the second data unit to be predicted can be randomly determined. If the prediction order is used to indicate that the second data unit to be predicted is predicted after the first data unit to be predicted, then after obtaining the first data unit to be predicted, a fourth embedding vector and a fifth embedding vector can be obtained. The fourth embedding vector is used to represent the first data unit to be predicted and its second position in the target data, and the fifth embedding vector is used to represent the third position of the second data unit to be predicted in the target data. The target encoder processes the M first embedding vectors and the fourth embedding vector to obtain M known data units and M+1 second output vectors corresponding to the first data unit to be predicted. The target prediction network processes the M+1 second output vectors and the fifth embedding vector to obtain the second data unit to be predicted.
[0155] The second output vector corresponding to each known data unit is generated based on the M first embedding vectors; the second output vector corresponding to the first data unit to be predicted is generated based on the M first embedding vectors and the fourth embedding vector.
[0156] The following description uses text data as an example to illustrate the data processing method in this application: Reference Figure 10 The preprocessing module can process the input word vector sequence and input the processing results into the autoregressive word vector encoding module and the query module. The output results of the autoregressive word vector encoding module and the query module can be input into the prediction module, and the prediction module can output the predicted label.
[0157] In this embodiment, the autoregressive word vector encoding module can be the target encoder in the above embodiment, the query module is used to generate the second embedding vector and the fifth embedding vector, the prediction module can be the target prediction network in the above embodiment, and the prediction label can be the data unit to be predicted in the above embodiment.
[0158] Reference Figure 11 The preprocessing module performs sequence rearrangement, block segmentation, and information extraction on the input word vector sequence. Sequence rearrangement reconstructs the order of the input word vectors, and reconstruction methods include, but are not limited to, maintaining the original order, randomizing the order, and reversing the order. Sequence segmentation divides the rearranged sentence into two blocks. In subsequent modeling, the information of each word in the first block is visible to all words, while the information of each word in the second block is only visible to the words following it after the rearrangement. The information extraction module extracts three parts of information based on the sequence rearrangement: the rearranged word vector sequence, the attention matrix (which defines which words in the sentence are visible during the modeling process of the autoregressive word vector encoding module for each word's vector representation; this matrix is obtained based on block segmentation and can be the attention information in the above embodiment), and auxiliary information for the predicted label (i.e., the second or fifth embedding vector in the above embodiment). The auxiliary information defines the position information of the predicted word in the original sentence. The first two parts of information are output to the autoregressive word vector encoding module, and the third part is output to the query module. It should be understood that the operation of the preprocessing module is manually defined and does not contain any learnable components.
[0159] The autoregressive word vector encoding module can learn the context information corresponding to each word, and finally learn a word vector sequence containing its context information for each word in the sentence (that is, the output vector in the above embodiment).
[0160] The autoregressive word vector encoding module can, for example, Figure 12As shown in the left figure, the module includes several layers of autoregressive word vector encoders (each layer can be a sub-transformer layer in the above embodiment). Each layer receives the word vectors output by the previous layer and calculates the dependencies between word vectors, integrating the context information of each word vector into the output word vector. Figure 12 The right figure shows the computation process of the i-th layer of the autoregressive word vector encoder. Each box in the figure represents a word vector. The bottom row represents the word vectors input to the layer, and the top row represents the word vectors output by the layer. Each arrow ai represents the dependency of each output word vector on the input word vector. Whether this dependency exists can be determined by the attention matrix.
[0161] Reference Figure 13 After obtaining the word vector sequence containing its context information and the query information, the word vector sequence containing its context information and the query information can be input into the prediction module. The prediction module will then make a prediction based on the word vector sequence containing its context information and the query information to obtain the predicted label.
[0162] Figure 14 This demonstrates an example of random prediction order. The original sentence is "the cat sat on the mat". The preprocessing module randomly rearranges the original sentence, resulting in the sequence "the on mat sat the cat". The rearranged sentence is then divided into blocks: the first block is "the on mat sat", where any word is visible to all other words in the sentence; the second block is "the cat", where any word is only visible to the words following it (within the rearranged sequence). The words in the second block will be predicted in this example, and because the original sentence was randomly rearranged, the prediction order of the words in the second block is random. The module obtains an attention matrix based on the sentence blocks. If the element in the i-th row and j-th column of this matrix is 1 (white), it indicates that in subsequent modeling, the j-th word in the rearranged sequence is visible to the i-th word; otherwise, it is invisible. This module outputs the rearranged word vector sequence (M first embedding vectors) and the attention matrix (i.e., the attention information in the above embodiment) to the autoregressive word vector encoding module (i.e., the target encoder in the above embodiment). It then outputs the auxiliary information of the label to be predicted (i.e., the second position indicated by the second embedding vector and the third position indicated by the fifth embedding vector in the above embodiment) to the autoregressive word vector encoding module (i.e., the second position indicated by the second embedding vector and the third position indicated by the fifth embedding vector in the above embodiment). Figure 14 In the example, the auxiliary information is position 1 and position 2, which means that the model will predict the words at positions 1 and 2 in the original sequence (the and cat, respectively) and output them to the query module. The query module can generate a second embedding vector and a fifth embedding vector.
[0163] In this embodiment, a random order method is used for prediction, which makes full use of the order information of the data units to be predicted and explicitly integrates the order information into the output vector.
[0164] It should be understood that the above description of the prediction method for the word to be predicted uses text data as an example. The data processing method of this application embodiment can also be applied to the fields of computer vision or speech. Specifically, the target text can be replaced by a sequence of images or speech. This sequence is then processed by the preprocessing module through operations such as scrambling and segmentation to obtain the rearranged vector sequence of image or speech units and the position information of the position to be predicted. This information is then fed into the autoregressive encoding module and the query module, and finally the prediction module obtains the image or speech unit at the corresponding position to be predicted.
[0165] This application's embodiments can also be presented as a cloud-based service or software, see reference... Figure 14 The service or software may have the function of obtaining the data unit to be predicted based on the known data units in the target data. Specifically, the service includes, but is not limited to, predicting and restoring content at any position in text data (sentences, paragraphs, chapters, etc.), restoring fuzzy speech or missing audio sequences in speech data, and restoring fuzzy / damaged pixels in image / video data.
[0166] This application provides a data processing method, comprising: acquiring M first embedding vectors and a second embedding vector; wherein each first embedding vector represents a known data unit in target data and a first position of the known data unit in the target data, and the second embedding vector represents a second position of a first data unit to be predicted in the target data; M is a positive integer; processing the M first embedding vectors through a target encoder to obtain M first output vectors corresponding to the M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors; and processing the M first output vectors and the second embedding vectors through a target prediction network to obtain the first data unit to be predicted. Through the above method, for the M first embedding vectors corresponding to the M known data units, the target encoder can use the M first embedding vectors as input, wherein the first embedding vectors include position information and data information of the known data units, without needing to separately set additional M position information as input to the target encoder. Furthermore, the number of latent variables in the intermediate outputs of the target encoder is consistent with the number of input embedding vectors, reducing the computational load and memory consumption of the target encoder.
[0167] The inference process of the model has been described above. Next, from the perspective of model training, the data processing method provided in the embodiments of this application will be described, with reference to... Figure 15 and Figure 16 , Figure 15 and Figure 16 This is a flowchart illustrating a data processing method provided in an embodiment of this application, such as... Figure 15 As shown, the data processing method provided in this application embodiment includes: 1501. Obtain a first encoder, a first prediction network, M first embedding vectors, and a second embedding vector; wherein each first embedding vector is used to represent a known data unit in the target data and a first position of the known data unit in the target data, and the second embedding vector is used to represent a second position of a first data unit to be predicted in the target data; M is a positive integer; In this embodiment of the application, the first encoder and the first prediction network are neural network models to be trained.
[0168] In one possible implementation, the target data is text data, the known data unit is a known word in the text data, and the first data unit to be predicted is the word to be predicted in the text data; or... The target data is speech data, the known data unit is a known audio sequence in the speech data, and the first data unit to be predicted is the audio sequence to be predicted in the speech data; or... The target data is image data, the known data unit is a known pixel in the image data, and the first data unit to be predicted is a pixel to be predicted in the image data.
[0169] In one possible implementation, the first position is used to indicate the relative positional relationship between the known data unit and other known data units, as well as between the known data unit and the first data unit to be predicted; the second position is used to indicate the relative positional relationship between the first data unit to be predicted and each known data unit in the target data.
[0170] In one possible implementation, the first encoder is a first transformer layer and the first prediction network is a second transformer layer.
[0171] For more details on step 1501, please refer to the description of step 601, which will not be repeated here.
[0172] 1502. The first encoder processes the M first embedding vectors to obtain M first output vectors corresponding to the M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors; In one possible implementation, the first transformer layer includes a plurality of serial sub-transformer layers; each sub-transformer layer can process the data output by the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and output the M intermediate vectors to the next sub-transformer layer adjacent to it; wherein, if the sub-transformer layer is the transformer layer closest to the input side among the plurality of sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the plurality of sub-transformer layers, then the output data of the sub-transformer layer is the M first output vectors.
[0173] For more details on step 1502, please refer to the description of step 602. The similarities will not be repeated here.
[0174] 1503. The first prediction network is used to process the M first output vectors and the second embedding vector to obtain a third prediction data unit. The third prediction data unit is the result of the prediction made by the first prediction network.
[0175] For more details on step 1503, please refer to the description of step 603. The similarities will not be repeated here.
[0176] 1504. Based on the difference between the third prediction data unit and the first prediction data unit, update the first encoder and the first prediction network to obtain the target encoder and the target prediction network.
[0177] The third prediction data unit is the result of the prediction by the first prediction network. Therefore, a loss needs to be constructed based on the difference between the third prediction data unit and the first data unit to be predicted. The first encoder and the first prediction network are then updated based on the constructed loss to obtain the target encoder and the target prediction network. It should be understood that other network structures such as the embedding layer can also be updated based on the above loss, and this is not a limitation.
[0178] In one possible implementation, the target data further includes a second data unit to be predicted; before processing the M first embedding vectors through the first encoder to obtain a first output vector corresponding to each known data unit, the prediction order of the first and second data units to be predicted can be randomly determined; if the prediction order is used to indicate that the second data unit to be predicted is predicted after the first data unit to be predicted, then after obtaining the third predicted data unit, a fourth and a fifth embedding vector are obtained, the fourth embedding vector being used to represent the first data unit to be predicted and its second position in the target data, and the fifth embedding vector being used to represent the first data unit to be predicted and its second position in the target data. The third position of the second data unit to be predicted in the target data is indicated by the first encoder. The first encoder processes the M first embedding vectors and the fourth embedding vector to obtain a second output vector corresponding to each known data unit and the first data unit to be predicted. The first prediction network processes the M+1 second output vectors and the fifth embedding vector to obtain the fourth data unit to be predicted. Then, based on the difference between the third predicted data unit and the first data unit to be predicted, and the difference between the fourth predicted data unit and the second data unit to be predicted, the first encoder and the first prediction network can be updated to obtain the target encoder and the target prediction network.
[0179] In one possible implementation, the second output vector corresponding to each known data unit is generated based on the M first embedding vectors; the second output vector corresponding to the first data unit to be predicted is generated based on the M first embedding vectors and the fourth embedding vector.
[0180] Taking text data as the target data as an example, parameter tuning during the training phase can be performed using the standard backpropagation algorithm in deep learning. The loss function for this phase can be: ; in Here, represents all the parameters of the model (including transformer parameters, position vector parameters, and classifier parameters), x is the entire input sequence containing several elements, y represents the sequence of all words to be predicted (i.e., the original word corresponding to each predicted position), and S represents the set of positions of all words in y. This represents the word that needs to be predicted at the i-th position.
[0181] Reference Figure 17 This application also provides a data processing method, the method comprising: 1701. Obtain M first embedding vectors and a second embedding vector; wherein each first embedding vector is used to represent a data unit in the target data and the first position of the data unit in the target data, and the second embedding vector is used to indicate the target processing task; M is a positive integer; and the above Figure 6a Unlike the corresponding embodiments, the second embedding vector in this embodiment is used to indicate the target processing task, which includes, but is not limited to: short text classification, long text classification, natural language inference, text similarity matching, sentiment classification, etc.
[0182] In one possible implementation, the first position is used to indicate the relative positional relationship between the data unit and other data units.
[0183] In one possible implementation, the target data is text data, and the data unit is a word in the text data; or, The target data is speech data, and the known data unit is an audio sequence within the speech data; or... The target data is image data, and the known data unit is a pixel in the image data.
[0184] For a more detailed description of step 1701, please refer to the description of step 601 in the above embodiments, which will not be repeated here.
[0185] 1702. The M first embedding vectors are processed by the target encoder to obtain M output vectors corresponding to the M data units, wherein the output vector corresponding to each data unit is generated based on the M first embedding vectors.
[0186] The target encoder in this embodiment can be... Figure 6a In the corresponding embodiment, the target encoder is obtained by fine-tuning the model as a pre-trained model for the target processing task.
[0187] In one possible implementation, the target encoder is a first transformation transformer layer.
[0188] In one possible implementation, the first transformer layer includes a plurality of serial sub-transformer layers; each sub-transformer layer can process the data output by the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and output the M intermediate vectors to the next sub-transformer layer adjacent to it; wherein, if the sub-transformer layer is the transformer layer closest to the input side among the plurality of sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the plurality of sub-transformer layers, then the output data of the sub-transformer layer is the M output vectors.
[0189] In one possible implementation, the target encoder includes an attention head that can acquire attention information, which indicates that when the attention head processes the M first embedding vectors, there is an attention association between any two of the M first embedding vectors; and the M first embedding vectors are processed by the target encoder according to the attention information.
[0190] For a more detailed description of step 1702, please refer to the description of step 602 in the above embodiments, which will not be repeated here.
[0191] 1703. Through the task network, the M output vectors and the second embedding vector are processed according to the target processing task to obtain the task processing result.
[0192] In one possible implementation, the prediction network is a second transformer layer.
[0193] The following section uses multi-task text classification and reading comprehension as examples to illustrate the data processing method provided in this application.
[0194] Reference Figure 18 , Figure 18Here's an example implementation of text sentiment classification. In this example, the target data is "the cat sat on the mat." An attention matrix is obtained by segmenting the sentence. If the element in the i-th row and j-th column of this matrix is 1 (white), it indicates that the j-th word in the rearranged sequence is visible to the i-th word in the subsequent modeling process; otherwise, it is not. This module outputs the rearranged word vector sequence and the attention matrix (obtained from the segmentation) to the autoregressive word vector encoding module, and outputs auxiliary information about the label to be predicted (task type, e.g., sentiment classification) to the query module.
[0195] The autoregressive module can use a transformer layer as an autoregressive word vector encoder. This module adds each word vector in the rearranged word vector sequence to its corresponding position vector (each position corresponds to a position vector, which is part of the model's parameters). Furthermore, during the modeling process, it uses an attention matrix provided by the preprocessing module. This matrix defines whether each word is visible to other words during the transformer's word representation modeling process. Figure 18 As shown by the solid lines, the transformer ultimately generates a word vector representation for each word that incorporates contextual information, and outputs it to the prediction module. The query module outputs the task vector corresponding to the task type and sends it to the prediction module. The prediction module still uses the transformer model, which models the vector representation of the sentence. Finally, each modeled word vector is passed through a classifier to predict the corresponding word.
[0196] During the fine-tuning phase of training, the model predicts the labels corresponding to sentences. Parameter tuning in this phase can be performed using the standard backpropagation algorithm in deep learning. The loss function for this phase can be:
[0197] in y represents all the parameters of the model (including Transformer parameters, position vector parameters, task encoding parameters, and classifier parameters), x is the entire input sequence containing several elements, and y represents the label corresponding to the sentence.
[0198] Figure 19This example demonstrates a reading comprehension task involving span extraction. Given the question "who sat on the mat?" and the paragraph "the catsat on the mat", the task is to find the answer span within the paragraph, specifically the start and end positions ("the" and "cat"). In this example, an attention matrix is obtained based on sentence segmentation. If the element in the i-th row and j-th column of this matrix is 1 (white), it indicates that the j-th word in the rearranged sequence is visible to the i-th word in subsequent modeling; otherwise, it is not. This module outputs the rearranged word vector sequence and the attention matrix (obtained from the segmentation) to the autoregressive word vector encoding module, and outputs auxiliary information for the predicted label (the position information of each word in the paragraph) to the query module.
[0199] The autoregressive module can use a Transformer as an autoregressive word vector encoder. This module adds each word vector in the rearranged word vector sequence to its corresponding position vector (each position corresponds to a position vector, which is part of the model's parameters). During modeling, it uses an attention matrix provided by the preprocessing module. This matrix defines whether each word is visible to other words during the Transformer's word representation modeling process; the solid lines in the diagram represent visibility. The Transformer ultimately obtains a word vector representation for each word that incorporates contextual information and outputs it to the prediction module. The query module outputs the task vector corresponding to the task type and sends it to the prediction module. The prediction module still uses the Transformer model, which models the vector representation of the sentence. Finally, each modeled word vector is processed by two classifiers (which output the probabilities of each word being START and END, respectively). Figure 19 (As shown in the table below), used to predict the corresponding START and END positions.
[0200] During the fine-tuning phase of training, the model predicts the probabilities of START and END for each word in the passage. Parameter tuning in this phase employs the standard backpropagation algorithm in deep learning. The loss function for this phase can be: ; in represents all the parameters of the model (including Transformer parameters, position vector parameters, task encoding parameters, and classifier parameters), and x is the entire input sequence containing several elements. This indicates the probability that the model predicts the word at the START position in the answer as START. This indicates the probability that the model predicts the word at the END position in the answer as END.
[0201] During the inference phase, the over-tuned model can be used for prediction in downstream tasks. Taking text classification and reading comprehension as examples, the prediction method of this model is the same as in the fine-tuning phase, obtaining sentence or word labels through four modules and a classifier. In the reading comprehension task, the model takes the word with the highest predicted probability by the START classifier as the starting word of the span, and then takes the word with the highest END probability after the starting position as the ending word of the span.
[0202] exist Figures 1 to 19 Based on the corresponding embodiments, in order to better implement the above-described solutions of this application, related equipment for implementing the above solutions is also provided below. See details. Figure 20 , Figure 20 This is a schematic diagram of a data processing apparatus 2000 provided in an embodiment of this application. The data processing apparatus 2000 may be a terminal device or a server. The data processing apparatus 2000 includes: The acquisition module 2001 is used to acquire M first embedding vectors and second embedding vectors; wherein each first embedding vector is used to represent a known data unit in the target data and a first position of the known data unit in the target data, and the second embedding vector is used to represent a second position of a first data unit to be predicted in the target data; M is a positive integer; For a detailed description of the acquisition module 2001, please refer to the description of step 601 in the above embodiments, which will not be repeated here.
[0203] The encoding module 2002 is used to process the M first embedding vectors through the target encoder to obtain M first output vectors corresponding to the M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors; For a detailed description of the encoding module 2002, please refer to the description of step 602 in the above embodiments, which will not be repeated here.
[0204] The prediction module 2003 is used to process the M first output vectors and the second embedding vector through the target prediction network to obtain the first data unit to be predicted.
[0205] For a detailed description of the prediction module 2003, please refer to the description of step 603 in the above embodiments, which will not be repeated here.
[0206] In one possible implementation, the first position is used to indicate the relative positional relationship between the known data unit and other known data units, as well as between the known data unit and the first data unit to be predicted; the second position is used to indicate the relative positional relationship between the first data unit to be predicted and each known data unit in the target data.
[0207] In one possible implementation, the target encoder is a first transformer layer and the target prediction network is a second transformer layer.
[0208] In one possible implementation, the first transformer layer includes a series of sub-transformer layers; the step of processing the M first embedding vectors through the target encoder to obtain M first output vectors corresponding to the M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M first output vectors.
[0209] In one possible implementation, the target encoder includes an attention head and an encoding module for acquiring attention information, the attention information being used to indicate that when the attention head processes the M first embedding vectors, there is an attention association between any two of the M first embedding vectors; Based on the attention information, the M first embedding vectors are processed by the target encoder.
[0210] In one possible implementation, the device further includes: An embedding module is used to embed M known data units in the target data through an embedding layer to obtain M third embedding vectors; Obtain the position vector of each of the M known data units, the position vector being used to indicate the first position; Each of the M third embedding vectors is fused with its corresponding position vector to obtain the M first embedding vectors.
[0211] In one possible implementation, the target data further includes a second data unit to be predicted, and the order in which the second data unit to be predicted and the first data unit to be predicted are randomly determined.
[0212] In one possible implementation, if the second data unit to be predicted is predicted after the first data unit to be predicted, the method further includes: Obtain a fourth embedding vector and a fifth embedding vector, wherein the fourth embedding vector is used to represent the first data unit to be predicted and the second position of the first data unit to be predicted in the target data, and the fifth embedding vector is used to represent the third position of the second data unit to be predicted in the target data; The target encoder processes the M first embedding vectors and the fourth embedding vector to obtain M known data units and M+1 second output vectors corresponding to the first data unit to be predicted. The target prediction network processes the M+1 second output vectors and the fifth embedding vector to obtain the second data unit to be predicted.
[0213] In one possible implementation, the second output vector corresponding to each known data unit is generated based on the M first embedding vectors; the second output vector corresponding to the first data unit to be predicted is generated based on the M first embedding vectors and the fourth embedding vector.
[0214] In one possible implementation, the target data is text data, the known data unit is a known word in the text data, and the first data unit to be predicted is the word to be predicted in the text data; or... The target data is speech data, the known data unit is a known audio sequence in the speech data, and the first data unit to be predicted is the audio sequence to be predicted in the speech data; or... The target data is image data, the known data unit is a known pixel in the image data, and the first data unit to be predicted is a pixel to be predicted in the image data.
[0215] See Figure 21 , Figure 21 This is a schematic diagram of a data processing device 2100 provided in an embodiment of this application. The data processing device 2100 may be a terminal device or a server. The data processing device 2100 includes: The acquisition module 2101 is used to acquire M first embedding vectors and second embedding vectors; wherein each first embedding vector is used to represent a data unit in the target data and a first position of the data unit in the target data, and the second embedding vector is used to indicate the target processing task; M is a positive integer; For a detailed description of the acquisition module 2101, please refer to the description of step 1701 in the above embodiments, which will not be repeated here.
[0216] The encoding module 2102 is used to process the M first embedding vectors through the target encoder to obtain M output vectors corresponding to the M data units, wherein the output vector corresponding to each data unit is generated based on the M first embedding vectors; For a detailed description of the encoding module 2102, please refer to the description of step 1702 in the above embodiments, which will not be repeated here.
[0217] The task processing module 2103 is used to process the M output vectors and the second embedded vector through the task network to obtain the task processing result.
[0218] For a detailed description of the task processing module 2103, please refer to the description of step 1703 in the above embodiments, which will not be repeated here.
[0219] In one possible implementation, the first position is used to indicate the relative positional relationship between the data unit and other data units.
[0220] In one possible implementation, the target encoder is a first transformer layer and the task network is a second transformer layer.
[0221] In one possible implementation, the first transformer layer includes a series of sub-transformer layers; the step of processing the M first embedding vectors through the target encoder to obtain M first output vectors corresponding to the M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M output vectors.
[0222] In one possible implementation, the target encoder includes an attention head and an encoding module for acquiring attention information, the attention information being used to indicate that when the attention head processes the M first embedding vectors, there is an attention association between any two of the M first embedding vectors; Based on the attention information, the M first embedding vectors are processed by the target encoder.
[0223] In one possible implementation, the target data is text data, and the data unit is a word in the text data; or, The target data is speech data, and the known data unit is an audio sequence within the speech data; or... The target data is image data, and the known data unit is a pixel in the image data.
[0224] In one possible implementation, the target processing task includes short text classification, long text classification, natural language inference, text similarity matching, or text sentiment classification.
[0225] See Figure 22 , Figure 22 This is a schematic diagram of a data processing apparatus 2200 provided in an embodiment of this application. The data processing apparatus 2200 may be a terminal device or a server. The data processing apparatus 2200 includes: The acquisition module 2201 is used to acquire a first encoder, a first prediction network, M first embedding vectors, and a second embedding vector; wherein each first embedding vector is used to represent a known data unit in the target data and a first position of the known data unit in the target data, and the second embedding vector is used to represent a second position of a first data unit to be predicted in the target data; M is a positive integer; For a detailed description of the acquisition module 2201, please refer to the description of step 1501 in the above embodiments, which will not be repeated here.
[0226] The encoding module 2202 is used to process the M first embedding vectors through the first encoder to obtain M first output vectors corresponding to M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors; For a detailed description of the encoding module 2202, please refer to the description of step 1502 in the above embodiment, which will not be repeated here.
[0227] The prediction module 2203 is used to process the M first output vectors and the second embedding vector through the first prediction network to obtain a third prediction data unit. For a detailed description of the prediction module 2203, please refer to the description of step 1503 in the above embodiments, which will not be repeated here.
[0228] The model training module 2204 is used to update the first encoder and the first prediction network based on the difference between the third prediction data unit and the first prediction data unit to obtain the target encoder and the target prediction network.
[0229] For a detailed description of the model training module 2204, please refer to the description of step 1504 in the above embodiment, which will not be repeated here.
[0230] In one possible implementation, the first position is used to indicate the relative positional relationship between the known data unit and other known data units, as well as between the known data unit and the first data unit to be predicted; the second position is used to indicate the relative positional relationship between the first data unit to be predicted and each known data unit in the target data.
[0231] In one possible implementation, the first encoder is a first transformer layer and the first prediction network is a second transformer layer.
[0232] In one possible implementation, the first transformer layer includes a series of sub-transformer layers; the step of processing the M first embedding vectors through the first encoder to obtain M first output vectors corresponding to M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M first output vectors.
[0233] In one possible implementation, the target data further includes a second data unit to be predicted, and the order in which the second data unit to be predicted and the first data unit to be predicted are randomly determined.
[0234] In one possible implementation, if the second data unit to be predicted is predicted after the first data unit to be predicted, the method further includes: Obtain a fourth embedding vector and a fifth embedding vector. The fourth embedding vector is used to represent the first data unit to be predicted and the second position of the first data unit to be predicted in the target data. The fifth embedding vector is used to represent the third position of the second data unit to be predicted in the target data. The first encoder processes the M first embedding vectors and the fourth embedding vector to obtain M known data units and M+1 second output vectors corresponding to the first data unit to be predicted. The first prediction network processes the M+1 second output vectors and the fifth embedding vector to obtain the fourth data unit to be predicted. The step of updating the first encoder and the first prediction network based on the difference between the third prediction data unit and the first data unit to be predicted, to obtain the target encoder and the target prediction network, includes: Based on the difference between the third prediction data unit and the first prediction data unit, and the difference between the fourth prediction data unit and the second prediction data unit, the first encoder and the first prediction network are updated to obtain the target encoder and the target prediction network.
[0235] In one possible implementation, the second output vector corresponding to each known data unit is generated based on the M first embedding vectors; the second output vector corresponding to the first data unit to be predicted is generated based on the M first embedding vectors and the fourth embedding vector.
[0236] In one possible implementation, the target data is text data, the known data unit is a known word in the text data, and the first data unit to be predicted is the word to be predicted in the text data; or... The target data is speech data, the known data unit is a known audio sequence in the speech data, and the first data unit to be predicted is the audio sequence to be predicted in the speech data; or... The target data is image data, the known data unit is a known pixel in the image data, and the first data unit to be predicted is a pixel to be predicted in the image data.
[0237] The following describes an execution device provided in an embodiment of this application. Please refer to [link / reference]. Figure 23 , Figure 23 This is a schematic diagram of an execution device provided in an embodiment of this application. The execution device 2300 can specifically be a virtual reality (VR) device, a mobile phone, a tablet, a laptop, a smart wearable device, a monitoring data processing device, or a server, etc., and is not limited thereto. Specifically, the execution device 2300 includes: a receiver 2301, a transmitter 2302, a processor 2303, and a memory 2304 (wherein the execution device 2300 may have one or more processors 2303). Figure 23 (Taking a processor as an example), the processor 2303 may include an application processor 23031 and a communication processor 23032. In some embodiments of this application, the receiver 2301, transmitter 2302, processor 2303, and memory 2304 may be connected via a bus or other means.
[0238] Memory 2304 may include read-only memory and random access memory, and provides instructions and data to processor 2303. A portion of memory 2304 may also include non-volatile random access memory (NVRAM). Memory 2304 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0239] The processor 2303 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together through a bus system, which may include not only the data bus, but also power buses, control buses, and status signal buses. However, for clarity, all buses in the diagram are referred to as the bus system.
[0240] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 2303. The processor 2303 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 2303 or by instructions in software form. The processor 2303 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 2303 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 2304. Processor 2303 reads the information in memory 2304 and, in conjunction with its hardware, completes the steps of the above method.
[0241] Receiver 2301 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the execution device. Transmitter 2302 can be used to output digital or character information through the first interface; transmitter 2302 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 2302 may also include a display device such as a display screen.
[0242] In one embodiment of this application, the processor 2303 is configured to execute... Figure 6a , Figure 17 The data processing method described in the corresponding embodiment.
[0243] This application also provides a training device; please refer to [link / reference]. Figure 24 , Figure 24 This is a schematic diagram of a training device provided in an embodiment of this application. Specifically, the training device 2400 is implemented by one or more servers. The training device 2400 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 2424 (e.g., one or more processors) and memory 2432, and one or more storage media 2430 (e.g., one or more mass storage devices) for storing application programs 2442 or data 2444. The memory 2432 and storage media 2430 can be temporary or persistent storage. The program stored in the storage media 2430 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the training device. Furthermore, the CPU 2424 may be configured to communicate with the storage media 2430 and execute the series of instruction operations in the storage media 2430 on the training device 2400.
[0244] The training device 2400 may also include one or more power supplies 2426, one or more wired or wireless network interfaces 2450, one or more input / output interfaces 2458; or, one or more operating systems 2441, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0245] In this embodiment of the application, the central processing unit 2424 is used to execute... Figure 15 The data processing method described in the corresponding embodiment.
[0246] This application also provides a computer program product that, when run on a computer, causes the computer to perform steps as performed by the aforementioned execution device, or causes the computer to perform steps as performed by the aforementioned training device.
[0247] This application also provides a computer-readable storage medium storing a program for signal processing, which, when run on a computer, causes the computer to perform steps as performed by the aforementioned execution device, or causes the computer to perform steps as performed by the aforementioned training device.
[0248] The execution device, training device, or terminal device provided in this application embodiment can specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip within the execution device to execute the data processing method described in the above embodiments, or to cause the chip within the training device to execute the data processing method described in the above embodiments. Optionally, the storage unit can be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0249] For details, please refer to Figure 25 , Figure 25 This is a schematic diagram of a chip structure provided in an embodiment of this application. Figure 6a , Figure 15 and Figure 17 The data processing method described in the corresponding embodiment can be used in... Figure 25 The chip shown is used for implementation. Specifically, the chip can be represented as a neural network processor (NPU) 2500. The NPU 2500 is mounted as a coprocessor on the host CPU, and tasks are assigned by the host CPU. The core of the NPU is the arithmetic circuit 2503, and the controller 2504 controls the arithmetic circuit 2503 to retrieve data from the memory (weight memory or input memory) and perform calculations.
[0250] The above Figure 6a , Figure 15 and Figure 17 The data processing method described in the corresponding embodiments can be derived from... Figure 25 The main CPU and NPU in the chip shown work together to complete this task.
[0251] In some implementations, the arithmetic circuit 2503 internally includes multiple processing engines (PEs). In some implementations, the arithmetic circuit 2503 is a two-dimensional pulsating array. The arithmetic circuit 2503 can also be a one-dimensional pulsating array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 2503 is a general-purpose matrix processor.
[0252] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 2502 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 2501 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is stored in the accumulator 2508.
[0253] Unified memory 2506 is used to store input and output data. Weight data is directly transferred to weight memory 2502 via Direct Memory Access Controller (DMAC) 2505. Input data is also transferred to unified memory 2506 via DMAC.
[0254] BIU stands for Bus Interface Unit, which is used for interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 2509.
[0255] The Bus Interface Unit (BIU) 2510 is used by the instruction fetch memory 2509 to fetch instructions from external memory, and also by the memory access controller 2505 to fetch the original data of the input matrix A or the weight matrix B from external memory.
[0256] The DMAC is mainly used to move input data from external memory DDR to unified memory 2506, or to weight data to weight memory 2502, or to input data to input memory 2501.
[0257] The vector computation unit 2507 includes multiple processing units that further process the output of the computation circuits as needed, such as vector multiplication, vector addition, exponential operations, logarithmic operations, size comparisons, etc. It is primarily used for computations in non-convolutional / fully connected layers of neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.
[0258] In some implementations, the vector computation unit 2507 can store the processed output vector in the unified memory 2506. For example, the vector computation unit 2507 can apply a linear function, or a nonlinear function, to the output of the computation circuit 2503, such as performing linear interpolation on the feature planes extracted by the convolutional layer, or, for example, accumulating a vector of values to generate activation values. In some implementations, the vector computation unit 2507 generates normalized values, pixel-level summed values, or both. In some implementations, the processed output vector can be used as an activation input to the computation circuit 2503, for example, for use in subsequent layers of the neural network.
[0259] The instruction fetch buffer 2509 connected to the controller 2504 is used to store the instructions used by the controller 2504; The unified memory 2506, input memory 2501, weighted memory 2502, and instruction fetch memory 2509 are all on-chip memories. External memory is proprietary to this NPU hardware architecture.
[0260] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of the above program.
[0261] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0262] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0263] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0264] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A data processing method, characterized in that, The method includes: M first embedding vectors and a second embedding vector are obtained; wherein each first embedding vector represents a known data unit in the target data and a first position of the known data unit in the target data, and the second embedding vector represents a second position of a first data unit to be predicted in the target data; M is a positive integer, and the target data is text data, or the target data is speech data, or the target data is image data; the M first embedding vectors are processed by a target encoder to obtain M first output vectors corresponding to the M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors; the M first output vectors and the second embedding vectors are processed by a target prediction network to obtain the first data unit to be predicted; The target data further includes a second data unit to be predicted, and the order in which the second data unit to be predicted and the first data unit to be predicted are randomly determined.
2. The method according to claim 1, characterized in that, The first position is used to indicate the relative positional relationship between the known data unit and other known data units, as well as between the known data unit and the first data unit to be predicted; the second position is used to indicate the relative positional relationship between the first data unit to be predicted and each known data unit in the target data.
3. The method according to claim 1 or 2, characterized in that, The target encoder is a first transformer layer, and the target prediction network is a second transformer layer.
4. The method according to claim 3, characterized in that, The first transformation transformer layer includes multiple serial sub-transformer layers; the step of processing the M first embedding vectors through the target encoder to obtain M first output vectors corresponding to M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M first output vectors.
5. The method according to any one of claims 1 to 2, characterized in that, The target encoder includes an attention head, and the processing of the M first embedding vectors by the target encoder includes: Obtain attention information, which is used to indicate that when the attention head processes the M first embedding vectors, there is an attention association between any two first embedding vectors among the M first embedding vectors; Based on the attention information, the M first embedding vectors are processed by the target encoder.
6. The method according to any one of claims 1 to 2, characterized in that, The method further includes: The target data is embedded by an embedding layer to obtain M third embedding vectors; Obtain the position vector of each of the M known data units, the position vector being used to indicate the first position; Each of the M third embedding vectors is fused with its corresponding position vector to obtain the M first embedding vectors.
7. The method according to claim 1, characterized in that, If the second data unit to be predicted is predicted after the first data unit to be predicted, the method further includes: Obtain a fourth embedding vector and a fifth embedding vector, wherein the fourth embedding vector is used to represent the first data unit to be predicted and the second position of the first data unit to be predicted in the target data, and the fifth embedding vector is used to represent the third position of the second data unit to be predicted in the target data; The target encoder processes the M first embedding vectors and the fourth embedding vector to obtain M known data units and M+1 second output vectors corresponding to the first data unit to be predicted. The target prediction network processes the M+1 second output vectors and the fifth embedding vector to obtain the second data unit to be predicted.
8. The method according to claim 7, characterized in that, The second output vector corresponding to each known data unit is generated based on the M first embedding vectors; the second output vector corresponding to the first data unit to be predicted is generated based on the M first embedding vectors and the fourth embedding vector.
9. The method according to any one of claims 1 to 2, characterized in that, The target data is text data, the known data unit is a known word in the text data, and the first data unit to be predicted is the word to be predicted in the text data; or... The target data is speech data, the known data unit is a known audio sequence in the speech data, and the first data unit to be predicted is the audio sequence to be predicted in the speech data; or... The target data is image data, the known data unit is a known pixel in the image data, and the first data unit to be predicted is a pixel to be predicted in the image data.
10. A data processing method, characterized in that, The method includes: A first encoder, a first prediction network, M first embedding vectors, and a second embedding vector are obtained; wherein each first embedding vector is used to represent a known data unit in the target data and a first position of the known data unit in the target data, and the second embedding vector is used to represent a second position of a first data unit to be predicted in the target data; M is a positive integer, and the target data is text data, or the target data is speech data, or the target data is image data; The first encoder processes the M first embedding vectors to obtain M first output vectors corresponding to the M known data units, wherein the first output vector corresponding to each known data unit is generated based on the M first embedding vectors. The first prediction network processes the M first output vectors and the second embedding vector to obtain the third prediction data unit. Based on the difference between the third prediction data unit and the first prediction data unit, the first encoder and the first prediction network are updated to obtain the target encoder and the target prediction network. The target data further includes a second data unit to be predicted, and the order in which the second data unit to be predicted and the first data unit to be predicted are randomly determined.
11. The method according to claim 10, characterized in that, The first position is used to indicate the relative positional relationship between the known data unit and other known data units, as well as between the known data unit and the first data unit to be predicted; the second position is used to indicate the relative positional relationship between the first data unit to be predicted and each known data unit in the target data.
12. The method according to claim 10 or 11, characterized in that, The first encoder is a first transformer layer, and the first prediction network is a second transformer layer.
13. The method according to claim 12, characterized in that, The first transformation transformer layer includes multiple serial sub-transformer layers; the step of processing the M first embedding vectors through the first encoder to obtain M first output vectors corresponding to M known data units includes: Each of the sub-transformer layers processes the data output from the previous sub-transformer layer adjacent to it to obtain M intermediate vectors, and outputs the M intermediate vectors to the next sub-transformer layer adjacent to it. Wherein, if the sub-transformer layer is the transformer layer closest to the input side among the multiple sub-transformer layers, then the input data of the sub-transformer layer is the M first embedding vectors; if the sub-transformer layer is the transformer layer closest to the output side among the multiple sub-transformer layers, then the output data of the sub-transformer layer is the M first output vectors.
14. The method according to claim 13, characterized in that, If the second data unit to be predicted is predicted after the first data unit to be predicted, the method further includes: Obtain a fourth embedding vector and a fifth embedding vector. The fourth embedding vector is used to represent the first data unit to be predicted and the second position of the first data unit to be predicted in the target data. The fifth embedding vector is used to represent the third position of the second data unit to be predicted in the target data. The first encoder processes the M first embedding vectors and the fourth embedding vector to obtain M known data units and M+1 second output vectors corresponding to the first data unit to be predicted. The first prediction network processes the M+1 second output vectors and the fifth embedding vector to obtain the fourth data unit to be predicted. The step of updating the first encoder and the first prediction network based on the difference between the third prediction data unit and the first data unit to be predicted, to obtain the target encoder and the target prediction network, includes: Based on the difference between the third prediction data unit and the first prediction data unit, and the difference between the fourth prediction data unit and the second prediction data unit, the first encoder and the first prediction network are updated to obtain the target encoder and the target prediction network.
15. The method according to claim 14, characterized in that, The second output vector corresponding to each known data unit is generated based on the M first embedding vectors; the second output vector corresponding to the first data unit to be predicted is generated based on the M first embedding vectors and the fourth embedding vector.
16. The method according to any one of claims 10 to 11, characterized in that, The target data is text data, the known data unit is a known word in the text data, and the first data unit to be predicted is the word to be predicted in the text data; or... The target data is speech data, the known data unit is a known audio sequence in the speech data, and the first data unit to be predicted is the audio sequence to be predicted in the speech data; or... The target data is image data, the known data unit is a known pixel in the image data, and the first data unit to be predicted is a pixel to be predicted in the image data.
17. A data processing apparatus, characterized in that, The device includes a memory and a processor; the memory stores code, and the processor is configured to retrieve the code and perform the method as described in any one of claims 1 to 16.
18. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions that, when executed by one or more computers, cause the one or more computers to perform the method of any one of claims 1 to 16.
19. A computer program product, characterized in that, The computer program product includes code that, when executed, performs the steps of the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Constraint double-cluster mining and missing value prediction method based on order preserving submatrix
CN110222089A
Model compression method and device
CN112257858A