Data processing method and system

By storing big data in external memory and processing it in batches, the problem that the amount of data processing by computer exceeds the memory capacity is solved, and efficient processing of big data is achieved.

CN120234149APending Publication Date: 2025-07-01ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510376917.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing computers have difficulty processing data whose amount of data is larger than their memory, and ordinary computers usually limit the amount of data they process to avoid memory overflow.

Method used

By storing the target data in external memory and reading it into memory in batches according to processing instructions to perform operations, the processor can process data whose data volume is larger than its memory.

Benefits of technology

It enables computers to process data beyond their memory capacity, improving the flexibility and efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234149A_ABST
    Figure CN120234149A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and system. In the method, target data with the data size larger than the memory capacity of the computing device is stored in an external memory of the computing device, and when the computing device obtains a processing instruction corresponding to the target data, the target data is read into the memory from the external memory in batches based on the processing instruction to execute operation. The data volume of each batch of data in the target data is smaller than the capacity of the memory, and the operation mode executed on each batch of data is related to the processing instruction. Large-scale data is stored through the external memory and loaded into the memory in batches for processing, so that the computing device can process data with the data size larger than that of the memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of data processing, and in particular, to a data processing method and system. Background Art

[0002] In the era of big data, the amount of data that a computer needs to process is increasing. For example, long texts with millions of words, videos with extremely long durations, etc. Processing such large-scale data requires extremely large-capacity memory, and ordinary computers such as 8C16G (8 CPU cores and 16GB of memory) or 16C32G (16 CPU cores and 32GB of memory) cannot handle it. Ordinary computers generally limit the amount of data to be processed. For example, they limit the text length, so as to process data with a data volume smaller than its memory, such as short texts, and processing large-scale data is a difficult problem.

[0003] Therefore, a data processing method is needed to enable a computer to process data with a data volume larger than its memory. The content in the background art section is only the information known to the inventor personally, and does not represent that the above information has entered the public domain before the filing date of this disclosure, nor does it represent that it can become the prior art of this disclosure. Summary of the Invention

[0004] The data method and system provided in this specification enable a computing device to process data with a data volume larger than its memory.

[0005] In a first aspect, this specification provides a data processing method, which is applied to a computing device and includes: obtaining a processing instruction corresponding to target data, where the target data is stored in the external memory of the computing device, and the data volume of the target data is larger than the capacity of the memory of the computing device; based on the processing instruction, reading the target data in batches from the external memory into the memory for execution of operations to obtain a processing result corresponding to the target data, where the data volume of each batch of data in the target data is smaller than the capacity of the memory, and the operation method executed on each batch of data is related to the processing instruction.

[0006] In some embodiments, the processing instruction represents performing a normalization operation on the target data.

[0007] In some embodiments, the target data is data that has been processed by a Q matrix and a K matrix respectively, and the Q matrix and the K matrix are the Q matrix and the K matrix in a Transformer model.

[0008] In some embodiments, the Q matrix and the K matrix are generated based on a plurality of data units, the target data corresponds to a target data unit among the plurality of data units, the target data includes a plurality of associated data, and each associated data represents the degree of association between each data unit and the target data unit.

[0009] In some embodiments, the operation of reading the target data in batches from the external memory to the memory for execution based on the processing instruction includes: based on the processing instruction, performing two rounds of traversal on the target data, and the process of each round of traversal includes reading the target data in batches from the external memory to the memory for execution. Among them, the operation mode corresponding to the first round of traversal at least includes addition operation, and the operation mode corresponding to the second round of traversal at least includes division operation.

[0010] In some embodiments, the operation mode corresponding to the first round of traversal further includes an exponential operation, and the process of the first round of traversal includes: reading the current batch of data from the external memory to the memory; for each associated data in the current batch of data, performing an exponential operation on the associated data in the memory to obtain corresponding exponential data; and performing the addition operation on the exponential data and the aggregation value in the current batch of data in the memory to obtain an updated aggregation value, where the aggregation value is initialized to 0 before the first round of traversal.

[0011] In some embodiments, the exponential data includes data obtained by performing the exponential operation on the difference between the associated data and the maximum value in the target data.

[0012] In some embodiments, the operation mode corresponding to the second round of traversal further includes an exponential operation, and the process of the second round of traversal includes: reading the current batch of data from the external memory to the memory; and for each associated data in the current batch of data: performing the exponential operation on the current associated data in the memory to obtain corresponding exponential data, and performing the division operation on the exponential data and the aggregation value after the first round of traversal to obtain the processing result corresponding to the current associated data.

[0013] In some embodiments, the Transformer model further includes a V matrix, and the processing result includes: a plurality of normalized data corresponding to the plurality of associated data, and the processing result is used for the enhancement representation process of the target data unit. The enhancement representation process includes: using the plurality of normalized data as weights to perform a weighted sum on the V matrix to obtain an enhanced representation result of the target data unit, and the enhanced representation result integrates the information of the plurality of data units.

[0014] In some embodiments, the target data unit is one of text, image, or audio.

[0015] In some embodiments, the external memory includes built-in external memory and / or removable external memory.

[0016] In some embodiments, the memory includes a target storage space for storing batches of data, and the data volume of each batch of data is less than or equal to the capacity of the target storage space.

[0017] In some embodiments, the method further includes: applying for a storage space in the memory as the target storage space; and determining the data volume of each batch of data based on the capacity of the target storage space.

[0018] In some embodiments, applying for a storage space in the memory as the target storage space includes: predicting, based on the processing instruction, a first buffer capacity required to buffer other data except each batch of data during the execution of the processing instruction; obtaining the current free capacity of the memory, and taking the difference between the current free capacity and the first buffer capacity as a second buffer capacity; and applying for a storage space less than or equal to the second buffer capacity in the memory as the target storage space.

[0019] In a second aspect, this specification further provides a data processing system, including: at least one storage medium storing at least one instruction set for implementing data processing; and at least one processor communicatively connected to the at least one storage medium, wherein when the data processing system runs, the at least one processor reads the at least one instruction set and implements the data processing method according to any one of the first aspect.

[0020] As can be seen from the above technical solutions, in the data processing method provided in this specification, target data with a data volume greater than the memory capacity of the computing device is stored in the external memory of the computing device. When the computing device obtains a processing instruction corresponding to the target data, the target data is read in batches from the external memory to the memory for execution of operations based on the processing instruction, and a processing result corresponding to the target data is obtained. Among them, the data volume of each batch of data in the target data is less than the capacity of the memory, and the operation mode performed on each batch of data is related to the processing instruction. By storing large-scale data in the external memory and loading it into the memory in batches for processing, the computing device can process data with a data volume greater than its memory.

[0021] Other functions of the data processing method and system provided in this specification will be partially listed in the following description. According to the description, the content introduced by the following numbers and examples will be obvious to those of ordinary skill in the art. The creative aspects of the data processing method and system provided in this specification can be fully explained by practicing or using the methods, devices, and combinations described in the detailed examples below. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 FIG. shows a schematic diagram of a data processing scenario provided according to some embodiments of this specification;

[0024] Figure 2 FIG. shows a hardware structure diagram of a computing device provided according to some embodiments of this specification;

[0025] Figure 3 FIG. shows a flowchart of a data processing method provided according to some embodiments of this specification; and

[0026] Figure 4 FIG. shows a schematic diagram of performing operations on batches of data provided according to some embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various partial modifications to the disclosed embodiments are obvious, and the general principles defined here can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the disclosed embodiments, but rather to the broadest scope consistent with the claims.

[0028] The terms used herein are for the purpose of describing particular example embodiments only and are not limiting. For example, unless the context clearly dictates otherwise, as used herein, the singular forms "a", "an" and "the" may also include the plural forms. When used in this specification, the terms "comprising", "including" and / or "containing" mean that the associated integers, steps, operations, elements and / or components are present, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups in the system / method.

[0029] In view of the following description, these features of the present specification and other features, as well as the operations and functions of the related elements of the structure, and the economy of the combination and manufacture of the components can be significantly improved. Referring to the accompanying drawings, all of which form a part of this specification. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0030] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments of this specification. It should be clearly understood that the operations of the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.

[0031] In this specification, the expression "X includes at least one of A, B or C" means that X includes at least A, or X includes at least B, or X includes at least C. That is, X may include only any one of A, B, C, or may include any combination of A, B, C and other possible contents / elements at the same time. Any combination of A, B, C may be A, B, C, AB, AC, BC, or ABC.

[0032] In this specification, unless otherwise explicitly stated, the associated relationship generated between structures may be a direct associated relationship or an indirect associated relationship. For example, when describing "A is connected to B", unless it is explicitly stated that A is directly connected to B, it should be understood that A may be directly connected to B or indirectly connected to B; for another example, when describing "A is above B", unless it is explicitly stated that A is directly above B (A and B are adjacent and A is above B), it should be understood that A may be directly above B or A may be indirectly above B (there are other elements between A and B and A is above B). And so on.

[0033] For the convenience of description, this specification explains the terms that will appear in the following description:

[0034] Softmax: A mathematical function used to implement normalization operations, commonly used in the fields of machine learning and deep learning. It can transform a K-dimensional vector containing arbitrary real numbers into another vector of the same length, where each normalized data value ranges between (0, 1) and the sum of all normalized data values equals 1.

[0035] Large scale: The scale is so large that it cannot be stored in memory.

[0036] The data processing method provided in this specification can be applied to scenarios where target data with a large data volume is processed. The form of the target data can be a matrix, a vector, or a tensor, or it can be a graph composed of nodes and edges used to represent the relationship between entities. The embodiments of this specification do not limit the form of the target data. The target data can be the data obtained after the computing device processes the target data unit. The target data unit can be one of text, image, or audio. For example, the target data is a matrix obtained after the computing device processes the target data unit using the Q matrix and K matrix in the Transformer model. The processing of the target data can be normalization operations (such as Softmax operation, Sigmod operation, Min-Max operation), addition operation, multiplication operation, weighted summation operation, or other operations, etc.

[0037] Taking the application of the method in this specification to the Transformer model as an example, the Transformer model can be used in various scenarios such as processing text (such as text translation, text generation), processing images (such as image classification), and processing audio (such as speech recognition). The most core part of the Transformer model is the Attention mechanism (attention mechanism), such as the Self-Attention mechanism (self-attention mechanism). When the Attention mechanism processes multiple data units, such as a long text sequence, the processing of a word at a position can simultaneously pay attention to all words in the text sequence and assign different weights to different positions, so as to better capture semantic relationships. The method in this specification can be used to normalize the attention score matrix in the Attention mechanism. The method in this specification can also be used to perform weighted summation processing on the normalized score matrix after normalization in the Attention mechanism.

[0038] It should be noted that the above example scenario is only one of the multiple usage scenarios provided in this specification. The data processing method provided in this specification can not only be applied to the above scenarios, but can be used when large-scale data needs to be processed in any scenario. Those skilled in the art should understand that the application of the data processing method described in this specification to other usage scenarios is also within the protection scope of this specification.

[0039] Figure 1 FIG. 1 is a schematic diagram of a data processing scenario 001 provided according to some embodiments of this specification. Figure 1 As shown, the scenario 001 may include a computing device 100 and an external memory 101 of the computing device 100. It should be noted that the user data obtained in this specification is authorized by the user and does not involve user privacy.

[0040] In some embodiments, the external memory 101 may be a device for storing data in the computing device 100. The external memory 101 may provide a larger storage space than the internal memory, and may store a large amount of data and files. The data in the external memory 101 may remain intact after the computing device 100 is powered off, so the external memory 101 may ensure the security and persistence of the data. In some embodiments, the external memory 101 may be a built-in external memory of the computing device 100. The built-in external memory may be permanently or semi-permanently installed inside the computing device 100, and may not be easily removed or replaced by the user. The built-in external memory may be, for example, a disk (such as a hard disk drive (HDD), a solid state drive (SSD), a network attached storage (NAS), etc. In some embodiments, the external memory 101 may be a removable external memory. The removable external memory is convenient for the user to insert and unplug at any time, and is convenient for carrying and sharing data between different devices. The removable external memory may be, for example, an optical drive, a USB flash drive, an external hard disk drive, etc.

[0041] In some embodiments, the computing device 100 may include a memory. The memory is a high-speed storage device in the computing device 100 for temporarily storing data and program instructions. The memory directly interacts with the central processing unit (CPU) of the computing device 100, which may also be referred to as a processor. The memory can respond quickly to requests from the processor to ensure that the data required by the processor can be obtained in a timely manner. Compared to the external memory 101, the capacity of the memory is usually smaller. Examples of the memory include dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), cache memory, etc.

[0042] In some embodiments, the data processing method may be executed on a computing device 100, specifically by a processor in the computing device 100. At this time, the computing device 100 may store data or instructions for executing the data processing method described in this specification, and may execute or be used to execute the data or instructions. The computing device 100 may include a hardware device with data information processing capabilities and a program required to drive the operation of the hardware device.

[0043] In some embodiments, the computing device 100 may be a device acting as a terminal. The terminal may include a mobile device, a tablet computer, a laptop computer, a built-in device of a motor vehicle or the like, a vending machine, a vending cabinet, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, a virtual reality device, an augmented reality device or the like, or any combination thereof. In some embodiments, the smart home device may include a smart TV, a desktop computer, etc., or any combination. In some embodiments, the smart mobile device may include a smart phone, a personal digital assistant, a gaming device, a navigation device, etc., or any combination thereof. In some embodiments, the virtual reality device or the augmented reality device may include a virtual reality headset, virtual reality glasses, a virtual reality patch, an augmented reality headset, augmented reality glasses, an augmented reality patch or the like, or any combination thereof. For example, the virtual reality device or the augmented reality device may include Google Glasses, a head-mounted display, VR, etc. In some embodiments, the built-in device in the motor vehicle may include an in-vehicle computer, an in-vehicle TV, etc.

[0044] In some embodiments, one or more application programs (APPs) may be installed on the terminal. The APP is capable of executing the data processing method. The APP includes, but is not limited to: web browser type APP programs, search type APP programs, chat type APP programs, shopping type APP programs, video type APP programs, financial management type APP programs, instant messaging tools, email clients, social platform software, and so on. In some embodiments, a target APP may be installed on the terminal. In some embodiments, the user may trigger a data processing request through the target APP. The target APP may obtain a processing instruction corresponding to the target data in response to the data processing request, and thus execute the data processing method.

[0045] In some embodiments, the computing device 100 may be a device acting as a server. The server may be a cloud server or a local server. When the server adopts a cluster architecture, the computing device 100 may be any device / node in the cluster system.

[0046] In some embodiments, the computing device 100 may be communicatively connected to other devices via a network and transmit information or data to each other via the network. In some embodiments, the network may be any type of wired or wireless network, or a combination thereof. For example, the network may include a cable network, a wired network, an optical fiber network, a telecommunication network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, a near field communication (NFC) network, or a similar network. In some embodiments, the network may include one or more network access points. For example, the network may include a wired or wireless network access point, such as a base station or an Internet exchange point, through which one or more components of the computing device 100 and other devices can be connected to the network to exchange data or information.

[0047] It should be understood that Figure 1 the number of computing devices and external memories in

[0048] Figure 2 The hardware structure diagram of a computing device 100 provided according to some embodiments of this specification is shown. The computing device 100 may execute the data processing method described in this specification. The data processing method is introduced in other parts of this specification.

[0049] As Figure 2 shown, the computing device 100 may include at least one storage medium 130 and at least one processor 120. In some embodiments, the computing device 100 may further include a communication port 150 and an internal communication bus 110. At the same time, the computing device 100 may further include I / O components 160.

[0050] The internal communication bus 110 may connect different system components, including the storage medium 130, the processor 120, and the communication port 150.

[0051] The I / O components 160 support input / output between the computing device 100 and other components.

[0052] The communication port 150 is used for data communication between the computing device 100 and the outside world. For example, the communication port 150 may be used for data communication between the computing device 100 and the network. The communication port 150 may be a wired communication port or a wireless communication port.

[0053] The storage medium 130 may include a data storage device. The data storage device may be a non-transitory storage medium or a transitory storage medium. The data storage device may be memory or external storage 101. The external storage 101 is, for example, a disk. For example, the data storage device may include one or more of a disk 101, a read-only storage medium (ROM) 103, or a random access storage medium (RAM) 105. The storage medium 130 may store at least one set of instruction sets for implementing data processing. The instructions are computer program codes, and the computer program codes may include programs, routines, objects, components, data structures, processes, modules, etc. for executing the data processing method provided in this specification. The storage medium 130 may also store a data processing model for implementing the data processing method. At this time, the model may be one or more instruction sets that execute corresponding instructions stored in the storage medium 130 and are executed by the processor 120 in the computing device 100. Of course, the model may also be a part of the circuit, a hardware device, or a module in the computing device 100. For example, the data processing model may be a hardware device / module in the computing device 100 that implements data processing, and the data processing model may be a hardware device / module in the computing device 100 that implements data processing, and so on. At this time, the processor 120 may store at least one set of instruction sets for controlling one or the instruction sets of the model.

[0054] At least one processor 120 may be communicatively connected to at least one storage medium 130 and a communication port 150 via an internal communication bus 110. The at least one processor 120 is configured to execute the above-mentioned at least one instruction set. When the computing device 100 is running, the at least one processor 120 may read the at least one instruction set and, according to the instructions of the at least one instruction set, execute the data processing method provided in this specification. The processor 120 may execute all steps included in the data processing method. The processor 120 may be in the form of one or more processors. In some embodiments, the processor 120 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field-programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, etc., or any combination thereof. For illustrative purposes only, only one processor 120 is described in the computing device 100 in this specification. However, it should be noted that the computing device 100 in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be executed by one processor as described in this specification or jointly executed by multiple processors. For example, if the processor 120 of the computing device 100 executes steps A and B in this specification, it should be understood that steps A and B may also be jointly or separately executed by two different processors 120 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).

[0055] Figure 3 FIG. shows a flowchart of a data processing method P100 provided according to some embodiments of this specification. As described above, the computing device 100 may execute the data processing method P100 described in this specification. Specifically, the processor 120 may read the instruction set stored in its local storage medium and then, according to the provisions of the instruction set, execute the data processing method P100 described in this specification. As Figure 3 shown, the method P100 may include:

[0056] S120: Obtain a processing instruction corresponding to target data, where the target data is stored in the external memory of the computing device, and the data volume of the target data is greater than the capacity of the memory of the computing device.

[0057] In some embodiments, the target data may be data for describing a certain sequence. The above-mentioned sequence includes multiple data units, and each data unit is one of text, image, or audio. The multiple data units can be converted into input feature representations that the model can understand. For example, the text is converted into word embeddings, and the image is converted into a tensor. Thus, these input feature representations are used as the input of the model. The target data corresponds to the target data unit among the multiple data units, and specifically, it can be the data obtained after the model processes the input feature representation of the target data unit.

[0058] In some embodiments, the target data is the data respectively processed by the Q matrix and the K matrix, and the Q matrix and the K matrix are the Q matrix and the K matrix in the Transformer model.

[0059] Among them, the Transformer model can use the attention mechanism to process sequence data. In the attention mechanism, the multiple input feature representations corresponding to the multiple data units can generate three different matrix representations through linear transformation: the Q (Query) matrix, the K (Key) matrix, and the V (Value) matrix. That is, the Q matrix, the K matrix, and the V matrix are generated based on the multiple data units. The Q matrix contains each q vector corresponding to each input feature representation, and the q vector of each input feature representation is used to perform similarity matching with the k vectors of all input feature representations to determine the degree of association between all input feature representations and this input feature representation. The K matrix contains each k vector corresponding to each input feature representation, and each k vector provides an identifier for the corresponding input feature representation, enabling other input feature representations to search for the relevant information of this input feature representation through the q vector. The V matrix contains each v vector corresponding to each input feature representation, and each v vector represents the information that the corresponding input feature representation actually wants to transmit, and is used to aggregate the information of different input feature representations in combination with the attention weights.

[0060] The q vector of the target data unit in the Q matrix can perform similarity calculation (such as dot product) with each k vector in the K matrix to obtain an attention score matrix. This attention score matrix contains the importance scores of all data units for the target data unit, and can represent the degree of association between all data units and the target data unit. For example, the multiple data units are 4 million words, and the attention score matrix of the target word contains the importance scores of 4 million words for the target word respectively. Assuming that each importance score is 128 - dimensional, the dimension of the attention score matrix is 128 * 4 million. The attention score matrix can also perform a normalization operation to convert the attention scores into probabilities between 0 and 1, so as to obtain an attention weight matrix, and this attention weight matrix can also represent the degree of association between all data units and the target data unit.

[0061] The target data includes multiple associated data, and each piece of associated data represents the degree of association between each data unit and the target data unit. In some embodiments, the target data may be an attention score matrix corresponding to the target data unit. In some embodiments, the target data may be an attention weight matrix corresponding to the target data unit.

[0062] After each data unit is processed by the attention mechanism, the attention layer can output the enhanced representation results of each data unit, and the enhanced representation results can be further input into the feed-forward neural network in the Transformer model for processing. The data volume of multiple data units is large-scale (such as 4 million tokens), and the data volume of the corresponding multiple enhanced representation results is naturally also large-scale. Therefore, there is also a problem that the computing device 100 cannot process large-scale data based on the memory. Therefore, in some embodiments, the target data may also be multiple enhanced representation results, so as to be processed by the data processing method provided in this specification.

[0063] The data processing method provided in this specification can be applied to the scenario of the Transformer model, and can process the data after being processed by the Q matrix and the K matrix, solving the problem that the computing device 100 cannot process large-scale data in this scenario. When the target data includes multiple associated data, the method in this specification enables the computing device 100 to process the large-scale data in the attention mechanism, making the Transformer model and the attention mechanism applicable to various complex scenarios and improving the generalization ability.

[0064] In some embodiments, the target data may be data in other models, such as data in Bert (Bidirectional Encoder Representations from Transformers) or GPT (Generative Pretrained Transformer), and the embodiments of this specification do not limit this.

[0065] In some embodiments, the target data may not only be the data corresponding to the sequence, but also the data corresponding to other types such as tables, web pages, emails, meteorological data, network structure data, directory structure data, etc. In some embodiments, the target data may be the input feature representations of multiple data units. In some embodiments, the target data may be the unencoded original content, such as sequences, tables, web pages, emails, meteorological data, network structure data, directory structure data, etc. As long as the data volume of the target data is greater than the memory capacity of the computing device 100, it is within the protection scope of this specification.

[0066] In this specification, the data volume of the target data is greater than the capacity of the memory of the computing device 100. The computing device 100 cannot store the target data only with the memory. Therefore, the external memory 101 of the computing device 100 is used to store the target data for the processor 120 to process. For example, the data volume of the target data is several hundred GB, while the memory capacity of the computing device 100 is 16 GB. Among them, the external memory 101 includes an internal external memory and / or a removable external memory. The internal external memory and the removable external memory have been introduced above and will not be elaborated here.

[0067] In some embodiments, the processor 120 can gradually generate the associated data in the target data and gradually store the associated data in the external memory 101, thereby forming the target data in the external memory 101. For example, when the data volume of the generated associated data reaches the storage threshold of the memory (such as the memory capacity or a preset value smaller than the capacity), the processor 120 stores this part of the associated data in the external memory 101. In some embodiments, the target data can be stored by the user in the removable external memory, such as storing the target data from the database in the removable external memory. After connecting the removable external memory to the computing device 100, the processor 120 can process the target data in the removable external memory.

[0068] In some embodiments, the processing instruction represents performing a normalization operation on the target data. For example, softmax operation, min-max normalization operation (Min-Max Normalization), Logistic normalization operation, etc. By the embodiments of this specification, the problem of small memory capacity for the computing device 100 to perform normalization processing on large-scale data is solved. For example, when the computing device 100 needs to perform a softmax operation on a text sequence of 4 million words or even longer, it is necessary to perform summation and averaging in the dimension of 4 million, which far exceeds its memory capacity. Therefore, the method of this specification can be used for processing, which is simple and efficient. Of course, the processing instruction can also represent other operations, such as integral operation and differential operation, etc.

[0069] S140: Based on the processing instruction, read the target data in batches from the external memory into the memory for execution of the operation, and obtain the processing result corresponding to the target data, where the data volume of each batch of data in the target data is less than the capacity of the memory, and the operation mode performed on each batch of data is related to the processing instruction.

[0070] After obtaining the processing instruction corresponding to the target data, the processor 120 can perform at least one round of traversal on the target data to obtain a processing result. In each round of traversal, the target data can be read in batches from the external memory 101 into the memory for operation. The traversal refers to accessing each associated data in the target data. The number of rounds of traversal is related to the processing instruction. Different processing instructions may correspond to different numbers of rounds of traversal. For example, when the processing instruction is a normalization operation (such as softmax), the number of rounds of traversal is 2, while when the processing instruction is an addition operation, the number of rounds of traversal is 1. Different processing instructions may also correspond to the same number of rounds of traversal. For example, the number of rounds of traversal for both the addition operation and the multiplication operation is 1.

[0071] In some embodiments, after obtaining the processing instruction corresponding to the target data, the processor 120 can determine the number of rounds of traversal of the target data and the operation mode performed on each batch of data in each round of traversal based on the processing instruction. Furthermore, the processor 120 performs at least one round of traversal on the target data according to the number of rounds of traversal and the operation mode corresponding to each round of traversal until the processing result corresponding to the target data is obtained. That is to say, the data processing method provided in this specification can support multiple processing instructions. After obtaining the processing instruction, the processor 120 can adaptively analyze the number of rounds of traversal that needs to be performed and the operation mode corresponding to each round of traversal, improving the application universality of the data processing method provided in this specification.

[0072] In some embodiments, the processor 120 can pre-store the correspondence between the processing instruction and the execution mode. The above execution mode represents the number of rounds of traversal and the operation mode corresponding to each round of traversal. In this way, after obtaining the processing instruction, the processor 120 can learn the number of rounds of traversal that needs to be performed and the operation mode performed during each round of traversal by querying the above correspondence.

[0073] In some embodiments, the computing device 100 can be pre-deployed with a trained target model. The target model can analyze the processing instruction and identify the execution mode corresponding to the processing instruction (that is, the number of rounds of traversal that needs to be performed and the operation mode performed during each round of traversal). In this way, when the processor 120 obtains the processing instruction and inputs the processing instruction into the target model, the target model can output the execution mode corresponding to the processing instruction.

[0074] Each batch of data includes some associated data in the target data. Multiple batches of data together constitute the target data. The data volume of each batch of data in the target data is less than the capacity of the memory, so as to ensure that a batch of data can be stored in the memory. In some embodiments, the memory includes a target storage space for storing each batch of data. The capacity of the target storage space can be less than the memory capacity and greater than or equal to the data volume of a batch of data. The data volume of each batch of data in the target data can be determined based on the capacity of the target storage space in the memory. In some embodiments, the processor 120 can obtain the capacity of the target storage space and read the associated data with a data volume less than or equal to this capacity from the external memory 101 each time. The associated data read each time can be called a batch of data. For example, if the capacity of the target storage space is 5GB, the processor 120 can read the associated data from the external memory 101 in a data volume of 5GB each time. The data volumes of different batches of data can be the same. For example, when the processor 120 just reads the target data completely by reading the associated data with the same data volume each time, the data volumes of multiple batches of data are the same. Taking a 128 * 4 million matrix as an example, this matrix is split into 10,000 batches of data, and the data volume of each batch of data is 128 * 400. There may also be batches of data with different data volumes among multiple batches of data. For example, when the data volume of the associated data read by the processor 120 for the last time is not enough for the required data volume, the data volume of the last batch of data is different from that of other batches of data.

[0075] In some embodiments, the processor 120 can specifically apply for a storage space in the memory as the target storage space and determine the data volume of each batch of data based on the capacity of the target storage space. By specifically applying for a storage space to store each batch of data, it avoids confusing the target data with other data.

[0076] In some embodiments, the target storage space can be a storage space that is preset and allocated in the memory and is dedicated to storing a batch of data. In this case, the capacity of the target storage space is a preset value. In some embodiments, the target storage space can be a storage space dynamically allocated by the processor 120 in the memory based on the processing instruction. For example, after the processor 120 obtains the processing instruction corresponding to the target data, it can apply for a storage space in the memory as the target storage space. Furthermore, the processor 120 adaptively determines the data volume of each batch of data based on the capacity of the target storage space. Thus, it is ensured that each batch of data can be stored in the memory and the smooth execution of the data processing method is ensured.

[0077] An example of the way for the processor 120 to dynamically allocate the target storage space is described below. Those skilled in the art can understand that during the execution of processing instructions, the processor 120 usually needs to buffer some data and also buffer some intermediate result data during the execution of operations (such as the aggregation value and the maximum value mentioned in the following description). In some embodiments, the processor 120 can, based on the processing instruction, first predict the first buffer capacity required to buffer other data (such as the aggregation value and the maximum value, etc.) in addition to each batch of data during the execution of the processing instruction. Further, the processor 120 can obtain the current free capacity of the memory and use the difference between the current free capacity and the first buffer capacity as the second buffer capacity. Then, the processor 120 applies for a storage space in the memory that is less than or equal to the second buffer capacity as the target storage space.

[0078] In the above solution, the processor 120 can adaptively apply for a target storage space with an appropriate capacity based on the current free capacity of the memory and the processing instruction to be executed. On the one hand, this makes the data processing method provided in this specification applicable to more data processing scenarios and has wide application. On the other hand, since the target storage space is adaptively allocated, it can ensure the sequential execution of the data processing process and avoid the problem of insufficient memory space during the data processing process due to a relatively large pre-allocated target storage space. On this basis, the processor 120 can allocate as large a target storage space as possible within the limit of the second buffer capacity, thereby reducing the number of times of reading in batches from the external memory and improving the data processing efficiency.

[0079] In some embodiments, the data volume of each batch of data can also be determined according to other methods. For example, the user inputs a splitting instruction in the computing device 100 to indicate the data volume of the associated data read from the external memory 101 by the computing device 100 each time, and the processor 120 can determine the data volume of each batch of data based on this splitting instruction.

[0080] Figure 4 Shows a schematic diagram of performing operations on each batch of data according to some embodiments of this specification. As Figure 4 shown, the target data is split into n data batches (n batches of data), and the processor 120 performs the addition operation of SUM and the maximum value operation of MAX on these n data batches.

[0081] The operation mode performed on each data batch in each round of traversal is related to the processing instruction. For example, when the processing instruction is a normalization operation, the operation mode includes an addition operation and a division operation, and when the processing instruction is a weighted summation operation, the operation mode includes an addition operation and a multiplication operation.

[0082] In the embodiment of the present specification, when the memory cannot store large-scale data, but needs to directly process the large-scale data, the method of the present specification can store the large-scale data in the external memory 101, and load the data batch by batch from the external memory 101 into the internal memory to perform operations, thereby realizing the processing of large-scale data. In particular, for long text sequences, the corresponding large-scale matrix (such as the attention score matrix) can be stored in the external memory 101, and then loaded into the internal memory in batches for processing, so that the processing of each word takes into account the influence of the entire long text sequence on it, so the captured semantic relationship is more accurate. If the long text sequence is divided into multiple short texts, the computing device 100 memory can store a small-scale matrix corresponding to a short text, then, when the computing device 100 independently processes the small-scale matrix (such as the attention score matrix) corresponding to each short text, the processing of each word only takes into account the local influence of the short text to which it belongs, and does not consider the global influence, so the captured semantic relationship is poor.

[0083] The specific operation process of the data processing method P100 is introduced below. The processor 120 can perform two rounds of traversal on the target data (such as the attention score matrix) based on the processing instruction (such as the normalization operation), and the process of each round of traversal includes reading the target data from the external memory to the internal memory in batches to perform operations, wherein the operation method corresponding to the first round of traversal includes at least an addition operation, and the operation method corresponding to the second round of traversal includes at least a division operation.

[0084] In some embodiments, the operation method corresponding to the first round of traversal also includes exponential operation. The first round of traversal may include the following process. The processor 120 reads the current batch of data from the external memory 101 into the internal memory. For each associated data in the current batch of data, the processor 120 performs an exponential operation in the internal memory based on the associated data to obtain the corresponding exponential data. Furthermore, the processor 120 performs an addition operation in the internal memory based on each exponential data in the current batch of data and the aggregate value to obtain an updated aggregate value.

[0085] Among them, the aggregated value is initialized to 0 before the first round of traversal. During the traversal process, the aggregated value is continuously updated. The current batch of data refers to any batch of data that the processor 120 is processing. In some embodiments, for each exponential data corresponding to the current batch of data, the processor 120 can add each exponential data to the aggregated value one by one. Each time the aggregated value is added, it is updated until the addition operation is performed on each exponential data corresponding to the current batch of data, that is, the updated aggregated value corresponding to the current batch of data is obtained. Of course, the processor 120 can also add each exponential data and add the result of adding each exponential data to the aggregated value, so as to obtain the updated aggregated value corresponding to the current batch of data. It should be noted that the aggregated value that performs the addition operation with the exponential data of each batch of data is the updated aggregated value corresponding to the previous batch of data. If each associated data is a vector containing multiple values, the exponential data and the aggregated value are both vectors with the same dimension as the associated data. The aggregated value obtained after the first round of traversal is the result of adding all the exponential data.

[0086] In the embodiments of this specification, the aggregated value obtained after the first round of traversal is the result of aggregating all associated data, integrating the features of all data units. Therefore, when processing the target data unit based on the aggregated value, the influence of all data units can be considered, so that richer information can be captured, making the processing result of the target data unit have strong expressive power.

[0087] In some embodiments, the exponential data includes the data obtained by performing the exponential operation on the difference between the maximum value of the associated data and the target data. When each associated data is a vector containing multiple values, the difference between the associated data and the maximum value in the target data means: the difference between the current value in the associated data and the maximum value in its dimension, and this dimension refers to the dimension in the target data that contains the values of different associated data. For example, the target data is a 128 * 4 million matrix, and each column is an associated data, that is, each associated data is a vector with a dimension of 128 * 1. For a certain value of a certain associated data, its corresponding maximum value is the maximum value in its row, and there are 4 million associated data values in this row. Correspondingly, the aggregated value obtained after the first round of traversal can be represented by where, refers to the exponential data obtained by performing the exponential operation on the difference between the jth associated data z j and the maximum value z max in the target data, represents performing an addition operation on the exponential data corresponding to n associated data. Such as Figure 4As shown, the operations performed on each data batch include the operation of taking the maximum value MAX. By subtracting the maximum value from the associated data, it is possible to prevent overflow errors caused by the exponential data being too large to exceed the range that the computing device 100 can represent, thereby improving numerical stability.

[0088] In some embodiments, the operation mode corresponding to the second round of traversal further includes an exponential operation, and the second round of traversal may include the following process. The processor 120 reads the current batch of data from the external memory 101 into the memory. For each associated data in the current batch of data, the processor 120 performs the exponential operation on the current associated data in the memory to obtain the corresponding exponential data, and performs the division operation based on the exponential data and the aggregated value after the first round of traversal to obtain the processing result (such as a normalization result (such as an attention weight)) corresponding to the current associated data.

[0089] The exponential operation performed in the second round of traversal may be the same as the exponential operation performed in the first round of traversal. The exponential data may be the data obtained by performing an exponential operation on the associated data, or the data obtained by performing an exponential operation on the difference between the associated data and the maximum value of the target data. The maximum value of the target data has been introduced above and will not be elaborated here.

[0090] In some embodiments, the processor 120 may obtain the normalized data through formula (1):

[0091]

[0092] where z i represents the i-th associated data, z max represents the maximum value in the target data, represents the exponential data obtained by performing an exponential operation on the difference between the i-th associated data and the maximum value of the target data, represents the aggregated value after the first round of traversal, and Softmax(z i ) represents the normalization result of the i-th associated data.

[0093] The data volume of the exponential data corresponding to the large-scale data is also very large. Therefore, when the processor 120 calculates the exponential data of each associated data in the first round of traversal, it does not store it, but recalculates the exponential data in the second round of traversal. The speed at which the processor 120 performs calculations by itself is much faster than the speed of reading data from the external memory 101. Therefore, compared with storing the large-scale exponential data in the external memory 101 and reading it from the external memory 101 when needed, calculating the exponential data as needed is faster and more efficient. Of course, in order to reduce the computing pressure on the processor 120, the processor 120 can also store the large-scale exponential data calculated in the first round of traversal in the external memory 101 and read it from the external memory 101 when used in the second round of traversal. The embodiments of this specification do not limit this.

[0094] In some embodiments, as described above, the Transformer model further includes a V matrix. The processing result can be multiple normalized data corresponding to multiple associated data. The processing result is used for the enhanced representation process of the target data unit, and the enhanced representation process includes: using the multiple normalized data as weights (such as attention weights) to perform a weighted sum on the V matrix to obtain the enhanced representation result of the target data unit, and the enhanced representation result fuses the information of the multiple data units.

[0095] Taking the task of translating a text sequence of 4 million words processed by the attention mechanism in the Transformer architecture as an example, the 4 million words can be converted into 4 million input feature representations, such as 4 million word embeddings. The 4 million word embeddings can generate a Q matrix, a K matrix, and a V matrix through linear transformation, and each matrix includes the vectors corresponding to the 4 million word embeddings. For one target word (target data unit), the q vector of the target word is respectively dot-producted with the 4 million k vectors to obtain 4 million attention scores, that is, the attention score matrix. The processor 120 performs a softmax operation on the attention score matrix to obtain 4 million attention weights, that is, the attention weight matrix. Each normalized data is a number between 0 and 1, and the sum of the 4 million normalized data is 1. The processor 120 can use the 4 million attention weights as weights and perform weighted summation on the 4 million v vectors to obtain the enhanced representation result of the target word. The information of 4 million words is fused in the enhanced representation result of the target word. The processor 120 can process the 4 million words according to the processing process of the target word to obtain 4 million enhanced representation results. Furthermore, the processor 120 can execute the subsequent processes in the Transformer architecture based on the enhanced representation results, such as inputting the 4 million enhanced representation results into a feed-forward neural network for processing, and finally obtaining the translated text sequence. Since the information of the entire text sequence is fused in the enhanced representation result of each word, the translation of the text sequence is more accurate.

[0096] In some embodiments, the data processing method P100 can be executed only by a single computing device 100. That is, this specification can implement the processing of large-scale data through a single machine and external storage, without the need for a network, reducing the cost caused by network communication, which is simple and efficient.

[0097] In summary, for the data processing method P100 and system provided in this specification, target data with a data volume larger than the memory capacity of the computing device is stored in the external storage 101 of the computing device. When the computing device obtains a processing instruction corresponding to the target data, based on the processing instruction, the target data is read in batches from the external storage into the memory for operation to obtain the processing result corresponding to the target data. By storing large-scale data in the external storage 101 and loading it into the memory in batches for processing, the computing device can process data with a data volume larger than its memory.

[0098] On the other hand, this specification provides a non-transitory storage medium storing at least one set of executable instructions for data processing. When the executable instructions are executed by a processor, the executable instructions direct the processor to perform the steps of the data processing method P100 described in this specification. In some possible implementation manners, each aspect of this specification can also be implemented in the form of a program product, which includes program code. When the program product runs on a data processing system, the program code is used to cause the data processing system to perform the steps of the data processing method P100 described in this specification. The program product for implementing the above method can adopt a portable compact disc read-only memory (CD-ROM) including program code and can run on a data processing system. However, the program product of this specification is not limited thereto. In this specification, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system. The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and the readable medium can send, propagate, or transmit a program used by or combined with an instruction execution system, apparatus, or device. The program code included on the readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification can be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, C++, etc., and also including conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the data processing system, partially on the data processing system, executed as an independent software package, partially on the data processing system and partially on a remote computing device, or entirely on the remote computing device.

[0099] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0100] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and is not necessarily limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to embrace various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.

[0101] In addition, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.

[0102] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, drawing, or its description. However, this does not mean that the combination of these features is necessary, and it is entirely possible for those skilled in the art, when reading this specification, to mark out some of the devices as separate embodiments for understanding. That is to say, the embodiments in this specification can also be understood as an integration of multiple sub - embodiments. And it also holds when the content of each sub - embodiment contains less than all the features of a single foregoing disclosed embodiment.

[0103] Each patent, patent application, published patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, items, etc., except for any historical prosecution documents associated therewith, any identical ones that may be inconsistent or in conflict with this document, or any identical historical prosecution documents that may have a limiting effect on the broadest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter associated with this document. Further, if there is any inconsistency or conflict between the description, definition, and / or use of a term associated with any of the included materials and the terms, descriptions, definitions, and / or uses associated with this document, the terms of this document shall control.

[0104] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art may adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.

Claims

1. A data processing method, applied to a computing device, comprising: Obtaining a processing instruction corresponding to target data, wherein the target data is stored in an external memory of the computing device, and the amount of the target data is greater than the capacity of the memory of the computing device; as well as Based on the processing instructions, the target data is read from the external memory into the internal memory in batches to perform operations, and processing results corresponding to the target data are obtained. The data volume of each batch of data in the target data is smaller than the capacity of the memory, and the operation method performed on each batch of data is related to the processing instruction.

2. The method of claim 1, wherein: The processing instruction represents a normalization operation to be performed on the target data.

3. The method of claim 1, wherein: The target data is data processed by the Q matrix and the K matrix respectively, and the Q matrix and the K matrix are the Q matrix and the K matrix in the Transformer model.

4. The method of claim 3, wherein: The Q matrix and the K matrix are generated based on multiple data units, the target data corresponds to a target data unit among the multiple data units, and the target data includes multiple associated data, each associated data represents the degree of association between each data unit and the target data unit.

5. The method of claim 4, wherein: The step of reading the target data from the external memory to the internal memory in batches to perform operations based on the processing instructions includes: Based on the processing instructions, two rounds of traversal are performed on the target data, and the process of each round of traversal includes reading the target data from the external memory into the internal memory in batches to perform operations, wherein the operation method corresponding to the first round of traversal includes at least addition operations, and the operation method corresponding to the second round of traversal includes at least division operations.

6. The method of claim 5, wherein: The operation method corresponding to the first round of traversal also includes exponential operation, and the process of the first round of traversal includes: Reading the current batch of data from the external memory into the internal memory; For each associated data in the current batch of data, performing an exponential operation based on the associated data in the memory to obtain corresponding exponential data; and The addition operation is performed in the memory based on each index data in the current batch data and the aggregate value to obtain an updated aggregate value, wherein the aggregate value is initialized to 0 before the first round of traversal.

7. The method of claim 6, wherein: The index data includes data obtained by performing the exponential operation on a difference between the associated data and a maximum value in the target data.

8. The method of claim 6, wherein: The operation method corresponding to the second round of traversal also includes exponential operation, and the process of the second round of traversal includes: Reading the current batch of data from the external memory into the internal memory; and For each associated data in the current batch of data: performing the exponential operation based on the current associated data in the memory to obtain corresponding exponential data, and The division operation is performed based on the index data and the aggregate value after the first round of traversal is completed to obtain a processing result corresponding to the current associated data.

9. The method of claim 4, wherein: The Transformer model also includes a V matrix, The processing result includes: a plurality of normalized data corresponding to the plurality of associated data, The processing result is used for the enhanced representation process of the target data unit, and the enhanced representation process includes: using the multiple normalized data as weights, performing weighted summation on the V matrix, and obtaining the enhanced representation result of the target data unit, and the enhanced representation result integrates the information of the multiple data units.

10. The method of claim 4, wherein: The target data unit is one of text, image or audio.

11. The method of claim 1, wherein: The external memory includes a built-in external memory and / or a removable external memory.

12. The method of claim 1, wherein: The memory includes a target storage space for storing each batch of data, and the data volume of each batch of data is less than or equal to the capacity of the target storage space.

13. The method of claim 12, wherein: The method further comprises: Applying for a storage space in the memory as the target storage space; and The data volume of each batch of data is determined based on the capacity of the target storage space.

14. The method of claim 13, wherein: Applying for a storage space in the memory as the target storage space includes: Based on the processing instruction, predicting a first buffer capacity required for buffering other data except the batches of data during the execution of the processing instruction; Obtaining a current free capacity of the memory, and using a difference between the current free capacity and the first buffer capacity as a second buffer capacity; and A storage space smaller than or equal to the second buffer capacity is applied for in the memory as the target storage space.

15. A data processing system comprising: at least one storage medium storing at least one instruction set for implementing data processing; as well as at least one processor, in communication with the at least one storage medium, When the data processing system is running, the at least one processor reads the at least one instruction set and implements the data processing method according to any one of claims 1 to 14.