Data processing device and method, NPU, equipment and storage medium
By batch reading and processing of data, using the difference value and variance calculation methods, the problems of low data processing efficiency and poor stability in the prior art are solved, and efficient and stable data calculation is achieved.
Patent Information
- Application Number
- CN202510333458.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-19
AI Technical Summary
In the prior art, since data cannot be fully loaded into the on-chip memory of the chip, data needs to be read from off-chip memory multiple times, resulting in low data processing efficiency and unstable calculation results, especially when the variance is small, the system error is large.
By batch reading of data and processing in on-chip memory, using the difference value and variance calculation method, k data are read from off-chip memory at a time, and iterative calculations are performed based on these differences value and intermediate mean and variance to reduce the number of reads and improve stability.
It improves the efficiency and stability of the data processor, reduces the number of off-chip memory reads, ensures the stability and accuracy of the calculation results, and avoids system errors.
Smart Images

Figure CN120277026A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a data processing device, method, NPU, device, and storage medium. Background Art
[0002] Currently, in the field of artificial intelligence technology, when calculating a large amount of data, since the data cannot be fully loaded into the on-chip memory of the chip, the data must be read from the off-chip memory and then calculated.
[0003] In the related art, both the mean and variance are calculated using the defined formula. Then it is necessary to first read the data to calculate the mean, and then read the data again to calculate the variance after obtaining the specific value of the mean. In this way, it is necessary to read the off-chip memory twice, that is, to read the data multiple times during the process of processing a large amount of data, resulting in low efficiency in processing data. In addition, the variance calculation is performed in the manner of E(X^2)-(EX)^2, and the data can be read only once, calculate the mean of the data and the mean of the square of the data, and then calculate the variance. Although this calculation method only needs to read the data once, when the actual variance is small, the systematic error in the calculation will cause a large error in the final calculation result, and thus the calculated value is very unstable. Summary of the Invention
[0004] Embodiments of this application provide a data processing device, method, NPU, device, and storage medium, which can batch-read data and batch-process data, thereby reducing the number of times of reading data from the off-chip memory, and having stable values, being convenient for hardware implementation, and thus improving the data processing efficiency and stability of the processor or chip. The technical solution is as follows:
[0005] According to the first aspect of the embodiments of this application, a data processing device is provided. The device includes:
[0006] An off-chip memory for storing a plurality of data;
[0007] A read-write module for reading k of the data from the off-chip memory for the (i + 1)-th read;
[0008] An on-chip memory for storing k of the data;
[0009] A processor, configured to obtain k first differences based on the k data; each of the first differences indicates the difference between one of the data and the i-th intermediate mean value; obtain the (i + 1)-th intermediate mean value based on the k first differences and the i-th intermediate mean value; obtain k second differences; each of the second differences indicates the difference between one of the data and the (i + 1)-th intermediate mean value; obtain the (i + 1)-th intermediate variance based on the k first differences, the k second differences and the i-th intermediate variance; k and i are integers greater than or equal to 1.
[0010] The on-chip memory is further configured to store the (i + 1)-th intermediate mean value and the (i + 1)-th intermediate variance.
[0011] In a possible implementation, the obtaining the (i + 1)-th intermediate mean value based on the k first differences and the i-th intermediate mean value includes:
[0012] Obtain the sum of the (i + 1)-th first differences; the sum of the (i + 1)-th first differences is the sum of the k first differences;
[0013] Obtain the (i + 1)-th average difference; the (i + 1)-th average difference is obtained based on the sum of the first differences and ((i + 1) × k);
[0014] Obtain the i-th intermediate mean value from the on-chip memory;
[0015] Obtain the (i + 1)-th intermediate mean value based on the i-th intermediate mean value and the (i + 1)-th average difference;
[0016] Write the (i + 1)-th intermediate mean value into the on-chip memory.
[0017] In a possible implementation, the obtaining the (i + 1)-th intermediate mean value based on the i-th intermediate mean value and the (i + 1)-th average difference includes:
[0018] Obtain the (i + 1)-th intermediate mean value based on the sum of the i-th intermediate mean value and the (i + 1)-th average difference.
[0019] In a possible implementation, the obtaining the (i + 1)-th intermediate variance based on the k first differences, the k second differences and the i-th intermediate variance includes:
[0020] Obtain k difference products; each of the difference products is the product of one of the first differences and one of the second differences;
[0021] Obtain the (i + 1)-th difference product sum; the (i + 1)-th difference product sum is the sum of k difference products;
[0022] Read the i-th intermediate variance from the on-chip memory;
[0023] Based on the i-th intermediate variance and the (i + 1)-th difference product sum, obtain the (i + 1)-th intermediate variance;
[0024] Write the (i + 1)-th intermediate variance into the on-chip memory.
[0025] In a possible implementation manner, the obtaining the (i + 1)-th intermediate variance based on the i-th intermediate variance and the (i + 1)-th difference product sum includes:
[0026] Based on the sum of the i-th intermediate variance and the (i + 1)-th difference product sum, obtain the (i + 1)-th intermediate variance.
[0027] In a possible implementation manner, the processor is further configured to:
[0028] Based on the (i + 1)-th intermediate variance, obtain a target intermediate variance;
[0029] Based on the quotient of the target intermediate variance and the total quantity, obtain a target variance; the total quantity indicates the total quantity of the data.
[0030] In a possible implementation manner, the processor is further configured to write the target mean value and the target variance into the on-chip memory;
[0031] The reading and writing module is further configured to read the target variance and the target mean value from the on-chip memory, and write the target variance and the target mean value into the off-chip memory.
[0032] According to a second aspect of the embodiments of the present application, a data processing method is provided. The method includes:
[0033] For the (i + 1)-th read, read k pieces of the data from the off-chip memory; write the k pieces of the data into the on-chip memory;
[0034] Based on the k pieces of data, obtain k first differences; each of the first differences indicates the difference between one piece of the data and the i-th intermediate mean; based on the k first differences and the i-th intermediate mean, obtain the (i + 1)-th intermediate mean; obtain k second differences; each of the second differences indicates the difference between one piece of the data and the (i + 1)-th intermediate mean; based on the k first differences, the k second differences, and the i-th intermediate variance, obtain the (i + 1)-th intermediate variance; k and i are integers greater than or equal to 1.
[0035] Store the (i + 1)-th intermediate mean and the (i + 1)-th intermediate variance into the on-chip memory.
[0036] According to the third aspect of the embodiments of the present application, an NPU is provided. The NPU includes:
[0037] A read-write module, configured to read k pieces of data from an off-chip memory for the (i + 1)-th read;
[0038] An on-chip memory, configured to store the k pieces of data;
[0039] A processor, configured to based on the k pieces of data, obtain k first differences; each of the first differences indicates the difference between one piece of the data and the i-th intermediate mean; based on the k first differences and the i-th intermediate mean, obtain the (i + 1)-th intermediate mean; obtain k second differences; each of the second differences indicates the difference between one piece of the data and the (i + 1)-th intermediate mean; based on the k first differences, the k second differences, and the i-th intermediate variance, obtain the (i + 1)-th intermediate variance; k and i are integers greater than or equal to 1.
[0040] The on-chip memory is further configured to store the (i + 1)-th intermediate mean and the (i + 1)-th intermediate variance.
[0041] According to the fourth aspect of the embodiments of the present application, a computer device is provided. The computer device includes a processor and a memory. The memory is configured to store at least one program, and the at least one program is loaded and executed by the processor to perform a data processing method.
[0042] According to the fifth aspect of the embodiments of the present application, a computer-readable storage medium is provided. At least one program is stored in the computer-readable storage medium, and the at least one program is loaded and executed by a processor to implement a data processing method.
[0043] In an embodiment of the present application, an embodiment of the present application provides a data processing method. Each time, a batch of data is read from an off-chip memory and loaded into an on-chip memory. The processor obtains k data from the on-chip memory each time and stably processes these k data to obtain a stable mean and variance. Thus, data can be read in batches, processed in batches, and then iteratively processed without an upper limit, reducing the number of times of reading data from the off-chip memory and processing data, improving the data processing efficiency of the processor, and at the same time ensuring that the number of operation instructions does not expand. Moreover, since the values in the iterative process are stable, the error of the final result is also within the allowable range. In addition, the embodiment of the present application calculates the variance of a batch of data each time through an iterative method, so it is not necessary to calculate through the variance formula E(X^2)-(EX)^2, thereby solving the problem of large systematic errors, improving the stability of the calculated target variance, facilitating hardware implementation, and improving the calculation performance and stability of the processor or chip. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0045] Figure 1 is a schematic diagram of an implementation environment provided according to an embodiment of the present application;
[0046] Figure 2 is a schematic structural diagram of a data processing device provided according to an embodiment of the present application;
[0047] Figure 3 is a schematic flowchart of a data processing method provided according to an embodiment of the present application;
[0048] Figure 4 is a schematic structural diagram of an NPU provided according to an embodiment of the present application;
[0049] Figure 5 is a schematic structural diagram of a terminal provided according to an embodiment of the present application;
[0050] Figure 6 is a schematic structural diagram of a server provided according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.
[0052] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.
[0053] In the present application, terms such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms.
[0054] These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first action can be called the second action, and similarly, the second action can also be called the first action. The first action and the second action can both be actions, and in some cases, they can be separate and different actions.
[0055] Among them, at least one means one or more than one. For example, at least one action can be one action, two actions, three actions, etc., any integer greater than or equal to one. And multiple means two or more than two. For example, multiple actions can be two actions, three actions, etc., any integer greater than or equal to two.
[0056] Figure 1 is a schematic diagram of an implementation environment provided according to an embodiment of the present application. The implementation environment may include a terminal 101 and a server 102.
[0057] In the terminal 101, a data processing device for data processing is provided. The data processing device at least includes an on-chip memory, an off-chip memory, a read-write module, and a processor, etc.
[0058] The terminal 101 can be a smart phone with a data processing device, a wearable device, a personal computer, a laptop computer, a tablet computer, a smart TV, a vehicle-mounted terminal, etc.
[0059] The server 102 can be a single server, a server cluster composed of multiple servers, or, alternatively, a cloud processing center.
[0060] The terminal 101 is connected to the server 102 through a wired or wireless network.
[0061] In some embodiments, a wireless network or a wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to any combination of a Local Area Network (LAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a mobile, wired or wireless network, a private network, or a virtual private network. In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged through the network. Additionally, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or proprietary data communication technologies can be used to replace or supplement the above data communication technologies.
[0062] In the related art, the intermediate mean is mainly calculated by the following two methods. The first is to calculate according to the definition formula, that is, calculate the sum of all data divided by the data volume, and the intermediate mean is obtained. The second is to calculate by the Welford algorithm. During the calculation process, the intermediate mean is iteratively updated, and each iteration processes one data.
[0063] In addition, the intermediate variance is mainly calculated by the following methods. The first is to calculate according to the definition; the second is to calculate according to the variance formula E(X^2)-(EX)^2; the third is to iteratively calculate by the Welford algorithm, and each calculation processes one data. Among them, when using the Welford algorithm for iterative calculation, although it can calculate the intermediate mean and intermediate variance of the data by only reading the data once under the condition of numerical stability, each iteration can only incrementally calculate one data, which is not suitable for calculation by a Neural Processing Unit (NPU).
[0064] To solve the above technical problems, the present application provides a data processing device, and the specific technical solution is as follows:
[0065] Figure 2 FIG. 13 is a schematic structural diagram of a data processing device 200 provided according to an embodiment of the present application. The device includes:
[0066] An off-chip memory 201, a read-write module 202, an on-chip memory, and a processor 204 that are electrically connected in sequence. Among them, the off-chip memory 201, the read-write module 202, the on-chip memory, and the processor 204 are all hardware structures.
[0067] In one example, the on-chip memory 203 (On-Chip Memory) refers to a memory integrated in the processor 204 or the chip, which has the characteristics of high-speed access and low latency. For example: In the embodiments of the present application, the on-chip memory 203 is implemented based on a Static Random Access Memory (SRAM).
[0068] In one example, the off-chip memory 201 refers to those storage devices separated from the main processor 204 or the main memory, and they are usually used to expand the storage capacity of the system or store data that needs to be retained for a long time. For example: In the embodiments of the present application, the off-chip memory 201 is implemented based on a hard disk.
[0069] In one example, the read-write module 202 is used to read data from the off-chip memory 201 and write the data into the on-chip memory. Or the read-write module 202 is used to read data from the on-chip memory and write the data into the off-chip memory 201. Among them, in the embodiments of the present application, the read-write module 202 is a hardware with read-write functions.
[0070] In one example, the processor 204 is used for streaming computing. Among them, streaming computing can be understood as calculating k data in the order in which k data flow through the processor 204. The specific functions of the processor 204 include four arithmetic operations and accumulation functions, etc. The processor 204 is used to read data from the on-chip memory, process the data, and write the result of the processed data into the on-chip memory. Among them, in the embodiments of the present application, the processor 204 is a hardware with functions such as supporting four arithmetic operations and accumulation.
[0071] In some embodiments, since a large amount of data stored in the off-chip memory 201 cannot be fully loaded into the on-chip memory, it is necessary to read a part of the data from the off-chip memory 201 each time. Specifically, the read-write module 202 is configured to read k data from the off-chip memory 201 for the (i + 1)-th read and store the k data in the on-chip memory. The processor 204 is configured to obtain k first differences based on the k data; each first difference indicates the difference between a data and the i-th intermediate mean; obtain the (i + 1)-th intermediate mean based on the k first differences and the i-th intermediate mean; obtain k second differences; each second difference indicates the difference between a data and the (i + 1)-th intermediate mean. Obtain the (i + 1)-th intermediate variance based on the k first differences, the k second differences, and the i-th intermediate variance; store the (i + 1)-th intermediate mean and the (i + 1)-th intermediate variance in the on-chip memory. k and i are integers greater than or equal to 1. Among them, the i-th intermediate mean can be understood as the mean of the first i×k data. The i-th intermediate variance can be understood as the variance of the first i×k data. Combining the above embodiments, it can be seen that due to the technical solution of the present application, a batch of data is read and processed each time, compared with the related art where one data is read and processed each time, the number of times of reading data is greatly reduced, thereby improving the efficiency of data processing. For example, k is 1000 or 10000.
[0072] In one example, the data processing device is started, and the intermediate mean, the intermediate variance, and i are initialized to 0. For example, the data involved in the embodiments of the present application is image data, audio data, video data, etc.
[0073] In some embodiments, obtaining the (i + 1)-th intermediate mean based on the k first differences and the i-th intermediate mean can be implemented in the following manner:
[0074] Obtain the sum of the (i + 1)-th first differences; the sum of the (i + 1)-th first differences is the sum of the k first differences. Obtain the (i + 1)-th average difference; the (i + 1)-th average difference is obtained based on the sum of the first differences and ((i + 1)×k). Obtain the i-th intermediate mean from the on-chip memory. Obtain the (i + 1)-th intermediate mean based on the i-th intermediate mean and the (i + 1)-th average difference. Write the (i + 1)-th intermediate mean into the on-chip memory. Among them, the (i + 1)-th average difference = sum of the first differences / ((i + 1)×k).
[0075] Combining the above implementation manner, it can be seen that the processor 204 obtains k data each time and stably processes the k data to obtain a stable mean, thereby iteratively processing the data without an upper limit. And since the values in the iterative process are stable, the error of the final result is also within the allowable range.
[0076] In some embodiments, the (i + 1)-th intermediate mean can be obtained based on the i-th intermediate mean and the (i + 1)-th average difference, and can be implemented by the following implementation method:
[0077] The (i + 1)-th intermediate mean is obtained based on the sum of the i-th intermediate mean and the (i + 1)-th average difference.
[0078] In one example, the (i + 1)-th intermediate mean can be calculated by the following formula:
[0079]
[0080] In the above formula, x m is the m-th data; is the i-th intermediate mean; is the sum of the (i + 1)-th average differences; is a first difference; is the (i + 1)-th average difference.
[0081] In some embodiments, the (i + 1)-th intermediate variance can be obtained based on k first differences, k second differences, and the i-th intermediate variance, and can be implemented by the following implementation method:
[0082] Obtain k products of differences; each product of differences is the product of a first difference and a second difference. Obtain the sum of the (i + 1)-th products of differences; the sum of the (i + 1)-th products of differences is the sum of k products of differences. Read the i-th intermediate variance from the on-chip memory. Based on the i-th intermediate variance and the sum of the (i + 1)-th products of differences, obtain the (i + 1)-th intermediate variance. Write the (i + 1)-th intermediate variance into the on-chip memory. Optionally, each product of differences is the product of a first difference and a second difference of the same data.
[0083] In some embodiments, the (i + 1)-th intermediate variance can be obtained based on the i-th intermediate variance and the sum of the (i + 1)-th products of differences, and can be implemented by the following implementation method:
[0084] The (i + 1)-th intermediate variance is obtained based on the sum of the i-th intermediate variance and the sum of the (i + 1)-th products of differences.
[0085] In one example, the (i + 1)-th intermediate variance can be calculated by the following formula:
[0086]
[0087] In the above formula, is the (i + 1)-th intermediate mean; is the sum of the (i + 1)-th products of differences; is a second difference value.
[0088] Through the above analysis, it can be seen that the embodiments of the present application calculate the variance of a batch of data iteratively each time, so that it is not necessary to calculate through the variance formula E(X^2)-(EX)^2, thereby solving the problem of large systematic errors and improving the stability of the target variance calculated by the processor or chip.
[0089] In some embodiments, the processor 204 is further configured to:
[0090] Obtain a target intermediate variance based on the (i + 1)-th intermediate variance;
[0091] Obtain a target variance based on the quotient of the target intermediate variance and the total quantity; the total quantity indicates the total quantity of data.
[0092] In one example, after the variance calculation of the last batch of data is completed, the last intermediate variance, that is, the target intermediate variance, is obtained. Calculating the quotient of the target intermediate variance and the total quantity can obtain the target variance. That is, the embodiments of the present application only perform one calculation of the quotient of the target intermediate variance and the total quantity.
[0093] In one example, the target variance can be calculated through the following formula.
[0094] Target variance = Target intermediate variance / Total quantity.
[0095] In one example, for the last data reading, if the quantity of data to be read is less than k, after reading the remaining data, "0" is supplemented so that the quantity of read data is equal to k. However, during the process of calculating the mean and variance of this batch of data, the quantity of data involved is calculated according to the actual quantity of data, that is, the quantity of "0" is not considered, so as to ensure the accuracy and stability of the mean and variance.
[0096] In some embodiments, the processor 204 is further configured to write the target mean and the target variance into the on-chip memory;
[0097] The read / write module 202 is further configured to read the target variance and the target mean from the on-chip memory and write the target variance and the target mean into the off-chip memory 201. Wherein, the target mean is the mean obtained after completing the mean calculation of all data.
[0098] In one example, the read-write module 202 waits for a preset duration when writing the last batch of data into the on-chip memory. The processor 204 calculates the target mean and the target variance, and writes the target mean and the target variance into the on-chip memory. After the waiting duration of the read-write module 202 reaches the preset duration, the read-write module 202 reads the target mean and the target variance from the on-chip memory, and writes the target mean and the target variance into the off-chip memory 201.
[0099] It should be noted that when the data processing device provided in the above embodiment executes the corresponding steps, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0100] In the embodiment of the present application, a batch of data is read from the off-chip memory each time, and the batch of data is loaded into the on-chip memory. The processor obtains k data from the on-chip memory each time and stably processes the k data to obtain a stable mean and variance. Thus, data can be read in batches and processed in batches, and then data can be iteratively processed without an upper limit, reducing the number of times of reading data from the off-chip memory and processing data, improving the data processing efficiency of the processor, and at the same time ensuring that the number of operation instructions does not expand. Moreover, since the values in the iterative process are stable, the error of the final result is also within the allowable range. In addition, the embodiment of the present application calculates the variance of a batch of data each time through an iterative method, so it is not necessary to calculate through the variance formula E(X^2)-(EX)^2, thereby solving the problem of large system error, improving the stability of the calculated target variance, facilitating hardware implementation, and improving the calculation performance and stability of the processor or chip.
[0101] Figure 3 is a schematic flowchart of a data processing method provided according to an embodiment of the present application. As Figure 3 shown, in the embodiment of the present application, it is described by taking an application to a terminal with a data processing device as an example. The method includes the following steps:
[0102] In step 301, for the (i + 1)-th read, the terminal reads k data from the off-chip memory; writes the k data into the on-chip memory.
[0103] In step 302, the terminal obtains k first differences based on the k data; each first difference indicates the difference between a data and the i-th intermediate mean; obtains k second differences; and obtains the (i + 1)-th intermediate variance based on the k first differences, the k second differences, and the i-th intermediate variance.
[0104] Among them, the (i + 1)-th intermediate mean is obtained based on k first differences and the i-th intermediate mean; each second difference indicates the difference between a data and the (i + 1)-th intermediate mean; k and i are integers greater than or equal to 1.
[0105] In step 303, the terminal stores the (i + 1)-th intermediate mean and the (i + 1)-th intermediate variance into the on-chip memory.
[0106] In some embodiments, obtaining the (i + 1)-th intermediate mean based on k first differences and the i-th intermediate mean includes:
[0107] Obtain the sum of the (i + 1)-th first differences; the sum of the (i + 1)-th first differences is the sum of k first differences;
[0108] Obtain the (i + 1)-th average difference; the (i + 1)-th average difference is obtained based on the sum of the first differences and ((i + 1)×k);
[0109] Obtain the i-th intermediate mean from the on-chip memory;
[0110] Based on the i-th intermediate mean and the (i + 1)-th average difference, obtain the (i + 1)-th intermediate mean;
[0111] Write the (i + 1)-th intermediate mean into the on-chip memory.
[0112] In some embodiments, obtaining the (i + 1)-th intermediate mean based on the i-th intermediate mean and the (i + 1)-th average difference includes:
[0113] Based on the sum of the i-th intermediate mean and the (i + 1)-th average difference, obtain the (i + 1)-th intermediate mean.
[0114] In some embodiments, obtaining the (i + 1)-th intermediate variance based on k first differences, k second differences and the i-th intermediate variance includes:
[0115] Obtain k products of differences; each product of differences is the product of a first difference and a second difference;
[0116] Obtain the sum of the (i + 1)-th products of differences; the sum of the (i + 1)-th products of differences is the sum of k products of differences;
[0117] Read the i-th intermediate variance from the on-chip memory;
[0118] Based on the i-th intermediate variance and the sum of the (i + 1)-th products of differences, obtain the (i + 1)-th intermediate variance;
[0119] Write the (i + 1)-th intermediate variance into the on-chip memory.
[0120] In some embodiments, obtaining the (i + 1)-th intermediate variance based on the i-th intermediate variance and the sum of the (i + 1)-th product of differences includes:
[0121] Obtaining the (i + 1)-th intermediate variance based on the i-th intermediate variance and the sum of the (i + 1)-th product of differences.
[0122] In some embodiments, the method further includes:
[0123] Obtaining a target intermediate variance based on the (i + 1)-th intermediate variance;
[0124] Obtaining a target variance based on the quotient of the target intermediate variance and the total quantity; the total quantity indicates the total quantity of data.
[0125] In some embodiments, the method further includes:
[0126] Writing the target mean value and the target variance into the on-chip memory;
[0127] Reading the target variance and the target mean value from the on-chip memory and writing the target variance and the target mean value into the off-chip memory.
[0128] It should be noted that the data processing method provided in the above embodiments and the embodiments of the data processing device belong to the same concept. For the specific implementation process, please refer to the device embodiments, which will not be elaborated here.
[0129] In the embodiments of the present application, a batch of data is read from the off-chip memory each time, and the batch of data is loaded into the on-chip memory. The processor obtains k data from the on-chip memory each time and stably processes these k data to obtain a stable mean value and variance. Thus, data can be read in batches and processed in batches, and then data can be iteratively processed without an upper limit, reducing the number of times of reading data from the off-chip memory and processing data, improving the data processing efficiency of the processor, and at the same time ensuring that the number of operation instructions does not expand. Moreover, since the values in the iterative process are stable, the error of the final result is also within the allowable range. In addition, the embodiments of the present application calculate the variance of a batch of data each time through an iterative method, so it is not necessary to calculate through the variance formula E(X^2)-(EX)^2, thereby solving the problem of large systematic errors, improving the stability of the calculated target variance, facilitating hardware implementation, and improving the calculation performance and stability of the processor or chip.
[0130] Figure 4 FIG. 32 is a schematic structural diagram of an NPU 400 according to an embodiment of the present application. The NPU 400 includes:
[0131] A reading and writing module 202, configured to read k data from the off-chip memory 201 for the (i + 1)-th reading;
[0132] On-chip memory 203 for storing k data;
[0133] Processor 204 for obtaining k first differences based on the k data; each first difference indicates the difference between a data and the i-th intermediate mean; obtaining the (i + 1)-th intermediate mean based on the k first differences and the i-th intermediate mean; obtaining k second differences; each second difference indicates the difference between a data and the (i + 1)-th intermediate mean; obtaining the (i + 1)-th intermediate variance based on the k first differences, the k second differences and the i-th intermediate variance; k and i are integers greater than or equal to 1;
[0134] The on-chip memory 203 is further configured to store the (i + 1)-th intermediate mean and the (i + 1)-th intermediate variance.
[0135] In some embodiments, obtaining the (i + 1)-th intermediate mean based on the k first differences and the i-th intermediate mean includes:
[0136] Obtaining the sum of the (i + 1)-th first differences; the sum of the (i + 1)-th first differences is the sum of the k first differences;
[0137] Obtaining the (i + 1)-th average difference; the (i + 1)-th average difference is obtained based on the sum of the first differences and ((i + 1)×k);
[0138] Obtaining the i-th intermediate mean from the on-chip memory;
[0139] Obtaining the (i + 1)-th intermediate mean based on the i-th intermediate mean and the (i + 1)-th average difference;
[0140] Writing the (i + 1)-th intermediate mean into the on-chip memory.
[0141] In some embodiments, obtaining the (i + 1)-th intermediate mean based on the i-th intermediate mean and the (i + 1)-th average difference includes:
[0142] Obtaining the (i + 1)-th intermediate mean based on the sum of the i-th intermediate mean and the (i + 1)-th average difference.
[0143] In some embodiments, obtaining the (i + 1)-th intermediate variance based on the k first differences, the k second differences and the i-th intermediate variance includes:
[0144] Obtaining k products of differences; each product of differences is the product of a first difference and a second difference;
[0145] Obtaining the sum of the (i + 1)-th products of differences; the sum of the (i + 1)-th products of differences is the sum of the k products of differences;
[0146] Read the i-th intermediate variance from the on-chip memory;
[0147] Based on the i-th intermediate variance and the (i + 1)-th difference product sum, obtain the (i + 1)-th intermediate variance;
[0148] Write the (i + 1)-th intermediate variance into the on-chip memory.
[0149] In some embodiments, based on the i-th intermediate variance and the (i + 1)-th difference product sum, obtaining the (i + 1)-th intermediate variance includes:
[0150] Based on the sum of the i-th intermediate variance and the (i + 1)-th difference product sum, obtain the (i + 1)-th intermediate variance.
[0151] In some embodiments, the processor is further configured to:
[0152] Based on the (i + 1)-th intermediate variance, obtain the target intermediate variance;
[0153] Based on the quotient of the target intermediate variance and the total quantity; the total quantity indicates the total quantity of data. Obtain the target variance;
[0154] In some embodiments, the processor is further configured to write the target mean and the target variance into the on-chip memory;
[0155] The read / write module is further configured to read the target variance and the target mean from the on-chip memory, and write the target variance and the target mean into the off-chip memory.
[0156] It should be noted that the NPU provided in the above embodiments belongs to the same concept as the data processing device embodiments. The specific implementation process can be found in the device embodiments and will not be elaborated here.
[0157] In the embodiments of the present application, each time a batch of data is read from the off-chip memory, the batch of data is loaded into the on-chip memory. The processor obtains k data from the on-chip memory each time and stably processes these k data to obtain a stable mean and variance. Thus, data can be read in batches, processed in batches, and then iteratively processed without an upper limit, reducing the number of times of reading and processing data from the off-chip memory, improving the data processing efficiency of the processor, and at the same time ensuring that the number of operation instructions does not expand. Moreover, since the values in the iterative process are stable, the error of the final result is also within the allowable range. In addition, the embodiments of the present application calculate the variance of a batch of data each time through an iterative method, so it is not necessary to calculate through the variance formula E(X^2)-(EX)^2, thereby solving the problem of large systematic errors, improving the stability of the calculated target variance, facilitating hardware implementation, and improving the computing performance and stability of the processor or chip.
[0158] An embodiment of the present application also provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the above method is implemented.
[0159] Taking the computer device as an example of a terminal, Figure 5 is a schematic structural diagram of a terminal provided by an embodiment of the present application. Refer to Figure 5 , the terminal 500 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The terminal 500 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0160] Generally, the terminal 500 includes: a processor 501 and a memory 502.
[0161] The processor 501 may include one or more processing cores, such as a 4-core processor, a 5-core processor, etc. The processor 501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 501 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 501 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0162] The memory 502 may include one or more computer-readable storage media, which may be non-transitory. The memory 502 may further include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 502 is used to store at least one program code, and the at least one program code is used to be executed by the processor 501 to implement the process executed by the terminal in the method provided in the method embodiments of the present application for the above-mentioned method.
[0163] In some embodiments, the terminal 500 may further optionally include: a peripheral device interface 503 and at least one peripheral device. The processor 501, the memory 502, and the peripheral device interface 503 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 503 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a display screen 504, a camera assembly 505, an audio circuit 506, and a power supply 507.
[0164] The peripheral device interface 503 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 501 and the memory 502. In some embodiments, the processor 501, the memory 502, and the peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 501, the memory 502, and the peripheral device interface 503 may be implemented on a separate chip or circuit board, and the embodiments of the present application do not limit this.
[0165] The display screen 504 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 504 is a touch display screen, the display screen 504 also has the ability to collect touch signals on or above the surface of the display screen 504. The touch signals can be input as control signals to the processor 501 for processing. At this time, the display screen 504 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 504, which is provided on the front panel of the terminal 500; in other embodiments, there may be at least two display screens 504, which are respectively provided on different surfaces of the terminal 500 or in a folding design; in other embodiments, the display screen 504 may be a flexible display screen, which is provided on the curved surface or folding surface of the terminal 500. Even, the display screen 504 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 504 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0166] The camera module 505 is used to collect images or videos. In some embodiments, the camera module 505 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 505 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0167] The audio circuit 506 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 501 for processing. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 500. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 501 into sound waves. The speaker may be a traditional thin-film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 506 may also include a headphone jack.
[0168] The power supply 507 is used to supply power to each component in the terminal 500. The power supply 507 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 507 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0169] Those skilled in the art can understand that Figure 5 the structure shown in
[0170] does not constitute a limitation on the terminal 500, and may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements. Figure 6 Taking a computer device as a server as an example,
[0171] Embodiments of the present application also provide a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method as described above. Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, optical data storage device, etc.
[0172] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0173] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A data processing device, characterized in that, Including: An off-chip memory for storing multiple data; A read / write module for reading k of the data from the off-chip memory for the (i + 1)-th read; An on-chip memory for storing k of the data; A processor for obtaining k first differences based on k of the data; each of the first differences indicates the difference between one of the data and the i-th intermediate mean; obtaining the (i + 1)-th intermediate mean based on k of the first differences and the i-th intermediate mean; obtaining k second differences; each of the second differences indicates the difference between one of the data and the (i + 1)-th intermediate mean; obtaining the (i + 1)-th intermediate variance based on k of the first differences, k of the second differences, and the i-th intermediate variance; k and i are integers greater than or equal to 1; The on-chip memory is further configured to store the (i + 1)-th intermediate mean and the (i + 1)-th intermediate variance.
2. The device according to claim 1, characterized in that, The obtaining the (i + 1)-th intermediate mean based on k of the first differences and the i-th intermediate mean includes: Obtaining the sum of the (i + 1)-th first differences; the sum of the (i + 1)-th first differences is the sum of k of the first differences; Obtaining the (i + 1)-th average difference; the (i + 1)-th average difference is obtained based on the sum of the first differences and ((i + 1)×k); Obtaining the i-th intermediate mean from the on-chip memory; Obtaining the (i + 1)-th intermediate mean based on the i-th intermediate mean and the (i + 1)-th average difference; Writing the (i + 1)-th intermediate mean into the on-chip memory.
3. The device according to claim 2, characterized in that, The obtaining the (i + 1)-th intermediate mean based on the i-th intermediate mean and the (i + 1)-th average difference includes: Obtaining the (i + 1)-th intermediate mean based on the sum of the i-th intermediate mean and the (i + 1)-th average difference.
4. The device according to claim 1, characterized in that, The obtaining the (i + 1)-th intermediate variance based on k of the first differences, k of the second differences, and the i-th intermediate variance includes: Obtaining k difference products; each of the difference products is the product of one of the first differences and one of the second differences; Obtaining the sum of the (i + 1)-th difference products; the sum of the (i + 1)-th difference products is the sum of k of the difference products; Reading the i-th intermediate variance from the on-chip memory; Obtaining the (i + 1)-th intermediate variance based on the i-th intermediate variance and the sum of the (i + 1)-th difference products; Writing the (i + 1)-th intermediate variance into the on-chip memory.
5. The device according to claim 4, characterized in that, The obtaining the (i + 1)-th intermediate variance based on the i-th intermediate variance and the (i + 1)-th difference includes: Obtaining the (i + 1)-th intermediate variance based on the sum of the i-th intermediate variance and the sum of the (i + 1)-th difference products.
6. The device according to claim 5, characterized in that, The processor is further configured to: Obtain a target intermediate variance based on the (i + 1)-th intermediate variance; Obtain a target variance based on the quotient of the target intermediate variance and the total quantity; The total quantity indicates the total quantity of the data.
7. The device according to claim 6, characterized in that, The processor is further configured to write the target mean value and the target variance into the on-chip memory; The reading and writing module is further configured to read the target variance and the target mean value from the on-chip memory and write the target variance and the target mean value into the off-chip memory.
8. A data processing method, characterized in that Comprising: For the (i + 1)-th reading, read k of the data from the off-chip memory; write the k data into the on-chip memory; Based on the k data, obtain k first differences; each of the first differences indicates the difference between one of the data and the i-th intermediate mean value; based on the k first differences and the i-th intermediate mean value, obtain the (i + 1)-th intermediate mean value; obtain k second differences; each of the second differences indicates the difference between one of the data and the (i + 1)-th intermediate mean value; based on the k first differences, the k second differences and the i-th intermediate variance, obtain the (i + 1)-th intermediate variance; k and i are integers greater than or equal to 1; Store the (i + 1)-th intermediate mean value and the (i + 1)-th intermediate variance in the on-chip memory.
9. An NPU, characterized in that, Comprising: A reading and writing module, configured to read k data from the off-chip memory for the (i + 1)-th reading; An on-chip memory, configured to store the k data; A processor, configured to obtain k first differences based on the k data; each of the first differences indicates the difference between one of the data and the i-th intermediate mean value; based on the k first differences and the i-th intermediate mean value, obtain the (i + 1)-th intermediate mean value; obtain k second differences; each of the second differences indicates the difference between one of the data and the (i + 1)-th intermediate mean value; based on the k first differences, the k second differences and the i-th intermediate variance, obtain the (i + 1)-th intermediate variance; k and i are integers greater than or equal to 1; The on-chip memory is further configured to store the (i + 1)-th intermediate mean value and the (i + 1)-th intermediate variance.
10. A computer device, characterized in that, The computer device includes a processor and a memory, the memory is configured to store at least one program, and the at least one program is loaded and executed by the processor to perform the data processing method according to claim 8.
11. A computer-readable storage medium, characterized in that, At least one program is stored in the computer-readable storage medium, and the at least one program is loaded and executed by a processor to implement the data processing method according to claim 8.
Citation Information
Patent Citations
Data standardization processing method and device, electronic equipment and storage medium
CN114356235A
Abnormality judgment basis processing method and abnormality judgment method and device
CN115690681A
Variance calculation method, device, circuit and equipment for true random number generator
CN116820402A
Hardware acceleration circuit, data processing acceleration method, chip and accelerator
CN117391158A