A sparse data processing method, apparatus, device, and storage medium
By grouping sparse data and processing it in turn, the problems of waste of resources and inefficiency during sparse data processing are solved, and more efficient memory utilization and data processing efficiency are achieved.
Patent Information
- Application Number
- CN202111162414.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-09-30
AI Technical Summary
When processing sparse data, directly performing calculations in memory lead to waste of resources and inefficiency. Traditional methods require loading all sparse data, occupying a large amount of memory.
By grouping sparse data and reading it in sequence, the input data points of each packet determine the output value and output location according to the processing strategy, and update the output data points in memory, and finally transfer the processed data to external memory.
It reduces the use of memory resources, improves the utilization efficiency of memory resources, and improves data processing efficiency through batch processing.
Smart Images

Figure CN113849290B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of data processing, and in particular, to a sparse data processing method, apparatus, device, and storage medium. Background Art
[0002] Due to the large sparsity of sparse data, that is, there are a large number of data points with a value of zero or missing, if calculations are directly performed based on the original sparse data, the system will consume a lot of useless work on the huge sparse data. At the same time, in the process of directly calculating based on the original sparse data, all sparse data needs to be loaded into the memory space, which not only has high requirements for hardware but also causes waste of system memory resources. Summary of the Invention
[0003] Embodiments of the present disclosure provide a sparse data processing method, apparatus, device, and storage medium.
[0004] In a first aspect, a sparse data processing method is provided, including: sequentially reading input data points of each of a plurality of groups corresponding to sparse data from an external memory to a first memory space; for the input data points of each group, based on the processing strategy corresponding to the group, determining an output value and an output position corresponding to the input data point, and using the output value to update a data value of an output data point corresponding to the output position in a second memory space; and transmitting the output data points in the second memory space as processed data to the external memory.
[0005] In some embodiments, the processing strategy includes a weight mapping strategy and an output mapping strategy; the determining an output value and an output position corresponding to the input data point based on the processing strategy corresponding to the group, and using the output value to update a data value of an output data point corresponding to the output position in a second memory space includes: generating at least one processing task based on the processing strategy corresponding to the group; the number of processing tasks is related to the number of output feature points corresponding to the input data points in the group; for each processing task, determining a sub-output value and a sub-output position corresponding to the input data point in the processing task based on a processing sub-strategy corresponding to the processing task; and using the sub-output value to update a data value of an output data point corresponding to the sub-output position in the second memory space.
[0006] In some embodiments, the processing sub - strategy includes a weight mapping strategy and an output mapping strategy; determining the sub - output value and sub - output position corresponding to the input data point in the processing task based on the processing sub - strategy corresponding to the processing task includes: obtaining the target weight corresponding to the processing task from the pre - loaded weight operators based on the weight mapping strategy corresponding to the processing task; the weight operators include multiple weights for processing the sparse data; determining the sub - output position based on the output mapping strategy corresponding to the processing task and the first position of the input data point in the sparse data; and determining the sub - output value based on the target weight and the data value of the input data point.
[0007] In some embodiments, updating the data value of the output data point corresponding to the sub - output position in the second memory space using the sub - output value includes: adding the sub - output value to the original data at the sub - output position in the second memory space to update the data value of the output data at the sub - output position in the second memory space.
[0008] In the embodiments of the present disclosure, since for the processing strategy corresponding to each group, the processing process of the input data points in the group is divided into at least one processing task, and each processing task can use the same processing sub - strategy to process all the input data points in the group. When the processing sub - strategy includes a weight mapping strategy and an output mapping strategy, full reuse of the target weight can be achieved, avoiding resource consumption caused by multiple weight readings; at the same time, the effect of batch - determining the sub - output positions corresponding to all the input data points in the group by the same output mapping strategy can also be achieved, improving the data processing efficiency.
[0009] In some embodiments, the method further includes: traversing the sparse data, grouping multiple input data points in the sparse data to obtain multiple groups; the processing strategies corresponding to the input data points in each group are the same; and storing the input data points corresponding to the multiple groups in the external memory.
[0010] In some embodiments, traversing the sparse data and grouping multiple input data points in the sparse data to obtain multiple groups includes: traversing the sparse data to determine the data information of each input data point; the data information includes the first position of the input data point in the sparse data; for each input data point, based on the first position of the input data point and a weight operator, determining the weight position corresponding to the input data point and output information; the weight position is the relative position of the target weight corresponding to the input feature point in the weight operator, and the output information includes the number of output data points corresponding to the input data point, and the mapping relationship between the first position of the input data point and the second position of each output data point; based on the weight position and output information of each input data point, grouping each input data point to obtain the multiple groups; the weight positions and output information corresponding to the input data points within each group are the same.
[0011] In the embodiments of the present disclosure, since before processing the sparse data, for each input data point in the sparse data, based on the relationship between the input data point and the weight position, the number of output data points corresponding to the input data point, and the mapping relationship between the first position of the input data point and the second position of each output data point, the input data points are grouped, it can be ensured that the input data points within the same group all correspond to the same processing strategy, providing a data basis for the subsequent batch processing of all input data points in each group, and indirectly improving the processing efficiency of subsequent processing of the sparse data.
[0012] In some embodiments, transferring the output data points in the second memory space as processed data to the external memory includes: in the case where a quantization process is required, quantizing the output data points in the second memory space and migrating them to a third memory space, and migrating the quantized output data points in the third memory space to the external memory; during the process of migrating the quantized output data points in the third memory space to the external memory, based on the second memory space, performing the processing process of the next sparse data; in the case where a quantization process is not required, while migrating the output data points in the second memory space to the external memory, based on the third memory space, performing the processing process of the next sparse data.
[0013] In some embodiments, the method further includes: partitioning the original data based on a preset partitioning rule to obtain at least one of the sparse data arranged in an ordered manner; there is no redundant data in the at least one sparse data; for each of the sparse data, the step of transmitting the output data points in the second memory space to the external memory as processed data includes: based on the boundary information of the sparse data, transmitting a first part of the data points in the second memory space to the external memory as processed data; the boundary information is determined by the partitioning rule; wherein, the output data of the second memory space further includes a second part of data points that have not been transmitted to the external memory; the data values of the second part of data points are updated during the process of processing the next adjacent sparse data of the sparse data.
[0014] In the embodiments of the present disclosure, on the basis of partitioning the original data, the sparse data corresponding to each partition is further grouped, which further reduces the occupation of memory resources in one processing process; at the same time, in the embodiments of the present disclosure, the first subspace and the second subspace in the third memory space alternately complete the two processes of reading data from the second memory space and transmitting data to the external memory, improving the data transmission efficiency.
[0015] In a second aspect, a sparse data processing device is provided, including: a reading module, configured to sequentially read input data points of each of a plurality of groups corresponding to sparse data from an external memory to a first memory space; a processing module, configured to, for the input data points of each group, determine an output value and an output position corresponding to the input data points based on a processing strategy corresponding to the group, and update the data value of the output data point corresponding to the output position in the second memory space by using the output value; a transmission module, configured to transmit the output data points in the second memory space to the external memory as processed data.
[0016] In some embodiments, the processing module is further configured to: generate at least one processing task based on the processing strategy corresponding to the group; the number of the processing tasks is related to the number of output feature points corresponding to the input data points in the group; for each of the processing tasks, determine a sub-output value and a sub-output position corresponding to the input data points in the processing task based on a processing sub-strategy corresponding to the processing task; update the data value of the output data point corresponding to the sub-output position in the second memory space by using the sub-output value.
[0017] In some embodiments, the processing sub-strategy includes a weight mapping strategy and an output mapping strategy; the processing module is further configured to: based on the weight mapping strategy corresponding to the processing task, obtain the target weight corresponding to the processing task from the pre-loaded weight operators; the weight operators include multiple weights for processing the sparse data; based on the output mapping strategy corresponding to the processing task, determine the sub-output position based on the first position of the input data point in the sparse data; determine the sub-output value based on the target weight and the data value of the input data point.
[0018] In some embodiments, the processing module is further configured to: accumulate the sub-output value onto the original data at the sub-output position in the second memory space to update the data value of the output data at the sub-output position in the second memory space.
[0019] In some embodiments, the sparse data processing device further includes a grouping module, and the grouping module is configured to: traverse the sparse data, group multiple input data points in the sparse data to obtain multiple groups; the processing strategies corresponding to the input data points in each group are the same; store the input data points corresponding to the multiple groups in the external memory.
[0020] In some embodiments, the grouping module is further configured to: traverse the sparse data, determine the data information of each input data point; the data information includes the first position of the input data point in the sparse data; for each input data point, determine the weight position and output information corresponding to the input data point based on the first position of the input data point and the weight operator; the weight position is the relative position of the target weight corresponding to the input feature point in the weight operator, and the output information includes the number of output data points corresponding to the input data point, and the mapping relationship between the first position of the input data point and the second position of each output data point; group each input data point based on the weight position and output information of each input data point to obtain the multiple groups; the weight positions and output information corresponding to the input data points within each group are the same.
[0021] In some embodiments, the transmission module is further configured to: when a quantization process is required, quantize the output data points in the second memory space and migrate them to the third memory space, and migrate the quantized output data points in the third memory space to the external memory; during the process of migrating the quantized output data points in the third memory space to the external memory, process the next sparse data based on the second memory space; when a quantization process is not required, while migrating the output data points in the second memory space to the external memory, process the next sparse data based on the third memory space.
[0022] In some embodiments, the grouping module is further configured to: partition the original data based on a preset chunking rule to obtain at least one of the sparse data arranged in an orderly manner; there is no redundant data in the at least one sparse data; for each sparse data, the transmission module is further configured to: based on the boundary information of the sparse data, transfer a first part of the data points in the second memory space as processed data to the external memory; the boundary information is determined by the chunking rule; wherein, the output data of the second memory space further includes a second part of the data points that have not been transferred to the external memory; the data values of the second part of the data points are updated during the process of processing the next adjacent sparse data of the sparse data.
[0023] In a third aspect, a sparse data processing device is provided, including: a memory and a processor, the memory stores a computer program that can run on the processor, and when the processor executes the computer program, the steps in the above method are implemented.
[0024] In a fourth aspect, a computer storage medium is provided, the computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the above method.
[0025] In the embodiments of the present disclosure, since during the process of processing sparse data, according to multiple groupings of the sparse data, the input data points of each grouping are sequentially read and processed, it is possible to process sparse data with a large data volume using a smaller memory. Compared with the conventional solution of loading all sparse data into the memory, the occupation of memory resources can be reduced, and the utilization efficiency of memory resources can be improved; at the same time, since the input data points in each grouping correspond to the same processing strategy, during the process of sequentially processing the input data points corresponding to each grouping, the same processing strategy can be used to batch process all the input data points of the current grouping, improving the processing efficiency. Description of the Drawings
[0026] Figure 1Schematic diagram of a sparse data processing system structure provided by an embodiment of the present disclosure;
[0027] Figure 2a Schematic flow diagram of a sparse data processing method provided by an embodiment of the present disclosure;
[0028] Figures 2b to 2e Schematic diagram of a pattern classification provided by an embodiment of the present disclosure;
[0029] Figure 3 Schematic flow diagram of a sparse data processing method provided by an embodiment of the present disclosure;
[0030] Figure 4 Schematic flow diagram of a sparse data processing method provided by an embodiment of the present disclosure;
[0031] Figure 5 Schematic diagram of the composition structure of a sparse data processing device provided by an embodiment of the present disclosure;
[0032] Figure 6 Schematic diagram of the hardware entity of a sparse data processing device provided by an embodiment of the present disclosure. Detailed implementation manners
[0033] The technical solutions of the present disclosure will be described in detail below through embodiments in combination with the accompanying drawings. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0034] It should be noted that in the examples of the present disclosure, "first", "second", etc. are used to distinguish similar objects, and do not have to be used to describe the order or sequence of the targets. In addition, the technical solutions described in the embodiments of the present disclosure can be combined arbitrarily without conflict.
[0035] Figure 1 Schematic flow diagram of a sparse data processing method provided by an embodiment of the present disclosure, which will be described in combination with Figure 1 the steps shown.
[0036] S101. Read the input data points of each of the multiple groups corresponding to the sparse data from the external memory into the first memory space in sequence;
[0037] Among them, the sparse data is data including a large number of data points with missing values or values of zero. Exemplarily, the sparse data can be image data collected by an imaging device, point cloud data collected by a radar device, text data obtained during text mining, or historical purchase records used in the process of constructing a customer portrait, etc.
[0038] It should be noted that the external memory (external storage) in the embodiments of the present disclosure refers to the memory in an electronic device other than the processor cache and the memory (local internal memory), and the external memory can still store data after power-off. Exemplarily, the external memory can be a magnetic memory, an optical memory, or an erasable flash memory. The memory (local internal memory) in the embodiments of the present disclosure refers to the random access memory (RAM), whose function is to temporarily store the operation data of the processor and store the data exchanged with the external memory. During the process of the electronic device processing sparse data, the processor will transfer the data to be operated to the memory for operation, and after the operation is completed, the processor will transmit the operation result.
[0039] In some embodiments, the sparse data has been divided into multiple groups and stored in the external memory. Among them, the sparse data can include multiple non-zero input data points, and after grouping, each group can include at least one input data point. During the process of processing the sparse data, since the sparse data is large, each group can be read sequentially and the data input points in the group can be processed to avoid occupying a large amount of memory.
[0040] In some embodiments, the memory of the electronic device can include a first memory space for storing the sparse data to be processed. Further, the first memory data is used to sequentially store the input data points of each group.
[0041] For example, the sparse data stored in the external memory has been divided into 3 groups, including the first group, the second group, and the third group. In the embodiments of the present disclosure, the electronic device will process the input data points of each group sequentially, that is, the electronic device will first read the input data points corresponding to the first group from the external memory into the first memory space, and after completing the processing of the first group, then read the input data points corresponding to the second group from the external memory into the first memory space; after completing the processing of the second group, then read the input data points corresponding to the third group from the external memory into the first memory space, and after completing the processing of the third group, the processing process of the sparse data is completed.
[0042] S102. For each input data point of each group, based on the processing strategy corresponding to the group, determine the output value and output position corresponding to the input data point, and use the output value to update the data value of the output data point corresponding to the output position in the second memory space;
[0043] In some embodiments, after reading the input data points of each group into the first memory space, the input data points of the first group can be processed based on the processing strategy corresponding to the group. Among them, the processing process of the input data points of the first group includes determining the output value and output position corresponding to the input data points; the output position is the relative position of the output data points affected by the input data points in the output data.
[0044] For example, after reading the input data point I1 into the first memory space, when the current processing process is convolution processing, and the convolution kernel size is 3×3 and the stride is 2, the output value corresponding to the input data point I1 can be determined based on the value of the input data point I1 and the convolution kernel; at the same time, based on the position (X, Y) of the input data point I1 in the sparse data, it can be determined that the output data point affected by the input data point I1 is one, and the relative position (X / 2, Y / 2) of the output data point O1 in the output data can also be determined.
[0045] It should be noted that the number of input data points in a group can be one or more. And, for an input data point, in the process of determining the output value and output position corresponding to the input data point, taking the input data point I1 as the above input data point, the output data point O1 corresponding to the input data point I1 can not only be affected by the output value corresponding to the input data point I1, but also may be affected by other input data points. Exemplarily, when the current processing process is convolution processing, and the convolution kernel size is 3×3 and the stride is 2, the output data point O1 may also be affected by 8 input data points around the input data point I1.
[0046] Therefore, after obtaining the output value and output position of an input data point in S102, only the data value of the output data point corresponding to the output position in the second memory space is updated using the output value, rather than directly using the output value as the data value of the output data point corresponding to the output position. Taking the above convolution processing or fully connected processing as an example, after obtaining the output value and output position of the input data point, the output value can be added to the original data value of the output data point corresponding to the output position in the second memory space to obtain a sum, and the sum is used as the data value of the output data point corresponding to the output position in the second memory space.
[0047] In some embodiments, the second memory space is used to store the intermediate data of each output data point in the output data during the process of processing the sparse data.
[0048] S103. Transmit the output data points in the second memory space to the external memory as the processed data.
[0049] Among them, after processing the input data points of each of the multiple groups, it is characterized that all non-zero input data points in the sparse data have updated the data values of the output data points at the corresponding output positions through the output values corresponding to the input data points. At this time, the processing process of the sparse data has been completed, and the output data points in the second memory space can be transferred to the external memory as the processed data.
[0050] In some embodiments, the transferring the output data points in the second memory space to the external memory as the processed data includes: in the case where a quantization process is required, quantizing the output data points in the second memory space and migrating them to a third memory space, and migrating the quantized output data points in the third memory space to the external memory; during the process of migrating the quantized output data points in the third memory space to the external memory, processing the next sparse data based on the second memory space; in the case where a quantization process is not required, while migrating the output data points in the second memory space to the external memory, processing the next sparse data based on the third memory space.
[0051] Among them, in the case where a quantization process is required, during the process of migrating the data in the second memory space to the third memory space, it is necessary to perform quantization processing on it. Compared with the processed data (the data in the third memory space), the data in the second memory space has a wider bit width and a larger size. Therefore, the bit widths and / or sizes of the third memory space and the second memory space are not the same; in the case where a quantization process is not required, the third memory space is equivalent to a second second memory space. While the second memory space is used to execute the processing process of the current sparse data, the third memory space is used to execute the process of transferring the previous sparse data to the external memory; correspondingly, while the second memory space is used to execute the process of transferring the current sparse data to the external memory, the third memory space is used to execute the processing process of the next sparse data.
[0052] In the embodiments of the present disclosure, since in the process of processing sparse data, according to multiple groups of the sparse data, the input data points of each group are sequentially read and processed, it is possible to process sparse data with a large data volume using a smaller memory. Compared with the conventional solution of loading all sparse data into the memory, it can reduce the occupation of memory resources and improve the utilization efficiency of memory resources; at the same time, since the input data points in each group correspond to the same processing strategy, in the process of sequentially processing the input data points corresponding to each group, the same processing strategy can be used to batch-process all input data points of the current group, improving the processing efficiency.
[0053] See Figure 2a , Figure 2ais an optional process schematic diagram of the sparse data processing method provided by the embodiments of the present disclosure. Based on Figure 1 , Figure 1 in S102 can be updated to S201 to S202, which will be described in combination with Figure 2a the steps shown.
[0054] S201. Generate at least one processing task based on the processing strategy corresponding to the grouping; the number of the processing tasks is related to the number of output feature points corresponding to the input data points in the grouping.
[0055] In some embodiments, for the convenience of understanding the present solution, the following processing process will be described by taking the convolution processing as an example. Currently, the embodiments of the present disclosure can also be applied to other processing methods (operators) such as fully connected processing and elewise processing.
[0056] Please refer to Figures 2b to 2e the pattern classification schematic diagram shown. Taking the convolution kernel size of 3×3 and the stride of 2 as an example, based on the position of the input data points in the sparse data (original feature map), the input data points can be grouped.
[0057] Among them, Figure 2b the input data point I1 in has the position of (2, 2) in the current feature map. In the case of the convolution kernel size of 3×3 and the stride of 2, this input data point I1 only affects one output data point O1(1, 1) in the output data (output feature map); Figure 2c the input data point I2 in has the position of (3, 2) in the current feature map. In the case of the convolution kernel size of 3×3 and the stride of 2, this input data point I2 not only affects the output data point O1(1, 1) in the output feature map, but also affects the output data point O2(2, 1) in the output feature map; Figure 2d the input data point I3 in has the position of (2, 3) in the current feature map. In the case of the convolution kernel size of 3×3 and the stride of 2, this input data point I3 not only affects the output data point O1(1, 1) in the output feature map, but also affects the output data point O3(1, 2) in the output feature map; Figure 2e the input data point I4 in has the position of (3, 3) in the current feature map. In the case of the convolution kernel size of 3×3 and the stride of 2, this input data point I4 affects 4 output data points in the output feature map, including the output data point O1(1, 1), the output data point O3(1, 2), the output data point O2(2, 1), and the output data point O4(2, 2) in the output feature map.
[0058] Based on the above Figures 2b to 2eThe schematic diagram of pattern classification shown can be seen that all input data points that are non-zero in the sparse data can be divided into four groups based on the above four types of input feature points. Among them, the first group includes input data points with the same pattern type as the input data point I1; the second group includes input data points with the same pattern type as the input data point I2; the third group includes input data points with the same pattern type as the input data point I3; the fourth group includes input data points with the same pattern type as the input data point I4.
[0059] In the above S201, for different groups, different numbers of processing tasks can be generated based on the processing strategy corresponding to the grouping strategy. Among them, the number of processing tasks corresponding to the group is related to the number of output feature points affected by the input data points in the group. For example, for the first group, since the input data points in the first group only affect one output data point, the number of processing tasks corresponding to the first group is one; for the second group and the third group, since the input data points in the second group and the third group affect two output data points, the number of processing tasks corresponding to the second group and the third group is two; for the fourth group, since the input data points in the fourth group affect four output data points, the number of processing tasks corresponding to the fourth group is four.
[0060] S202. For each of the processing tasks, based on the processing sub-strategy corresponding to the processing task, determine the sub-output value and sub-output position corresponding to the input data point in the processing task; use the sub-output value to update the data value of the output data point corresponding to the sub-output position in the second memory space.
[0061] It should be noted that for one processing task, the processing sub-strategy of the processing task for the input feature points is the same. For the input data points within the group, through different processing tasks, the output positions and corresponding output values of different output data points affected by the output data point can be obtained.
[0062] In some embodiments, the above-mentioned determining the sub-output value and sub-output position corresponding to the input data point in the processing task based on the processing sub-strategy corresponding to the processing task can be implemented through S2021 to S2023.
[0063] S2021. Based on the weight mapping strategy corresponding to the processing task, obtain the target weight corresponding to the processing task from the pre-loaded weight operators; the weight operators include multiple weights for processing the sparse data.
[0064] Among them, taking the convolution processing of the sparse data as an example, the weight operator is the convolution kernel for processing the sparse data.
[0065] S2022. Based on the output mapping strategy corresponding to the processing task, determine the sub-output position based on the first position of the input data point in the sparse data.
[0066] S2023. Determine the sub-output value based on the target weight and the data value of the input data point.
[0067] In some embodiments, the above-mentioned updating the data value of the output data point corresponding to the sub-output position in the second memory space by using the sub-output value can be implemented through S2024.
[0068] S2024. Accumulate the sub-output value to the original data at the sub-output position in the second memory space to update the data value of the output data at the sub-output position in the second memory space.
[0069] Among them, taking the above-mentioned fourth group as an example, since the input data points in this fourth group can affect four output data points, therefore, for this fourth group, four processing tasks can be generated. For the input data point I4, the corresponding first processing task is used to update the data value of the output data point O1(1, 1) in the output feature map, the second processing task is used to update the data value of the output data point O3(1, 2) in the output feature map, the third processing task is used to update the data value of the output data point O2(2, 1) in the output feature map, and the fourth processing task is used to update the data value of the output data point O4(2, 2) in the output feature map.
[0070] It should be noted that for the multiple input data points in this fourth group, during the execution of each processing task, it is necessary to process all the input data points in this fourth group based on this processing task. That is, when the fourth group includes input data points I4, I5, and I6, it is necessary to first process input data points I4, I5, and I6 based on the first processing task; after all the input data points in the fourth group have completed the first processing task, then process input data points I4, I5, and I6 based on the second processing task, and so on.
[0071] In some embodiments, the processing sub-strategy corresponding to each processing task includes a weight mapping strategy and an output mapping strategy. Correspondingly, the processing strategy corresponding to the group includes the processing sub-strategies corresponding to each processing task.
[0072] For each processing task, taking the first processing task as an example, the processing process of the input data point I4 includes:
[0073] Obtain the weight mapping strategy corresponding to the first processing task. The weight mapping strategy includes the relative position of the target weight in the weight operator. The relative position of the target weight required by the input data point I4 in the weight operator (convolution kernel) is (3, 3), that is, the weight at the relative position (3, 3) in the weight operator (convolution kernel) is used as the target weight. Correspondingly, in the process of processing other input data points in the fourth group based on the first processing task, the target weight is shared.
[0074] Meanwhile, the output mapping strategy corresponding to the first processing task can also be obtained. Among them, the output mapping strategy represents the mapping relationship between the first position of the input data point in the sparse data and the corresponding sub-output position. For example, the output mapping strategy corresponding to the first processing task can be: when the first position is (X, Y), the corresponding sub-output position is ((X - 1) / 2, (Y - 1) / 2). For the input data point I4, the first position of the input data point I4 in the sparse data is (3, 3). Based on the above output mapping strategy corresponding to the first processing task, it can be determined that in the process of processing the first processing task, the sub-output position corresponding to the input data point I4 is (1, 1), that is, the input data point affects the output data point O1 at the position (1, 1) in the output data.
[0075] In the embodiments of the present disclosure, since for the processing strategy corresponding to each group, the processing process of the input data points in the group is divided into at least one processing task, and each processing task can use the same processing sub-strategy to process all the input data points in the group. When the processing sub-strategy includes a weight mapping strategy and an output mapping strategy, the full reuse of the target weight can be achieved, avoiding the resource consumption caused by multiple weight readings. At the same time, the effect of batch determining the sub-output positions corresponding to all the input data points in the group by the same output mapping strategy can also be achieved, improving the data processing efficiency.
[0076] See Figure 3 , Figure 3 is an optional flowchart of the sparse data processing method provided by the embodiments of the present disclosure. Based on any of the above embodiments, taking the embodiment corresponding to Figure 1 as an example, before S101 in Figure 1 , the method may further include S301 to S302, which will be described in conjunction with the steps shown in Figure 3 .
[0077] S301. Traverse the sparse data, group multiple input data points in the sparse data to obtain multiple groups; the processing strategies corresponding to the input data points in each group are the same.
[0078] In some embodiments, traversing the sparse data as described above, grouping multiple input data points in the sparse data, and obtaining multiple groups can be achieved through S3011 to S3013.
[0079] S3011. Traverse the sparse data to determine the data information of each input data point; the data information includes the first position of the input data point in the sparse data.
[0080] S3012. For each input data point, based on the first position of the input data point and the weight operator, determine the weight position and output information corresponding to the input data point; the weight position is the relative position of the target weight corresponding to the input feature point in the weight operator, and the output information includes the number of output data points corresponding to the input data point, and the mapping relationship between the first position of the input data point and the second position of each output data point.
[0081] S3013. Based on the weight position and output information of each input data point, group each input data point to obtain the multiple groups; the weight positions and output information corresponding to the input data points within each group are the same.
[0082] S302. Store the input data points corresponding to the multiple groups into the external memory.
[0083] In the embodiments of the present disclosure, since before processing the sparse data, for each input data point in the sparse data, based on the relationship between the input data point and the weight position, the number of output data points corresponding to the input data point, and the mapping relationship between the first position of the input data point and the second position of each output data point, the input data point is grouped, which can make the input data points within the same group all correspond to the same processing strategy, provide a data basis for the subsequent batch processing of all input data points in each group, and indirectly improve the processing efficiency of subsequent processing of the sparse data.
[0084] See Figure 4 , Figure 4 is an optional flowchart of the sparse data processing method provided by the embodiments of the present disclosure. Based on any of the above embodiments, taking the embodiment corresponding to Figure 1 as an example, before S101 in Figure 1 , the method may further include S401, and step S103 may be updated to S402, which will be described in combination with the steps shown in Figure 4 .
[0085] S401. Chunk the original data based on a preset chunking rule to obtain at least one of the sparse data arranged in an orderly manner; there is no redundant data in the at least one sparse data;
[0086] In some embodiments, the corresponding chunking rule can be set according to the data form of the original data. For example, when the original data is in the NHWC form, where N is the number of feature maps in a batch, H is the number of data points in the vertical height direction, W is the number of data points in the horizontal width direction, and C is the number of channels. The original data can be initially chunked based on the N dimension, and then the data obtained after the initial chunking can be further chunked based on the C dimension. At this time, data in the HW form can be obtained; if further chunking is required, based on the data amounts corresponding to H and W, further chunking can be selected based on the H dimension and / or the W dimension to obtain the above at least one sparse data.
[0087] It should be noted that in the embodiments of the present disclosure, there will be no duplicate data points among the at least one sparse data obtained, that is, there is no redundant data in the at least one sparse data.
[0088] To facilitate understanding of this solution, taking the case where the original data is 1×6×6×3 as an example, since there is only one batch of this original data, therefore, the original data can be directly divided into 3 pieces of 1×6×6×1 data according to the channel dimension. At this time, if further chunking is required, each piece of 1×6×6×1 data can be divided into 4 pieces of 1×3×3×1 sparse data.
[0089] In some embodiments, the original data can be stored in the form of a sparse matrix, that is, the original data contains the coordinate values of each non-zero data point in the original data. During the process of chunking the original data based on the above chunking rule, each non-zero data point can be directly divided into the corresponding sparse data according to the division expression value of the data point coordinate value in the chunking rule. For example, based on the above example, when the original data is 1×6×6×3, if the coordinate value of a non-zero data point is (1, 1, 1, 1), this non-zero data point can be divided into the first two-dimensional feature map (sparse data) with a size of (3×3) under the first channel of the original data.
[0090] S402. Based on the boundary information of the sparse data, transfer the first part of the data points in the second memory space to the external memory as processed data; the boundary information is determined by the chunking rule.
[0091] Based on the above processing scheme for each sparse data, during the process of processing each sparse data, the output data points corresponding to the input data points at the boundary of the sparse data have not completed all calculations. Therefore, during the process of transferring the data in the second memory space to the external memory as the processed data, only the data of the output data points that have undergone all processing is transferred to the external memory.
[0092] In some embodiments, the boundary information is determined by the partitioning rule. Taking the example of dividing each 1×6×6×1 data into 4 1×3×3×1 sparse data, if the size of the corresponding convolution operator is 3×3 and the stride is 1, during the process of performing convolution calculation on the original 1×6×6×1, the size of the obtained output data is 1×4×4×1. Among them, the output data point (1, 1, 1, 1) in the output data can be determined only by the first sparse data. That is to say, during the process of processing the first sparse data, the output data point corresponding to (1, 1, 1, 1) in the second memory space has completed the calculation. During the process of transferring the first part of the data points in the second memory space to the external memory as the processed data, this output data point can be used as the first part of the data points; at the same time, during the process of processing the first sparse data, the data points (1, 1, 2, 1), (1, 1, 3, 1), (1, 2, 1, 1), (1, 2, 2, 1), (1, 2, 3, 1), (1, 3, 1, 1), (1, 3, 2, 1), (1, 3, 3, 1) in the second memory space have not completed the calculation yet, so they are retained in the second memory space; during the process of processing the second sparse data, the data points (1, 1, 2, 1), (1, 1, 3, 1) in the second memory space have completed the calculation and will be used as the first part of the data points corresponding to the second sparse data and transferred to the external memory as the processed data.
[0093] In some embodiments, the above process of transferring the first part of the data points in the second memory space to the external memory as the processed data can be implemented through steps S4021 to S4022.
[0094] S4021. Transfer the first part of the data points in the second memory space to the first subspace in the third memory space; at the same time, transfer the historical data points in the second subspace of the third memory space to the external memory; the historical data points in the second subspace are the first part of the data points obtained when processing the previous adjacent sparse data of the sparse data.
[0095] S4022. Migrate the first part of the data in the first subspace to the external memory; and migrate the new data points in the second memory space to the second subspace; the new data points in the second memory space are the first part of the data points obtained when processing the next adjacent sparse data of the sparse data.
[0096] In the embodiments of the present disclosure, on the basis of partitioning the original data, the sparse data corresponding to each partition is further grouped, which further reduces the occupation of memory resources in one processing process; at the same time, in the embodiments of the present disclosure, the first subspace and the second subspace in the third memory space alternately complete the two processes of reading data from the second memory space and transmitting data to the external memory, improving the data transmission efficiency.
[0097] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.
[0098] The input data (feature map) of the vehicle-mounted radar generally has a large sparsity. If calculations are directly performed based on the original input data, the system will consume a lot of useless work on the huge feature map. In the prior art, there is no implementation method for a system without a sparse operator, or it depends on a unique hardware accelerator and is not universal. The inventors have found through research that if the characteristics of sparse data are used for calculation, the CPU load can be relatively easily reduced. However, for sparse data, such as vehicle-mounted radar, due to its characteristics such as an ultra-large feature map, the system bandwidth often becomes a bottleneck in calculation. Therefore, the optimization of the system bandwidth is the optimization direction of the embodiments of the present disclosure. To solve the above problems, the embodiments of the present disclosure provide a data processing method applicable to sparse data.
[0099] In the embodiments of the present disclosure, the memory of the computing device can be simply classified into device memory (hereinafter referred to as memory for short) and device external memory (hereinafter referred to as external memory for short). The latter is a memory shared by the entire system and belongs to a peripheral device relative to the computing device. Generally speaking, the computing device has a dedicated route to access the memory. To simplify the problem, it is considered that memory access does not consume system bandwidth or is not a bandwidth bottleneck. Therefore, the optimization goal is simplified to minimize external memory access. Exemplarily, taking the convolution operator as an example, in the process of sparse data processing, external memory access includes: reading the feature map, reading the weights, and outputting the feature Figure 3 three parts. Among them, there is no special optimization method for reading weights and outputting the feature map for sparse data. The inventors have found that the optimization goal can be further clarified as the part of reading the feature map while ensuring the highest reuse rate of the weight data.
[0100] For a sparse feature map, only the data and coordinates of its non-zero elements need to be read (e.g., the most common COO format), and its calculation update for the output can be completed. To standardize the operator, different calculation methods of the input for the output need to be decomposed into standard calculation tasks (tasks). The present disclosure innovatively introduces a pattern classification method. Through two dimensions of pattern grouping and the serial numbers of the calculation tasks decomposed therefrom, the one-to-many mapping relationships of the two groups of output data and weights are decomposed into multiple one-to-one mapping relationships. At the same time, for most operators, complete reuse of weights can be achieved through task decomposition.
[0101] For ease of understanding, please refer to Figures 2b to 2e the schematic diagram of pattern classification shown, taking the convolution kernel size of 3×3 and the stride of 2 as an example. Based on the positions of the input data points in the sparse data, the input data points can be grouped.
[0102] Among them, Figure 2b the input data point I1 in has the position of (2, 2) in the current feature map. In the case of a convolution kernel size of 3×3 and a stride of 2, this input data point I1 only affects one output data point O1(1, 1) in the output data (output feature map); Figure 2c the input data point I2 in has the position of (3, 2) in the current feature map. In the case of a convolution kernel size of 3×3 and a stride of 2, this input data point I2 not only affects the output data point O1(1, 1) in the output feature map, but also affects the output data point O2(2, 1) in the output feature map; Figure 2d the input data point I3 in has the position of (2, 3) in the current feature map. In the case of a convolution kernel size of 3×3 and a stride of 2, this input data point I3 not only affects the output data point O1(1, 1) in the output feature map, but also affects the output data point O3(1, 2) in the output feature map; Figure 2e the input data point I4 in has the position of (3, 3) in the current feature map. In the case of a convolution kernel size of 3×3 and a stride of 2, this input data point I4 affects 4 output data points in the output feature map, including the output data point O1(1, 1), the output data point O3(1, 2); the output data point O2(2, 1), and the output data point O4(2, 2) in the output feature map.
[0103] In a further embodiment, since the memory of a general computing device is limited, the embodiments of the present disclosure also introduce a tiling method to solve the memory bottleneck and is more conducive to hiding the system bandwidth overhead.
[0104] For the offline processing process, it may include a storage management processing process, a pattern offline processing process, and a data arrangement process.
[0105] In the embodiments of the present disclosure, in the storage management processing, the system memory can be pre-divided into an input data memory (corresponding to the first memory space in the above embodiments), an output data memory (corresponding to the third memory space in the above embodiments), a weight data memory, and an accumulator memory (corresponding to the second memory space in the above embodiments).
[0106] Among them, for operators whose output needs to be quantized, the element size of the accumulator memory is set according to the accumulator. That is to say, for operators that need to be quantized, the accumulator memory needs to be set with a wider bit width and a larger size to ensure that there is no overflow during the accumulation process; correspondingly, since the quantization operation has been performed, the output data memory corresponding to this accumulator memory can be set with a smaller space to save memory resources. For example, for an 8-bit quantization scheme, the data output memory can be set to 8 bits, and the corresponding accumulator can be set to 32 bits or 40 bits.
[0107] For non-quantized operators, the accumulator memory and the output data memory have the same size, and the increase of this memory is only for data synchronization requirements, similar to a ping-pong buffer.
[0108] In the embodiments of the present disclosure, in the pattern offline processing, according to the influence of the input feature map on the weights and output positions, it can be divided into several pattern groups, such as the aforementioned convolution example, and a weight mapping table (corresponding to the weight mapping strategy in the above embodiments) and an output mapping table (corresponding to the output mapping strategy in the above embodiments) are obtained.
[0109] According to the positions of the data points in the input feature map and the convolution kernel information (convolution kernel size and stride), the feature map is divided into at least one pattern group, and a weight mapping table and an output mapping table corresponding to each pattern group are established. Among them, the weight mapping table is used to obtain the weights corresponding to the data points in the pattern group, and the output mapping table is used to determine the output positions of the data points in the pattern group.
[0110] In the embodiments of the present disclosure, in the data arrangement process, data arrangement preprocessing of constants such as weights can be performed according to the operator characteristics. For example, for a convolution operator, generally speaking, at this time, the input feature map and the output feature map are better processed according to NHWC, where N: the number of feature maps in a batch. H: the number of data points in the vertical height direction. W: the number of data points in the horizontal width direction. C: the number of channels. Correspondingly, its weights can be arranged according to HWcoci, where co is the output channel and ci is the input channel.
[0111] For the online processing process, it may include a data preprocessing process and a data processing process.
[0112] In some embodiments, the data preprocessing process includes: (1) traversing the feature map and converting it into a sparse matrix data format, such as the COO format. (2) During the process of traversing the feature map, grouping the feature map according to the dimensions of pattern and tiling, and storing it in external memory. It should be noted that the steps of data preprocessing are implemented during the generation of the feature map without consuming additional computing resources;
[0113] In some embodiments, the data processing process includes: (1) for the data points within a block, sequentially performing the processing process of the data points within each pattern group; (2) for the data points within each pattern group, reading the data points within the pattern group from external memory to the input data memory, and splitting the processing task corresponding to the pattern group into at least one processing subtask; (3) for each processing subtask, based on the processing subtask, reading the corresponding weight from the weight mapping table to the weight memory; (4) for each data point in the processing subtask, using the data value of the data point, the weight in the weight memory, and the data processing operator to determine the output value of the data point in the processing subtask, based on the output mapping table, determining the position of the output data point corresponding to the output value, and based on the position of the output data point, updating the value of the output data point in the accumulator memory; (5) after completing the processing process of the data points within all pattern groups, according to the current quantization state, adopting the data migration strategy corresponding to the quantization state to migrate the data in the accumulator memory to external memory.
[0114] Among them, the quantization state includes two states: quantization valid and quantization invalid. The above-mentioned migrating the data in the accumulator memory to external memory according to the current quantization state by adopting the data migration strategy corresponding to the quantization state includes:
[0115] In the case of quantization being valid, performing quantization calculation on the data in the accumulator memory and then migrating it to the output data memory. While the data output memory migrates the quantized data to external memory, it uses the accumulator memory to execute the processing process of the next block.
[0116] In the case of quantization being invalid, since there is no need to perform the quantization process, the bit widths and sizes of the accumulator memory and the output data memory are the same, and the output data memory is equivalent to a second accumulator memory. During the processing process, directly migrate the data in the accumulator memory to external memory, and at the same time use the output data memory (equivalent to the second accumulator memory) to execute the processing process of the next block.
[0117] Based on the sparse data processing method provided in the above embodiments, after the processing of the current level, if the processing of the next level is still required, since sparse data with a large size may still be sparse data after one level of processing. For example, in the case where the processing process is a convolution process, when the bias is 0 or a negative number, the output data is also sparse. Therefore, if the sparse data is still sparse after the first level of processing, during the data output process of the first level of processing, only the non-zero output data points in the output data in the output data memory or the accumulator memory can be saved and stored in the external memory in the form of a sparse matrix. This process is equivalent to the data preprocessing process in the second level of processing, which can not only reduce the transmission bandwidth but also reduce the occupation of the external memory space.
[0118] The technical problems that can be solved by the sparse data processing method provided by the above embodiments include: (1) accelerating the operators of huge sparse feature maps (typically, such as convolution operators); (2) establishing a general framework to extend this solution to most operators (fully connected, elewise, etc.).
[0119] Figure 5 The following is a schematic diagram of the composition structure of a sparse data processing device provided by an embodiment of the present disclosure. As Figure 5 shown, the sparse data processing device 500 includes:
[0120] A reading module 501, configured to sequentially read the input data points of each group corresponding to the sparse data from the external memory to the first memory space;
[0121] A processing module 502, configured to, for the input data points of each group, determine the output value and output position corresponding to the input data points based on the processing strategy corresponding to the group, and update the data value of the output data point corresponding to the output position in the second memory space by using the output value;
[0122] A transmission module 503, configured to transmit the output data points in the second memory space to the external memory as the processed data.
[0123] In some embodiments, the processing strategy includes a weight mapping strategy and an output mapping strategy; the processing module 502 is further configured to: generate at least one processing task based on the processing strategy corresponding to the group; the number of processing tasks is related to the number of output feature points corresponding to the input data points in the group; for each processing task, determine the sub-output value and sub-output position corresponding to the input data points in the processing task based on the processing sub-strategy corresponding to the processing task; and update the data value of the output data point corresponding to the sub-output position in the second memory space by using the sub-output value.
[0124] In some embodiments, the processing sub-strategy includes a weight mapping strategy and an output mapping strategy; the processing module 502 is further configured to: obtain a target weight corresponding to the processing task from pre-loaded weight operators based on the weight mapping strategy corresponding to the processing task; the weight operators include multiple weights for processing the sparse data; determine the sub-output position based on the output mapping strategy corresponding to the processing task and the first position of the input data point in the sparse data; and determine the sub-output value based on the target weight and the data value of the input data point.
[0125] In some embodiments, the processing module 502 is further configured to: accumulate the sub-output value onto the original data at the sub-output position in the second memory space to update the data value of the output data at the sub-output position in the second memory space.
[0126] In some embodiments, the sparse data processing apparatus 500 further includes a grouping module, and the grouping module is configured to: traverse the sparse data, group multiple input data points in the sparse data to obtain multiple groups; the processing strategies corresponding to the input data points in each group are the same; and store the input data points corresponding to the multiple groups into the external memory.
[0127] In some embodiments, the grouping module is further configured to: traverse the sparse data to determine the data information of each input data point; the data information includes the first position of the input data point in the sparse data; for each input data point, determine the weight position and output information corresponding to the input data point based on the first position of the input data point and the weight operator; the weight position is the relative position of the target weight corresponding to the input feature point in the weight operator, and the output information includes the number of output data points corresponding to the input data point and the mapping relationship between the first position of the input data point and the second position of each output data point; group each input data point based on the weight position and output information of each input data point to obtain the multiple groups; the weight positions and output information corresponding to the input data points within each group are the same.
[0128] In some embodiments, the transmission module 503 is further configured to, when a quantization process is required, quantize the output data points in the second memory space and migrate them to the third memory space, and then migrate the quantized output data points in the third memory space to the external memory; during the process of migrating the quantized output data points in the third memory space to the external memory, process the next sparse data based on the second memory space; when a quantization process is not required, while migrating the output data points in the second memory space to the external memory, process the next sparse data based on the third memory space.
[0129] In some embodiments, the grouping module is further configured to: divide the original data into blocks based on a preset block division rule to obtain at least one of the sparse data arranged in an orderly manner; there is no redundant data in the at least one sparse data; for each sparse data, the transmission module 503 is further configured to: based on the boundary information of the sparse data, transfer the first part of the data points in the second memory space to the external memory as processed data; the boundary information is determined by the block division rule; wherein, the output data of the second memory space further includes a second part of data points that have not been transferred to the external memory; the data values of the second part of data points are updated during the process of processing the next adjacent sparse data of the sparse data.
[0130] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.
[0131] It should be noted that in the embodiments of the present disclosure, if the above sparse data processing method is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a device to execute all or part of the methods of the various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present disclosure are not limited to any combination of hardware and software for a specific purpose.
[0132] Figure 6 FIG. is a schematic diagram of the hardware entity of a sparse data processing device provided for the embodiments of the present disclosure, as Figure 6As shown, the hardware entities of the sparse data processing device 600 include: a processor 601 and a memory 602. Among them, the memory 602 stores a computer program that can run on the processor 601. When the processor 601 executes the program, it implements the steps in the method of any of the above embodiments. In some embodiments, the device 600 for collecting and paying out game coins on the gaming table can be the detection device described in any of the above embodiments.
[0133] The memory 602 stores a computer program that can run on the processor. The memory 602 is configured to store instructions and applications executable by the processor 601, and can also cache data to be processed or already processed by the processor 601 and each module in the sparse data processing device 600 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0134] When the processor 601 executes the program, it implements the steps of the sparse data processing method of any of the above. The processor 601 generally controls the overall operation of the sparse data processing device 600.
[0135] The embodiments of the present disclosure provide a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the sparse data processing method of any of the above embodiments.
[0136] It should be pointed out here that: the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects to the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present disclosure, please refer to the descriptions of the method embodiments of the present disclosure for understanding.
[0137] The above-mentioned processor may be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that the electronic device implementing the functions of the above-mentioned processor may also be other types, and the embodiments of the present disclosure do not make specific limitations.
[0138] The above-mentioned computer storage medium / memory may be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface memory, an optical disc, or a Compact Disc Read-Only Memory (CD-ROM), etc.; it may also be various terminals including one or any combination of the above-mentioned memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.
[0139] It should be understood that the "one embodiment" or "an embodiment" or "an embodiment of the present disclosure" or "the foregoing embodiment" or "some embodiments" mentioned throughout the specification means that the target features, structures or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, the "in one embodiment" or "in an embodiment" or "an embodiment of the present disclosure" or "the foregoing embodiment" or "some embodiments" that appear throughout the specification do not necessarily refer to the same embodiment. In addition, these target features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present disclosure, the magnitude of the serial numbers of the above processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0140] Without special instructions, when the detection device executes any step in the embodiments of the present disclosure, it can be the processor of the detection device that executes this step. Unless otherwise specified, the embodiments of the present disclosure do not limit the order of execution of the following steps by the detection device. In addition, the methods used to process data in different embodiments can be the same method or different methods. It should also be noted that any step in the embodiments of the present disclosure can be independently executed by the detection device, that is, when the detection device executes any step in the above embodiments, it can be independent of the execution of other steps.
[0141] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0142] The units described as separate components above may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0143] In addition, in each embodiment of the present disclosure, each functional unit may be entirely integrated into one processing unit, or each unit may be separately regarded as one unit, or two or more units may be integrated into one unit; the above-mentioned integrated unit may be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0144] The methods disclosed in several method embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new method embodiments.
[0145] The features disclosed in several product embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new product embodiments.
[0146] The features disclosed in several method or device embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0147] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the aforementioned storage medium includes: removable storage devices, read-only memories (ROMs), magnetic disks, or optical discs, etc., which can store program codes.
[0148] Alternatively, if the above-mentioned integrated unit of the present disclosure is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure, in essence, or the part that contributes to the related art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a detection device, or a network device, etc.) to execute all or part of the methods described in various embodiments of the present disclosure. And the aforementioned storage medium includes: removable storage devices, ROMs, magnetic disks, or optical discs, etc., which can store program codes.
[0149] In the embodiments of the present disclosure, the descriptions of the same steps and the same content in different embodiments can be referred to each other. In the embodiments of the present disclosure, the term "and" does not affect the order of steps.
[0150] As described above, it is only the implementation mode of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claimed rights.
Claims
1. A sparse data processing method, characterized in that, the method includes: sequentially reading input data points of each group corresponding to sparse data from an external memory into a first memory space; for the input data points of each group, based on the processing strategy corresponding to the group, determining an output value and an output position corresponding to the input data point, and using the output value to update a data value of an output data point corresponding to the output position in a second memory space; transferring the output data points in the second memory space as processed data to the external memory; the determining an output value and an output position corresponding to the input data point, and using the output value to update a data value of an output data point corresponding to the output position in a second memory space based on the processing strategy corresponding to the group includes: generating at least one processing task based on the processing strategy corresponding to the group; the number of the processing tasks is related to the number of output feature points corresponding to the input data points in the group; for each processing task, determining a sub-output value and a sub-output position corresponding to the input data point in the processing task based on a processing sub-strategy corresponding to the processing task; using the sub-output value to update a data value of an output data point corresponding to the sub-output position in the second memory space.
2. The method according to claim 1, characterized in that, the processing sub-strategy includes a weight mapping strategy and an output mapping strategy; the determining a sub-output value and a sub-output position corresponding to the input data point in the processing task based on the processing sub-strategy corresponding to the processing task includes: obtaining a target weight corresponding to the processing task from pre-loaded weight operators based on the weight mapping strategy corresponding to the processing task; the weight operators include multiple weights for processing the sparse data; determining the sub-output position based on a first position of the input data point in the sparse data based on the output mapping strategy corresponding to the processing task; determining the sub-output value based on the target weight and a data value of the input data point.
3. The method according to claim 1 or 2, characterized in that, the using the sub-output value to update a data value of an output data point corresponding to the sub-output position in the second memory space includes: accumulating the sub-output value onto original data at the sub-output position in the second memory space to update a data value of the output data at the sub-output position in the second memory space.
4. The method according to claim 1 or 2, characterized in that, the method further includes: traversing the sparse data, grouping multiple input data points in the sparse data to obtain multiple groups; the processing strategies corresponding to the input data points in each group are the same; storing the input data points corresponding to the multiple groups to the external memory.
5. The method according to claim 4, characterized in that, the traversing the sparse data, grouping multiple input data points in the sparse data to obtain multiple groups includes: Traverse the sparse data to determine the data information of each input data point; the data information includes the first position of the input data point in the sparse data. For each input data point, based on the first position of the input data point and the weight operator, determine the weight position and output information corresponding to the input data point; the weight position is the relative position of the target weight corresponding to the input feature point in the weight operator, and the output information includes the number of output data points corresponding to the input data point, and the mapping relationship between the first position of the input data point and the second position of each output data point. Based on the weight position and output information of each input data point, group each input data point to obtain the multiple groups; the weight position and output information corresponding to the input data points within each group are the same.
6. The method according to claim 1, wherein, the step of transmitting the output data points in the second memory space to the external memory as the processed data includes: in the case where a quantization process is required, quantize the output data points in the second memory space and migrate them to the third memory space, and migrate the quantized output data points in the third memory space to the external memory; during the process of migrating the quantized output data points in the third memory space to the external memory, perform the processing process of the next sparse data based on the second memory space; in the case where a quantization process is not required, while migrating the output data points in the second memory space to the external memory, perform the processing process of the next sparse data based on the third memory space.
7. The method according to claim 1, wherein, the method further includes: partition the original data based on a preset partitioning rule to obtain at least one sparse data arranged in order; there is no redundant data in the at least one sparse data; for each sparse data, the step of transmitting the output data points in the second memory space to the external memory as the processed data includes: based on the boundary information of the sparse data, transmit the first part of the data points in the second memory space to the external memory as the processed data; the boundary information is determined by the partitioning rule; wherein, the output data in the second memory space further includes a second part of data points that have not been transmitted to the external memory; the data values of the second part of data points are updated during the process of processing the next adjacent sparse data of the sparse data.
8. A sparse data processing device, wherein, it includes: a reading module, configured to sequentially read each input data point in each group corresponding to the sparse data from the external memory to the first memory space; a processing module, configured to, for each input data point in each group, determine the output value and output position corresponding to the input data point based on the processing strategy corresponding to the group, and update the data value of the output data point corresponding to the output position in the second memory space by using the output value. A transmission module, configured to transmit the output data points in the second memory space as processed data to the external memory; A processing module, further configured to generate at least one processing task based on the processing strategy corresponding to the group; the number of the processing tasks is related to the number of output feature points corresponding to the input data points in the group; for each of the processing tasks, determine a sub-output value and a sub-output position corresponding to the input data point in the processing task based on the processing sub-strategy corresponding to the processing task; and update the data value of the output data point corresponding to the sub-output position in the second memory space with the sub-output value.
9. A sparse data processing device, characterized in that it comprises: a memory and a processor, the memory stores a computer program that can run on the processor, and when the processor executes the computer program, the steps in the method according to any one of claims 1 to 7 are implemented.
10. A computer storage medium, characterized in that the computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image processing method and device based on convolutional neural network
CN110555847A