Downsampling pooling method, device, equipment and program product
By comparing or accumulating pooled data one by one in cache and RAM space, the problem of limited cache space is solved, and the pooling operation of large pooled cores is realized, and the dimensionality reduction efficiency and feature extraction capabilities are improved.
Patent Information
- Application Number
- CN202510400752.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
AI Technical Summary
In the case of limited cache space and RAM space, it is difficult for the prior art to pool the large pooling core, which affects the dimensionality reduction efficiency and feature extraction.
By comparing or accumulating the data to be pooled one by one by one, the cache is updated by comparing or accumulating the data to be pooled one by one, the local results are updated to the RAM space one by one, the final results are determined, and the pooling operation of large-size pooled cores is realized.
Without increasing additional resource consumption, the dimensionality reduction efficiency of pooling operations is improved and a larger range of features can be extracted.
Smart Images

Figure CN120297352A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a pooling method, device, equipment, and program product for downsampling. Background Art
[0002] The pooling layer (fully known as Pooling Layer in English) is an important component in a convolutional neural network (abbreviated as CNN in English, fully known as Convolutional Neural Networks), mainly used to perform downsampling on the input feature map, that is, to compress the spatial dimensions (height and width) of the feature map while retaining important features.
[0003] The pooling operation aggregates the local areas of the input feature map through a sliding window. Each sliding window operation needs to cache all the data within the window for aggregation calculation. If the input feature map contains multiple channels (such as the 3 channels of an RGB image), the pooling operation for each channel needs to be performed independently, and the cache requirement will double according to the number of channels. Since the resource consumption of the pooling operation grows exponentially with the kernel size, a relatively small pooling kernel (such as 2×2, 3×3) is usually selected in practical applications. In the case of limited cache space and RAM space, it is not conducive to the pooling operation with a large pooling kernel, not conducive to extracting features in a larger range, and affects the dimensionality reduction efficiency. Summary of the Invention
[0004] In view of this, embodiments of this application provide a pooling method, device, equipment, and program product for downsampling to solve the problem in the prior art that in the case of limited cache space and RAM space, it is not conducive to the pooling operation with a large pooling kernel, not conducive to extracting features in a larger range, and affects the dimensionality reduction efficiency.
[0005] The first aspect of the embodiments of this application provides a pooling method for downsampling, and the method includes:
[0006] Receiving data to be pooled;
[0007] Updating the cache by comparing or accumulating the data to be pooled row by row to generate a local result after row pooling;
[0008] Updating the local result to the RAM space by comparing or accumulating to determine the final result of the data to be pooled.
[0009] In combination with the first aspect, in the first possible implementation manner of the first aspect, before updating the cache by comparing or accumulating the data to be pooled row by row to generate a local result after row pooling, the method further includes:
[0010] Obtain the pooled downstream demand information;
[0011] Determine the pooling mode according to the downstream demand information.
[0012] Combined with the first possible implementation manner of the first aspect, in the second possible implementation manner of the first aspect, the pooling mode is a two-dimensional pooling mode;
[0013] Update the cache in the way of comparing or accumulating each row of the data to be pooled one by one, and generate the local result after row pooling, including:
[0014] Write the first data of each row in the data to be pooled into the cache;
[0015] Compare or accumulate the data after the first data of each row in the data to be pooled with the data in the cache respectively, and update the result of comparison or accumulation to the cache. After the result of comparison or accumulation of the last data of each row is updated, the local result after row pooling is obtained.
[0016] Combined with the second possible implementation manner of the first aspect, in the third possible implementation manner of the first aspect, the data to be pooled is multi-channel data to be pooled;
[0017] Update the local result to the RAM space by comparing or accumulating one by one to determine the final result of the data to be pooled, including:
[0018] Write the local result of the first row data of each channel into the corresponding RAM space;
[0019] Compare or accumulate the local results of the other row data after the first row data of each channel with the data stored in the RAM space in sequence, and update the result of comparison or accumulation to the RAM space until the local result of the last row data is compared or accumulated with the data stored in the RAM space, and update the result of comparison or accumulation to the RAM space to obtain the pooling result corresponding to each channel.
[0020] Combined with the first possible implementation manner of the first aspect, in the fourth possible implementation manner of the first aspect, the pooling mode is a large-scale dimensionality reduction pooling mode, the data to be pooled is single-channel single-row data, and the method further includes:
[0021] Write the first data of the single-channel single-row data into the cache;
[0022] Sequentially compare or accumulate the data after the first data of the single-channel single-row data with the cached data, and update the comparison or accumulation result to the cache. After the comparison or accumulation result of the last data of the single-channel single-row data is updated, determine that the cached data is the pooling result of the single-channel single-row data.
[0023] Combined with the first possible implementation manner of the first aspect, in the fifth possible implementation manner of the first aspect, the pooling mode is a large-scale dimensionality reduction pooling mode, the data to be pooled is single-channel single-column data, and the method further includes:
[0024] Divide the single-channel single-column data into multiple data segments, and determine the corresponding relationship between the data in the data segments and the RAM space;
[0025] Store the first data segment of the single-column data into the corresponding RAM space respectively according to the corresponding relationship;
[0026] Compare or accumulate the data in the data segments after the first data segment with the data stored in the RAM space respectively according to the corresponding relationship, and update the data stored in the RAM space according to the comparison or accumulation result;
[0027] After completing the comparison or accumulation of the multiple segmented data corresponding to the data to be pooled, compare or accumulate the data stored in the multiple RAM spaces, and determine the pooling result of the single-column data according to the comparison or accumulation result.
[0028] Combined with the first possible implementation manner of the first aspect, in the sixth possible implementation manner of the first aspect, the pooling mode is a large-scale dimensionality reduction pooling mode, the data to be pooled is single-channel single-column data, and the method further includes:
[0029] Divide the single-channel single-column data into multiple data segments, and determine the corresponding relationship between the data segments and the RAM space;
[0030] Store the first data of each data segment into the corresponding RAM space respectively according to the corresponding relationship;
[0031] Compare or accumulate the data after the first data of each data segment with the data stored in the RAM space respectively according to the corresponding relationship, and update the data stored in the RAM space according to the comparison or accumulation result;
[0032] After completing the comparison or accumulation of the multiple segmented data corresponding to the data to be pooled, compare or accumulate the data stored in the multiple RAM spaces, and determine the pooling result of the single-column data according to the comparison or accumulation result.
[0033] Combined with the fifth or sixth possible implementation manner of the first aspect, in the seventh possible implementation manner of the first aspect, dividing the single-channel single-column data into multiple data segments and determining the correspondence between the data in the data segments and the RAM space includes:
[0034] Determine the data segment length according to the number of the RAM spaces, and divide the single-column data into multiple data segments according to the data segment length;
[0035] Determine the correspondence between the data in the data segment and the RAM space according to the sequence number of the data in the data segment and the sequence of the RAM space.
[0036] A second aspect of the embodiments of the present application provides a downsampling pooling device, and the device includes:
[0037] A data receiving unit, configured to receive data to be pooled;
[0038] A local result generating unit, configured to update a cache by comparing or accumulating the data to be pooled row by row, and generate a local result after row pooling;
[0039] A final result generating unit, configured to update the local result to the RAM space by comparing or accumulating one by one, and determine the final result of the data to be pooled.
[0040] Combined with the second aspect, in the first possible implementation manner of the second aspect, the device further includes:
[0041] A downstream demand information obtaining unit, configured to obtain downstream demand information after pooling;
[0042] A pooling mode determining unit, configured to determine a pooling mode according to the downstream demand information.
[0043] Combined with the first possible implementation manner of the second aspect, in the second possible implementation manner of the second aspect, the pooling mode is a two-dimensional pooling mode;
[0044] A local result obtaining subunit, configured to compare or accumulate the data after the first data of each row in the data to be pooled with the cached data respectively, and update the result of the comparison or accumulation to the cache, and obtain the local result after row pooling after the result of the comparison or accumulation of the last data of each row is updated.
[0045] Combined with the second possible implementation manner of the second aspect, in the third possible implementation manner of the second aspect, the data to be pooled is multi-channel data to be pooled;
[0046] The final result generating unit includes:
[0047] The first partial result writing subunit is used to write the partial results of the first-row data of each channel into the corresponding RAM space;
[0048] The calculation subunit for sequential update is used to compare or accumulate the partial results of the other-row data after the first-row data of each channel with the data stored in the RAM space in sequence, update the results of the comparison or accumulation to the RAM space, until the partial result of the last-row data is compared or accumulated with the data stored in the RAM space, and update the results of the comparison or accumulation to the RAM space, so as to obtain the pooling results corresponding to each channel.
[0049] Combined with the first possible implementation manner of the second aspect, in the fourth possible implementation manner of the second aspect, the pooling mode is a large-scale dimensionality reduction pooling mode, the data to be pooled is a single-row data of a single channel, and the device further includes:
[0050] The cache writing unit is used to write the first data of the single-row data of the single channel into the cache;
[0051] The pooling result determination unit is used to compare or accumulate the data after the first data of the single-row data of the single channel with the data in the cache in sequence, update the results of the comparison or accumulation to the cache, and after the results of the comparison or accumulation of the last data of the single-row data of the single channel are updated, determine the data in the cache as the pooling result of the single-row data of the single channel.
[0052] Combined with the first possible implementation manner of the second aspect, in the fifth possible implementation manner of the second aspect, the pooling mode is a large-scale dimensionality reduction pooling mode, the data to be pooled is a single-row data of a single channel, and the device further includes:
[0053] The segmentation unit is used to divide the single-column data of the single channel into multiple data segments, and determine the corresponding relationship between the data in the data segments and the RAM space;
[0054] The segmented storage unit is used to store the first data segment of the single-column data into the corresponding RAM space respectively according to the corresponding relationship;
[0055] The RAM processing unit is used to compare or accumulate the data in the data segments after the first data segment with the data stored in the RAM space respectively according to the corresponding relationship, and update the data stored in the RAM space according to the results of the comparison or accumulation;
[0056] The global comparison or accumulation unit is used to compare or accumulate the data stored in multiple RAM spaces after the comparison or accumulation of the multiple segmented data corresponding to the data to be pooled is completed, and determine the pooling result of the single-column data according to the results of the comparison or accumulation.
[0057] Combined with the fifth possible implementation manner of the second aspect, in the sixth possible implementation manner of the second aspect, the segmentation unit includes:
[0058] A splitting subunit, configured to determine the data segmentation length according to the number of the RAM spaces, and split the single-column data into multiple data segments according to the data segmentation length;
[0059] A correspondence determination subunit, configured to determine the correspondence between the data in the data segment and the RAM space according to the sequence number of the data in the data segment and the sequence of the RAM spaces.
[0060] A third aspect of the embodiments of the present application provides a downsampling pooling device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the downsampling pooling device implements the method according to any one of the first aspect.
[0061] A fourth aspect of the embodiments of the present application provides a computer program product, which when running on a computer, causes the computer to execute the method in the above first aspect or its various implementation manners.
[0062] A fifth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of the first aspect are implemented.
[0063] A sixth aspect of the embodiments of the present application provides a chip for implementing the methods in the various implementation manners in the above first aspect. Specifically, the above chip includes: a processor, configured to call and run a computer program from a memory, so that a device installed with the above chip executes the method according to the above first aspect or its various implementation manners.
[0064] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: After receiving the data to be pooled, the embodiments of the present application first fill the data to be pooled into the cache row by row to generate a local result after row pooling, write the local result corresponding to each row into the RAM space, and then calculate the final result after the pooling of the data to be pooled through the local results of each row stored in the RAM space. Since this method can implement the pooling operation of a larger-size pooling kernel through the associated cooperation of logical devices without additional resource consumption, it is beneficial to extract features in a larger range and improve the dimensionality reduction efficiency. Description of the Drawings
[0065] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for describing the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0066] Figure 1 It is a schematic flowchart of an implementation of a downsampling pooling method provided by an embodiment of the present application;
[0067] Figure 2 It is a schematic diagram of row pooling in a two-dimensional matrix pooling process provided by an embodiment of the present application;
[0068] Figure 3 It is a schematic diagram of column pooling in a two-dimensional matrix pooling process provided by an embodiment of the present application;
[0069] Figure 4 It is a schematic diagram of single-row pooling for large-scale downsampling provided by an embodiment of the present application;
[0070] Figure 5 It is a schematic diagram of single-column pooling for large-scale downsampling provided by an embodiment of the present application;
[0071] Figure 6 It is a schematic diagram of a downsampling pooling device provided by an embodiment of the present application;
[0072] Figure 7 It is a schematic diagram of a downsampling pooling device provided by an embodiment of the present application. Detailed implementation manners
[0073] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0074] To illustrate the technical solutions described in the present application, the following will be described through specific embodiments.
[0075] The pooling layer is an important component in convolutional neural networks (CNNs). It is mainly used to downsample the input feature map, that is, to compress the spatial dimensions (height and width) of the feature map while retaining important features. The pooling operation aggregates local regions of the input feature map through a sliding window. Each sliding window operation needs to cache all the data within the window for aggregation calculation. If the input feature map contains multiple channels (such as the 3 channels of an RGB image), the pooling operation for each channel needs to be performed independently, and the cache requirement will increase by a factor equal to the number of channels.
[0076] Since the resource consumption of the pooling operation grows quadratically with the kernel size, in practical applications, a relatively small pooling kernel (such as 2×2, 3×3) is usually selected. For example, a 2×2 kernel needs to process 4 data points, while a 3×3 kernel needs to process 9 data points, and the resource consumption is 2.25 times that of the 2×2 kernel. In the case of limited cache space and RAM space, a large pooling kernel (such as 5×5) will cause a sharp increase in cache requirements, making it difficult to deploy. At the same time, it is not conducive to extracting features in a larger range and affects the dimensionality reduction efficiency.
[0077] To solve the above problems, the present application proposes a pooling method for downsampling. Figure 1 The schematic diagram of the implementation process of this method is as follows:
[0078] In S101, the data to be pooled is received.
[0079] In a deep learning network, especially in a convolutional neural network (CNN), the pooling operation is often used to reduce the dimensionality and computational complexity of data while retaining important features. The data to be pooled can be the feature map output from the convolutional layer, which exists in the form of a two-dimensional matrix, and each element represents a feature value.
[0080] Among them, the data to be pooled may come from the output of the previous neural network layer, such as the feature map after convolution. For example, assume there is a three-dimensional tensor with the shape [height, width, number of channels], where each channel corresponds to a two-dimensional feature map. The number of channels is not limited to 3 and can also be determined according to the amount of RAM space. To improve the pooling efficiency, multiple two-dimensional feature maps can be pooled simultaneously. For example, when the amount of RAM space is 8, two-dimensional feature maps with 8 channels can be pooled simultaneously.
[0081] The data to be pooled can be in integer or floating-point format, depending on the design of the network and the precision requirements of the data. For example, the data to be pooled can be 8-bit integers or 16-bit floating-point numbers, etc. The format of the data to be pooled is related to the resources occupied. The wider the number of bits, the more resources are occupied by the pooling, and the smaller the choice of the pooling kernel.
[0082] The data to be pooled in the embodiments of the present application is the data determined according to the pooling kernel. For example, when the pooling kernel is a 9*9 two-dimensional matrix, the determined data to be pooled (the data corresponding to a single pooling) is a 9*9 data matrix. When the pooling kernel is 256*1, the corresponding data to be pooled is a 256*1 data matrix.
[0083] In S102, the cache is updated by comparing or accumulating the data to be pooled row by row to generate a local result after row pooling.
[0084] The cache in the embodiments of the present application refers to a high-speed storage unit in hardware, which is used to temporarily store data for quick access. In the pooling operation, filling the data into the cache row by row is for in-row pooling calculations, such as max pooling or average pooling.
[0085] The cache can quickly store and access the row data being processed, reduce the access frequency to the slower RAM space, and improve the processing speed.
[0086] For the data to be pooled, the embodiments of the present application can allocate a cache unit for each row of data, and the number of rows of the data to be pooled is the same as the number of cache units. For example, in the case of the pooling mode being the conventional pooling mode, that is, the two-dimensional pooling mode, if the data to be pooled is a 9*9 two-dimensional matrix, 9 cache units can be allocated, and the pooling of 9 rows of data is completed through 9 cache units.
[0087] When performing row pooling, the first data of each row can be written into the corresponding cache unit. Then, the second data of each row is compared or accumulated with the data in the cache. After determining the result of the comparison or accumulation, it is updated to the cache unit. Then, the third data of each row is compared or accumulated with the data in the cache. After determining the result of the comparison or accumulation, it is updated to the cache unit... and so on in a loop until the last data of each row is compared or accumulated with the data in the cache. After determining the result of the comparison or accumulation, it is updated to the cache unit. The data in the cache is the local result after row pooling, and this result can be written into the RAM space.
[0088] For example Figure 2In the schematic diagram of row pooling in the two-dimensional matrix pooling process shown, the data to be pooled is one row of data in a 9*9 two-dimensional array, which is [4, 2, 5, 1, 3, 6, 2, 4, 9]. During row pooling, the first data is first stored in the cache unit corresponding to that row. In the case of max pooling, the second data is compared with the data stored in the cache unit, and the comparison result is 4, so the data in the cache unit does not need to be updated. The third data is compared with the data stored in the cache unit, and the comparison result is 5, and 5 is updated to the cache unit... and so on until the 9th data is compared. After comparison, the comparison result is 9, and the cache unit is updated to 9. After completing the pooling of the row data, the data in the cache unit can be written into the RAM space.
[0089] Among them, the depth of the RAM space can be related to the number of rows of the pooling kernel of the two-dimensional array. Each RAM space can store the local results of row pooling of multiple rows of data in the same channel. The deeper the depth of the RAM space, the more local results can be stored, that is, the more rows the pooling kernel can have.
[0090] In the embodiments of the present application, the number of channels used for pooling calculation simultaneously can be determined according to the number of RAM spaces. The number of channels is also the number of two-dimensional arrays for parallel pooling calculation. For example, for an image, it is the number of pictures for parallel pooling calculation, or when each picture has 3 channels, it is the total number of channels of the pictures for parallel pooling calculation.
[0091] In S103, the local result is updated to the RAM space by means of one-by-one comparison or accumulation to determine the final result of the data to be pooled.
[0092] In the normal mode, after obtaining the local result of row pooling, according to the row order, the first local result of row pooling can be written into the RAM space first, and then the subsequent data is compared or accumulated with the data in the RAM space one by one, and the result of comparison or accumulation is updated to the RAM space until the local result of the last row of data is compared or accumulated with the data stored in the RAM space, and the result of comparison or accumulation is updated to the RAM space to obtain the pooling result corresponding to each channel.
[0093] For the two-dimensional data of channel 1, after row pooling operation, local results of N rows and one column (for a 9*9 pooling kernel, it is 9 rows and one column of data) are obtained. The first local result in the local results of this channel can be written into the RAM space, and then the second local result in the cache unit of row pooling of this channel is compared or accumulated with the data stored in this RAM space to obtain the result of comparison or accumulation, and it is updated to the RAM space until the comparison or accumulation update of the local result of the last row in the RAM space is completed.
[0094] Such as Figure 3In the schematic diagram of column pooling in the two-dimensional matrix pooling process shown, for a single channel, the data obtained after row pooling is [3, 4, 7, 2, 7, 2, 1, 6, 3] T , write the first local result into the RAM space, compare or accumulate the local results after the first local result with the data stored in the RAM space, and update the result of the comparison or accumulation to the RAM space until the comparison or accumulation update process of all local results of row pooling is completed to obtain the pooling result of a single channel.
[0095] Process the local results of other channels in parallel in the same way to obtain the pooling results of multiple channels.
[0096] If the pooling requirement is average pooling, the result of average pooling can be obtained by summing the product of the accumulated value and the weight coefficient.
[0097] In a deep learning network, the downstream of the pooling operation usually includes subsequent network layers or processing modules, and these modules have specific requirements for the pooled data. It can be obtained through network configuration files, user input, or communication with downstream modules. For example, the input requirements of subsequent layers are specified in the network configuration file; users specify requirements during training or deployment of the model; or information is exchanged with downstream modules through network protocols.
[0098] The types of downstream demand information can include pooling modes (such as max pooling, average pooling), pooling window size, stride, etc. For example, the downstream module may require max pooling, the window size is 9×9, the stride is 1, etc., or the window is a large regular reduction pooling mode (in the embodiments of the present application, it refers to the pooling kernel of a single row or a single column of a single channel, and the length of the single row or single column is greater than a predetermined value, such as greater than 9), such as 256*1 or 1*256, etc.
[0099] When the downstream demand information is a large-scale dimensionality reduction pooling mode, the present application can perform pooling processing through caching or the RAM space according to the type of pooling kernel.
[0100] When the data to be pooled is the row data of a single channel, a pooling kernel of (1*N, N is greater than a predetermined value), such as Figure 4 shown, when the single-row data is 1*256 (respectively D1, D2, D3, D4... D253, D254, D255, D256), the first one in the single-row data can be written into the cache, and the subsequent data is compared or accumulated with the data in the cache in turn, and the result of the comparison or accumulation is updated to the cache until the result update of the comparison or accumulation of the last data is completed to determine the pooling result of the single-row data of a single channel.
[0101] In a possible implementation, single-line data can be segmented into multiple data segments, and the cache units corresponding to each data segment are determined respectively. The first data in the data segment is written into the corresponding cache unit, and then the subsequent data in each data segment are compared or accumulated with the data in the corresponding cache unit in sequence, and the result of the comparison or accumulation is updated to the corresponding cache unit to obtain the segmented pooling results of multiple data segments. The segmented pooling results can be updated to the RAM space by comparing or accumulating one by one to determine the pooling result of the single-line data. For example, the maximum pooling value or minimum pooling value can be obtained from the data finally stored in the RAM space, or the sum of accumulation can be obtained according to the finally stored data, and then the mean pooling result can be obtained by multiplying the sum of accumulation by the weight coefficient.
[0102] When the data to be pooled is column data of a single channel, for example Figure 5 as shown, when the single-column data is 256*1 (D1, D2... D255, D256 respectively), the single-column data can be first divided into multiple data segments, and the corresponding relationship between each data segment and the RAM space is determined. For example Figure 5 in the first data segment D1-D32 corresponds to the first RAM space, and the 8th data segment D225-D256 corresponds to the 8th RAM space. Then based on this corresponding relationship, the first data in each data segment is written into the corresponding RAM space. For example, D1 in the first data segment is stored in the 1st RAM space, and D225 in the 8th data segment is stored in the 8th RAM space. The data after the first data in each data segment is compared with the data stored in the corresponding RAM space respectively, and the data stored in the RAM space is updated according to the result of the comparison or accumulation; after the comparison or accumulation of the multiple segmented data corresponding to the data to be pooled is completed, the data stored in the multiple RAM spaces is compared or accumulated, and the pooling result of the single-column data is determined according to the result of the comparison or accumulation.
[0103] Among them, the number of data segments can match the number of RAM spaces, and the length of the data segment is determined according to the data in the RAM space. When the single-column data is 256*1 and the number of RAM spaces is 8, the length of the data segment can be 256 / 8 = 32, that is, each data segment includes 32 rows of single-column data. When establishing the corresponding relationship, the corresponding relationship between the data in the data segment and the RAM space can be determined according to the sequence number of the data in the data segment and the sequence of the RAM space. The sequence number of the RAM space can be the same as the sequence number of the data segment, or can be set arbitrarily.
[0104] In a possible implementation, the corresponding relationship between each piece of data in each data segment and the RAM space can be determined. Based on this corresponding relationship, first, each piece of data in the first data segment is stored in the corresponding RAM space. For the data in the data segments after the first data segment, according to the above corresponding relationship, each piece of data is respectively compared or accumulated with the data previously stored in the corresponding RAM space, and the data stored in the RAM space is updated according to the comparison or accumulation result. For example, the data in the second data segment is compared or accumulated with the data in the first data segment in the corresponding RAM space to determine the data stored in the RAM space. Then, the data in the third data segment is respectively compared or accumulated with the data stored in the corresponding RAM space, and the pooling result of the single-column data is determined according to the comparison or accumulation result.
[0105] Among them, the number of data segments can match the number of RAM spaces, and the length of the data segment is determined according to the data in the RAM space. When the single-column data is 256*1 and the number of RAM spaces is 8, the length of the data segment can be 8, that is, each data segment includes 8 rows of single-column data. When establishing the corresponding relationship, the corresponding relationship between each piece of data in the data segment and the RAM space can be determined according to the serial number of the data in the data segment. For example, the data segment includes 8 pieces of data, and their serial numbers correspond to the sequences of 8 RAM spaces respectively.
[0106] Through the associated pooling operation of the cache and the RAM space in the embodiments of the present application, without additional resource consumption, a very large-scale downsampling (reduce) operation of size 256 can be implemented on the logic device, which is beneficial to the next-step calculation of the neural network. For example, the softmax activation function can utilize the large reduce operation to improve the processing ability.
[0107] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0108] Figure 6 The following is a schematic diagram of a downsampling pooling device proposed in the embodiments of the present application. The device includes:
[0109] A data receiving unit 601, configured to receive data to be pooled.
[0110] A local result generating unit 602, configured to update the cache in a manner of comparing or accumulating each row of the data to be pooled one by one, and generate a local result after row pooling.
[0111] A final result generation unit 603 is configured to update the local result to the RAM space by means of comparison one by one or accumulation, and determine the final result of the data to be pooled.
[0112] Figure 6 The downsampling pooling device shown and Figure 1 corresponds to the downsampling pooling method.
[0113] Figure 7 FIG. is a schematic diagram of a downsampling pooling device provided by an embodiment of the present application. As Figure 7 shown, the downsampling pooling device 7 of this embodiment includes: a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70, such as a downsampling pooling program. When the processor 70 executes the computer program 72, the steps in the above-mentioned various embodiments of the downsampling pooling method are implemented. Alternatively, when the processor 70 executes the computer program 72, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0114] Exemplarily, the computer program 72 may be divided into one or more modules / units. The one or more modules / units are stored in the memory 71 and executed by the processor 70 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 72 in the downsampling pooling device 7.
[0115] The downsampling pooling device 7 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The downsampling pooling device may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art can understand that Figure 7 is only an example of a downsampling pooling device 7, and does not constitute a limitation on the downsampling pooling device 7. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the downsampling pooling device may further include an input / output device, a network access device, a bus, etc.
[0116] The so-called processor 70 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0117] The memory 71 may be an internal storage unit of the downsampling pooling device 7, such as the hard disk or memory of the downsampling pooling device 7. The memory 71 may also be an external storage device of the downsampling pooling device 7, such as a plug-in hard disk equipped on the downsampling pooling device 7, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 71 may also include both the internal storage unit of the downsampling pooling device 7 and the external storage device. The memory 71 is used to store the computer program and other programs and data required by the downsampling pooling device. The memory 71 may also be used to temporarily store the data that has been output or will be output.
[0118] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0119] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0120] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0121] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0122] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0123] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0124] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by hardware related to computer program instructions. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0125] In addition, the embodiment of this application also provides a computer program product, which when running on a computer, causes the computer to execute the methods in the above various implementation manners.
[0126] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing various embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A downsampling pooling method, characterized in that The method includes: Receiving the data to be pooled; Updating the cache by comparing or accumulating the data to be pooled row by row to generate a partial result after row pooling; Updating the partial result to the RAM space by comparing or accumulating one by one to determine the final result of the data to be pooled.
2. The method according to claim 1, characterized in that, Before updating the cache by comparing or accumulating the data to be pooled row by row to generate a partial result after row pooling, the method further includes: Obtaining the downstream requirement information after pooling; Determining the pooling mode according to the downstream requirement information.
3. The method according to claim 2, wherein The pooling mode is a two-dimensional pooling mode; Updating the cache by comparing or accumulating the data to be pooled row by row to generate a partial result after row pooling, including: Writing the first data of each row in the data to be pooled into the cache; Comparing or accumulating the data after the first data of each row in the data to be pooled with the data in the cache respectively, and updating the result of comparison or accumulation to the cache. After the result of comparison or accumulation of the last data of each row is updated, a partial result after row pooling is obtained.
4. The method according to claim 3, wherein The data to be pooled is multi-channel data to be pooled; Updating the partial result to the RAM space by comparing or accumulating one by one to determine the final result of the data to be pooled, including: Writing the partial result of the first row data of each channel into the corresponding RAM space; Comparing or accumulating the partial results of the other row data after the first row data of each channel with the data stored in the RAM space in sequence, and updating the result of comparison or accumulation to the RAM space until the partial result of the last row data is compared or accumulated with the data stored in the RAM space and the result of comparison or accumulation is updated to the RAM space to obtain the pooling result corresponding to each channel.
5. The method according to claim 2, wherein The pooling mode is a large-scale dimensionality reduction pooling mode, the data to be pooled is single-channel single-row data, and the method further includes: Writing the first data of the single-channel single-row data into the cache; Comparing or accumulating the data after the first data of the single-channel single-row data with the data in the cache in sequence, and updating the result of comparison or accumulation to the cache. After the result of comparison or accumulation of the last data of the single-channel single-row data is updated, it is determined that the data in the cache is the pooling result of the single-channel single-row data.
6. The method according to claim 2, wherein The pooling mode is a large-scale dimensionality reduction pooling mode, the data to be pooled is single-channel single-column data, and the method further includes: Dividing the single-channel single-column data into multiple data segments, and determining the corresponding relationship between the data in the data segments and the RAM space; Storing the first data segment of the single-column data into the corresponding RAM space according to the corresponding relationship; Comparing or accumulating the data in the data segments after the first data segment with the data stored in the RAM space according to the corresponding relationship, and updating the data stored in the RAM space according to the result of comparison or accumulation; After completing the comparison or accumulation of multiple segmented data corresponding to the data to be pooled, compare or accumulate the data stored in multiple RAM spaces, and determine the pooling result of the single-column data according to the result of the comparison or accumulation.
7. The method according to claim 2, wherein The pooling mode is a large-scale dimensionality reduction pooling mode, and the data to be pooled is a single-channel single-column data. The method further includes: Divide the single-channel single-column data into multiple data segments, and determine the correspondence between the data segments and the RAM spaces; Store the first data of each data segment into the corresponding RAM space respectively according to the correspondence; For the data after the first data of each data segment, compare or accumulate them with the data stored in the RAM space respectively according to the correspondence, and update the data stored in the RAM space according to the result of the comparison or accumulation; After completing the comparison or accumulation of multiple segmented data corresponding to the data to be pooled, compare or accumulate the data stored in multiple RAM spaces, and determine the pooling result of the single-column data according to the result of the comparison or accumulation.
8. The method according to claim 6 or 7, characterized in that, Dividing the single-channel single-column data into multiple data segments and determining the correspondence between the data in the data segments and the RAM spaces includes: Determine the data segment length according to the number of RAM spaces, and divide the single-column data into multiple data segments according to the data segment length; Determine the correspondence between the data in the data segments and the RAM spaces according to the sequence number of the data in the data segments and the sequence of the RAM spaces.
9. A downsampling pooling device, characterized in that, The device includes: A data receiving unit, configured to receive data to be pooled; A local result generating unit, configured to update the cache in a manner of comparing or accumulating the data to be pooled row by row, and generate a local result after row pooling; A final result generating unit, configured to update the local result to the RAM space in a manner of comparing or accumulating one by one, and determine the final result of the data to be pooled.
10. A downsampling pooling device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the downsampling pooling device implements the method according to any one of claims 1-8.
11. A computer program product comprising computer program instructions, characterized in that, When the computer program is run, the method according to any one of claims 1-8 is executed.