Maximum pooling method and device suitable for vector processor, equipment and medium

By determining the maximum value of sub-blocks based on the pooling pane and preset step-by-step block data in the vector processor, the problem of inefficient maximum pooling caused by the limit on the number of vector registers is solved, and the computing efficiency is improved.

CN120295966APending Publication Date: 2025-07-11NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510431617.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The maximum pooling efficiency in vector processors is inefficient due to the limitation of vector registers, especially when processing larger specification data, the vector processing unit utilization is low.

Method used

By synchronizing the pending data in the dynamic random memory at double rate to the vector processor, and determining the chunking pane based on the pooling pane and the preset step distance in the vector processor, chunking the data into the second specification data, using the preset step distance and the chunking pane to determine the maximum value of each sub-block, and finally determining the maximum pooling value of the first specification data.

Benefits of technology

Improves the computing efficiency of vector processors when processing larger pooling panes, and avoids the inefficiency problem caused by the limit on the number of vector registers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295966A_ABST
    Figure CN120295966A_ABST
Patent Text Reader

Abstract

The invention discloses a maximum pooling method and device suitable for a vector processor, equipment and a medium, and relates to the field of deep learning, and the method comprises the steps: transmitting to-be-processed data in a double-rate synchronous dynamic random access memory to the vector processor; in the vector processor, determining block panes based on the pooling panes and a preset step pitch; wherein the block panes are smaller than the pooling panes; partitioning the to-be-processed data based on the pooling panes to obtain first specification data, and partitioning the first specification data into second specification data based on the partitioning panes; and determining the maximum value of each sub-block of the second specification data based on the preset step pitch, the block pane and the pooling pane, so as to determine the maximum pooling value of the first specification data based on each sub-block maximum value. Therefore, the problem of low maximum pooling efficiency caused by limitation of the vector register can be avoided, and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and particularly to a maximum pooling method, device, equipment, and medium applicable to a vector processor. Background Art

[0002] In a vector processor, Max Pooling is an important operation. The traditional method uses 64 vector registers to load the data of the pooling pane to find the maximum value, and at most supports Max Pooling with a data specification of 8x8. For Max Pooling of larger specifications such as 9x9 and 13x13, due to the limitation of the number of vector registers, all data cannot be loaded at one time. Currently, the method of taking data row by row to find the maximum value of each row and then obtaining the unique maximum value is adopted. For example, 9x9s1 will generate 9 row maximum values, and 13x13s1 will generate 13, where s1 represents that the preset step (i.e., stride) is 1. As the pooling pane increases, using vector registers to read data in this way will be more and more frequent, and the amount of data read each time is much smaller than the number of vector register resources, resulting in low utilization of the vector processing unit, thus affecting the overall operation efficiency.

[0003] In summary, how to avoid the low efficiency of maximum pooling caused by the limitation of the number of vector registers is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a maximum pooling method applicable to a vector processor, which can avoid the low efficiency of maximum pooling caused by the limitation of the number of vector registers. The specific scheme is as follows:

[0005] In a first aspect, the present application provides a maximum pooling method applicable to a vector processor, including:

[0006] Transmitting the data to be processed in the double data rate synchronous dynamic random access memory to the vector processor;

[0007] In the vector processor, determining a sub-pane based on the pooling pane and a preset step; wherein, the sub-pane is smaller than the pooling pane;

[0008] Dividing the data to be processed based on the pooling pane to obtain data of a first specification, and dividing the data of the first specification into data of a second specification based on the sub-pane;

[0009] Determining the maximum value of each sub-block of the data of the second specification based on the preset step, the sub-pane, and the pooling pane, so as to determine the maximum pooling value of the data of the first specification based on the maximum values of each sub-block.

[0010] Optionally, the transmission of the data to be processed in the double data rate synchronous dynamic random access memory to the vector processor includes:

[0011] Obtaining the data to be processed in the double data rate synchronous dynamic random access memory; wherein, the data volume of the data to be processed is greater than the data volume of the first specification data;

[0012] Using the direct memory access component in the vector processor to transmit the data to be processed to the preset initial data storage space in the vector processor.

[0013] Optionally, after the first specification data is blocked into the second specification data based on the blocking pane, it further includes:

[0014] Using the vector memory access unit to transmit the second specification data in the preset initial data storage space to the vector register.

[0015] Optionally, the determination of the blocking pane based on the pooling pane and the preset step includes:

[0016] When the size corresponding to the pooling pane is odd, determining a blocking cross-pane based on the pooling pane and the preset step; wherein, there is data intersection between the sub-blocks corresponding to the blocking cross-pane.

[0017] Optionally, the blocking of the first specification data into the second specification data based on the blocking pane includes:

[0018] Blocking the first specification data based on the blocking pane and obtaining each sub-block; wherein, the data in each sub-block is the second specification data.

[0019] Optionally, the determination of the maximum value of each sub-block of the second specification data based on the preset step, the blocking pane, and the pooling pane includes:

[0020] Calculating a preset number of sub-block maximum values based on the preset maximum value calculation rule and the preset step, and using the vector memory access unit to save the sub-block maximum values to the preset target data storage space in the vector processor.

[0021] Optionally, the determination of the maximum pooling value of the first specification data based on the maximum value of each sub-block includes:

[0022] Determining the target sub-block interval using the size difference between the pooling pane and the blocking pane;

[0023] In the preset target data storage space, using the target sub-block interval to locate the target position of the maximum value sub-block, so as to determine each target sub-block maximum value among the maximum values of each sub-block based on the target position;

[0024] Compare the maximum values of each of the target sub - blocks, and determine the maximum pooling value of the first specification data based on the maximum value among the obtained maximum values of each of the target sub - blocks.

[0025] In a second aspect, the present application provides a maximum pooling device applicable to a vector processor, including:

[0026] A data transmission module, configured to transmit the data to be processed in a double - data - rate synchronous dynamic random access memory to the vector processor;

[0027] A pane determination module, configured to determine a segmented pane in the vector processor based on a pooling pane and a preset step size; wherein, the segmented pane is smaller than the pooling pane;

[0028] A data segmentation module, configured to segment the data to be processed based on the pooling pane to obtain first - specification data, and segment the first - specification data into second - specification data based on the segmented pane;

[0029] A pooling value determination module, configured to determine the maximum value of each sub - block of the second - specification data based on the preset step size, the segmented pane, and the pooling pane, and determine the maximum pooling value of the first - specification data based on the maximum values of each sub - block.

[0030] In a third aspect, the present application provides an electronic device, including:

[0031] A memory, configured to store a computer program;

[0032] A processor, configured to execute the computer program to implement the foregoing maximum pooling method applicable to a vector processor.

[0033] In a fourth aspect, the present application provides a computer - readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the foregoing maximum pooling method applicable to a vector processor is implemented.

[0034] In this application, the data to be processed in a double data rate synchronous dynamic random access memory is transmitted to a vector processor; in the vector processor, a partitioned pane is determined based on a pooling pane and a preset stride; wherein, the partitioned pane is smaller than the pooling pane; the data to be processed is partitioned based on the pooling pane to obtain first - specification data, and the first - specification data is further partitioned into second - specification data based on the partitioned pane; the maximum value of each sub - block of the second - specification data is determined based on the preset stride, the partitioned pane, and the pooling pane, so as to determine the maximum pooling value of the first - specification data based on the maximum values of each sub - block. As can be seen from the above, the data to be processed is transmitted to the vector processor; in the vector processor, a partitioned pane is determined according to the pooling pane and the preset stride; the data to be processed is partitioned by the pooling pane to obtain first - specification data, and then the first - specification data is divided into second - specification data according to the partitioned pane; the maximum value of each sub - block of the second - specification data is determined according to the preset stride, the partitioned pane, and the pooling pane, and then the maximum pooling value of the first - specification data is obtained. In this way, this application can process the maximum pooling operation of a relatively large pooling pane in the vector processor, improving the operation efficiency of the maximum pooling. Description of the Drawings

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0036] Figure 1 It is a flowchart of a maximum pooling method applicable to a vector processor disclosed in this application;

[0037] Figure 2 It is a schematic diagram of the internal structure of a vector processor disclosed in this application;

[0038] Figure 3 It is a schematic diagram of data partitioning when the kernel (i.e., the pooling pane) is 12x12 disclosed in this application;

[0039] Figure 4 It is a schematic diagram of the determination process of the maximum value of the target sub - block when the kernel is 9x9 and the stride (i.e., the preset stride) is 1 disclosed in this application;

[0040] Figure 5 It is a schematic diagram of the determination process of the maximum value of the target sub - block when the kernel is 9x9 and the stride is 2 disclosed in this application;

[0041] Figure 6Schematic diagram of a max pooling device structure applicable to a vector processor disclosed in this application;

[0042] Figure 7 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners

[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Currently, due to the limitation of the number of vector registers, all data cannot be loaded at one time. Currently, the method of fetching data row by row to find the maximum value of each row and then obtaining the unique maximum value is adopted. For example, 9x9s1 will generate 9 row maximum values, and 13x13s1 will generate 13, where s1 represents a step size of 1. As the pooling window increases, using vector registers to read data based on this method will be more and more frequent, and the amount of data read each time is much smaller than the number of vector register resources, resulting in low utilization of the vector processing unit, thus affecting the overall operation efficiency. For this reason, this application provides a max pooling scheme applicable to a vector processor, which can avoid the problem of low max pooling efficiency caused by the limitation of the number of vector registers.

[0045] See Figure 1 As shown, an embodiment of the present invention discloses a max pooling method applicable to a vector processor, including:

[0046] Step S11: Transmit the data to be processed in the double data rate synchronous dynamic random access memory to the vector processor.

[0047] First of all, it should be noted that in this embodiment, Figure 2 is the internal structure of the vector processor. The vector processor includes an SPU (i.e., Scalar Processing Unit, scalar processing unit) for scalar operations, a VPU (i.e., Vector Processing Unit, vector processing unit) for vector operations, and a DMA (i.e., Direct Memory Access, direct memory access) component for data transmission, etc. Among them, the SPU is composed of an SPE (i.e., Scalar Processing Element, scalar processing component) and an SM (i.e., Scalar Memory, scalar memory). And by Figure 2It can be known that there are a total of L VPEs (i.e., Vector Processing Elements) in the vector processor. Therefore, the VPU is composed of L VPEs and AM (i.e., Array Memory). The L VPEs operate collaboratively in the SIMD (i.e., Single Instruction Multiple Data) manner, supporting the turning on and off of specified VPE components, but not supporting data interaction between multiple VPEs. A single VPE can process 1 8-byte data at a time. Among them, the 8-byte data includes but is not limited to FP64 (i.e., double-precision floating-point format, occupying 64-bit storage space), Int64 (i.e., integer format, occupying 64-bit storage space), etc. Also, a single VPE can process 2 4-byte data at a time. Among them, the 4-byte data includes but is not limited to FP32, Int32, etc. In addition, the 768KB (i.e., Kilobyte) AM space is divided into two equal parts: AM0 (i.e., the preset initial data storage space) and AM1 (i.e., the preset target data storage space). The input initial data (i.e., the data to be processed) and intermediate data (i.e., the maximum value of the target sub-block) are stored in these two spaces respectively.

[0048] In this embodiment, the data to be processed is obtained in a double data rate synchronous dynamic random access memory (i.e., DDR, Double Data Rate Synchronous Dynamic Random Access Memory); among them, the data volume of the data to be processed is greater than the data volume of the first specification data, and the first specification data is the data obtained by dividing the data to be processed based on the pooling pane.

[0049] Furthermore, the direct memory access component in the vector processor is used to transfer the data to be processed to the preset initial data storage space in the vector processor. That is to say, DMA can achieve efficient and fast data transfer between the DDR and the vector processor during the data transfer process. Through the DMA component, the data to be processed is transferred to the preset initial data storage space in the vector processor.

[0050] Step S12: In the vector processor, determine the block pane based on the pooling pane and the preset step; among them, the block pane is smaller than the pooling pane.

[0051] In this embodiment, the pooling pane is the basic unit for performing the pooling operation, which determines the scope and effect of the pooling operation. The preset step is the parameter that controls the movement of the pooling pane, which determines the interval distance of the pooling pane during each movement. By setting these two parameters, the block pane can be determined.

[0052] Step S13: Based on the pooling pane, the data to be processed is chunked to obtain first - specification data, and based on the chunking pane, the first - specification data is chunked into second - specification data.

[0053] It can be understood that in this embodiment, the first - specification data is chunked based on the chunking pane to obtain each sub - chunk; where the data in each sub - chunk is second - specification data. When the size corresponding to the pooling pane is odd, the chunking cross - pane is determined based on the pooling pane and the preset step size. That is, an odd - sized pooling pane will cause special cases during chunking, that is, an odd - sized pooling pane will generate a corresponding chunking cross - pane, resulting in data crossover between the sub - chunks corresponding to the chunking cross - pane. Among them, there is data crossover between the sub - chunks corresponding to the chunking cross - pane.

[0054] As Figure 3 shown, the first - specification data can be sliced into 4 or 9 equal parts. In the figure, kernel (that is, the pooling pane) is 12x12, and block (that is, the chunking pane) being 4x4 and block being 6x6 can both complete the corresponding Max Pooling calculation. The difference between the two is only in the division of the kernel.

[0055] In addition, it should be emphasized that in this embodiment, the value of the preset step size must be set to "1" when determining the chunking pane based on the pooling pane and the preset step size. However, during the calculation process of max - pooling, the value of the preset step size can be adjusted according to the actual situation and current requirements. For example, as Figure 4 shown is a schematic diagram of the calculation process when the preset step size is set to 2 during the max - pooling calculation.

[0056] It should be noted that in this embodiment, the second - specification data in the preset initial data storage space needs to be transferred to the vector register (that is, VRx, where x represents the serial number of the vector register, and ), so as to perform corresponding calculations on the data in the vector register. And the vector memory access unit is an important component for data interaction between the vector processor and the external storage device. It can quickly read the data stored in the preset initial data storage space into the vector register, providing data support for subsequent calculation operations. It should be noted that all operations related to max - pooling calculation in this embodiment need to be performed in the vector register. In addition, it should be noted that when the preset step size is 1, most of the data in 2 calculation regions will overlap, and generally, the maximum value of 2 consecutive sub - chunks is calculated at one time. Correspondingly, when transferring the second - specification data to the vector register, the amount of data transferred at one time is block x (block + 1).

[0057] Step S14: Determine the maximum value of each target sub-block of the second specification data based on a preset step size, a block window pane, and a pooling window pane, so as to determine the maximum pooling value of the first specification data based on the maximum value of each target sub-block.

[0058] In this embodiment, calculate the maximum value of a preset number of sub-blocks based on a preset maximum value calculation rule and the preset step size, and use the vector memory access unit to save the maximum value of the sub-blocks to a preset target data storage space in the vector processor.

[0059] Further, in this embodiment, determine the target sub-block interval by using the size difference between the pooling window pane and the block window pane. In the preset target data storage space, use the target sub-block interval to locate the target position of the maximum value sub-block, so as to determine each target sub-block maximum value among the maximum values of each sub-block based on the target position.

[0060] Finally, compare the maximum values of each target sub-block, so as to determine the maximum pooling value of the first specification data based on the maximum value among the obtained maximum values of each target sub-block.

[0061] As can be seen from the above, transfer the data to be processed to the vector processor; in the vector processor, determine the block window pane according to the pooling window pane and the preset step size; use the pooling window pane to block the data to be processed to obtain the first specification data, and then divide it into the second specification data according to the block window pane; determine the maximum value of each sub-block of the second specification data according to the preset step size, the block window pane, and the pooling window pane, and then obtain the maximum pooling value of the first specification data. In this way, the present application can process the maximum pooling operation of a larger pooling window pane in the vector processor, improving the operation efficiency of the maximum pooling.

[0062] Next, in combination with Figure 4 and Figure 5 shown in the schematic diagram, the technical solution of the embodiment of the present application will be specifically described.

[0063] In a specific embodiment, in Figure 5 , kernel is 9x9. When calculating MAX POOL with kernel 9x9, according to the step size of 1, it can be divided into 4 5x5 blocks for calculation. Obviously, there is data intersection among these 4 5x5 blocks, that is, Figure 5 in the fifth row and fifth column of the first table from left to right in

[0064] Specifically, Figure 5The first table from left to right in the figure is 9x9 data (i.e., the first specification data). Then, perform the 5x5s1 operation on the 9x9 numbers in the table to obtain the second table from left to right in the figure. Among them, 5x5 in 5x5s1 represents the block (i.e., the divided pane) is 5x5, and s1 represents the preset step size is 1.

[0065] It should be noted that the maximum value in the square from the first row to the fifth row and the first column to the fifth column in the first table is the value in the first row and the first column of the second table. Since the step size is 1, the divided pane can be translated 1 column to the right in the first table, that is, the maximum value in the square from the first row to the fifth row and the second column to the sixth column in the first table is the value in the first row and the second column of the second table. In this way, the maximum values of each sub-block are determined in turn, that is, the data in the second table.

[0066] After obtaining the maximum values of each sub-block, the position of the maximum value sub-block can be determined based on the size difference between the pooling pane and the divided pane. For example, assuming that the position of the value in the first row and the first column of the second table has been located, and the size difference between the pooling pane and the divided pane is 4, then on the basis of the position in the first row and the first column, translate 4 columns to the right to determine the position of another sub-block maximum value. By this way of taking numbers at intervals, the positions of the maximum values of each sub-block can be determined, and then 4 maximum values are determined, that is, the data in the third table from left to right in the figure.

[0067] In this embodiment, after obtaining the data in the third table, perform the MaxPooling operation on the data in the third table, that is, compare the data in the third table to determine the maximum value in the table. Obviously, 9.9 is the maximum value in the third table, so 9.9 is the maximum value in the first table.

[0068] Similarly, referring to Figure 4 As shown, when the specification of the data to be processed is 11x11, and the kernel is 9x9 and the step size is set to 2, the maximum values of each target sub-block can also be correctly obtained. In addition, in this way, when the block is 7x7, if the kernel can be divided into 2x2 blocks, the range that the kernel can cover is 8x8 to 14x14. Assuming that the kernel can be divided into 3x3 blocks, then the kernel will support the specifications of 11x11, 13x13, 15x15, 17x17, 19x19, and 21x21.

[0069] Correspondingly, referring to Figure 6 As shown, an embodiment of the present application provides a maximum pooling device applicable to a vector processor, including:

[0070] A data transmission module 11 for transmitting data to be processed in a double data rate synchronous dynamic random access memory to a vector processor;

[0071] A pane determination module 12 for determining a partitioned pane in the vector processor based on a pooling pane and a preset stride; wherein, the partitioned pane is smaller than the pooling pane;

[0072] A data partitioning module 13 for partitioning the data to be processed based on the pooling pane to obtain first-specification data, and partitioning the first-specification data into second-specification data based on the partitioned pane;

[0073] A pooling value determination module 14 for determining the maximum value of each sub-block of the second-specification data based on the preset stride, the partitioned pane, and the pooling pane, so as to determine the maximum pooling value of the first-specification data based on the maximum value of each sub-block

[0074] As can be seen from the above, the data to be processed is transmitted to the vector processor; in the vector processor, a partitioned pane is determined according to the pooling pane and the preset stride; the data to be processed is partitioned by the pooling pane to obtain first-specification data, and then the first-specification data is divided into second-specification data according to the partitioned pane; the maximum value of each target sub-block of the second-specification data is determined according to the preset stride, the partitioned pane, and the pooling pane, and then the maximum pooling value of the first-specification data is obtained. In this way, the present application can process the maximum pooling operation of a larger pooling pane in the vector processor, improving the operation efficiency of the maximum pooling.

[0075] In some specific embodiments, the data transmission module 11 specifically includes:

[0076] A data acquisition unit for acquiring data to be processed in a double data rate synchronous dynamic random access memory; wherein, the data volume of the data to be processed is greater than the data volume of the first-specification data;

[0077] A first data transmission unit for transmitting the data to be processed to a preset initial data storage space in the vector processor by using a direct memory access component in the vector processor.

[0078] In some specific embodiments, the data partitioning module 13 specifically further includes:

[0079] A second data transmission unit for transmitting the data to be processed in the preset initial data storage space to a vector register by using a vector memory access unit.

[0080] In some specific embodiments, the pane determination module 12 specifically includes:

[0081] A pane determination unit, configured to determine a segmented cross-pane based on the pooling pane and a preset stride when the size corresponding to the pooling pane is odd; wherein, there is data crossing between sub-blocks corresponding to the segmented cross-pane.

[0082] In some specific embodiments, the data segmentation module 13 specifically includes:

[0083] A data segmentation unit, configured to segment the first specification data based on the segmentation pane and obtain each sub-block; wherein, the data in each sub-block is second specification data.

[0084] In some specific embodiments, the pooling value determination module 14 specifically includes:

[0085] A maximum value storage unit, configured to calculate maximum values of a preset number of sub-blocks based on a preset maximum value calculation rule and the preset stride, and store the sub-block maximum values into a preset target data storage space in the vector processor by using the vector memory access unit.

[0086] In some specific embodiments, the pooling value determination module 14 specifically includes:

[0087] An interval determination unit, configured to determine a target sub-block interval by using a size difference between the pooling pane and the segmentation pane;

[0088] A maximum value determination unit, configured to locate a target position of a maximum value sub-block in the preset target data storage space by using the target sub-block interval, so as to determine each target sub-block maximum value among the sub-block maximum values based on the target position;

[0089] A pooling value determination unit, configured to compare each target sub-block maximum value, so as to determine a maximum pooling value of the first specification data based on the maximum value among the obtained target sub-block maximum values.

[0090] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 7 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be considered as any limitation to the usage scope of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Wherein, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement relevant steps in the maximum pooling method applicable to a vector processor disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0091] In this embodiment, the power supply 23 is used to provide operating voltages for the various hardware devices on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and specific limitations thereof are not provided herein; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type can be selected according to specific application requirements, and specific limitations are not provided herein.

[0092] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be transient storage or permanent storage.

[0093] Among them, the operating system 221 is used to manage and control the various hardware devices and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the maximum pooling method applicable to the vector processor executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.

[0094] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the maximum pooling method applicable to the vector processor disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.

[0095] In this specification, the various embodiments are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0096] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0097] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules may be located in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well known in the art.

[0098] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0099] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this document to illustrate the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A max pooling method applicable to a vector processor, characterized in that Including: Transmitting the data to be processed in the double data rate synchronous dynamic random access memory to the vector processor; In the vector processor, determining a block window based on a pooling window pane and a preset step size; wherein, the block window is smaller than the pooling window pane; Chunking the data to be processed based on the pooling window pane to obtain first-specification data, and chunking the first-specification data into second-specification data based on the block window; Determining the maximum value of each sub-block of the second-specification data based on the preset step size, the block window pane, and the pooling window pane, so as to determine the maximum pooling value of the first-specification data based on the maximum value of each sub-block.

2. The maximum pooling method applicable to a vector processor according to claim 1, wherein The transmitting the data to be processed in the double data rate synchronous dynamic random access memory to the vector processor includes: Obtaining the data to be processed in the double data rate synchronous dynamic random access memory; wherein, the data volume of the data to be processed is greater than the data volume of the first-specification data; Using the direct memory access component in the vector processor to transmit the data to be processed to a preset initial data storage space in the vector processor.

3. The max pooling method applicable to a vector processor according to claim 2, characterized in that After chunking the first-specification data into second-specification data based on the block window pane, further including: Using the vector memory access unit to transmit the second-specification data in the preset initial data storage space to the vector register.

4. The max pooling method applicable to a vector processor according to any one of claims 1 to 3, characterized in that The determining the block window based on the pooling window pane and the preset step size includes: When the size corresponding to the pooling window pane is odd, determining a block cross window based on the pooling window pane and the preset step size; wherein, there is data crossing between the sub-blocks corresponding to the block cross window.

5. The maximum pooling method applicable to a vector processor according to claim 1, wherein The chunking the first-specification data into second-specification data based on the block window pane includes: Chunking the first-specification data based on the block window pane and obtaining each sub-block; wherein, the data in each sub-block is second-specification data.

6. The max pooling method applicable to a vector processor according to claim 3, wherein The determining the maximum value of each sub-block of the second-specification data based on the preset step size, the block window pane, and the pooling window pane includes: Calculating the maximum values of a preset number of sub-blocks based on a preset maximum value calculation rule and the preset step size, and using the vector memory access unit to save the maximum values of the sub-blocks to a preset target data storage space in the vector processor.

7. The max pooling method applicable to a vector processor according to claim 6, wherein The determining the maximum pooling value of the first-specification data based on the maximum value of each sub-block includes: Determining a target sub-block interval using the size difference between the pooling window pane and the block window pane; In the preset target data storage space, using the target sub-block interval to locate the target position of the maximum value sub-block, so as to determine each target sub-block maximum value among the maximum values of each sub-block based on the target position; Comparing the maximum values of each target sub-block, so as to determine the maximum pooling value of the first-specification data based on the maximum value among the obtained maximum values of each target sub-block.

8. A maximum pooling device applicable to a vector processor, characterized in that, Including: A data transmission module, configured to transmit the data to be processed in the double data rate synchronous dynamic random access memory to the vector processor; A window pane determination module, configured to determine a block window based on a pooling window pane and a preset step size in the vector processor; wherein, the block window is smaller than the pooling window pane; A data chunking module, configured to chunk the data to be processed based on the pooling pane to obtain first-specification data, and chunk the first-specification data into second-specification data based on the chunking pane; A pooling value determination module, configured to determine the maximum value of each sub-block of the second-specification data based on the preset step, the chunking pane, and the pooling pane, so as to determine the maximum pooling value of the first-specification data based on the maximum value of each sub-block.

9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the maximum pooling method applicable to a vector processor according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, the maximum pooling method applicable to a vector processor according to any one of claims 1 to 7 is implemented.