Convolution operation filling value generation method, convolution operation filling value application method, convolution operation filling value generation device, convolution operation filling value application device, medium, equipment and product
By generating the filling value of the convolution operation, the input data blocks are automatically filled with preset values in the convolution operation, which solves the problem of high data handling bandwidth in the convolution operation and improves the utilization of storage space and processor performance.
Patent Information
- Application Number
- CN202510412954.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
During the convolutional operation, the bandwidth required for data handling is high, resulting in pressure on the performance of artificial intelligence processors, and the prior art is difficult to effectively reduce bandwidth without adding additional hardware.
By generating the method of filling values for convolution operations, the input data blocks are automatically filled with preset values during convolution operations, and there is no need to prestore multiple preset value data in the storage space, thereby improving the utilization rate of the storage space and reducing the bandwidth of data handling.
It realizes automatic filling of preset values in convolution operations, improves the utilization rate of storage space, reduces the bandwidth of data handling, and improves the performance of artificial intelligence processors.
Smart Images

Figure CN119939095A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method for generating a convolution operation fill value, an application method, a device, a computer-readable storage medium, an electronic device, and a computer program product. Background Art
[0002] Artificial intelligence processors usually need to perform a lot of calculations when training and reasoning large models, and the efficiency of data transfer is one of the key factors affecting their performance. Especially in the process of convolution operations, a large amount of data reuse is involved, and the demand for data transfer bandwidth increases significantly, which puts pressure on the performance of artificial intelligence processors. Therefore, how to effectively reduce the bandwidth required for data transfer in convolution operations without adding additional hardware is an urgent problem to be solved. Summary of the invention
[0003] The purpose of the embodiments of the present invention is to provide a method for generating a convolution operation fill value, an application method, an apparatus, a computer-readable storage medium, an electronic device and a computer program product, which can automatically fill the input data block with a preset value during the convolution operation, without the need to pre-store multiple preset value data in the storage space, thereby effectively improving the utilization of the storage space and reducing the bandwidth of data transfer.
[0004] The first aspect of the present invention provides a method for generating a convolution operation padding value, comprising: Acquire an input data block and corresponding original attribute information from the first storage space; wherein the original attribute information includes: size data of the input data block in the original space structure and the space coordinates of each data point; Get the coordinates of the convolution starting point and the position offset of each weight in the convolution kernel and the center of the convolution kernel; According to the convolution starting point coordinates, the size data and the position offset corresponding to each weight, a mask sequence corresponding to each weight is obtained; wherein the mask sequence is used to indicate the filling position of the preset value of the corresponding weight in the convolution operation; According to the mask sequence, a convolution cover sequence corresponding to each of the weights is obtained.
[0005] Optionally, the original spatial structure is an N-dimensional tensor, wherein N≥1.
[0006] Optionally, the input data block is a one-dimensional array obtained by continuously storing the original spatial structure in row priority order.
[0007] Optionally, acquiring a mask sequence corresponding to each weight according to the convolution starting point coordinates, the size data and the position offset corresponding to each weight includes: Starting from the data point corresponding to the coordinates of the convolution starting point, the center of the convolution kernel sequentially traverses the remaining data points in the input data block; When the center of the convolution kernel traverses to the current data point, taking the spatial coordinates of the current data point as a reference, combined with the size data and the position offset, it is determined whether each of the weights needs to be masked to the preset value at the current coverage position, so as to obtain a mask sequence corresponding to each of the weights; wherein the mask sequence is used to indicate the filling position of the preset value of the corresponding weight in the convolution operation.
[0008] Optionally, the mask value in the mask sequence is obtained by the following formula: ; in, The mask sequence The mask value, used to represent the convolution kernel center traversed to the When there are data points, whether the corresponding weight needs to be masked to a preset value at the current coverage position, Indicates that the mask needs to be the default value. Indicates that the mask does not need to be the preset value; The corresponding weight is The position offset from the center of the convolution kernel in the dimensional direction; For the Data points in Coordinate values in the dimensional direction; The original spatial structure is Dimensional data in the dimensional direction; ; N is the number of dimensions of the original spatial structure.
[0009] Optionally, acquiring a convolution cover sequence corresponding to each of the weights according to the mask sequence includes: According to the convolution starting point coordinates and the position offset of each weight, the mask starting point position of the corresponding mask sequence in the second storage space is obtained; wherein the second storage space is a temporary storage area where the input data block is stored after being taken out from the first storage space; Based on each of the mask starting positions, the corresponding mask sequence is applied to the corresponding data interval in the second storage space to obtain a convolution cover sequence corresponding to each of the weights.
[0010] A second aspect of the present invention provides an application method of a convolution operation padding value, comprising: Obtain a convolution cover sequence corresponding to each weight in the convolution kernel; wherein the convolution cover sequence is obtained by adopting the method for generating a convolution operation filling value described in any one of the first aspects above; Calculate the product between each of the weights and the corresponding convolution cover sequence to obtain a corresponding weight response sequence; All the weight response sequences are summed up at corresponding positions to obtain a convolution result sequence.
[0011] A third aspect of the present invention provides a device for generating a convolution operation filling value, comprising: A data acquisition module, used to acquire an input data block and corresponding original attribute information from the first storage space; wherein the original attribute information includes: size data of the input data block in the original space structure and the space coordinates of each data point; The convolution positioning module is used to obtain the coordinates of the convolution starting point and the position offset of each weight in the convolution kernel and the center of the convolution kernel; A mask generation module, used to obtain a mask sequence corresponding to each weight according to the convolution starting point coordinates, the size data and the position offset corresponding to each weight; wherein the mask sequence is used to indicate the filling position of the preset value of the corresponding weight in the convolution operation; The mask application module is used to obtain the convolution cover sequence corresponding to each of the weights according to the mask sequence.
[0012] A fourth aspect of the present invention provides an application device for a convolution operation padding value, comprising: A sequence acquisition module, used to obtain a convolution cover sequence corresponding to each weight in the convolution kernel; wherein the convolution cover sequence is obtained by adopting the method for generating a convolution operation filling value described in any one of the first aspects above; A weight response module, used to calculate the product between each of the weights and the corresponding convolution cover sequence to obtain a corresponding weight response sequence; The result output module is used to sum the elements of all the weight response sequences at corresponding positions to obtain a convolution result sequence.
[0013] The fifth aspect embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located controls the method for generating a convolution operation fill value as described in any one of the first aspect above, or the method for applying a convolution operation fill value as described in the second aspect above.
[0014] The sixth aspect embodiment of the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the method for generating a convolution operation fill value as described in any one of the first aspect above, or the method for applying a convolution operation fill value as described in the second aspect above.
[0015] The seventh aspect embodiment of the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the method for generating the convolution operation fill value described in any one of the first aspect above, or the method for applying the convolution operation fill value described in the second aspect above when executing the computer program.
[0016] Compared with the prior art, the embodiments of the present invention provide a method for generating a convolution operation filling value, an application method, an apparatus, a computer-readable storage medium, an electronic device, and a computer program product, which have the following beneficial effects: the embodiments of the present invention directly obtain the input data block and the corresponding original attribute information from the first storage space, and obtain the mask sequence corresponding to each weight according to the convolution starting point coordinates, size data, and the position offset corresponding to each weight; then, the filling position mask indicated in each mask sequence is a preset value, and the convolution cover sequence corresponding to each weight is obtained. Therefore, the present invention can automatically fill the input data block with a preset value during the convolution operation, without the need to pre-store multiple preset value data in the storage space, thereby effectively improving the utilization of the storage space and reducing the bandwidth of data transfer. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flowchart of an embodiment of a method for generating a convolution operation filling value provided by the present invention; Figure 2 It is a schematic diagram of an embodiment of generation and application of a mask sequence corresponding to a one-dimensional convolution kernel provided by the present invention; Figure 3 It is a schematic diagram of an embodiment of generating and applying a mask sequence corresponding to a two-dimensional convolution kernel provided by the present invention; Figure 4 It is a flowchart of an embodiment of the application method of the convolution operation filling value provided by the present invention; Figure 5 It is a structural schematic diagram of an embodiment of a device for generating a convolution operation filling value provided by the present invention; Figure 6 It is a structural schematic diagram of an embodiment of an application device for convolution operation filling value provided by the present invention; Figure 7 is a structural schematic diagram of an embodiment of an electronic device provided by the present invention; Figure 8 It is a structural diagram of an embodiment of the artificial intelligence processor provided by the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this technical field without creative work are within the scope of protection of the present invention.
[0019] The artificial intelligence processor involved in the present invention can be any one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), a NPU (Neural network Processing Unit), a DPU (Deeplearning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit), which can be determined when the embodiments of the present invention are applied to specific products or technologies.
[0020] At present, there are two main methods for implementing convolution operations: the Image to Column method and the hardware implementation method based on the systolic array. In the process of implementing the convolution operation, the inventors found that these two methods have the following defects: (1) The Image to Column method converts multi-dimensional convolution operations into two-dimensional matrix multiplication operations, but it still needs to store a large amount of reused data, resulting in high bandwidth usage for data transfer.
[0021] (2) Although the hardware implementation method based on systolic arrays can accelerate convolution operations at the physical level, its implementation control is complex and the cost is high. In addition, the large array area easily leads to a high scrap rate, which limits the large-scale application of systolic arrays.
[0022] Therefore, how to effectively reduce the bandwidth required for data transfer in convolution operations without adding additional hardware is an urgent problem to be solved.
[0023] See also Figure 1 , is a flow chart of an embodiment of a method for generating a convolution operation filling value provided by the present invention.
[0024] To solve the above problem, a first aspect of the present invention provides a method for generating a convolution operation padding value, including steps S11 to S14, which are as follows: Step S11: acquiring an input data block and corresponding original attribute information from a first storage space; wherein the original attribute information includes: size data of the input data block in the original space structure and the space coordinates of each data point; Step S12: Obtain the coordinates of the convolution starting point and the position offset between each weight in the convolution kernel and the center of the convolution kernel; Step S13: acquiring a mask sequence corresponding to each weight according to the convolution starting point coordinates, the size data and the position offset corresponding to each weight; wherein the mask sequence is used to indicate the filling position of the preset value of the corresponding weight in the convolution operation; Step S14: According to the mask sequence, obtain the convolution cover sequence corresponding to each of the weights.
[0025] It should be noted that the original spatial structure is the original organization state of the input data block before the storage conversion. For example, the original spatial structure of a grayscale image is a two-dimensional matrix, and the original spatial structure of a color image is a three-dimensional matrix. The convolution cover sequence is the input sequence after filling with preset values, that is, the coverage data of the corresponding weights in the convolution kernel in the convolution sliding, including: original data and / or preset values for filling. In an embodiment of the present invention, the preset value for filling can be zero or any other arbitrary value, depending on the specific filling operation used. For example, the preset value for filling in the zero filling operation is 0; the preset value for filling in the constant filling operation is a specified constant value; the preset value for filling in the repeated filling operation is the edge value of the input data block; the preset value for filling in the mirror filling operation is the filling value generated by reflecting the edge value.
[0026] In the specific implementation, firstly, the input data block involved in the convolution operation and its original attribute information are obtained from the first storage space (such as a mechanical hard disk, solid-state hard disk, and memory, etc.); wherein the original attribute information includes: the size data of the input data block in the original spatial structure (i.e., the size scale in each dimension), and the spatial coordinates of each data point in the original spatial structure. Subsequently, the coordinates of the convolution starting point are obtained, i.e., the starting coordinates of the convolution kernel for the sliding window operation on the input data block. At the same time, the position offset of each weight in the convolution kernel relative to the convolution center is also obtained, which is used to determine the specific coverage position point of the weight in the subsequent convolution sliding window process; for example, the position offset of a 3×3 convolution kernel is: ;in, The vertical position offset between the corresponding weight and the center of the convolution kernel is -1, and the horizontal position offset is 1. Next, according to the convolution starting point coordinates, size data and the position offset corresponding to each weight, the mask sequence corresponding to each weight is obtained; finally, the filling position mask indicated in each mask sequence is a preset value, and the convolution cover sequence corresponding to each weight (that is, the cover data including the preset value) is obtained; for example, for a The convolution kernel contains 9 weights, and finally 9 convolution coverage sequences are obtained.
[0027] It is worth noting that the convolution operation is a spatial correlation calculation, so the original attribute information involved in the embodiment of the present invention is also stored by default in the prior art (such as for determining the location of the convolution starting point), and no additional memory overhead is introduced. In addition, compared with the prior art that requires pre-storing multiple preset value data in the storage space, the embodiment of the present invention can automatically fill the input data block with preset values during the convolution operation, thereby effectively improving the utilization of the storage space and reducing the bandwidth of data transfer.
[0028] In an optional embodiment, the original spatial structure is an N-dimensional tensor, wherein N≥1.
[0029] It should be noted that the embodiment of the present invention obtains a convolution cover sequence corresponding to each weight of the convolution kernel. When calculating the convolution result, there is no need to consider the local sliding window operation of the convolution kernel in the original spatial structure. Specifically, it is only necessary to multiply each weight by its corresponding convolution cover sequence to obtain a weight response sequence; and then sum all weight response sequences at the corresponding positions to obtain a convolution result sequence. It is no longer necessary to perform local calculations on the receptive field area of the convolution kernel each time the sliding window operation is performed, but the sliding window process of the convolution kernel in the original spatial structure is decoupled into the position calculation of the parallel convolution cover sequence. Based on this, the embodiment of the present invention has extremely strong versatility and adaptability, and can be widely used in various original spatial structures from one dimension to high dimensions; for example, one-dimensional time series data, two-dimensional grayscale images, three-dimensional color images, and multi-dimensional sensor data, etc. At the same time, whether it is a small-sized convolution kernel or a large-sized convolution kernel, the embodiment of the present invention can be flexibly adapted.
[0030] In an optional embodiment, the input data block is a one-dimensional array obtained by continuously storing the original spatial structure in row-priority order.
[0031] It should be noted that the layout of the input data block in the first storage space is a one-dimensional array obtained by continuously storing the data points in the original spatial structure in row priority order, that is, the N-dimensional data is expanded into one-dimensional data by row and continuously stored in the first storage space. This matches the default row priority order of the convolution kernel sliding, thereby maximizing the continuity of memory access to reduce access latency.
[0032] In an optional embodiment, acquiring a mask sequence corresponding to each weight according to the convolution starting point coordinates, the size data and the position offset corresponding to each weight includes: Starting from the data point corresponding to the coordinates of the convolution starting point, the center of the convolution kernel sequentially traverses the remaining data points in the input data block; When the center of the convolution kernel traverses to the current data point, taking the spatial coordinates of the current data point as a reference, combined with the size data and the position offset, it is determined whether each of the weights needs to be masked to the preset value at the current coverage position, so as to obtain a mask sequence corresponding to each of the weights; wherein the mask sequence is used to indicate the filling position of the preset value of the corresponding weight in the convolution operation.
[0033] It should be noted that starting from the data point corresponding to the coordinates of the starting point of the convolution, each remaining data point in the input data block is traversed in turn according to the convolution sliding window order. During the traversal of the convolution kernel center, for each weight, the current coverage position can be determined by its position offset. If the coverage position exceeds any boundary in the original spatial structure, it is determined that the coverage position needs to be masked to a preset value. When the convolution kernel center traverses to the last data point of the input data block, the mask sequence corresponding to each weight can be obtained.
[0034] In an optional embodiment, the mask value in the mask sequence is obtained by the following formula: ; in, The mask sequence The mask value, used to represent the convolution kernel center traversed to the When there are data points, whether the corresponding weight needs to be masked to a preset value at the current coverage position, Indicates that the mask needs to be the default value. Indicates that the mask does not need to be the preset value; The corresponding weight is The position offset from the center of the convolution kernel in the dimensional direction; For the Data points in Coordinate values in the dimensional direction; The original spatial structure is Dimensional data in the dimensional direction; ; N is the number of dimensions of the original spatial structure.
[0035] For example, Figure 2 FIG. 1 is a schematic diagram of an embodiment of the generation and application of a mask sequence corresponding to a one-dimensional convolution kernel provided by the present invention. Figure 2 In (a), the default padding value is 0, and the 1×3 convolution kernel contains three weights w0, w1, and w2. The position offsets (filter_offset(1), filter_offset(2)) are (0,-1), (0,0), and (0,1), respectively. Among them, filter_offset(1) is the position offset of the corresponding weight and the center of the convolution kernel in the vertical direction (d=1), and filter_offset(2) is the position offset in the horizontal direction (d=2). Taking the mask sequence of weight w0 as an example, when the coordinates of the convolution kernel center at the starting point of the convolution are (0,0), the above mask value calculation formula is used to obtain mask(1)=1, and the mask point at the position (coor(i,1)+filter_offset(1),coor(i,2)+filter_offset(2)) needs to be the preset value, that is, (0,-1); when the coordinates covered by the convolution kernel center are (0,1), the above mask value calculation formula is used to obtain mask(1)=0; and so on, until the convolution center traverses to the last data point of the input data block, and finally the mask sequence of w0 is obtained, as shown in Figure 2 (b) shown.
[0036] In an optional embodiment, obtaining the convolution cover sequence corresponding to each of the weights according to the mask sequence includes: According to the convolution starting point coordinates and the position offset of each weight, the mask starting point position of the corresponding mask sequence in the second storage space is obtained; wherein the second storage space is a temporary storage area where the input data block is stored after being taken out from the first storage space; Based on each of the mask starting positions, the corresponding mask sequence is applied to the corresponding data interval in the second storage space to obtain a convolution cover sequence corresponding to each of the weights.
[0037] It should be noted that in the second storage space, the data located before and after the input data block, such as Figure 2 The ① and ② positions in Figure 3 In The data at the position can come from the additional data taken out from the first storage space. For example, in the process of batch processing image data, the data of multiple pictures are usually read continuously. Then, for the input data block corresponding to a certain picture, the data at the front and back positions in the second storage space are other picture data read from the first storage space. On the other hand, it can also come from data generated by other instructions, which are stored in the front and back positions of the input data block.
[0038] In order to more clearly describe the technical solution provided by the embodiments of the present invention, some specific embodiments are provided below for reference: like Figure 2 As shown in the figure, taking the weight w0 in the 1×3 convolution kernel as an example, assuming that the coordinates of the convolution starting point are (0,0) and the position offset of w0 is (0,-1), then the mask starting position corresponding to w0 is (0,-1), which is the first position of the convolution starting point (marked as position ①); the mask sequence of w0 will start masking from position ① until the subsequent 25th data point (spatial coordinates are (4,3)), thereby obtaining the convolution coverage sequence corresponding to w0; this masking process does not occupy additional memory space. Figure 2 It can be seen that for weight w0, the leftmost 0 value needs to be filled in the original spatial structure, which is obtained by masking the rightmost data in the previous row to 0. Similarly, the position offset of weight w2 is (0,1), and the corresponding mask starting point coordinates are (0,1). The mask sequence of w2 will start masking from the last bit of the convolution starting point until it ends at position ②, thereby obtaining the convolution coverage sequence corresponding to w2.
[0039] like Figure 3 As shown, it is a schematic diagram of an embodiment of the generation and application of the mask sequence corresponding to the two-dimensional convolution kernel provided by the present invention. Taking the weight q0 in the 3×3 convolution kernel as an example, assuming that the coordinates of the convolution starting point are (0,0) and the position offset of q0 is (-1,-1), then the mask starting position corresponding to q0 is (-1,-1), that is, in the first 6 bits of the convolution starting point (marked x1 position); the mask sequence of q0 will start masking from the x1 position until the subsequent 25th data point (spatial coordinates are (3,3)), thereby obtaining the convolution coverage sequence corresponding to q0.
[0040] See also Figure 4 , is a flow chart of an embodiment of the application method of the convolution operation filling value provided by the present invention.
[0041] The second aspect of the present invention provides a method for applying a convolution operation padding value, including steps S21 to S23, which are specifically as follows: Step S21: obtaining a convolution cover sequence corresponding to each weight in the convolution kernel; wherein the convolution cover sequence is obtained by adopting the method for generating a convolution operation filling value described in any embodiment of the first aspect above; Step S22: Calculate the product between each weight and the corresponding convolution cover sequence to obtain a corresponding weight response sequence; Step S23: summing up the elements of all the weight response sequences at corresponding positions to obtain a convolution result sequence.
[0042] Specifically, in combination with the above embodiments, after obtaining the convolution covering sequence corresponding to each weight in the convolution kernel through any embodiment of the first aspect above, each weight is multiplied element by element with its corresponding convolution covering sequence to obtain the corresponding weight response sequence. Then, all weight response sequences are element-summed at corresponding positions to finally obtain an accurate convolution result sequence, which is equivalent to the convolution operation result of the input data after padding and the convolution kernel. The embodiment of the present invention decouples the sliding window process of the convolution kernel in the original spatial structure into the position calculation of the parallel convolution covering sequence, and can efficiently and accurately complete the convolution operation.
[0043] See also Figure 5 , is a structural diagram of an embodiment of a device for generating a convolution operation filling value provided by the present invention.
[0044] A third aspect of the present invention provides a device for generating a convolution operation filling value, including: The data acquisition module 11 is used to acquire the input data block and the corresponding original attribute information from the first storage space; wherein the original attribute information includes: the size data of the input data block in the original space structure and the space coordinates of each data point; The convolution positioning module 12 is used to obtain the coordinates of the convolution starting point and the position offset between each weight in the convolution kernel and the center of the convolution kernel; The mask generation module 13 is used to obtain a mask sequence corresponding to each weight according to the convolution starting point coordinates, the size data and the position offset corresponding to each weight; wherein the mask sequence is used to indicate the filling position of the preset value of the corresponding weight in the convolution operation; The mask application module 14 is used to obtain a convolution cover sequence corresponding to each of the weights according to the mask sequence.
[0045] It should be noted that the device for generating convolution operation filling values provided in the embodiment of the third aspect of the present invention can implement all processes of the method for generating convolution operation filling values described in any embodiment of the first aspect above. The functions of each module and unit in the device and the technical effects achieved are respectively the same as the functions and technical effects achieved by the method for generating convolution operation filling values described in any embodiment of the first aspect above, and will not be repeated here.
[0046] See also Figure 6 , is a structural diagram of an embodiment of an application device for convolution operation filling value provided by the present invention.
[0047] A fourth aspect of the present invention provides an application device for a convolution operation padding value, comprising: A sequence acquisition module 21, used to obtain a convolution cover sequence corresponding to each weight in the convolution kernel; wherein the convolution cover sequence is obtained by adopting the method for generating a convolution operation filling value described in any embodiment of the first aspect above; A weight response module 22, used for calculating the product between each of the weights and the corresponding convolution cover sequence to obtain a corresponding weight response sequence; The result output module 23 is used to sum the elements of all the weight response sequences at corresponding positions to obtain a convolution result sequence.
[0048] It should be noted that the application device of convolution operation filling value provided in the fourth aspect of the embodiment of the present invention can realize all processes of the application method of convolution operation filling value described in the above-mentioned second aspect embodiment. The functions of each module and unit in the device and the technical effects achieved are respectively the same as the functions and technical effects achieved by the application method of convolution operation filling value described in the above-mentioned second aspect embodiment, and they will not be repeated here.
[0049] The fifth aspect embodiment of the present invention also provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located controls the method for generating a convolution operation filling value as described in any embodiment of the first aspect above, or the method for applying a convolution operation filling value as described in any embodiment of the second aspect above.
[0050] The sixth aspect of the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the method for generating a convolution operation fill value as described in any embodiment of the first aspect, or the method for applying a convolution operation fill value as described in any embodiment of the second aspect.
[0051] See also Figure 7, is a schematic structural diagram of an embodiment of an electronic device provided by the present invention.
[0052] The seventh aspect embodiment of the present invention also provides an electronic device, including a processor 31, a memory 32, and a computer program stored in the memory 32 and configured to be executed by the processor 31, wherein the processor 31 implements the method for generating the convolution operation filling value described in any embodiment of the first aspect when executing the computer program, or the method for applying the convolution operation filling value described in any embodiment of the second aspect.
[0053] Preferably, the computer program can be divided into one or more modules / units (such as computer program 1, computer program 2, ...), and the one or more modules / units are stored in the memory 32 and executed by the processor 31 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0054] The processor 31 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor 31 may be any conventional processor. The processor 31 is the control center of the electronic device, and various parts of the electronic device are connected using various interfaces and lines.
[0055] The memory 32 mainly includes a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function, etc., and the data storage area can store related data, etc. In addition, the memory 32 can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, and a flash card (Flash Card), etc., or the memory 32 can also be other volatile solid-state storage devices.
[0056] It should be noted that the above electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 7 The structural block diagram shown is only a structural example of the above-mentioned electronic device, and does not constitute a structural limitation of the above-mentioned electronic device. The above-mentioned electronic device may include more or less components than shown in the figure, or combine certain components, or different components.
[0057] See also Figure 8 , is a schematic diagram of the structure of an embodiment of an artificial intelligence processor provided by the present invention; The artificial intelligence processor provided by the embodiment of the present invention includes: an instruction unit 41, an instruction parsing unit 42, a mask calculation unit 43, a second memory 45 and a result calculation unit 46; wherein the instruction unit 41 is used to generate and send a convolution operation instruction; the instruction parsing unit 42 is used to receive and parse the convolution operation instruction, obtain the control information in the instruction, and send the control information to the mask calculation unit 43; the mask calculation unit 43 is used to execute the method for generating the convolution operation filling value described in any embodiment of the first aspect after receiving the control information; the result calculation unit 46 is used to receive the convolution cover sequence transmitted by the mask calculation unit 43, and execute the method for applying the convolution operation filling value described in the embodiment of the second aspect; the second memory 45 is an intermediate storage device, used to store the input data block loaded by the mask calculation unit 43 from the first memory.
[0058] It is worth mentioning that Figure 8 The first memory 44 in the embodiment may be located in the artificial intelligence processor, or in an external memory or other external device, and is used to store input data blocks.
[0059] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for generating a convolution operation filling value, characterized in that: include: Acquire an input data block and corresponding original attribute information from the first storage space; wherein the original attribute information includes: size data of the input data block in the original space structure and the space coordinates of each data point; Get the coordinates of the convolution starting point and the position offset of each weight in the convolution kernel and the center of the convolution kernel; According to the convolution starting point coordinates, the size data and the position offset corresponding to each weight, a mask sequence corresponding to each weight is obtained; wherein the mask sequence is used to indicate the filling position of the preset value of the corresponding weight in the convolution operation; According to the mask sequence, a convolution cover sequence corresponding to each of the weights is obtained.
2. The method for generating a convolution operation filling value according to claim 1, wherein: The original spatial structure is an N-dimensional tensor, wherein N≥1.
3. The method for generating a convolution operation filling value according to claim 1, wherein: The input data block is a one-dimensional array obtained by continuously storing the original spatial structure in row priority order.
4. The method for generating a convolution operation filling value according to claim 1, wherein: The step of obtaining a mask sequence corresponding to each weight according to the convolution starting point coordinates, the size data and the position offset corresponding to each weight comprises: Starting from the data point corresponding to the coordinates of the convolution starting point, the center of the convolution kernel sequentially traverses the remaining data points in the input data block; When the center of the convolution kernel traverses to the current data point, taking the spatial coordinates of the current data point as a reference, combined with the size data and the position offset, it is determined whether each of the weights needs to be masked to the preset value at the current coverage position to obtain a mask sequence corresponding to each of the weights.
5. The method for generating a convolution operation filling value according to claim 4, wherein: The mask value in the mask sequence is obtained by the following formula: ; in, The mask sequence The mask value, used to represent the convolution kernel center traversed to the When there are data points, whether the corresponding weight needs to be masked to a preset value at the current coverage position, Indicates that the mask needs to be the default value. Indicates that the mask does not need to be the preset value; The corresponding weight is The position offset from the center of the convolution kernel in the dimensional direction; For the Data points in Coordinate values in the dimensional direction; The original spatial structure is Dimensional data in the dimensional direction; ; N is the number of dimensions of the original spatial structure.
6. The method for generating a convolution operation filling value according to claim 4, wherein: The step of obtaining, according to the mask sequence, a convolution cover sequence corresponding to each of the weights comprises: According to the convolution starting point coordinates and the position offset of each weight, the mask starting point position of the corresponding mask sequence in the second storage space is obtained; wherein the second storage space is a temporary storage area where the input data block is stored after being taken out from the first storage space; Based on each of the mask starting positions, the corresponding mask sequence is applied to the corresponding data interval in the second storage space to obtain a convolution cover sequence corresponding to each of the weights.
7. A method for applying a convolution operation fill value, characterized in that: include: Obtain a convolution cover sequence corresponding to each weight in the convolution kernel; wherein the convolution cover sequence is obtained by adopting the method for generating a convolution operation filling value according to any one of claims 1 to 6; Calculate the product between each of the weights and the corresponding convolution cover sequence to obtain a corresponding weight response sequence; All the weight response sequences are summed up at corresponding positions to obtain a convolution result sequence.
8. A device for generating a convolution operation filling value, characterized in that: include: A data acquisition module, used to acquire an input data block and corresponding original attribute information from the first storage space; wherein the original attribute information includes: size data of the input data block in the original space structure and the space coordinates of each data point; The convolution positioning module is used to obtain the coordinates of the convolution starting point and the position offset of each weight in the convolution kernel and the center of the convolution kernel; A mask generation module, used to obtain a mask sequence corresponding to each weight according to the convolution starting point coordinates, the size data and the position offset corresponding to each weight; wherein the mask sequence is used to indicate the filling position of the preset value of the corresponding weight in the convolution operation; The mask application module is used to obtain the convolution cover sequence corresponding to each of the weights according to the mask sequence.
9. An application device for convolution operation filling value, characterized in that: include: A sequence acquisition module, used to obtain a convolution cover sequence corresponding to each weight in the convolution kernel; wherein the convolution cover sequence is obtained by adopting the method for generating a convolution operation filling value according to any one of claims 1 to 6; A weight response module, used to calculate the product between each of the weights and the corresponding convolution cover sequence to obtain a corresponding weight response sequence; The result output module is used to sum the elements of all the weight response sequences at corresponding positions to obtain a convolution result sequence.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program; wherein, when the computer program is run, it controls the device where the computer-readable storage medium is located to execute the method for generating a convolution operation filling value as described in any one of claims 1 to 6, or the method for applying a convolution operation filling value as described in claim 7.
11. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the method for generating a convolution operation filling value as described in any one of claims 1 to 6, or the method for applying a convolution operation filling value as described in claim 7.
12. An electronic device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method for generating a convolution operation filling value as described in any one of claims 1 to 6, or the method for applying a convolution operation filling value as described in claim 7.
Citation Information
Patent Citations
Deformable convolution accelerator and deformable convolution acceleration method
CN113516235A
Method of performing convolution calculation, computing device, medium
CN117763275A
Processor operating method and device, electronic device and program product
CN119005274A
Convolutional network masking method, system and device and storage medium
CN119129655A
Subflow type sparse convolution rule table generation method, convolution operation method and device
CN119493949A
Cited By
Pixel parallel depth operation implementation method and device, medium, equipment and product
CN120298196A
Sliding input method, device, equipment and medium
CN121187495A