Data Shift Processing Method, Device and Equipment Based on Convolutional Neural Network

By using multiple convolutional computing layers to process data in a convolutional neural network, the problem that data shift tasks cannot be executed on adapted neural network devices is solved, efficient data shift operations are achieved, and hardware resource utilization and computing efficiency are improved.

CN114154623BActive Publication Date: 2025-07-11GUANGZHOU XIAOPENG CONNECTIVITY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111485792.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-07-11
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

In the prior art, data shift processing tasks cannot be directly executed on the operating processing equipment adapted to the neural network structure, resulting in low computing efficiency, high latency and low hardware resource utilization.

Method used

By using multiple different convolutional computing layers to process the preprocessed data in a convolutional neural network, the logical shift and cyclic shift operations of the data are realized, including converting the input data into matrix data blocks and convolution processing through multiple convolutional computing layers, and finally converting it into a serial data stream.

Benefits of technology

It improves hardware resource utilization, improves computing efficiency, reduces delay, and can efficiently perform data shift tasks on operation and processing devices that are adapted to neural network structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114154623B_ABST
    Figure CN114154623B_ABST
Patent Text Reader

Abstract

The present application relates to a data shift processing method, apparatus, and device based on a convolutional neural network. The data shift processing method based on a convolutional neural network is applied to a processor, and the method includes: obtaining preprocessed data by preprocessing input data; according to different shift operation requirements, processing the preprocessed data through a plurality of different convolutional calculation layers of a shift unit in the convolutional neural network to obtain shifted operation data; and obtaining output data by postprocessing the shifted operation data. The solution provided by the present application can implement the execution of data shift tasks on a running processing device adapted to the neural network structure, which can improve the operation efficiency, reduce the latency, and improve the utilization rate of hardware resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technologies, and in particular, to a data shift processing method, apparatus, and device based on a convolutional neural network. Background Art

[0002] In the field of data processing, it is usually necessary to rearrange tensor data in sequence, which is generally implemented using algorithms such as Reshape and Transpose. Among them, tensor data can be a multi-dimensional array, and a common way to implement data rearrangement is: the shift operation of data.

[0003] With the rapid development of deep learning algorithms based on convolutional neural networks, convolutional neural networks have been widely applied in different technical fields. Convolutional neural networks can run on operation processing devices adapted to the neural network structure. The operation processing devices adapted to the neural network structure include neural network dedicated chips (such as convolutional neural network inference chips, ASIC (Application Specific Integrated Circuit) chips, etc.), general-purpose processors, and image processors. These operation processing devices that can adapt to the neural network structure have excellent performance and have been widely used in the market.

[0004] However, currently, data shift processing tasks are generally implemented by externally connecting an ARM (Advanced RISC Machine) chip or by using DMA (Direct Memory Access), and cannot directly run on operation processing devices adapted to the neural network structure, which in turn leads to problems such as low operation efficiency, high latency, and low hardware resource utilization. Summary of the Invention

[0005] To solve or partially solve the problems existing in the related technologies, this application provides a data shift processing method, apparatus, and device based on a convolutional neural network, which can implement the execution of data shift tasks on operation processing devices adapted to the neural network structure, improve operation efficiency, reduce latency, and improve hardware resource utilization.

[0006] The first aspect of this application provides a data shift processing method based on a convolutional neural network, which is applied to a processor. The method includes:

[0007] Obtain preprocessed data by preprocessing the input data;

[0008] According to different shift operation requirements, process the preprocessed data through multiple different convolutional calculation layers of a shift unit in the convolutional neural network to obtain shifted operation data;

[0009] The shifted operation data is post - processed to obtain output data.

[0010] In one embodiment, the pre - processing of the input data to obtain pre - processed data includes: converting the input serial data stream into a matrix data block through a splicing operation;

[0011] The post - processing of the shifted operation data to obtain output data includes: converting the shifted operation data in matrix data block format into a serial data stream.

[0012] In one embodiment, the processing of the pre - processed data through multiple different convolutional calculation layers by a shift unit in a convolutional neural network to obtain shifted operation data includes:

[0013] After performing convolutional processing on the pre - processed data in the first convolutional calculation layer of the logical shift unit to output two - channel convolutional data, performing convolutional processing through three different intermediate convolutional calculation layers, and then performing convolutional data channel merging processing through the fifth convolutional calculation layer to obtain the shifted operation data of logical shift; or,

[0014] After performing convolutional processing on the pre - processed data in the first convolutional calculation layer of the cyclic shift unit to output three - channel convolutional data, performing convolutional processing through four different intermediate convolutional calculation layers, and then performing convolutional data channel merging processing through the sixth convolutional calculation layer to obtain the shifted operation data of cyclic shift.

[0015] In one embodiment, the processing of the pre - processed data in the first convolutional calculation layer of the logical shift unit to output two - channel convolutional data, performing convolutional processing through three different intermediate convolutional calculation layers, and then performing convolutional data channel merging processing through the fifth convolutional calculation layer to obtain the shifted operation data of logical shift includes:

[0016] After performing convolutional processing on the pre - processed data in the first convolutional calculation layer of the logical shift unit to output two - channel convolutional data, performing convolutional processing through three different intermediate convolutional calculation layers set according to the logical left - shift operation rule or the logical right - shift operation rule, and then performing convolutional data channel merging processing through the fifth convolutional calculation layer to obtain the shifted operation data of logical left - shift or logical right - shift respectively.

[0017] In one embodiment, the processing of the pre - processed data in the first convolutional calculation layer of the cyclic shift unit to output three - channel convolutional data, performing convolutional processing through four different intermediate convolutional calculation layers, and then performing convolutional data channel merging processing through the sixth convolutional calculation layer to obtain the shifted operation data of cyclic shift includes:

[0018] After the preprocessed data is convolved in the first convolutional calculation layer of the cyclic shift unit to output three-channel convolutional data, it is convolved through four different intermediate convolutional calculation layers set according to the rules of cyclic left shift operation or cyclic right shift operation, and then the convolutional data channel merging process is performed through the sixth convolutional calculation layer to obtain the shift operation data of cyclic left shift or cyclic right shift respectively.

[0019] In one embodiment, each of the three different intermediate convolutional calculation layers includes two different convolutional kernels, where each convolutional kernel includes two channel parameters; or,

[0020] Each of the four different intermediate convolutional calculation layers includes three different convolutional kernels, where each convolutional kernel includes three channel parameters.

[0021] In one embodiment, the byte length of the input serial data stream is a square number or a non-square number of any positive integer;

[0022] When it is a non-square number, the serial data stream is segmented into data segments with a byte length of a square number of any positive integer including overlap for preprocessing, and after obtaining the shift operation data, splicing and merging are performed according to the principle of giving priority to the backend data segment in the shift direction.

[0023] The second aspect of the present application provides a data shift processing device based on a convolutional neural network, which is applied to a processor. The device includes:

[0024] A preprocessing module for preprocessing the input data to obtain preprocessed data;

[0025] A shift operation module for processing the preprocessed data of the preprocessing module through a plurality of different convolutional calculation layers in the convolutional neural network according to different shift operation requirements to obtain shift operation data;

[0026] A postprocessing module for postprocessing the shift operation data obtained by the shift operation module to obtain output data.

[0027] In one embodiment, the shift operation module includes:

[0028] A logical shift unit for convolving the preprocessed data in the first convolutional calculation layer to output two-channel convolutional data, then convolving through three different intermediate convolutional calculation layers, and then performing convolutional data channel merging processing through the fifth convolutional calculation layer to obtain the shift operation data of logical shift; or,

[0029] A cyclic shift unit is used to perform convolution processing on preprocessed data in the first convolutional calculation layer to output three-channel convolutional data, then perform convolution processing through four different intermediate convolutional calculation layers, and then perform convolution data channel merging processing through the sixth convolutional calculation layer to obtain shifted operation data of cyclic shift.

[0030] The third aspect of the present application provides an artificial intelligence chip, including the data shift processing device based on convolutional neural network as described above.

[0031] The fourth aspect of the present application provides a computing device, and the computing device includes the above artificial intelligence chip.

[0032] The fifth aspect of the present application provides a board card, and the board card includes: a storage device, an interface device, a control device, and the above artificial intelligence chip;

[0033] wherein, the artificial intelligence chip is respectively connected to the storage device, the control device, and the interface device;

[0034] The storage device is used to store data;

[0035] The interface device is used to realize data transmission between the artificial intelligence chip and an external device;

[0036] The control device is used to monitor the state of the artificial intelligence chip.

[0037] The sixth aspect of the present application provides a computing device, including:

[0038] a processor; and

[0039] a memory, on which executable code is stored, and when the executable code is executed by the processor, the processor is enabled to execute the method as described above.

[0040] The seventh aspect of the present application provides a computer-readable storage medium, on which executable code is stored, and when the executable code is executed by a processor of a computing device, the processor is enabled to execute the method as described above.

[0041] The technical solution provided by the present application may include the following beneficial effects:

[0042] The method provided by this application is applied to a processor. After the processor preprocesses the input data to obtain preprocessed data, according to different shift operation requirements, in a convolutional neural network, the preprocessed data can be processed through multiple different convolutional calculation layers of a shift unit to obtain shifted operation data; then the shifted operation data is post-processed to obtain output data. In this way, multiple different convolutional calculation layers can be used to construct shift operation calculations, enabling the execution of data shift tasks on a running processing device adapted to the neural network structure, thereby improving the utilization rate of hardware resources, enhancing the operation efficiency, and reducing the latency.

[0043] Further, in the method provided by this application, the processor can convert a serial data stream into a matrix data block, enabling the input data to be received and processed by the convolutional neural network, and finally converting the shifted operation data in the matrix data block format back into a serial data stream.

[0044] Further, in the method provided by this application, the processor can respectively construct processing methods such as logical left shift, logical right shift, circular left shift, and circular right shift according to different shift operation requirements, which can meet different data shift processing tasks of the processor.

[0045] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] By describing the exemplary embodiments of this application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of this application will become more obvious. Among them, in the exemplary embodiments of this application, the same reference numerals generally represent the same components.

[0047] Figure 1 is a schematic flowchart of a data shift processing method based on a convolutional neural network shown in an embodiment of this application;

[0048] Figure 2 is another schematic flowchart of a data shift processing method based on a convolutional neural network shown in an embodiment of this application;

[0049] Figure 3 is another schematic flowchart of a data shift processing method based on a convolutional neural network shown in an embodiment of this application;

[0050] Figure 4 is a schematic diagram of the conversion between a serial data stream and a matrix data block shown in an embodiment of this application;

[0051] Figure 5 is a schematic diagram of the logical left shift processing process shown in an embodiment of this application;

[0052] Figure 6It is a schematic diagram of the logical right shift processing process shown in the embodiments of the present application;

[0053] Figure 7 It is a schematic diagram of the circular left shift processing process shown in the embodiments of the present application;

[0054] Figure 8 It is a schematic diagram of the circular right shift processing process shown in the embodiments of the present application;

[0055] Figure 9 It is a schematic diagram of the structure of the data shift processing device based on a convolutional neural network shown in the embodiments of the present application;

[0056] Figure 10 It is another schematic diagram of the structure of the data shift processing device based on a convolutional neural network shown in the embodiments of the present application;

[0057] Figure 11 It is a structural block diagram of the artificial intelligence chip shown in the embodiments of the present application;

[0058] Figure 12 It is a structural block diagram of the board shown in the embodiments of the present application;

[0059] Figure 13 It is a schematic diagram of the structure of the computing device shown in the embodiments of the present application. Detailed implementation manners

[0060] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0061] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0062] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0063] In the related art, data shift processing tasks are generally implemented by an external ARM (Advanced RISC Machine) chip or by using DMA (Direct Memory Access), and cannot be directly run on a running processing device adapted to a neural network structure, which in turn leads to problems such as low computing efficiency, high latency, and low utilization rate of hardware resources.

[0064] To address the above problems, an embodiment of this application provides a data shift processing method based on a convolutional neural network, which can implement the execution of data shift tasks on a running processing device adapted to a neural network structure, improve computing efficiency, reduce latency, and improve the utilization rate of hardware resources.

[0065] The data processing method based on a convolutional neural network in this application can be applied to a processor, which can be a general-purpose processor, such as a CPU (Central Processing Unit), or an artificial intelligence processor for performing artificial intelligence operations. Artificial intelligence operations may include machine learning operations, brain-like operations, etc. Among them, machine learning operations include neural network operations, k-means operations, support vector machine operations, etc. The artificial intelligence processor includes, for example, one or a combination of an artificial intelligence chip processor, a GPU (Graphics Processing Unit), an NPU (Neural-Network Processing Unit), a DSP (Digital Signal Process), and a field-programmable gate array (FPGA) chip.

[0066] The artificial intelligence processor may be a processor applied to an artificial intelligence chip. The artificial intelligence chip may be, for example, a neural network chip or other chips. The neural network chip may be, for example, a convolutional neural network inference chip, an ASIC chip, etc. The present application does not limit the specific type of the processor.

[0067] In a possible implementation, the processor mentioned in the present application may include multiple processing units. Each processing unit can independently run various tasks assigned to it, such as: convolution operation tasks, pooling tasks, or fully connected tasks, etc. The present application does not limit the processing units and the tasks run by the processing units. The multiple processing units in the processor can share part of the storage space. For example, they can share part of the RAM storage space and the register bank, and can also have their own storage spaces simultaneously.

[0068] The technical solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0069] Figure 1 It is a schematic flowchart of a data shift processing method based on a convolutional neural network shown in the embodiments of the present application. This method can be applied to a processor, which may include a general-purpose processor, an artificial intelligence processor, etc. The artificial intelligence processor may be an artificial intelligence chip processor or a GPU, etc.

[0070] See Figure 1 , this method includes:

[0071] S101. Obtain preprocessed data from the input data through preprocessing.

[0072] In this step, the input serial data stream can be converted into a matrix data block through a splicing operation. That is to say, the input data may be the input serial data stream, and the preprocessed data may be the matrix data block.

[0073] Among them, the serial data stream (Serial) is a data stream composed of "0" and "1", such as 110010101, 101001110, etc.

[0074] Among them, the byte length of the serial data stream can be the square of any positive integer. That is to say, the byte length of the serial data stream can be a serial data stream with a byte length of 4 bits, 9 bits, 16 bits, etc. For example, a serial data stream with a 9-bit byte length is received, such as 110010101.

[0075] In this step, the input serial data stream can be converted in form into a matrix data block (Matrix) in matrix data form. For example, a serial data stream with a 9-bit byte length can be converted into a matrix data block in the form of a 3×3 matrix.

[0076] S102. According to different shifting operation requirements, process the preprocessed data through multiple different convolutional calculation layers of a shifting unit in a convolutional neural network to obtain shifted operation data.

[0077] This step may include:

[0078] After performing convolutional processing on the preprocessed data in the first convolutional calculation layer of the logical shifting unit to output two-channel convolutional data, perform convolutional processing through three different intermediate convolutional calculation layers, and then perform convolutional data channel merging processing through the fifth convolutional calculation layer to obtain the shifted operation data of logical shifting; or,

[0079] After performing convolutional processing on the preprocessed data in the first convolutional calculation layer of the cyclic shifting unit to output three-channel convolutional data, perform convolutional processing through four different intermediate convolutional calculation layers, and then perform convolutional data channel merging processing through the sixth convolutional calculation layer to obtain the shifted operation data of cyclic shifting.

[0080] Among them, after performing convolutional processing on the preprocessed data in the first convolutional calculation layer of the logical shifting unit to output two-channel convolutional data, performing convolutional processing through three different intermediate convolutional calculation layers set according to the logical left shift operation rule or the logical right shift operation rule, and then performing convolutional data channel merging processing through the fifth convolutional calculation layer to obtain the shifted operation data of logical shifting, including: after performing convolutional processing on the preprocessed data in the first convolutional calculation layer of the logical shifting unit to output two-channel convolutional data, performing convolutional processing through three different intermediate convolutional calculation layers set according to the logical left shift operation rule or the logical right shift operation rule, and then performing convolutional data channel merging processing through the fifth convolutional calculation layer to respectively obtain the shifted operation data of logical left shift or logical right shift.

[0081] Among them, after performing convolutional processing on the preprocessed data in the first convolutional calculation layer of the cyclic shifting unit to output three-channel convolutional data, performing convolutional processing through four different intermediate convolutional calculation layers set according to the cyclic left shift operation rule or the cyclic right shift operation rule, and then performing convolutional data channel merging processing through the sixth convolutional calculation layer to obtain the shifted operation data of cyclic shifting, including: after performing convolutional processing on the preprocessed data in the first convolutional calculation layer of the cyclic shifting unit to output three-channel convolutional data, performing convolutional processing through four different intermediate convolutional calculation layers set according to the cyclic left shift operation rule or the cyclic right shift operation rule, and then performing convolutional data channel merging processing through the sixth convolutional calculation layer to respectively obtain the shifted operation data of cyclic left shift or cyclic right shift.

[0082] S103. Post-process the shifted operation data to obtain output data.

[0083] In this step, the shifted operation data in the matrix data block format can be converted into a serial data stream.

[0084] As can be seen from this example, in the method provided by this application, which is applied to a processor, after the processor preprocesses the input data to obtain preprocessed data, according to different shift operation requirements, the preprocessed data can be processed through multiple different convolutional calculation layers of a shift unit in a convolutional neural network to obtain shifted operation data; and then the shifted operation data is post-processed to obtain output data. In this way, multiple different convolutional calculation layers can be used to construct a shift operation, enabling the execution of data shift tasks on a running processing device adapted to the neural network structure, thereby improving the utilization rate of hardware resources, enhancing the operation efficiency, and reducing the latency.

[0085] Figure 2 It is another schematic flowchart of the data shift processing method based on a convolutional neural network according to an embodiment of this application. Figure 2 The processing method of logical shift is described therein.

[0086] It should be noted that logical shift includes logical left shift and logical right shift. Logical left shift means that a serial data stream with a set byte length is shifted left by one bit, and zeros are filled at the right end of the data stream. For example, after the serial data stream 101001110 is logically shifted left, it becomes 010011100. Logical right shift means that a serial data stream with a set byte length is shifted right by one bit, and zeros are filled at the left end of the data stream. For example, after the serial data stream 101001110 is logically shifted left, it becomes 010100111.

[0087] See Figure 2 , this method includes:

[0088] S201. Receive an input serial data stream.

[0089] Among them, a serial data stream (Serial) is a data stream composed of "0"s and "1"s, such as 110010101, 101001110, etc. The number of bits of the byte length of the serial data stream can be the square of any positive integer. That is to say, the number of bits of the byte length of the serial data stream can be a serial data stream with a byte length of 4 bits, 9 bits, 16 bits, etc., and this application does not make any limitations in this regard. In this step, the processor can receive a serial data stream with a set byte length, for example, receive a serial data stream with a 9-bit byte length, such as 110010101.

[0090] S202. Perform preprocessing to convert the input serial data stream into a matrix data block.

[0091] In this step, the processor can perform a format conversion on the input serial data stream to convert it into a matrix data block in matrix data form. For example, a serial data stream with a 9-bit byte length can be converted into a matrix data block in the form of a 3×3 matrix.

[0092] Such as Figure 4As shown, in a 9-bit byte-length serial data stream, when sorted from right to left, they are the first bit, the second bit, up to the ninth bit. First, take the first three bytes and place them in the top row of a 3×3 matrix in the order from left to right; then take the middle three bytes and place them in the middle row of the 3×3 matrix in the order from left to right; finally, take the last three bytes and place them in the bottom row of the 3×3 matrix in the order from left to right. It can be understood that sorting can also be performed in other orders to obtain a matrix data block in the form of a 3×3 matrix.

[0093] For another example Figure 5 As shown, taking the serial data stream 101001110 as an example, the splicing conversion method can be to place the 9-bit byte-length serial data stream into a matrix in the 3×3 form to obtain a matrix data block in the form of a 3×3 matrix. This process is as Figure 5 identified by the "S / M" conversion process:

[0094] Convert 101001110 into

[0095] In this way, the converted matrix data block can be adapted to the convolution operation in the convolutional neural network.

[0096] The above operation of converting the serial data stream into a matrix data block can be called a data preprocessing operation, which can be represented by "S / M".

[0097] S203. Process the matrix data block obtained by preprocessing through multiple different convolutional calculation layers in the convolutional neural network according to the logical shift requirements to obtain the shift operation data of logical left shift or logical right shift.

[0098] In this step, in the convolutional neural network, the processor can perform convolutional processing on the preprocessing data in the first convolutional calculation layer to output two-channel convolutional data, then perform convolutional processing through three different intermediate convolutional calculation layers, and then perform convolutional data channel merging processing through the fifth convolutional calculation layer to obtain the shift operation data of logical shift.

[0099] Among them, the logical shift requirements can include: "logical left shift" and "logical right shift" operation requirements. Different logical shift units can be "logical left shift unit", "logical right shift unit", etc. These logical shift units can also be called logical shift circuits.

[0100] Among them, in the convolutional neural network, the convolutional calculation layer in the convolutional neural network can be constructed, and the convolutional parameters such as the convolutional kernel, padding, stride, and bias during convolutional processing can be preset in advance, so as to adapt to the corresponding logical shift requirements and realize the logical shift processing of the matrix data block.

[0101] Among them, the preprocessed data can be a single-channel matrix data block. In this embodiment, the single-channel matrix data block is taken as an example but not limited thereto, and the single-channel matrix data block can be represented by A.

[0102] Among them, the three different intermediate convolution calculation layers can include the second convolution calculation layer, the third convolution calculation layer, and the fourth convolution calculation layer.

[0103] Among them, the first convolution calculation layer can be understood as being used to implement the "copy" operation on the data, the second convolution calculation layer, the third convolution calculation layer, and the fourth convolution calculation layer can be understood as being used to implement the "move" operation on the data, and the fifth convolution calculation layer can be understood as being used to implement the "rearrangement" operation on the data.

[0104] Among them, the first convolution calculation layer is provided with two convolution kernels, both of which are single-channel 3×3 (i.e., kernel = 3x3) convolution kernels, the stride is set to 1 (i.e., stride = 1), the padding is 1 (i.e., padding = 1), and the bias value is set to 0 (i.e., Bias = 0).

[0105] The second convolution calculation layer, the third convolution calculation layer, and the fourth convolution calculation layer are all provided with two convolution kernels, and both of the two convolution kernels are two-channel 3×3 (i.e., kernel = 3x3) convolution kernels; the strides of the second convolution calculation layer, the third convolution calculation layer, and the fourth convolution calculation layer are all set to 1 (i.e., stride = 1), the padding is all 1 (i.e., padding = 1), and the bias values are all set to 0 (i.e., Bias = 0).

[0106] The fifth convolution calculation layer is provided with one convolution kernel, and this convolution kernel is a two-channel 3×3 (i.e., kernel = 3x3) convolution kernel, the stride is set to 1 (i.e., stride = 1), the padding is 1 (i.e., padding = 1), and the bias value is set to 0 (i.e., Bias = 0).

[0107] The processing procedures for performing "logical left shift" and "logical right shift" are introduced separately below.

[0108] (1) Logical left shift operation

[0109] After the processor performs convolution processing on the preprocessed data in the first convolution calculation layer to output two-channel convolution data, it performs convolution processing through three different intermediate convolution calculation layers according to the logical left shift operation rule, and then performs convolution data channel merging processing through the fifth convolution calculation layer to obtain the shift operation data of the logical left shift.

[0110] That is to say, in order to implement the "logical left shift" operation of data, after the data is preprocessed, it is processed by different convolution kernels through the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer respectively, and then through the post-processing operation, the output result after the "logical left shift" operation can be obtained.

[0111] Please also refer to Figure 5 , in one embodiment, the processing processes of the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer include:

[0112] 1), in the first convolutional calculation layer, the input single-channel matrix data block is subjected to convolutional processing according to the convolutional parameters.

[0113] In this convolutional processing, the two convolutional kernels are the same, and the channel parameters of each convolutional kernel are

[0114] The input single-channel matrix data block is: matrix data block

[0115] Among them, matrix data block A is obtained by converting the serial input data stream 101001110.

[0116] After matrix data block A undergoes padding processing with padding = 1, it is

[0117] After the two convolutional kernels in the first convolutional calculation layer perform convolutional processing according to the convolutional principle, the output two-channel matrix data is obtained And

[0118] It can be found that in this step, the "duplication" of the input single-channel matrix data block is realized.

[0119] 2), in the second convolutional calculation layer, the data output by the first convolutional calculation layer is subjected to convolutional processing according to the convolutional parameters.

[0120] In this convolutional processing, the channel parameter of one of the convolutional kernels is And The channel parameter of the other convolutional kernel is And

[0121] For example, the two-channel matrix data And After undergoing padding processing with padding = 1, they are respectively subjected to convolutional processing through the two convolutional kernels in the second convolutional calculation layer, and respectively obtain:

[0122] and

[0123] After fitting the matrix data, the output two-channel matrix data is obtained:

[0124] and

[0125] It can be found that in this step, the upward "shift" of the one-channel matrix data is realized to obtain the leftward "shift" of the other-channel matrix data to obtain

[0126] 3), In the third convolution calculation layer, the data output by the second convolution calculation layer is convolved according to the convolution parameters.

[0127] In this convolution process, the channel parameter of one of the convolution kernels is and the channel parameter of the other convolution kernel is and

[0128] For example, the two-channel matrix data and After the padding process with padding = 1, they are respectively convolved through two convolution kernels in the third convolution calculation layer, and the following are obtained respectively:

[0129] and

[0130] After fitting the matrix data, the output two-channel matrix data is obtained:

[0131] and

[0132] It can be found that in this step, the rightward "shift" of the one-channel matrix data is realized to obtain the "copy" of the other-channel matrix data to obtain

[0133] 4), In the fourth convolution calculation layer, the data output by the third convolution calculation layer is convolved according to the convolution parameters.

[0134] In this convolution process, the channel parameter of one of the convolution kernels is and the channel parameter of the other convolution kernel is and

[0135] For example, the dual-channel matrix data and after being padded with padding = 1 are respectively convolved through two convolutional kernels in the fourth convolutional calculation layer, and the following are obtained respectively:

[0136] and

[0137] After fitting the matrix data, the output two-channel matrix data is obtained:

[0138] and

[0139] It can be found that in this step, a one-channel matrix data is "shifted" to the right to obtain and another one-channel matrix data is "copied" to obtain

[0140] 5), In the fifth convolutional calculation layer, the data output by the fourth convolutional calculation layer is convolved according to the convolutional parameters.

[0141] In this convolutional process, the channel parameters of the convolutional kernel in the fifth convolutional calculation layer are:

[0142] and

[0143] For example, the dual-channel matrix data and after being padded with padding = 1 are convolved through the convolutional kernel in the fifth convolutional calculation layer,

[0144] and the following is obtained:

[0145] It can be found that in this step, by merging the two-channel matrix data, the effect of "rearranging" the data is achieved.

[0146] It can be seen that during the "logical left shift" operation, the convolutional operation characteristics of the convolutional neural network are utilized. By using different convolutional kernels of the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer to perform different operations on the input data, the logical left shift operation of the input data is achieved, and the matrix data block of the logical left shift operation result is obtained.

[0147] (2) Logical right shift operation

[0148] After the preprocessed data is subjected to convolution processing in the first convolutional calculation layer to output two-channel convolutional data, it is then subjected to convolution processing through three different intermediate convolutional calculation layers for the logical right shift operation rule, and then the convolutional data channel merging processing is performed through the fifth convolutional calculation layer to obtain the shift operation data of the logical right shift.

[0149] That is to say, in order to implement the "logical right shift" operation of data, after the data is preprocessed, it passes through the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer respectively, and is processed with different convolutional kernels, and then through the post-processing operation, the output result after the "logical right shift" operation can be obtained.

[0150] Please also refer to Figure 6 , in one embodiment, the processing processes of the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer include:

[0151] 1), In the first convolutional calculation layer, the input single-channel matrix data block is subjected to convolution processing according to the convolution parameters.

[0152] In this convolution processing, the two convolutional kernels are the same, and the channel parameters of each convolutional kernel are

[0153] The input single-channel matrix data block is: matrix data block

[0154] Among them, the matrix data block A is obtained by converting the serial input data stream 101001110.

[0155] After the matrix data block A undergoes padding processing with padding = 1, it is

[0156] After the two convolutional kernels in the first convolutional calculation layer perform convolution processing according to the convolution principle, the output two-channel matrix data is obtained And

[0157] It can be found that in this step, the "duplication" of the input single-channel matrix data block is realized.

[0158] 2), In the second convolutional calculation layer, the data output by the first convolutional calculation layer is subjected to convolution processing according to the convolution parameters.

[0159] In this convolution processing, the channel parameter of one of the convolutional kernels is And The channel parameter of the other convolutional kernel is And

[0160] For example, after padding = 1 is performed on the dual-channel matrix data and they are respectively convolved through two convolutional kernels in the second convolutional calculation layer, and the following are obtained respectively:

[0161] and

[0162] After fitting the matrix data, the output dual-channel matrix data is obtained:

[0163] and

[0164] It can be found that in this step, a "shift down" of the single-channel matrix data is achieved to obtain and a "shift right" of the other single-channel matrix data is achieved to obtain

[0165] 3) In the third convolutional calculation layer, the data output from the second convolutional calculation layer is convolved according to the convolutional parameters.

[0166] In this convolutional process, the channel parameter of one convolutional kernel is and the channel parameter of the other convolutional kernel is and

[0167] For example, after padding = 1 is performed on the dual-channel matrix data and they are respectively convolved through two convolutional kernels in the third convolutional calculation layer, and the following are obtained respectively:

[0168] and

[0169] After fitting the matrix data, the output dual-channel matrix data is obtained:

[0170] and

[0171] It can be found that in this step, a "shift left" of the single-channel matrix data is achieved to obtain and a "duplication" of the other single-channel matrix data is achieved to obtain

[0172] 4), In the fourth convolutional calculation layer, the data output by the third convolutional calculation layer is subjected to convolutional processing according to convolutional parameters.

[0173] In this convolutional processing, the channel parameter of one convolutional kernel is and the channel parameter of another convolutional kernel is and

[0174] For example, the two-channel matrix data and After the padding process with padding = 1, they are respectively subjected to convolutional processing through two convolutional kernels in the fourth convolutional calculation layer, and the following are obtained respectively:

[0175] and

[0176] After fitting the matrix data, the output two-channel matrix data is obtained:

[0177] and

[0178] It can be found that in this step, the "shift to the left" of the one-channel matrix data is realized to obtain and the "copy" of the other-channel matrix data is realized to obtain

[0179] 5), In the fifth convolutional calculation layer, the data output by the fourth convolutional calculation layer is subjected to convolutional processing according to convolutional parameters.

[0180] In this convolutional processing, the channel parameters of the convolutional kernels in the fifth convolutional calculation layer are:

[0181] and

[0182] For example, the two-channel matrix data and After the padding process with padding = 1, they are subjected to convolutional processing through the convolutional kernels in the fifth convolutional calculation layer,

[0183] and the following is obtained:

[0184] It can be found that in this step, by merging the two-channel matrix data, the effect of "rearrangement" of the data is realized.

[0185] It can be seen that during the "logical right shift" operation, the convolution operation characteristics of the convolutional neural network are utilized. By using different convolutional kernels of the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer to perform different operations on the input data, the logical right shift operation of the input data is achieved, and the matrix data block of the logical right shift operation result is obtained.

[0186] In summary, it can be found that the difference between the processing of "logical left shift" and "logical right shift" lies in the parameter settings of the convolutional kernels of the three intermediate convolutional layers (i.e., the second convolutional calculation layer, the third convolutional calculation layer, and the fourth convolutional calculation layer) in the convolutional neural network. By setting different parameters for the convolutional kernels of the three intermediate convolutional layers, different data "transfer" operations are achieved to correspondingly implement the processing of data "logical left shift" and "logical right shift".

[0187] S204. Perform post-processing to convert the shifted operation data in the form of a matrix data block into a serial data stream and output it.

[0188] In this step, the form conversion of the matrix data block can be achieved, and it is converted into a serial data stream. For example, a matrix data block in the form of a 3×3 matrix can be converted into a serial data stream with a 9-bit byte length. This operation process can be called a data post-processing operation, which is the inverse operation of the data pre-processing operation (i.e., converting the serial data stream into a matrix data block), and can be represented by "M / S". For example Figure 5 as shown:

[0189] can be converted to 010011100;

[0190] For example Figure 6 as shown:

[0191] can be converted to 010100111.

[0192] As can be seen from this embodiment, the method provided by the embodiments of the present application can, for different logical shift requirements, such as logical left shift and logical right shift, construct different types of convolutional neural networks according to different logical shift rules and in combination with the operating principle of the convolutional neural network to implement the processing of the above-mentioned "logical left shift" and "logical right shift". Thus, the logical shift task of data can be executed on an operating processing device adapted to the neural network structure (for example, a neural network dedicated chip, a general-purpose processor, an image processor, etc.), and the excellent computing acceleration function of the neural network dedicated chip can be fully utilized, so that it is no longer necessary to externally connect an ARM chip or use a DMA chip to implement the logical shift. Therefore, the hardware computing resources can be fully utilized, the neural network inference task and the logical shift task can be executed simultaneously, the operation efficiency can be improved, and the latency can be reduced.

[0193] Figure 3 It is another schematic flowchart of the data shift processing method based on the convolutional neural network according to the embodiments of the present application. Figure 3 The processing method of circular shift is described in

[0194] It should be noted that circular shift includes circular left shift and circular right shift. Circular left shift means that the serial data stream with a set byte length is shifted left by one bit, and a character at the left end of the original serial data stream is moved to the right end. For example, after the serial data stream 101001110 is circularly left-shifted, it becomes 010011101. Circular right shift means that the serial data stream with a set byte length is shifted right by one bit, and a character at the right end of the original serial data stream is moved to the left end. For example, after the serial data stream 101001110 is circularly right-shifted, it becomes 010100111.

[0195] See Figure 3 , the method includes:

[0196] S301. Receive the input serial data stream.

[0197] This step can refer to the description in step S201 and will not be elaborated here.

[0198] S302. Perform preprocessing to convert the input serial data stream into a matrix data block.

[0199] This step can refer to the description in step S202 and will not be elaborated here.

[0200] S303. Process the matrix data block obtained by preprocessing through multiple different convolutional calculation layers in the convolutional neural network according to the circular shift requirement to obtain the shift operation data of circular left shift or circular right shift.

[0201] In this step, in the convolutional neural network, after the processor performs convolution processing on the preprocessed data in the first convolutional calculation layer to output three-channel convolution data, it performs convolution processing through four different intermediate convolutional calculation layers, and then performs convolution data channel merging processing through the sixth convolutional calculation layer to obtain the shifted operation data with cyclic shift, and obtains the shifted operation data with cyclic shift.

[0202] Among them, the cyclic shift requirements may include: "circular left shift" and "circular right shift" operation requirements. Different logical shift units may be "circular logical left shift unit", "circular logical right shift unit", etc., and these logical shift units may also be referred to as circular logical shift circuits.

[0203] Among them, in the convolutional neural network, by constructing the convolutional calculation layers in the convolutional neural network, the convolutional parameters such as the convolutional kernels, padding, stride, and bias during convolution processing in the convolutional calculation layers can be preset, so as to adapt to the corresponding cyclic shift requirements and realize the cyclic shift processing of matrix data blocks.

[0204] Among them, the preprocessed data may be a single-channel matrix data block. In this embodiment, a single-channel matrix data block is taken as an example but not limited thereto, and the single-channel matrix data block can be represented by A.

[0205] Among them, the four different intermediate convolutional calculation layers may include the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer.

[0206] Among them, the first convolutional calculation layer can be understood as being used to implement the "copy" operation on the data, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer can be understood as being used to implement the "move" operation on the data, and the sixth convolutional calculation layer can be understood as being used to implement the "rearrangement" operation on the data.

[0207] Among them, the first convolutional calculation layer is provided with three convolutional kernels, and all three convolutional kernels are single-channel 3×3 (i.e., kernel = 3x3) convolutional kernels, the stride is set to 1 (i.e., stride = 1), the padding is 1 (i.e., padding = 1), and the bias value is set to 0 (i.e., Bias = 0).

[0208] The second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer are all provided with three convolutional kernels, and all three convolutional kernels are three-channel 3×3 (i.e., kernel = 3x3) convolutional kernels; the strides of the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer are all set to 1 (i.e., stride = 1), the padding is all 1 (i.e., padding = 1), and the bias values are all set to 0 (i.e., Bias = 0).

[0209] The sixth convolutional calculation layer is provided with a convolutional kernel, which is a three-channel 3×3 (i.e., kernel = 3x3) convolutional kernel, the stride is set to 1 (i.e., stride = 1), the padding is 1 (i.e., padding = 1), and the bias value is set to 0 (i.e., Bias = 0).

[0210] The processing procedures for performing "circular left shift" and "circular right shift" are introduced separately below.

[0211] (3) Circular left shift operation

[0212] After the processor performs convolutional processing on the preprocessed data in the first convolutional calculation layer to output three-channel convolutional data, it performs convolutional processing through four different intermediate convolutional calculation layers according to the circular left shift operation rule, and then performs convolutional data channel merging processing through the sixth convolutional calculation layer to obtain the shifted operation data of the circular left shift.

[0213] That is to say, in order to implement the "circular left shift" operation of data, after the data is preprocessed, it passes through the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, the fifth convolutional calculation layer, and the sixth convolutional calculation layer respectively, and is processed by different convolutional kernels, and then through the post-processing operation, the output result after the "circular left shift" operation can be obtained.

[0214] Please refer to Figure 7 In an embodiment, the processing procedures of the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, the fifth convolutional calculation layer, and the sixth convolutional calculation layer include:

[0215] 1) In the first convolutional calculation layer, the input single-channel matrix data block is subjected to convolutional processing according to the convolutional parameters.

[0216] In this convolutional processing, the three convolutional kernels are the same, and the channel parameters of each convolutional kernel are all

[0217] The input single-channel matrix data block is: matrix data block

[0218] where the matrix data block A is obtained by converting the serial input data stream 101001110.

[0219] After the matrix data block A undergoes padding processing with padding = 1, it is

[0220] After passing through the three convolutional kernels in the first convolutional calculation layer for convolutional processing according to the convolutional principle, the output three-channel matrix data is obtained and

[0221] It can be found that in this step, the "copy" of the input single-channel matrix data block is realized.

[0222] 2), In the second convolution calculation layer, the data output by the first convolution calculation layer is convolved according to the convolution parameters.

[0223] In this convolution process,

[0224] The channel parameters of the first convolution kernel are and

[0225] The channel parameters of the second convolution kernel are and

[0226] The channel parameters of the third convolution kernel are and

[0227] For example, the three-channel matrix data and After the padding process with padding = 1, they are respectively convolved through the three convolution kernels in the second convolution calculation layer to obtain:

[0228]

[0229] and

[0230] After fitting the matrix data, the output three-channel matrix data is obtained:

[0231] and

[0232] It can be found that in this step, the rightward "shift" of the first-channel matrix data is obtained to get The rightward "shift" of the second-channel matrix data is obtained to get The leftward "shift" of the third-channel matrix data is obtained to get

[0233] 3), In the third convolution calculation layer, the data output by the second convolution calculation layer is convolved according to the convolution parameters.

[0234] In this convolution process,

[0235] The channel parameters of the first convolution kernel are and

[0236] The channel parameters of the second convolution kernel are and

[0237] The channel parameters of the third convolution kernel are and

[0238] For example, the three-channel matrix data and

[0239] After padding processing with padding = 1, they are respectively convolved through the three convolution kernels in the third convolution calculation layer to obtain:

[0240]

[0241] and

[0242] After fitting the matrix data, the output three-channel matrix data is obtained:

[0243] and

[0244] It can be found that in this step, the "shift to the right" of the first-channel matrix data is performed to obtain the "shift to the right" of the second-channel matrix data is performed to obtain the "copy" of the third-channel matrix data is performed to obtain

[0245] 4) In the fourth convolution calculation layer, the data output from the third convolution calculation layer is convolved according to the convolution parameters.

[0246] In this convolution process,

[0247] the channel parameters of the first convolution kernel are and

[0248] the channel parameters of the second convolution kernel are and

[0249] the channel parameters of the third convolution kernel are and

[0250] For example, the three-channel matrix data and After padding processing with padding = 1, they are respectively convolved through the three convolution kernels in the fourth convolution calculation layer to obtain:

[0251]

[0252] and

[0253] After fitting the matrix data, the output three-channel matrix data is obtained:

[0254] and

[0255] It can be found that in this step, the downward "shift" of the first-channel matrix data is realized to obtain the upward "shift" of the second-channel matrix data to obtain the "copy" of the third-channel matrix data to obtain

[0256] 5), In the fifth convolutional calculation layer, the data output by the fourth convolutional calculation layer is convolved according to the convolutional parameters.

[0257] In this convolution process,

[0258] the channel parameters of the first convolutional kernel are and

[0259] the channel parameters of the second convolutional kernel are and

[0260] the channel parameters of the third convolutional kernel are and

[0261] For example, the three-channel matrix data and After the padding process with padding = 1, they are respectively convolved through the three convolutional kernels in the fifth convolutional calculation layer to obtain:

[0262]

[0263] and

[0264] After fitting the matrix data, the output three-channel matrix data is obtained:

[0265] and

[0266] It can be found that in this step, the downward "shift" of the first-channel matrix data is realized is obtained by "shifting down" For the second-channel matrix data is obtained by "copying" For the third-channel matrix data is obtained by "copying"

[0267] 6) In the sixth convolutional calculation layer, the data output by the fifth convolutional calculation layer is subjected to convolutional processing according to convolutional parameters.

[0268] In this convolutional processing, the channel parameters of the convolutional kernel in the sixth convolutional calculation layer are:

[0269] and

[0270] For example, the three-channel matrix data and After being subjected to padding processing with padding = 1, it is subjected to convolutional processing through the convolutional kernel in the sixth convolutional calculation layer,

[0271] to obtain:

[0272] It can be found that in this step, by merging the three-channel matrix data, the effect of "rearranging" the data is achieved.

[0273] It can be seen that during the "circular left shift" operation, the convolutional operation characteristics of the convolutional neural network are utilized. By using different convolutional kernels in the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, the fifth convolutional calculation layer, and the sixth convolutional calculation layer to perform different operations on the input data, the circular left shift operation of the input data is achieved, and the matrix data block of the circular left shift operation result is obtained.

[0274] (4) Circular right shift operation

[0275] After the processor performs convolutional processing on the preprocessed data in the first convolutional calculation layer and outputs three-channel convolutional data, it performs convolutional processing through four different intermediate convolutional calculation layers according to the circular right shift operation rule, and then performs convolutional data channel merging processing through the sixth convolutional calculation layer to obtain the shifted operation data of the circular right shift.

[0276] That is to say, in order to implement the "circular right shift" operation of the data, after the data is preprocessed, it is respectively processed by different convolutional kernels in the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, the fifth convolutional calculation layer, and the sixth convolutional calculation layer, and then through the post-processing operation, the output result after the "circular right shift" operation can be obtained.

[0277] Please also refer to Figure 8 , in one embodiment, the processing processes of the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, the fifth convolutional calculation layer and the sixth convolutional calculation layer include:

[0278] 1), in the first convolutional calculation layer, the input single-channel matrix data block is subjected to convolutional processing according to convolutional parameters.

[0279] In this convolutional processing, the three convolutional kernels are the same, and the channel parameters of each convolutional kernel are all

[0280] The input single-channel matrix data block is: matrix data block wherein, the matrix data block A is obtained by converting the serial input data stream 101001110.

[0281] After the matrix data block A undergoes padding processing with padding = 1, it is

[0282] After the three convolutional kernels in the first convolutional calculation layer perform convolutional processing according to the convolutional principle, the output three-channel matrix data is obtained and

[0283] It can be found that in this step, the "duplication" of the input single-channel matrix data block is realized.

[0284] 2), in the second convolutional calculation layer, the data output by the first convolutional calculation layer is subjected to convolutional processing according to convolutional parameters.

[0285] In this convolutional processing,

[0286] the channel parameters of the first convolutional kernel are and

[0287] the channel parameters of the second convolutional kernel are and

[0288] the channel parameters of the third convolutional kernel are and

[0289] For example, the three-channel matrix data and After undergoing padding processing with padding = 1 and passing through the three convolutional kernels in the second convolutional calculation layer for convolutional processing respectively, the following are obtained:

[0290]

[0291] and

[0292] After fitting the matrix data, the output three-channel matrix data is obtained:

[0293] and

[0294] It can be found that in this step, the "shift to the left" of the first-channel matrix data is obtained to the "shift to the left" of the second-channel matrix data is obtained to the "shift to the right" of the third-channel matrix data is obtained to

[0295] 3) In the third convolutional calculation layer, the data output by the second convolutional calculation layer is convolved according to the convolutional parameters.

[0296] In this convolution process,

[0297] the channel parameters of the first convolutional kernel are and

[0298] the channel parameters of the second convolutional kernel are and

[0299] the channel parameters of the third convolutional kernel are and

[0300] For example, the three-channel matrix data and after being padded with padding = 1 are respectively convolved through the three convolutional kernels in the third convolutional calculation layer to obtain:

[0301]

[0302] and

[0303] After fitting the matrix data, the output three-channel matrix data is obtained:

[0304] and

[0305] It can be found that in this step, the "shift to the left" of the first-channel matrix data is the "shift to the left" of the second-channel matrix data The "shift to the right" of For the third-channel matrix data The "copy" of

[0306] 4), In the fourth convolutional calculation layer, the data output by the third convolutional calculation layer is subjected to convolutional processing according to the convolutional parameters.

[0307] In this convolutional processing,

[0308] The channel parameters of the first convolutional kernel are and

[0309] The channel parameters of the second convolutional kernel are and

[0310] The channel parameters of the third convolutional kernel are and

[0311] For example, the three-channel matrix data and After being subjected to padding processing with padding = 1, they are respectively subjected to convolutional processing through three convolutional kernels in the fourth convolutional calculation layer, and the following are obtained:

[0312]

[0313] and

[0314] After fitting the matrix data, the output three-channel matrix data is obtained:

[0315] and

[0316] It can be found that in this step, the upward "shift" of the first-channel matrix data is obtained to The downward "shift" of the second-channel matrix data is obtained to The "copy" of the third-channel matrix data is obtained to

[0317] 5), In the fifth convolutional calculation layer, the data output by the fourth convolutional calculation layer is subjected to convolutional processing according to the convolutional parameters.

[0318] In this convolutional processing,

[0319] The channel parameters of the first convolutional kernel are and

[0320] The channel parameters of the second convolutional kernel are and

[0321] The channel parameters of the third convolutional kernel are and

[0322] For example, the three-channel matrix data and After the padding process with padding = 1, they are respectively convolved through the three convolutional kernels in the fifth convolutional calculation layer, and the following are obtained:

[0323]

[0324] and

[0325] After fitting the matrix data, the output three-channel matrix data is obtained:

[0326] and

[0327] It can be found that in this step, the upward "shift" of the first-channel matrix data is achieved to obtain the "copy" of the second-channel matrix data is achieved to obtain the "copy" of the third-channel matrix data is achieved to obtain

[0328] 6) In the sixth convolutional calculation layer, the data output by the fifth convolutional calculation layer is convolved according to the convolutional parameters.

[0329] In this convolutional process, the channel parameters of the convolutional kernel in the sixth convolutional calculation layer are:

[0330] and

[0331] For example, the three-channel matrix data and After the padding process with padding = 1, they are convolved through the convolutional kernel in the sixth convolutional calculation layer,

[0332] and the following are obtained:

[0333] It can be found that in this step, by merging the three-channel matrix data, the effect of "rearranging" the data is achieved.

[0334] It can be seen that during the "circular right shift" operation, the convolutional operation characteristics of the convolutional neural network are utilized. By using different convolutional kernels of the first convolutional calculation layer, the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, the fifth convolutional calculation layer, and the sixth convolutional calculation layer to perform different operations on the input data, the circular right shift operation of the input data is realized, and the matrix data block of the circular right shift operation result is obtained.

[0335] In summary, it can be found that the difference between the processing of "circular left shift" and "circular right shift" lies in the parameter settings of the convolutional kernels of the four intermediate convolutional layers (i.e., the second convolutional calculation layer, the third convolutional calculation layer, the fourth convolutional calculation layer, and the fifth convolutional calculation layer) in the convolutional neural network. By setting different parameters of the convolutional kernels of the four intermediate convolutional layers, different data "transfer" operations are realized to correspondingly implement the processing of data "circular left shift" and "circular right shift".

[0336] S304. Perform post-processing to convert the shifted operation data in the form of a matrix data block into a serial data stream and output it.

[0337] In this step, the form conversion of the matrix data block can be realized and converted into a serial data stream. For example, a matrix data block in the form of a 3×3 matrix can be converted into a serial data stream with a 9-bit byte length. This operation process can be called a data post-processing operation, which is the inverse operation of the data preprocessing operation (i.e., converting a serial data stream into a matrix data block) and can be represented by "M / S". For example Figure 7 as shown:

[0338] It can be converted into 010011101, that is, the serial input data stream 101001110 is circularly right-shifted to obtain 010011101;

[0339] For example Figure 8 as shown: It can be converted into 010100111, that is, the serial input data stream 101001110 is circularly right-shifted to obtain 010100111.

[0340] As can be seen from this embodiment, for the method provided by the embodiments of the present application, for different cyclic shift requirements, such as cyclic left shift and cyclic right shift, different types of convolutional neural networks can be constructed according to different cyclic shift rules in combination with the operating principle of the convolutional neural network to implement the processing of the above "cyclic left shift" and "cyclic right shift". Thus, the cyclic shift task of data can be executed on a running processing device adapted to the neural network structure (for example, a neural network dedicated chip, a general-purpose processor, an image processor, etc.), which can give full play to the excellent computing acceleration function of the neural network dedicated chip, so that there is no need to externally connect an ARM chip or use a DMA chip to implement cyclic shift. Therefore, the hardware computing resources can be fully utilized, and the neural network inference task and the cyclic shift task can be executed simultaneously, which can improve the operation efficiency and reduce the latency.

[0341] Among them, different types of convolutional neural networks can be implemented by constructing different types of convolutional calculation layers, and different types of convolutional calculation layers can be implemented by constructing different types of convolutional kernels. In the related art, parameter customization of convolutional kernels can be implemented on training platforms such as Caffe, TensorFlow, and PyTorch, so as to construct different types of convolutional calculation layers, and further implement the construction of different types of convolutional neural networks to adapt to the processing requirements of different data shift operations.

[0342] It should be noted that in the related art, the shift of data is implemented by externally connecting an ARM (Advanced RISC Machine) chip or by using DMA (Direct Memory Access). However, the methods of externally connecting an ARM and using DMA will lead to low operation efficiency and certain latency due to the transmission of data I / O (input / output) in the memory (Memory). The method provided by the embodiments of the present application uses multiple different convolutional calculation layers to construct a shift operation calculation and implements the data shift task on a running processing device adapted to the neural network structure. Therefore, the shift processing of data can be executed on a neural network dedicated chip, a general-purpose processor, an image processor, etc., without externally connecting an ARM and using DMA for data shift processing, thereby improving the utilization rate of hardware resources, improving the operation efficiency, and reducing the latency.

[0343] It can be understood that the above content of the present application is an example of performing "logical left shift", "logical right shift", "cyclic left shift", and "cyclic right shift" operations on the input data of a 9-bit byte-length serial data stream, but it is not limited thereto. On this basis, the size (length and width) and parameters of the convolutional kernels in each convolutional calculation layer can be modified to adapt to the shift operations of receiving and processing serial data streams of different byte lengths.

[0344] Furthermore, on this basis, the channel depth of the convolution kernels in each convolution calculation layer can be increased to expand the dimension of each convolution calculation layer, so that multiple input data can be received and processed simultaneously, realizing an expansion of the capacity of the processable data (i.e., expanding the Batch Size) and improving the shifting processing efficiency of the input data.

[0345] Furthermore, for a serial data stream with a set byte length as the input data, and when the number of bits of its set byte length is not the square of any positive integer (i.e., the length is not a square number), the serial data stream can be split and then shifted separately; after the shifting process is completed, the overlapping parts can be selectively retained and spliced according to the principle of "prioritizing the rear data segment in the shifting direction" to complete its shifting operation process. For example, the splicing and merging can be completed according to the principle of "fully retaining the rear data segment in the shifting direction and discarding the last byte of the front data segment", and its shifting operation process can be completed. For example, for a logical left shift process of a serial data stream with a 7-bit byte length, the 7-bit byte length serial data stream can be split into two serial data streams A and B each composed of 4-bit byte lengths. The A serial data stream takes the first four byte data, and the B serial data stream takes the last four byte data. After Figure 5 performing the logical left shift process on A and B respectively in the embodiment, A1 and B1 are obtained correspondingly. According to the principle of "fully retaining the rear data segment in the shifting direction and discarding the last byte of the front data segment", so the B1 data is retained. For the "front data segment in the shifting direction" A1, the first 3 bytes are retained, and after splicing and merging, a serial data stream with a 7-bit byte length is obtained, realizing the logical left shift process of the original input serial data stream with a 7-bit byte length.

[0346] It should also be noted that for more generalized situations, such as non-fixed-length data segments, the convolution operations of left (right), up (down) single shifting of data in the solution of this application can be referred to, and the logical (circular) left (right) shifting operations of non-constant data can be adjusted by increasing the number of layers of left (right), up (down) operations. Such generalized operations can also be used for operations such as shifting multiple times.

[0347] Corresponding to the foregoing method embodiments for implementing application functions, the present application also provides a data shifting processing device, a computing device, a chip, a board card based on a convolutional neural network, and corresponding embodiments.

[0348] Figure 9 It is a schematic structural diagram of a data shifting processing device based on a convolutional neural network shown in the embodiments of the present application.

[0349] See Figure 9, a data shift processing device 90 based on a convolutional neural network, which is applied to a processor. The data shift processing device 90 based on the convolutional neural network includes: a preprocessing module 91, a shift operation module 92, and a postprocessing module 93.

[0350] The preprocessing module 91 is used to obtain preprocessed data by preprocessing the input data. The preprocessing module 91 can convert the input serial data stream into a matrix data block through a splicing operation. That is to say, the input data can be the input serial data stream, and the preprocessed data can be the matrix data block. Among them, the serial data stream (Serial) is a data stream composed of "0" and "1", such as 110010101, 101001110, etc.

[0351] The shift operation module 92 is used to process the preprocessed data of the preprocessing module 91 through multiple different convolutional calculation layers in the convolutional neural network according to different shift operation requirements to obtain shift operation data. After the shift operation module 92 performs convolutional processing on the preprocessed data in the first convolutional calculation layer and outputs two-channel convolutional data, it performs convolutional processing through three different intermediate convolutional calculation layers, and then performs convolutional data channel merging processing through the fifth convolutional calculation layer to obtain the shift operation data of logical shift; or, after the shift operation module 92 performs convolutional processing on the preprocessed data in the first convolutional calculation layer and outputs three-channel convolutional data, it performs convolutional processing through four different intermediate convolutional calculation layers, and then performs convolutional data channel merging processing through the sixth convolutional calculation layer to obtain the shift operation data of cyclic shift.

[0352] The postprocessing module 93 is used to obtain output data by postprocessing the shift operation data obtained by the shift operation module 92. The postprocessing module 93 can convert the shift operation data in the matrix data block format into a serial data stream.

[0353] The device provided in this application is applied to a processor. After the processor obtains preprocessed data by preprocessing the input data, it can process the preprocessed data through multiple different convolutional calculation layers in the convolutional neural network according to different shift operation requirements to obtain shift operation data; then postprocess the shift operation data to obtain output data. In this way, multiple different convolutional calculation layers can be used to construct a shift operation, so as to perform a data shift task on a running processing device adapted to the neural network structure, thereby improving the utilization rate of hardware resources, improving the operation efficiency, and reducing the latency.

[0354] Figure 10 It is another structural schematic diagram of the data shift processing device based on the convolutional neural network shown in the embodiment of this application.

[0355] See Figure 10, a data shift processing device 90 based on a convolutional neural network, which is applied to a processor. The data shift processing device 90 based on the convolutional neural network includes: a preprocessing module 91, a shift operation module 92, and a postprocessing module 93. Among them, the shift operation module 92 includes: a logical shift unit 921 and a circular shift unit 922.

[0356] The logical shift unit 921 is used to perform convolution processing on the preprocessed data in the first convolutional calculation layer to output two-channel convolution data, then perform convolution processing through three different intermediate convolutional calculation layers, and then perform convolution data channel merging processing through the fifth convolutional calculation layer to obtain the shift operation data of the logical shift.

[0357] The logical shift unit 921 can perform convolution processing on the preprocessed data in the first convolutional calculation layer to output two-channel convolution data, then perform convolution processing through three different intermediate convolutional calculation layers set according to the logical left shift operation rule or the logical right shift operation rule, and then perform convolution data channel merging processing through the fifth convolutional calculation layer to obtain the shift operation data of the logical left shift or the logical right shift respectively. Among them, each of the three different intermediate convolutional calculation layers includes two different convolution kernels, and each convolution kernel includes two channel parameters.

[0358] The circular shift unit 922 is used to perform convolution processing on the preprocessed data in the first convolutional calculation layer to output three-channel convolution data, then perform convolution processing through four different intermediate convolutional calculation layers, and then perform convolution data channel merging processing through the sixth convolutional calculation layer to obtain the shift operation data of the circular shift. The circular shift unit 922 can perform convolution processing on the preprocessed data in the first convolutional calculation layer to output three-channel convolution data, then perform convolution processing through four different intermediate convolutional calculation layers set according to the circular left shift operation rule or the circular right shift operation rule, and then perform convolution data channel merging processing through the sixth convolutional calculation layer to obtain the shift operation data of the circular left shift or the circular right shift respectively. Among them, each of the four different intermediate convolutional calculation layers includes three different convolution kernels, and each convolution kernel includes three channel parameters.

[0359] Regarding the device in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0360] It should be understood that the above device embodiments are merely illustrative, and the devices of the present application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.

[0361] In addition, without special instructions, in each embodiment of the present application, each functional unit / module can be integrated in one unit / module, or each unit / module can exist physically alone, or two or more units / modules can be integrated together. The above integrated unit / module can be implemented in the form of hardware or in the form of a software program module.

[0362] When the integrated unit / module is implemented in the form of hardware, the hardware can be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Without special instructions, the processor can be any suitable hardware processor, such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), DSP (Digital Signal Processor), and ASIC (Application-Specific Integrated Circuit), etc. Without special instructions, the storage unit can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory RRAM (Resistive Random Access Memory), dynamic random access memory DRAM (Dynamic Random Access Memory), static random access memory SRAM (Static Random-Access Memory), enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory), high-bandwidth memory HBM (High-Bandwidth Memory), hybrid memory cube HMC (Hybrid Memory Cube), etc.

[0363] When the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this disclosure. The aforementioned memory includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), mobile hard disks, magnetic disks, or optical discs.

[0364] Figure 11 It is the structural block diagram of the artificial intelligence chip shown in the embodiments of this application.

[0365] See Figure 11 , this application also provides an artificial intelligence chip 120, which includes the above-mentioned data processing device 90 based on a convolutional neural network. The structure of the data processing device 90 based on a convolutional neural network can be seen in Figure 9 and Figure 10 for the description. The artificial intelligence chip 120 can be, for example, a neural network chip or other chips. The neural network chip can be, for example, a convolutional neural network inference chip, an ASIC chip, etc.

[0366] This application also provides a board card, which includes a storage device, an interface device, a control device, and the above-mentioned artificial intelligence chip; among them, the artificial intelligence chip is respectively connected to the storage device, the control device, and the interface device; the storage device is used to store data; the interface device is used to realize data transmission between the artificial intelligence chip and external devices; the control device is used to monitor the state of the artificial intelligence chip.

[0367] Figure 12 It is the structural block diagram of the board card shown in the embodiments of this application. Refer to Figure 12 , in addition to including the above-mentioned artificial intelligence chip 1289, the above-mentioned board card can also include other supporting components. The supporting components include, but are not limited to: a storage device 1290, an interface device 1291, and a control device 1292;

[0368] The storage device 1290 is connected to the artificial intelligence chip 1289 via a bus and is used to store data. The storage device 1290 may include multiple groups of storage units 1293. Each group of storage units 1293 is connected to the artificial intelligence chip 1289 via a bus. It can be understood that each group of storage units 1293 may be a DDR SDRAM (Double Data Rate SDRAM).

[0369] DDR can double the speed of SDRAM without increasing the clock frequency. DDR allows data to be read on both the rising and falling edges of the clock pulse. The speed of DDR is twice that of standard SDRAM. In one embodiment, the storage device may include 4 groups of storage units 1293. Each group of storage units 1293 may include multiple DDR4 dies (chips). In one embodiment, the artificial intelligence chip 1289 may internally include 4 72-bit DDR4 controllers. Among the above 72-bit DDR4 controllers, 64 bits are used for data transmission and 8 bits are used for ECC check. It can be understood that when DDR4-3200 dies are used in each group of storage units 1293, the theoretical bandwidth of data transmission can reach 25600MB / s.

[0370] In one embodiment, each group of storage units 1293 includes multiple double data rate synchronous dynamic random access memories arranged in parallel. DDR can transfer data twice within one clock cycle. A controller for controlling DDR is provided in the chip to control the data transmission and data storage of each storage unit 1293.

[0371] The interface device 1291 is electrically connected to the artificial intelligence chip 1289. The interface device 1291 is used to implement data transmission between the artificial intelligence chip 1289 and an external device (such as a server or a computer). For example, in one embodiment, the interface device 1291 may be a standard PCIE interface. For instance, the data to be processed is transferred from the server to the chip through the standard PCIE interface to achieve data transfer. Preferably, when using a PCIE 3.0X 16 interface for transmission, the theoretical bandwidth can reach 16000MB / s. In another embodiment, the interface device 1291 may also be other interfaces. The present application does not limit the specific form of the above other interfaces, as long as the interface unit can implement the transfer function. In addition, the calculation result of the artificial intelligence chip 1289 is still transmitted back to the external device (such as a server) by the interface device.

[0372] The control device 1292 is electrically connected to the artificial intelligence chip 1289. The control device 1292 is used to monitor the state of the artificial intelligence chip 1289. Specifically, the artificial intelligence chip 1289 and the control device 1292 can be electrically connected through an SPI interface. The control device 1292 can include a microcontroller unit (MCU). The artificial intelligence chip 1289 can include multiple processing chips, multiple processing cores, or multiple processing circuits, and can drive multiple loads. Therefore, the artificial intelligence chip 1289 can be in different working states such as multi-load and light-load. Through the control device 1292, the working states of multiple processing chips, multiple processors, or multiple processing circuits in the artificial intelligence chip can be regulated.

[0373] In a possible implementation manner, the present application further provides a computing device, which includes the above artificial intelligence chip. The computing device includes a data processing device, a robot, a computer, a printer, a scanner, a tablet computer, a smart terminal, a mobile phone, a driving recorder, a navigator, a sensor, a camera, a server, a cloud server, a camera, a video camera, a projector, a watch, a headset, a mobile storage device, a wearable device, a vehicle, a household appliance, and / or a medical device. The vehicle includes an airplane, a ship, and / or a vehicle; the household appliance includes a television, an air conditioner, a microwave oven, a refrigerator, a rice cooker, a humidifier, a washing machine, a light, a gas stove, a range hood; the medical device includes a nuclear magnetic resonance instrument, a B-ultrasound instrument, and / or an electrocardiogram instrument.

[0374] Figure 13 It is a schematic structural diagram of the computing device shown in the embodiment of the present application.

[0375] See Figure 13 , the computing device 1000 includes a memory 1010 and a processor 1020.

[0376] The processor 1020 can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0377] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage device can be a readable and writable storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 1010 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 1010 can include removable storage devices that are readable and / or writable, such as compact discs (CDs), read-only digital versatile discs (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray discs, ultra-density discs, flash memory cards (such as SD cards, min SD cards, Micro-SD cards, etc.), magnetic floppy disks, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or by wire.

[0378] Executable code is stored on the memory 1010, and when the executable code is processed by the processor 1020, it can cause the processor 1020 to execute some or all of the methods described above.

[0379] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the above steps of the method according to the present application.

[0380] Alternatively, the present application can also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium), on which executable code (or computer program or computer instruction code) is stored. When the executable code (or computer program or computer instruction code) is executed by a processor of an electronic device (or a server, etc.), it causes the processor to execute some or all of the steps of the method according to the present application.

[0381] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A data shift processing method based on a convolutional neural network, characterized in that Applied to a processor, which is a processor applied to a neural network chip. The processor includes a plurality of processing units, and each processing unit independently runs the assigned tasks, including convolution operation tasks. The method includes: Obtaining preprocessed data by preprocessing the input data; According to different shifting operation requirements, processing the preprocessed data through a plurality of different convolution calculation layers of a shifting unit in a convolutional neural network to obtain shifted operation data; Post-processing the shifted operation data to obtain output data; Wherein the shifting unit includes a logical shifting unit or a cyclic shifting unit; The logical shifting unit includes a first convolution calculation layer, three different intermediate convolution calculation layers, and a fifth convolution calculation layer; The cyclic shifting unit includes a first convolution calculation layer, four different intermediate convolution calculation layers, and a sixth convolution calculation layer; Wherein each of the three different intermediate convolution calculation layers includes two mutually different convolution kernels, and each convolution kernel includes two channel parameters; or, Each of the four different intermediate convolution calculation layers includes three mutually different convolution kernels, and each convolution kernel includes three channel parameters.

2. The method according to claim 1, wherein: The obtaining preprocessed data by preprocessing the input data includes: converting the input serial data stream into a matrix data block through a splicing operation; The post-processing the shifted operation data to obtain output data includes: converting the shifted operation data in the form of a matrix data block into a serial data stream.

3. The method according to claim 1, characterized in that The processing the preprocessed data through a plurality of different convolution calculation layers of a shifting unit in a convolutional neural network to obtain shifted operation data includes: After performing convolution processing on the preprocessed data in the first convolution calculation layer of the logical shifting unit to output two-channel convolution data, performing convolution processing through three different intermediate convolution calculation layers, and then performing convolution data channel merging processing through the fifth convolution calculation layer to obtain the shifted operation data of logical shifting; or, After performing convolution processing on the preprocessed data in the first convolution calculation layer of the cyclic shifting unit to output three-channel convolution data, performing convolution processing through four different intermediate convolution calculation layers, and then performing convolution data channel merging processing through the sixth convolution calculation layer to obtain the shifted operation data of cyclic shifting.

4. The method according to claim 3, wherein The after performing convolution processing on the preprocessed data in the first convolution calculation layer of the logical shifting unit to output two-channel convolution data, performing convolution processing through three different intermediate convolution calculation layers, and then performing convolution data channel merging processing through the fifth convolution calculation layer to obtain the shifted operation data of logical shifting includes: After performing convolution processing on the preprocessed data in the first convolution calculation layer of the logical shifting unit to output two-channel convolution data, performing convolution processing through three different intermediate convolution calculation layers set according to the logical left shift operation rule or the logical right shift operation rule, and then performing convolution data channel merging processing through the fifth convolution calculation layer to obtain the shifted operation data of logical left shift or logical right shift respectively.

5. The method according to claim 3, characterized in that, After the preprocessed data is convolved in the first convolutional calculation layer of the cyclic shift unit to output three-channel convolutional data, it is convolved through four different intermediate convolutional calculation layers, and then the convolutional data channel merging process is performed through the sixth convolutional calculation layer to obtain the shift operation data of the cyclic shift, including: After the preprocessed data is convolved in the first convolutional calculation layer of the cyclic shift unit to output three-channel convolutional data, it is convolved through four different intermediate convolutional calculation layers set according to the rules of cyclic left shift operation or cyclic right shift operation, and then the convolutional data channel merging process is performed through the sixth convolutional calculation layer to obtain the shift operation data of cyclic left shift or cyclic right shift respectively.

6. The method according to any one of claims 2 to 5, characterized in that: The byte length of the input serial data stream is a square number or a non-square number of any positive integer; When it is a non-square number, the serial data stream is segmented into data segments with a byte length of a square number of any positive integer including overlap for preprocessing, and after obtaining the shift operation data, splicing and merging are performed according to the principle of priority of the backend data segment in the shift direction.

7. A data shift processing device based on a convolutional neural network, characterized in that, Applied to a processor, the processor is a processor applied to a neural network chip, the processor includes a plurality of processing units, and each processing unit independently runs the assigned tasks, including convolutional operation tasks, and the device includes: A preprocessing module for preprocessing the input data to obtain preprocessed data; A shift operation module for processing the preprocessed data of the preprocessing module through a plurality of different convolutional calculation layers in the convolutional neural network according to different shift operation requirements to obtain shift operation data; A post-processing module for post-processing the shift operation data obtained by the shift operation module to obtain output data; Wherein the shift operation module includes a logical shift unit or a cyclic shift unit; The logical shift unit includes a first convolutional calculation layer, three different intermediate convolutional calculation layers, and a fifth convolutional calculation layer; The cyclic shift unit includes a first convolutional calculation layer, four different intermediate convolutional calculation layers, and a sixth convolutional calculation layer; Wherein each of the three different intermediate convolutional calculation layers includes two different convolutional kernels, and each convolutional kernel includes two channel parameters; or, Each of the four different intermediate convolutional calculation layers includes three different convolutional kernels, and each convolutional kernel includes three channel parameters.

8. The device according to claim 7, characterized in that: The logical shift unit is used to convolve the preprocessed data in the first convolutional calculation layer to output two-channel convolutional data, then convolve it through three different intermediate convolutional calculation layers, and then perform the convolutional data channel merging process through the fifth convolutional calculation layer to obtain the shift operation data of the logical shift; or, The cyclic shift unit is configured to perform convolution processing on the preprocessed data in the first convolutional calculation layer to output three-channel convolutional data, then perform convolution processing through four different intermediate convolutional calculation layers, and then perform convolution data channel merging processing through the sixth convolutional calculation layer to obtain the shifted operation data of cyclic shift.

9. An artificial intelligence chip, characterized in that, It includes the data shift processing device based on convolutional neural network according to any one of claims 7 to 8.

10. A computing device, characterized in that, It includes: A processor; And A memory storing executable code thereon, which when executed by the processor, causes the processor to execute the method according to any one of claims 1-6.

11. A computer-readable storage medium storing executable code thereon, which when executed by a processor of a computing device, causes the processor to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Neural network convolution operation device and method

    CN108229654A

  • Convolutional neural network generating classification for input image and computer implementation method

    CN108734269A