Processor, data processing method and computer equipment
By introducing a follow-up processing module into the processor to perform partial vector computing, the problems of large processing burden, low execution efficiency and high power consumption of the vector processing module are solved, and more efficient calculations and lower power consumption are achieved.
Patent Information
- Application Number
- CN202311812319.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
The vector processing module in existing processors has a high processing burden, low execution efficiency, and high power consumption.
The processor is introduced to the channel-accompanying processing module, which is responsible for the first type of vector operation, and the vector processing module is responsible for the second type of vector operation, and the output data of the tensor processing module is first passed into the channel-accompanying processing module for partial vector operation, reducing the burden on the vector processing module.
Through the calculation of the on-road processing module, the execution efficiency of the processor is improved, the processing burden of the vector processing module is reduced, and power consumption is saved.
Smart Images

Figure CN120216847A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a processor, a data processing method, and a computer device. Background Art
[0002] With the rapid development of computer technologies, processors are widely integrated into various chips. Among them, a processor usually includes a tensor processing module and a vector processing module. The tensor processing module is used to perform matrix operations, such as implementing the operations of the convolutional layer in a neural network. The vector processing module is used to perform vector operations, such as implementing the operations of the activation layer in a neural network. The output data of the tensor processing module is input into the vector processing module for processing.
[0003] Since various vector operations all need to rely on the computing power of the vector processing module, and the data flow in the vector processing module involves multiple levels such as buffers, arithmetic units, and registers, the processing burden of the vector processing module is relatively large and the execution efficiency is relatively low. Summary of the Invention
[0004] Embodiments of the present application provide a processor, a data processing method, and a computer device, which can reduce the processing burden of the vector processing module and save power consumption. The technical solutions are as follows:
[0005] On the one hand, a processor is provided. The processor includes a tensor processing module, an on-path processing module, and a vector processing module. The tensor processing module is connected to the on-path processing module, and the on-path processing module is connected to the vector processing module;
[0006] The tensor processing module is configured to perform a matrix operation on first data according to a data processing request to obtain second data, and transmit the second data to the on-path processing module;
[0007] The on-path processing module is configured to perform a first type of vector operation on the second data according to the data processing request to obtain third data. When the data processing request indicates to perform a second type of vector operation, the third data is transmitted to the vector processing module. When the data processing request does not indicate to perform a second type of vector operation, the third data is transmitted to a memory;
[0008] The vector processing module is configured to, when receiving the third data, perform a second type of vector operation on the third data according to the data processing request to obtain fourth data, and transmit the fourth data to the memory.
[0009] On the other hand, a data processing method is provided, which is executed by a processor. The processor includes a tensor processing module, a follow-up processing module, and a vector processing module. The tensor processing module is connected to the follow-up processing module, and the follow-up processing module is connected to the vector processing module. The method includes:
[0010] Through the tensor processing module, perform matrix operations on the first data according to a data processing request to obtain second data, and transmit the second data to the follow-up processing module;
[0011] Through the follow-up processing module, perform a first type of vector operation on the second data according to the data processing request to obtain third data. In the case where the data processing request indicates to perform a second type of vector operation, transmit the third data to the vector processing module. In the case where the data processing request does not indicate to perform a second type of vector operation, transmit the third data to the memory;
[0012] In the case where the vector processing module receives the third data, through the vector processing module, perform a second type of vector operation on the third data according to the data processing request to obtain fourth data, and transmit the fourth data to the memory.
[0013] Optionally, the target operation unit includes a rectified linear unit. The performing a vector operation on the second data through the target operation unit to obtain the third data includes:
[0014] Through the rectified linear unit, receive the second data and perform rectified linear transformation on the second data to obtain the third data.
[0015] Optionally, the transmitting the third data to the vector processing module in the case where the data processing request indicates to perform a second type of vector operation and transmitting the third data to the memory in the case where the data processing request does not indicate to perform a second type of vector operation includes:
[0016] Through the connection control unit, receive the third data. In the case where the data processing request indicates to perform a second type of vector operation, transmit the third data to the vector processing module. In the case where the data processing request does not indicate to perform a second type of vector operation, transmit the third data to the memory.
[0017] Optionally, the in-path processing module further includes a reshaping unit, and the reshaping unit is connected to the connection control unit; in the case that the data processing request indicates to execute a second type of vector operation, the third data is passed into the vector processing module, and in the case that the data processing request does not indicate to execute a second type of vector operation, passing the third data into the memory includes:
[0018] Receiving the third data through the connection control unit and passing the third data into the reshaping unit;
[0019] Reshaping the third data through the reshaping unit to obtain the reshaped third data, and in the case that the data processing request indicates to execute a second type of vector operation, passing the reshaped third data into the vector processing module, and in the case that the data processing request does not indicate to execute a second type of vector operation, passing the reshaped third data into the memory.
[0020] On the other hand, a computer device is provided, and the computer device includes a processor, and the processor is used to implement the operations performed by the processor as described in the above aspect.
[0021] On the other hand, a computer-readable storage medium is provided, and at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the processor as described in the above aspect.
[0022] On the other hand, a computer program product is provided, including a computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the processor as described in the above aspect.
[0023] In the solution provided by the embodiments of the present application, an in-path processing module is additionally added between the tensor processing module and the vector processing module in the processor. The in-path processing module is responsible for the first type of vector operation, and the vector processing module is responsible for the second type of vector operation. The output data of the tensor processing module will first be passed into the in-path processing module, and the in-path processing module performs the first type of vector operation on the data. If it is not necessary to execute the second type of vector operation, the in-path processing module can directly pass the operation result into the memory. If it is necessary to execute the second type of vector operation, the in-path processing module passes the operation result into the vector processing module for further operation. Therefore, in the process of the tensor processing module transmitting data to the vector processing module in the present application, the in-path processing module first completes a part of the vector operation in-path. On the one hand, the execution efficiency is improved through in-path calculation, and on the other hand, the processing burden of the vector processing module is also reduced, saving power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0025] Figure 1 is a schematic structural diagram of a processor provided by an embodiment of the present application;
[0026] Figure 2 is a schematic diagram of the position of an on-path processing module in a processor provided by an embodiment of the present application;
[0027] Figure 3 is a schematic structural diagram of another processor provided by an embodiment of the present application;
[0028] Figure 4 is a schematic structural diagram of another processor provided by an embodiment of the present application;
[0029] Figure 5 is a schematic diagram of a reshaping process provided by an embodiment of the present application;
[0030] Figure 6 is a schematic structural diagram of a quantization bias unit provided by an embodiment of the present application;
[0031] Figure 7 is a schematic structural diagram of a function operation unit provided by an embodiment of the present application;
[0032] Figure 8 is a schematic structural diagram of a number system conversion unit provided by an embodiment of the present application;
[0033] Figure 9 is a schematic diagram of an on-path processing module provided by an embodiment of the present application;
[0034] Figure 10 is a schematic diagram of a data flow provided by an embodiment of the present application;
[0035] Figure 11 is a schematic diagram of another data flow provided by an embodiment of the present application;
[0036] Figure 12 is a schematic diagram of the delay situation of an on-path processing module provided by an embodiment of the present application;
[0037] Figure 13 is a flowchart of a data processing method provided by an embodiment of the present application;
[0038] Figure 14 is a schematic structural diagram of a terminal provided by an embodiment of the present application;
[0039] Figure 15 It is a schematic structural diagram of a server provided by an embodiment of the present application. Specific embodiments
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0041] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the present application, the first type may be referred to as the second type, and similarly, the second type may be referred to as the first type.
[0042] Among them, at least one means one or more than one. For example, at least one arithmetic unit may be one arithmetic unit, two arithmetic units, three arithmetic units, etc., any integer greater than or equal to one. A plurality means two or more than two. For example, a plurality of arithmetic units may be two arithmetic units, three arithmetic units, etc., any integer greater than or equal to two. Each means each of at least one. For example, each arithmetic unit refers to each arithmetic unit among a plurality of arithmetic units. If there are 3 arithmetic units in a plurality of arithmetic units, then each arithmetic unit refers to each of the 3 arithmetic units.
[0043] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in the present application are all fully authorized by users or relevant parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0044] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0045] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0046] Machine Learning (ML) is an interdisciplinary subject across multiple fields, involving several disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration. The pre-trained model is the latest development result of deep learning, integrating the above technologies.
[0047] Computer Vision Technology (CV) is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as target recognition and measurement in machine vision, and further performing graphic processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as Swin-Transformer, ViT (Vision Transformer), V-MoE (Vision MoE), and MAE (Masked Auto Encoder) can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (3 Dimensions) technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies.
[0048] The processor provided by the embodiment of the present application can be applied to the above artificial intelligence technology and computer vision technology to implement the operations of the convolutional layer and the activation layer in the neural network model.
[0049] The embodiment of the present application provides a processor, which includes a tensor processing module, a follow-up processing module, and a vector processing module. The tensor processing module is connected to the follow-up processing module, and the follow-up processing module is connected to the vector processing module. In some embodiments, the processor is disposed in a computer device. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto.
[0050] Figure 1 It is a schematic structural diagram of a processor provided by the embodiment of the present application. As Figure 1 shown, the processor includes a tensor processing module 101, a follow-up processing module 102, and a vector processing module 103. The tensor processing module 101 is connected to the follow-up processing module 102, and the follow-up processing module 102 is connected to the vector processing module 103.
[0051] Refer to Figure 1 , the tensor processing module 101 is configured to perform matrix operations on the first data according to a data processing request to obtain second data, and transmit the second data to the follow-up processing module 102; the follow-up processing module 102 is configured to perform a first type of vector operation on the second data according to the data processing request to obtain third data. In the case where the data processing request indicates to perform a second type of vector operation, the third data is transmitted to the vector processing module 103. In the case where the data processing request does not indicate to perform a second type of vector operation, the third data is transmitted to the memory; the vector processing module 103 is configured to perform a second type of vector operation on the third data according to the data processing request when receiving the third data to obtain fourth data, and transmit the fourth data to the memory.
[0052] Among them, the processor in the embodiments of the present application can be an AI (Artificial Intelligence) processor or an AI chip. The AI processor is also called an AI accelerator or an AI computing card, and the AI processor is a processor used to process large computing tasks in the field of artificial intelligence. The tensor processing module 101 can be an MPU (Matrix Process Unit), and the tensor processing module 101 can implement the operations of the convolutional layer in the neural network. The vector processing module 103 can be a VPU (Vector Process Unit), and the vector processing module 103 can implement the operations of the activation layer in the neural network.
[0053] Generally, through the cooperation between the tensor processing module 101 and the vector processing module 103 in the processor, the processing of each layer in the neural network can be completed. In the embodiments of the present application, an additional on-path processing module 102 is added between the tensor processing module 101 and the vector processing module 103. The on-path processing module 102 can be an OLPU (Online Process Unit). The on-path processing module 102 is used to perform the first type of vector operations, and the vector processing module 103 is used to perform the second type of vector operations. It can be understood that a part of the computing power of the vector processing module 103 is transferred to the on-path processing module 102, and the on-path processing module 102 shares a part of the work of vector operations. The first type of vector operations includes various vector operations, and the second type of vector operations includes various vector operations, and the first type of vector operations is different from the second type of vector operations. Optionally, the second type of vector operations is more complex than the first type of vector operations. For example, the first type of vector operations includes vector quantization bias operations, number system conversion operations, etc., and the second type of vector operations includes vector shift operations, etc. The embodiments of the present application do not make specific limitations on the first type of vector operations and the second type of vector operations.
[0054] In the embodiment of the present application, the data processing request is used to request the processing of the first data. The first data can be the input data or intermediate data of the neural network model. The data processing request instructs to perform matrix operations, vector operations, etc. on the first data according to the architecture of the neural network model. First, the tensor processing module 101 obtains the first data, performs matrix operations on the first data according to the data processing request, obtains the second data, and passes the second data to the in-path processing module 102. The in-path processing module 102 receives the second data passed by the tensor processing module 101, and continues to perform the first type of vector operation on the second data according to the data processing request, obtaining the third data. If the data processing request instructs to perform the second type of vector operation, then the third data is not the final execution result of the data processing request, but only an intermediate execution result, that is, the operation is not yet completed and the vector processing module 103 needs to continue processing. Then the in-path processing module 102 passes the third data to the vector processing module 103, and the vector processing module 103 continues to perform the second type of vector operation on the third data according to the data processing request, obtaining the fourth data. The fourth data is the final execution result of the data processing request. Therefore, the vector processing module 103 can pass the fourth data to the memory. If the data processing request does not instruct to perform the second type of vector operation, then the third data is the final execution result of the data processing request, that is, the operation is completed. Then the in-path processing module 102 can pass the third data to the memory.
[0055] It should be noted that the above-mentioned tensor processing module 101 is used to perform at least one matrix operation, the above-mentioned in-path processing module 102 is used to perform at least one first type of vector operation, and the above-mentioned vector processing module 103 is used to perform at least one second type of vector operation. The data processing request indicates which operations to perform. Therefore, the tensor processing module 101 can determine which matrix operation needs to be performed on the first data according to the data processing request, the in-path processing module 102 can determine which first type of vector operation needs to be performed on the second data according to the data processing request, and the vector processing module 103 can determine which second type of vector operation needs to be performed on the third data according to the data processing request.
[0056] Figure 2 It is a schematic diagram of the position of an in-path processing module provided in the embodiment of the present application, as Figure 2As shown in the figure, the in-path processing module 102 is on the transmission path between the tensor processing module 101 and the vector processing module 103. The data output by the tensor processing module 101 needs to pass through the in-path processing module 102 first. If the in-path processing module 102 has completed all the vector operation processes of the data, the in-path processing module 102 can directly output the data to the memory through DMA (Direct Memory Access). If the in-path processing module 102 has only completed a part of the vector operation processes of the data, the in-path processing module 102 will pass the data into the vector processing module 103 to complete the remaining vector operation processes.
[0057] In the related art, the data output by the tensor processing module is directly input into the vector processing module. The vector processing module includes an L1 buffer (level 1 cache), a VRF (Vector Register File), and a VALU (Vector Arith Logic Unit). The data received by the vector processing module will first be stored in the L1 buffer, and then the data will be loaded from the L1 buffer to the VRF through a vector load instruction. Then, the VALU reads the data in the VRF through a vector calculation instruction and completes the operation. The operation result is written into the VRF, and finally, the operation result is written into the L1 buffer through a vector store instruction. Therefore, all the data that needs to perform vector operations needs to go through multiple levels such as the L1 buffer, VRF, and VALU in the vector processing module, consuming a large amount of execution time and having a high power consumption. However, the embodiment of the present application proposes an in-path processing module that couples the tensor processing module and the vector processing module. On the path from the tensor processing module to the vector processing module, the in-path processing module performs online data processing. The in-path processing module can first complete a part of the vector operations, sharing a part of the work of the vector processing module, releasing the data processing pressure of the vector processing module to a certain extent, reducing the processing burden of the vector processing module, and enabling the vector processing module to focus on more complex vector processing scenarios. Moreover, since the operation process of the in-path processing module is online and in-path, it does not need to go through multiple levels in the vector processing module, so it is beneficial to improve the data processing efficiency and reduce the power consumption.
[0058] In summary, in the embodiment of the present application, an on-the-fly processing module 102 is additionally added between the tensor processing module 101 and the vector processing module 103 in the processor. The on-the-fly processing module 102 is responsible for the first type of vector operations, and the vector processing module 103 is responsible for the second type of vector operations. The output data of the tensor processing module 101 will first be transmitted to the on-the-fly processing module 102, and the on-the-fly processing module 102 will perform the first type of vector operations on the data. If there is no need to perform the second type of vector operations, the on-the-fly processing module 102 can directly transmit the operation result to the memory. If it is necessary to perform the second type of vector operations, the on-the-fly processing module 102 will transmit the operation result to the vector processing module 103 for further operations. Therefore, in the process of the tensor processing module 101 transmitting data to the vector processing module 103 in the present application, the on-the-fly processing module 102 first completes a part of the vector operations on the fly. On the one hand, the execution efficiency is improved through on-the-fly calculation, and on the other hand, the processing burden of the vector processing module 103 is reduced, saving power consumption.
[0059] In some embodiments, referring to Figure 3 the structural schematic diagram of the processor shown, the on-the-fly processing module 102 includes a connection control unit 112 and a plurality of arithmetic units 122. The plurality of arithmetic units 122 are used to perform different vector operations, and the vector operations performed by the plurality of arithmetic units 122 all belong to the first type of vector operations. The connection control unit 112 is used to determine a target arithmetic unit 122 among the plurality of arithmetic units 122, and turn on the connection between the target arithmetic unit 122 and the connection control unit 112. The target arithmetic unit 122 is used to perform the vector operation indicated by the data processing request; the connection control unit 112 is further used to receive the second data transmitted by the tensor processing module 101 and transmit the second data to the target arithmetic unit 122; the target arithmetic unit 122 is used to perform a vector operation on the second data to obtain third data and transmit the third data to the connection control unit 112.
[0060] Among them, the connection control unit 112 is respectively connected to a plurality of arithmetic units 122, and the connection control unit 112 can control the opening and closing of the connection between the connection control unit 112 and each arithmetic unit 122. The connection control unit 112 is also connected to the tensor processing module 101, and the second data output by the tensor processing module 101 is input to the connection control unit 112. The connection control unit 112 can determine the target arithmetic unit 122 to be used according to the data processing request, and the target arithmetic unit 122 can execute the first type of vector operation indicated by the data processing request. Therefore, the connection control unit 112 opens the connection between the target arithmetic unit 122 and the connection control unit 112, and passes the received second data into the target arithmetic unit 122 through the connection between the target arithmetic unit 122 and the connection control unit 112. After receiving the second data passed in by the connection control unit, the target arithmetic unit 122 can perform a vector operation on the second data to obtain the third data, and pass the third data back to the connection control unit through the connection between the target arithmetic unit 122 and the connection control unit 112. Optionally, the connection control unit 112 is connected to a plurality of arithmetic units 122 through a multiplexer, and the multiplexer can be used to implement the opening and closing of the connection between the connection control unit 112 and the arithmetic unit 122. Therefore, the connection control unit 112 can open the connection with the target arithmetic unit 122 through the multiplexer and close the connection with other arithmetic units 122, so that the connection control unit 112 can pass the second data into the target arithmetic unit 122.
[0061] In the embodiment of the present application, the connection control unit 112 in the in-path processing module 102 can control the opening and closing of the connection with a plurality of arithmetic units 122, so as to flexibly select which arithmetic unit to use to execute vector operations according to different data processing requests, and can flexibly respond to different calculation scenarios and calculation requirements.
[0062] In a possible implementation manner, as Figure 3 shown, the connection control unit 112 is configured to receive the third data, and pass the third data into the vector processing module 103 when the data processing request indicates the execution of the second type of vector operation, and pass the third data into the memory when the data processing request does not indicate the execution of the second type of vector operation. Among them, the connection control unit 112 is also connected to the memory and the vector processing module 103. When all the vector operation processes indicated by the data processing request have been completed, the connection control unit 112 can pass the third data into the memory without further processing by the vector processing module 103. When all the vector operation processes indicated by the data processing request have not been completed, the connection control unit 112 passes the third data into the vector processing module 103 to continue the vector operation.
[0063] In a possible implementation, referring to Figure 4 the schematic structural diagram of the processor shown, the on-path processing module 102 further includes a reshaping unit 132, and the reshaping unit 132 is connected to the connection control unit 112. The connection control unit 112 is configured to receive the third data and transmit the third data to the reshaping unit 132; the reshaping unit 132 is configured to reshape the third data to obtain the reshaped third data, and when the data processing request indicates to execute the second type of vector operation, transmit the reshaped third data to the vector processing module 103, and when the data processing request does not indicate to execute the second type of vector operation, transmit the reshaped third data to the memory.
[0064] Wherein, the reshaping unit 132 can be referred to as a Reshape unit, and the reshaping unit 132 is configured to complete the specification conversion of the input data, such as adjusting the number of rows, columns or dimensions of the data, etc.
[0065] Figure 5 is a schematic diagram of a reshaping process provided by an embodiment of the present application. As Figure 5 shown, the input data of the reshaping unit 132 is an image. The input image before reshaping is 4 rows and 2 columns. After being processed by the reshaping unit 132, the output image after reshaping is 2 rows and 4 columns.
[0066] In a possible implementation, as Figure 4 shown, the processor further includes a control module 104, and the on-path processing module 102 further includes a configuration unit 142, and the configuration unit 142 is connected to the connection control unit 112. Wherein, the control module 104 is respectively connected to the tensor processing module 101, the on-path processing module 102 and the vector processing module 103. The control module 104 is configured to generate configuration information based on the data processing request and transmit the configuration information to the configuration unit 142. The configuration information includes a target identifier, and the target identifier indicates the target arithmetic unit 122; the configuration unit 142 is configured to receive the configuration information and transmit the target identifier in the configuration information to the connection control unit 112; the connection control unit 112 is configured to determine the target arithmetic unit 122 indicated by the target identifier and turn on the connection between the target arithmetic unit 122 and the connection control unit 112.
[0067] Among them, the control module 104 receives a data processing request, generates configuration information by parsing the data processing request, and the configuration information includes a target identifier for indicating the target arithmetic unit 122. The control module 104 passes the configuration information to the configuration unit 142 in the in-path processing module 102. The configuration unit 142 parses the configuration information to obtain the target identifier and passes the target identifier to the connection control unit 112. Then, the connection control unit 112 can determine the target arithmetic unit 122 according to the target identifier.
[0068] Optionally, the connection control unit 112 includes an internal register. After the configuration unit 142 parses the configuration information to obtain the target identifier, it writes the target identifier into the internal register in the connection control unit 112. Optionally, the connection control unit 112 is connected to multiple arithmetic units 122 through a multiplexer. The connection control unit 112 passes the target identifier in the internal register to the multiplexer, and the multiplexer opens the connection between the connection control unit 112 and the target arithmetic unit 122 indicated by the target identifier.
[0069] In a possible implementation, as Figure 4 shown, the configuration information further includes configuration parameters corresponding to the target identifier. The configuration unit 142 is also connected to multiple arithmetic units 122. The configuration unit 142 is further configured to pass the configuration parameters to the target arithmetic unit 122 indicated by the corresponding target identifier. The target arithmetic unit 122 is configured to perform a vector operation on the second data based on the configuration parameters to obtain a third data, and pass the third data to the connection control unit 112.
[0070] Among them, the configuration parameters corresponding to the target identifier refer to the configuration parameters of the target arithmetic unit 122 indicated by the target identifier. For example, when the target arithmetic unit 122 is a quantization bias unit, the configuration parameters include quantization parameters and bias parameters. When the target arithmetic unit 122 is a function arithmetic unit, the configuration parameters include function identifiers, etc. The process of the target arithmetic unit 122 performing a vector operation on the second data based on the configuration parameters can be seen in the description of parts (1)-(4) in the following embodiments, and will not be elaborated here.
[0071] Optionally, the arithmetic unit includes an internal register. The configuration information further includes the register address of the internal register in the target arithmetic unit 122 indicated by the target identifier, and the register address is stored corresponding to the configuration parameters. The configuration unit 142 passes the configuration parameters into the internal register indicated by the corresponding register address. The target arithmetic unit 122 receives the second data passed by the connection control unit 112, obtains the configuration parameters in the internal register, and thus performs a vector operation on the second data based on the configuration parameters.
[0072] It should be noted that the above embodiments are only described by taking the control module 104 sending configuration information to the configuration unit 142 in the in-path processing module 102 as an example. As Figure 4 shown, the control module 104 is also connected to the tensor processing module 101 and the vector processing module 103. The control module 104 can generate configuration information corresponding to the tensor processing module 101, configuration information corresponding to the in-path processing module 102, and configuration information corresponding to the vector processing module 103 respectively based on data processing requests. The configuration information corresponding to the tensor processing module 101 is used to guide the tensor processing module 101 to perform matrix operations on the first data. For example, this configuration information indicates what kind of matrix operations to execute, and this configuration information includes the configuration parameters required by the tensor processing module 101. The configuration information corresponding to the in-path processing module 102 is used to guide the in-path processing module 102 to perform a first type of vector operation on the second data. For example, this configuration information indicates what kind of first type of vector operation to execute, and this configuration information includes the configuration parameters required by the in-path processing module 102. The configuration information corresponding to the vector processing module 103 is used to guide the vector processing module 103 to perform a second type of vector operation on the third data. For example, this configuration information indicates what kind of second type of vector operation to execute, and this configuration information includes the configuration parameters required by the vector processing module 103, etc.
[0073] In some embodiments, the multiple operation units 122 in the in-path processing module 102 include at least one of a quantization bias unit, a function operation unit, a number system conversion unit, or a rectified linear unit. The target operation unit 122 can be any one of the multiple operation units 122, and the specific content is as follows.
[0074] (1) The target operation unit 122 includes a quantization bias unit. The quantization bias unit is used to receive configuration parameters and the second data. The configuration parameters include quantization parameters and bias parameters; the quantization bias unit is used to multiply the second data by the quantization parameters and add the result of the multiplication to the bias parameters to obtain the third data. Optionally, the configuration parameters (quantization parameters and bias parameters) in this quantization bias unit are passed in by the above-mentioned configuration unit 142, and the second data in this quantization bias unit is passed in by the above-mentioned connection control unit 112.
[0075] Among them, this quantization bias unit can be called QBU (Quantization Bias Unit), the quantization parameter can be called q_coef (quantization coefficient), and the bias parameter can be called b_coef (bias coefficient). This quantization bias unit is used to perform quantization operations and bias operations on vectors. The quantization operation refers to multiplying the input data by the quantization parameter, and the bias operation refers to adding the input data to the bias parameter.
[0076] Figure 6 It is a schematic structural diagram of a quantization bias unit provided by an embodiment of the present application. As Figure 6 shown, the quantization bias unit includes a multiplication operator, an addition operator, and an internal register. The internal register stores the quantization parameters and bias parameters passed in by the configuration unit 142. The connection control unit 112 passes the input data (second data) into the quantization bias unit. The quantization bias unit multiplies the input data by the quantization parameters through the multiplication operator to obtain the multiplication result, and adds the multiplication result and the bias parameters through the addition operator to obtain the output data (third data).
[0077] (2) The target operation unit 122 includes a function operation unit, and the function operation unit is used to perform a variety of different function operations. The function operation unit is used to receive configuration parameters and second data, and the configuration parameters include a function identifier; the function operation unit is used to query the operation coefficients corresponding to the target function indicated by the function identifier, substitute the operation coefficients into a preset function template to obtain a fitting function, and the fitting function is used to fit the target function; the function operation unit is used to input the second data into the fitting function to obtain the third data. Optionally, the configuration parameters (function identifier) in the function operation unit are passed in by the above-mentioned configuration unit 142, and the second data in the function operation unit is passed in by the above-mentioned connection control unit 112.
[0078] Among them, the function operation unit can be called an SFU (Special Function Unit), and the function operation unit is used to perform a variety of different function operations. For example, the variety of different functions include exponential function (exp), logarithmic function (ln), and reciprocal function (1 / x), etc. The function identifier is used to indicate which function to use. Optionally, the function operation unit includes an operation coefficient table corresponding to each function, and the operation coefficient table includes at least one operation coefficient corresponding to the function. The function operation unit includes a multiplexer (mux, multiplexer). The input of the multiplexer is the operation coefficient tables corresponding to multiple functions and the function identifier, and the output of the multiplexer is at least one operation coefficient in the operation coefficient table corresponding to the target function indicated by the function identifier. The function operation unit further includes a preset function template. Substituting at least one operation coefficient output by the multiplexer into the preset function template can obtain a fitting function, and the fitting function is used to fit the target function. Therefore, the third data obtained after inputting the second data into the fitting function can be approximately the result of operating on the second data using the target function.
[0079] Figure 7 It is a schematic structural diagram of a function operation unit provided by an embodiment of the present application. As Figure 7As shown, the function operation unit includes an operation coefficient table corresponding to the exponential function, an operation coefficient table corresponding to the logarithmic function, an operation coefficient table corresponding to the reciprocal function, a multiplexer, a preset function template, and an internal memory. The internal memory stores the function identifier passed in by the configuration unit 142. The preset function template is ax 2 +bx + c, where a, b, and c are operation coefficients. The three operation coefficient tables and the function identifier are input into the multiplexer. The multiplexer outputs the operation coefficients in the operation coefficient table corresponding to the target function indicated by the function identifier. The output operation coefficients are substituted into the preset function template to obtain a fitting function. The input data (second data) is input into the fitting function to obtain the output data (third data). For example, if the target function is a logarithmic function, the operation coefficients corresponding to the logarithmic function are substituted into the preset function template to obtain a fitting function, and this fitting function is used to fit the logarithmic function, that is, the operation result of this fitting function is approximately the same as the operation result of the logarithmic function.
[0080] (3) The target operation unit 122 includes a number system conversion unit, which is used to perform various different number system conversion operations; the number system conversion unit is used to receive configuration parameters and second data, and the configuration parameters include a conversion identifier; the number system conversion unit is used to perform the number system conversion operation indicated by the conversion identifier on the second data to obtain third data. Optionally, the configuration parameter (conversion identifier) in this number system conversion unit is passed in by the above-mentioned configuration unit 142, and the second data in this number system conversion unit is passed in by the above-mentioned connection control unit 112.
[0081] Among them, this number system conversion unit can be called CVT (Convert, number system conversion unit). The number system conversion unit is used to perform various different number system conversion operations, and the conversion identifier is used to indicate which number system conversion operation to use. For example, the various number system conversion operations include FP2INT (conversion from floating point number to fixed point number), FP2FP (conversion between different types of floating point numbers), INT2FP (conversion from fixed point number to floating point number), and INT2INT (conversion between different types of fixed point numbers). Among them, FP (Floating-Point) refers to a floating point number, and INT (Integer) refers to a fixed point number.
[0082] Figure 8 is a schematic structural diagram of a number system conversion unit provided by an embodiment of the present application, as Figure 8As shown in the figure, the number system conversion unit includes operators corresponding to four number system conversion operations: FP2INT, FP2FP, INT2FP, and INT2INT. The number system conversion unit further includes a multiplexer and an internal memory. The internal memory stores the conversion identifier passed in by the configuration unit 142. The multiplexer determines the operator corresponding to the number system conversion operation indicated by the conversion identifier, and performs a number system conversion on the input data (the second data) using the operator corresponding to the number system conversion operation indicated by the conversion identifier to obtain the output data (the third data).
[0083] (4) The target operation unit 122 includes a rectified linear unit. The rectified linear unit is configured to receive the second data and perform a rectified linear operation on the second data to obtain the third data.
[0084] Among them, the rectified linear unit can be referred to as RELU (Rectified Linear Unit), and the rectified linear unit is an activation function of a neural network.
[0085] The above embodiments are described from the perspective that the number of target operation units is 1. In some embodiments, the number of target operation units 122 is n, where n is an integer greater than 1. Then refer to Figure 3 or Figure 4 For the structural schematic diagram, the process of the follow-up processing module 102 performing the first type of vector operation on the second data is as follows.
[0086] The connection control unit 112 is configured to determine the first target operation unit 122 to the nth target operation unit 122 according to the execution order indicated by the data processing request, open the connection between the first target operation unit 122 and the connection control unit 112, open the connection between the ith target operation unit 122 and the (i + 1)th target operation unit 122, and open the connection between the nth target operation unit 122 and the connection control unit 112, where i is a positive integer less than n; the connection control unit 112 is further configured to receive the second data passed in by the tensor processing module 101 and pass the second data to the first target operation unit 122; the (i + 1)th target operation unit 122 is configured to receive the ith intermediate result passed in by the ith target operation unit 122, perform a vector operation on the ith intermediate result to obtain the (i + 1)th intermediate result; the (i + 1)th target operation unit 122 is configured to pass the (i + 1)th intermediate result to the (i + 2)th target operation unit 122 when i + 1 is less than n, and pass the (i + 1)th intermediate result to the connection control unit 112 as the third data when i + 1 is equal to n.
[0087] Among them, when the data processing request indicates to sequentially execute n vector operations of the first type, the in-path processing module 102 needs to use the corresponding n target operation units to execute vector operations. First, the connection control unit 112 in the in-path processing module 102 determines the 1st target operation unit 122 to the nth target operation unit 122 among multiple operation units according to the execution order indicated by the data processing request. The 1st target operation unit 122 to the nth target operation unit 122 are sorted according to the execution order. For example, multiple operation units 122 include a quantization bias unit, a function operation unit, a number system conversion unit, and a rectified linear unit. If the data processing request indicates to sequentially execute a quantization bias operation, a number system conversion, and a rectified linear operation, then the quantization bias unit is the 1st target operation unit 122, the number system conversion unit is the 2nd target operation unit 122, and the rectified linear unit is the 3rd target operation unit.
[0088] After determining the 1st target operation unit 122 to the nth target operation unit 122, the connection control unit 112 sequentially turns on the connections for the 1st target operation unit 122 to the nth target operation unit 122, that is, turns on the connections between any two adjacent target operation units 122, and turns on the connections between the connection control unit 112 and the 1st target operation unit 122 and the nth target operation unit 122 (that is, the last target operation unit 122).
[0089] After turning on the connections, the connections between the connection control unit 112 and the 1st target operation unit 122 to the nth target operation unit 122 are the data flow directions. First, the connection control unit 112 passes the second data into the 1st target operation unit 122. The 1st target operation unit 122 performs a vector operation on the second data to obtain the 1st intermediate result, passes the 1st intermediate result into the 2nd target operation unit 122. The 2nd target operation unit 122 performs a vector operation on the 1st intermediate result to obtain the 2nd intermediate result, passes the 2nd intermediate result into the 3rd target operation unit 122, and so on, until the nth target operation unit performs a vector operation on the n - 1th intermediate result to obtain the nth intermediate result. The nth intermediate result is the third data.
[0090] In the embodiments of the present application, the connection control unit 112 in the in-path processing module 102 can adjust the order between different operation units, can flexibly select which vector operation to execute first and which to execute later, and can flexibly respond to different calculation scenarios and calculation requirements.
[0091] Combined with the above various embodiments, Figure 9 is a schematic diagram of an in-path processing module provided by the embodiments of the present application, as Figure 9As shown, the in-path processing module includes a configuration unit, a connection control unit, a reshaping unit, and multiple arithmetic units. The multiple arithmetic units include a quantization bias unit, a function arithmetic unit, a number system conversion unit, and a rectified linear unit. The structure of the quantization bias unit is the same as that of Figure 6 Similarly, the structure of the function arithmetic unit is the same as that of Figure 7 Similarly, the structure of the number system conversion unit is the same as that of Figure 8 Similarly, it will not be elaborated here.
[0092] Among them, the configuration unit is connected to the connection control unit, the quantization bias unit, the function arithmetic unit, and the number system conversion unit. The configuration unit is used to receive configuration information and complete the conversion of the configuration information into internal register write requests for each connected unit. That is, the configuration parameters in the connection control unit, the quantization bias unit, the function arithmetic unit, and the number system conversion unit all come from the configuration unit. For example, the configuration unit writes the quantization parameters and bias parameters in the configuration information into the internal register of the quantization bias unit, writes the function identifier in the configuration information into the internal register of the function arithmetic unit, and writes the conversion identifier in the configuration information into the internal register of the number system conversion unit.
[0093] Among them, the connection control unit is connected to the quantization bias unit, the function arithmetic unit, the number system conversion unit, and the rectified linear unit. The connection control unit is used to control the opening and closing of the connections with each arithmetic unit. The connection control unit is also connected to the configuration unit and the reshaping unit. The second data output by the tensor processing module will be input into the connection control unit of this in-path processing module, and the reshaping unit of this in-path processing module outputs the third data.
[0094] The following Figure 10 and Figure 11 embodiments, based on the structure schematic diagram of the in-path processing module shown in Figure 9 respectively show the data flow in the in-path processing module when the number of target arithmetic units is 1 and 2.
[0095] Figure 10 is a schematic diagram of a data flow provided by an embodiment of the present application. As shown in Figure 10 , the number of target arithmetic units is 1, and the target arithmetic unit is the function arithmetic unit. Then the data flow in the in-path processing module is as follows.
[0096] (1a) The configuration unit configures the connection control unit according to the configuration information passed in by the control module. For example, it passes the target identifier into the internal register of the connection control unit, and the target identifier indicates the function arithmetic unit. The connection control unit opens the connection between the connection control unit and the function arithmetic unit, so that the input data of the connection control unit flows to the entrance of the function arithmetic unit, and the output data of the function arithmetic unit flows to the exit of the connection control unit.
[0097] (1b) The configuration unit configures the function operation unit according to the configuration information passed in by the control module. For example, it passes the function identifier into the internal register of the function operation unit, and this function identifier indicates the logarithmic function.
[0098] (2) The connection control unit receives the input second data and passes the second data into the function operation unit.
[0099] (3) The function operation unit completes the function operation on the second data based on the function identifier configured in the internal register, and obtains the third data.
[0100] (4) The function operation unit passes the third data into the connection control unit.
[0101] (5) The connection control unit passes the third data into the reshaping unit, and the reshaping unit completes the reshaping of the third data and outputs the reshaped third data.
[0102] Figure 11 It is a schematic diagram of another data flow direction provided by an embodiment of the present application. As Figure 11 shown, the number of target operation units is 2, and the target operation units are the function operation unit and the number system conversion unit. Then the data flow direction in the follow-up processing module is as follows.
[0103] (1a) The configuration unit configures the connection control unit according to the configuration information passed in by the control module. For example, it passes the target identifier into the internal register of the connection control unit, and this target identifier indicates the function operation unit and the number system conversion unit. The connection control unit enables the connection between the connection control unit and the function operation unit, enables the connection between the function operation unit and the number system conversion unit, and enables the connection between the number system conversion unit and the connection control unit, so that the input data of the connection control unit flows to the entrance of the function operation unit, the output data of the function operation unit flows to the entrance of the number system conversion unit, and the output data of the number system conversion unit flows to the exit of the connection control unit.
[0104] (1b) The configuration unit configures the function operation unit according to the configuration information passed in by the control module. For example, it passes the function identifier into the internal register of the function operation unit, and this function identifier indicates the logarithmic function.
[0105] (1c) The configuration unit configures the number system conversion unit according to the configuration information passed in by the control module. For example, it passes the conversion identifier into the internal register of the number system conversion unit, and this conversion identifier indicates FP2FP.
[0106] (2) The connection control unit receives the input second data and passes the second data into the function operation unit.
[0107] (3) The function operation unit performs a function operation on the second data based on the function identifier configured in the internal register to obtain an intermediate result.
[0108] (4) The function operation unit passes the intermediate result to the number system conversion unit.
[0109] (5) The number system conversion unit performs a number system conversion on the intermediate result based on the conversion identifier configured in the internal register to obtain the third data.
[0110] (6) The number system conversion unit passes the third data to the connection control unit.
[0111] (7) The connection control unit passes the third data to the reshaping unit, and the reshaping unit completes the reshaping of the third data and outputs the reshaped third data.
[0112] In some embodiments, the in-path processing module can process the data of multiple data processing requests in the pipeline order. And when the vector operations indicated by two or more adjacent data processing requests are different from each other and the data indicated for processing are independent of each other, the in-path processing module can also process the data corresponding to two or more adjacent data processing requests in parallel. For example, if the second data processing request indicates to perform a quantization bias operation and the third data processing request indicates to perform a function operation, the in-path processing module can simultaneously perform the quantization bias operation on the data corresponding to the second data processing request and perform the function operation on the data corresponding to the third data processing request.
[0113] Since the vector operation process of the in-path processing module is in-path, only the delay will affect the execution efficiency. Figure 12 is a schematic diagram of the delay situation of an in-path processing module provided by an embodiment of the present application. As Figure 12 shown, IN0-IN6 are the input data corresponding to different data processing requests, OUT0-OUT6 are the output data corresponding to different data processing requests, and IN0-IN6 and OUT0-OUT6 correspond one by one. After inputting IN0, after a period of delay, the in-path processing module outputs OUT0. Among them, the duration of the delay is related to the number of arithmetic units required for the data processing request. The larger the number of arithmetic units required, the longer the duration of the delay.
[0114] Figure 13 is a flowchart of a data processing method provided by an embodiment of the present application. The embodiment of the present application is executed by a processor. The processor includes a tensor processing module, an in-path processing module, and a vector processing module. The tensor processing module is connected to the in-path processing module, and the in-path processing module is connected to the vector processing module. Refer to Figure 13 , the method includes:
[0115] 1301. Through the tensor processing module, perform matrix operations on the first data according to the data processing request to obtain the second data, and transmit the second data to the in-path processing module.
[0116] 1302. Through the in-path processing module, perform vector operations of the first type on the second data according to the data processing request to obtain the third data. In the case where the data processing request indicates to perform vector operations of the second type, transmit the third data to the vector processing module. In the case where the data processing request does not indicate to perform vector operations of the second type, transmit the third data to the memory.
[0117] 1303. In the case where the vector processing module receives the third data, through the vector processing module, perform vector operations of the second type on the third data according to the data processing request to obtain the fourth data, and transmit the fourth data to the memory.
[0118] In the method provided by the embodiment of this application, an in-path processing module is additionally added between the tensor processing module and the vector processing module in the processor. The in-path processing module is responsible for vector operations of the first type, and the vector processing module is responsible for vector operations of the second type. The output data of the tensor processing module will first be transmitted to the in-path processing module, and the in-path processing module performs vector operations of the first type on the data. If there is no need to perform vector operations of the second type, the in-path processing module directly transmits the operation result to the memory. If it is necessary to perform vector operations of the second type, the in-path processing module transmits the operation result to the vector processing module for further operations. Therefore, in the process of transmitting data from the tensor processing module to the vector processing module in this application, part of the vector operations are completed in-path by using the in-path processing module. On the one hand, the execution efficiency is improved through in-path calculation, and on the other hand, the processing burden of the vector processing module is reduced, saving power consumption.
[0119] In a possible implementation manner, the in-path processing module includes a connection control unit and multiple operation units. The multiple operation units are used to perform different vector operations, and the vector operations performed by the multiple operation units belong to the first type. The step of performing vector operations of the first type on the second data according to the data processing request through the in-path processing module to obtain the third data includes: through the connection control unit, determine the target operation unit among the multiple operation units, and enable the connection between the target operation unit and the connection control unit. The target operation unit is used to perform the vector operations indicated by the data processing request; through the connection control unit, receive the second data transmitted by the tensor processing module and transmit the second data to the target operation unit; through the target operation unit, perform vector operations on the second data to obtain the third data, and transmit the third data to the connection control unit.
[0120] In a possible implementation, the number of target arithmetic units is n, where n is an integer greater than 1. The connection control unit determines the target arithmetic units among multiple arithmetic units and enables the connection between the target arithmetic units and the connection control unit, including: through the connection control unit, determining the first target arithmetic unit to the nth target arithmetic unit according to the execution order indicated by the data processing request, enabling the connection between the first target arithmetic unit and the connection control unit, enabling the connection between the ith target arithmetic unit and the (i + 1)th target arithmetic unit, and enabling the connection between the nth target arithmetic unit and the connection control unit, where i is a positive integer less than n. The connection control unit receives the second data passed in by the tensor processing module and passes the second data to the target arithmetic unit, including: through the connection control unit, receiving the second data passed in by the tensor processing module and passing the second data to the first target arithmetic unit. The target arithmetic unit performs vector operations on the second data to obtain the third data and passes the third data to the connection control unit, including: through the (i + 1)th target arithmetic unit, receiving the ith intermediate result passed in by the ith target arithmetic unit, performing vector operations on the ith intermediate result to obtain the (i + 1)th intermediate result; when i + 1 is less than n, passing the (i + 1)th intermediate result to the (i + 2)th target arithmetic unit through the (i + 1)th target arithmetic unit, and when i + 1 is equal to n, passing the (i + 1)th intermediate result as the third data to the connection control unit through the (i + 1)th target arithmetic unit.
[0121] In a possible implementation, the processor further includes a control module, and the in-path processing module further includes a configuration unit, and the configuration unit is connected to the connection control unit. The method further includes: through the control module, generating configuration information based on the data processing request and passing the configuration information to the configuration unit, where the configuration information includes a target identifier indicating the target arithmetic unit; through the configuration unit, receiving the configuration information and passing the target identifier in the configuration information to the connection control unit. The connection control unit determines the target arithmetic units among multiple arithmetic units and enables the connection between the target arithmetic units and the connection control unit, including: through the connection control unit, determining the target arithmetic unit indicated by the target identifier and enabling the connection between the target arithmetic unit and the connection control unit.
[0122] In a possible implementation, the configuration information further includes configuration parameters corresponding to the target identifier, and the configuration unit is further connected to multiple arithmetic units. The method further includes: through the configuration unit, further used to pass the configuration parameters to the target arithmetic unit indicated by the corresponding target identifier. The target arithmetic unit performs vector operations on the second data to obtain the third data and passes the third data to the connection control unit, including: through the target arithmetic unit, performing vector operations on the second data based on the configuration parameters to obtain the third data and passing the third data to the connection control unit.
[0123] In a possible implementation manner, the target arithmetic unit includes a quantization bias unit. By means of the target arithmetic unit, vector operations are performed on the second data to obtain third data, including: receiving, by the quantization bias unit, a configuration parameter and the second data, where the configuration parameter includes a quantization parameter and a bias parameter; multiplying, by the quantization bias unit, the second data by the quantization parameter, and adding the result of the multiplication to the bias parameter to obtain the third data.
[0124] In a possible implementation manner, the target arithmetic unit includes a function arithmetic unit, and the function arithmetic unit is used to perform a variety of different function operations. By means of the target arithmetic unit, vector operations are performed on the second data to obtain third data, including: receiving, by the function arithmetic unit, a configuration parameter and the second data, where the configuration parameter includes a function identifier; querying, by the function arithmetic unit, the operation coefficients corresponding to the target function indicated by the function identifier, substituting the operation coefficients into a preset function template to obtain a fitting function, where the fitting function is used to fit the target function; inputting, by the function arithmetic unit, the second data into the fitting function to obtain the third data.
[0125] In a possible implementation manner, the target arithmetic unit includes a number system conversion unit, and the number system conversion unit is used to perform a variety of different number system conversion operations. By means of the target arithmetic unit, vector operations are performed on the second data to obtain third data, including: receiving, by the number system conversion unit, a configuration parameter and the second data, where the configuration parameter includes a conversion identifier; performing, by the number system conversion unit, the number system conversion operation indicated by the conversion identifier on the second data to obtain the third data.
[0126] In a possible implementation manner, the target arithmetic unit includes a rectified linear unit. By means of the target arithmetic unit, vector operations are performed on the second data to obtain third data, including: receiving, by the rectified linear unit, the second data and performing rectified linear transformation on the second data to obtain the third data.
[0127] In a possible implementation manner, when the data processing request indicates to execute the second type of vector operation, the third data is passed into the vector processing module, and when the data processing request does not indicate to execute the second type of vector operation, the third data is passed into the memory, including: receiving, by the connection control unit, the third data, and when the data processing request indicates to execute the second type of vector operation, passing the third data into the vector processing module, and when the data processing request does not indicate to execute the second type of vector operation, passing the third data into the memory.
[0128] In a possible implementation, the in-path processing module further includes a reshaping unit, and the reshaping unit is connected to the connection control unit. When the data processing request indicates the execution of the second type of vector operation, the third data is passed into the vector processing module, and when the data processing request does not indicate the execution of the second type of vector operation, the third data is passed into the memory, including: receiving the third data through the connection control unit and passing the third data into the reshaping unit; reshaping the third data through the reshaping unit to obtain the reshaped third data, and when the data processing request indicates the execution of the second type of vector operation, passing the reshaped third data into the vector processing module, and when the data processing request does not indicate the execution of the second type of vector operation, passing the reshaped third data into the memory.
[0129] It should be noted that the data processing method provided in the embodiments of the present application and the processor provided in the above embodiments belong to the same inventive concept. The specific implementation of the data processing method provided in the embodiments of the present application can refer to the embodiments of the above processor, and will not be elaborated herein.
[0130] The embodiments of the present application further provide a computer device, which includes a processor for implementing the operations performed by the processor in the above embodiments. In some embodiments, the computer device further includes a memory, and at least one computer program is stored in the memory and loaded and executed by the processor to implement the operations performed by the processor in the above embodiments.
[0131] Optionally, the computer device is provided as a terminal. Figure 14 FIG. 11 shows a schematic structural diagram of a terminal 1400 provided by an exemplary embodiment of the present application. The terminal 1400 includes a processor 1401 and a memory 1402.
[0132] The processor 1401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning. In some embodiments, the processor 1401 is used to implement the operations performed by the processor in the above embodiments.
[0133] The memory 1402 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1402 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1402 is used to store at least one computer program, and the at least one computer program is used to be executed by the processor 1401 to implement the data processing method provided in the method embodiments of the present application.
[0134] In some embodiments, the terminal 1400 may further optionally include: a peripheral device interface 1403 and at least one peripheral device. The processor 1401, the memory 1402, and the peripheral device interface 1403 are connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1403 through a bus, signal lines, or a circuit board. Optionally, the peripheral device includes at least one of a radio frequency circuit 1404 or a display screen 1405.
[0135] The peripheral device interface 1403 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1401 and the memory 1402. In some embodiments, the processor 1401, the memory 1402, and the peripheral device interface 1403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1401, the memory 1402, and the peripheral device interface 1403 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.
[0136] The radio frequency circuit 1404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1404 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1404 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 1404 can communicate with other devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network.
[0137] The display screen 1405 is used to display a UI (User Interface). The UI can include graphics, text, icons, videos, and any combination thereof. When the display screen 1405 is a touch display screen, the display screen 1405 also has the ability to collect touch signals on or above the surface of the display screen 1405. The touch signal can be input to the processor 1401 as a control signal for processing. At this time, the display screen 1405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard.
[0138] Those skilled in the art can understand that Figure 14 the structure shown in
[0139] does not constitute a limitation on the terminal 1400, and may include more or fewer components than shown in the figure, or combine some components, or adopt a different component layout. Figure 15FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1500 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1501 and one or more memories 1502. In some embodiments, the processor 1501 is used to implement the operations performed by the processor in the above embodiments. In some embodiments, at least one computer program is stored in the memory 1502, and the at least one computer program is loaded and executed by the processor 1501 to implement the operations performed by the processor in the above embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0140] An embodiment of the present application also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the operations performed by the above-mentioned processor.
[0141] An embodiment of the present application also provides a computer program product, including a computer program, which is loaded and executed by a processor to implement the operations performed by the processor in the above embodiments.
[0142] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0143] The above are only optional embodiments of the embodiments of the present application, and are not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the present application.
Claims
1. A processor, characterized in that, The processor includes a tensor processing module, an in-path processing module, and a vector processing module. The tensor processing module is connected to the in-path processing module, and the in-path processing module is connected to the vector processing module; The tensor processing module is configured to perform matrix operations on first data according to a data processing request to obtain second data, and transmit the second data to the in-path processing module; The in-path processing module is configured to perform a first type of vector operation on the second data according to the data processing request to obtain third data. When the data processing request indicates to perform a second type of vector operation, the third data is transmitted to the vector processing module. When the data processing request does not indicate to perform a second type of vector operation, the third data is transmitted to the memory; The vector processing module is configured to, when receiving the third data, perform a second type of vector operation on the third data according to the data processing request to obtain fourth data, and transmit the fourth data to the memory.
2. The processor according to claim 1, wherein The in-path processing module includes a connection control unit and a plurality of operation units. The plurality of operation units are configured to perform different vector operations, and the vector operations performed by the plurality of operation units belong to the first type; The connection control unit is configured to determine a target operation unit among the plurality of operation units, and activate the connection between the target operation unit and the connection control unit. The target operation unit is configured to perform the vector operation indicated by the data processing request; The connection control unit is further configured to receive the second data transmitted by the tensor processing module and transmit the second data to the target operation unit; The target operation unit is configured to perform a vector operation on the second data to obtain the third data, and transmit the third data to the connection control unit.
3. The processor according to claim 2, wherein The number of the target operation units is n, where n is an integer greater than 1; The connection control unit is configured to determine the first target operation unit to the nth target operation unit according to the execution order indicated by the data processing request, activate the connection between the first target operation unit and the connection control unit, activate the connection between the ith target operation unit and the (i + 1)th target operation unit, and activate the connection between the nth target operation unit and the connection control unit, where i is a positive integer less than n; The connection control unit is further configured to receive the second data transmitted by the tensor processing module and transmit the second data to the first target operation unit; The (i + 1)th target operation unit is configured to receive the ith intermediate result transmitted by the ith target operation unit, and perform a vector operation on the ith intermediate result to obtain the (i + 1)th intermediate result; The (i + 1)th target operation unit is configured to, when i + 1 is less than n, transmit the (i + 1)th intermediate result to the (i + 2)th target operation unit, and when i + 1 is equal to n, transmit the (i + 1)th intermediate result as the third data to the connection control unit.
4. The processor according to claim 2, wherein The processor further includes a control module, and the in-path processing module further includes a configuration unit, and the configuration unit is connected to the connection control unit; The control module is configured to generate configuration information based on the data processing request, and transmit the configuration information to the configuration unit, and the configuration information includes a target identifier, and the target identifier indicates the target arithmetic unit; The configuration unit is configured to receive the configuration information, and transmit the target identifier in the configuration information to the connection control unit; The connection control unit is configured to determine the target arithmetic unit indicated by the target identifier, and activate the connection between the target arithmetic unit and the connection control unit.
5. The processor according to claim 4, wherein The configuration information further includes configuration parameters corresponding to the target identifier, and the configuration unit is further connected to the plurality of arithmetic units; The configuration unit is further configured to transmit the configuration parameters to the target arithmetic unit indicated by the corresponding target identifier; The target arithmetic unit is configured to perform a vector operation on the second data based on the configuration parameters to obtain the third data, and transmit the third data to the connection control unit.
6. The processor according to claim 2, wherein The target arithmetic unit includes a quantization bias unit; The quantization bias unit is configured to receive configuration parameters and the second data, and the configuration parameters include quantization parameters and bias parameters; The quantization bias unit is configured to multiply the second data by the quantization parameters, and add the result of the multiplication to the bias parameters to obtain the third data.
7. The processor according to claim 2, wherein The target arithmetic unit includes a function operation unit, and the function operation unit is configured to perform a variety of different function operations; The function operation unit is configured to receive configuration parameters and the second data, and the configuration parameters include a function identifier; The function operation unit is configured to query operation coefficients corresponding to a target function indicated by the function identifier, substitute the operation coefficients into a preset function template to obtain a fitting function, and the fitting function is used to fit the target function; The function operation unit is configured to input the second data into the fitting function to obtain the third data.
8. The processor according to claim 2, wherein The target arithmetic unit includes a number system conversion unit, and the number system conversion unit is configured to perform a variety of different number system conversion operations; The number system conversion unit is configured to receive configuration parameters and the second data, and the configuration parameters include a conversion identifier; The number system conversion unit is configured to perform a number system conversion operation indicated by the conversion identifier on the second data to obtain the third data.
9. The processor according to claim 2, wherein The target arithmetic unit includes a rectified linear unit; The rectified linear unit is configured to receive the second data, and perform rectified linear transformation on the second data to obtain the third data.
10. The processor according to claim 2, characterized in that, The connection control unit is configured to receive the third data, and in the case where the data processing request indicates to perform a second type of vector operation, transmit the third data to the vector processing module, and in the case where the data processing request does not indicate to perform a second type of vector operation, transmit the third data to the memory.
11. The processor according to claim 2, wherein, The in-path processing module further includes a reshaping unit, and the reshaping unit is connected to the connection control unit; The connection control unit is configured to receive the third data and pass the third data to the reshaping unit; The reshaping unit is configured to reshape the third data to obtain the reshaped third data. When the data processing request indicates to perform a second type of vector operation, the reshaped third data is passed to the vector processing module. When the data processing request does not indicate to perform a second type of vector operation, the reshaped third data is passed to the memory.
12. A data processing method, characterized in that, Executed by a processor, the processor includes a tensor processing module, a follow - through processing module, and a vector processing module. The tensor processing module is connected to the follow - through processing module, and the follow - through processing module is connected to the vector processing module. The method includes: Through the tensor processing module, perform a matrix operation on the first data according to the data processing request to obtain the second data, and pass the second data to the follow - through processing module; Through the follow - through processing module, perform a first type of vector operation on the second data according to the data processing request to obtain the third data. When the data processing request indicates to perform a second type of vector operation, the third data is passed to the vector processing module. When the data processing request does not indicate to perform a second type of vector operation, the third data is passed to the memory; When the vector processing module receives the third data, through the vector processing module, perform a second type of vector operation on the third data according to the data processing request to obtain the fourth data, and pass the fourth data to the memory.
13. The data processing method according to claim 12, wherein The follow - through processing module includes a connection control unit and a plurality of operation units. The plurality of operation units are configured to perform different vector operations, and the vector operations performed by the plurality of operation units belong to the first type; The step of performing, through the follow - through processing module, a first type of vector operation on the second data according to the data processing request to obtain the third data includes: Through the connection control unit, determine a target operation unit among the plurality of operation units, and activate the connection between the target operation unit and the connection control unit. The target operation unit is configured to perform the vector operation indicated by the data processing request; Through the connection control unit, receive the second data passed by the tensor processing module and pass the second data to the target operation unit; Through the target operation unit, perform a vector operation on the second data to obtain the third data, and pass the third data to the connection control unit.
14. The data processing method according to claim 13, wherein The number of the target operation units is n, where n is an integer greater than 1; Determining a target arithmetic unit among the multiple arithmetic units through the connection control unit and enabling the connection between the target arithmetic unit and the connection control unit includes: through the connection control unit, determining the first target arithmetic unit to the nth target arithmetic unit according to the execution order indicated by the data processing request, enabling the connection between the first target arithmetic unit and the connection control unit, enabling the connection between the ith target arithmetic unit and the (i + 1)th target arithmetic unit, and enabling the connection between the nth target arithmetic unit and the connection control unit, where i is a positive integer less than n; Receiving second data passed in by the tensor processing module through the connection control unit and passing the second data into the target arithmetic unit includes: through the connection control unit, receiving the second data passed in by the tensor processing module and passing the second data into the first target arithmetic unit; Performing a vector operation on the second data through the target arithmetic unit to obtain the third data and passing the third data into the connection control unit includes: through the (i + 1)th target arithmetic unit, receiving the ith intermediate result passed in by the ith target arithmetic unit, performing a vector operation on the ith intermediate result to obtain the (i + 1)th intermediate result; when i + 1 is less than n, passing the (i + 1)th intermediate result into the (i + 2)th target arithmetic unit through the (i + 1)th target arithmetic unit, and when i + 1 is equal to n, passing the (i + 1)th intermediate result as the third data into the connection control unit through the (i + 1)th target arithmetic unit.
15. The data processing method according to claim 13, wherein The processor further includes a control module, and the in-line processing module further includes a configuration unit, and the configuration unit is connected to the connection control unit; the method further includes: Generating configuration information based on the data processing request through the control module and passing the configuration information into the configuration unit, where the configuration information includes a target identifier that indicates the target arithmetic unit; Receiving the configuration information through the configuration unit and passing the target identifier in the configuration information into the connection control unit; Determining a target arithmetic unit among the multiple arithmetic units through the connection control unit and enabling the connection between the target arithmetic unit and the connection control unit includes: Determining the target arithmetic unit indicated by the target identifier through the connection control unit and enabling the connection between the target arithmetic unit and the connection control unit.
16. The data processing method according to claim 15, wherein The configuration information further includes configuration parameters corresponding to the target identifier, and the configuration unit is further connected to the multiple arithmetic units; the method further includes: The configuration unit is further configured to pass the configuration parameters into the target arithmetic unit indicated by the corresponding target identifier; Performing a vector operation on the second data through the target arithmetic unit to obtain the third data and passing the third data into the connection control unit includes: Through the target operation unit, perform vector operations on the second data based on the configuration parameters to obtain the third data, and transmit the third data to the connection control unit.
17. The data processing method according to claim 13, wherein The target operation unit includes a quantization bias unit; Performing vector operations on the second data through the target operation unit to obtain the third data includes: Through the quantization bias unit, receive the configuration parameters and the second data, where the configuration parameters include quantization parameters and bias parameters; Through the quantization bias unit, multiply the second data by the quantization parameters, and add the result of the multiplication to the bias parameters to obtain the third data.
18. The data processing method according to claim 13, wherein The target operation unit includes a function operation unit, and the function operation unit is used to perform various different function operations; performing vector operations on the second data through the target operation unit to obtain the third data includes: Through the function operation unit, receive the configuration parameters and the second data, where the configuration parameters include a function identifier; Through the function operation unit, query the operation coefficients corresponding to the target function indicated by the function identifier, substitute the operation coefficients into a preset function template to obtain a fitting function, and the fitting function is used to fit the target function; Through the function operation unit, input the second data into the fitting function to obtain the third data.
19. The data processing method according to claim 13, characterized in that The target operation unit includes a number system conversion unit, and the number system conversion unit is used to perform various different number system conversion operations; performing vector operations on the second data through the target operation unit to obtain the third data includes: Through the number system conversion unit, receive the configuration parameters and the second data, where the configuration parameters include a conversion identifier; Through the number system conversion unit, perform the number system conversion operation indicated by the conversion identifier on the second data to obtain the third data.
20. A computer device, characterized in that, The computer device includes a processor, and the processor is used to implement the operations performed by the processor according to any one of claims 1 to 11.