Processor, data processing method, and computer device

By dispersing the operation submodules in the processor and introducing the design of interactive submodules, the problem of executing module area expansion and layout and routing difficulties is solved, and a more efficient processor design is achieved.

WO2025077427A9PCT designated stage expired Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/112218
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-11
Filing Date
2024-08-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The area expansion of the execution module in the processor leads to increased difficulty in layout and routing.

Method used

The sharding processing and interaction of data is realized by dispersing multiple operational submodules in the subprocessor on multiple execution modules and deploying interactive submodules in each execution module.

Benefits of technology

The area of ​​the execution module is reduced, the difficulty of layout and routing is reduced, and the design efficiency of the processor is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024112218_30052025_PF_FP_ABST
    Figure CN2024112218_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, and discloses a processor, a data processing method, and a computer device. The processor comprises a sub-processor, the sub-processor comprises a memory and a plurality of execution modules connected to each other, and each execution module comprises an interaction sub-module and a plurality of operation sub-modules; the memory is used for dividing first data into K pieces of first sub-data, and transmitting the K pieces of first sub-data into K operation sub-modules, respectively; the operation sub-modules are used for performing operation on the first sub-data transmitted from the memory to obtain intermediate sub-data, and transmitting the intermediate sub-data into the interaction sub-module in the same execution module; and the interaction sub-module is used for receiving the intermediate sub-data transmitted from the operation sub-modules, transmitting the intermediate sub-data into other interaction sub-modules, receiving the intermediate sub-data transmitted from the other interaction sub-modules, and processing the obtained plurality of pieces of intermediate sub-data to obtain second data. The present invention can reduce the area of the execution modules, and reduce the difficulty of layout and wiring of the execution modules.
Need to check novelty before this filing date? Find Prior Art

Description

Processor, data processing method and computer equipment

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on October 11, 2023, with application number 2023113157880 and application name “Processor, Data Processing Method and Computer Device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of computer technology, and in particular to processor technology and data processing technology. Background Art

[0003] A processor typically includes multiple different types of sub-processors, each of which typically includes multiple computing units. These multiple computing units can perform operations on data in parallel, thereby ensuring the parallelism of the sub-processors. These multiple computing units are arranged on an execution module, which is the smallest module for layout and routing of the sub-processor. Since multiple computing units in the sub-processors need to be arranged on one execution module, the execution module needs to stack a large number of computing units, which leads to the expansion of the execution module area and makes layout and routing in the execution module more difficult.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a processor, a data processing method, and a computer device that can reduce the area of ​​an execution module and reduce the difficulty of layout and routing of the execution module. The technical solution is as follows:

[0006] In one aspect, a processor is provided, comprising a sub-processor, the sub-processor comprising a memory and a plurality of execution modules connected to each other, the execution module comprising an interaction sub-module and a plurality of operation sub-modules, the interaction sub-module being connected to the plurality of operation sub-modules; the total number of the operation sub-modules in the sub-processor is K, where K is an integer greater than 1;

[0007] The memory is configured to divide the first data into K first sub-data and transmit the K first sub-data to the K operation sub-modules respectively;

[0008] The operation submodule is configured to operate on the first subdata inputted from the memory to obtain intermediate subdata, and to input the intermediate subdata to the interaction submodule in the same execution module;

[0009] The interaction sub-module is used to receive the intermediate sub-data transmitted by the operation sub-module in the same execution module, and transmit the intermediate sub-data to the interaction sub-module in different execution modules; is used to receive the intermediate sub-data transmitted by the interaction sub-module in different execution modules; and is used to process the received multiple intermediate sub-data to obtain second data.

[0010] In another aspect, a data processing method is provided, the method being executed by a processor, the processor comprising a sub-processor, the sub-processor comprising a memory and a plurality of execution modules connected to each other, the execution module comprising an interaction sub-module and a plurality of operation sub-modules, the interaction sub-module being connected to the plurality of operation sub-modules; the total number of the operation sub-modules in the sub-processor being K, where K is an integer greater than 1; the method comprising:

[0011] Dividing the first data into K first sub-data through the memory, and respectively transmitting the K first sub-data to the K operation sub-modules;

[0012] The operation submodule operates on the first subdata input from the memory to obtain intermediate subdata, and transmits the intermediate subdata to the interaction submodule in the same execution module;

[0013] Through the interaction sub-module, the intermediate sub-data transmitted from the operation sub-module in the same execution module is received, the intermediate sub-data is transmitted to the interaction sub-module in different execution modules, the intermediate sub-data transmitted from the interaction sub-module in different execution modules is received, and the received multiple intermediate sub-data are processed to obtain the second data.

[0014] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, the at least one computer program is loaded and executed by the processor, and the processor is used for the above-mentioned processing to execute the above-mentioned data processing method.

[0015] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor, which is the above-mentioned processor, to implement the above-mentioned data processing method.

[0016] On the other hand, a computer program product is provided, comprising a computer program, wherein the computer program is loaded and executed by a processor, wherein the processor is the above-mentioned processor, to implement the above-mentioned data processing method.

[0017] The solution provided by the embodiment of the present application is that the sub-processor includes multiple execution modules, and each execution module includes multiple operation sub-modules, which is equivalent to deploying all the operation sub-modules in the sub-processor on different execution modules respectively, and the memory divides the first data into multiple first sub-data, and respectively transmits them to the operation sub-modules of different execution modules for operation. In addition, an interactive sub-module is also deployed on the execution module, and the intermediate sub-data obtained by the operation sub-modules in different execution modules can be concentrated in the interactive sub-module, and the interactive sub-module performs overall processing on the intermediate sub-data obtained by the operation sub-modules of all the operation sub-modules, thereby obtaining the second data. Since the multiple operation sub-modules are dispersed on multiple different execution modules in the present application, the number of operation sub-modules required to be laid out on an execution module is greatly reduced, which is conducive to reducing the area of ​​the execution module and reducing the difficulty of layout and wiring of the execution module. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] FIG1 is a schematic diagram of the structure of a processor provided in an embodiment of the present application;

[0020] FIG2 is a schematic diagram of the structure of another processor provided in an embodiment of the present application;

[0021] FIG3 is a schematic diagram of a flow chart of a shift operation provided in an embodiment of the present application;

[0022] FIG4 is a schematic diagram of the structure of a sub-processor provided in an embodiment of the present application;

[0023] FIG5 is a schematic diagram of the structure of an execution module provided in an embodiment of the present application;

[0024] FIG6 is a schematic diagram of the structure of another execution module provided in an embodiment of the present application;

[0025] FIG7 is a schematic diagram of the structure of another processor provided in an embodiment of the present application;

[0026] FIG8 is a schematic diagram of the structure of another sub-processor provided in an embodiment of the present application;

[0027] FIG9 is a schematic structural diagram of an interaction submodule provided in an embodiment of the present application;

[0028] FIG10 is a flow chart of a data processing method provided in an embodiment of the present application;

[0029] FIG11 is a schematic structural diagram of a terminal provided in an embodiment of the present application;

[0030] FIG12 is a schematic diagram of the structure of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0032] It is understood that the terms "first," "second," and the like used herein may be used to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are used solely to distinguish one concept from another. For example, a first execution module may be referred to as a second execution module, and similarly, a second execution module may be referred to as a first execution module, without departing from the scope of this application.

[0033] Here, "at least one" refers to one or more than one. For example, the at least one execution module can be one execution module, two execution modules, three execution modules, or any other integer greater than or equal to one. "Multiple" refers to two or more than two. For example, the multiple execution modules can be two execution modules, three execution modules, or any other integer greater than or equal to two. "Each" refers to each of at least one or more. For example, "each execution module" refers to each execution module in the multiple execution modules. If the multiple execution modules are three execution modules, "each execution module" refers to each execution module in the three execution modules.

[0034] It should be noted that the information (including but not limited to user device information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this application are all fully authorized by users or relevant parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0035] An embodiment of the present application provides a processor, which includes a sub-processor, which includes a memory and multiple execution modules connected to each other, each execution module includes multiple operation sub-modules and interaction sub-modules, and the interaction sub-module is connected to the multiple operation sub-modules. In some embodiments, the processor is set in a computer device. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, smart voice interaction device, smart home appliance and car terminal, etc., but is not limited to this.

[0036] FIG1 is a schematic diagram of the structure of a processor provided in an embodiment of the present application. As shown in FIG1 , the processor includes a sub-processor, which includes a memory 10 and multiple interconnected execution modules 20. Each execution module 20 includes an interaction sub-module 202 and multiple operation sub-modules 201, and the interaction sub-module 202 is connected to the multiple operation sub-modules 201.

[0037] Among them, the processor can be an AI (Artificial Intelligence) processor or an AI chip. The AI ​​processor is also called an AI accelerator or AI computing card. The AI ​​processor is a processor for processing computing tasks in the field of artificial intelligence. Among them, the sub-processor can be a vector operator sub-processor in the AI ​​processor. The vector operator sub-processor provides hardware support for single instruction multiple data operations (SIMD). Parallelism is used to describe the number of vector data that can be calculated in one cycle. The higher the parallelism, the higher the computing power of the vector operator sub-processor, but at the same time, the more hardware computing resources need to be stacked in the vector operator sub-processor, which may cause the area of ​​the chip in the vector operator sub-processor to expand sharply.

[0038] The execution module 20 in the sub-processor is a Harden, which refers to the smallest module for layout and routing in the back-end process. The memory 10 in the sub-processor is interconnected with each execution module 20 respectively, and any two execution modules 20 in the multiple execution modules 20 are interconnected, and any two execution modules 20 are connected to each other through their respective interaction sub-modules 202. The memory 10 can be a vector L1 memory. Each execution module 20 includes multiple operation sub-modules 201 and interaction sub-modules 202. In an execution module 20, the interaction sub-module 202 is interconnected with each operation sub-module in the execution module 20. Among them, the interaction sub-module 202 can be called CROSS, which is a component for data interaction between different modules.

[0039] In the related art, a sub-processor includes multiple operation sub-modules, and multiple operation sub-modules are integrated into one execution module. When the number of integrated operation sub-modules is large, the area of ​​the execution module will expand rapidly, and the difficulty of layout and routing of the execution module will also increase. The embodiment of the present application provides a solution for splitting the execution module, splitting the original execution module in the sub-processor into multiple execution modules, each of which includes multiple operation sub-modules. By splitting multiple execution modules, the number of operation sub-modules required to be deployed on each execution module can be reduced, which is conducive to reducing the area of ​​the execution module, and can reduce the complexity of layout and routing of the back-end process, simplify the back-end layout and routing process, and alleviate resource occupation pressure.

[0040] 1 , when a sub-processor includes K (K is an integer greater than 1) operation sub-modules 201, that is, when the total number of operation sub-modules 201 included in all execution modules 20 in the sub-processor is K, the memory 10 is configured to divide the first data into K first sub-data and transmit the K first sub-data to the K operation sub-modules 201, that is, the memory 10 transmits one first sub-data to each operation sub-module 201. For example, if the sub-processor has a processing bit number of 4096 bits, the operation sub-module 201 has a processing bit number of 256 bits, and the sub-processor includes a total of 16 operation sub-modules 201, if the first data has a bit number of 4096 bits, the memory 10 may divide the 4096-bit first data into 16 256-bit first sub-data, and transmit one 256-bit first sub-data to each operation sub-module 201, respectively.

[0041] 1 , the operation submodule 201 is configured to operate on the first subdata input from the memory 10 to obtain intermediate subdata, and then transmit the intermediate subdata to the interaction submodule 202 within the same execution module 20. For any operation submodule 201, after the memory 10 transmits the first subdata to the operation submodule 201, the operation submodule 201 operates on the first subdata to obtain intermediate subdata, and then transmits the intermediate subdata to the interaction submodule 202 within the same execution module 20 as the operation submodule 201.

[0042] Referring to Figure 1 , the interaction submodule 202 is configured to receive intermediate subdata transmitted from the operation submodule 201 in the same execution module 20 and transmit the intermediate subdata to the interaction submodule 202 in a different execution module 20; further configured to receive intermediate subdata transmitted from the interaction submodule 202 in a different execution module; and further configured to process the received multiple intermediate subdata to obtain second data. Specifically, for any interaction submodule 202 in an execution module 20, the interaction submodule 202 receives intermediate subdata transmitted from the operation submodule 201 in the same execution module 20, as well as intermediate subdata transmitted from interaction submodules 202 in different execution modules 20. The interaction submodule 202 transmits the intermediate subdata transmitted from the operation submodule 201 in the execution module 20 to which it belongs to the interaction submodule 202 in a different execution module 20. That is, the interaction submodule 202 in each execution module 20 can obtain the intermediate subdata provided by all operation submodules 201. The interaction submodule 202 processes the received multiple intermediate subdata to obtain second data. Since the second data is obtained by processing the intermediate sub-data corresponding to each first sub-data in the first data, that is, the operation process of the entire first data is assigned to multiple operation sub-modules 201 for execution, and finally the interactive sub-module 202 processes the intermediate sub-data obtained by the multiple operation sub-modules 201 to obtain the second data, the second data can be regarded as the data obtained by operating the entire first data.

[0043] In summary, in the solution provided by the embodiment of the present application, the sub-processor includes multiple execution modules, and each execution module includes multiple operation sub-modules, which is equivalent to deploying all the operation sub-modules in the sub-processor on different execution modules respectively, and the memory divides the first data into multiple first sub-data, and respectively transmits them to the operation sub-modules of different execution modules for operation. In addition, an interactive sub-module is also deployed on the execution module, and the intermediate sub-data obtained by the operation sub-modules in different execution modules can be concentrated in the interactive sub-module, and the interactive sub-module performs overall processing on the intermediate sub-data obtained by the operation sub-modules of all the operation sub-modules, thereby obtaining the second data. Since the multiple operation sub-modules are dispersed on multiple different execution modules in the present application, the number of operation sub-modules required to be laid out on an execution module is greatly reduced, which is conducive to reducing the area of ​​the execution module and reducing the difficulty of layout and wiring the execution module.

[0044] In some embodiments, referring to FIG2 , the operator module 201 includes a first operation unit 211 and a register 221. The register 221 is connected to the first operation unit 211, and the memory 10 is connected to the register 221. The first operation unit 211 may be a vector arithmetic logic unit (VALU), and the register 221 may be a vector register file (VRF).

[0045] Referring to Figure 2, the memory 10 is used to transfer each of the K first sub-data to the register 221 in the corresponding operation sub-module. The register 221 is used to cache the first sub-data. The first operation unit 211 is used to read the first sub-data from the register 221, perform operations on the first sub-data, and obtain intermediate sub-data. After the memory 10 divides the first data into K first sub-data, each first sub-data corresponds to an operation sub-module 201. For any first sub-data, the memory 10 can use a Load instruction (data loading instruction) to load the first sub-data into the register 221 of the corresponding operation sub-module 201, and the register 221 caches the first sub-data. In the same operation sub-module 201, the first operation unit 211 can read the first sub-data from the register 221.

[0046] In the related art, a register is deployed in the sub-processor, and each operation sub-module reads data from the same register. In the embodiment of the present application, the register is also distributedly split, and the number of split registers is equal to the number of first operation units in the sub-processor, that is, each first operation unit corresponds to a register, and each first operation unit and the corresponding register constitute an operation sub-module. Each first operation unit can read data from its corresponding register, so that the register and the first operation unit can be tightly coupled, and each first operation unit is closer to its corresponding register, which is conducive to reducing the timing pressure of the first operation unit reading data or writing back data when executing internal operation logic.

[0047] In one possible implementation, the length of the first data is M bits, the length of the first sub-data is N bits, the ratio of M to N is equal to K, M and N are both integers greater than 1, and N is less than M. The memory 10 is used to transfer P first sub-data of consecutive bit lengths to the registers 221 of P operation sub-modules 201 in the same execution module 20, where P is an integer greater than 1, that is, the number of first sub-data loaded in an execution module is equal to the number of operation sub-modules included in the execution module. For example, if the length of the first data is 4096 bits, K is 16, the length of the first sub-data is 256 bits, and the number of execution modules is two, the memory 10 can transfer the first sub-data corresponding to bits 1 to 2048 of the first data to the registers 221 of multiple operation sub-modules 201 in the first execution module 20, and transfer the first sub-data corresponding to bits 2049 to 4096 of the first data to the registers 221 of multiple operation sub-modules 201 in the second execution module 20.

[0048] In an embodiment of the present application, the number of bits of multiple first sub-data loaded into the registers 221 of multiple operation sub-modules 201 in the same execution module 20 is continuous, so as to ensure that the number of bits of the first sub-data processed by multiple operation sub-modules 201 in the same execution module 20 is continuous, thereby reducing the burden of data interaction.

[0049] In one possible implementation, the interaction sub-module 202 is also used to obtain an indication signal, which is used to indicate the target number of bits that the execution module 20 where the interaction sub-module 202 is located needs to process, and to determine the second sub-data at the target number of bits in the second data; the interaction sub-module 202 is also used to, when the second data is the processing result of the first data, write the second sub-data back to the register 221 in the execution module 20 where the interaction sub-module 202 is located; when the second data is not the processing result of the first data, pass the second sub-data into the first operation unit 211 in the execution module 20 where the interaction sub-module 202 is located.

[0050] Since the sub-processor includes multiple execution modules 20, each execution module 20 only needs to process part of the first sub-data in the first data. The number of bits corresponding to the first sub-data that each execution module 20 needs to process is different, so it is necessary to use an indication signal to distinguish which bits each execution module needs to process. For example, the sub-processor includes two execution modules 20, the number of bits of the first data is 4096 bits, the number of bits that one execution module 20 needs to process is the 1st to the 2048th bit, and the number of bits that the other execution module 20 needs to process is the 2049th to the 4096th bit. When the value of the indication signal obtained by the interaction sub-module 202 is 0, it indicates that the target number of bits that the execution module 20 where the interaction sub-module is located needs to process is the 1st to the 2048th bit. When the value of the indication signal obtained by the interaction sub-module 202 is 1, it indicates that the target number of bits that the execution module 20 where the interaction sub-module is located needs to process is the 2049th to the 4096th bit.

[0051] After the interaction submodule 202 obtains the second data, if the second data is the result of processing the first data, it indicates that the first data has been processed and no further processing of the second data is required. The interaction submodule 202 writes back the second subdata at the target bit number in the second data to the register 221 in the execution module 20 where the interaction submodule 202 is located. The register 221 caches the second subdata. Subsequently, the memory 10 reads and stores the second subdata in the register 221 through the Store instruction (data save instruction). For example, if the number of bits to be processed by the execution module 20 where the interaction submodule 202 is located is the 2049th to the 4096th bit, the interaction submodule 202 writes back the second subdata at the 2049th to the 4096th bit in the second data to the register 221. Optionally, the interaction sub-module 202 writes the second sub-data back to the register 221 in the execution module 20 where the interaction sub-module 202 is located, which means that after dividing the second sub-data into multiple pieces, the second sub-data are written back to the register 221 in each operation sub-module 201 in the execution module 20, respectively, and the number of bits of the second sub-data written back to each register 221 is the same as the number of bits of the first sub-data cached in the register 221. The second data is the processing result of the first data, which can be understood as the second data being the final processing result of the request to process the first data. For example, when a shift operation is requested to be performed on the first data, the second data is the result of the shift operation on the first data, and thus the second data is the processing result of the first data.

[0052] After the interaction submodule 202 obtains the second data, if the second data is not the result of processing the first data, it indicates that the first data has not been processed completely and needs to be processed again. The interaction submodule 202 writes the second subdata at the target bit number in the second data back to the first operation unit 211 in the execution module 20 where the interaction submodule 202 is located. The first operation unit 211 then continues to operate on the second subdata. If the operation result does not need to be interacted with the operation results of other first operation units 211 again, the first operation unit 211 writes the operation result back to the register 221. If the operation result needs to be interacted with the operation results of other first operation units 211 again, the first operation unit 211 transfers the operation result to the interaction submodule 202. For example, if the number of bits to be processed by the execution module 20 where the interaction submodule 202 is located is bits 2049 to 4096, the interaction submodule 202 transfers the second subdata at bits 2049 to 4096 in the second data to the first operation unit 211. Optionally, the interaction sub-module 202 transmitting the second sub-data to the first operation unit 211 in the execution module 20 where the interaction sub-module 202 is located means dividing the second sub-data into multiple pieces and transmitting them respectively to the first operation unit 211 in each operation sub-module 201 in the execution module 20, and the number of bits of the second sub-data transmitted to each first operation unit 211 is the same as the number of bits of the first sub-data transmitted to the first operation unit 211. The fact that the second data is not the processing result of the first data can be understood as meaning that the second data is not the final processing result of the request to process the first data.

[0053] In an embodiment of the present application, after obtaining the second data, the interaction sub-module only transfers the sub-data on the target number of bits in the second data to the operation sub-module (into the first operation unit or register). The target number of bits is the number of bits that the execution module where the interaction sub-module is located is responsible for, thereby ensuring that each operation sub-module is only responsible for processing data on a specific number of bits, avoiding errors in the operation process.

[0054] In some embodiments, referring to FIG2 , the interaction submodule 202 includes a splicing unit 212, a second operation unit 222, an input port 232, and an output port 242. The input port 232 is used to receive intermediate sub-data inputted from the interaction submodules 202 in different execution modules 20; the splicing unit 212 is used to splice the intermediate sub-data inputted from the operation submodule 201 in the same execution module 20 with the intermediate sub-data inputted from the interaction submodules 202 in different execution modules 20 to obtain intermediate data; the second operation unit 222 is used to operate on the intermediate data to obtain second data; and the output port 242 is used to transmit the intermediate sub-data inputted from the operation submodule 201 in the same execution module 20 to the interaction submodule 202 in different execution modules 20.

[0055] Among them, the input port 232 in each interaction sub-module 202 is connected to the output port 242 in the interaction sub-module 202 in different execution modules 20. Through the connection between the input port 232 and the output port 242, the interaction sub-modules 202 in different execution modules 20 can transmit intermediate sub-data to each other, that is, the interaction sub-module 202 in one execution module 20 can obtain the intermediate sub-data provided by the interaction sub-module 202 in another execution module 20.

[0056] In this way, by splitting the interaction sub-module into a splicing unit, a second operation unit, an input port and an output port, the interaction sub-module can use the corresponding units or ports to perform corresponding tasks, ensuring the reliable execution of the operation tasks and data transmission tasks in the interaction sub-module.

[0057] In one possible implementation, the splicing unit in the interaction sub-module 202 further obtains an indication signal, which is used to indicate the target number of bits to be processed by the execution module 20 in which the interaction sub-module 202 is located. Based on the target number of bits, the splicing unit splices the intermediate sub-data input from the operation sub-module 201 in the same execution module 20 and the intermediate sub-data input from the interaction sub-module 202 in different execution modules 20 to obtain intermediate data, so that the intermediate sub-data input from the operation sub-module 201 in the same execution module 20 is within the target number of bits of intermediate data.

[0058] Since each execution module 20 is responsible for a different number of bits, when splicing the intermediate sub-data provided by the operation sub-modules 201 in multiple execution modules 20, it is necessary to consider the number of bits of each intermediate sub-data in the first data. Therefore, it is necessary to use an indication signal to distinguish which bits each execution module 20 needs to process. For example, the sub-processor includes two execution modules 20, the number of bits of the first data is 4096 bits, the number of bits that one execution module 20 needs to process is the 1st to the 2048th bit, and the number of bits that the other execution module 20 needs to process is the 2049th to the 4096th bit. The intermediate sub-data provided by the current execution module 20 obtained by the splicing unit is the intermediate sub-data a. The intermediate sub-data provided by the other execution module 20 is the intermediate sub-data b. Then, when the value of the indication signal obtained by the splicing unit is 0, it means that the target number of bits that the execution module 20 where the interactive sub-module is located needs to process is from the 1st to the 2048th bit, so the intermediate sub-data a is from the 1st to the 2048th bit, and the intermediate sub-data b is from the 2049th to the 4096th bit, so the intermediate data obtained by splicing is {intermediate sub-data a, intermediate sub-data b}. When the value of the indication signal obtained by the splicing unit is 1, it means that the target number of bits that the execution module 20 where the splicing unit is located needs to process is from the 2049th to the 4096th bit, so the intermediate sub-data a is from the 2049th to the 4096th bit, and the intermediate sub-data b is from the 1st to the 2048th bit, so the intermediate data obtained by splicing is {intermediate sub-data b, intermediate sub-data a}.

[0059] Therefore, the splicing unit splices the intermediate sub-data transmitted from the operation sub-module in the same execution module and the intermediate sub-data transmitted from different execution modules based on the obtained indication signal, so as to ensure that the spliced ​​intermediate data accurately corresponds to the input first data, that is, each bit in the intermediate data accurately corresponds to each bit in the first data.

[0060] In one possible implementation, the second operation unit 222 determines a target number of shifts in response to a shift instruction for the intermediate data, performs a first shift on the intermediate data, and obtains the third data. The second operation unit 222 obtains an indication signal indicating a target number of bits to be processed by the execution module 20 where the interaction sub-module 202 resides, determines the third sub-data within the target number of bits in the third data, and shifts the third sub-data until the current number of shifts reaches the target number of shifts, thereby obtaining the second data.

[0061] Since the interaction sub-module 202 only needs to input the data corresponding to the number of bits that the current execution module 20 is responsible for into the operation sub-module 201, in order to reduce the amount of calculation of the interaction sub-module 202, in the second operation unit 222 of the interaction sub-module 202, only the data on the target number of bits that the current execution module 20 is responsible for can be processed, thereby reducing the amount of calculation of the shift operation, which is conducive to improving processing efficiency.

[0062] Taking the shift instruction as an example, the intermediate data is shifted in units of byte, and the shift information is a 9-bit array, which is used to indicate whether the shift is performed this time. As shown in Figure 3, (1) in Figure 3 is the shift process provided by the relevant technology, and (2) in Figure 3 is the shift process provided by the embodiment of the present application. Referring to (1) in Figure 3, in the relevant technology, according to the 9-bit shift information, 9 shift operations are performed respectively to obtain shift data 8, which is the second data obtained by shifting the intermediate data, and then the second sub-data is screened out in the shift data 8 according to the target number of bits indicated by the indicator signal. Since the shift operation needs to be performed on the complete shift data each time, the amount of calculation is large and the processing efficiency is low. Referring to (2) in FIG3 , in the embodiment of the present application, the first shift operation is first performed according to the 9th bit shift information [8] in the shift information to obtain shift data 0, and then according to the target number of bits indicated by the indication signal, the shift sub-data 0 is screened out from the shift data 0. The shift data 0 is also the third data mentioned above. The shift sub-data 0 is also the third sub-data corresponding to the target number of bits mentioned above. Then, according to the remaining 8 bits of shift information in the shift information, 8 shift operations are performed on the third sub-data to obtain shift sub-data 8. The shift sub-data 8 is also the second data. Among them, the second data is the same as the second sub-data in (1), so there is no need to screen the second data again. Since, except for the first shift operation, the remaining shift operations are all only for shifting the data on the target number of bits that the current execution module is responsible for, the amount of operation of the shift operation is reduced, which is conducive to improving processing efficiency.

[0063] In one possible implementation, the multiple execution modules 20 in the sub-processor include a first execution module 20a and a second execution module 20b. The first execution module 20a includes an interaction sub-module 202a, and the second execution module 20b includes an interaction sub-module 202b. To reduce the difficulty of layout and routing, the first execution module 20a and the second execution module 20b are arranged in a left-right mirrored manner. Referring to FIG4 , the first input port 232a and the first output port 242a of the interaction sub-module 202a are located on the right side, with the first input port 232a located above the first output port 242a. The second input port 232b and the second output port 242b of the interaction sub-module 202b are located on the left side, with the second input port 232b located above the second output port 242b. Consequently, the wiring between the first input port 232a and the second output port 242b, and the wiring between the first output port 242a and the second input port 232b, will intersect.

[0064] In some embodiments, the interaction sub-module 202 is located in the central area of ​​the execution module 20 , and the multiple operation sub-modules 201 in the execution module 20 are distributed around the interaction sub-module 202 .

[0065] Optionally, the subprocessor includes two execution modules, and the number of operation submodules in the subprocessor is 16. In this case, each execution module includes an interaction submodule and eight operation submodules. Referring to Figures 5 and 6 , the execution module includes operation submodules 0 through 7 and an interaction submodule. The interaction submodule is located in the center of the execution module, and operation submodules 0 through 7 are distributed around the interaction submodule.

[0066] In the embodiment of the present application, since each operation sub-module needs to interact with the interaction sub-module for data, the interaction sub-module is arranged in the central area of ​​the execution module, and the operation sub-module is arranged around the interaction sub-module to minimize the distance between the operation sub-module and the interaction sub-module, which helps to reduce the wiring pressure between the operation sub-module and the interaction sub-module.

[0067] Furthermore, each register in the operator module includes a load port and a store port. The load port is used to load data from the memory into the register, and the store port is used to store data from the register into the memory. When placing and routing the registers, the load port and store port in the register are placed close together to reduce the routing pressure on the load port and store port.

[0068] The embodiments of the present application provide a solution for splitting an execution module. When splitting an execution module, one issue that needs to be considered is the number of execution modules to be split into. If the execution module is not split, the actual number of minimum units in an execution module is between 8000k and 9000k, while the expected number of minimum units in an execution module is between 4000k and 5000k. Therefore, the execution module can be split into two to meet the expected number of minimum units in an execution module.

[0069] FIG7 is a schematic diagram of the structure of another processor provided in an embodiment of the present application. As shown in FIG7 , the processor includes a sub-processor, which includes a memory 10 and a first execution module 20a and a second execution module 20b connected to each other. The first execution module 20a includes multiple operation sub-modules 201a and an interaction sub-module 202a, and the interaction sub-module 202a is connected to the multiple operation sub-modules 201a. The second execution module 20b includes multiple operation sub-modules 201b and an interaction sub-module 202b, and the interaction sub-module 202b is connected to the multiple operation sub-modules 201b.

[0070] The interaction submodule 202a includes a splicing unit, a second operation unit, a first input port, and a first output port. The interaction submodule 202b includes a splicing unit, a second operation unit, a second input port, and a second output port. The first input port is connected to the second output port, and the second input port is connected to the first output port.

[0071] The memory 10 divides the first data into K first sub-data, and transmits the K first sub-data to K operation sub-modules respectively, where K is the total number of operation sub-modules in the sub-processor, and K is an integer greater than 1.

[0072] In the first execution module 20a, the operation submodule 201a in the first execution module 20a operates on the first subdata input from the memory 10 to obtain intermediate subdata, which are then passed to the interaction submodule 202a in the first execution module 20a. The first input port in the interaction submodule 202a receives the intermediate subdata input from the interaction submodule 202b in the second execution module 20b via a connection between the first input port and the second output port. The concatenation unit in the interaction submodule 202a concatenates the intermediate subdata input from the operation submodule 201a with the intermediate subdata input from the interaction submodule 202b to obtain intermediate data. The second operation unit in the interaction submodule 202a operates on the intermediate data to obtain second data. The first output port in the interaction submodule 202a transmits the intermediate subdata input from the operation submodule 201a to the interaction submodule 202b in the second execution module 20b via a connection between the first output port and the second input port.

[0073] In the second execution module 20b, the operation submodule 201b in the second execution module 20b operates on the first subdata input from the memory 10 to obtain intermediate subdata, which are then passed to the interaction submodule 202b in the second execution module 20b. The second input port in the interaction submodule 202b receives the intermediate subdata input from the interaction submodule 202a in the first execution module 20a through its connection to the first output port. The concatenation unit in the interaction submodule 202b concatenates the intermediate subdata input from the operation submodule 201b with the intermediate subdata input from the interaction submodule 202a to obtain intermediate data. The second operation unit in the interaction submodule 202b operates on the intermediate data to obtain second data. The second output port in the interaction submodule 202b transmits the intermediate subdata input from the operation submodule 201b to the interaction submodule 202a in the first execution module 20a through its connection to the first input port.

[0074] In one possible implementation, referring to FIG8 , the first execution module 20a and the second execution module 20b have the same structure. The first input port includes a first left input port and a first right input port, and the first output port includes a first left output port and a first right output port. The first left output port is located below the first left input port, and the first right input port is located below the first right output port. The second input port includes a second left input port and a second right input port, and the second output port includes a second left output port and a second right output port. The second left output port is located below the second left input port, and the second right input port is located below the second right output port. The first right input port is connected to the second left output port, and the second left input port is connected to the first right output port.

[0075] For the interaction sub-module 202a in the first execution module 20a, the first right input port is used to receive the intermediate sub-data transmitted from the second left output port of the interaction sub-module 202b in the second execution module 20b through the connection between itself and the second left output port; the first right output port is used to transmit the intermediate sub-data of the operation sub-module 201a received by the interaction sub-module 202a to the second left input port of the interaction sub-module 202b in the second execution module 20b through the connection between the second left input port and itself.

[0076] For the interaction sub-module 202b in the second execution module 20b, the second left input port is used to receive the intermediate sub-data transmitted from the first right output port of the interaction sub-module 202a in the first execution module 20a through the connection between itself and the first right output port; the second left output port is used to transmit the intermediate sub-data of the operation sub-module 201b received by the interaction sub-module 202b to the first right input port of the interaction sub-module 202a in the first execution module 20a through the connection between the first right input port and itself.

[0077] In a possible implementation, the interaction submodule 202a and the interaction submodule 202b further include an OR operation unit.

[0078] Referring to Figures 8 and 9, for the interaction submodule 202a in the first execution module 20a, the first left input port transmits the input preset signal to the OR operation unit, where the value of the preset signal is 0. The first right input port receives the intermediate sub-data input from the second left output port of the interaction submodule 202b in the second execution module 20b, and transmits the intermediate sub-data input from the second left output port to the OR operation unit. The OR operation unit performs an OR operation on the preset signal and the input intermediate sub-data, and transmits the OR-operated intermediate sub-data to the splicing unit. Since the value of the preset signal is 0, the data obtained after the OR operation unit performs an OR operation on the preset signal and the input intermediate sub-data is still intermediate sub-data.

[0079] Referring to Figures 8 and 9, for the interaction submodule 202b in the second execution module 20b, the second right input port transmits the input preset signal to the OR operation unit, where the value of the preset signal is 0. The second left input port receives the intermediate sub-data input from the first right output port of the interaction submodule 202a in the first execution module 20a, and transmits the intermediate sub-data input from the first right output port to the OR operation unit. The OR operation unit performs an OR operation on the preset signal and the input intermediate sub-data, and transmits the OR-operated intermediate sub-data to the splicing unit. Since the value of the preset signal is 0, the data obtained after the OR operation unit performs an OR operation on the preset signal and the input intermediate sub-data is still the intermediate sub-data.

[0080] In one possible implementation, the interaction submodule 202a and the interaction submodule 202b further include a replication unit.

[0081] Referring to Figures 8 and 9 , for the interaction sub-module 202a in the first execution module 20a, the replication unit replicates the intermediate sub-data inputted by the operation sub-module 201a, generating two intermediate sub-data. One intermediate sub-data is transmitted to the first left output port, and the other intermediate sub-data is transmitted to the first right output port. The first left output port is open-connected, meaning that the intermediate sub-data is not outputted to other components via the first left output port. The intermediate sub-data can be transmitted to the second left input port via the first right output port.

[0082] Referring to Figures 8 and 9 , for the interaction sub-module 202b in the second execution module 20b, the replication unit replicates the intermediate sub-data input from the operation sub-module 201b, generating two intermediate sub-data. One intermediate sub-data is transmitted to the second right output port, and the other intermediate sub-data is transmitted to the second left output port. The second right output port is a no-connect. That is, the intermediate sub-data is not output to other components via the second right output port, but can be transmitted to the first right input port via the second left output port.

[0083] In the embodiment of the present application, as shown in Figures 8 and 9, although a set of redundant input ports and output ports are added to the execution module, the structural consistency of the first execution module and the second execution module is maintained, and the data interaction ports between the first execution module and the second execution module are aligned, the connection is simple, and no wiring crossover occurs. The first execution module and the second execution module can be placed back to back and then directly connected, avoiding the consumption of area by the channel occupied by the wiring, and further reducing the complexity of layout and wiring.

[0084] The processor provided in the embodiments of the present application splits the execution module into two structurally identical first and second execution modules, reducing the complexity of the back-end synthesis process's layout and routing. Furthermore, because the first and second execution modules have identical structures, a single Harden instance can be used to generate both the first and second execution modules, further simplifying the layout and routing process and improving its efficiency.

[0085] FIG10 is a flow chart of a data processing method provided by an embodiment of the present application. The embodiment of the present application is executed by a processor, which includes a sub-processor, which includes a memory and multiple execution modules connected to each other. The execution module includes an interaction sub-module and multiple operation sub-modules, and the interaction sub-module is connected to the multiple operation sub-modules. The total number of operation sub-modules in the sub-processor is K, where K is an integer greater than 1. The processor may be an AI processor or an AI chip, etc. Referring to FIG10, the method includes:

[0086] 1001. Divide first data into K first sub-data via a memory, and transmit the K first sub-data to K operation sub-modules respectively.

[0087] 1002. Perform operation on the first sub-data inputted from the memory through the operation sub-module to obtain intermediate sub-data, and input the intermediate sub-data to the interaction sub-module in the same execution module.

[0088] 1003. Receive the intermediate sub-data transmitted from the operation sub-module in the same execution module through the interaction sub-module, and transmit the intermediate sub-data to the interaction sub-module in different execution modules, receive the intermediate sub-data transmitted from the interaction sub-module in different execution modules, process the received multiple intermediate sub-data, and obtain the second data.

[0089] In the method provided by the embodiment of the present application, the sub-processor includes multiple execution modules, and each execution module includes multiple operation sub-modules, which is equivalent to deploying all the operation sub-modules in the sub-processor on different execution modules respectively, and the memory divides the first data into multiple first sub-data, and respectively transmits them to the operation sub-modules of different execution modules for operation. In addition, an interactive sub-module is also deployed on the execution module, and the intermediate sub-data obtained by the operation sub-modules in different execution modules can be concentrated in the interactive sub-module, and the interactive sub-module performs overall processing on the intermediate sub-data obtained by the operation sub-modules of all the operation sub-modules, thereby obtaining the second data. Since the multiple operation sub-modules are dispersed on multiple different execution modules in the present application, the number of operation sub-modules required to be laid out on an execution module is greatly reduced, which is conducive to reducing the area of ​​the execution module and reducing the difficulty of layout and wiring of the execution module.

[0090] In one possible implementation, the operation submodule includes a first operation unit and a register, the register being connected to the first operation unit, and the memory being connected to the register. The first data is divided into K first sub-data via the memory, and the K first sub-data are respectively transmitted to the K operation submodules, including: transmitting each first sub-data via the memory to the register in the corresponding operation submodule. The operation submodule performs operations on the first sub-data transmitted from the memory to obtain intermediate sub-data, including: caching the first sub-data via the register; and reading the first sub-data from the register via the first operation unit, performing operations on the first sub-data to obtain intermediate sub-data.

[0091] In one possible implementation, the length of the first data is M bits, the length of the first sub-data is N bits, the ratio of M to N is equal to K, M and N are both integers greater than 1, and N is less than M. Loading each first sub-data into a corresponding register via a memory includes: transferring P first sub-data having consecutive bits into registers of P operation sub-modules in the same execution module via the memory, where P is an integer greater than 1.

[0092] In one possible implementation, the method further includes: obtaining an indication signal through the interaction sub-module, the indication signal being used to indicate the target number of bits that the execution module where the interaction sub-module is located needs to process, and determining the second sub-data at the target number of bits in the second data; through the interaction sub-module, if the second data is the processing result of the first data, writing the second sub-data back to the register in the execution module where the interaction sub-module is located; if the second data is not the processing result of the first data, passing the second sub-data into the first operation unit in the execution module where the interaction sub-module is located.

[0093] In one possible implementation, the interaction submodule includes a splicing unit, a second operation unit, an input port, and an output port. The interaction submodule receives intermediate subdata from an operation submodule in the same execution module, transmits the intermediate subdata to an interaction submodule in a different execution module, receives intermediate subdata from interaction submodules in different execution modules, and processes the received multiple intermediate subdata to obtain second data, including:

[0094] The intermediate sub-data transmitted from the interactive sub-modules in different execution modules are received through the input port; the intermediate sub-data transmitted from the operation sub-module in the same execution module and the intermediate sub-data transmitted from the interactive sub-modules in different execution modules are spliced ​​through the splicing unit to obtain intermediate data; the intermediate data are operated on by the second operation unit to obtain second data; and the intermediate sub-data transmitted from the operation sub-module in the same execution module are transmitted to the interactive sub-modules in different execution modules through the output port.

[0095] In one possible implementation, the multiple execution modules include a first execution module and a second execution module; the interaction submodule in the first execution module includes a first input port and a first output port, and the interaction submodule in the second execution module includes a second input port and a second output port, wherein the first input port is connected to the second output port, and the second input port is connected to the first output port. Receiving intermediate sub-data from interaction submodules in different execution modules via the input port includes: receiving intermediate sub-data from the interaction submodule in the second execution module via the first input port based on the connection between the first input port and the second output port. Transmitting intermediate sub-data from the operation submodule in the same execution module to the interaction submodule in a different execution module via the output port includes: transmitting intermediate sub-data to the interaction submodule in the second execution module via the first output port based on the connection between the first output port and the second input port.

[0096] In one possible implementation, the first execution module and the second execution module have the same structure, wherein the first input port includes a first left input port and a first right input port, the first output port includes a first left output port and a first right output port, the first left output port is located below the first left input port, and the first right input port is located below the first right output port; the second input port includes a second left input port and a second right input port, the second output port includes a second left output port and a second right output port, the second left output port is located below the second left input port, and the second right input port is located below the second right output port; the first right input port is connected to the second left output port, and the second left input port is connected to the first right output port. Receiving intermediate sub-data inputted by the interaction sub-module in the second execution module via the first input port based on the connection between the first input port and the second output port includes: receiving intermediate sub-data inputted by the second left output port in the second execution module via the first right input port based on the connection between the first right input port and the second left output port. The intermediate sub-data is transmitted to the interactive sub-module in the second execution module through the first output port based on the connection between the first output port and the second input port, including: the intermediate sub-data is transmitted to the second left input port in the second execution module through the first right output port based on the connection between the second left input port and the first right output port.

[0097] In one possible implementation, the interaction submodule further includes an operation unit. The method further includes:

[0098] The preset signal is input into the or operation unit through the first left input port, and the value of the preset signal is 0;

[0099] The intermediate sub-data input from the second left output port is input to the OR operation unit through the first right input port; the OR operation unit performs an OR operation on the preset signal and the intermediate sub-data, and inputs the intermediate sub-data after the OR operation to the splicing unit.

[0100] In one possible implementation, the interaction submodule further includes a replication unit. The method further includes:

[0101] The intermediate sub-data inputted by the operation sub-module is copied through the copy unit to obtain two intermediate sub-data, one of which is inputted into the first left output port, and the other is inputted into the first right output port; wherein the first left output port is an empty connection.

[0102] In one possible implementation, the method further includes:

[0103] An indication signal is obtained through a splicing unit, the indication signal being used to indicate a target number of bits to be processed by the execution module where the interaction sub-module is located. The intermediate sub-data inputted from the operation sub-module in the same execution module and the intermediate sub-data inputted from the interaction sub-module in different execution modules are spliced ​​through the splicing unit to obtain intermediate data, including: splicing through the splicing unit, based on the target number of bits, the intermediate sub-data inputted from the operation sub-module in the same execution module and the intermediate sub-data inputted from the interaction sub-module in different execution modules to obtain the intermediate data, wherein the intermediate sub-data inputted from the operation sub-module is within the target number of bits of the intermediate data.

[0104] In one possible implementation, the intermediate data is operated on by the second operation unit to obtain the second data, including: determining the target number of shifts by the second operation unit in response to a shift instruction for the intermediate data, performing a first shift on the intermediate data to obtain third data; obtaining an indication signal by the second operation unit, the indication signal being used to indicate the target number of bits to be processed by the execution module where the interactive sub-module is located, and determining the third sub-data at the target number of bits in the third data; shifting the third sub-data by the second operation unit until the current number of shifts reaches the target number of shifts to obtain the second data.

[0105] In a possible implementation, the interaction submodule is located in the central area of ​​the execution module, and the multiple operation submodules in the execution module are distributed around the interaction submodule.

[0106] It should be noted that the data processing method provided in the embodiment of the present application and the processor provided in the above embodiment belong to the same inventive concept. The specific implementation method of the data processing method provided in the embodiment of the present application can refer to the embodiment of the above processor, and the embodiment of the present application will not be repeated here.

[0107] An embodiment of the present application also provides a computer device, which includes a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the processor of the above embodiment. The structure of the processor is the processor structure in the above embodiment.

[0108] Optionally, the computer device is provided as a terminal. FIG11 shows a schematic diagram of the structure of a terminal 1100 provided in an exemplary embodiment of the present application. The terminal 1100 includes: a processor 1101 and a memory 1102.

[0109] The processor 1101 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0110] The memory 1102 may include one or more computer-readable storage media, which may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more magnetic disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1102 is used to store at least one computer program, which is used to be stored by the processor 1101 to implement the various embodiments described above.

[0111] In some embodiments, terminal 1100 may also optionally include a peripheral device interface 1103 and at least one peripheral device. Processor 1101, memory 1102, and peripheral device interface 1103 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 1103 via a bus, signal lines, or circuit boards. Optionally, the peripheral device includes at least one of a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, and a power supply 1108.

[0112] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0113] The RF circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1104 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1104 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the RF circuit 1104 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The RF circuit 1104 can communicate with other devices via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, metropolitan area networks, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1104 may also include circuits related to Near Field Communication (NFC), which is not limited in this application.

[0114] The display screen 1105 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch screen display, the display screen 1105 also has the ability to collect touch signals on the surface or above the surface of the display screen 1105. The touch signal can be input as a control signal to the processor 1101 for processing. At this time, the display screen 1105 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there can be one display screen 1105, which is set on the front panel of the terminal 1100; in other embodiments, there can be at least two display screens 1105, which are respectively set on different surfaces of the terminal 1100 or in a folding design; in other embodiments, the display screen 1105 can be a flexible display screen, which is set on the curved surface or folding surface of the terminal 1100. Even more, the display screen 1105 can be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 1105 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0115] The camera assembly 1106 is used to capture images or videos. Optionally, the camera assembly 1106 includes a front camera and a rear camera. The front camera is arranged on the front panel of the terminal 1100, and the rear camera is arranged on the back of the terminal 1100. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1106 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0116] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 1101 for processing, or input into the radio frequency circuit 1104 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 1100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as distance measurement. In some embodiments, the audio circuit 1107 may also include a headphone jack.

[0117] Power supply 1108 is used to power various components in terminal 1100. Power supply 1108 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1108 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0118] Those skilled in the art will understand that the structure shown in FIG11 does not constitute a limitation on the terminal 1100 , and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0119] Optionally, the computer device is provided as a server. Figure 12 is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server 1200 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1201 and one or more memories 1202, wherein the memory 1202 stores at least one computer program, and the at least one computer program is loaded and executed by the processor 1201 to implement the above-mentioned embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server may also include other components for realizing device functions, which will not be described in detail here.

[0120] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one computer program. The at least one computer program is loaded and executed by a processor to implement the operations performed by the processor of the above embodiment. The structure of the processor is the processor structure in the above embodiment.

[0121] An embodiment of the present application also provides a computer program product, including a computer program, which is loaded and executed by a processor to implement the operations performed by the processor in the above embodiment, and the structure of the processor is the processor structure in the above embodiment.

[0122] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0123] The above description is only an optional embodiment of the embodiment of the present application and is not intended to limit the embodiment of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiment of the present application should be included in the scope of protection of the present application.

Claims

1. A processor, the processor comprising a subprocessor, the subprocessor comprising a memory and a plurality of execution modules connected to each other, the execution module comprising an interaction submodule and a plurality of operation submodules, the interaction submodule being connected to the plurality of operation submodules; the total number of the operation submodules in the subprocessor is K, and K is an integer greater than 1; The memory is used to divide the first data into K first sub-data, and transfer the K first sub-data to the K operation sub-modules respectively; The operation submodule is used to operate the first subdata transmitted from the memory to obtain intermediate subdata, and transmit the intermediate subdata to the interaction submodule in the same execution module; The interaction submodule is used to receive the intermediate subdata transmitted by the operation submodule in the same execution module, and transmit the intermediate subdata to the interaction submodule in the different execution module; Used to receive the intermediate sub-data transmitted by the interactive sub-modules in different execution modules; used to process the received multiple intermediate sub-data to obtain second data.

2. The processor according to claim 1, wherein the operation submodule comprises a first operation unit and a register, the register is connected to the first operation unit, and the memory is connected to the register; The memory is used to transfer each of the first sub-data to a register in the corresponding operation sub-module; The register is used to cache the first sub-data; The first operation unit is used to read the first sub-data from the register, and perform operation on the first sub-data to obtain the intermediate sub-data.

3. The processor according to claim 2, wherein the length of the first data is M bits, the length of the first sub-data is N bits, the ratio of M to N is equal to K, M and N are both integers greater than 1, and N is less than M; The memory is used to transfer P first sub-data with consecutive bits to the registers of P operation sub-modules in the same execution module respectively, where P is an integer greater than 1.

4. The processor according to claim 2 or 3, wherein the interaction submodule is further used to obtain an indication signal, wherein the indication signal is used to indicate the target number of bits to be processed by the execution module where the interaction submodule is located, and to determine the second subdata at the target number of bits in the second data; The interaction sub-module is also used to, when the second data is the processing result of the first data, write the second sub-data back to the register in the execution module where the interaction sub-module is located; and when the second data is not the processing result of the first data, pass the second sub-data to the first operation unit in the execution module where the interaction sub-module is located.

5. The processor according to any one of claims 1 to 4, wherein the interaction submodule comprises a splicing unit, a second operation unit, an input port and an output port; The input port is used to receive the intermediate sub-data transmitted by the interaction sub-modules in different execution modules; The splicing unit is used to splice the intermediate sub-data transmitted by the operation sub-module in the same execution module and the intermediate sub-data transmitted by the interaction sub-module in different execution modules to obtain intermediate data; The second operation unit is used to operate on the intermediate data to obtain the second data; The output port is used to transmit the intermediate sub-data transmitted by the operation sub-module in the same execution module to the interaction sub-module in the different execution module.

6. The processor according to any one of claims 1 to 5, wherein the plurality of execution modules include a first execution module and a second execution module; the interaction submodule in the first execution module includes a first input port and a first output port, the interaction submodule in the second execution module includes a second input port and a second output port, the first input port is connected to the second output port, and the second input port is connected to the first output port; The first input port is used to receive the intermediate sub-data transmitted by the interaction sub-module in the second execution module through the connection between the first input port and the second output port; The first output port is used to transfer the intermediate sub-data to the interaction sub-module in the second execution module through the connection between the first output port and the second input port.

7. The processor according to claim 6, wherein the first execution module and the second execution module have the same structure, the first input port includes a first left input port and a first right input port, the first output port includes a first left output port and a first right output port, the first left output port is located below the first left input port, and the first right input port is located below the first right output port; the second input port includes a second left input port and a second right input port, the second output port includes a second left output port and a second right output port, the second left output port is located below the second left input port, and the second right input port is located below the second right output port; the first right input port is connected to the second left output port, and the second left input port is connected to the first right output port; The first right input port is used to receive the intermediate sub-data transmitted from the second left output port in the second execution module through the connection between the first right input port and the second left output port; The first right output port is used to transfer the intermediate sub-data to the second left input port in the second execution module through the connection between the second left input port and the first right output port.

8. The processor according to claim 7, wherein the interaction submodule further comprises an OR operation unit; The first left input port is used to transmit an input preset signal to the OR operation unit, and the value of the preset signal is 0; The first right input port is used to transfer the intermediate sub-data transferred from the second left output port to the OR operation unit; The OR operation unit is used to perform an OR operation on the preset signal and the intermediate sub-data, and transfer the intermediate sub-data after the OR operation to the splicing unit in the interaction sub-module.

9. The processor according to claim 7 or 8, wherein the interaction submodule further comprises a replication unit; The copying unit is used to copy the intermediate sub-data transmitted by the operation sub-module to obtain two intermediate sub-data, transmit one of the intermediate sub-data to the first left output port, and transmit the other intermediate sub-data to the first right output port; wherein, The first left output port is an empty connection.

10. The processor according to any one of claims 5 to 9, wherein the splicing unit is further used to obtain an indication signal, wherein the indication signal is used to indicate the target number of bits that the execution module where the interaction submodule is located needs to process; The splicing unit is used to splice the intermediate sub-data transmitted by the operation sub-module in the same execution module and the intermediate sub-data transmitted by the interaction sub-module in different execution modules based on the target number of bits to obtain the intermediate data, wherein the intermediate sub-data transmitted by the operation sub-module is within the target number of bits of the intermediate data.

11. The processor according to any one of claims 5 to 10, wherein the second operation unit is configured to determine a target shifting number in response to a shift instruction for the intermediate data, and perform a first shift on the intermediate data to obtain third data; The second operation unit is used to obtain an indication signal, the indication signal is used to indicate the target number of bits to be processed by the execution module where the interaction submodule is located, and determine the third sub-data at the target number of bits in the third data; The second operation unit is used to shift the third sub-data until the current shift times reaches the target shift times to obtain the second data.

12. According to the processor according to any one of claims 1 to 11, the interaction submodule is located in the central area of ​​the execution module, and the multiple operation submodules in the execution module are distributed around the interaction submodule.

13. A data processing method, executed by a processor, the processor comprising a subprocessor, the subprocessor comprising a memory and a plurality of execution modules connected to each other, the execution module comprising an interaction submodule and a plurality of operation submodules, the interaction submodule being connected to the plurality of operation submodules; The total number of the operation submodules in the subprocessor is K, and K is an integer greater than 1; the method comprises: Dividing the first data into K first sub-data through the memory, and transmitting the K first sub-data to the K operation sub-modules respectively; The first sub-data transmitted from the memory are operated by the operation sub-module to obtain intermediate sub-data, and the intermediate sub-data are transmitted to the interaction sub-module in the same execution module; Through the interaction sub-module, the intermediate sub-data transmitted from the operation sub-module in the same execution module is received, and the intermediate sub-data is transmitted to the interaction sub-module in different execution modules; the intermediate sub-data transmitted from the interaction sub-module in different execution modules is received, and the received multiple intermediate sub-data are processed to obtain second data.

14. The method according to claim 13, wherein the operation submodule comprises a first operation unit and a register, the register is connected to the first operation unit, and the memory is connected to the register; The first data is divided into K first sub-data by the memory, and the K first sub-data are respectively transmitted to the K operation sub-modules, including: Transferring each of the first sub-data to a register in a corresponding operation sub-module through the memory; The step of performing operation on the first sub-data transmitted from the memory to obtain intermediate sub-data by the operation sub-module includes: caching the first sub-data through the register; The first sub-data is read from the register through the first operation unit, and the first sub-data is operated to obtain the intermediate sub-data.

15. The method according to claim 14, wherein the length of the first data is M bits, the length of the first sub-data is N bits, the ratio of M to N is equal to K, M and N are both integers greater than 1, and N is less than M; The step of transferring each of the first sub-data to a register in a corresponding operation sub-module through the memory includes: Through the memory, P first sub-data with consecutive bits are respectively transferred to the registers of P operation sub-modules in the same execution module, where P is an integer greater than 1.

16. The method according to claim 14 or 15, further comprising: Obtaining an indication signal through the interaction submodule, the indication signal being used to indicate the target number of bits to be processed by the execution module where the interaction submodule is located, and determining second sub-data at the target number of bits in the second data; By means of the interaction submodule, when the second data is a result of processing the first data, writing the second subdata back to a register in the execution module where the interaction submodule is located; In the case where the second data is not a result of processing the first data, the second sub-data is transmitted to the first computing unit in the execution module where the interaction sub-module is located.

17. According to the method described in any one of claims 13 to 16, the interaction submodule comprises a splicing unit, a second operation unit, an input port and an output port; the interaction submodule receives the intermediate subdata transmitted by the operation submodule in the same execution module, and transmits the intermediate subdata to the interaction submodule in a different execution module; Receiving the intermediate sub-data transmitted by the interaction sub-modules in different execution modules, and processing the received plurality of the intermediate sub-data to obtain second data, including: Receiving the intermediate sub-data transmitted by the interaction sub-modules in different execution modules through the input port; The intermediate sub-data transmitted from the operation sub-module in the same execution module and the intermediate sub-data transmitted from the interaction sub-module in different execution modules are spliced ​​by the splicing unit to obtain intermediate data; The second computing unit performs computing on the intermediate data to obtain the second data; The intermediate sub-data transmitted from the operation sub-module in the same execution module are transmitted to the interaction sub-module in the different execution modules through the output port.

18. The method according to any one of claims 13 to 17, wherein the plurality of execution modules include a first execution module and a second execution module; the interaction submodule in the first execution module includes a first input port and a first output port, the interaction submodule in the second execution module includes a second input port and a second output port, the first input port is connected to the second output port, and the second input port is connected to the first output port; The receiving, through the input port, the intermediate sub-data transmitted by the interaction sub-modules in different execution modules includes: Receiving intermediate sub-data transmitted by the interaction sub-module in the second execution module through the first input port based on the connection between the first input port and the second output port; The step of transmitting the intermediate sub-data transmitted by the operation sub-module in the same execution module to the interaction sub-module in the different execution modules through the output port includes: The intermediate sub-data is transmitted to the interaction sub-module in the second execution module through the first output port based on the connection between the first output port and the second input port.

19. The method according to claim 18, wherein the first execution module and the second execution module have the same structure, the first input port includes a first left input port and a first right input port, the first output port includes a first left output port and a first right output port, the first left output port is located below the first left input port, and the first right input port is located below the first right output port; the second input port includes a second left input port and a second right input port, the second output port includes a second left output port and a second right output port, the second left output port is located below the second left input port, and the second right input port is located below the second right output port; the first right input port is connected to the second left output port, and the second left input port is connected to the first right output port; The receiving, through the first input port and based on the connection between the first input port and the second output port, the intermediate sub-data transmitted by the interaction sub-module in the second execution module includes: Receiving, through the first right input port, the intermediate sub-data transmitted from the second left output port in the second execution module based on the connection between the first right input port and the second left output port; The step of transferring the intermediate sub-data to the interaction sub-module in the second execution module through the first output port based on the connection between the first output port and the second input port includes: The intermediate sub-data is transmitted to the second left input port in the second execution module through the first right output port based on the connection between the second left input port and the first right output port.

20. A computer device, comprising a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor, the processor is the processor according to any one of claims 1 to 12, and the processor is used to execute the data processing method according to any one of claims 13 to 19.

21. A computer-readable storage medium, wherein at least one computer program is stored in the computer-readable storage medium, wherein the computer program is loaded and executed by a processor to implement the data processing method according to any one of claims 13 to 19, and the processor is the processor according to any one of claims 1 to 12.

22. A computer program product, comprising a computer program, wherein the computer program is stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device to implement the data processing method according to any one of claims 13 to 19, wherein the processor is the processor according to any one of claims 1 to 12.