Data processing method, device, chip and medium
By setting target circuits between calculation cores in the multi-core chip, obtaining the working state and controlling the synchronization processing of idle cores, the inefficiency problem caused by global synchronization is solved, and more efficient local synchronous data processing is achieved.
Patent Information
- Application Number
- CN202210134715.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-14
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-02-14
AI Technical Summary
The data processing process of the computing core in the multi-core chip is globally synchronized, resulting in low data processing efficiency. If a computing core fails to complete processing, subsequent processing cannot be carried out.
Set up target circuits between calculation cores, obtain the working status of each calculation core through the circuit, determine the calculation core area of the idle state, and control these cores to synchronize data processing to achieve local synchronization.
Improve data processing efficiency, avoid waiting for global synchronization, and enhance the flexibility and efficiency of the processing process.
Smart Images

Figure CN114546640B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of data processing technology, and in particular to a data processing method, device, chip, and medium. Background Art
[0002] With the increasing demand for high-performance computing, the research on multi-core chips, which have a large number of computing cores and strong computing capabilities, has gradually become one of the important research directions in the field of chip research.
[0003] In related technologies, when a multi-core chip processes data through multiple computing cores, the data processing processes of these multiple computing cores are globally synchronized. That is, when processing through these multiple computing cores, it is necessary to wait for the data processing processes of all computing cores to be completed before continuing the subsequent data processing process. Once there is a computing core among these multiple computing cores that has not completed the data processing process, the subsequent data processing process will be unable to proceed, resulting in low data processing efficiency. Summary of the Invention
[0004] To overcome the problems existing in the related art, this specification provides a data processing method, device, terminal and medium.
[0005] According to a first aspect of an embodiment of this specification, there is provided a data processing method applied to a many-core chip, the many-core chip including multiple computing cores, a target circuit being provided between two adjacent computing cores on the many-core chip, the method comprising:
[0006] Based on the target circuits set between the computing cores, obtaining the working status of each computing core;
[0007] Determining, based on the acquired working status of each computing core, that all computing cores in the first target area are in an idle state;
[0008] The computing cores in the first target area are controlled to perform data processing synchronously.
[0009] In some embodiments of this specification, during the process of controlling the computing cores in the first target area to synchronously process data, the method further includes:
[0010] When it is determined that all computing cores in the second target area are in an idle state, controlling the computing cores in the second target area to synchronously process data;
[0011] Furthermore, the computing cores in the second target area partially overlap or do not overlap with the computing cores in the first target area.
[0012] In some embodiments of this specification, determining that all computing cores in the first target area are in an idle state according to the acquired working status of each computing core includes:
[0013] Obtaining the working status of any two computing cores connected to the target circuit through the target circuit;
[0014] When any two connected computing cores are in an idle state, the two connected computing cores are determined as first sub-region cores;
[0015] The plurality of connected computing cores of the first sub-region cores are determined as computing cores within the first target region.
[0016] In some embodiments of the present specification, the target circuit is an OR circuit.
[0017] In some embodiments of the present specification, the OR circuit is an OR circuit including a switch.
[0018] In some embodiments of the present specification, for any one of the two computing cores connected to the target circuit, when the computing core is in an idle state, first indication information is sent to the target circuit; and when the computing core is in a non-idle state, second indication information is sent to the target circuit;
[0019] The first indication information is used to indicate that the computing core is in an idle state, and the second indication information is used to indicate that the computing core is in a non-idle state.
[0020] In some embodiments of the present specification, the method further includes any of the following:
[0021] When the indication information sent by the two computing cores is both the first indication information, determining that the sub-region core formed by the two computing cores is in an idle state;
[0022] When the indication information sent by any computing core is the second indication information, it is determined that the sub-region cores formed by the two computing cores are in a non-idle state.
[0023] According to a second aspect of an embodiment of this specification, a data processing device is provided, which is applied to a many-core chip, wherein a target circuit is provided between two adjacent computing cores on the many-core chip, and the device includes:
[0024] an acquisition module, configured to acquire the working status of each computing core based on a target circuit set between the computing cores;
[0025] a determination module, configured to determine, based on the acquired working status of each computing core, that all computing cores in the first target area are in an idle state;
[0026] The control module is used to control the computing cores in the first target area to synchronously process data.
[0027] In some embodiments of this specification, the control module, while controlling the computing cores in the first target area to synchronously process data, is further configured to:
[0028] When it is determined that all computing cores in the second target area are in an idle state, controlling the computing cores in the second target area to synchronously process data;
[0029] Furthermore, the computing cores in the second target area partially overlap or do not overlap with the computing cores in the first target area.
[0030] In some embodiments of the present specification, the acquisition module, when used to determine that all computing cores in the first target area are in an idle state based on the acquired working status of each computing core, is specifically used to:
[0031] Obtaining the working status of any two computing cores connected to the target circuit through the target circuit;
[0032] When any two connected computing cores are in an idle state, the two connected computing cores are determined as first sub-region cores;
[0033] The plurality of connected computing cores of the first sub-region cores are determined as computing cores within the first target region.
[0034] In some embodiments of the present specification, the target circuit is an OR circuit.
[0035] In some embodiments of the present specification, the OR circuit is an OR circuit including a switch.
[0036] In some embodiments of the present specification, for any one of the two computing cores connected to the target circuit, when the computing core is in an idle state, first indication information is sent to the target circuit; and when the computing core is in a non-idle state, second indication information is sent to the target circuit;
[0037] The first indication information is used to indicate that the computing core is in an idle state, and the second indication information is used to indicate that the computing core is in a non-idle state.
[0038] In some embodiments of this specification, the determining module is further used for any of the following:
[0039] When the indication information sent by the two computing cores is both the first indication information, determining that the sub-region core formed by the two computing cores is in an idle state;
[0040] When the indication information sent by any computing core is the second indication information, it is determined that the sub-region cores formed by the two computing cores are in a non-idle state.
[0041] According to a third aspect of an embodiment of this specification, a multi-core chip is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the operations performed by the above-mentioned data processing method when executing the computer program.
[0042] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a program is stored, and the program is used by a processor to execute the operations performed by the above-mentioned data processing method.
[0043] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program, which implements the operations performed by the above-mentioned data processing method when executed by a processor.
[0044] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:
[0045] By setting a target circuit between two adjacent computing cores, the working status of each computing core can be obtained through the target circuit set between the computing cores. According to the working status of each computing core, when all the computing cores in the first target area are in an idle state, the computing cores in the first target area are controlled to perform data processing synchronously, thereby realizing local synchronization of the computing cores in the many-core chip. Compared with the global synchronization of the computing cores in the many-core chip, the processing process is more flexible, thereby making the data processing efficiency higher.
[0046] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.
[0048] Figure 1 This is a flow chart of a data processing method according to an exemplary embodiment of the present specification.
[0049] Figure 2 This is a schematic diagram of the connection relationship of a computing core according to an exemplary embodiment of this specification.
[0050] Figure 3 This is a schematic diagram of a target circuit to which computing cores at various positions are connected according to an exemplary embodiment of this specification.
[0051] Figure 4 This is a schematic diagram of a computing core included in a many-core chip according to an exemplary embodiment of this specification.
[0052] Figure 5 FIG. 1 is a schematic diagram of a target area division according to an exemplary embodiment of the present specification.
[0053] Figure 6 It is a schematic diagram of a data processing process according to an exemplary embodiment of this specification.
[0054] Figure 7 It is a block diagram of a data processing device according to an exemplary embodiment of the present specification. DETAILED DESCRIPTION
[0055] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of this specification, as detailed herein.
[0056] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. The singular forms "a," "the," and "the" used in this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0057] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information without departing from the scope of this specification. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."
[0058] This application provides a data processing method that can be applied to a many-core chip, so that the many-core chip can process the data to be processed using the data processing method provided by this application. The many-core chip is a chip that includes multiple computing cores, so that the many-core chip can use these multiple computing cores to process the data to be processed.
[0059] Optionally, the many-core chip is a computing chip, or a sensing chip, etc. This application does not limit the chip type of the many-core chip. The many-core chip can be used in various types of computer devices, for example, the many-core chip can be used in a server, or the many-core chip can be used in a terminal, such as a desktop computer, a portable computer, a tablet computer, a smartphone, a smartwatch, etc. This application does not limit this either.
[0060] The above is an introduction to the application scenarios of this application. Next, the data processing method provided by this application is described in detail in conjunction with the embodiments of this specification.
[0061] like Figure 1 As shown, Figure 1 This is a flow chart of a data processing method according to an exemplary embodiment of this specification, which is applied to a many-core chip. The many-core chip includes multiple computing cores. A target circuit is provided between two adjacent computing cores on the many-core chip. The method includes the following steps:
[0062] Step 101: Based on the target circuits set between the computing cores, obtain the working status of each computing core.
[0063] Among them, the computing core is used to execute the computing process corresponding to the neurons included in the neural network mapped in the multi-core chip. One computing core can process one or more neurons, or multiple computing cores can process one neuron, thereby realizing the processing of the data to be processed.
[0064] It should be noted that if a computing core is executing the computing process mapped in the corresponding neuron, that is, it is processing the data to be processed, then the computing core is in a non-idle state (or working state), and if the computing core is not executing the computing process at this time, that is, it is not processing the data to be processed, then the computing core is in an idle state.
[0065] Step 102: Determine, based on the acquired working status of each computing core, whether all computing cores in the first target area are in an idle state.
[0066] Step 103: Control the computing cores in the first target area to perform data processing synchronously.
[0067] Optionally, the computing core in the first target area can implement the data processing process through various types of processing operations, wherein the processing operation can be a convolution operation, a mapping operation, etc., and the present application does not limit the specific type of the processing operation.
[0068] In addition, the types of processing operations performed by the computing cores in the first target area may be the same or different, and this application does not impose any limitation on this.
[0069] The data processing method provided by the present application sets a target circuit between two adjacent computing cores so as to obtain the working status of each computing core through the target circuit set between the computing cores. Thus, according to the working status of each computing core, when all computing cores in the first target area are in an idle state, the computing cores in the first target area are controlled to synchronously process data, thereby achieving local synchronization of computing cores in a multi-core chip. Compared with the global synchronization of computing cores in a multi-core chip, the processing process is more flexible, thereby making data processing more efficient. In addition, the data processing method provided by the present application does not require routing communication through the multi-core chip, but directly utilizes the set target circuit to achieve synchronization of local computing cores on the multi-core chip by setting a dedicated circuit. Compared with the routing synchronization method, it is more direct and efficient.
[0070] After introducing the basic implementation process of the present application, various non-limiting implementation methods of the present application are described in detail below.
[0071] above Figure 1 The embodiment shown is explained by taking the example of setting a target circuit between two adjacent computing cores. In more possible implementations, the target circuit can also be set according to a set interval, that is, a target circuit can be set every set number of computing cores, for example, a target circuit can be set every 2 computing cores.
[0072] Regardless of the method used to set the target circuit, it can be used to obtain the working status of the computing core. In some embodiments, for the above step 101, when obtaining the working status of each computing core based on the target circuit set between the computing cores, the many-core chip can send a control instruction to each computing core after receiving the data to be processed, so that each computing core sends indication information to its corresponding target circuit based on the received control instruction. This indication information can be used to indicate the working status of the computing core, so that the target circuit can determine the working status of the two computing cores to which it is connected based on the received indication information after receiving the indication information sent by the two computing cores.
[0073] Optionally, each computing core may also proactively send indication information to its corresponding target circuit. For example, each computing core may send indication information indicating its own working status to its corresponding target circuit at preset intervals.
[0074] The indication information may include first indication information and second indication information. The first indication information may be used to indicate that the computing core is in an idle state, and the second indication information may be used to indicate that the computing core is in a non-idle state.
[0075] In some embodiments, an OR circuit can be used as the target circuit, and the connection relationship between the two computing cores can be seen in Figure 2 , Figure 2 This is a schematic diagram of the connection relationship of a computing core shown in this specification according to an exemplary embodiment. Two computing cores are connected through an OR circuit. The two computing cores can send indication information indicating the working status to the OR circuit based on their own working status, so that the OR circuit can determine the working status of the two computing cores based on the received indication information.
[0076] It should be noted that the computing cores included in the many-core chip can be distributed on the many-core chip in the form of a computing core array. For computing cores at different positions in the many-core chip, the number of target circuits connected to them is different. Figure 3 , Figure 3 This is a schematic diagram of target circuits connected to computing cores at various positions according to an exemplary embodiment of the present specification. For computing cores in the middle of the array, such computing cores can be connected to four OR circuits; for computing cores at the edge of the array, such computing cores can be connected to three OR circuits; and for computing cores at the vertex of the array, such computing cores can be connected to two OR circuits.
[0077] The following combination Figure 4 , further explain the target circuit connection of the computing core at each location, Figure 4 is a schematic diagram of a computing core included in a many-core chip according to an exemplary embodiment of this specification. Figure 4The 9 computing cores in the array form a 3*3 computing core array, wherein computing core 405 is the computing core in the middle of the array, and computing core 405 is connected to four target circuits, namely target circuit 3, target circuit 4, target circuit 8 and target circuit 11; computing cores 402, 404, 406 and 408 are the computing cores at the edge of the array, and computing cores 402, 404, 406 and 408 are connected to three target circuits respectively, wherein the three target circuits connected to computing core 402 are target circuit 1, target circuit 2 and target circuit 8, the three target circuits connected to computing core 404 are target circuit 3, target circuit 7 and target circuit 10, and the three target circuits connected to computing core 406 are target circuit 1, target circuit 2 and target circuit 8. The three target circuits connected to computing core 408 are target circuit 5, target circuit 6 and target circuit 11 respectively; computing core 401, computing core 403, computing core 407 and computing core 409 are computing cores at the vertices of the array, and computing core 401, computing core 403, computing core 407 and computing core 409 are respectively connected to two target circuits, among which the two target circuits connected to computing core 401 are target circuit 1 and target circuit 7 respectively, the two target circuits connected to computing core 403 are target circuit 2 and target circuit 9 respectively, the two target circuits connected to computing core 407 are target circuit 5 and target circuit 10 respectively, and the two target circuits connected to computing core 409 are target circuit 6 and target circuit 12 respectively.
[0078] When an OR circuit is used as the target circuit, the operating status of the sub-region cores composed of each two computing cores can be determined through the logic of the OR operation, thereby determining whether all computing cores in the first target region are in an idle state. In other words, in step 102 above, when determining that all computing cores in the first target region are in an idle state based on the acquired operating status of each computing core, the following steps can be included:
[0079] Step 1021: Obtain the working status of any two computing cores connected to the target circuit through the target circuit.
[0080] For any one of the two computing cores connected to the target circuit, when the computing core is in an idle state, first indication information is sent to the target circuit; when the computing core is in a non-idle state, second indication information is sent to the target circuit.
[0081] In one possible implementation, 0 can be used as the first indication information and 1 can be used as the second indication information. For any computing core, when the computing core is in an idle state, 0 can be sent to its corresponding target circuit, so that the target circuit can determine that the computing core is in an idle state based on the received 0; when the computing core is in a non-idle state (that is, a working state), 1 can be sent to its corresponding target circuit, so that the target circuit can determine that the computing core is in a non-idle state based on the received 1.
[0082] Step 1022: When any two connected computing cores are in an idle state, determine the two connected computing cores as first sub-region cores.
[0083] In one possible implementation, if the indication information sent by any two connected computing cores is the first indication information, the sub-region cores formed by these two computing cores are determined to be in an idle state. Still taking 0 as an example, if the first indication information is represented by 0, both computing cores connected to the target circuit output 0, then the sub-region cores formed by these two computing cores can be determined to be in an idle state.
[0084] In another possible implementation, when the indication information sent by any computing core is the second indication information, it is determined that the sub-region cores formed by the two computing cores are in a non-idle state.
[0085] The indication information sent by any computing core as the second indication information may include the following two situations:
[0086] 1. The instruction information sent by one computing core is the first instruction information, and the instruction information sent by another computing core is the second instruction information;
[0087] 2. The indication information sent by the two computing cores is both the second indication information.
[0088] Still taking 0 to represent the first indication information and 1 to represent the second indication information as an example, if one computing core outputs 0 and the other computing core outputs 1, it can be determined that the sub-region core composed of these two computing cores is in a non-idle state; or, if both computing cores output 1, it can be determined that the sub-region core composed of these two computing cores is in a non-idle state.
[0089] The target circuit may be an OR circuit, for example, an OR circuit including a switch. Optionally, it may also be an OR circuit including an OR gate. The present application does not limit the specific implementation of the OR circuit.
[0090] Step 1023: Determine the computing cores of the multiple connected first sub-region cores as computing cores within the first target region.
[0091] Among them, for any two connected sub-region cores, the two sub-region cores include an identical computing core, that is, when there is an identical computing core in the two computing cores respectively included in the two sub-region cores, the two sub-region cores can be determined to be connected sub-region cores.
[0092] It should be noted that the first target area may be a pre-set area, or may be an area dynamically determined according to restriction conditions.
[0093] When determining the first target area based on the restriction conditions, it can be achieved in the following ways:
[0094] Based on the amount of data to be processed, a target number of computing cores for processing the data to be processed is determined, and an area corresponding to a sub-area core that is in an idle state and includes computing cores that meet the target number is determined as a first target area.
[0095] The data processing capabilities of each computing core are limited, or in other words, the amount of data that each computing core can process at one time has an upper limit. Therefore, when determining the target number, the target number of computing cores for the data to be processed can be determined based on the upper limit of the amount of data that each computing core can process at one time.
[0096] Optionally, the upper limit of the amount of data that can be processed at one time by each computing core is the same, or the upper limit of the amount of data that can be processed at one time by each computing core is different. Based on this, the process of determining the first target area according to the constraint condition can be implemented in the following two specific ways:
[0097] In one possible implementation, if the upper limit of the amount of data that each computing core can process at one time is the same, the amount of data to be processed and the upper limit of the amount of data that each computing core can process at one time can be divided, and the resulting value can be used as the target number of computing cores for processing the data to be processed, and then multiple connected first sub-region cores are determined from the computing cores in the idle state, and the number of computing cores included in the determined first sub-region cores must meet the target number.
[0098] In another possible implementation, if the upper limits of the amount of data that each computing core can process at one time are different, the upper limits of the amount of data that each connected computing core can process at one time can be accumulated until the accumulated data value is greater than or equal to the amount of data to be processed. The number of computing cores corresponding to the accumulated data value is the target number, thereby obtaining a first target area composed of at least one first sub-area core that is connected and includes a number of computing cores that meets the target number.
[0099] For example, the maximum amount of data that any idle first computing core can process at one time can be added to the amount of data that a second computing core connected to the first computing core can process at one time. Similarly, the amount of data that each connected computing core can process at one time can be added together until the accumulated amount of data reaches the amount of data to be processed. This accumulation method ensures that the computing cores that meet the data volume requirement also meet the connection relationship requirement, thereby obtaining a first target region consisting of at least one first sub-region core that is connected and includes a target number of computing cores.
[0100] In one possible implementation, the division of the first target area can be achieved using an OR circuit including a switch. For example, after determining the first sub-area core constituting the first target area, a control instruction can be sent to the corresponding target circuit to control the OR switch included in the target circuit to close or open, thereby achieving control over whether the target circuit is available or unavailable.
[0101] Among them, when sending a control instruction, a first control instruction is sent to the target circuit corresponding to the first sub-region core in an idle state, so that the target circuit corresponding to the first sub-region core in an idle state is controlled to be in an available state through the first control instruction; and a second control instruction is sent to the target circuit corresponding to the first sub-region core in a non-idle state, so that the target circuit corresponding to the first sub-region core in a non-idle state is controlled to be in an unavailable state through the second control instruction.
[0102] See also Figure 5 , Figure 5 is a schematic diagram of a target area division according to an exemplary embodiment of this specification. Figure 5 The many-core chip shown includes 36 computing cores, which are divided into 4 target areas and can process 4 groups of data to be processed respectively. Figure 5 The circles in the figure represent computing cores, and circles with different shading represent computing cores in different target areas. In other words, the computing cores corresponding to circles with different shading constitute different target areas, and the switches included in target circuits 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, and 517 are in the open state, and the switches included in other target circuits in the figure are in the closed state, thereby realizing the division of the four target areas.
[0103] The above process is described using the example of dividing the target area by controlling the closing or opening of switches. In other possible implementations, the target area can be divided using other methods. For example, each target circuit in the many-core chip can be in a closed state, so that the many-core chip sends control information to each target circuit to control the target circuits with closed switches to divide the target area based on the received control information. The control information is used to indicate whether the two computing cores connected to the target circuit can be used as computing cores in the same target area.
[0104] By grouping the computing cores through switches, some of the computing cores in the many-core chip can be organized into a first target area, so that the computing cores in the first target area can process data synchronously without the need for control through routing and main control, thereby reducing the processing pressure of the many-core chip and improving the processing speed of the many-core chip.
[0105] Optionally, the data to be processed can be multiple types of data. For example, the data to be processed can be image data, text data, behavioral data, etc. Specifically, the data to be processed can be image features, text, user behavior data (such as click-through rate, etc.). This application does not limit the specific type of data to be processed.
[0106] In addition, when the first target area is a pre-set area, the set first target area can also be adjusted based on the amount of data to be processed in a manner similar to the above process to ensure that the processing capacity of the computing core included in the first target area can meet the data processing requirements of the data to be processed.
[0107] In some embodiments, during the process of controlling the computing cores in the first target area and synchronously processing data through step 103, the multi-core chip may also receive other data to be processed. At this time, other data to be processed can be processed by other computing cores other than the computing cores included in the first target area.
[0108] In one possible implementation, when it is determined that all computing cores in the second target area are in an idle state, the computing cores in the second target area are controlled to perform data processing synchronously; and the computing cores in the second target area may partially overlap or not overlap with the computing cores in the first target area.
[0109] The second target area may be a pre-set area or an area dynamically determined according to restriction conditions. For an introduction to the second target area, reference may be made to the above introduction to the first target area, which will not be repeated here.
[0110] Optionally, the type of data to be processed processed by the computing cores included in the first target area may be the same as or different from the type of data to be processed processed by the computing cores included in the second target area, and this application does not impose any limitation on this.
[0111] In addition, the type of processing operation performed by the computing cores included in the first target area may be the same as or different from the type of processing operation performed by the computing cores included in the second target area, and this application does not impose any limitation on this.
[0112] It should be noted that the computing cores in the first target area and the computing cores in the second target area can perform calculations simultaneously. After some of the computing cores in the first target area complete their computing tasks, the first target area can release the computing cores that have completed the computing tasks, so that the computing cores that have completed the tasks can be determined as the computing cores included in the second target area, so that the computing tasks corresponding to the second target area can continue to be executed through these computing cores. In other words, the first target area and the second target area may not be fixed, and the division of the areas may be dynamically adjusted as needed during the calculation process. For example, the division of the first target area and the second target area may be dynamically adjusted by increasing or decreasing the number of computing cores included in the area.
[0113] For example, see Figure 6 , Figure 6 This is a schematic diagram of a data processing process according to an exemplary embodiment of the present specification. Taking a many-core chip including 6 computing cores as an example, computing cores 601, 602, 603, 604, 605, and 606 can be used as computing cores in a first target area to synchronously execute computing tasks. When time t1 is reached, the computing tasks of computing cores 601, 602, and 603 have been completed, while the computing tasks of computing cores 604, 605, and 606 have not been completed. At this time, the computing cores included in the first target area can be used as the computing cores in the first target area. Computing cores 601, 602, and 603 are released from the computing core so that they can serve as computing cores in the second target area to synchronously execute the next computing task. At the same time, computing cores 604, 605, and 606 still serve as computing cores in the first target area to execute their unfinished computing tasks. Until time t2 is reached, the computing tasks of computing cores 604, 605, and 606 are completed, and other computing tasks can be continued to be executed by computing cores 604, 605, and 606.
[0114] During the execution of the above tasks, computing cores 601, 602, and 603 in the first target area can continue to execute other data processing tasks after completing the current data processing tasks, without having to wait for computing cores 604, 605, and 606 in the first target area to complete the data processing tasks, thereby realizing local synchronization of the data processing process of the computing cores, so that the many-core chip can start subsequent data processing tasks without waiting for the data processing tasks of all computing cores to be completed, thereby improving the data processing efficiency of the many-core chip and fully utilizing the computing power of the many-core chip.
[0115] Corresponding to the aforementioned method embodiments, this specification also provides embodiments of a device and a chip used therein.
[0116] like Figure 7 As shown, Figure 7 This is a block diagram of a data processing device according to an exemplary embodiment of the present specification. The many-core chip includes multiple computing cores, and a target circuit is provided between two adjacent computing cores on the many-core chip. The data processing device includes:
[0117] An acquisition module 701 is configured to acquire the working status of each computing core based on a target circuit set between the computing cores;
[0118] A determination module 702 is configured to determine, based on the acquired working status of each computing core, that all computing cores in the first target area are in an idle state;
[0119] The control module 703 is used to control the computing cores in the first target area to synchronously process data:
[0120] In some embodiments of this specification, the control module 703, while controlling the computing cores in the first target area to synchronously process data, is further configured to:
[0121] When it is determined that all computing cores in the second target area are in an idle state, controlling the computing cores in the second target area to synchronously process data;
[0122] Furthermore, the computing cores in the second target area partially overlap or do not overlap with the computing cores in the first target area.
[0123] In some embodiments of the present specification, when the acquisition module 701 is used to determine that all computing cores in the first target area are in an idle state based on the acquired working status of each computing core, it is specifically used to:
[0124] Obtaining the working status of any two computing cores connected to the target circuit through the target circuit;
[0125] When any two connected computing cores are in an idle state, the two connected computing cores are determined as first sub-region cores;
[0126] The plurality of connected computing cores of the first sub-region cores are determined as computing cores within the first target region.
[0127] In some embodiments of the present specification, the target circuit is an OR circuit.
[0128] In some embodiments of the present specification, the OR circuit is an OR circuit including a switch.
[0129] In some embodiments of the present specification, for any one of the two computing cores connected to the target circuit, when the computing core is in an idle state, first indication information is sent to the target circuit; and when the computing core is in a non-idle state, second indication information is sent to the target circuit;
[0130] The first indication information is used to indicate that the computing core is in an idle state, and the second indication information is used to indicate that the computing core is in a non-idle state.
[0131] In some embodiments of this specification, the determining module 702 is further configured to:
[0132] When the indication information sent by the two computing cores is both the first indication information, determining that the sub-region core formed by the two computing cores is in an idle state;
[0133] When the indication information sent by any computing core is the second indication information, it is determined that the sub-region cores formed by the two computing cores are in a non-idle state.
[0134] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0135] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0136] The present application also provides a many-core chip, which includes a memory and multiple computing cores, and a computer program stored in the memory and run on the computing cores, wherein the computing cores implement the operations performed by the data processing method provided in any embodiment when executing the program.
[0137] The present application also provides a computer-readable storage medium, which can be in various forms. For example, in different examples, the computer-readable storage medium can be: RAM (Radom Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as a hard disk drive), solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage medium, or a combination thereof. In particular, the computer-readable storage medium can also be paper or other suitable medium capable of printing programs. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the computing core, the data processing method provided in any embodiment of the present application is implemented.
[0138] The present application also provides a computer program product, including a computer program, which implements the data processing method provided in any embodiment of the present application when executed by a computing core.
[0139] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, apparatus, chip, computer-readable storage medium, or computer program product. Thus, one or more embodiments of this specification may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0140] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from the other embodiments. In particular, the chip embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.
[0141] The foregoing description of specific embodiments of this specification is provided. Other embodiments are within the scope of this application. In some cases, the actions or steps described herein may be performed in an order different from that shown in the embodiments and still achieve the desired results. Additionally, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0142] Embodiments of the subject matter and functional operations described in this specification may be implemented in the following: digital electronic circuits, tangibly embodied computer software or firmware, hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier to be executed by a data processing device or to control the operation of a data processing device. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information and transmit it to a suitable receiver device for execution by a data processing device. The computer-readable storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0143] Although this specification includes many specific implementation details, these should not be interpreted as limiting the scope of any invention or the scope of protection claimed, but are mainly used to describe the features of specific embodiments of specific inventions. Certain features described in multiple embodiments within this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may work in certain combinations as described above and even initially claimed as such, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may point to a sub-combination or a variation of the sub-combination.
[0144] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.
[0145] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of this application. In some cases, the actions described in this application can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.
[0146] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the present invention and practicing the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this specification and include common knowledge or customary techniques in the art that are not claimed herein. That is, this specification is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof.
[0147] The above description is only an optional embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.
Claims
1. A data processing method, characterized in that: Applied to a many-core chip, the many-core chip including multiple computing cores, with a target circuit disposed between two adjacent computing cores on the many-core chip, the method comprising: Based on the target circuits set between the computing cores, obtaining the working status of each computing core; determining, based on the acquired working status of each computing core, that all computing cores in the first target area are in an idle state, wherein any two computing cores in the first target area are connected via the target circuit; controlling computing cores within the first target area to synchronously process data; The target circuit is an OR circuit including a switch, and the target circuit is a circuit different from an on-chip router; and controlling the computing cores in the first target area to synchronously perform data processing includes: A first control instruction is sent to the target circuit corresponding to the computing core within the first target area, and a second control instruction is sent to the target circuit corresponding to the computing core outside the first target area; the first control instruction is used to control the target circuit corresponding to the computing core within the first target area to be in an available state, and the second control instruction is used to control the target circuit corresponding to the computing core outside the first target area to be in an unavailable state.
2. The method according to claim 1, characterized in that During the process of controlling the computing cores in the first target area to synchronously process data, the method further includes: When it is determined that all computing cores in the second target area are in an idle state, controlling the computing cores in the second target area to synchronously process data, wherein any two computing cores in the second target area are connected via the target circuit; Furthermore, the computing cores in the second target area may partially overlap or not overlap with the computing cores in the first target area.
3. The method according to claim 1, characterized in that The determining, based on the acquired working status of each computing core, that all computing cores in the first target area are in an idle state includes: Obtaining, through the target circuit, the working status of any two computing cores connected to the target circuit; When any two connected computing cores are in an idle state, the two connected computing cores are determined as first sub-region cores; The computing cores of the plurality of connected first sub-region cores are determined as computing cores within the first target region.
4. The method according to claim 1, wherein For any one of the two computing cores connected to the target circuit, when the computing core is in an idle state, sending first indication information to the target circuit; when the computing core is in a non-idle state, sending second indication information to the target circuit; The first indication information is used to indicate that the computing core is in an idle state, and the second indication information is used to indicate that the computing core is in a non-idle state.
5. The method according to claim 4, characterized in that The method further comprises any of the following: When the indication information sent by the two computing cores is both the first indication information, determining that the sub-region core formed by the two computing cores is in an idle state; In a case where the indication information sent by any computing core is the second indication information, it is determined that the sub-region core composed of the two computing cores is in a non-idle state.
6. A data processing device, characterized in that: Applied to a many-core chip, the many-core chip including multiple computing cores, with a target circuit provided between two adjacent computing cores on the many-core chip, the device comprising: an acquisition module, configured to acquire the working status of each computing core based on a target circuit set between the computing cores; a determining module, configured to determine, based on the acquired working status of each computing core, that all computing cores in a first target area are in an idle state, wherein any two computing cores in the first target area are connected via the target circuit; a control module, configured to control the computing cores in the first target area to synchronously process data; Wherein, the target circuit is an OR circuit including a switch, and the target circuit is a circuit different from an on-chip routing circuit; The data processing device is also used to send a first control instruction to the target circuit corresponding to the computing core within the first target area, and to send a second control instruction to the target circuit corresponding to the computing core outside the first target area; the first control instruction is used to control the target circuit corresponding to the computing core within the first target area to be in an available state, and the second control instruction is used to control the target circuit corresponding to the computing core outside the first target area to be in an unavailable state.
7. A many-core chip, characterized in that: The many-core chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the operations performed by the data processing method according to any one of claims 1 to 5 when executing the program.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, and the processor uses the program to execute the operations performed by the data processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Task allocation method, task allocation device and on-chip network
CN104156267A