Many-core device

By designing autonomous dynamic routing computing cores in the multi-core equipment, the problem of inflexible routing in the existing technology is solved, the processing speed and quality of neural networks are improved, and power consumption is reduced.

CN113836077BActive Publication Date: 2025-06-24LYNXI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111050498.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-08
Publication Date
2025-06-24
Estimated Expiration
2041-09-08

AI Technical Summary

Technical Problem

In the prior art, the calculation inter-core routing paths of neural networks need to be planned in advance and controlled by SoC in real time, resulting in inflexible routing. Due to the processing speed and bandwidth of SoC, it affects the processing speed and processing quality of the neural network, and increases power consumption.

Method used

A multi-core device is designed. Each computing core includes a data processing unit and a routing unit. The routing unit stores a routing index table, obtains index information through processing results, dynamically determines the routing target, and realizes autonomous and dynamic routing.

Benefits of technology

Improve routing flexibility, eliminates the limitations of SoC processing speed and bandwidth, improves the processing speed and processing quality of neural networks, and reduces the power loss due to frequent data transmission between SoC and computing cores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113836077B_ABST
    Figure CN113836077B_ABST
Patent Text Reader

Abstract

A many-core device is disclosed, which includes a plurality of computing cores, and each of the computing cores includes a data processing unit and a routing unit. Among them, the data processing unit is used to receive data to be processed, process the data to be processed, and send the processing result to a routing target, where the routing target is at least one of the plurality of computing cores. A routing index table is stored in the routing unit, and the routing unit is used to obtain first index information according to the processing result, and determine the routing target according to the first index information and the routing index table. The routing index table is used to store at least one second index information and at least one computing core corresponding to each second index information. Thus, autonomous dynamic routing is realized in the many-core structure, and the flexibility of routing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of multi-core data processing, and more particularly, to a many-core device. Background Art

[0002] This section aims to provide background or context for the embodiments of the present disclosure recited in the claims. The description herein is not admitted to be prior art by including it in this section.

[0003] With the booming development of big data and neural network technologies, data processing systems include more and more processing nodes.

[0004] At the hardware implementation level, each neuron node of a neural network can correspond to one or more computing cores, and the routing paths between the computing cores need to be pre-planned and controlled in real time by an SoC (System on Chip), and this total control routing planning method is very inflexible. Summary of the Invention

[0005] In view of this, at least one many-core device is provided in the embodiments of the present disclosure to autonomously and dynamically determine the next routing by each computing core in a many-core structure, implement dynamic routing in the many-core structure, and improve routing efficiency.

[0006] The present disclosure provides a many-core device, including a plurality of computing cores, and each of the computing cores includes a data processing unit and a routing unit; wherein, the data processing unit is configured to receive data to be processed, process the data to be processed, and send a processing result to a routing target, where the routing target is at least one of the plurality of computing cores; a routing index table is stored in the routing unit, and the routing unit is configured to obtain first index information according to the processing result, and determine the routing target according to the first index information and the routing index table, and the routing index table is configured to store at least one second index information and at least one computing core corresponding to each second index information.

[0007] In some embodiments, the routing unit includes a computing subunit, and the computing subunit is configured to: compare the first index information with the at least one second index information in the routing index table; and determine, as the routing target, one or more computing cores corresponding to the one or more second index information having the highest similarity with the first index information among the at least one second index information.

[0008] In some embodiments, the plurality of computing cores are divided into at least one computing core region, and each of the computing core regions includes at least one of the computing cores.

[0009] In some embodiments, each of the computing core regions is configured to process the data to be processed received by the computing core region according to a preset algorithm to obtain a corresponding processing result; wherein, each computing core in the computing core region is configured to process the sub-data to be processed obtained by the computing core according to a preset sub-algorithm to obtain a sub-processing result; the sub-data to be processed is the data to be processed, or a sub-processing result generated by a computing core other than the computing core in the computing core region.

[0010] In some embodiments, the routing index table is divided into a plurality of index blocks, and the index blocks are divided according to the at least one second index information, and different index blocks correspond to different computing core regions.

[0011] In some embodiments, the index blocks are divided by a sorting result obtained by sorting the at least one second index information according to a preset rule.

[0012] In some embodiments, the index blocks are obtained by obtaining the similarity between every two second index information in the at least one second index information and dividing at least two second index information with a similarity higher than a set threshold into the same index block.

[0013] In some embodiments, the routing index table further includes third index information corresponding to the index block, and the third index information is determined according to the second index information in the corresponding index block; the computing subunit is specifically configured to: compare the first index information with the third index information in the routing index table; determine the computing core region corresponding to the index block corresponding to one or more third index information with the highest similarity to the first index information as the routing target.

[0014] In some embodiments, when each computing core sends the processing result to the routing target, the routing target and the computing core sending the processing result are in the same computing core region or in different computing core regions.

[0015] In some embodiments, the computing subunit includes a computing module implemented by hardware.

[0016] In some embodiments, the processing result includes the first index information.

[0017] In the embodiments of the present disclosure, a many-core device includes a plurality of computing cores, and each of the computing cores includes a data processing unit and a routing unit; wherein, the data processing unit is configured to receive data to be processed, process the data to be processed, and send a processing result to a routing target, where the routing target is at least one of the plurality of computing cores; a routing index table is stored in the routing unit, and the routing unit is configured to obtain first index information according to the processing result, and determine the routing target according to the first index information and the routing index table, and the routing index table is configured to store at least one second index information and at least one computing core corresponding to each second index information. Each computing core can autonomously determine the next routing, realizing dynamic routing in the many-core architecture, thereby improving routing flexibility, eliminating system limitations caused by the processing speed and bandwidth of the SoC, improving the processing speed and quality of neural networks, and in addition, helping to reduce power consumption caused by frequent data transmission between the SoC and the computing cores. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown by way of illustration and not limitation, in which:

[0019] Figure 1 FIG. shows a schematic diagram of a typical neural network structure in the prior art.

[0020] Figure 2 FIG. shows a schematic diagram of the network connection of a many-core structure in the prior art.

[0021] Figure 3 FIG. schematically shows a structural diagram of a many-core device according to an embodiment of the present disclosure.

[0022] Figure 4 FIG. schematically shows a structural diagram of a computing core according to an embodiment of the present disclosure.

[0023] Figure 5 FIG. schematically shows a structural diagram of an exemplary computing core according to an embodiment of the present disclosure.

[0024] Figure 6 FIG. schematically shows a structural diagram of an exemplary many-core device according to an embodiment of the present disclosure.

[0025] Figure 7 FIG. schematically shows a schematic diagram of an exemplary routing index table according to an embodiment of the present disclosure.

[0026] Figure 8A schematic diagram schematically shows an exemplary routing index table according to an embodiment of the present disclosure.

[0027] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed implementation manners

[0028] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present disclosure, rather than limiting the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0029] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, a device, an apparatus, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0030] In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0031] The principles and spirit of the present disclosure will be elaborated below with reference to several representative embodiments of the present disclosure.

[0032] With the booming development of big data and neural network technologies, more and more processing nodes are included in the data processing system. Figure 1 A typical neural network structure in the prior art is shown. Each circle in the figure represents a neuron node, and the connection relationships between the respective neuron nodes are pre-planned.

[0033] At the hardware implementation level, each neuron node of the above neural network can correspond to one or more computing cores, and the routing paths between the respective computing cores need to be pre-planned and controlled by a SoC (System on Ship) 201, as Figure 2 shown. After each of the computing cores 202 to 210 obtains a processing result, it needs to be sent to a corresponding routing destination for further processing according to the routing instruction of the SoC 201, and the routing destination is one or more of the computing cores 202 to 210 specified by the SoC 201. The processing speed and bandwidth of the SoC undertaking the main control function are limited, which limits the number of computing cores that the entire many-core system 200 can support, thus severely restricting the processing speed and processing quality of the neural network. In addition, the frequent data transmission between the SoC 201 and the computing cores 202 to 210 also brings additional power consumption losses.

[0034] Therefore, an embodiment of the present disclosure provides a many-core device, aiming to autonomously determine the next routing by each computing core, implement dynamic routing in the many-core architecture, thereby improving routing flexibility, eliminating system limitations caused by the processing speed and bandwidth of the SoC, improving the processing speed and quality of neural networks, and in addition, helping to reduce power consumption caused by frequent data transmission between the SoC and the computing core.

[0035] Figure 3 Schematically shows a schematic structural diagram of a many-core device 300 according to an embodiment of the present disclosure, as Figure 3 shown, the many-core device 300 may include a plurality of computing cores, such as the computing cores 301 to 316 shown in the figure. As those skilled in the art understand, Figure 3 it is only used to schematically illustrate the many-core device 300 including a plurality of computing cores, and is not used to limit the number, arrangement position, connection relationship, etc. of the computing cores. The structure of each of the computing cores 301 to 316 may be as Figure 4 shown.

[0036] Figure 4 Schematically shows a schematic structural diagram of a computing core 400 according to an embodiment of the present disclosure, as Figure 4 shown, the computing core 400 may include a data processing unit 401 and a routing unit 402.

[0037] Among them, the data processing unit 401 is used to receive data to be processed, process the data to be processed, and send the processing result to the routing target. Among them, the routing target is at least one of the plurality of computing cores. In an embodiment of the present disclosure, the data to be processed may be video data, image data, audio data, text data, and the like.

[0038] Among them, a routing index table 4022 is stored in the routing unit 402. The routing unit 402 can receive its processing result from the data processing unit 401. The routing unit 402 obtains first index information according to the received processing result, and determines the above routing target according to the first index information and the routing index table 4022. The routing index table 4022 is used to store at least one second index information and at least one computing core corresponding to each second index information. Different second index information may correspond to different computing cores.

[0039] In some embodiments, the processing result obtained by the data processing unit 401 may include first index information, and the routing unit 402 can directly extract the first index information therefrom. For example, after the data processing unit 401 processes the data to be processed, the eigenvalue of the data to be processed is obtained as the processing result, and the routing unit 402 can directly extract the eigenvalue as the first index information. In the routing index table 4022, different eigenvalues may correspond to different computing cores, that is, different eigenvalues point to different routing targets.

[0040] In the embodiments of the present disclosure, the routing unit 402 can compare the first index information with each second index information in the routing index table 4022, and use the computing core corresponding to the second index information that matches the first index information as the routing target. The second index information that matches the first index information may be one or more, and each computing core corresponding to the second index information may be one or more.

[0041] Figure 5 Schematically shows a schematic structural diagram of an exemplary computing core 500 according to an embodiment of the present disclosure, as Figure 5 shown, the computing core 500 may include a data processing unit 501 and a routing unit 502, wherein the routing unit 502 includes a computing subunit 5021 and a routing index table 5022. The computing subunit 5021 can be used to compare the first index information with each second index information in the routing index table 5022, and determine the computing core corresponding to one or more second index information with the highest similarity to the first index information among each second index information as the routing target.

[0042] For example, the computing core corresponding to one second index information with the highest similarity to the first index information can be determined as the routing target; it is also possible to determine all computing cores corresponding to a predetermined number of second index information with the highest similarity to the first index information as the routing target, such as determining all computing cores corresponding to the first two second index information with the highest similarity to the first index information as the routing target; it is also possible to determine all computing cores corresponding to all second index information within a preset range of difference from the first index information as the routing target, and so on.

[0043] In some embodiments, the computing subunit 5021 may include a computing module implemented by hardware. That is, the computing function can be solidified in the routing unit 502 through hardware. For example, a general hardware computing unit can be used, or a non-general hardware computing unit can be used. Implementing the computing function by hardware, such as implementing cosine calculation by hardware for cosine comparison, can significantly improve the computing speed and reduce power consumption.

[0044] In some embodiments, multiple computing cores in a many-core device are divided into at least one computing core region, and each of the computing core regions includes at least one of the computing cores.

[0045] In some embodiments, each computing core region is configured to process the data to be processed received by the computing core region according to a preset algorithm, and obtain a corresponding processing result. Among them, each computing core in the computing core region is configured to process the sub-data to be processed obtained by the computing core according to a preset sub-algorithm, and obtain a sub-processing result; the sub-data to be processed is the data to be processed, or is a sub-processing result generated by a computing core other than the computing core in the computing core region.

[0046] In some embodiments, when each computing core sends a processing result to a routing target, the routing target and the computing core sending the processing result are in the same computing core region or in different computing core regions.

[0047] Figure 6 Schematically shows a schematic structural diagram of an exemplary many-core device 600 according to an embodiment of the present disclosure, as Figure 6 shown, the multiple computing cores of the many-core device 600 are divided into four computing core regions A, B, C, and D. The functions of each computing core region and the computing cores in the region can be planned in advance. For example, the computing core region A is set to extract features from the data to be processed received by it, and obtain corresponding feature values, such as feature values identifying a human face, etc.; the computing core region B is set to identify "animal features" for the received data to be processed, and determine whether the object is a cat, a dog, or a cow, etc.; the computing core region C is set to identify "human features" for the received data to be processed, and determine whether the object is an old person, a young person, or a child, etc.; the computing core region D is set to identify "vehicle features" for the received data to be processed, and determine whether the object is a truck, a car, or an engineering vehicle, etc.

[0048] Each computing core in each computing core region also processes the sub-data to be processed obtained by it according to a preset sub-algorithm, and obtains a sub-processing result. For example, the computing core 603 in the computing core region B processes the data received from it according to a preset sub-algorithm, and extracts the animal feature values therein; the computing core 604 processes the received data according to a preset sub-algorithm for cats; the computing core 607 processes the received data according to a preset sub-algorithm for dogs, and so on.

[0049] In one example, the computing core 603 receives data from the computing core area A, processes the received data to extract animal feature values, and compares the animal feature values as the first index information with the routing index table stored in the computing core 603. It is determined that the similarity with the second index information representing a cat is the highest, and the second index information representing a cat corresponds to the computing core 604. Then, the computing core 604 is determined as the routing target. The computing core 603 sends the processing sub-result to the computing core 604, and the computing core 604 receives the data as its current sub-data to be processed.

[0050] The above is only an example of partitioning. Those skilled in the art can partition multiple computing cores as needed to achieve partition multi-level processing. For example, multiple computing cores can be divided into multiple functional areas corresponding to the human brain structure. According to the many-core device of the present disclosure, multiple computing cores can be divided into multiple computing core areas, and these computing core areas can respectively correspond to the visual function, auditory function, somatosensory function, thinking function, and mental function of the human brain, etc.

[0051] In the above embodiment, by dividing multiple computing cores into computing core areas, partition multi-level processing of the many-core device can be achieved. Between computing core areas and between individual computing cores within a computing core area, autonomous dynamic routing can be achieved, further improving the flexibility and efficiency of the many-core device.

[0052] In some embodiments, the above routing index table can be divided into multiple index blocks, and the index blocks are divided according to the above at least one second index information, and different index blocks correspond to different computing core areas.

[0053] According to the information stored in the routing index table, the target routing can point to the corresponding computing core area or directly to a specific computing core in the corresponding computing core area.

[0054] Figure 7 Schematically shows a schematic diagram of an exemplary routing index table according to an embodiment of the present disclosure. As Figure 7 shown, the first index block indicates that the second index information 1a, 1b, and 1c corresponds to the computing core areas B and C, and the second index block indicates that the second index information 2a and 2b corresponds to the computing core area D, and so on.

[0055] In some embodiments, the index blocks are divided by the sorting result obtained by sorting at least one second index information according to a preset rule.

[0056] In one example, the routing index table can be partitioned in the following manner.

[0057] First, based on a preset rule, at least one piece of second index information stored in the routing index table is sorted.

[0058] The preset rule can be set according to the characteristics represented by the second index information. For example, when the second index information is a face feature value, sorting can be performed according to the age of a person. For example, when the age represented by the second index information is smaller, the sorting is more forward.

[0059] Sorting can also be performed according to other rules. For example, sorting can be performed according to the number of times the second index information is searched. For example, the more times a certain second index information is searched within a preset time period, the more forward its sorting is. Thus, when indexing, this second index information will be compared earlier.

[0060] Next, the routing index table is divided into multiple index blocks according to the sorting result.

[0061] For second index information with similar sorting results, they can be divided into the same index block.

[0062] In the embodiments of the present disclosure, by dividing the index block according to the sorting result of the second index information, the routing indexes with similar second index information can be divided into the same index block.

[0063] In some embodiments, the index block can also be obtained by obtaining the similarity between every two pieces of the at least one second index information and dividing at least two pieces of second index information with a similarity higher than a set threshold into the same index block.

[0064] In one example, the routing index table can be divided into blocks in the following manner.

[0065] First, the similarity between every two pieces of second index information in the routing index table is obtained.

[0066] Among them, the similarity between two pieces of second index information can be determined according to the Euclidean distance between the feature vectors corresponding to the second index information, or other methods can be used for calculation. The present disclosure does not limit the calculation method of the similarity.

[0067] Next, at least two pieces of second index information with a similarity higher than a set threshold are divided into the same index block.

[0068] Among them, for the second index information divided into the same index block, the similarity between every two pieces of second index information can be higher than the set threshold, or the similarity between one piece of second index information and at least one other piece of second index information can be higher than the set threshold.

[0069] In the embodiments of the present disclosure, by partitioning the index block according to the similarity of the second index information, processing results with similar second index information can be sent to the same routing destination.

[0070] In some embodiments, the routing index table further includes third index information corresponding to the index block, and the third index information is determined according to the second index information in the corresponding index block. Wherein, the calculation subunit is specifically configured to: compare the first index information with the third index information in the routing index table; and determine the calculation core area corresponding to the index block corresponding to one or more third index information with the highest similarity to the first index information as the routing destination.

[0071] Figure 8 Schematically shows a schematic diagram of an exemplary routing index table according to an embodiment of the present disclosure. For example, the third index information of the index block can be obtained by concatenating each second index information in the index block; for another example, the corresponding third index information can be obtained by averaging or weighted averaging each second index information.

[0072] For example, input a cat face image into the many-core device according to the present disclosure. The calculation core area A extracts features from the input data to be processed, obtains a feature value, and uses this feature value as the first index information to compare with the third index information in the local routing index table as shown in Figure 8 If it is determined that the third index information 1 has the highest similarity to the extracted first index information, the next routing destination is autonomously determined as the calculation core areas B and C, so as to send the processing results obtained by the calculation core area A to the calculation core areas B and C.

[0073] After receiving the data to be processed from the calculation core area A, the calculation core area B further extracts animal features by a designated calculation core, and uses the extracted feature value as the first index information to compare with the local routing index table. If it is determined that the second index information representing a cat has the highest similarity to it, the calculation core corresponding to the second index information is autonomously determined as the routing destination, and this calculation core is a calculation core in this calculation area that has a sub-algorithm for cats preset.

[0074] After receiving the data to be processed from the calculation core area A, the calculation core area C further extracts human features by a designated calculation core, and uses the extracted feature value as the first index information to compare with the local routing index table. If it is determined that the second index information representing non-human has the highest similarity to it, the calculation core corresponding to the second index information is autonomously determined as the routing destination, and this calculation core is a calculation core in this calculation area that has a sub-algorithm for non-human preset.

[0075] In an embodiment of the present disclosure, in the computing core area A, by comparing the first index information with the third index information in the local routing index table, the computing core area corresponding to the third index information with the highest similarity is used as the routing target. Compared with comparing the first index information with each second index information in the local routing table one by one, the amount of data to be processed is reduced and the routing speed is increased.

[0076] At least one embodiment of this specification also provides an electronic device, including a memory and a processor. The memory is used to store computer instructions that can run on the processor, and the processor is used to execute the computer instructions to implement the methods involved in the many-core device described in any embodiment of this specification.

[0077] At least one embodiment of this specification also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the methods involved in the many-core device described in any embodiment of this specification.

[0078] It should be noted that although several units / modules or sub-units / modules of the many-core device and computing cores are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple unit / modules.

[0079] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a device, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0080] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiment of the data information processing device, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

[0081] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures need not be in the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0082] Embodiments of the subject matter and the functional operations described in this specification can be implemented in: digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or one or more combinations of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier to be executed by, or to control the operation of, a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, generated to encode and transmit information to a suitable receiver apparatus for execution by the data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0083] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform the corresponding functions by operating on input data and generating output. The processes and logical flows can also be performed by, or the apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0084] Computers suitable for executing computer programs include, for example, general and / or special purpose microprocessors, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, etc., or the computer will be operably coupled to such mass storage devices to receive data therefrom or transfer data thereto, or both. However, a computer is not necessarily required to have such devices. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning device (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few examples.

[0085] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as including semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0086] Although this specification contains many specific implementation details, these should not be construed as limiting the scope of any invention or the scope of what is claimed, but rather as primarily describing the features of specific embodiments of particular inventions. Certain features described in multiple embodiments in this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may operate in certain combinations as described above and even be claimed as such initially, one or more features from a claimed combination may in some cases be removed from that combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0087] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood to require that the operations be performed in the particular order shown or sequentially, or that all illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various device modules and components in the above embodiments should not be understood to be required in all embodiments, and it should be understood that the described program components and devices can generally be integrated together in a single software product or packaged into multiple software products.

[0088] Accordingly, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. In addition, the processes depicted in the figures are not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0089] The above are only the preferred embodiments of one or more embodiments of this specification, and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope protected by one or more embodiments of this specification.

Claims

1. A many-core device, characterized in that, comprises a plurality of computing cores, each of the computing cores including a data processing unit and a routing unit; wherein, the data processing unit is configured to receive data to be processed, process the data to be processed, and send a processing result to a routing target, wherein the routing target is at least one of the plurality of computing cores; a routing index table is stored in the routing unit, and the routing unit is configured to receive the processing result from the data processing unit, obtain first index information according to the received processing result, compare the first index information with each index information in the routing index table, and use the computing core corresponding to the index information that matches the first index information as the routing target, and the routing index table is used to store at least one second index information and at least one computing core corresponding to each second index information.

2. The device according to claim 1, characterized in that, The routing unit includes a computing subunit, and the computing subunit is configured to: compare the first index information with the at least one second index information in the routing index table; determine, as the routing target, the computing core or cores corresponding to one or more second index information having the highest similarity with the first index information among the at least one second index information.

3. The many-core device according to claim 1, characterized in that, The plurality of computing cores are divided into at least one computing core region, and each computing core region includes at least one of the computing cores.

4. The device according to claim 3, wherein: each computing core region is configured to process the data to be processed received by the computing core region according to a preset algorithm to obtain a corresponding processing result; wherein each computing core in the computing core region is configured to process the sub-data to be processed obtained by the computing core according to a preset sub-algorithm to obtain a processing sub-result; the sub-data to be processed is the data to be processed, or is a processing sub-result generated by a computing core other than the computing core in the computing core region.

5. The device according to claim 3, characterized in that, The routing index table is divided into a plurality of index blocks, the index blocks are divided according to the at least one second index information, and different index blocks correspond to different computing core regions.

6. The device according to claim 5, characterized in that, The index blocks are divided by a sorting result obtained by sorting the at least one second index information according to a preset rule.

7. The device according to claim 5, characterized in that, The index blocks are obtained by obtaining the similarity between every two second index information in the at least one second index information, and dividing at least two second index information with a similarity higher than a set threshold into the same index block.

8. The device according to claim 5, wherein: the routing index table further includes third index information corresponding to the index block, and the third index information is determined according to the second index information in the corresponding index block; the routing unit includes a computing subunit, and the computing subunit is specifically configured to: compare the first index information with the third index information in the routing index table; determine, as the routing target, the computing core region corresponding to the index block corresponding to one or more third index information having the highest similarity with the first index information.

9. The device according to any one of claims 3 to 8, characterized in that, When each of the computing cores sends the processing result to the routing target, the routing target is in the same computing core area or in a different computing core area from the computing core that sends the processing result.

10. The device according to claim 2, characterized in that, The computing subunit includes a computing module implemented by hardware.

11. The device according to claim 1, characterized in that, The processing result includes the first index information.

Citation Information

Patent Citations

  • Intelligent semantic retrieval method and system and electronic equipment

    CN112035598A

  • Streaming data processing method based on many-core processor and computing equipment

    CN112114942A