A new type of computing architecture

By dividing large-scale data matrices into submatrices in the new computing architecture and building neural network models, parallel computing processing is realized, and the inefficiency of the existing brain-like computing system architecture in a variety of pattern recognition problems is solved, and the accuracy and efficiency of pattern recognition are improved.

CN117933325BActive Publication Date: 2025-06-03NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311845231.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-06-03
Estimated Expiration
2043-12-28

AI Technical Summary

Technical Problem

The existing brain-like computing system architecture is inefficient among various pattern recognition problems, which is manifested as a decrease in accuracy, large delay, and insufficient sensitivity in pattern recognition.

Method used

Using a new computing architecture, we use a large-scale data matrix to divide it into multiple submatrixes and build a neural network model based on the number of submatrixes, including cortical columns, neuron groups and neurons, and realize parallel computing processing and multiplication accumulation calculation to reduce the calculation delay.

Benefits of technology

The accuracy and efficiency of pattern recognition are improved, the calculation delay is reduced, and the inefficiency of brain-like computing system architecture in the prior art in various pattern recognition problems is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117933325B_ABST
    Figure CN117933325B_ABST
Patent Text Reader

Abstract

The present invention relates to a novel computing architecture, belonging to the technical field of computer architecture. The architecture includes: a controller and an information processor; the controller uses a data splitting unit to divide a large-scale data matrix to be processed into multiple small-scale sub-matrices as needed; the information processor uses a neural network controller to obtain the number of corresponding elements in the neural network model to be constructed based on the number of sub-matrices obtained after division, and constructs the neural network model based on the number of corresponding elements in the neural network model to be constructed, distributes the sub-matrices obtained after division to the neural network model for parallel computing processing, and performs multiplication and accumulation calculation on the parallel computing processing results to obtain the final calculation result. The architecture provided by this application enables the basic sub-matrices corresponding to the large-scale data matrix to be calculated in parallel, reduces the calculation latency, and solves the problem of low calculation speed of the large-scale data matrix on ordinary servers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer architecture, and in particular to a new computing system architecture. Background Art

[0002] At present, most computers in the existing technology use the von Neumann architecture, which encodes the program into binary form and stores it in an external memory together with the data, and executes the program through the "store and execute" method. This architecture facilitates the software and hardware design of the computer. However, with the development of semiconductor technology and integrated circuit technology, the read and write speed of the external memory has far lagged behind the computing speed of the CPU.

[0003] With the rapid development of neuromorphic engineering in the past decade, integrated circuit chips and systems that draw on the computing principles of the brain have made high-density, low-energy, and real-time information processing architectures possible, which is expected to break the storage performance bottleneck caused by the current von Neumann architecture that separates storage and computing. Brain-like computing system architecture refers to a new type of information processing architecture that is inspired by the brain and provides efficient solutions to the computing power problems of artificial intelligence through a large-scale parallel computing system of neural networks with neurons as basic elements. Compared with the traditional von Neumann architecture, the brain-like computing system architecture draws on the working principles of human brain neurons. The actual computing unit realizes the integration of computing and storage, gets rid of the dependence on external memory, has high robustness to data, low energy consumption, high efficiency, and alleviates the problem of mismatch between computing and storage rates.

[0004] However, the current brain-like computing system architecture, based on a single neuron + neural network structure, is very inefficient for multiple pattern recognition problems, with characteristics such as reduced accuracy of pattern recognition, large delays, and insufficient sensitivity. Summary of the invention

[0005] The present invention is intended to provide a new computing system architecture to solve the deficiencies in the prior art. The technical problem to be solved by the present invention is achieved through the following technical solutions.

[0006] The novel computing system structure provided by the present invention comprises:

[0007] Controllers and information processors;

[0008] The controller specifically includes an input and output control unit, a data segmentation unit and a storage control unit;

[0009] The information processor includes a neural network controller and a neural network model;

[0010] The data splitting unit is used to divide the large-scale data matrix to be processed into multiple sub-matrices of small scales according to requirements;

[0011] The neural network controller obtains the number of corresponding elements in the neural network model to be constructed based on the number of sub-matrices obtained after division, constructs the neural network model based on the number of corresponding elements in the neural network model to be constructed, distributes the sub-matrices obtained after division to the constructed neural network model for parallel computing processing, and performs multiplication and accumulation calculations on the parallel computing processing results to obtain the final calculation result.

[0012] In the above solution, the neural network model includes a network router, neurons, neuron groups, and cortical columns. Among them, the neurons include a multiplier-accumulator and a memory, the neuron groups are composed of multiple neurons, the cortical columns are composed of multiple neuron groups, and the network router is used to construct the neural network model under the control of the neural network controller and perform information transmission between neurons, neuron groups, and cortical columns.

[0013] In the above solution, the neural network controller obtains the number of cortical columns in the neural network model to be constructed based on the number of sub-matrices obtained after division, and constructs the cortical columns.

[0014] In the above solution, during the process of constructing the cortical columns, multiple neuron groups are constructed, and a cortical column is formed by the multiple constructed neuron groups; during the process of constructing the neuron groups, multiple neurons are constructed, and a neuron group is formed by the multiple constructed neurons.

[0015] In the above solution, the sub-matrices obtained after division are distributed to the constructed neural network model for parallel computing processing, and performing multiplication and accumulation calculations on the parallel computing processing results to obtain the final calculation result includes:

[0016] The sub-matrices to be processed are further divided into multiple secondary sub-matrices, and each secondary sub-matrix is further divided into multiple basic sub-matrices. The multiplier-accumulator in the constructed neurons is used to perform multiplication and accumulation calculations on each basic sub-matrix, and the memory in the constructed neurons is used to store the calculation results of each basic sub-matrix, where one neuron corresponds to one basic sub-matrix;

[0017] The neuron groups are used to perform multiplication and accumulation calculations on the calculation results of each basic sub-matrix to obtain the calculation results of each secondary sub-matrix;

[0018] The cortical columns are used to perform multiplication and accumulation calculations on the calculation results of each secondary sub-matrix to obtain the calculation results of the sub-matrices;

[0019] Multiply and accumulate the calculation results of each sub - matrix to obtain the calculation result of the large - scale data matrix to be processed, and record the calculation result of the large - scale data matrix to be processed as the result matrix.

[0020] In the above - mentioned solution, the number of elements in the result matrix is the number of neurons.

[0021] In the above - mentioned solution, the number of secondary sub - matrices is the number of neuron groups, and the number of basic sub - matrices is the number of neurons.

[0022] In the above - mentioned solution, the structure further includes: an input module, an input storage / cache module, an output cache module, and an output module;

[0023] The input module, the input storage / cache module, the output cache module, and the output module all communicate with the controller.

[0024] The input storage / cache module is connected to the input module, the information processor is connected to the input module, the output cache module is connected to the information processor, and the output module is connected to the output module;

[0025] The controller receives the data to be processed, and uses the input - output control unit to control the input module to input the data to be processed into the input storage / cache module, and uses the storage control unit to control the input storage / cache module to perform cache synchronization on the data to be processed after segmentation processing;

[0026] The controller uses the storage control unit to control the output cache module to cache the final calculation result, and controls the output cache module to transmit the final calculation result to the output module. The controller controls the output module to output the calculation result to an external memory or a network memory.

[0027] In the above - mentioned solution, the input module uses a virtual input memory, the output module uses a virtual output memory, and the input module and the output module can be infinitely extended.

[0028] In the above - mentioned solution, when the output data needs to be reused, the controller uses the storage control unit to control the output cache module to write the cached calculation result back to a specified position in the input storage / cache module.

[0029] The embodiments of the present invention have the following advantages:

[0030] The novel computing architecture provided by the embodiments of the present invention divides a large-scale data matrix into multiple sub-matrices, obtains the number of cortical columns in the neural network model to be constructed based on the number of sub-matrices obtained after division, further divides the sub-matrices to obtain secondary sub-matrices, divides the secondary sub-matrices into multiple basic sub-matrices, and processes the basic sub-matrices through the neurons in the cortical columns, realizing that the basic sub-matrices corresponding to the large-scale data matrix can be calculated in parallel, reducing the calculation latency, and solving the problem of low calculation speed of the large-scale data matrix on ordinary servers; moreover, the neural network model including cortical columns, neuron groups, and neurons solves the problems of reduced accuracy of pattern recognition, large latency, and insufficient sensitivity in the existing brain-inspired computing system architecture mode. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a schematic structural diagram of a novel computing architecture in an embodiment of the present invention;

[0032] Figure 2 is a schematic structural diagram of the neural network model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0034] As Figure 1 shown, the present invention provides a novel computing architecture, which includes:

[0035] a controller, an input module, an input storage / cache module, an information processor, an output cache module, and an output module;

[0036] The input module, the input storage / cache module, the information processor, the output cache module, and the output module are all in communication with the controller;

[0037] The input storage / cache module is connected to the input module, the information processor is connected to the input module, the output cache module is connected to the information processor, and the output module is connected to the output module;

[0038] Among them, the controller specifically includes an input / output control unit, a data segmentation unit, and a storage control unit;

[0039] The information processor includes a neural network controller and a neural network model;

[0040] The controller receives the data to be processed, uses the input / output control unit to control the input module to input the data to be processed into the input storage / cache module, uses the storage control unit to control the input storage / cache module to perform cache synchronization on the data to be processed, the controller uses the data splitting unit to split the data to be processed that has been cache-synchronized by the storage / cache module, and controls the storage / cache module to send the split data to be processed to the neural network controller in the information processor in sequence. The neural network controller constructs a neural network model based on the split data to be processed, and distributes the split data to be processed into the constructed neural network model for parallel computing, and performs multiplication and accumulation calculations on the parallel computing results to obtain the final calculation result. After the neural network model finishes calculating the data to be processed, the controller uses the storage control unit to control the output cache module to cache the final calculation result, and controls the output cache module to transmit the final calculation result to the output module. The controller controls the output module to output the calculation result to an external memory or a network memory.

[0041] Specifically, in the embodiment of the present invention, the input module uses a virtual input memory, the output module uses a virtual output memory, and the input module and the output module can be infinitely expanded.

[0042] Specifically, in the embodiment of the present invention, when the output data needs to be reused, the controller uses the storage control unit to control the output cache module to write the cached calculation result back to the specified position in the input storage / cache module.

[0043] Such as Figure 2As shown, in an embodiment of the present invention, the neural network model includes: network routing 1, neurons 2, neuron groups 3, and cortical columns 4. Among them, neuron 2 is the smallest computing unit in the neural network model. During the training of the neural network model, neuron 2 continuously adjusts the weight parameters through calculation to learn relevant recognition information. The neuron group 3 is composed of 100 - 200 neurons 2. The interconnection structure among the neurons in the neuron group 3 is relatively fixed, and its function is equivalent to a low-level pattern recognizer, which can store the surface information obtained from the trained neurons. When receiving the same surface information during the recognition process, it can recognize faster than neurons. The cortical column 4 is composed of 60,000 neurons and 300 neuron groups. The function of the cortical column 4 is equivalent to a high-level pattern recognizer. It obtains pattern recognition information from multiple neuron groups and processes it into cognitive concept class information. When receiving the same information during the recognition process, it can recognize higher-level cognitive abilities faster than neurons and neuron groups. The network routing 1 is used to construct the neural network model under the control of the neural network controller to achieve the purpose of information transmission between neurons, neuron groups, and cortical columns. The weight parameters of the neurons 2 in the neural network model are variable and can change continuously during the training process. The neuron group 3 and the cortical column 4 are fixed pattern recognizers. After the training is completed, they can learn surface class information and cognitive concept class information from neurons respectively, and the recognition speed is faster than that of neurons, thereby accelerating the calculation speed of the entire model.

[0044] Specifically, the neuron 2 includes a multiply-accumulator and a memory.

[0045] Specifically, in an embodiment of the present invention, during the process of constructing the neural network model by the information processor, the neurons, neuron groups, and cortical columns can be infinitely expanded.

[0046] The above neural network model subverts the traditional von Neumann architecture and can solve problems such as the memory wall problem of multi-core accessing shared memory and the system insecurity problem that program data in the same area may be accidentally modified by data in the von Neumann architecture.

[0047] Specifically, in an embodiment of the present invention, the data to be processed is represented as a large-scale data matrix, and the data segmentation unit in the controller divides the large-scale data matrix to be processed into multiple small-scale sub-matrices according to needs.

[0048] Specifically, in an embodiment of the present invention, the neural network controller in the information processor constructs the same number of cortical columns as the number of sub-matrices to be processed based on the number of sub-matrices to be processed.

[0049] Specifically, during the construction of the cortical column, 300 neuron groups are constructed, and a cortical column is formed by the 300 constructed neuron groups; during the construction of the neuron group, 200 neurons are constructed, and a neuron group is formed by the 200 constructed neurons.

[0050] Specifically, the sub-matrix to be processed is further divided into 300 secondary sub-matrices, and each secondary sub-matrix is further divided into 200 basic sub-matrices. The multiplication accumulator in the constructed neurons is used to perform multiplication and accumulation calculations on each basic sub-matrix, and the memory in the constructed neurons is used to store the calculation results of each basic sub-matrix. Here, one neuron corresponds to one basic sub-matrix; the neuron group is used to perform multiplication and accumulation calculations on the calculation results of each basic sub-matrix to obtain the calculation results of each secondary sub-matrix, and the cortical column is used to perform multiplication and accumulation calculations on the calculation results of each secondary sub-matrix to obtain the calculation result of the sub-matrix. Finally, the calculation results of each sub-matrix are multiplied and accumulated to obtain the calculation result of the large-scale data matrix to be processed. The calculation result of the large-scale data matrix to be processed is recorded as the result matrix. The number of elements in the result matrix is the number of basic sub-matrices, which is the number of neurons. Thus, the basic sub-matrices corresponding to the large-scale data matrix can be calculated in parallel, reducing the calculation latency.

[0051] In an embodiment of the present invention, the data to be processed is represented as a large-scale data matrix. The data splitting unit in the controller divides the large-scale data matrix to be processed into 2 small-scale sub-matrices as needed: sub-matrix A and sub-matrix B. The neural network controller in the information processor constructs 2 cortical columns based on the number of sub-matrices to be processed. Then, the calculation result matrices obtained by processing sub-matrix A and sub-matrix B through the 2 cortical columns are respectively st and B tr , where s is the number of rows in the calculation result matrix corresponding to sub-matrix A, t is the number of columns in the calculation result matrix corresponding to sub-matrix A, and is the number of rows in the calculation result matrix corresponding to sub-matrix B, r is the number of columns in the calculation result matrix corresponding to sub-matrix B. The calculation result of the large-scale data matrix to be processed is recorded as the result matrix C. Then, the expression of the result matrix C is:

[0052] where i = 1, 2, …, s; j = 1, 2, …, r.

[0053] It should be noted that the above detailed description is exemplary and is intended to provide further explanation for the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0054] It should be noted that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of the features, steps, operations, devices, components, and / or combinations thereof.

[0055] It should be noted that the terms "first", "second", etc. in the description, claims, and above-mentioned drawings of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that these terms can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.

[0056] In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0057] For ease of description, spatial relative terms such as "above", "on top of", "on the upper surface", "above", etc. may be used herein to describe the spatial positional relationship of one device or feature to other devices or features as shown in the figures. It should be understood that the spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is inverted, the device described as "above" or "on top of" other devices or structures will then be positioned "below" or "beneath" other devices or structures. Thus, the exemplary term "above" can include both the orientation of "above" and "below". The device can also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and corresponding interpretations of the spatial relative descriptions used herein will be made.

[0058] In the detailed description above, reference has been made to the accompanying drawings, which form a part hereof. In the drawings, like reference numerals typically identify like components, unless the context indicates otherwise. The illustrated embodiments described in the detailed description, the drawings, and the claims are not meant to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.

[0059] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A computing system, characterized in that, it includes: a controller and an information processor; the controller specifically includes an input / output control unit, a data segmentation unit, and a storage control unit; the information processor includes a neural network controller and a neural network model; the data segmentation unit is used to divide a large-scale data matrix to be processed into multiple sub-matrices of small scales according to needs; the neural network controller obtains the number of corresponding elements in the neural network model to be constructed based on the number of sub-matrices obtained after division, constructs a neural network model based on the number of corresponding elements in the neural network model to be constructed, allocates the sub-matrices obtained after division to the constructed neural network model for parallel computing processing, and performs multiplication and accumulation calculations on the parallel computing processing results to obtain the final calculation result; the neural network model includes network routing, neurons, neuron groups, and cortical columns. Among them, the neurons include multiplication accumulators and memories, the neuron groups are composed of multiple neurons, the cortical columns are composed of multiple neuron groups, and the network routing is used to construct a neural network model under the control of the neural network controller and perform information transmission between neurons, neuron groups, and cortical columns.

2. The computing system according to claim 1, characterized in that, the neural network controller obtains the number of cortical columns in the neural network model to be constructed based on the number of sub-matrices obtained after division, and constructs cortical columns.

3. The computing system according to claim 2, characterized in that, during the construction of cortical columns, multiple neuron groups are constructed, and a cortical column is formed by the constructed multiple neuron groups; during the construction of neuron groups, multiple neurons are constructed, and a neuron group is formed by the constructed multiple neurons.

4. The computing system according to claim 1, characterized in that, allocating the sub-matrices obtained after division to the constructed neural network model for parallel computing processing, and performing multiplication and accumulation calculations on the parallel computing processing results to obtain the final calculation result includes: further dividing the sub-matrices to be processed into multiple secondary sub-matrices, and further dividing each secondary sub-matrix into multiple basic sub-matrices, using the multiplication accumulators in the constructed neurons to perform multiplication and accumulation calculations on each basic sub-matrix, and using the memories in the constructed neurons to store the calculation results of each basic sub-matrix, where one neuron corresponds to one basic sub-matrix; using neuron groups to perform multiplication and accumulation calculations on the calculation results of each basic sub-matrix to obtain the calculation results of each secondary sub-matrix; using cortical columns to perform multiplication and accumulation calculations on the calculation results of each secondary sub-matrix to obtain the calculation results of sub-matrices; performing multiplication and accumulation calculations on the calculation results of each sub-matrix to obtain the calculation result of the large-scale data matrix to be processed, and recording the calculation result of the large-scale data matrix to be processed as the result matrix.

5. The computing system according to claim 4, characterized in that, the number of elements in the result matrix is the number of neurons.

6. The computing system according to claim 4, wherein, the number of the secondary sub - matrices is the number of neuron groups, and the number of the basic sub - matrices is the number of neurons.

7. The computing system according to claim 1, wherein, the system further includes: an input module, an input storage / cache module, an output cache module, and an output module; the input module, the input storage / cache module, the output cache module, and the output module communicate with the controller; the input storage / cache module is connected to the input module, the information processor is connected to the input module, the output cache module is connected to the information processor, and the output module is connected to the output module; The controller receives the data to be processed, and uses the input - output control unit to control the input module to input the data to be processed into the input storage / cache module, and uses the storage control unit to control the input storage / cache module to perform cache synchronization on the data to be processed after segmentation; The controller uses the storage control unit to control the output cache module to cache the final calculation result, and controls the output cache module to transmit the final calculation result to the output module, and the controller controls the output module to output the calculation result to an external memory or a network memory.

8. The computing system according to claim 7, wherein, the input module uses a virtual input memory, the output module uses a virtual output memory, and the input module and the output module can be infinitely expanded.

9. The computing system according to claim 7, wherein, when the output data needs to be reused, the controller uses the storage control unit to control the output cache module to write the cached calculation result back to a specified location in the input storage / cache module.

Citation Information

Patent Citations

  • Computing array based neural network processor

    CN107918794A

  • Parallel method based on convolutional neural network training

    CN112396154A