PORTABLE MEMORY PROCESSING DEVICE BASED ON P-TYPE SPIKING NEURAL NETWORKS WITH UNARY CODING FOR EMBEDDED SYSTEMS.

MX431690BActive Publication Date: 2026-02-25NATIONAL POLYTECHNIC INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
MX2020012864
Authority / Receiving Office
MX · MX
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-11-27
Publication Date
2026-02-25
Estimated Expiration
2040-11-27

AI Technical Summary

Technical Problem

Existing memory systems for computers face significant area consumption due to separate processors and serial operations, while parallel computing is needed for high-performance applications, and no neural arithmetic device performs arithmetic operations and stores results in the same unit.

Method used

A portable memory processing device using a P-type spiking neural network with unary coding for parallel performance of addition, subtraction, multiplication, and division, incorporating a processing system, storage, and an execution method, capable of local or global operation through a communications network.

Benefits of technology

Enables efficient parallel arithmetic operations and storage within a single unit, reducing area consumption and enhancing computing performance for complex applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure MX431690B0
    Figure MX431690B0
  • Figure MX431690B1
    Figure MX431690B1
Patent Text Reader

Abstract

The present invention relates to a portable memory processing device for performing operations such as addition, subtraction, multiplication, and division of multiple digits in parallel, with the result stored in the same unit. The model and processing mode of the memory processing unit are based on P-type spiking neural networks. First, the numbers processed by the memory processing unit are stored in a unary system, in which the data are represented by spikes (or ones). Subsequently, the user enables the desired operation by activating the firing rules of the neurons. Finally, the resulting spikes, generated by the neurons, are stored in the neuron's soma.The activation mechanism of neurons and the storage process are performed in the same unit, eliminating the need for separate ALUs and memory, thus increasing processing speed. In current computing systems, data stored in memory is transferred to the ALU. This creates a bottleneck and generates high energy consumption. Therefore, our solution optimizes storage and processing within the same unit using unconventional mechanisms (neurons) to perform both operations.
Need to check novelty before this filing date? Find Prior Art

Description

The present invention focuses on the area of ​​digital design. In particular, this invention presents the design of a portable memory processing device for embedded systems. This portable memory processing device performs operations such as addition, subtraction, multiplication, and division in parallel, and the result is stored in the same device. BACKGROUND Today, applications such as video processing, cryptography, and others generate an immense amount of data that must be stored in memory. Therefore, in these applications, memory access time significantly impacts computer performance, as a massive amount of data needs to be transferred between memory and the ALU. This bottleneck becomes critical because CPU speeds have increased by a factor of two every two years, according to Moore's Law, while memory access times have decreased at a slower rate. In recent years, some designers have proposed memories that contain an embedded microprocessor or processor to speed up write, read, and instruction execution operations within the same storage unit. Basavaraj I. Pawate, Kenneth A. Poteet, and Joe H. Neal have filed a patent application, US5678021A, for an invention entitled “Apparatus and method for a memory unit with a processor integrated therein.” The intelligent memory includes data storage and a processing core for executing instructions stored in the data storage area. The intelligent memory is accessible as a standard memory device. In one mode of operation, the intelligent memory is a data storage device associated with a central processing unit (CPU). In a second mode of operation, the intelligent memory is a storage device for both the processing core and the CPU, thus executing instructions simultaneously. The CPU controls the mode of operation and determines which instructions are executed by the processing core.The memory contains a data bus of a large number of bits, which is available with an integrated processor / storage function, allowing certain processing operations to be offloaded to the intelligent memory where the processing operations can be performed more efficiently. Basavaraj I. Pawatc and Bctty Prince filed a patent application, US6000027A, for the invention entitled “Method and apparatus for improved graphics / image processing using a processor and a memory.” The intelligent video memory includes data storage, serial access memory, and a processing core for executing instructions stored in the data storage area. Externally, the intelligent memory is directly accessible as a standard video memory device. Kazumasa Kishi, Shigcki Masumura, Hidco Nakamura, K.ouki Noguchi, Shumpei Kawasaki, and Yasushi Akao filed a patent application for a "Semiconductor Integrated Circuit Having CPU and Multiplier" under number US5832248A. The Logic Integrated Circuit (LSI) chip includes a CPU, a bus, memory, and a multiplier. Furthermore, the LSI chip includes a command signal line for transferring data from the CPU to the multiplier and a command to specify the multiplication instruction related to data reading. Specifically, while data is being read from memory, the multiplier can retrieve data directly from the bus. As the CPU reads data from memory, a multiplication instruction command related to the data read is transferred from the CPU to the multiplier.A cycle control circuit receives a signal when the multiplier is performing a repetitive operation, and the bus cycle control circuit responds to the status signal by instructing the CPU to delay issuing the next command to the multiplier. Basavaraj I. Pawate, Gene A. Frantz, and Rajan Chirayil filed a patent application, US5638530A, for the invention entitled “Direct memory access scheme using memory with an integrated processor having communication with external devices.” The inventors presented a method and system for improving processing between a host computer and a logic unit. Data instructions are stored in multiple memory locations. Data is processed in response to instructions from the logic unit, which is integrated with the memory within a single integrated circuit. The memory locations are directly accessible without bus arbitration by an external device coupled to the integrated circuit, through an external interface that controls the processing speed of the logic. Analyzing previous inventions, it can be observed that the memories proposed to date have a separate processor and memory, resulting in a significant area consumption. Furthermore, the memory performs operations serially. Currently, parallel computing has become a cutting-edge area of ​​technological research and development. Therefore, software and hardware engineers are collaborating on extraordinary efforts to develop parallel computing systems that can meet the current high-performance computing needs of advanced and complex applications.Inspired by neural biology, a new branch of computer science is making extraordinary efforts to develop novel arithmetic circuits (adders, testers, dividers, and multipliers) based on P-type spiking neural networks. These networks have intrinsic parallel computation and use a single-digit encoding (one = 1 spike). It should be noted that to date, no arithmetic neural device has been proposed that performs arithmetic operations and stores the result in the same unit. The proposed invention uses a processing unit, which contains a processing system, storage and execution method for processing numbers under a unique encoding. BRIEF DESCRIPTION OF THE FIGURES The following is a description of the figures that have been used to describe the portable memory processing device. Figure 1.- Diagram of the connection network of portable memory processing devices. Figure 2.- Portable memory processing scheme. Figure 3 - Schematic of the memory module with processing. Figure 4.- Schematic of the neuronal processing and storage submodule. Figure 5 - Schematic of the neuronal memory submodule. Figure 6,- Schematic of the submodule with axonal delay (16395). DETAILED DESCRIPTION OF THE INVENTION The present invention relates to a portable memory processing device (1) for performing addition, subtraction, multiplication, and division of multiple uyv numbers in parallel and in a distributed manner. Therefore, this device can operate locally or globally. In the case of a global connection, the device is connected to an array of devices (1) via the communications network, as shown in Figure 1. Each user sends the uyv numbers from their computer (2) to their respective portable memory processing device (1) via the PCIe data bus (pcierx) or the USB data bus (usbrx). Once the data is received, each portable memory processing device (1) performs different arithmetic operations.If a portable memory processing device (1) requires more memory space, it sends a space request via the (cthtx) buses to available portable memory processing devices (l) on the global communications network. If a portable memory processing device (1) is available, it generates a response, which is sent across the global communications network via the (cthtx) bus. Therefore, each portable memory processing device (1) can extend its processing and storage capacity by distributing data to other portable memory processing devices (1) as long as they are connected to the global communications network.Once processing is complete on each portable memory processing device (1), the generated data can be sent to its respective computer via the PCIe data bus (pcietx) or the USB data bus (iisbtx). The portable memory processing device (1) is explained in detail below. Portable memory processing device (1) The portable memory processing device (1) contains a PCIe port control module (11), an Ethernet port control module (12), a USB port control module (13), a multiplexer (14), a data control module (15), a memory processing module (16), and a multiplexer (17), as shown in Figure 2. In general, the PCIe (11), Ethernet (12), and USB (13) port control modules are responsible for establishing communication between the memory processing module (16) and the computer (2). Specifically, the use of these communication protocols (Ethernet, PCIe, USB) allows the user to read or write data from the computer to any memory processing device (1), either locally or globally (see Fig. 1).The user enables the data control module (15) using the (ssel) signal to tell the multiplexer (14) which data bus to use (del) or (d_e_2) or (d_e_3) and which data bus to use (d_sj) or (d_s_2) or (d_s_3). The user can also enable and disable data writing and reading to the portable memory processing device (1) using the (shab) signal. If the processing memory module (16) does not have enough memory space, this module generates the ( / ') signal to tell the computer (2) to distribute data to other processing memory modules (16) connected to the global communications network. Once the data buses are defined, the data control module (15) generates the control signals (s crol' and scrol'') to establish the input (de') and output (d_s_1) data flows. d_s_2, .... d_s_n) through their respective multiplexers (14 and 17).In particular, the multiplexer (17) sends the output data (d_s_l, d_s_2.....dsji), which represent the numbers processed by the processing memory module (16), to the multiplexer (14) by means of the signal (d_s') to subsequently send them to the computer (2). Memory module with processing (16) The processing memory module (16) is composed of a configuration submodule (161), n ​​data distribution submodules (162), nxm neural processing and storage submodules (163), nxm multiplexers (164), and nxm multiplexers (165), where nym are defined by the user, as shown in Figure 3. In general, the configuration submodule (161) is responsible for segmenting the input data (de ') and subsequently sending it to the n columns, which are controlled by the n data distribution submodules (162), through the data buses (b_d_l.....b dji) and the address buses (ba l.....b_a_n) in parallel. It should be noted that the data distribution submodules (162) are responsible for configuring and controlling the input and output data flow of the neural processing and storage submodules (163) by means of the control signals (e_m_l, i_d_l, sel, d_c_l, l_g_l, .... e_m_n, idji, sen, den, l_g_n).Furthermore, it manages the available memory space. If all the processing and storage neural submodules (163) contain data, the data distribution submodule (162) sends signal (f) to the configuration submodule (161) to indicate that the processing and storage neural submodules (163) are busy. Additionally, the processing and storage neural submodules (163) perform mathematical operations such as: Addition, subtraction, multiplication, and division are performed in parallel using a unique encoding. Finally, the multiplexers (164) along with the multiplexers (165) are controlled by the data distribution submodules (162) to feed data (r1, ..., rp) between the neural processing and storage submodules (163), thus establishing local communication. Additionally, data (dsl, ..., dsn) can be sent to the computer (2), establishing global communication. The neural processing and storage submodule (163) is explained in detail below. Neural processing and storage submodule (163) Each of the neuronal processing and storage submodules (163) has the same configuration and structure. The neuronal processing and storage submodule (163) is composed of a multiplexer (1631), multiplexers (1632), a comparator (1633), a register (1634), p shift registers (1635), p multiplexers (1636), a multiplexer (1637), a multiplexer (1638), and a neuronal memory submodule (1639), where p is user-defined, as shown in Figure 4. It should be noted that each neuronal processing and storage submodule (163) has an identification number, which is configured by means of the data distribution submodule (162) using the signal (i_d_l). Once the identification number is configured, the data distribution submodule (162) controls the neural processing and storage submodule (163) to perform three operations: Case 1 (writing the numbers uyv)·. To write the number u to the neural processing and storage submodule (163), the data distribution submodule (162) enables the signal (e_ni_l) so that the multiplexers (1632) allow the input of the signals (til, u_2, u_p), which represent the digits of the number u and are encoded by pulse trains (spikes) or a sequence of ones. Subsequently, the output signals (u_l, u_2, u_p) are recorded in their respective shift registers (1635) when the comparator (1633) generates the signal (i_d_l') with a value of one. It should be noted that the value of the signal (i_d_l') is equal to one when the signal (c), which is stored in register (1634), and the signal (i_d_l) are identical. Furthermore, to write the number v to the neural memory submodule (1639), the data distribution submodule (162) enables the control signal (sel j) and the data signal (d_c_l) to distribute the digits of the number v and the control signals (evr, ev) via the multiplexer (1631). Additionally, the multiplexer (1631) is enabled by the j signal when the output of the comparator (1633) is one, that is, when the (c) signal and the (i_d_I) signal are identical. Specifically, the data distribution submodule (162) sends the write signal (vv) via the multiplexer (1637) and also sends the digits of v via the (v) signal through the multiplexer (1638). It should be noted that the writing of the digits of v (y_l v_2 v_3 '.....v_p') are sent sequentially to the neuronal memory submodule (1639) by means of the multiplexers (1637 and 1638), which are controlled by means of their respective signals (e_vv, e_v).It should be noted that the neural memory submodule (1639) is composed of p registers (16391), p multiplexers (16392), where p is user-defined, an adder (16393), a multiplexer (16394), and an axonal delay submodule (16395), as shown in Figure 5. For example, if the digit (v_ / ') needs to be written to the first register (16391), the data distribution submodule (162) together with the multiplexer (1637) (see Figure 4) activates the signal {e_1} with a value of one to enable writing to that register. Therefore, the remaining registers (16391 > 1) are enabled by means of the signals {e_2, e_3, ..., e_p} to write the digits (v_2 ..., v_p).Once the data distribution submodule (162) finishes writing the uyv digits to the processing and storage neural submodule (163), it continues writing the uyv digits to the next processing and storage neural submodule (163) until it finishes writing the uyv digits to the last processing and storage neural submodule (163). This procedure is performed identically for each of the columns of the array of processing and storage neural submodules (163) by its respective data distribution submodule (162). Case 2 (processing: addition, subtraction, multiplication, and division): Once the digits of u and v are written in each neuronal processing and storage submodule (163), the data distribution submodule (162) together with the multiplexer (1631) controls the flow of the digits of u to perform the following arithmetic operations: To perform the addition, the digits of u, which are represented by spikes or pulses, are sent through the signals (u_l, u_2', u_3', ..., u_p') to their respective multiplexers (16392) (see Figure 5). For example, if the pulse {u_l'} is one, the multiplexer (16392) transfers the value of register (16391) to the adder (16393) using the signal (r_p_l). If there is no pulse for any digit of {u_l}, the multiplexer (16392) transfers a zero to the adder (16393). Therefore, the adder (16393) adds the values ​​of the registers (16391) as long as their respective digits (u_l', u_p') are equal to one. - To perform the subtraction, each of the values ​​in registers (16391) contains the sign bit equal to one. Therefore, adder (16393) performs the sum of the negative values ​​in registers (16391) as long as their respective digits (n_¡u_p”) are equal to one. - To perform the multiplication, each digit of (u) is represented by a train of n pulses, where n is equivalent to the value of the multiplier. In this case, the adder (16393) performs the sum of the multiplicand of each register (16391) n times, as long as the digits (u_l tij)') are equal to one. In the present invention, the multiplication is based on successive sums. Therefore, the adder (16393) performs the sum of the multiplier n times, where the value n corresponds to the value of the multiplicand (u). - To perform the division, each digit of (u) is represented by a train of n pulses, where n is equivalent to the value of the dividend. In addition, registers (16391) must contain the value one. When the value of (i) is equal to one, the multiplexer (16392) sends the signal (r_p_l j) to the adder (16393). Once the adder (16393) finishes adding, the resulting value is transferred by the signal (rj') to the axonal delay submodule (16395) through the multiplexer (16394) as long as the signal (la) is equal to one (see Figure 5). It should be noted that the axonal delay submodule (16395) is composed of a register (163951), a multiplexer (163952), a comparator (163953), a subtractor (163954), a multiplexer (163955), a shift register (163956), and a multiplexer (163957) (see Figure 6).To generate the result of the division, the data distribution submodule (162) sends the divisor value to register (163951) through the multiplexer (1631) (see Figure 4) using the signal (um). Once the divisor value is stored in register (163951), the subtractor (163954) (see Figure 5) subtracts the contents of register (163951) from the result (r^ / jy) as long as the remainder (res) is greater than the divisor. This condition is indicated by the comparator (163953) using the signal (s_c), which controls the output of the multiplexer (163952). If the remainder (res) is greater than the divisor, the subtractor (163954) subtracts the contents of register (163951) from the result (rj); otherwise, the subtractor (163954) subtracts 0.In the present invention, division is based on successive subtractions; therefore, the subtractor (163954) performs several subtractions as long as the remainder (res) is greater than the divisor and the signal (en_d) is equal to one. Case 3 (reading the result and local or global storage): The result of the addition, subtraction, multiplication, or division performed by each neural processing and storage submodule (163) is sent to its respective multiplexer (164) via the signal (ni_p). It should be noted that the axonal delay submodule (16395) generates the output signal (ni_p), which represents the result of the mathematical operation with its corresponding sign. This submodule first generates the result, which is then sent via the signal (a) (see Figure 6). Therefore, the multiplexer (163957) enables the input signal (a) when the control signal (s_o) is one. The sign pulse (s_a) is sent to the multiplexer (163955) when the control signal (d) is equal to one. In this way, the multiplexer (163955) sends the signal (s_a') to the shift register (163956). Subsequently, if the control signal (x_o) is zero and the signal (en_s) is equal to one, the output of the multiplexer (163957) sends the sign pulse (s_a'') through the signal (ni_p).It should be noted that the signal (en_s) enables the shift register (163956) to perform a right shift. In this way, the sign pulse is sent to the multiplexer (163957). The final result can be distributed locally or globally as described below: Local distribution: The output of each neuronal processing and storage submodule (163) is fed back through its respective signal (ni_p'} (see Figure 3). To perform distribution and local storage in each neuronal processing and storage submodule (163), the signal (ni_p) is sent to its respective multiplexer (164). Subsequently, each multiplexer (164) generates the output signal (ni..., ni_p_ni) if any of the control signals (l_g_i, ..., l_g_n) is equal to 0. In this way, each signal (w_p_l, ..., ni_p_tn) is sent to the multiplexers (165) to generate the feedback input signals (r_l, r_2, r_3, ..., r_p) to each of the neuronal processing and storage submodules (163) in order to perform the three functions mentioned above. Global distribution: In this distribution mode, each signal (l_g_l.....l_S_n) must be equal to 1 to enable the multiplexers (164), and in this way the signals (ni_p_l' m_p_m') from each column are sent to each data distribution submodule (162) (see Figure 3). Subsequently, the data distribution submodule (162) sends the output data (d_s_l.....d_s_n) to the multiplexer (17) (see Figure 2) to send the data to the computer (2) or to the global communications network.

Claims

1. A portable memory processing device (1) operating with a P-type spiking neural network with finger encoding, characterized in that it contains a system with a communication module, a processing module and a storage module, wherein said communication module preferably contains communication protocols for establishing local communication, said communication allowing interaction between the device (1) and a computer (2) and global communication for communicating the device (1) with other devices (1) connected to the global communication network, said processing module preferably contains neural processing and storage submodules (163) for performing the addition, subtraction, multiplication and division of multiple numbers,This storage module preferably contains axonal delay submodules (16395) to store the result of arithmetic operations in the same device.

2. The portable memory processing device (1) according to claim 1, characterized in that it uses a PCIe port control module (11) and a USB port control module (13) to establish local communication between the device (1) and the computer (2).

3. The portable memory processing device (1) according to claim 1, characterized in that it uses an Ethernet port control module (12) to establish global communication between the device (1) and other devices (1). 20 4. The portable memory processing device (1) according to claim 1, characterized in that it uses neural processing and storage submodules (163) to perform the addition, subtraction, multiplication and division of multiple numbers.

5. The neuronal processing and storage submodule (163) according to claim 4, uses a nail-like encoding to represent the numbers uyv in trains of 25 pulses.

6. The neuronal processing and storage submodule (163) according to claim 4, uses synaptic weights to perform addition, subtraction, multiplication and division of multiple numbers.

7. The axonal delay submodule (16395) according to claim 1, uses the subtractor 30 (163954) to generate the result of the arithmetic operations.

8. The axonal delay submodule (16395) according to claim 1, uses shift registers (163956) to generate the sign bit of the result of the mathematical operations.