A control method, device, storage system and electronic device

CN122838352APending Publication Date: 2026-09-29SHANGHAI ZHICUN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610855892.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-12
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

随着大数据和人工智能等技术的发展,数据处理量和存储量迅速增长,算力需求和存储压力制约着该类技术的发展,传统的计算模式难以满足处理能力的需求,存储器的传统应用难以满足技术发展的需要

Benefits of technology

[0028]本申请提供的一种控制方法、装置、存储系统、电子设备及云端设备,利用第一数据分布特征进行第一序列和第一映射关系的设计,能够有效平衡乘累加运算精度与计算密度,提高存内计算的可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838352A_ABST
    Figure CN122838352A_ABST
Patent Text Reader

Abstract

This application provides a control method, apparatus, storage system, and electronic device, relating to the field of electronic technology, for effectively balancing the precision and computational density of multiply-accumulate operations and improving the reliability of in-memory computation. The control method includes: encoding first data according to a first sequence and a first mapping relationship to obtain first encoded data, wherein the first sequence and the first mapping relationship are related to the distribution characteristics of the first data; storing the first encoded data in a storage circuit, wherein the first encoded data is used for in-memory computation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technology, and in particular to a control method, apparatus, storage system and electronic device. Background Technology

[0002] In traditional computing models, such as the von Neumann architecture, storage and computation are physically separated. When processing data using this model, data is frequently transferred between memory and the processor, resulting in data transfer latency and energy consumption. With the development of technologies such as big data and artificial intelligence, the amount of data processed and stored is growing rapidly. The demand for computing power and the pressure on storage are constraining the development of these technologies. Traditional computing models are struggling to meet the demands for processing power, and traditional applications of memory are failing to meet the needs of technological development.

[0003] In recent years, analog compute-in-memory (ACIM) technology for implementing multiply-and-accumulate (MAC) functions has received widespread attention. In ACIM arrays implementing MAC, the larger the absolute value of the MAC operation, the more analog devices involved in the operation, and the more significant the accumulated error. Therefore, achieving a good trade-off between the accuracy and computational density of MAC operations remains a key technical challenge. Summary of the Invention

[0004] This application provides a control method, apparatus, storage system, and electronic device for effectively balancing the accuracy and computational density of multiply-accumulate operations and improving the reliability of in-memory computation.

[0005] In a first aspect, embodiments of this application provide a control method, which includes: First encoded data is obtained by encoding the first data according to the first sequence and the first mapping relationship, wherein the first sequence and the first mapping relationship are related to the distribution characteristics of the first data; The first encoded data is stored in the storage circuit, and the first encoded data is used for in-memory calculation.

[0006] As one possible implementation, the length of the first sequence ranges from 10 to 14.

[0007] As one possible implementation, the first sequence and / or the first mapping relationship is obtained based on a first cost function, which is related to the distribution characteristics of the first data.

[0008] As one possible implementation, the first sequence includes W weight bits, and the first cost function is related to the probability that the first encoded data takes the value 1 at the i-th weight bit of the first sequence, where i is an integer greater than or equal to 0 and less than W.

[0009] As one possible implementation, the absolute value of the weight value of the i-th weight bit is greater than the first threshold.

[0010] As one possible implementation method, The first cost function is expressed as: ; Among them, the For the first cost function, the The probability that the first encoded data takes the value 1 at the i-th weight bit of the first sequence. The probability that the first encoded data takes the value 1 at the j-th weight bit of the first sequence. Let j be the weight value of the i-th weight position in the first sequence, where i is an integer greater than or equal to 0 and less than W, j is an integer greater than or equal to 0 and less than W, and W is the number of weight positions in the first sequence.

[0011] As one possible implementation, the control method further includes: The second data is encoded according to the second sequence and the second mapping relationship to obtain the second encoded data. The second sequence and the second mapping relationship are related to the distribution characteristics of the second data. The second encoded data is input into the storage circuit, and the second encoded data is used to perform in-memory calculations with the first encoded data.

[0012] As one possible implementation, the length of the second sequence ranges from 10 to 14.

[0013] As one possible implementation, the second sequence and / or the second mapping relationship is obtained based on a second cost function, which is related to the distribution characteristics of the second data.

[0014] In one possible implementation, the second sequence includes X weight bits, and the second cost function is related to the probability that the second encoded data takes the value 1 at the nth weight bit of the second sequence, where n is an integer greater than or equal to 0 and less than X.

[0015] As one possible implementation, the absolute value of the weight of the nth weight bit is greater than the second threshold.

[0016] As one possible implementation, the second cost function is expressed as ; Among them, the For the second cost function, the The probability that the second encoded data takes the value 1 at the nth weight bit of the second sequence. The probability that the second encoded data takes the value 1 at the m-th weight bit of the second sequence. Let m be the weight value of the nth weight position in the second sequence, where n is an integer greater than or equal to 0 and less than X, m is an integer greater than or equal to 0 and less than X, and X is the number of weight positions in the second sequence.

[0017] As one possible implementation, one or more of the first sequence, the first mapping relationship, the second sequence, and the second mapping relationship are obtained based on a third cost function, which is related to the distribution characteristics of the first data and the distribution characteristics of the second data.

[0018] As one possible implementation, the third cost function is related to the first probability and the second probability; The first sequence includes W weight bits, and the first probability is the probability that the first encoded data takes the value 1 at the a-th weight bit of the first sequence, where a is an integer greater than or equal to 0 and less than W. The second sequence includes X weight bits, and the second probability is the probability that the second encoded data takes the value 1 at the b-th weight bit of the second sequence, where b is an integer greater than or equal to 0 and less than X.

[0019] As one possible implementation, the absolute value of the weight value of the a-th weight bit is greater than the third threshold; the absolute value of the weight value of the b-th weight bit is greater than the fourth threshold.

[0020] As one possible implementation, the third cost function is expressed as ; Among them, the For the third cost function, The probability that the second encoded data takes the value 1 at the b-th weight bit of the second sequence. The probability that the first encoded data takes the value 1 at the a-th weight bit of the first sequence. This is the weight value of the b-th weight position in the second sequence. Let b be the weight value of the a-th weight position in the first sequence, where a is an integer greater than or equal to 0 and less than W, b is an integer greater than or equal to 0 and less than X, X is the number of weight positions in the second sequence, and W is the number of weight positions in the first sequence.

[0021] In one possible implementation, the probability that the second encoded data is 1 at the k-th weight bit of the second sequence is equal to the probability that the first encoded data is 1 at the k-th weight bit of the first sequence, where k is an integer greater than or equal to 0.

[0022] In a second aspect, embodiments of this application provide a control device, which includes a processing circuit and an interface circuit. The interface circuit is used to signal connect with a storage circuit, and the processing circuit is used to perform the method as described in any of the first aspects.

[0023] Thirdly, embodiments of this application provide a storage system including a storage circuit and a control circuit, the control circuit being configured to perform the method as described in any of the first aspects.

[0024] Fourthly, embodiments of this application also provide an electronic device, including a control device as described in the second aspect or a storage system as described in the third aspect.

[0025] Fifthly, embodiments of this application also provide a cloud device, including a processor and a memory, the memory being used to store a program executable by the processor, the processor being used to read the program in the memory and execute the method as described in any of the first aspects.

[0026] In a sixth aspect, embodiments of this application also provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, is used to implement the steps of any of the methods in the first aspect described above.

[0027] In a seventh aspect, this application provides a computer program product comprising: computer program code that, when run on a computer, causes the computer to perform the method of any one of the first aspects.

[0028] The control method, apparatus, storage system, electronic device, and cloud device provided in this application utilize the first data distribution characteristics to design a first sequence and a first mapping relationship, which can effectively balance the accuracy and computational density of multiplication and accumulation operations and improve the reliability of in-memory computation. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 A schematic diagram of a storage system provided in an embodiment of this application; Figure 2 A schematic diagram of yet another storage system provided in an embodiment of this application; Figure 3 A flowchart illustrating a specific implementation of a control method provided in this application embodiment; Figure 4 This application provides a schematic diagram of the value distribution of first data after quantization, as shown in an embodiment. Figure 5 A probability diagram showing that the first encoded data takes the value "1" at each weight bit, as provided in the embodiments of this application; Figure 6 This application provides a schematic diagram illustrating the relationship between the output signal-to-noise ratio and the input signal-to-noise ratio of a first data source. Figure 7 This is a schematic diagram of a control device provided in an embodiment of this application. Detailed Implementation

[0031] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0032] To keep the drawings concise, the figures in this application only schematically show the parts related to the corresponding embodiments, and they do not represent the actual structure of the product. In addition, to make the drawings concise and easy to understand, some figures only schematically show some structures or components, and there may actually be more or fewer identical or similar structures or components.

[0033] The business scenarios described in the embodiments of this application are for illustrative purposes only and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0034] In this application, unless otherwise expressly specified and limited, "connection" includes direct or indirect connection between objects: connected objects may be directly connected through a medium (e.g., wires, traces, etc.), or indirectly connected through other components, or may be an internal connection. "Coupling" includes signal connection between objects, which may be achieved directly through a medium (e.g., wires, traces, etc.), or through other components. "Grounding" includes direct grounding or indirect grounding, with indirect grounding including, for example, grounding through other components.

[0035] In this application, unless otherwise expressly specified and limited, ordinal numbers, such as "first," "second," etc., are used only to distinguish the objects being described and should not be construed as indicating or implying the relative importance or order between the objects being described. Furthermore, ordinal numbers do not represent the quantity of the objects being described. "Multiple" includes two or more, and other quantifiers are similar. "Or," "and / or," etc., are used to describe the relationship between objects, indicating a non-exclusive inclusion. For example, "A and / or B," "A or B" can include: "A alone," "B alone," or "A and B." Similarly, "A, B, and / or C," "A, B, or C" can include: "A alone," "B alone," "C alone," "A and B," "A and C," "B and C," or "A, B, and C." Additionally, the " / " in this application is used to indicate an "or" relationship between preceding and following objects. The meaning of "one or more of A and B" or "at least one of A and B" in this application is the same as the meaning of "A and / or B" or "A or B" above. "One or more of A, B and C" or "at least one of A, B and C" has the same meaning as "A, B and / or C" or "A, B or C" above.

[0036] In this application, the elements of a material can be represented by the element symbols in a chemical formula, and the subscripts of the elements in the chemical formula are omitted. The values ​​of the subscripts of the elements in the chemical formula can be determined based on factors such as the valence of the elements. Unless explicitly described in the context, this application does not limit the stoichiometry, near-stoichiometry, or impurity / doped materials of the material.

[0037] The positional relationships of the elements, features, or structures shown in the schematic diagrams of the embodiments of this application are for illustrative purposes only. The elements, features, or structures shown in the figures do not necessarily mean that they are at substantially the same height in a certain direction in the device, or necessarily exist in a single cross section of the device, simply because they are shown in the figures.

[0038] With the development of technologies such as big data and artificial intelligence (AI), the volume of data processing and storage is growing rapidly. Computing power requirements and storage pressure have become important factors restricting the development of these technologies. Currently, the application of memory mainly alleviates storage pressure by increasing storage density, but the frequent transfer of data between memory and processor still restricts the improvement of computing power.

[0039] For example, in traditional computing architectures (Von Neumann architecture), storage and computation are physically separated. When using this architecture for data processing, data is frequently transferred between memory and processor, resulting in data transmission latency and energy consumption. With the development of technologies such as big data and artificial intelligence, the amount of data processing is growing rapidly, and the demand for data transmission is also increasing rapidly. The resulting transmission latency and energy consumption are becoming increasingly prominent, restricting the development of data processing capabilities and making it difficult for traditional computing architectures to meet computing power requirements.

[0040] The compute-in-memory (CIM) architecture physically integrates storage and computing, which can reduce transmission latency and energy consumption, thus greatly improving data processing efficiency.

[0041] For ease of understanding, Figure 1 A schematic diagram of a storage system according to an exemplary embodiment of this application is shown. Figure 1 As shown, the storage system 100 may include a storage circuit 110 and a control device 120. The storage circuit 110 is used to store data. For example, in the field of artificial intelligence, the storage circuit can store weight data such as weight factors that characterize the AI ​​model. The control device 120 is used to control the operation of the storage circuit 110, such as controlling one or more operations such as writing (or programming), reading (or retrieving), refreshing, erasing, calculating, or reading the calculation results.

[0042] The storage system 100 can be used in an in-memory computing architecture, and when used in an in-memory computing architecture, the storage system 100 can also be called an in-memory computing system. For example, the storage system 100 can support near-memory computing and / or in-memory computing. For example, in near-memory computing, the control device 120 controls the reading of data stored in the storage circuit 110. In one implementation, the control device 120 integrates a computing circuit that performs calculations based on the read data. In another implementation, the control device 120 can provide the read data to a near-integrated processor for calculation. For example, the chip containing the storage system 100 and the processor performing the calculation are integrated through an interposer. As another example, in in-memory computing, the control device 120 can control the storage circuit 110 to perform in-memory calculations based on the stored data and control the reading of the calculation results. Optionally, the control device 120 can also process the calculation results, such as performing post-processing like compensation. The control device 120 may integrate one or more processing circuits. For example, the control device 120 may include a first processing circuit for controlling the operation of the storage circuit; this first processing circuit may be referred to as a control circuit. Optionally, the control device 120 may include a second processing circuit for performing calculations based on the read data; this second processing circuit may be referred to as a calculation circuit. Optionally, the control device 120 may include an interface circuit configured to provide data to a processor closely integrated with the storage system 100. Optionally, the control device 120 may include a third processing circuit for processing the results of in-memory calculations performed by the storage circuit 110.

[0043] The storage circuit 110 can store data in units of storage cells. For example, the storage circuit 110 may include a storage cell array (hereinafter referred to as the storage array), which includes multiple storage cells arranged in an array. A storage cell may include a semiconductor device, and the storage of data is achieved by utilizing the conductance capability of the semiconductor device, such as conductivity or transconductance. For example, a storage cell may include a resistive storage device or a transistor storage device. For example, data storage can be achieved by controlling the conductivity of a resistive storage device, or by controlling the transconductance of a transistor storage device. Alternatively, the storage cell can utilize the energy stored in an energy storage element to store data, such as by utilizing the charge stored in a capacitor; this energy storage element can be connected to the semiconductor device, and the stored energy can act on the semiconductor device, causing the semiconductor device to generate a corresponding conductance capability.

[0044] When the storage circuit 110 supports in-memory computation, the output terminals of multiple storage cells can be connected collinearly, allowing the output signals of multiple storage cells to be accumulated and output, thus achieving equivalent multiply-accumulate computation. Multiple storage cells with collinearly connected output terminals can be called a computation group. Within the computation group, the multiple data stored in the multiple storage cells can be equivalent to a first data vector / matrix (e.g., a weighting factor vector / matrix). The input data corresponding to the input signal received by the computation group can be equivalent to a second data vector (or input vector). The output of the computation group can be equivalent to the product of the first data vector / matrix and the second data vector. In programming mode, data is written into the storage cells, which is equivalent to writing the first data vector / matrix into the storage circuit 110. In computation mode, the storage circuit 110 receives input signals, and the conduction capability of the storage cells can change the input signal to obtain an output signal. By accumulating and outputting the output signals of multiple storage cells in the computation group, equivalent multiply-accumulate computation can be achieved. The storage array may include a one-dimensional array, a two-dimensional array, or a three-dimensional array, etc. The computing group may include multiple storage units located in the same row or column of the storage array, or multiple storage units located in multiple rows or columns, etc., and the multiple storage units may be output in a collinear manner.

[0045] In some embodiments, the storage system 100 may further include a readout circuit 130. The readout circuit 130 can convert and output the output signal S; for clarity, the converted output signal may be referred to as output data D. This conversion may include one or more conversions such as signal type conversion, signal magnitude conversion, etc., for example, one or more conversions such as current-to-voltage conversion, analog-to-digital conversion, amplification, etc. The control device 120 can also control the operating timing of the readout circuit 130. In some embodiments, by controlling the operating timing, the resources of the readout circuit 130 can be multiplexed from multiple storage arrays or multiple outputs of a storage array, thereby reducing the hardware overhead of peripheral circuits. The output data D can be provided to subsequent circuits (e.g., processing circuits) to process the output data. The processing circuit can be integrated into the control device 120 or can be independent of the control device 120; this application does not impose any limitations. Furthermore, this application does not limit the processing content of the subsequent circuits; different processing operations may exist in different application systems or application devices.

[0046] This application does not limit the type of storage medium used in the storage circuit. The storage medium may include, but is not limited to, non-volatile memory (NVM) or volatile memory (VM). Volatile memory may include, but is not limited to, static random access memory (SRAM) or dynamic random access memory (DRAM). Non-volatile memory may include, but is not limited to, flash memory, resistive random access memory (RRAM), magnetic random access memory (MRAM), ferroelectric memory (FeRAM), or phase change memory (PCM).

[0047] As an example, Figure 2 A schematic diagram of yet another storage system provided in an embodiment of this application is shown. For example... Figure 2 As shown, the storage system 200 includes a storage circuit 210 and a control device 220. The storage circuit 210 may include multiple storage arrays, such as storage arrays A1, A2, etc. Each storage array may include multiple storage cells. For example, storage array A1 is described as an example; other storage arrays are similar. Figure 2 As shown, the storage array A1 includes multiple storage cells S 11 -S mn Where m is the number of rows in storage array A1, and n is the number of columns in storage array A1. Storage cell S ij It can store data, to store data W ij For example, i [1,m],j [1,n]. The control device 220 can control the memory array A1 to be in a programming state in order to send signals to the memory cells S. ij Write data W ij The control device 220 can control the storage array A1 to be in a read state, so as to put the storage cell S... ij Stored data W ij Alternatively, the control device 220 can control the storage array A1 to be in a computing state and read out the computing results of the computing group. In in-memory computing, the storage cell S ij The conduction capability can be based on the stored data W ijIn response to changes, the control device 220 can control the storage array A1 to be in a computing state. The control device 220 can send signals to the storage unit S through the input terminal of the storage array A1. ij Provide input signal V i For example, it can be input via a collinear input line IN; the input signal V i Acting on storage unit S ij Storage unit S ij The stored data W ij This makes the storage unit S ij It has a corresponding conduction capability, under which current flows to or into the output terminal. Multiple memory cells (e.g., S...) 1j -S mj The output terminals of the memory cell can be collinear; for example, they can be collinearly output via the output line OUT. According to Kirchhoff's laws, the current generated by multiple memory cells accumulates to obtain the output signal I. j Satisfy the following formula:

[0048] As can be seen, a storage array can be used to perform multiplication and accumulation calculations, and the calculation results can be output at the output terminal. The multiple storage units with collinear outputs can be called a calculation group.

[0049] The storage system 200 may further include a readout circuit 230 for reading data stored in the storage circuit 210, and / or for reading the above calculation results. For example, the readout circuit 230 may include a first conversion circuit 231 and a second conversion circuit 232. The first conversion circuit 231 performs a first conversion on the output signal of the storage circuit 210 (…). Figure 2 (Identified as Conversion 1), for example, the output signal may include a current signal, and the first conversion circuit 231 can convert the current signal into a voltage signal. The second conversion circuit 232 performs a second conversion on the output of the first conversion (…). Figure 2 The first conversion circuit 231, also known as a sampling circuit, is used to sample the current signal output by the storage circuit 210 as a voltage signal. The second conversion circuit 232, also known as a decision circuit, is used to convert the waveform parameters of the analog signal into digital signals, such as the amplitude, pulse width, or area of ​​the analog signal. This application does not limit the accuracy of the conversion or decision; for example, it may include 1-bit or multi-bit conversion accuracy. The control device 220 can control the operation of the readout circuit 230.

[0050] In some embodiments, different memory arrays may reuse the resources of the first conversion circuit 231 and / or the second conversion circuit 232; or, within the same memory array, different output terminals may reuse the resources of the first conversion circuit 231 and / or the second conversion circuit 232 to reduce the hardware cost of the readout circuit and reduce its footprint on the chip. For example, memory array A1 and memory array A2 may reuse the first conversion circuit 231 and / or the second conversion circuit 232, or some sub-circuits of the first conversion circuit 231 and / or the second conversion circuit 232; as another example, different output terminals in memory array A1 may reuse some sub-circuits of the first conversion circuit 231 and / or the second conversion circuit 232.

[0051] In some embodiments, the storage circuitry may include a three-dimensional structure, such as stacked storage arrays, to achieve greater storage density or capacity. Storage systems with this three-dimensional storage circuitry can provide higher-density storage.

[0052] Figure 2 This is just an example illustrating how storage cells are connected in a storage array, except... Figure 2 Besides the connection method shown, other connection methods can also be used. For example, the input terminals of the memory cells can be connected in columns along a common line, and the output terminals of the memory cells can be connected in rows along a common line. Furthermore, the input terminals of the memory cells can include the gate of a transistor memory device, or the input terminals of the memory cells can include the source or drain of a transistor memory device; this application does not limit this. This application also does not limit the type of memory cell; for example, the memory cell can include, but is not limited to, transistors, memristors, magnetic tunnel junctions (MTJs), or phase-change structures. This application also does not limit the type of transistor, including, for example, metal-oxide-semiconductor field-effect transistors (MOSFETs), floating-gate transistors (FGTs), ferroelectric field-effect transistors (FeFETs), and thin-film transistors.

[0053] In some possible embodiments, the memory cell may include multiple transistors; for example, the memory cell may include a first transistor and a second transistor, wherein the drive terminal (e.g., gate) of the first transistor (which may be referred to as a "read transistor" or "read tube") and the non-drive terminal (e.g., source or drain) of the second transistor are connected, and the charge stored at the drive terminal of the first transistor or the voltage at the drive terminal of the first transistor can be used to characterize the stored data. Optionally, the memory cell may also include a capacitor, and the drive terminal of the first transistor may also be connected to the capacitor to increase the stability and duration of the stored charge or the voltage at the drive terminal. The following description is in conjunction with the accompanying drawings: The exemplary embodiments of this application utilize the aforementioned storage circuit to perform simulated in-memory computation. This storage circuit can serve as a hardware accelerator for multiply-and-accumulate (MAC) intensive operations in Deep Neural Networks (DNNs) and Large Language Models (LLMs). This storage circuit exhibits low data migration and high parallelism, demonstrating significant advantages in energy efficiency and throughput compared to traditional digital architectures. Optionally, the device used for simulating in-memory computation in the storage circuit can employ non-volatile memory (NVM) such as resistive random access memory (RRAM) or 3D NAND flash memory. 3D NAND flash memory offers advantages such as high storage density and controllable cost. To balance device reliability and computational efficiency, the device used for simulating in-memory computation in the storage circuit can also be an NVM-based compute-in-memory (CIM) device, for example, employing a single-level cell (SLC) architecture.

[0054] Taking a storage circuit based on SLC architecture as an example, the storage circuit includes... layer, Column, among which, For integers greater than or equal to 1, It is an integer greater than or equal to 1. The binary input vector of this storage circuit is represented as: The binary weight factors stored on the SLC device (such as model parameters of deep neural networks or large language models) are represented as follows: ;in, , All are integers; the output vector of the storage circuit after MAC operation is defined as... For example, this weighting factor It can be Figure 2 Storage cell S of the central storage array A1ij W written data ij The input vector It can be understood as Figure 2 The signal is transmitted to the memory cell S through the input terminal of the memory array A1. ij The provided input signal V i .

[0055] When implementing analog domain MAC operations using an NVM storage array, the input vector can be equivalent to: A synchronous control switch. When At that time, the corresponding number in the NVM array Rows are selected. In each selected row, only devices with a stored signal value of "1" will output current. The remaining devices remain off (output current value is 0). It is the desired output current. This indicates the error current introduced by device defects or other reasons.

[0056] The total current at each output tap of the storage circuit is And after ideal analog-to-digital conversion (ignoring the measurement error of the ADC), we get: Formula (1); in, It can represent the ideal output value of the read circuit connected to the storage circuit.

[0057] To avoid loss of generality and simplify the analysis, it is assumed that at the same time, all devices are subjected to [a certain condition]. They are mutually independent and follow a mean of 0 and a standard deviation of . The input signal-to-noise ratio follows a normal distribution. The relationship between the intensity of the effective signal and noise in the output current of a single device is characterized by the following definition: Formula (2); From formula (1), the expression for the output error vector E can be derived: Formula (3); Optionally, output signal-to-noise ratio can be used. To measure the accuracy of MAC operation results, this metric can be obtained through statistical analysis at the output.

[0058] To increase computational density, memory circuits can extend the MAC operation depth to 128 layers or more. However, the more devices involved in a single MAC operation, the greater the accumulated error at the output, leading to a decrease in operational reliability. Even with error control coding designed for analog in-memory computing scenarios, the coding will fail and cause a negative performance gain when the error rate exceeds a threshold. Weighted coding can mitigate this problem by reducing device activation rates. For example, the performance of weighted coding is related to coding redundancy and "1"-bit sparsity. The higher the coding redundancy of the coding scheme, the more storage and computational resources it consumes; if the "1"-bit sparsity is insufficient, the high-level bits of effective operations will be densely distributed, making it impossible to significantly reduce the number of actual working devices, making it difficult to adapt to analog domain MAC scenarios, and failing to achieve a good trade-off between multiply-accumulate operation accuracy and computational density.

[0059] Therefore, this application provides a control method that designs a sequence and mapping relationship corresponding to the weight encoding based on the distribution characteristics of the first data, thereby using the first sequence and the first mapping relationship to encode the first data to obtain the first encoded data, which takes into account both "1" bit sparsity and computational density, thereby achieving a good balance between computational accuracy and computational density.

[0060] Figure 3 A flowchart illustrating an implementation of a control method according to an exemplary embodiment of this application is shown. Figure 3 As shown, the specific implementation flow of the control method provided in this application is as follows: Step S300: Encode the first data according to the first sequence and the first mapping relationship to obtain the first encoded data, wherein the first sequence and the first mapping relationship are related to the distribution characteristics of the first data; In this embodiment, the first data can be a weighting factor, such as the model parameters of a neural network or a large model. Before storing the first data in the storage circuit, the first data is encoded, and the encoded first data is stored in the storage circuit. This allows the first encoded data to use fewer bit values ​​"1" to represent values ​​that appear more frequently in the first data, significantly reducing the number of actual working devices during in-memory calculations, reducing device-induced noise, and thus improving computational accuracy.

[0061] Optionally, the first sequence in this embodiment includes W weight bits, each weight bit corresponding to a coding bit. In one example, the first sequence is: ,in, This represents the weight value of the 0th weight position. This represents the weight value of the first weight position, and so on. This represents the weight value of the (W-1)th weight position in the first sequence. There are a total of W weight bits. Each weight bit corresponds to an encoding bit, and the encoding bit takes the value of a binary number {0, 1}. The first data can be represented by the values ​​of each encoding bit in the first sequence, and finally the values ​​of each encoding bit form the binary first encoded data. For example, the encoding bits corresponding to each weight bit in the first sequence are a0, a1, ..., a w-1 Where, the first data = ×a0+ ×a1+ … + ×a w-1 The first encoded data is (a0, a1, ..., a w-1 .

[0062] Optionally, in this embodiment, the first mapping relationship refers to the mapping relationship between the first data and the first encoded data. This first mapping relationship is determined based on the first sequence. The value of the first data is the result of the operation between the first encoded data and the first sequence. When the first sequence is determined, the first encoded data corresponding to different values ​​of the first data can be determined based on the value range of the first data. Optionally, for the same first sequence, the first data can correspond to multiple first encoded data, that is, the first data can obtain multiple first encoded data according to multiple mapping relationships. For example, taking a first sequence including 9 weight bits as an example, the value range of the first data is [-128, 127]. When the first data is -128, the corresponding first encoded data is 110011001. For example, when encoding the first data, the first encoded data corresponding to the current first data can be quickly found through this first mapping relationship.

[0063] The distribution characteristics of the first data may include one or more of the following: the range of values ​​of the first data, the numerical distribution, the central tendency, and the degree of dispersion. For example, the first sequence and the first mapping relationship are related to the frequency of occurrence of the values ​​in the first data.

[0064] For example, the first sequence and first mapping relationship in this application embodiment incorporate the distribution characteristics of the first data during design. When encoding the first data using the first sequence and first mapping relationship, the distribution characteristics of the first data are also introduced. By analyzing the distribution characteristics of the frequency of different values ​​within the range of the first data, fewer bits of "1" are used as much as possible when encoding higher frequency values, thus ensuring the sparsity of "1" in the first encoded data. This ensures the accuracy of the operation even with a fixed bit width of the encoded data.

[0065] Step S301: Store the first encoded data in the storage circuit. The first encoded data is used for in-memory calculation.

[0066] The embodiments of this application can store the first encoded data in the storage circuit to facilitate in-memory calculation. Since the first encoded data of this application can guarantee the sparsity of "1", the distribution of high-level bits of effective operation is sparse when performing in-memory calculation, which greatly reduces the number of actual working devices and is more suitable for analog domain MAC scenarios, achieving a good balance between the accuracy of multiply-accumulate operation and the calculation density.

[0067] In one possible example, in order to balance computational precision and computational density, the length of the first sequence in this embodiment of the application ranges from 10 to 14.

[0068] Optionally, the bit width (total number of encoded bits) of the first encoded data is the same as the length of the first sequence, and the bit width of the first encoded data ranges from 10 to 14. Compared with existing encoding schemes, the encoding method provided in this application does not improve computational accuracy by increasing the bit width of the encoded data. Instead, it designs the first sequence and the first mapping relationship by introducing the distribution characteristics of the first data, using fewer encoded bits "1" to represent the frequently occurring first data, thereby reducing encoding redundancy, improving computational accuracy, ensuring computational density, and effectively balancing computational accuracy and computational density.

[0069] In one possible example, the first sequence and / or the first mapping relationship in this application embodiment are obtained based on a first cost function, which is related to the distribution characteristics of the first data.

[0070] According to an embodiment of this application, a cost function is designed to optimize the first sequence and the first mapping relationship, and the distribution characteristics of the first data are introduced into the cost function so that the distribution characteristics of the first data are considered during the optimization process, thereby ensuring the sparsity of "1" in the first encoded data and improving the computational accuracy.

[0071] In one possible example, the first sequence in this application embodiment includes W weight bits, and the first cost function is related to the probability that the first encoded data takes the value 1 at the i-th weight bit of the first sequence, where i is an integer greater than or equal to 0 and less than W.

[0072] In this embodiment, the coding bits in the first coding data correspond one-to-one with the weight bits of the first sequence. One weight bit corresponds to one coding bit. The probability that the first coding data takes the value of 1 at the i-th weight bit of the first sequence can be understood as the probability that the i-th coding bit in the first coding data takes the value of 1. That is, the first cost function in this embodiment is related to the probability that the i-th coding bit in the first coding data takes the value of 1.

[0073] It should be noted that since the first encoded data is obtained by encoding the first data, the probability that the i-th encoded bit in the first encoded data is 1 is related to the distribution characteristics of the first data. Since the first cost function is related to the probability that the i-th encoded bit in the first encoded data is 1, the first cost function is also related to the distribution characteristics of the first data. For values ​​that appear frequently in the first data, when encoding with the first encoded data, try to use as few encoded bits as possible that are "1".

[0074] In one example, the absolute value of the weight value of the i-th weight bit in this embodiment is greater than a first threshold. Optionally, the design of the first threshold is related to the computational precision of in-memory computation. For example, if the first threshold is 8, the weight value will only have a significant impact on the computational precision of simulated in-memory computation when the absolute value of the weight value of the i-th weight bit in the first sequence is greater than 8.

[0075] In one example, the first cost function in this embodiment can be expressed as: Formula (4); In formula (4), For the first cost function, Let be the probability that the first encoded data takes the value 1 at the i-th weight bit of the first sequence. Let be the probability that the first encoded data takes the value 1 at the j-th weight bit of the first sequence. Let be the weight value of the i-th weight position in the first sequence, where i is an integer greater than or equal to 0 and less than W, and j is an integer greater than or equal to 0 and less than W, where W is the number of weight positions in the first sequence.

[0076] In one example, the first cost function in this embodiment of the application can also be expressed as: Formula (5); In formula (5), Indicates only when When, function The value is 1, otherwise the function The value is 0. This represents the first threshold.

[0077] When calculating the first cost function, only the probability of the value of the corresponding encoding bit at the weight bit with higher weight value in the first sequence is considered, which can reduce the amount of computation while ensuring the accuracy of the operation.

[0078] For example, in formulas (4) and (5) It can also be replaced with , This is the weight value of the j-th weight position in the first sequence.

[0079] Based on the same encoding principle as the first data described above, an exemplary embodiment of this application also provides an encoding method for the second data, which can be used as input to a storage circuit. The second data is encoded before being input, and the second encoded data is input to the storage circuit for in-memory calculation with the first encoded data stored in the storage circuit.

[0080] The encoding method in this application can be understood as encoding data through an encoding sequence and a mapping relationship. Optionally, the encoding method may also include obtaining the encoding sequence and mapping relationship through a cost function.

[0081] For example, the specific implementation process for the second data processing is as follows: The second data is encoded according to the second sequence and the second mapping relationship to obtain the second encoded data. The second sequence and the second mapping relationship are related to the distribution characteristics of the second data. The second encoded data is input into the storage circuit, and the second encoded data is used to perform in-memory calculations with the first encoded data.

[0082] In this embodiment, the second data can be input data of the storage circuit. The second data can be a binary input vector. The input second data is used to perform in-memory calculations (such as multiplication and accumulation) with the first data stored in the storage circuit. Before being input into the storage circuit, the second data is encoded so that the encoded second data uses fewer bit values ​​"1" to represent the values ​​that occur more frequently in the second data, which greatly reduces the number of actual working devices, reduces device-induced noise, and thus improves the calculation accuracy.

[0083] Optionally, the second sequence in this embodiment includes X weight bits, each weight bit corresponding to one encoding bit. In one example, the second sequence is: ,in, This represents the weight value of the 0th weight position. This represents the weight value of the first weight position, and so on. This represents the weight value of the (X-1)th weight position in the second sequence. There are a total of X weight bits. Each weight bit corresponds to an encoding bit, and the encoding bit takes the value of a binary number {0, 1}. The second data can be represented by the values ​​of each encoding bit in the second sequence, and finally the values ​​of each encoding bit form the binary second encoded data. For example, the encoding bits corresponding to each weight bit in the second sequence are b0, b1, ..., b X-1 ;wherein, the second data = ×b0+ ×b1+ … + ×b X-1The second encoded data is (b0, b1, ..., b X-1 .

[0084] Optionally, the second mapping relationship in this embodiment refers to the mapping relationship between the second data and the second encoded data. This second mapping relationship is determined based on the second sequence. The value of the second data is the result of the operation between the second encoded data and the second sequence. When the second sequence is determined, the second encoded data corresponding to different values ​​of the second data can be determined based on the value range of the second data. Optionally, for the same second sequence, the second data can correspond to multiple second encoded data, that is, the second data can obtain multiple second encoded data according to multiple mapping relationships. For example, taking a second sequence including 9 weight bits as an example, the value range of the second data is [-128, 127]. When the second data is -128, the corresponding second encoded data is 110011001. For example, when encoding the second data, the second encoded data corresponding to the current second data can be quickly found through this second mapping relationship.

[0085] The distribution characteristics of the second data may include one or more of the following: the range of values, numerical distribution, central tendency, and dispersion of the second data. For example, the second sequence and the second mapping relationship are related to the frequency of occurrence of the values ​​in the second data.

[0086] For example, the second sequence and the second mapping relationship in this application embodiment incorporate the distribution characteristics of the second data during design. Similarly, when encoding the second data using the second sequence and the second mapping relationship, the distribution characteristics of the second data are also introduced. By analyzing the distribution characteristics of the frequency of different values ​​within the range of the second data, fewer bits of "1" are used when encoding higher-frequency values, ensuring the sparsity of "1" in the second encoded data. This guarantees computational accuracy even with a fixed encoded data bit width.

[0087] Since both the first and second encoded data in this application can guarantee the sparsity of "1", when performing in-memory calculations on the first and second encoded data in the storage circuit, the number of high-level bits for effective operations can be reduced, thereby reducing the number of actual working devices and achieving a good balance between the accuracy of multiply-accumulate operations and the computational density.

[0088] In one possible example, in order to balance computational precision and computational density, the length of the second sequence in this embodiment of the application ranges from 10 to 14.

[0089] Optionally, the bit width (total number of encoded bits) of the second encoded data is the same as the length of the second sequence, and the bit width of the second encoded data ranges from 10 to 14. Compared with existing encoding schemes, the encoding method provided in this application does not improve computational accuracy by increasing the bit width of the encoded data. Instead, it designs the second sequence and the second mapping relationship by introducing the distribution characteristics of the second data, using fewer encoded bits "1" to represent the frequently occurring second data, thereby reducing encoding redundancy, improving computational accuracy, and ensuring computational density, thus effectively balancing computational accuracy and computational density.

[0090] In one possible example, the second sequence and / or the second mapping relationship in the embodiments of this application are obtained based on a second cost function, which is related to the distribution characteristics of the second data.

[0091] According to an embodiment of this application, the second sequence and the second mapping relationship are optimized by designing a cost function, and the distribution characteristics of the second data are introduced into the cost function so that the distribution characteristics of the second data are considered during the optimization process, thereby ensuring the sparsity of "1" in the second encoded data and improving the computational accuracy.

[0092] In one possible example, the second sequence in this application embodiment includes X weight bits, and the second cost function is related to the probability that the second encoded data takes the value 1 at the nth weight bit of the second sequence, where n is an integer greater than or equal to 0 and less than X.

[0093] In this embodiment, the coding bits in the second coding data correspond one-to-one with the weight bits of the second sequence. One weight bit corresponds to one coding bit. The probability that the second coding data takes the value of 1 at the nth weight bit of the second sequence can be understood as the probability that the nth coding bit in the second coding data takes the value of 1. That is, the second cost function in this embodiment is related to the probability that the nth coding bit in the second coding data takes the value of 1.

[0094] It should be noted that since the second encoded data is obtained by encoding the second data, the probability of the nth encoded bit in the second encoded data being 1 is related to the distribution characteristics of the second data. Since the second cost function is related to the probability of the nth encoded bit in the second encoded data being 1, it is also related to the distribution characteristics of the second data. For values ​​that appear frequently in the second data, when encoding with the second encoded data, try to use as few encoded bits as possible that are "1".

[0095] In one example, the absolute value of the weight of the nth weight bit in this embodiment is greater than the second threshold. Optionally, the design of the second threshold is related to the computational precision of in-memory computation. The second threshold can be the same as or different from the first threshold. For example, if the second threshold is 8, the impact of the weight value on the computational precision of the simulated in-memory computation will be significant only when the absolute value of the weight of the nth weight bit in the second sequence is greater than 8.

[0096] In one example, the second cost function in this embodiment can be expressed as: Formula (6); In formula (6), For the second cost function, The probability that the second encoded data takes the value 1 at the nth weight bit of the second sequence. Let m be the probability that the second encoded data takes the value 1 at the m-th weight bit of the second sequence. This represents the weight value of the nth weight position in the second sequence, where n is an integer greater than or equal to 0 and less than X, m is an integer greater than or equal to 0 and less than X, and X is the number of weight positions in the second sequence.

[0097] In one example, the second cost function in this embodiment of the application can also be expressed as: Formula (7); In formula (7), Indicates only when When, function The value is 1, otherwise the function The value is 0. This indicates the second threshold.

[0098] When calculating the second cost function, only the probability of the value of the corresponding encoding bit at the weight bit with higher weight value in the second sequence is considered, which can reduce the amount of computation while ensuring the accuracy of the operation.

[0099] For example, in formulas (6) and (7) It can also be replaced with , This is the weight value of the m-th weight position in the second sequence.

[0100] In the embodiments of this application, the first data can be encoded only according to the first sequence and the first mapping relationship as described above, or the second data can be encoded only according to the second sequence and the second mapping relationship as described above, or both the first data and the second data can be encoded according to the first sequence and the first mapping relationship as described above. It can also be understood that, in the embodiments of this application, the first sequence and the first mapping relationship can be obtained solely based on the first cost function and used to encode the first data; the second sequence and the second mapping relationship can be obtained solely based on the second cost function and used to encode the second data; or the first data can be encoded after obtaining the first sequence and the first mapping relationship based on the first cost function, and the second data can be encoded after obtaining the second sequence and the second mapping relationship based on the second cost function.

[0101] Optionally, when both the first data and the second data are encoded according to the first sequence and the first mapping relationship as described above, this application can also obtain one or more of the first sequence, the first mapping relationship, the second sequence, and the second mapping relationship based on the third cost function through a joint encoding method.

[0102] An exemplary embodiment of this application also provides a joint encoding method, which jointly encodes first data and second data, stores the encoded first coded data in a storage circuit, inputs the encoded second coded data into the storage circuit, and the storage circuit performs in-memory calculations on the input second coded data and first coded data before outputting the result. The specific implementation process is as follows: First encoded data is obtained by encoding first data according to a first sequence and a first mapping relationship, and second encoded data is obtained by encoding second data according to a second sequence and a second mapping relationship. The first sequence and the first mapping relationship are related to the distribution characteristics of the first data; the second sequence and the second mapping relationship are related to the distribution characteristics of the second data. The first encoded data is stored in the storage circuit, and the second encoded data is input into the storage circuit. The second encoded data is used to perform in-memory calculations with the first encoded data.

[0103] Optionally, the control storage circuit performs in-memory calculations on the first encoded data and the second encoded data.

[0104] In one possible example, one or more of the first sequence, the first mapping relationship, the second sequence, and the second mapping relationship in the embodiments of this application are obtained based on a third cost function, which is related to the distribution characteristics of the first data and the distribution characteristics of the second data.

[0105] According to the embodiments of this application, the first sequence and the first mapping relationship, the second sequence and the second mapping relationship are optimized by designing a cost function, and the distribution characteristics of the first data and the second data are introduced into the cost function, so that the distribution characteristics of the first data and the second data are considered in the optimization process, thereby ensuring the sparsity of "1" in the first coded data and the second coded data used for in-memory computation, improving the computational accuracy, ensuring the computational density, and effectively balancing the computational accuracy and computational density.

[0106] In one possible example, the third cost function in this application embodiment is related to the first probability and the second probability; The first sequence includes W weight bits, and the first probability is the probability that the first encoded data takes the value 1 at the a-th weight bit of the first sequence, where a is an integer greater than or equal to 0 and less than W. The second sequence includes X weight bits, and the second probability is the probability that the second encoded data takes the value 1 at the b-th weight bit of the second sequence, where b is an integer greater than or equal to 0 and less than X.

[0107] In this embodiment, the coded bits in the first coded data correspond one-to-one with the weight bits of the first sequence, with one weight bit corresponding to one coded bit. The probability that the first coded data takes a value of 1 at the a-th weight bit of the first sequence can be understood as the probability that the a-th coded bit in the first coded data takes a value of 1. That is, the third cost function in this embodiment is related to the probability that the a-th coded bit in the first coded data takes a value of 1. Similarly, in this embodiment, the coded bits in the second coded data correspond one-to-one with the weight bits of the second sequence. The probability that the second coded data takes a value of 1 at the b-th weight bit of the second sequence can be understood as the probability that the b-th coded bit in the second coded data takes a value of 1. That is, the third cost function in this embodiment is related to the probability that the b-th coded bit in the second coded data takes a value of 1.

[0108] In one example, the absolute value of the weight value of the a-th weight bit in this embodiment is greater than the third threshold; the absolute value of the weight value of the b-th weight bit is greater than the fourth threshold. Optionally, the third threshold and the first threshold may be the same or different, and the fourth threshold and the second threshold may be the same or different.

[0109] In one example, the third cost function in this embodiment can be expressed as: Formula (8); In formula (8), For the third cost function, Let b be the probability that the second encoded data takes the value 1 at the b-th weight bit of the second sequence. Let be the probability that the first encoded data takes the value 1 at the a-th weight bit in the first sequence. Let b be the weight value of the b-th weight position in the second sequence. Let be the weight value of the a-th weight position in the first sequence, where a is an integer greater than or equal to 0 and less than W, b is an integer greater than or equal to 0 and less than X, X is the number of weight positions in the second sequence, and W is the number of weight positions in the first sequence.

[0110] In one example, the third cost function in this embodiment can also be expressed as: Formula (9); In formula (9), Indicated as when and When, function The value is 1, otherwise the function The value is 0. This represents the third threshold. This represents the fourth threshold.

[0111] In one possible example, the probability that the second encoded data has a value of 1 at the k-th weight bit of the second sequence is equal to the probability that the first encoded data has a value of 1 at the k-th weight bit of the first sequence, where k is an integer greater than or equal to 0. Since the probability distribution of different encoded data having a value of 1 at the k-th weight bit of their respective sequences is approximately the same, to reduce the computational difficulty of joint optimization, when calculating the first cost function of the first data, it is assumed that the second data uses the same encoding scheme as the first data. Based on this, the third cost function can also be expressed as: Formula (10); Formula (10) above can be understood as optimizing the encoding methods of the first and second data respectively, or as designing the first sequence and the first mapping relationship through the cost function, and designing the second sequence and the second mapping relationship respectively. The first term in formula (10) Equivalent to the first cost function mentioned above, it can be described by combining the descriptions of formulas (4) and (5); for example, in the first term... It can also be replaced with The second term in formula (10) Equivalent to the second cost function described above, it can be described by combining the descriptions of formulas (6) and (7); for example, in the second term... It can also be replaced with .

[0112] In this embodiment, the approximate expected value of the MAC operation of the higher weight bit in the first and second sequences is used as the output to evaluate the average error level of the MAC operation. For example, for an L-layer analog MAC circuit using SLC memory mode, the expected value of the MAC operation can be expressed as follows: Formula (11); In formula (11), the constant L only has a fixed scaling effect on the output of the expression. Other variables can be found in the explanation in formula (8), and will not be repeated here. Optionally, the constant L can be ignored.

[0113] It should be noted that the different subscripts used for the same parameter in different formulas in the embodiments of this application are merely examples, representing different parameter labels. The physical meaning expressed by using different subscripts for the same parameter is equivalent. For example, in formula (11) and in formula (5) The physical meanings are equivalent, and the formula (11) contains... and in formula (6) The physical meanings are equivalent.

[0114] After determining the cost function (first cost function, second cost function, third cost function) through any of the methods described above in this application, a first sequence and / or a second sequence can be constructed based on the cost function. The construction methods include, but are not limited to, Differential Evolution (DE) algorithm, mathematical model, neural network model, etc. The embodiments of this application do not impose too many limitations on this.

[0115] Based on the aforementioned cost function, an exemplary embodiment of this application provides a method for constructing sequence and mapping relationships based on the idea of ​​differential evolution. Taking the encoding of the first data as an example, the construction process of this method can be represented in pseudocode as follows: initialization: Set the initial first sequence The initial first sequence ; Establish the first mapping relationship between integer values ​​in the interval [-128, 127] and the first sequence; The first cost function associated with the initial first sequence is calculated using formula (5). .

[0116] The iterative process is as follows: For ; Traversing the first sequence For all legal values, create the corresponding first mapping relationship and calculate the first cost function associated with the first sequence; Record the value that minimizes the first cost function. and the second smallest ; with probability choose (in terms of probability) choose ),replace The Middle The weight of each weight position.

[0117] End

[0118] Based on the updated first sequence Create the latest first mapping relationship.

[0119] In the above construction process, the first sequence A valid value is one that is different from the first sequence. The first sequence is selected from the other elements in the interval [-128, 127] to ensure that it completely represents all integers within the interval. Furthermore, when constructing the first mapping relationship, if multiple feasible mapping schemes exist for a given integer, the first cost function corresponding to each scheme can be calculated separately, and the first sequence and first mapping relationship with the smallest first cost function can be selected as the output. Finally, through the above iterations, the optimal first sequence and first mapping relationship are output, and the first data is encoded using this optimal first sequence and first mapping relationship to obtain the first encoded data.

[0120] For example, to avoid the algorithm getting trapped in local optima, when choosing replacements... When determining the weight of the i-th weight position, it is necessary to select it with probability p. (Choose with probability (1-p)) For example, another strategy to escape local optima is to use several significantly different initial first sequences. Make multiple attempts to improve global search capabilities.

[0121] The construction process of the second sequence in this embodiment can refer to the construction method based on the differential evolution idea described above. It is similar to the construction process of the first sequence and will not be repeated here.

[0122] For example, if the first data and the second data are jointly encoded, when constructing the first sequence and the second sequence based on the third cost function, the second sequence can be fixed first, and the optimal first sequence and the first mapping relationship can be found using the above differential evolution idea. Then, the first sequence can be fixed, and the optimal second sequence and the second mapping relationship can be found using the above differential evolution idea.

[0123] In one possible example, in the first encoded data of this application embodiment, as the bit order increases, the probability of a 1 appearing in the bit order decreases at least partially, where the bit order is the position of the binary value in the first encoded data; or, in the second encoded data, as the bit order increases, the probability of a 1 appearing in the bit order decreases at least partially, where the bit order is the position of the binary value in the second encoded data. It is easy to understand that the bit order in the embodiments of this application refers to the arrangement order of the encoded bits in the first or second encoded data. For example, if the first encoded data is 110011001000, then as the bit order increases in this first encoded data, the probability of a 1 appearing in the bit order decreases. This decreasing trend is not absolutely monotonically decreasing, but rather mostly decreasing.

[0124] Simulations were performed based on a fundamental model of the MAC circuit, with the number of storage circuit layers set to l=128, and the influence of ADC measurement errors was ignored. Both the first and second data points were taken from the up-dimensional portion of the FFN (FeedForward Network) in the 26th layer of the Llama 2.7B model's Transformer. After quantization, both exhibited distribution characteristics typical of deep Transformer structures. Figure 4 The diagram illustrates the value distribution of quantized first data according to an exemplary embodiment of this application. The horizontal axis represents the value of the quantized first data, which is in the range of [-128, 127]. The vertical axis represents the probability of the corresponding value occurring. The quantized first data follows a normal-like distribution, with the vast majority of quantized values ​​concentrated around 0, which conforms to the classic distribution law of deep learning weights.

[0125] An exemplary embodiment of this application provides an example of a first sequence. Taking the lengths of the first sequence as 9, 10, and 12 as examples, first data codes with bit widths of 9, 10, and 12 are generated using the first sequence and a first mapping relationship, respectively. The corresponding first sequences are: (-1, -2, 4, 8, -12, -27, 30, 85, -86), (-1, -2, 4, 8, -12, 12, -26, 28, 86, -88), and (-1, 2, 4, -4, 8, -8, 16, -16, 32, 32, 67, -67). This application embodiment provides examples of different encoding schemes. For example, the first data code bit width includes 9 bits, 10 bits, and 12 bits, and the value range of the first data code is within the interval [-128, 127]. In the first mapping relationship, the mapped first encoded data is represented in unsigned hexadecimal format. To ensure uniform compatibility with the 12-bit bit width specification, three and two binary zeros need to be added to the end of the 9-bit and 10-bit binary codes, respectively. For example, the 9-bit binary code for the value -128 is 110011001. After adding three zeros to the end, it becomes 110011001000, and the corresponding hexadecimal value of this binary number is CC8.

[0126] The encoding method used in this application embodiment makes the value range of the first data with medium to high amplitude significantly smaller than that of the traditional weight sequence constructed based on powers of 2. Statistical analysis shows that when using these three encoding schemes, the probability of the first encoded data taking a value of "1" at each weight bit is as follows: Figure 5 As shown, the horizontal axis represents the bit sequence number (bit order) of the first encoded data binary, and the vertical axis represents the statistical probability that the value of the bit in that bit order is 1. Figure 5 The paper also provides the classic 8-bit two's complement representation as a baseline for comparison. Under the 8-bit two's complement scheme, the probability of all weighted bits being "1" is close to 50%, showing a balanced distribution. In contrast, in the 10-bit and 12-bit encoding schemes, the probability of a bit being "1" in the first encoded data decreases significantly as the bit order increases.

[0127] Furthermore, in the 8-bit two's complement scheme, each integer requires an average of approximately 4 bits for encoding, while the three encoding schemes with bit widths of 9, 10, and 12 require only approximately 3.17, 2.98, and 2.54 bits, respectively. These encoding characteristics, combined with the smaller encoding amplitudes corresponding to the medium-to-high weight bits where the weight value is greater than the first threshold, effectively control the MAC operation error on the high-weight bits.

[0128] Figure 6This embodiment of the application illustrates the relationship between the output signal-to-noise ratio (SNR)_out and the input signal-to-noise ratio (SNR)_in when the first data adopts an encoding scheme with different bit widths, provided that the second data is fixed as 12-bit weighted encoding. Figure 6 The range of values ​​for SNR_in covers the typical intervals of NAND media health status from the early to the late stages. Figure 6 It can be seen that, compared to the 8-bit two's complement scheme, the SNR_out of the other encoding methods increases significantly with SNR_in. Combined with... Figure 5 Compared to the 8-bit two's complement scheme, in other bit encoding schemes (such as 10-bit and 12-bit), the probability of a 1 appearing in the first encoded data decreases significantly with increasing bit order. This means that other bit encoding schemes can significantly reduce the number of actual working devices, thereby ensuring computational accuracy and improving the output signal-to-noise ratio. Figure 6 Simulation results also verified that the output signal-to-noise ratio of the other bit encoding schemes increased significantly, indicating that the encoding scheme of this application can effectively reduce noise and improve computational accuracy.

[0129] Figure 7 A schematic diagram of a control device according to an exemplary embodiment of this application is shown. The control device 700 includes a processing circuit 710 and an interface circuit 720, the interface circuit 720 being signal-connected to a storage circuit, and the processing circuit 710 being used to execute any of the control methods provided in the above embodiments.

[0130] In some possible embodiments, the processing circuit is a circuit with signal processing capabilities; for example, it may be a circuit with instruction reading and execution capabilities. In other possible embodiments, the processing circuit can implement its functions through the logical relationships of hardware circuits, which are fixed or reconfigurable. For example, the processing circuit may be a hardware circuit implemented by an ASIC or PLD, such as a field-programmable gate array (FPGA). In a reconfigurable hardware circuit, the process of the processing circuit loading a configuration document and configuring the hardware circuit can be understood as the process of the processing circuit loading instructions to implement the functions of some or all of the above units. This application does not limit the type of processing circuit, including, for example, a central processing unit (CPU), a microcontroller unit (MCU), a graphics processing unit (GPU), or a digital signal processor. Alternatively, it may be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), or a deep learning processing unit (DPU).

[0131] In some possible embodiments, the units in the above control device may be integrated in whole or in part, or may be implemented independently. In some embodiments, these units are integrated together and implemented as a system on chip (SOC).

[0132] This application also provides a storage system, including a storage circuit and a control circuit, the control circuit being configured to execute any of the control methods provided in the above embodiments.

[0133] This application also provides an electronic device, which includes any of the control devices or storage systems provided in the above embodiments.

[0134] This application does not limit the type of electronic device. For example, according to some embodiments, the electronic device may include wearable devices. Wearable devices include, but are not limited to: head-mounted devices (e.g., helmets or hats), devices worn on the ears (e.g., headphones), devices worn on the wrist (e.g., watches), and devices worn on other parts of the body (e.g., electronic necklaces, medical monitoring devices, or glasses). According to some embodiments, the electronic device may include portable terminals. For example, the electronic device may include, but is not limited to, mobile phones, general-purpose computing devices (e.g., laptops or tablets), personal digital assistants, etc. According to some embodiments, the electronic device may include other types of edge devices, such as personal computers, in-vehicle computers or in-vehicle computing platforms, or smart home electronic products. According to some embodiments, the electronic device may also include devices such as servers.

[0135] This application also provides a cloud device, which includes a processor and a memory. The memory is used to store a program executable by the processor, and the processor is used to read the program in the memory and execute any of the control methods provided in the above embodiments.

[0136] This application also provides a computer program product, which includes instructions that, when executed by a processor, cause any of the control methods described in the above embodiments to be executed.

[0137] This application also provides a computer-readable medium storing instructions that, when executed by a processor, cause any of the control methods described in the above embodiments to be executed.

[0138] In the above method embodiments, the order of the process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0139] In the above embodiments, the descriptions of different embodiments each have their own emphasis. Parts not described in detail or recorded in a certain embodiment can be referred to in the relevant descriptions of other embodiments. Furthermore, the different embodiments described above can be freely combined as needed. Moreover, as technology evolves, the elements described in this application can be replaced by equivalent elements appearing after this application.

Claims

1. A control method, characterized in that, The control method includes: First encoded data is obtained by encoding the first data according to the first sequence and the first mapping relationship, wherein the first sequence and the first mapping relationship are related to the distribution characteristics of the first data; The first encoded data is stored in the storage circuit, and the first encoded data is used for in-memory calculation.

2. The method according to claim 1, characterized in that, The length of the first sequence ranges from 10 to 14.

3. The method according to claim 1 or 2, characterized in that, The first sequence and / or the first mapping relationship are obtained based on a first cost function, which is related to the distribution characteristics of the first data.

4. The method according to claim 3, characterized in that, The first sequence includes W weight bits, and the first cost function is related to the probability that the first encoded data takes the value 1 in the i-th weight bit of the first sequence, where i is an integer greater than or equal to 0 and less than W.

5. The method according to claim 4, characterized in that, The absolute value of the weight of the i-th weight bit is greater than the first threshold.

6. The method according to any one of claims 3-5, characterized in that, The first cost function is expressed as ; Among them, the For the first cost function, the The probability that the first encoded data takes a value of 1 at the i-th weight bit of the first sequence, the The probability that the first encoded data takes a value of 1 at the j-th weight bit of the first sequence, the Let j be the weight value of the i-th weight position in the first sequence, where i is an integer greater than or equal to 0 and less than W, j is an integer greater than or equal to 0 and less than W, and W is the number of weight positions in the first sequence.

7. The method according to claim 1, characterized in that, The control method further includes: The second data is encoded according to the second sequence and the second mapping relationship to obtain the second encoded data. The second sequence and the second mapping relationship are related to the distribution characteristics of the second data. The second encoded data is input into the storage circuit, and the second encoded data is used to perform in-memory calculations with the first encoded data.

8. The method according to claim 7, characterized in that, The length of the second sequence ranges from 10 to 14.

9. The method according to claim 7 or 8, characterized in that, The second sequence and / or the second mapping relationship are obtained based on the second cost function, which is related to the distribution characteristics of the second data.

10. The method according to claim 9, characterized in that, The second sequence includes X weight bits, and the second cost function is related to the probability that the second encoded data takes the value 1 at the nth weight bit of the second sequence, where n is an integer greater than or equal to 0 and less than X.

11. The method according to claim 10, characterized in that, The absolute value of the weight of the nth weight bit is greater than the second threshold.

12. The method according to any one of claims 9-11, characterized in that, The second cost function is expressed as ; Among them, the For the second cost function, the The probability that the second encoded data takes the value 1 at the nth weight bit of the second sequence, the The probability that the second encoded data takes the value 1 at the m-th weight bit of the second sequence, the The weight value of the nth weight position in the second sequence is given by m, where n is an integer greater than or equal to 0 and less than X, m is an integer greater than or equal to 0 and less than X, and X is the number of weight positions in the second sequence.

13. The method according to claim 7 or 8, characterized in that, One or more of the first sequence, the first mapping relationship, the second sequence, and the second mapping relationship are obtained based on a third cost function, which is related to the distribution characteristics of the first data and the distribution characteristics of the second data.

14. The method according to claim 13, characterized in that, The third cost function is related to the first probability and the second probability; The first sequence includes W weight bits, and the first probability is the probability that the first encoded data takes the value 1 at the a-th weight bit of the first sequence, where a is an integer greater than or equal to 0 and less than W. The second sequence includes X weight bits, and the second probability is the probability that the second encoded data takes the value 1 at the b-th weight bit of the second sequence, where b is an integer greater than or equal to 0 and less than X.

15. The method according to claim 14, characterized in that, The absolute value of the weight of the a-th weight position is greater than the third threshold; the absolute value of the weight of the b-th weight position is greater than the fourth threshold.

16. The method according to any one of claims 13-15, characterized in that, The third cost function is expressed as follows: ; Among them, the For the third cost function, the The probability that the second encoded data takes a value of 1 at the b-th weight bit of the second sequence, the The probability that the first encoded data takes a value of 1 at the a-th weight bit of the first sequence, the The weight value of the b-th weight position in the second sequence, Let b be the weight value of the a-th weight position in the first sequence, where a is an integer greater than or equal to 0 and less than W, b is an integer greater than or equal to 0 and less than X, X is the number of weight positions in the second sequence, and W is the number of weight positions in the first sequence.

17. The method according to claim 16, characterized in that, The probability that the second encoded data is 1 at the kth weight bit of the second sequence is equal to the probability that the first encoded data is 1 at the kth weight bit of the first sequence, where k is an integer greater than or equal to 0.

18. A control device, characterized in that, The control device includes a processing circuit and an interface circuit. The interface circuit is used to connect to the storage circuit via signals. The processing circuit is used to execute any one of the control methods of claims 1-17.

19. A storage system, characterized in that, It includes a storage circuit and a control circuit, the control circuit being configured to perform any one of the control methods of claims 1-17.

20. An electronic device, characterized in that, include: The control device as described in claim 18 or the storage system as described in claim 19.