Storage device for performing asymmetric compression and decompression and method of operating same

By combining reversible and irreversible neural network models, potential data representations and distribution parameters are generated, solving the speed problem of storage devices when processing asymmetric requests and improving the processing speed of read requests and processor performance.

CN121918752APending Publication Date: 2026-04-24SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-04-15
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing storage devices suffer from slow read request processing speeds when handling asymmetric write and read requests, impacting processor performance.

Method used

By combining reversible and irreversible neural network models, and generating latent data representations and distribution parameters, data compression and partial decompression are achieved, thereby optimizing the processing speed of storage devices.

Benefits of technology

This improves the speed at which storage devices process read requests, enhances processor performance, and enables efficient data storage and retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121918752A_ABST
    Figure CN121918752A_ABST
Patent Text Reader

Abstract

A memory controller includes one or more processors including processing circuitry, and a memory storing instructions. When the instructions are executed individually or collectively by the one or more processors, the memory controller is caused to: generate potential data representations corresponding to data values to be written to the storage space by using a reversible neural network model; generating a distribution parameter corresponding to the potential data representation by using an irreversible neural network model; compressing the potential data representation based on the distributed parameters; partially decompressing at least one potential data representation corresponding to at least one data value among the compressed potential data representations based on a command from the host device to provide the at least one data value among the data values written to the storage space; and providing the at least one data value to the host device.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims the benefit of priority to Korean Patent Application No. 10-2024-0145947 filed with the Korean Intellectual Property Office on October 23, 2024, and Korean Patent Application No. 10-2025-0034166 filed with the Korean Intellectual Property Office on March 17, 2025, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0002] This disclosure generally relates to storage devices, and more specifically, to storage devices for performing asymmetric compression and decompression, and methods of operating thereof. Background Technology

[0003] Computational Fast Link (CXL) can refer to a relatively high-speed interconnect technology that can be used for relatively high-performance computing. CXL can provide a relatively fast transfer environment between devices (such as, but not limited to, processors, memory, or accelerators). CXL-based memory can provide the ability to compress and / or store data.

[0004] Deep learning-based neural networks have recently been applied to various fields (such as, but not limited to, data compression). Deep learning-based neural networks can be trained using deep learning and can perform inference for a desired purpose by mapping input and output data that can have non-linear relationships with each other. This ability to generate such mappings through training can be referred to as the learning capability of deep learning-based neural networks. Summary of the Invention

[0005] According to one aspect of this disclosure, a memory controller includes one or more processors including processing circuitry, and a memory storing instructions. When the instructions are executed individually or jointly by the one or more processors, the memory controller causes to perform the following operations: generating a latent data representation corresponding to a data value to be written to memory using a reversible neural network model; generating distribution parameters corresponding to the latent data representation using an irreversible neural network model; compressing the latent data representation based on the distribution parameters; partially decompressing at least one latent data representation corresponding to the at least one data value in the compressed latent data representation based on a command from a host device for providing at least one data value among the data values ​​to be written to memory; and providing the at least one data value to the host device.

[0006] According to one aspect of this disclosure, a storage device includes a memory array configured to store compressed data values, one or more processors including processing circuitry, and a memory storing instructions. When the instructions are executed individually or jointly by the one or more processors, the storage device performs the following operations: generating a latent data representation corresponding to the data values ​​by performing domain transformations to reduce the entropy of the data values ​​to be written to the storage space using a reversible neural network model; generating distribution parameters corresponding to the latent data representation by using an irreversible neural network model; generating compressed data values ​​by performing entropy encoding on the latent data representation based on the distribution parameters; generating a first latent data representation in the latent data representation by decompressing the first compressed data value in the compressed data values ​​based on a read request received from a host device for reading a first data value among the data values, and based on a first distribution parameter in the distribution parameters corresponding to the first data value; generating a first data value corresponding to the first latent data representation by using a reversible neural network model; and providing the first data value to the host device.

[0007] According to one aspect of this disclosure, a method of operating a memory controller includes: generating a latent data representation corresponding to data values ​​to be written to a memory space using a reversible neural network model; generating distribution parameters corresponding to the latent data representation using an irreversible neural network model; compressing the latent data representation based on the distribution parameters; partially decompressing at least one latent data representation corresponding to the at least one data value in the compressed latent data representation based on a command from a host device for providing at least one data value among the data values ​​to be written to the memory space; and providing the at least one data value to the host device.

[0008] Additional aspects may be set forth in part in the description which follows, and may be apparent in part from the description, and / or may be learned by practice of the presented embodiments. Attached Figure Description

[0009] The above and other aspects, features, and advantages of certain embodiments of the present disclosure will become clearer from the following description taken in conjunction with the accompanying drawings, in which:

[0010] Figure 1 This is a diagram illustrating an example of the operation of a storage device in response to write and read requests from a processor, according to an embodiment.

[0011] Figure 2 This is a diagram illustrating an example configuration of a storage device according to an embodiment;

[0012] Figure 3This is a diagram illustrating an example of a data compression process based on a reversible neural network model and an irreversible neural network model according to an embodiment;

[0013] Figure 4 This is a diagram illustrating an example of the data flow during the compression process according to an embodiment;

[0014] Figure 5 This is a diagram illustrating an example of the process for generating distribution parameters according to an embodiment;

[0015] Figure 6 This is a diagram illustrating an example of a data decompression process based on a reversible neural network model and an irreversible neural network model according to an embodiment;

[0016] Figure 7 This is a diagram illustrating an example of the data flow during the decompression process according to an embodiment;

[0017] Figure 8 This is a diagram illustrating an example of the process for extracting distribution parameters according to an embodiment;

[0018] Figure 9 This is a diagram illustrating an example of an operation method of a memory controller according to an embodiment; and

[0019] Figure 10 This is a diagram illustrating an example configuration of an electronic device according to an embodiment. Detailed Implementation

[0020] The following descriptions of structure and / or function are provided as examples only, and various changes and modifications can be made to the embodiments. The examples used herein should not be construed as limiting the scope of this disclosure, and can be understood to include all changes, equivalents, and substitutions within the spirit and technical scope of this disclosure.

[0021] In this document, terms such as first, second, etc., may be used to describe various components. Each of these terms is not used to define the nature, order, or sequence of the corresponding component, but may only be used to distinguish the corresponding component from other components. For example, the first component may be referred to as the second component, and similarly, the second component may be referred to as the first component.

[0022] It should be understood that if the first component is described as being “connected,” “coupled,” and / or “joined” to the second component, then the third component may be “connected,” “coupled,” and / or “joined” between the first component and the second component, although the first component may be directly connected, coupled, and / or joined to the second component.

[0023] The singular forms “a” and “the” may be intended to also include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising / including” and / or “containing / including” as used herein may specify the presence of the said feature, integer, step, operation, element, and / or component, but do not preclude the presence and / or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0024] The phrases “at least one of A and B”, “at least one of A, B or C”, etc., used in this article can be included in any of the items listed together in the corresponding phrase, or all possible combinations thereof.

[0025] Unless otherwise defined, all terms used herein (including technical and scientific terms) may have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant field and should not be interpreted as having an idealized or overly formal meaning unless expressly defined herein.

[0026] Throughout this disclosure, references to “an embodiment,” “an embodiment,” “an exemplary embodiment,” or similar language may indicate that a particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of the solution. Therefore, the phrases “in an embodiment,” “in an embodiment,” “in an exemplary embodiment,” and similar language throughout this disclosure may, but not necessarily, refer to the same embodiment. The embodiments described herein are exemplary embodiments, and therefore, this disclosure is not limited thereto and may be implemented in various other forms.

[0027] It will be understood that the specific order or hierarchy of blocks in the disclosed process / flowchart is an illustration of exemplary methods. Based on design preferences, it should be understood that the specific order or hierarchy of blocks in the process / flowchart can be rearranged. Furthermore, some blocks can be combined or omitted. The appended claims present the elements of the various blocks in an exemplary order and are not intended to limit one to the specific order or hierarchy presented.

[0028] The embodiments described herein can be illustrated and described as blocks that perform one or more of the described functions, as shown in the accompanying drawings. These blocks (which may be referred to herein as cells or modules, or named as devices, logic, circuits, controllers, counters, comparators, generators, converters, etc.) can be physically implemented by analog and / or digital circuits including one or more of logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, etc.

[0029] In this disclosure, the article “a” is intended to include one or more items and may be used interchangeably with “one or more.” When intended for only one item, the term “a” or similar language is used. For example, the term “(a) processor” may refer to a single processor or multiple processors. When a processor is described as performing an operation, and a processor is mentioned as performing additional operations, those multiple operations may be performed by a single processor, or by any one or a combination of multiple processors.

[0030] In the following description, embodiments will be referenced to the accompanying drawings. When describing embodiments with reference to the accompanying drawings, the same reference numerals denote the same elements, and for the sake of brevity, repeated descriptions may be omitted.

[0031] Figure 1 This is a diagram illustrating an example of the operation of a storage device in response to write and read requests from a processor, according to an embodiment. (Refer to...) Figure 1 The processor 110 can send a write request 10 to the storage device 120 to write the data value 11 from the cache 111 to the storage device 120. The storage device 120 can respond to the write request 10 from the processor 110 and store the data value 11 in its storage space. The storage device 120 can compress the data value 11 and store it in its storage space. Data compression enables efficient use of storage space.

[0032] If data value 21 from data value 11 is needed, processor 110 may send a read request 20 to storage device 120 to read data value 21 from storage device 120. Data value 21 may be a portion of data value 11. For example, the size of data value 11 may be 4 kilobytes (KB), and the size of data value 21 may be 64 bytes (B); however, this disclosure is not limited thereto. For example, within the scope of this disclosure, the size of data value 11 and / or the size of data value 21 stored in storage device 120 may vary.

[0033] Storage device 120 can obtain data value 21 by decompressing the compressed data value of data value 21 within the compressed data value of data value 11. To obtain data value 21, partial decompression can be performed on the compressed data value. Storage device 120 can then send data value 21 to processor 110. Processor 110 can store data value 21 in cache 111.

[0034] The data size of write request 10 (e.g., 4KB) may be asymmetrical with the data size of read request 20 (e.g., 64B). Despite this asymmetry, storage device 120 can still compress the relatively large size of data for write request 10 and / or perform partial decompression for the relatively small size of data for read request 20. In an embodiment, if the smaller size of data for read request 20 is obtained from the decompressed data after all the compressed data of the larger size of write request 10 has been decompressed, the processing speed for read request 20 may not be fast enough. The processing speed for read request 20 may be directly related to the processing speed of processor 110. Therefore, slow processing speed for read request 20 can severely impact the processing performance of processor 110.

[0035] In an embodiment, storage device 120 can provide data value 21 to processor 110 by performing partial decompression and responding to read request 20 relatively quickly (e.g., within a timing threshold). Furthermore, as described below, storage device 120 can process read request 20 relatively quickly by potentially reducing the throughput used to process read request 20.

[0036] Figure 2 This is a diagram illustrating an example configuration of a storage device according to an embodiment. (Reference) Figure 2 The storage device 200 may include a memory controller 210 and a memory array 220. The memory controller 210 may compress the data value requested for writing and may store the compressed data value in the storage space of the memory array 220. Alternatively or additionally, the memory controller 210 may decompress the compressed data value corresponding to the data value requested for reading. The memory array 220 may include storage space. The memory array 220 may store the compressed data value in the storage space.

[0037] The memory controller 210 may include a domain converter 211, a parameter generator 212, and a compression processor 213. The domain converter 211, parameter generator 212, and / or compression processor 213 may be implemented as digital circuitry separate from the memory controller 210, and / or may be incorporated into the memory controller 210. In embodiments, the domain converter 211, parameter generator 212, and / or compression processor 213 may be physically implemented using analog and / or digital circuitry including one or more of logic gates, integrated circuits, microprocessors, microcontrollers, memory circuitry, passive electronic components, active electronic components, optical components, etc. For example, a field-programmable gate array (FPGA) may be used to implement custom logic that may include the functionality of at least one of the domain converter 211, parameter generator 212, or compression processor 213. As another example, the combination of processor and memory may be used to execute one or more instructions to perform the functionality of at least one of the domain converter 211, parameter generator 212, or compression processor 213. Alternatively or additionally, at least a portion of the functionality of at least one of the domain converter 211, parameter generator 212, or compression processor 213 may be incorporated into the memory controller 210, and / or may be implemented as instructions executed by the memory controller 210.

[0038] In response to receiving a write request for a data value to be written to the storage space, the domain converter 211 can generate a latent data representation corresponding to the data value using a reversible neural network model. The parameter generator 212 can generate distribution parameters corresponding to the latent data representation using an irreversible neural network model. The compression processor 213 can compress the latent data representation based on the distribution parameters. The memory array 220 can store the compressed latent data representation in its data space.

[0039] Domain converter 211 generates potential data representations by performing domain transformations using a domain transformation model to reduce the entropy of data values. In the example, the domain transformation model can be implemented as hardware (e.g., analog and / or digital circuits, microprocessors, microcontrollers, memory circuits, etc.). For example, the network parameters of the domain transformation model can be stored as parameter values ​​of a network operator. The domain transformation model can perform hardware-based network operations based on the input data values ​​and can generate potential data representations.

[0040] Domain transformation models may include neural network-based coupling layers. These coupling layers can perform the same type of domain transformation and / or different types of domain transformation. One coupling layer may divide the input data into a first sub-data and a second sub-data, generate a first processed sub-data by processing the first sub-data, and generate output data by combining the first processed sub-data with the second sub-data. The input data may be data values ​​that are written to, and the output data may be a latent data representation corresponding to those data values.

[0041] The coupling layer can process the first sub-data using a neural network-based processing layer. The processing layer can process the first sub-data using trained network parameters and quantize the processing result for combination with the second sub-data. Domain transformations can be performed on the data values ​​being written, as they pass through the coupling layer.

[0042] The processing layer can be trained to perform domain transformations by using network parameters, thereby reducing the entropy of the data values. The network parameters of the processing layer may include partitioning parameters and / or selection parameters. The processing layer can extract a first sub-data and a second sub-data from the input data by using the partitioning parameters and / or selection parameters. The network parameters of the processing layer can be determined during the training of the coupling layers.

[0043] In this embodiment, the reversible neural network model may include a normalizing flow. The reversible neural network model generates output data by sequentially transforming input data using layers. These layers can process the output data sequentially based on the reversible properties of the reversible neural network model to generate the input data. These reversible properties can be used for lossless compression.

[0044] Reversible neural network models can be configured to satisfy high levels of compression performance constraints to provide reversibility. Reversible neural network models may require a relatively large number of layers to potentially achieve high compression performance constraints. However, as the number of layers increases, the time required for compression and decompression also increases. According to embodiments, by limiting the number of layers in the reversible neural network model and using smaller-sized input data, compression performance can be improved compared to related reversible neural network models. For example, the irreversible neural network model can estimate the probability distribution of the potential data representation generated using the reversible neural network model. Based on the probability distribution estimated by the irreversible neural network model, the bit allocation of the potential data representation can be optimized compared to related storage devices, and the compression ratio can be improved.

[0045] The parameter generator 212 can generate distribution parameters corresponding to the latent data representation by using an irreversible neural network model. The distribution parameters can be and / or may include a probability distribution of the latent data representation. For example, the probability distribution can be a Gaussian distribution or a Laplace distribution. However, this disclosure is not limited thereto. The distribution parameters may include the mean and standard deviation.

[0046] Parameter generator 212 may include a parameter detection model, an entropy encoder, and an entropy decoder. According to embodiments, the irreversible neural network model may include a variational autoencoder (VAE). For example, the irreversible neural network model may include a hyperprior.

[0047] An irreversible neural network model may include a superencoder and a superdecoder. The superencoder can encode a latent data representation to generate a hyperlatent data representation. The entropy encoder can perform entropy encoding on the hyperlatent data representation using a trained probability distribution. For example, the probability distribution can be obtained through Gaussian modeling based on trained parameters; however, this disclosure is not limited thereto. Arithmetic encoding can be used for entropy encoding; however, this disclosure is not limited thereto. Encoded hyperlatent data representations can be generated based on entropy encoding.

[0048] An entropy decoder can perform entropy decoding on an encoded hyperlatent data representation using a trained probability distribution. A recovered hyperlatent data representation can be generated from the entropy decoding. Arithmetic decoding can be used for entropy decoding; however, this disclosure is not limited thereto. A superdecoder can decode the recovered hyperlatent data representation to generate distribution parameters of the latent data representation. The distribution parameters can be and / or can include feature vectors that express the latent data representation as a probability distribution. To quickly process read requests (e.g., within a timing threshold), the number of layers in the superdecoder can be less than the number of layers in the superencoder. For example, the superencoder can have five (5) layers, and the superdecoder can have one or two (2) layers; however, this disclosure is not limited thereto.

[0049] In this embodiment, the parameter detection model can be implemented as hardware (e.g., analog and / or digital circuits, microprocessors, microcontrollers, memory circuits, etc.). For example, the network parameters of the parameter detection model can be stored as parameter values ​​of network operators. The parameter detection model can perform hardware-based network operations based on inputs of latent data representations and can generate latent data distribution parameters.

[0050] Compression processor 213 can compress the underlying data representation based on distributed parameters. Compression processor 213 may include an entropy encoder and an entropy decoder. The entropy encoder and entropy decoder of compression processor 213 may be different from the entropy encoder and entropy decoder of parameter generator 212. The entropy encoder and entropy decoder of compression processor 213 may be referred to as a first entropy encoder and a first entropy decoder, respectively, and the entropy encoder and entropy decoder of parameter generator 212 may be referred to as a second entropy encoder and a second entropy decoder, respectively. In embodiments, the first entropy encoder and the second entropy encoder, as well as the first entropy decoder and the second entropy decoder, may be implemented as hardware (e.g., analog and / or digital circuits, microprocessors, microcontrollers, memory circuits, etc.).

[0051] The compression processor 213 can perform entropy encoding on the latent data representation based on distribution parameters. The compression processor 213 can perform entropy encoding by using a first entropy encoder. Compressed data values ​​can be generated as the result of the entropy encoding. For example, the compressed data values ​​can be and / or may include encoded data values.

[0052] In an embodiment, partial decompression can be performed in response to a read request from the processor. The processor can send a read request to read a specified data value and can perform partial decompression on the specified data value. For example, in response to a read request for reading a first data value from the data values ​​of a write request, the first data value can be generated from the first compressed data value in the latent data representation using a first distribution parameter in the distribution parameters corresponding to the first data value. The first data value of the write request can be a portion of the data values ​​of the write request. For example, the size of the data value can be 4KB, and the size of the first data value can be 64B; however, this disclosure is not limited thereto. The compressed data value of the latent data representation can be a value encoded by a first entropy encoder based on the latent data representation and the distribution parameters.

[0053] In response to receiving a read request for a first data value from the data values, the compression processor 213 can generate a first latent data representation by decompressing the first compressed data value from the compressed data values ​​of the latent data representation, based on a first distribution parameter corresponding to the first data value in the distribution parameters. As described below, the distribution parameter corresponding to the data value can be generated independently. Data values ​​can be compressed independently using independent distribution parameters. The first distribution parameter corresponding to the first data value in the independent distribution parameters can be selectively used for partial decompression of the first compressed data value.

[0054] Domain converter 211 can generate a first data value corresponding to a first latent data representation by using a domain conversion model. The first decompressed data can be the first latent data representation. The domain conversion model can generate a latent data representation by performing a domain conversion on the data value, and can generate a data value by performing an inverse domain conversion on the latent data representation. The domain conversion model can generate the first data value by performing a partial inverse domain conversion on the first latent data representation within the latent data representation. Domain conversion can refer to a forward domain conversion, and inverse domain conversion can refer to a reverse domain conversion. The domain conversion model can perform forward and / or reverse domain conversions based on reversibility.

[0055] Figure 3 This is a diagram illustrating an example of a data compression process based on a reversible neural network model and an irreversible neural network model according to an embodiment. (Reference) Figure 3 Based on the domain transformation 310 associated with the data value 301, a potential data representation 311 can be generated. The domain transformation 310 can be performed using a reversible neural network model. The number of data values ​​301 can be N, where N is a positive integer greater than zero (0). For example, if a write request for 4KB data values ​​301 is received, each of the data values ​​301 can be 64B, and N can be 64; however, this disclosure is not limited thereto.

[0056] N domain transformations 310 can be performed on N data values ​​301. These N domain transformations 310 can be performed independently. They can also be performed sequentially using a single domain transformation model, and / or in parallel using N sub-models of the domain transformation model. Sub-models can have independent network parameters, and / or can share network parameters. The number of potential data representations 311 generated from the N domain transformations 310 can be N.

[0057] As described above, reversible neural network models can be designed to meet relatively high levels of compression performance constraints to provide reversibility. According to embodiments, the number of layers in the reversible neural network model can be finite; however, N domain transformations 310 can be performed independently on the N data values ​​301. For example, instead of performing one (1) domain transformation on the integrated N data values ​​301, N independent domain transformations 310 can be performed on the N data values ​​301. Therefore, the limited expressive power of the reversible neural network model can be compensated.

[0058] As described above, compression performance can be improved when compared to the associated storage device by limiting the number of layers in the reversible neural network model and / or by using an irreversible neural network model. A merged latent data representation can be generated based on the reshaping operation 320 of the latent data representation 311. Based on the reshaping operation 320, the latent data representation 311 can be merged. The number of latent data representations 311 can be N, and the number of reshaped latent data representations can be one (1), however, this disclosure is not limited thereto. For example, based on the reshaping operation 320 of the latent data representation 311, M merged latent data representations can be generated. M can be a positive integer greater than one (1) and less than N (e.g., 1 < M < N).

[0059] The corresponding distribution parameters of the latent data representation 311 can be generated by performing parameter generation 330 based on the distribution of the merged latent data representation. Parameter generation 330 can be performed using a parameter generator, including an irreversible neural network model. The parameter generator can generate the corresponding distribution parameters of the latent data representation 311 by predicting the corresponding probability distribution of the latent data representation 311 based on the distribution of the merged latent data representation. A parameter generator (e.g., a parameter generation model) can be trained to predict the corresponding probability distribution of the latent data representation 311 based on the distribution of the merged latent data representation. Due to the wider range of input data, the accuracy of distribution prediction can be better when compared with a relevant reversible neural network model. When using the distribution of the merged latent data representation, the accuracy of the distribution parameters 331 can be improved compared with using the corresponding distribution of the latent data representation 311.

[0060] Based on the latent data representation 311 and distribution parameters 331, entropy coding 340 can be performed. According to the entropy coding 340, encoded data value 341 can be generated. Encoded data value 341 can be stored in memory array 350. Memory array 350 may include the referenced above. Figure 2 The memory array 220 described herein and / or may be similar in many respects to the reference above. Figure 2 The memory array 220 is described, and may include additional features not mentioned above. Therefore, for the sake of brevity, the memory array 350 referenced above can be omitted. Figure 2 The description is a repetitive description.

[0061] Entropy coding 340 can be performed using a first entropy encoder. The latent data representation 311 can have a one-to-one correspondence with the distribution parameters 331. The latent data representation 311 and the distribution parameters 331 can have data pairs with a one-to-one correspondence. For example, the distribution parameters 331 may include a first distribution parameter that indicates the probability distribution of the first latent data representation in the latent data representation 311. In this case, the first latent data representation and the first distribution parameter can form a first pair.

[0062] The first entropy encoder can generate encoded data value 341 by performing entropy encoding on each corresponding data value. Data value 301 can have a one-to-one correspondence with encoded data value 341. For example, encoded data value 341 can include a first encoded data value corresponding to a first data value in data value 301.

[0063] Figure 4 This is a diagram illustrating an example of the data flow during the compression process according to an embodiment. (Reference) Figure 4 It can generate latent data representations 420 (e.g., first latent data representation 421, second latent data representation 422, and third latent data representation 423) corresponding to data values ​​410 (e.g., first data value 411, second data value 412, and third data value 413) based on domain transformation, generate distribution parameters 430 (e.g., first distribution parameter 431, second distribution parameter 432, and third distribution parameter 433) corresponding to latent data representations 420 based on parameter generation, and generate encoded data values ​​440 (e.g., first encoded data value 441, second encoded data value 442, and third encoded data value 443) corresponding to latent data representations 420 and distribution parameters 430 based on entropy coding.

[0064] Data value 410, latent data representation 420, distribution parameter 430, and encoded data value 440 can have a one-to-one correspondence with each other. For example, a first latent data representation 421 can be generated based on the first data value 411, a first distribution parameter 431 can be generated based on the first latent data representation 421, and a first encoded data value 441 can be generated based on the first latent data representation 421 and the first distribution parameter 431. As another example, a second latent data representation 422 can be generated based on the second data value 412, a second distribution parameter 432 can be generated based on the second latent data representation 422, and a second encoded data value 442 can be generated based on the second latent data representation 422 and the second distribution parameter 432. As another example, a third latent data representation 423 can be generated based on the third data value 413, a third distribution parameter 433 can be generated based on the third latent data representation 423, and a third encoded data value 443 can be generated based on the third latent data representation 423 and the third distribution parameter 433.

[0065] Figure 5 This is a diagram illustrating an example of the process for generating distribution parameters according to an embodiment. (Reference) Figure 5 Based on the reshaping operation 520 of the latent data representation 511, a merged latent data representation can be generated. Distribution analysis 5301 can be performed on the merged latent data representation. Distribution analysis 5301 can be performed using a super encoder. A super latent data representation can be generated based on the distribution analysis 5301. An encoded super latent data representation can be generated based on the entropy encoding 5302 of the super latent data representation. Entropy encoding 5302 can be performed using a second entropy encoder. The encoded super latent data representation can be stored in a memory array 550.

[0066] Entropy decoding 5304 based on the encoded hyperlatent data representation can generate a recovered hyperlatent data representation. Entropy decoding 5304 can be performed using a second entropy decoder. Synthesis 5305 can be performed based on the recovered hyperlatent data representation. Synthesis 5305 can be performed using a superdecoder. Distribution parameters 531 can be generated from the synthesis 5305.

[0067] Encoded data value 541 can be generated based on entropy encoding 540 using latent data representation 511 and distribution parameters 531. The encoded data value 541 and the encoded hyperlatent data representation stored in memory array 550 can be used to process subsequent read requests. Memory array 550 may include the referenced data above. Figure 2 and Figure 3 The memory arrays 220 and 350 described may be similar in many respects to those referenced above. Figure 2 and Figure 3 The memory arrays 220 and 350 are described, and may include additional features not mentioned above. Therefore, for the sake of brevity, the reference to memory array 550 above can be omitted. Figure 2 and Figure 3 The description is a repetitive description.

[0068] In this embodiment, the distribution parameter 531 may be stored in the memory array 550. For example, the distribution parameter 531 (rather than the encoded hyperlatent data representation) may be stored in the memory array 550. In this case, read requests can be processed based on the encoded data value 541 and the distribution parameter 531 stored in the memory array 550. In this case, the processing speed for read requests can be improved compared to the associated storage device.

[0069] Figure 6 This is a diagram illustrating an example of a data decompression process based on a reversible neural network model and an irreversible neural network model according to an embodiment. (See reference) Figure 6A read request may include a first index value 602 indicating a first data value 601. The first index value 602 may indicate the first data value 601, which may be part of the data value requested in a write request. A first distribution parameter 631 may be selectively used from the distribution parameters based on the first index value 602.

[0070] The memory array 650 can store encoded hyperlatent data representations and / or distribution parameters. The memory array 650 may include the references mentioned above. Figure 2 , Figure 3 and Figure 5 The memory arrays 220, 350, and 550 described herein and / or may be similar in many respects to those referenced above. Figure 2 , Figure 3 and Figure 5 The memory arrays 220, 350, and 550 are described, and may include additional features not mentioned above. Therefore, for the sake of brevity, the reference to memory array 650 above can be omitted. Figure 2 , Figure 3 and Figure 5 The description is a repetitive description.

[0071] Based on the first index value 602, the first coded hyperlatent data representation corresponding to the first index value 602 in the coded hyperlatent data representation and / or the first distribution parameter 631 corresponding to the first index value 602 in the distribution parameters can be extracted from the memory array 650. If the first coded hyperlatent data representation is extracted, the first distribution parameter 631 can be extracted from the first coded hyperlatent data representation based on parameter extraction 630. Parameter extraction 630 can be performed using a superdecoder and a second entropy decoder.

[0072] Entropy decoding 640 can be performed based on the first encoded data value 651 and the first distribution parameter 631. A first latent data representation can be generated based on entropy decoding 640. Entropy decoding 640 can be performed using a first entropy decoder. A first domain transformation corresponding to the first latent data representation can be performed within domain transformation 610. For example, if the domain transformation model includes N sub-models, a first sub-model can be selected from the sub-models based on the first index value 602, and the first domain transformation can be performed using the first sub-model. A first data value 601 can be generated based on the first domain transformation. The first data value 601 can be provided to the processor in response to a read request.

[0073] Figure 7 This is a diagram illustrating an example of the data flow during the decompression process according to an embodiment. (Reference) Figure 7The data values ​​710 (e.g., first data value 711, second data value 712, and third data value 713), latent data representations 720 (e.g., first latent data representation 721, second latent data representation 722, and third latent data representation 723), distribution parameters 730 (e.g., first distribution parameter 731, second distribution parameter 732, and third distribution parameter 733), and encoded data values ​​740 (e.g., first encoded data value 741, second encoded data value 742, and third encoded data value 743) can have a one-to-one correspondence with each other. For example, a one-to-one correspondence can be formed between the first data values ​​711 to 713 of the data values ​​710, the first latent data representations 721 to 723 of the latent data representations 720, the first distribution parameters 731 to 733 of the distribution parameters 730, and the first encoded data values ​​741 to 743 of the encoded data values ​​740.

[0074] The processor can send a read request indicating a portion of data value 710. In this case, the read request may include an index value indicating a portion of data value 710. For example, the read request may include a first index value indicating a first data value 711 within data value 710. In this case, a first encoded data value 741 can be extracted from the memory array. A first distribution parameter 731 associated with the first encoded data value 741 can be extracted.

[0075] The memory array 650 can store coded hyperlatent data representations and / or distribution parameters 730. If the coded hyperlatent data representation is stored in the memory array 650, a first coded hyperlatent data representation can be extracted from the memory array 650. Based on the parameter extraction operation associated with the first coded hyperlatent data representation, a first distribution parameter 731 can be generated. If the distribution parameter 730 is stored in the memory array, the first distribution parameter 731 can be extracted from the memory array. According to the domain inverse transformation based on the first distribution parameter 731 and the first latent data representation 721, a first data value 711 can be generated.

[0076] Figure 8 This is a diagram illustrating an example of the process for extracting distribution parameters according to an embodiment. (Reference) Figure 8 Based on a read request for the first data value, the first encoded data value 841 and the first encoded hyperlatent data representation 851 can be extracted from the memory array 850. The memory array 850 may include the referenced above. Figure 2 , Figure 3 , Figure 5 and Figure 6 The memory arrays described, 220, 350, 550, and 650, and / or may be similar in many respects to those referenced above. Figure 2 , Figure 3 , Figure 5 and Figure 6 The memory arrays 220, 350, 550, and 650 are described, and may include additional features not mentioned above. Therefore, for brevity, the reference to memory array 850 above can be omitted. Figure 2 , Figure 3 , Figure 5 and Figure 6 The description is a repetitive description.

[0077] A first recovered hyperlatent data representation can be generated based on entropy decoding 8304 of the first encoded hyperlatent data representation 851. Entropy decoding 8304 can be performed using a second entropy decoder. A first distribution parameter 831 can be generated based on synthesis 8305 associated with the first recovered hyperlatent data representation. Synthesis 8305 can be performed using a hyperdecoder. A first latent data representation 811 can be generated based on entropy decoding 840 associated with the first encoded data value 841 and the first distribution parameter 831. Entropy decoding 840 can be performed using a first entropy decoder.

[0078] Figure 9 This is a diagram illustrating an example of an operation method of a memory controller according to an embodiment. (Reference) Figure 9 In operation 910, the memory controller can generate a latent data representation corresponding to the data values ​​to be written into the memory space by using a reversible neural network model. In operation 920, the memory controller can generate distribution parameters corresponding to the latent data representation by using an irreversible neural network model. In operation 930, the memory controller can compress the latent data representation based on the distribution parameters.

[0079] In response to a read request for a first data value among the data values, the memory controller can generate a first data value from the first compressed data value among the compressed data values ​​of the potential data representation by using a first distribution parameter among the distribution parameters corresponding to the first data value.

[0080] In response to receiving a read request for reading a first data value from the data values, the memory controller can generate a first latent data representation from the compressed data values ​​of the latent data representation by decompressing the first compressed data value from the compressed data values ​​of the latent data representation, based on a first distribution parameter in the distribution parameters corresponding to the first data value.

[0081] The memory controller can generate a first data value corresponding to the first potential data representation by using a domain transformation model.

[0082] The distribution parameters can be stored in storage space. A read request may include a first index value indicating a first data value. The first distribution parameter can be selectively used from the distribution parameters based on the first index value.

[0083] The memory controller can generate potential data representations by performing domain transformations using a domain transformation model to reduce the entropy of data values.

[0084] The memory controller can perform entropy encoding on the potential data representation based on distributed parameters.

[0085] The memory controller can reshape the latent data representation and generate a merged latent data representation, and can generate the corresponding distribution parameters of the latent data representation based on the distribution of the merged latent data representation.

[0086] Reversible neural network models can include normalized flow. Irreversible neural network models can include VAEs.

[0087] Figure 10 This is a diagram illustrating an example configuration of an electronic device according to an embodiment. (Reference) Figure 10 Electronic device 1000 may include a processor, a first-tier memory device 1040, and a second-tier memory device 1050. Electronic device 1000 may also include additional components (e.g., but not limited to, storage devices (e.g., disks), input / output (I / O) devices, communication interfaces, auxiliary processors (e.g., graphics processing units (GPUs), neural processing units (NPUs), or accelerators)). For example, electronic device 1000 may be and / or may include computing devices (e.g., but not limited to, desktop computers or servers); however, this disclosure is not limited thereto. In embodiments, processor 1010 may be and / or may include a host processor (e.g., a central processing unit (CPU)). For example, if electronic device 1000 includes an auxiliary processor, then processor 1010 may be the main processor. According to embodiments, second-tier memory device 1050 may be and / or may include a memory device that performs partial decompression based on size-asymmetric write and read requests.

[0088] In an embodiment, the first-level storage device 1040 may be faster than the second-level storage device 1050 (e.g., faster read and / or write speeds). In such embodiments, the second-level storage device 1050 may provide a larger capacity than the first-level storage device 1040. For example, the first-level storage device 1040 may be and / or may include system memory (e.g., dynamic random access memory (DRAM)), and the second-level storage device 1050 may be and / or may include additional memory, auxiliary memory, or remote memory (e.g., but not limited to, CXL-based memory). The first-level storage device 1040 may be referred to as fast memory, and the second-level storage device 1050 may be referred to as slow memory.

[0089] The processor 1010 and the first-level storage device 1040 can be connected via a first interface 1020. For example, the first-level storage device 1040 can be DRAM, and the first interface 1020 can be a dual in-line memory module (DIMM). The processor 1010 and the second-level storage device 1050 can be connected via a second interface 1030. For example, the second-level storage device 1050 can be a CXL-based memory, and the second interface 1030 can be a CXL interface. The first interface 1020 can be different from the second interface 1030. For example, the first interface 1020 and the second interface 1030 can use different communication standards and / or different connection methods.

[0090] Processor 1010 can access first-level storage device 1040 and second-level storage device 1050 respectively through first interface 1020 and second interface 1030. Processor 1010 can perform memory access using memory addresses. Memory access can include, but is not limited to, write requests and read requests. Data can be stored in the storage space of the memory address according to a write request. Data can be retrieved from the storage space according to a read request.

[0091] The units described herein can be implemented using hardware components, software components, and / or combinations thereof. The processing device can be implemented using one or more general-purpose or special-purpose computers, such as processors, controllers, and arithmetic logic units (ALUs), digital signal processors (DSPs), microcomputers, field-programmable gate arrays (FPGAs), programmable logic units (PLUs), microprocessors, or any other device capable of responding to and executing instructions in a defined manner. The processing device can run an operating system (OS) and one or more software applications running on the OS. The processing unit can also access, store, manipulate, process, and generate data in response to the execution of software. For simplicity, the description of processing units is used in the singular; however, those skilled in the art will understand that a processing unit can include multiple processing elements and various types of processing elements. For example, a processing unit can include multiple processors, or a single processor and a single controller. Furthermore, different processing configurations (e.g., parallel processors) are possible.

[0092] Software may include computer programs, code, instructions, or some combination thereof, to independently or jointly instruct and / or configure processing units to operate on demand. Software and data may be stored in any type of machine, component, physical or virtual device, or computer storage medium or device capable of providing instructions or data to the processing unit, or capable of providing instructions or data to be interpreted by the processing unit. Software may also be distributed across network-coupled computer systems, enabling the software to be stored and executed in a distributed manner. Software and data may be stored on one or more non-transitory computer-readable recording media.

[0093] The methods described in the examples above can be recorded in a non-transitory computer-readable medium that includes program instructions for implementing the various operations described in the examples. The medium may also include data files, data structures, etc., alone or in combination with the program instructions. The program instructions recorded on the medium may be program instructions specifically designed and constructed for the purposes of the examples, or these instructions may be program instructions well-known and available to those skilled in the art of computer software. Examples of non-transitory computer-readable media include magnetic media (e.g., hard disks, floppy disks, and magnetic tapes), optical media (e.g., CD-ROMs and DVDs), magneto-optical media (e.g., optical discs), and hardware devices (e.g., read-only memory (ROM), RAM, flash memory), etc., that can be specifically configured to store and execute program instructions. Examples of program instructions include both machine code (e.g., generated by a compiler) and / or files containing high-level code that can be executed by a computer using an interpreter.

[0094] The device described above can act as one or more software modules to perform the operations described above, or vice versa.

[0095] As described above, although examples have been described with reference to limited accompanying drawings, those skilled in the art can apply various technical modifications and variations based on them. For example, suitable results may be achieved if the described techniques are performed in a different order and / or if the components in the described system, architecture, device, or circuit are combined in a different manner and / or replaced or supplemented by other components or their equivalents.

[0096] Therefore, other embodiments are also within the scope of the appended claims.

Claims

1. A memory controller, comprising: One or more processors, including processing circuitry; as well as Memory, stored instructions When the instruction is executed individually or jointly by the one or more processors, the memory controller performs the following operations: By using a reversible neural network model, a potential data representation corresponding to the data value to be written to the storage space is generated; Distribution parameters corresponding to the latent data representation are generated by using an irreversible neural network model; Compress the latent data representation based on the distribution parameters; Based on a command from the host device for providing at least one data value among the data values ​​written to the storage space, partially decompress at least one latent data representation corresponding to the at least one data value in the compressed latent data representation; and Provide the host device with the at least one data value.

2. The memory controller according to claim 1, wherein, When the instructions are executed individually or jointly by the one or more processors, the memory controller is also caused to perform the following operations: Based on a command including a read request for reading a first data value among the data values, the first data value is generated from the first compressed data value among the compressed data values ​​of the potential data representation by using a first distribution parameter among the distribution parameters corresponding to the first data value.

3. The memory controller according to claim 1, wherein, When the instructions are executed individually or jointly by the one or more processors, the memory controller is also caused to perform the following operations: Based on a command including a read request for reading a first data value among the data values, and based on a first distribution parameter among the distribution parameters corresponding to the first data value, a first latent data representation is generated by decompressing the first compressed data value among the compressed data values ​​of the latent data representation.

4. The memory controller according to claim 3, wherein, When the instructions are executed individually or jointly by the one or more processors, the memory controller is also caused to perform the following operations: The first data value corresponding to the first potential data representation is generated by using the reversible neural network model.

5. The memory controller according to claim 3, wherein, The distribution parameters are stored in the storage space. The read request includes a first index value indicating the first data value, and When the instruction is executed individually or jointly by the one or more processors, the memory controller also performs the following operation: selecting the first distribution parameter among the distribution parameters based on the first index value.

6. The memory controller according to claim 1, wherein, When the instructions are executed individually or jointly by the one or more processors, the memory controller is also caused to perform the following operations: The potential data representation is generated by using the reversible neural network model and reducing the entropy of the data values ​​through domain transformation.

7. The memory controller according to claim 1, wherein, When the instructions are executed individually or jointly by the one or more processors, the memory controller is also caused to perform the following operations: Entropy encoding is performed on the potential data representation based on the distribution parameters.

8. The memory controller according to claim 1, wherein, When the instructions are executed individually or jointly by the one or more processors, the memory controller is also caused to perform the following operations: A merged potential data representation is generated by reshaping the potential data representation; as well as The corresponding distribution parameters of the potential data representation are generated based on the distribution of the merged potential data representation.

9. The memory controller according to claim 1, wherein, The reversible neural network model includes a normalized flow.

10. The memory controller according to claim 1, wherein, The irreversible neural network model includes a variational autoencoder (VAE).

11. A storage device, comprising: The memory array is configured to store compressed data values; One or more processors, including processing circuitry; as well as Memory, stored instructions Wherein, when the instructions are executed individually or jointly by the one or more processors, the storage device performs the following operations: By using a reversible neural network model, a potential data representation corresponding to the data value is generated by performing domain transformation to reduce the entropy of the data value to be written to the storage space. Distribution parameters corresponding to the latent data representation are generated by using an irreversible neural network model; The compressed data value is generated by performing entropy encoding on the potential data representation based on the distribution parameters. Based on a read request received from the host device for reading a first data value from the data values, and based on the first distribution parameter in the distribution parameters corresponding to the first data value, the first latent data representation in the latent data representation is generated by decompressing the first compressed data value from the compressed data values. The first data value corresponding to the first latent data representation is generated by using the reversible neural network model; and The first data value is provided to the host device.

12. The storage device according to claim 11, wherein, The read request includes a first index value indicating the first data value, and When the instruction is executed individually or jointly by the one or more processors, the storage device also selects the first distribution parameter among the distribution parameters based on the first index value.

13. The storage device according to claim 11, wherein, When the instructions are executed individually or jointly by the one or more processors, the storage device also performs the following operations: A merged data representation is generated by reshaping the underlying data representation; and The corresponding distribution parameters of the potential data representation are generated based on the distribution of the merged data representation.

14. A method of operating a memory controller, the method comprising: By using a reversible neural network model, a potential data representation corresponding to the data value to be written to the storage space is generated; Distribution parameters corresponding to the latent data representation are generated by using an irreversible neural network model; Compress the latent data representation based on the distribution parameters; Based on a command from the host device for providing at least one data value among the data values ​​written to the storage space, partially decompress at least one latent data representation corresponding to the at least one data value in the compressed latent data representation; and Provide the host device with the at least one data value.

15. The operating method according to claim 14, further comprising: Based on a command including a read request for reading a first data value among the data values, the first data value is generated from the first compressed data value among the compressed data values ​​of the potential data representation by using a first distribution parameter among the distribution parameters corresponding to the first data value.

16. The operating method according to claim 14, further comprising: Based on a command including a read request for reading a first data value among the data values, and based on a first distribution parameter among the distribution parameters corresponding to the first data value, a first latent data representation is generated by decompressing the first compressed data value among the compressed data values ​​of the latent data representation.

17. The operating method according to claim 16, further comprising: The first data value corresponding to the first potential data representation is generated by using the reversible neural network model.

18. The operating method according to claim 16, wherein, The distribution parameters are stored in the storage space. The read request includes a first index value indicating the first data value, and The operation method further includes: selecting the first distribution parameter among the distribution parameters based on the first index value.

19. The operating method according to claim 14, wherein, Generating the potential data representation includes: The potential data representation is generated by using the reversible neural network model and reducing the entropy of the data values ​​through domain transformation.

20. The operating method according to claim 14, wherein, The latent data representation is reshaped to generate a merged latent data representation, and The generation of the distribution parameters includes: generating the corresponding distribution parameters of the potential data representation based on the distribution of the merged potential data representation.

Citation Information

Patent Citations

  • Power supply apparatus of underground transmission line

    KR1020240145947A

  • Compounds for the treatment of neurodegenerative diseases

    KR1020250034166A