Operation method of host controlling computing device

By reducing and aligning weight data based on configuration information, the method enhances the efficiency of data transfer and processing in artificial intelligence systems, addressing inefficiencies in existing methods.

JP2025122634APending Publication Date: 2025-08-21SAMSUNG ELECTRONICS CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025013619
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-17
Filing Date
2025-01-30
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing methods for operating a computing device in artificial intelligence systems are inefficient due to the large amount of calculations required by neural networks, leading to suboptimal performance and resource utilization.

Method used

A method involving a host that receives configuration information from the computing device, performs weight reduction and alignment operations to generate lightweight data files, and loads them into multiple memory devices, allowing parallel access by an accelerator through direct memory access operations.

Benefits of technology

This approach enables efficient data transfer and processing by distributing lightweight data across multiple memory channels, improving the performance of the computing device and the artificial intelligence system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025122634000001_ABST
    Figure 2025122634000001_ABST
Patent Text Reader

Abstract

To provide an operation method of a host that controls a computing device with improved performance, and an operation method of an artificial intelligence system that includes the computing device and the host.SOLUTION: According to the present invention, an operation method of a host which controls a computing device performing an artificial intelligence computation includes: receiving configuration information from the computation device; generating lightening weight data by performing lightening on weight data based on the configuration information; generating a plurality of files by performing an aligning operation on the lightening weight data (LWD) based on the configuration information; and loading the plurality of files into a memory device of the computing device. The configuration information includes channel information about a plurality of channels between the memory device and an accelerator, which are included in the computing device, and the number of the plurality of files is equal to the number of the plurality of channels.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to artificial intelligence, and to a method of operating a host that controls a computing device, and a method of operating an artificial intelligence system that includes a computing device and a host. [Background technology]

[0002] Artificial intelligence (AI) is a field of computer science that attempts to simulate various human capabilities such as learning, reasoning, and perception. Recently, AI has been widely used in diverse fields such as natural language understanding, natural language translation, robotics, artificial vision, problem solving, learning, knowledge acquisition, and cognitive science by enabling computer systems to perform inferences that are impractical for humans to perform, or to make decisions without additional human input.

[0003] Artificial intelligence is implemented based on various algorithms. For example, a neural network is composed of a complex network in which nodes and synapses are repeatedly connected. As data moves from the current node to the next node, various signal processing operations may occur depending on the corresponding synapse, and such signal processing operations are called layers. In other words, a neural network may include various layers that are intricately connected to each other. Because the various layers included in a neural network require a large amount of calculations and data, various methods for optimizing this are being researched. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] U.S. Patent No. 10,402,120 [Patent Document 2] US Patent Application Publication No. 2022 / 0374348 [Patent Document 3] US Patent Application Publication No. 2022 / 0309011 [Patent Document 4] Chinese Patent Publication No. 113469350 [Patent Document 5] International Publication No. 2020 / 038551 Summary of the Invention [Problem to be solved by the invention]

[0005] The present invention has been made in view of the above-mentioned prior art, and an object of the present invention is to provide an operating method of a host that controls a computing device with improved performance, and an operating method of an artificial intelligence system including a computing device and a host. [Means for solving the problem]

[0006] According to an embodiment of the present invention, a method for operating a host configured to control a computing device that performs an artificial intelligence computation includes receiving configuration information from the computing device, performing a lightening operation on weight data based on the configuration information to generate light weight data, performing an aligning operation on the light weight data based on the configuration information to generate a plurality of files, and loading the plurality of files into a memory device of the computing device, wherein the configuration information includes channel information for a plurality of channels between the memory device and an accelerator included in the computing device, and the number of the plurality of files is the same as the number of the plurality of channels.

[0007] According to one embodiment of the present invention, a method for operating a host configured to control a computing device that performs artificial intelligence computing includes receiving configuration information from the computing device, performing a lightening operation on first weight data and second weight data based on the configuration information to generate first light weight data including first weight fragments and second light weight data including second weight fragments, performing an aligning operation on the first weight fragments and the second weight fragments to generate a plurality of files, and loading the plurality of files into the plurality of memories of the computing device, respectively. The configuration information includes channel information for a plurality of channels between the memory device and an accelerator included in the computing device, the number of the plurality of files being the same as the number of the plurality of channels, and the first weight fragments are distributed among the plurality of files and the second weight fragments are distributed among the plurality of files.

[0008] According to one embodiment of the present invention, a method for operating an artificial intelligence system including a computing device and a host includes receiving configuration information of the computing device by the host, performing weight reduction and aligning operations on weight data based on the configuration information by the host to generate a plurality of files, loading the plurality of files into a plurality of memories of the computing device by the host, performing DMA operations in parallel to each of the plurality of memories to read the plurality of files by the computing device, and performing an artificial intelligence operation based on the plurality of files by the computing device. The configuration information includes channel information for a plurality of channels between the plurality of memories included in the computing device and an accelerator, and the number of the plurality of files is the same as the number of the plurality of channels. [Effects of the Invention]

[0009] According to the present invention, a host configured to control a computing device can generate multiple files based on configuration information of the computing device so that an accelerator of the computing device can efficiently access an artificial intelligence model. The multiple files are loaded into multiple memories of the computing device, respectively. This allows the accelerator of the computing device to read the multiple files by performing direct memory access (DMA) operations in parallel from the multiple memories. Therefore, a method for operating a host that controls a computing device with improved performance and a method for operating an artificial intelligence system including a computing device and a host are provided. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a block diagram illustrating an artificial intelligence system according to an embodiment of the present invention. [Figure 2] 2 is a flowchart showing the operation of the host of FIG. 1; [Figure 3] FIG. 2 is a block diagram showing the arithmetic unit of FIG. 1. [Figure 4] 2 is a diagram for explaining the operation of a weight reduction module of the host in FIG. 1; [Figure 5A] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 5B] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 6A] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 6B] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 7A] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 7B] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 8] 2 is a flowchart showing the operation of the arithmetic device of FIG. 1. [Figure 9] 2 is a diagram for explaining the operation of a weight reduction module of the host in FIG. 1; [Figure 10A] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 10B] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 11A] FIG. 16 is a diagram for explaining the method of reading multiple files described with reference to FIGS. 10A and 10B. [Figure 11B] FIG. 16 is a diagram for explaining the method of reading multiple files described with reference to FIGS. 10A and 10B. [Figure 11C] FIG. 16 is a diagram for explaining the method of reading multiple files described with reference to FIGS. 10A and 10B. [Figure 12] 2 is a flowchart showing the operation of the host of FIG. 1; [Figure 13A] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 13B] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 14A] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 14B] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 15] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 16] 2 is a diagram for explaining the operation of an aligning module of the host in FIG. 1. FIG. [Figure 17] 1 is a block diagram illustrating an artificial intelligence system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0011] In the following, embodiments of the present invention will be described clearly and in detail so as to enable those skilled in the art to easily practice the present invention.

[0012] Terms such as “unit,” “module,” and the like, used in the detailed description or drawings, or functional blocks illustrated in the drawings, may be embodied in the form of hardware, software, or a combination thereof configured to perform specific functions. As an example, a “lightening module” may be hardware, software, or a combination thereof configured to support the data lightening operations or functions described herein. In at least some embodiments, the processing circuitry may more specifically include or be activated by, but is not limited to, a central processing unit (CPU), an arithmetic logic unit (ALU), a digital signal processor (DSP), a microcomputer, a field programmable gate array (FPGA), a system-on-chip (SoC), a programmable logic unit, a microprocessor, an application-specific integrated circuit (ASIC), or the like, and may comprise active elements such as transistors, resistors, capacitors, etc., manual electrical elements, or electronic circuitry including one or more of the above elements configured to perform or implement the functions of the functional blocks.

[0013] FIG. 1 is a block diagram illustrating an artificial intelligence system according to an embodiment of the present invention. Referring to FIG. 1, the artificial intelligence system 10 may include a host 100, an artificial intelligence computing device 200, and a storage device 300. In one embodiment, the artificial intelligence system 10 is configured to perform various artificial intelligence calculations, inferences, or learning. The artificial intelligence system 10 may include a graphics processing unit (GPU), a neural processing unit (NPU), or separate dedicated hardware. Alternatively, the artificial intelligence system 10 may be included in an application processor (AP) or a mobile device.

[0014] The host 100 may control the general operation of the artificial intelligence system 10. In one embodiment, the host 100 may include a processor configured to perform various operations in the artificial intelligence system 10 or to run various operating systems and various applications. The processor may control various components of the artificial intelligence system 10 according to the requirements of the various operating systems or various applications. The processor may include a central processing unit (CPU) or an application processor (AP) including one or more cores. For example, the artificial intelligence system 10 may be configured to perform voice recognition, text-to-speech (TTS), image recognition, image classification, or image processing using a neural network. For example, the artificial intelligence system 10 may be implemented in one of various types of electronic devices, such as a tablet device, a smart TV, an augmented reality (AR) device, an Internet of Things (IoT) device, an autonomous vehicle, a robot, a medical device, a drone, an advanced driver assistance system (ADAS), an image display device, a data processing server, or a measurement device.

[0015] The AI ​​computing device 200 (hereinafter referred to as the "computing device" for convenience of explanation) is configured to perform various AI computations under the control of the host 100. For example, the computing device 200 is configured to perform AI computations, inference, or learning based on an AI model or weight provided by the host 100.

[0016] In one embodiment, the artificial intelligence model or weights are generated by machine learning, which may include various learning methods such as supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, etc., although the scope of the present invention is not limited thereto.

[0017] In one embodiment, the artificial intelligence model is generated or trained by one or a combination of at least two or more of various neural networks, such as a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, etc. Alternatively, the artificial intelligence model may be a large language model (LLM) configured to perform various natural language processing (NLP) tasks, or may be trained or generated using a Transformer model. The artificial intelligence model may include multiple neural network layers, each of which may be configured to perform an artificial intelligence operation based on a trained model or weights.

[0018] The storage device 300 is configured to store various information or data used by the artificial intelligence system 10. For example, the storage device 300 may be configured to store weight data WD that causes the calculation device 200 to perform artificial intelligence calculation, inference, or learning. In one embodiment, the weight data WD may be information generated by learning by the calculation device 200. In one embodiment, the storage device 300 may be a mass storage medium such as a solid state drive (SSD). However, the scope of the present invention is not limited thereto, and the storage device 300 may be a mass storage medium such as a universal flash storage (UFS) card, an embedded UFS, a hard disk drive (HDD), etc., or an external storage server connected via a separate communication line.

[0019] In one embodiment, the weight data WD may be based on a large-scale language model. In this case, the size of the weight data WD may be very large. Due to limited resources of the computing device 200 (e.g., the capacity of the memory device 210), the entire weight data WD may not be loaded into the computing device 200. This allows the host 100 to reduce the size of the weight data WD before loading it into the computing device 200.

[0020] For example, the host 100 may include a weight reduction module 110 and an aligning module 120. The weight reduction module 110 may perform weight reduction on the weight data WD stored in the storage device 300. In one embodiment, weight reduction refers to an operation or technique for reducing the size or capacity of the weight data WD, such as pruning, quantization, compression, or knowledge distillation. Pruning refers to a technique for removing values ​​below a reference value from the weight data WD to reduce the capacity of the weight data WD. Quantization refers to a technique for converting the number of bits represented by each piece of weight data WD into a number of bits with relatively low precision to reduce the capacity of the weight data WD. Compression refers to a technique for compressing the weight data WD to reduce the capacity of the weight data WD. Knowledge distillation refers to a technique that uses weight data WD as a teacher model to train a lightweight student model, thereby reducing the size of the weight data WD.

[0021] The above-described lightening methods are merely examples, and the scope of the present invention is not limited thereto. The lightening module 110 is configured to generate lightened weight data LWD by lightening the weight data WD based on various lightening techniques according to the type of AI calculation or the type of AI model performed in the calculation device 200.

[0022] The aligning module 120 may align the lightweight weight data LWD based on the configuration information CONFIG of the computing device 200 to generate a lightweight weight data file FILE_LWD (hereinafter, for convenience of explanation, referred to as a “lightweight file”). For example, the computing device 200 may include a memory device 210 and an accelerator 220. The accelerator 220 may communicate with the memory device 210 through multiple channels. The aligning module 120 may align the lightweight weight data LWD based on the configuration information CONFIG so that the lightweight weight file FILE_LWD is efficiently transferred from the memory device 210 to the accelerator 220. In one embodiment, the configuration information CONFIG may include information regarding the number of channels and a quantization type between the memory device 210 and the accelerator 220 included in the computing device 200. The configuration information CONFIG may include quantization information indicating the number of weight fragments included in each of the multiple lightweight weight data LWD. The light weight data LWD may be loaded into the memory device 210. The aligning operation of the aligning module 120 will be described in further detail with reference to the following figures.

[0023] In one embodiment, the accelerator 220 may include an internal buffer 221. The accelerator 220 can store the lightweight weight file FILE_LWD received from the memory device 210 in the internal buffer 221. The accelerator 220 is configured to generate lightweight weight data LWD based on the lightweight weight file FILE_LWD stored in the internal buffer 221, and to perform an artificial intelligence calculation based on the lightweight weight data LWD.

[0024] As described above, the host 100 according to an embodiment of the present invention may generate a lightweight weight file FILE_LWD by aligning the lightweight weight data LWD based on the configuration information CONFIG of the computing device 200. The lightweight weight file FILE_LWD may be loaded into the memory device 210 of the computing device 200. In this case, the accelerator 220 of the computing device 200 may receive the lightweight weight file FILE_LWD from the memory device 210 through multiple channels, thereby improving the performance of the computing device 200.

[0025] 2 is a flowchart illustrating the operation of the host of FIG. 1. Referring to FIG. 1 and FIG. 2, in step S110, the host 100 may collect configuration information CONFIG for the computing device 200. For example, the host 100 may receive the configuration information CONFIG from the computing device 200 or the accelerator 220 of the computing device 200. In one embodiment, the configuration information CONFIG may be received from an initialization operation of the computing device 200. Alternatively, the configuration information CONFIG may be obtained from the computing device 200 in response to a request from the host 100. In one embodiment, the configuration information CONFIG may include information regarding the number and quantization type of channels between the memory device 210 and the accelerator 220 included in the computing device 200.

[0026] In step S120, the host 100 can read the weight data WD from the storage device 300.

[0027] In operation S130, the host 100 may perform lightening on the weight data WD to generate light weight data LWD. For example, the lightening module 110 of the host 100 may perform various lightening operations on the weight data WD. The lightening of the weight data WD is performed with reference to the configuration information CONFIG received from the calculation device 200. That is, the lightening module 110 may perform lightening on the weight data WD to generate light weight data LWD so as to satisfy a format required by the calculation device 200. In one embodiment, the light weight data LWD may include multiple weight fragments. The number or size of the multiple weight fragments is determined according to a quantization technique or compression technique supported by the calculation device 200.

[0028] In step S140, the host 100 may align the lightweight weight data LWD based on the configuration information CONFIG to generate a lightweight weight file FILE_LWD. For example, the host 100 may align the lightweight weight data LWD so that the lightweight weight file FILE_LWD is efficiently transferred from the memory device 210 to the accelerator 220. The operation of step S140 will be described in more detail with reference to the following drawings.

[0029] In step S150, the host 100 may load the lightweight weight file FILE_LWD into the memory device 210.

[0030] Figure 3 is a block diagram illustrating the computing device of Figure 1. Referring to Figures 1 and 3, the computing device 200 may include a memory device 210 and an accelerator 220.

[0031] The memory device 210 may include multiple memories MEM1-MEM4. Each of the multiple memories MEM1-MEM4 may refer to a physically or logically separated memory space. For example, each of the multiple memories MEM1-MEM4 may be implemented on a separate die, separate chip, separate device, or separate module. The multiple memories MEM1-MEM4 may communicate with the accelerator 220 through multiple channels CH1-CH4, respectively. In one embodiment, the multiple channels CH1-CH4 may refer to communication paths that allow communication independently of each other.

[0032] The accelerator 220 can communicate with a plurality of memories MEM1 to MEM4 through a plurality of channels CH1 to CH4, respectively. The accelerator 220 can include a plurality of memory controllers MCT1 to MCT4, an internal buffer 221, and a processing element (PE) 222.

[0033] The plurality of memory controllers MCT1-MCT4 can communicate with the plurality of memories MEM1-MEM4 through the plurality of channels CH1-CH4, respectively. In one embodiment, the plurality of memory controllers MCT1-MCT4 can read data or files stored in the plurality of memories MEM1-MEM4 through a direct memory access (DMA) operation. For example, the first memory controller MCT1 can perform a first DMA operation on the first memory MEM1 through the first channel CH1. The second memory controller MCT2 can perform a second DMA operation on the second memory MEM2 through the second channel CH2. The third memory controller MCT3 can perform a third DMA operation on the third memory MEM3 through the third channel CH3. The fourth memory controller MCT4 can perform a fourth DMA operation on the fourth memory MEM4 through the fourth channel CH4.

[0034] In one embodiment, each of the memory controllers MCT1-MCT4 may perform DMA operations independently or in parallel with one another. In one embodiment, the DMA operations for each of the memory controllers MCT1-MCT4 may be initiated or configured by a separate DMA controller.

[0035] The internal buffer 221 is configured to temporarily store data or files received via the multiple memory controllers MCT1 to MCT4.

[0036] The processing unit (processing element) 222 can perform artificial intelligence calculations based on the data temporarily stored in the internal buffer 221. For example, the processing unit 222 can perform a multiply and accumulate (MAC) calculation on the weight data and activation data stored in the internal buffer 221. In one embodiment, the calculation results of the processing unit 222 may be stored in the internal buffer 221 or a plurality of memories MEM1 to MEM4.

[0037] As described above, the memory device 210 and the accelerator 220 of the computing device 200 can communicate with each other through multiple channels CH1 to CH4 to improve transfer speed. However, a conventional host uploads weight data to the memory device 210 regardless of the structures of the memory device 210 and the accelerator 220 included in the computing device 200, which results in inefficient data transfer between the memory device 210 and the accelerator 220. For example, a conventional host can upload lightweight weight data to the memory device 210 in a single file format. In this case, even though the memory device 210 and the accelerator 220 communicate with each other through multiple channels CH1 to CH4, the accelerator 220 accesses the lightweight weight data from the memory device 210 through a single channel.

[0038] Meanwhile, according to an embodiment of the present invention, the aligning module 120 of the host 100 can align the light weight data LWD to multiple files based on the configuration information CONFIG of the arithmetic device 200. The multiple files may be loaded into multiple memories MEM1 to MEM4 of the memory device 210, respectively. In this case, the accelerator 220 can simultaneously access the multiple files from the multiple memories MEM1 to MEM4, respectively, independently or in parallel. This improves the speed at which the accelerator 220 accesses the light weight data from the memory device 210.

[0039] FIG. 4 is a diagram illustrating the operation of the weight reduction module of the host of FIG. 1. Referring to FIGS. 1 and 4, the weight data WD may include first to fourth weight data WD1, WD2, WD3, and WD4. Each of the first to fourth weight data WD1, WD2, WD3, and WD4 may correspond to one weight value. Alternatively, each of the first to fourth weight data WD1, WD2, WD3, and WD4 may include a predetermined number of weight values. For example, each of the first to fourth weight data WD1, WD2, WD3, and WD4 may include 16 weight values. However, the scope of the present invention is not limited thereto.

[0040] In one embodiment, each of the first to fourth weighting data WD1, WD2, WD3, and WD4 may include multiple weighting values ​​included in the corresponding artificial intelligence layer. For example, the artificial intelligence model used in the computing device 200 may include multiple layers. The multiple layers may use different weighting values. That is, the weighting values ​​for each of the multiple layers may be different, and each of the first to fourth weighting data WD1, WD2, WD3, and WD4 may include multiple weighting values ​​included in the corresponding artificial intelligence layer.

[0041] The above description of the weight data is provided for the purpose of easily explaining the embodiments of the present invention, and is not intended to limit the scope of the present invention. The number of weight data WD included in the artificial intelligence model may be expanded in various ways, and the weight value included in one weight data may be varied in various ways.

[0042] The lightening module 110 can perform lightening on the weight data WD based on the configuration information CONFIG to generate light weight data LWD. For example, the configuration information CONFIG can include information regarding the type of quantization applied to the arithmetic device 200. Based on the information regarding the type of quantization included in the configuration information CONFIG, the lightening module 110 can perform lightening on the first weight data WD1 to generate first light weight data LWD1, lightening on the second weight data WD2 to generate second light weight data LWD2, lightening on the third weight data WD3 to generate third light weight data LWD3, and lightening on the fourth weight data WD4 to generate fourth light weight data LWD4.

[0043] The first light weight data LWD1 may include weight fragments 1a, 1b, 1c, and 1d W1a, W1b, W1c, and W1d. The second light weight data LWD2 may include weight fragments 2a, 2b, 2c, and 2d W2a, W2b, W2c, and W2d. The third light weight data LWD3 may include weight fragments 3a, 3b, 3c, and 3d W3a, W3b, W3c, and W3d. The fourth light weight data LWD4 may include weight fragments 4a, 4b, 4c, and 4d W4a, W4b, W4c, and W4d.

[0044] For example, the information about quantization included in the configuration information CONFIG may indicate 4-bit non-uniform quantization. In this case, the lightening module 110 may perform 4-bit non-uniform quantization on the first to fourth weighting data WD1 to WD4. The results of the 4-bit non-uniform quantization on each of the first to fourth weighting data WD1 to WD4 are divided into four unit data. The lightening module 110 may perform compression on each unit data to generate multiple weight fragments W1a to W4d. In this case, each unit data may be compressed at a different compression rate. That is, each of the multiple weight fragments W1a to W4d may have a different size or length.

[0045] In one embodiment, the accelerator 220 may calculate first weight data WD1 or a value approximating the first weight data WD1 based on the first light weight data LWD1, calculate second weight data WD2 or a value approximating the second weight data WD2 based on the second light weight data LWD2, calculate third weight data WD3 or a value approximating the third weight data WD3 based on the third light weight data LWD3, and calculate fourth weight data WD4 or a value approximating the fourth weight data WD4 based on the fourth light weight data LWD4. The accelerator 220 may perform an artificial intelligence calculation based on the calculated light weight data LWD1 to LWD4. In this case, the appropriate value may be, for example, a value resulting from a transformation of the first light weight data LWD1 that provides a correct result in the artificial intelligence calculation.

[0046] 5A and 5B are diagrams for explaining the operation of the aligning module of the host in FIG. 1. For convenience of explanation, it is assumed that the light-weight data LWD is the light-weight data LWD described with reference to FIG. 4 (i.e., weight data to which 4-bit non-uniform quantization is applied). However, the scope of the present invention is not limited thereto.

[0047] 1, 3, 4, 5A, and 5B, the aligning module 120 can align the first to fourth light weight data LWD1 to LWD4 based on the configuration information CONFIG and generate multiple files FLa to FLb. For example, the aligning module 120 can generate the a-th file FLa based on the multiple weight fragments W1a to W1d of the first light weight data LWD1, the b-th file FLb based on the multiple weight fragments W2a to W2d of the second light weight data LWD2, the c-th file FLc based on the multiple weight fragments W3a to W3d of the third light weight data LWD3, and the d-th file FLd based on the multiple weight fragments W4a to W4d of the fourth light weight data LWD4. That is, the a-th file FLa contains weight fragments W1a to W1d, the b-th file FLb contains weight fragments W2a to W2d, the c-th file FLc contains weight fragments W3a to W3d, and the d-th file FLd contains weight fragments W4a to W4d.

[0048] The files FLa to FLd may refer to data units that are loaded into the memories MEM1 to MEM4, respectively. For example, the a-th file FLa is loaded into the first memory MEM1, the b-th file FLb is loaded into the second memory MEM2, the c-th file FLc is loaded into the third memory MEM3, and the d-th file FLd is loaded into the fourth memory MEM4.

[0049] The accelerator 220 can access a plurality of files FLa to FLd from a plurality of memories MEM1 to MEM4, obtain light weight data LWD1 to LWD4 based on the plurality of files FLa to FLd, and perform artificial intelligence calculations.

[0050] For example, as shown in FIG. 5B, the first memory controller MCT1 may perform a first DMA operation DMA1 on the first memory MEM1 to read the a-th file FLa. The second memory controller MCT2 may perform a second DMA operation DMA2 on the second memory MEM2 to read the b-th file FLb. The third memory controller MCT3 may perform a third DMA operation DMA3 on the third memory MEM3 to read the c-th file FLc. The fourth memory controller MCT4 may perform a fourth DMA operation DMA4 on the fourth memory MEM4 to read the d-th file FLd.

[0051] In this case, accelerator 220 can sequentially receive weight fragments 1a-1d W1a-W1d from first memory MEM1 through first DMA operation DMA1, and can generate first light weight data LWD1 at second time point t2 when weight fragments 1a-1d W1a-W1d have all been received. That is, after second time point t2, accelerator 220 can start calculations using first light weight data LWD1.

[0052] The accelerator 220 can sequentially receive the 2a-2d weight fragments W2a-W2d from the second memory MEM2 through the second DMA operation DMA2, and can generate the second light weight data LWD2 at the first time point t1 when all of the 2a-2d weight fragments W2a-W2d have been received. That is, after the first time point t1, the accelerator 220 can start calculations using the second light weight data LWD2.

[0053] Similarly, accelerator 220 can sequentially receive weight fragments W3a-W3d from third memory MEM3 through third DMA operation DMA3, and can sequentially receive weight fragments W4a-W4d from fourth memory MEM4 through fourth DMA operation DMA4. Accelerator 220 can generate third light weight data LWD3 and fourth light weight data LWD4 at third time point t3 and fourth time point t4, respectively, and can subsequently perform corresponding operations.

[0054] As described above, as the light weight data LWD1, LWD2, LWD3, and LWD4 are distributed and stored in multiple memories MEM1, MEM2, MEM3, and MEM4, the accelerator 220 can receive the light weight data LWD1, LWD2, LWD3, and LWD4 through individual DMA operations to the multiple memories MEM1, MEM2, MEM3, and MEM4.

[0055] 6A and 6B are diagrams for explaining the operation of the aligning module of the host in FIG. 1. For convenience of explanation, detailed descriptions of the aforementioned components will be omitted. Referring to FIG. 1, FIG. 2, FIG. 4, FIG. 6A, and FIG. 6B, the aligning module 120 can align the light weight data LWD to generate a plurality of files FLe, FLf, FLg, and FLh.

[0056] In the embodiment of FIG. 5A, compressed data of the same light weight data is included in the same file. For example, weight fragments W1a, W1b, W1c, and W1d of the first light weight data LWD1 are included in the same a-th file FLa. On the other hand, in the embodiment of FIG. 6A, weight fragments of the same light weight data may be included in different files. For example, weight fragments W1a, W1b, W1c, and W1d of the first light weight data LWD1 may be included in the e-th through h-th files FLe through FLh, respectively. Weight fragments W2a, W2b, W2c, and W2d of the second light weight data LWD2 may be included in the e-th through h-th files FLe through FLh, respectively. Weight fragments W3a, W3b, W3c, and W3d of the third light weight data LWD3 may be included in the e-th through h-th files FLe through FLh, respectively. Weight fragments W4a, W4b, W4c, and W4d of the fourth light weight data LWD4 may be included in the e-th to h-th files FLe to FLh, respectively.

[0057] In this case, the e-th file FLe may contain 1a, 2a, 3a, and 4a weight fragments W1a, W2a, W3a, and W4a, the f-th file FLf may contain 1b, 2b, 3b, and 4b weight fragments W1b, W2b, W3b, and W4b, the g-th file FLg may contain 1c, 2c, 3c, and 4c weight fragments W1c, W2c, W3c, and W4c, and the h-th file FLh may contain 1d, 2d, 3d, and 4d weight fragments W1d, W2d, W3d, and W4d. The e-th to h-th files FLe to FLh are stored in the first to fourth memories MEM1 to MEM4, respectively.

[0058] In one embodiment, as compressed data of the same light weight data is stored in different files, the point at which artificial intelligence calculations can be started can be accelerated in the accelerator 220. For example, as shown in FIG. 6B, the accelerator 220 can read the e-th to h-th files FLe to FLh in parallel from the first to fourth memories MEM1 to MEM4.

[0059] For example, the first memory controller MCT1 can sequentially read weight fragments W1a, W2a, W3a, and W4a of the e-th file FLe from the first memory MEM1 through a first DMA operation DMA1. The second memory controller MCT2 can sequentially read weight fragments W1b, W2b, W3b, and W4b of the h-th file FLh from the second memory MEM2 through a second DMA operation DMA2. The third memory controller MCT3 can sequentially read weight fragments W1c, W2c, W3c, and W4c of the f-th file FLf from the third memory MEM3 through a third DMA operation DMA3. The fourth memory controller MCT4 can sequentially read weight fragments W1d, W2d, W3d, and W4d of the g-th file FLg from the fourth memory MEM4 through a fourth DMA operation DMA4.

[0060] In this case, at a first time point t1 when all of the weight fragments W1a, W1b, W1c, and W1d corresponding to the first light weight data LWD1 have been received, the accelerator 220 can obtain the first light weight data LWD1 based on the weight fragments W1a, W1b, W1c, and W1d and start the artificial intelligence calculation for the first light weight data LWD1. Then, at a second time point t2 when all of the weight fragments W2a, W2b, W2c, and W2d corresponding to the second light weight data LWD2 have been received, the accelerator 220 can obtain the second light weight data LWD2 based on the weight fragments W2a, W2b, W2c, and W2d and start the artificial intelligence calculation for the second light weight data LWD2. Thereafter, at a third time point t3 when all of the weight fragments W3a, W3b, W3c, and W3d corresponding to the third light weight data LWD3 have been received, the accelerator 220 can obtain the third light weight data LWD3 based on the weight fragments W3a, W3b, W3c, and W3d, and can start the artificial intelligence calculation for the third light weight data LWD3. Thereafter, at a fourth time point t4 when all of the weight fragments W4a, W4b, W4c, and W4d corresponding to the fourth light weight data LWD4 have been received, the accelerator 220 can obtain the fourth light weight data LWD4 based on the weight fragments W4a, W4b, W4c, and W4d, and can start the artificial intelligence calculation for the fourth light weight data LWD4.

[0061] As described above, the aligning module 120 may perform an aligning operation such that weight fragments of the light weight data LWD are distributed across multiple files based on the configuration information CONFIG of the calculation device 200. In this case, the accelerator 220 of the calculation device 200 may access multiple files in parallel from multiple memories MEM1 to MEM4 and obtain the light weight data based on the multiple files accessed in parallel.

[0062] 7A and 7B are diagrams for explaining the operation of the aligning module of the host in FIG. 1. For convenience of explanation, detailed descriptions of the aforementioned components will be omitted. Referring to FIG. 1, FIG. 2, FIG. 4, FIG. 7A, and FIG. 7B, the aligning module 120 can align the light weight data LWD to generate multiple files FL1, FL2, FL3, and FL4.

[0063] In the embodiment of Figure 7A, weight fragments of the same light weight data may be included in different files. For example, weight fragments W1a, W1b, W1c, and W1d of the first light weight data LWD1 may be included in the first through fourth files FL1 through FL4, respectively. Weight fragments W2a, W2b, W2c, and W2d of the second light weight data LWD2 may be included in the first through fourth files FL1 through FL4, respectively. Weight fragments W3a, W3b, W3c, and W3d of the third light weight data LWD3 may be included in the first through fourth files FL1 through FL4, respectively. Weight fragments W4a, W4b, W4c, and W4d of the fourth light weight data LWD4 may be included in the first through fourth files FL1 through FL4, respectively.

[0064] In this case, the first file FL1 may contain weight fragments W1a, W2a, W3a, W4a, the second file FL2 may contain weight fragments W1b, W2b, W3b, W4b, the third file FL3 may contain weight fragments W1c, W2c, W3c, W4c, and the fourth file FL4 may contain weight fragments W1d, W2d, W3d, W4d.

[0065] In the embodiment of FIG. 7A, the aligning module 120 can add padding data so that the weight fragments contained in the first to fourth files FL1 to FL4 have equal lengths.

[0066] 6A and 6B, the weight fragments included in each of the files FLe-FLh have different lengths, so that individual DMA configurations are required for the first through fourth memory controllers MCT1-MCT4. In one embodiment, the DMA configurations may be performed by separate DMA controllers included in the computing device 220.

[0067] For example, in order to perform DMA operations in the first to fourth memory controllers MCT1 to MCT4, settings for the DMA start address, offset, etc. are required. In the embodiment of FIGS. 6A and 6B, when the first memory controller MCT1 accesses the 2a weight fragment W2a from the first memory MEM1, a first start address and a first offset for the 2a weight fragment W2a are set in the first memory controller MCT1. At the same time, when the second memory controller MCT2 accesses the 2b weight fragment W2b from the second memory MEM2, a second start address and a second offset for the 2b weight fragment W2b are set in the second memory controller MCT2. As shown in FIGS. 6A and 6B, because the weight fragments included in the e-th file FLe and the f-th file FLf have different sizes, the first and second start addresses and the first and second offsets should be different from each other. In addition, the time when the first memory controller MCT1 accesses the 2a weight fragment W2a and the time when the second memory controller MCT2 accesses the 2b weight fragment W2b should be different. In this case, the DMA settings for the first memory controller MCT1 and the second memory controller MCT2 should be set individually, which may increase the complexity of the computing device 200.

[0068] 7A and 7B, the aligning module 120 can add padding data to the weight fragments W1a-W4d of the first to fourth light weight data LWD1-LWD4 so that the weight fragments W1a-W4d of the first to fourth light weight data LWD1-LWD4 are stored at the same addresses in the memories MEM1-MEM4, respectively.

[0069] As a more detailed example, as shown in FIG. 7A, a first file FL1 includes 1a and 2a weight fragments W1a, W2a, a second file FL2 includes 1b and 2b weight fragments W1b, W2b, a third file FL3 includes 1c and 2c weight fragments W1c, W2c, and a fourth file FL4 includes 1d and 2d weight fragments W1d, W2d.

[0070] At this time, the aligning module 120 can add padding data to the remaining weight fragments W1b, W1c, and W1d based on the 1a weight fragment W1a, which has the largest size or longest length among the 1a, 1b, 1c, and 1d weight fragments W1a, W1b, W1c, and W1d.

[0071] Similarly, the first file FL1 may further include 3a and 4a weight fragments W3a, W4a, the second file FL2 may further include 3b and 4b weight fragments W3b, W4b, the third file FL3 may further include 3c and 4c weight fragments W3c, W4c, and the fourth file FL4 may further include 3d and 4d weight fragments W3d, W4d. The aligning module 120 may add padding data to the remaining weight fragments W2b, W2c based on the 2a or 2d weight fragment W2a or W2d having the largest size or longest length among the 2a, 2b, 2c, and 2d weight fragments W2a, W2b, W2c, W2d; may add padding data to the remaining weight fragments W3b, W3c, W3d based on the 3a weight fragment W3a having the largest size or longest length among the 3a, 3b, 3c, and 3d weight fragments W3a, W3b, W3c, W3d; and may add padding data to the remaining weight fragments W4a, W4b, W4d based on the 4c ​​weight fragment W4c having the largest size or longest length among the 4a, 4b, 4c, and 4d weight fragments W4a, W4b, W4c, W4d.

[0072] In this case, the first to fourth files FL1 to FL4 may be aligned as shown in FIG. 7A, and the first to fourth files FL1 to FL4 may be loaded into the first to fourth memories MEM1 to MEM4, respectively.

[0073] As shown in FIG. 7A, when the light weight data LWD is aligned and the first to fourth files FL1 to FL4 are generated, the DMA settings for the first to fourth memory controllers MCT1 to MCT4 of the accelerator 220 are simplified.

[0074] For example, weight fragments W1a, W1b, W1c, and W1d corresponding to the first light weight data LWD1 are read from the first to fourth memories MEM1 to MEM4 by the first to fourth memory controllers MCT1 to MCT4, respectively. At this time, as padding data is added to some of the weight fragments W1a, W1b, W1c, and W1d, they will have the same offsets. Therefore, the first to fourth memory controllers MCT1 to MCT4 are set to the same first start address and the same first offset, and the weight fragments W1a, W1b, W1c, and W1d are read from the first to fourth memories MEM1 to MEM4 during the period from the zeroth time point t0 to the first time point t1.

[0075] Weight fragments W2a, W2b, W2c, and W2d corresponding to the second light weight data LWD2 are read from the first through fourth memories MEM1 through MEM4 by the first through fourth memory controllers MCT1 through MCT4, respectively. At this time, as padding data is added to some of the weight fragments W2a, W2b, W2c, and W2d, they will have the same offsets. Furthermore, because the offsets corresponding to the previous weight fragments W1a, W1b, W1c, and W1d are equal to each other, the starting addresses of the weight fragments W2a, W2b, W2c, and W2d will be the same. Therefore, the first through fourth memory controllers MCT1 through MCT4 are set to the same second starting addresses and the same second offsets, and the weight fragments W2a, W2b, W2c, and W2d are read from the fourth memories MEM1 through MEM4 during the period from the first time point t1 to the second time point t2.

[0076] Weight fragments W3a, W3b, W3c, and W3d corresponding to the third light weight data LWD3 are read from the first to fourth memories MEM1 to MEM4 by the first to fourth memory controllers MCT1 to MCT4, respectively. At this time, as padding data is added to some of the weight fragments W3a, W3b, W3c, and W3d, they will have the same offsets. Furthermore, because the offsets corresponding to the previous weight fragments W2a, W2b, W2c, and W2d are equal to each other, the starting addresses of the weight fragments W3a, W3b, W3c, and W3d will be the same. Therefore, the first to fourth memory controllers MCT1 to MCT4 are set to the same third starting address and the same third offset, and the weight fragments W3a, W3b, W3c, and W3d are read from the first to fourth memories MEM1 to MEM4 during the period from the second time point t2 to the third time point t3.

[0077] Similarly, weight fragments W4a, W4b, W4c, and W4d corresponding to the fourth light weight data LWD4 are read from the first through fourth memories MEM1 through MEM4 by the first through fourth memory controllers MCT1 through MCT4, respectively. At this time, as padding data is added to some of the weight fragments W4a, W4b, W4c, and W4d, they will have the same offsets. Furthermore, because the offsets corresponding to the previous weight fragments W3a, W3b, W3c, and W3d are equal to each other, the starting addresses of the weight fragments W4a, W4b, W4c, and W4d will be the same. Therefore, the first through fourth memory controllers MCT1 through MCT4 are set to the same fourth starting address and the same fourth offset, and weight fragments W4a, W4b, W4c, and W4d are read from the first through fourth memories MEM1 through MEM4 during the period from the third time point t3 to the fourth time point t4.

[0078] As described above, according to an embodiment of the present invention, the aligning module 120 of the host 100 can align the lightweight weight data LWD based on the configuration information CONFIG of the arithmetic device 200 and generate the lightweight weight file FILE_LWD. In this case, the aligning module 120 can add padding data to the weight fragments so that the lengths of data accessed from each of the multiple memories MEM1 to MEM4 in the same DMA interval are equal to each other. In this case, the first to fourth memory controllers MCT1 to MCT4 of the accelerator 220 can successfully read the weight fragments W1a to W4d from the first to fourth memories MEM1 to MEM4 using the same start address and the same offset.

[0079] In one embodiment, the accelerator 220 may further include a separate device or structure for removing padding data read from the first to fourth memories MEM1 to MEM4.

[0080] For ease of explanation, the term "DMA interval" will be used below. A DMA interval refers to an interval of DMA operation performed by a memory controller through a single setting of a starting address and offset. For example, in the embodiment of FIG. 7B, DMA interval A may refer to an interval from time 0 t0 to time 1 t1. That is, in DMA interval A, the first memory controller MCT1 may read the 1a weight fragment W1a from the first memory MEM1, the second memory controller MCT2 may read the 1b weight fragment W1b from the second memory MEM2, the third memory controller MCT3 may read the 1c weight fragment W1c from the third memory MEM3, and the fourth memory controller MCT4 may read the 1d weight fragment W1d from the fourth memory MEM4. DMA interval B may refer to an interval from time 1 t1 to time t2. That is, in the Bth DMA period, the first memory controller MCT1 can read the 2a weight fragment W2a from the first memory MEM1, the second memory controller MCT2 can read the 2b weight fragment W2b from the second memory MEM2, the third memory controller MCT3 can read the 2c weight fragment W2c from the third memory MEM3, and the fourth memory controller MCT4 can read the 2d weight fragment W2d from the fourth memory MEM4.

[0081] As described above, the aligning module 120 can add padding data to the weight fragments so that the sizes of the data read from the first to fourth memories MEM1 to MEM4 in the same DMA interval are equal, thereby improving the performance of the computing device 200.

[0082] 8 is a flowchart showing the operation of the computing device of FIG. 1. Referring to FIGS. 1 and 8, in step S210, the computing device 200 may perform a DMA operation to read weight fragments. For example, the host 100 may generate a plurality of files based on the method described with reference to FIGS. 1 to 7B and load the generated files into a plurality of memories MEM1 to MEM4 of the memory device 210. The accelerator 220 of the computing device 200 may perform a DMA operation on each of the plurality of memories MEM1 to MEM4 to read weight fragments from the plurality of memories MEM1 to MEM4.

[0083] In step S220, the calculation device 200 may combine the weight fragments to generate weight data, light weight data, or approximate weight data. For example, as described with reference to FIGS. 1 to 7B, the first to fourth memory controllers MCT1 to MCT4 of the accelerator 220 may perform DMA operations on the first to fourth memories MEM1 to MEM4, respectively, to read weight fragments (e.g., W1a to W4d). The weight fragments (e.g., W1a to W4d) may be temporarily stored in an internal buffer 221 of the accelerator 220. The accelerator 220 may generate light weight data LWD1 to LWD4 based on the weight fragments (e.g., W1a to W4d) temporarily stored in the internal buffer 221.

[0084] As a more detailed example, accelerator 220 may combine 1a, 1b, 1c, and 1d weight fragments W1a, W1b, W1c, and W1d to generate first light-weight weight data LWD1; may combine 2a, 2b, 2c, and 2d weight fragments W2a, W2b, W2c, and W2d to generate second light-weight weight data LWD2; may combine 3a, 3b, 3c, and 3d weight fragments W3a, W3b, W3c, and W3d to generate third light-weight weight data LWD3; and may combine 4a, 4b, 4c, and 4d weight fragments W4a, W4b, W4c, and W4d to generate fourth light-weight weight data LWD4.

[0085] In step S230, the computing device 200 may perform an artificial intelligence operation using the weight data. For example, the computing device 200 may perform a MAC operation using the weight data. In one embodiment, the operation result may be stored back in the memory device 210.

[0086] 9 is a diagram for explaining the operation of the weight reduction module of the host in FIG. 1. Referring to FIG. 1 and FIG. 9, the weight data WD may include first to fourth weight data WD1, WD2, WD3, and WD4. Since the weight data WD has been described with reference to FIG. 4, that is, the weight data WD in FIG. 9 may be substantially the same as or similar to the weight data WD in FIG. 4, detailed description thereof will be omitted.

[0087] The lightening module 110 may perform lightening on the weight data WD based on the configuration information CONFIG to generate light weight data LWD-1. For example, the configuration information CONFIG may include information regarding the type of quantization applied to the arithmetic device 200. Based on the information regarding the type of quantization included in the configuration information CONFIG, the lightening module 110 may perform lightening on the first weight data WD1 to generate first light weight data LWD1-1, lightening on the second weight data WD2 to generate second light weight data LWD2-1, lightening on the third weight data WD3 to generate third light weight data LWD3-1, and lightening on the fourth weight data WD4 to generate fourth light weight data LWD4-1. For example, the information regarding quantization included in the configuration information CONFIG may indicate 3-bit non-uniform quantization. In this case, the lightening module 110 may perform 3-bit non-uniform quantization on the first to fourth weighting data WD1 to WD4. The results of the 3-bit non-uniform quantization on each of the first to fourth weighting data WD1 to WD4 are divided into three data chunks. The lightening module 110 may perform compression on the data chunks to generate multiple weight fragments W1a to W4c. In this case, each data chunk may be compressed at a different compression rate. That is, each of the multiple weight fragments W1a to W4d may have a different size or length.

[0088] In one embodiment, the first light weight data LWD1-1 includes weight fragments 1a, 1b and 1c W1a, W1b, W1c, the second light weight data LWD2-1 includes weight fragments 2a, 2b and 2c W2a, W2b, W2c, the third light weight data LWD3-1 includes weight fragments 3a, 3b and 3c W3a, W3b, W3c, and the fourth light weight data LWD4-1 includes weight fragments 4a, 4b and 4c W4a, W4b, W4c.

[0089] 10A and 10B are diagrams for explaining the operation of the aligning module of the host in FIG. 1. For convenience of explanation, the light weight data LWD-1 in FIG. 9, 10A, and 10B is provided as weight data to which 3-bit non-uniform quantization is applied. However, this is merely an example, and the scope of the present invention is not limited thereto.

[0090] Referring to Figures 1, 3, 9, 10A and 10B, the aligning module 110 can align the first to fourth light weight data LWD1-1 to LWD4-1 based on the configuration information CONFIG to generate multiple files FL1-1 to FL4-1.

[0091] For example, the configuration information CONFIG may include information about four channels CH1 to CH4 between the memory device 210 and the accelerator 220 of the arithmetic device 200. The aligning module 110 may generate four files FL1-1 to FL4-1 based on the configuration information CONFIG (i.e., information about the four channels CH1 to CH4). At this time, the aligning module 120 may distribute the weight fragments W1a to W4c of the first to fourth light weight data LWD1-1 to LWD4-1 to the four files FL1-1 to FL4-1. The four files FL1-1 to FL4-1 may be loaded into the first to fourth memories MEM1 to MEM4, respectively.

[0092] In one embodiment, each of the first to fourth light weight data LWD1-1 to LWD4-1 may include three weight fragments W1a to W1c, W2a to W2c, W3a to W3c, and W4a to W4c. That is, there may be 12 weight fragments in total, and in this case, each of the first to fourth files FL1-1 to FL4-1 may include three weight fragments. For example, a first file FL1-1 may include 1a, 2b and 3c weight fragments W1a, W2b, W3c, a second file FL2-1 may include 1b, 2c and 4a weight fragments W1b, W2c, W4a, a third file FL3-1 may include 1c, 3a and 4b weight fragments W1c, W3a, W4b, and a fourth file FL4-1 may include 2a, 3b and 4c weight fragments W2a, W3b, W4c.

[0093] In one embodiment, padding data may be added to the weight fragments included in each of the first to fourth files FL1-1 to FL4-1 based on the method described with reference to Figures 7A and 7B. For example, among the weight fragments included in each of the first to fourth files FL1-1 to FL4-1, the weight fragments 1a, 1b, 1c, and 2a W1a, W1b, W1c, and W2a accessed during the Ath DMA period, based on the weight fragment 1a W1a having the longest length or largest size, padding data may be added to the remaining weight fragments W1b, W1c, and W2a. Therefore, the sizes of the weight fragments 1a, 1b, 1c, and 2a W1a, W1b, W1c, and W2a accessed during the Ath DMA period may be equal.

[0094] The first through fourth memory controllers MCT1 through MCT4 of the accelerator 220 can perform DMA operations on the first through fourth memories MEM1 through MEM4, respectively, to access weight fragments W1a through W4c. For example, during a DMA interval A from time 0 to time 1, the accelerator 220 can read weight fragments W1a, W1b, W1c, and W2a from the first through fourth memories MEM1 through MEM4. During a DMA interval B from time 1 to time t2, the accelerator 220 can read weight fragments W1a, W1b, W1c, and W2a from the first through fourth memories MEM1 through MEM4. During the Cth DMA interval from the second time point t2 to the third time point t3, the accelerator 220 can read the 3c, 4a, 4b and 4c weight fragments W3c, W4a, W4b, W4c from the first to fourth memories MEM1 to MEM4.

[0095] Figures 11A to 11C are diagrams for explaining the method of reading out the multiple files described with reference to Figures 10A and 10B. For convenience of explanation, it is assumed that the first to fourth files FL1-1 to FL4-1 described with reference to Figures 10A and 10B are loaded into the multiple memories MEM1 to MEM4, respectively. That is, the first memory MEM1 may include a first file FL1-1 including weight fragments W1a, W2b, and W3c, the second memory MEM2 may include a second file FL2-1 including weight fragments W1b, W2c, and W4a including weight fragments W1b, W2c, and W4a, the third memory MEM3 may include a third file FL3-1 including weight fragments W1c, W3a, and W4b including weight fragments W1c, W3a, and W4b, and the third memory MEM3 may include a fourth file FL4-1 including weight fragments W2a, W3b, and W4c including weight fragments W2a, W3b, and W4c.

[0096] 10A, 10B, and 11A, during a DMA interval A from time 0 to time 1, the first memory controller MCT1 may read the 1a weight fragment W1a from the first memory MEM1, the second memory controller MCT2 may read the 1b weight fragment W1b from the second memory MEM2, the third memory controller MCT3 may read the 1c weight fragment W1c from the third memory MEM3, and the fourth memory controller MCT4 may read the 2a weight fragment W2a from the fourth memory MEM4. The accelerator 220 may generate first light weight data LWD1-1 based on the 1a, 1b, and 1c weight fragments W1a, W1b, and W1c. The 2a weight fragment W2a may be stored in an internal buffer 221 of the accelerator 220.

[0097] 10A, 10B, and 11B, during a DMA interval B from a first time point t1 to a second time point t2, the first memory controller MCT1 may read the 2b weight fragment W2b, the second memory controller MCT2 may read the 2c weight fragment W2c, the third memory controller MCT3 may read the 3a weight fragment W3a, and the fourth memory controller MCT4 may read the 3b weight fragment W3b. The accelerator 220 may generate second light weight data LWD2-1 based on the 2a weight fragment W2a read by the fourth memory controller MCT4 during the DMA interval A and the 2b and 2c weight fragments W2b and W2c read by the first and second memory controllers MCT1 and MCT2, respectively, during the DMA interval B. The 3a and 3b weight fragments W3a and W3b may be stored in an internal buffer 221 of the accelerator 220.

[0098] 10A, 10B, and 11C, during a DMA interval C from the second time point t2 to the third time point t3, the first memory controller MCT1 may read the 3c weight fragment W3c, the second memory controller MCT2 may read the 4a weight fragment W4a, the third memory controller MCT3 may read the 3a weight fragment W4b, and the fourth memory controller MCT4 may read the 4c ​​weight fragment W4c. The accelerator 220 may generate third light weight data LWD3-1 based on the 3a and 3b weight fragments W3a and W3b read by the third and fourth memory controllers MCT3 and MCT4 during the DMA interval B, and the 3c weight fragment W3c read by the first memory controller MCT1 during the DMA interval C. The accelerator 220 can generate fourth light weight data LWD4-1 based on the 4a-th to 4c-th weight fragments W4a-W4c read by the second to fourth memory controllers MCT2-MCT4, respectively, in the C-th DMA period.

[0099] As described above, according to an embodiment of the present invention, the host 100 can generate a lightweight weight file FILE_LWD by performing lightening and alignment on the weight data WD based on the configuration information CONFIG of the calculation device 200. The lightweight weight file FILE_LWD is loaded or stored in each of the memories MEM1 to MEM4 of the calculation device 200. In this case, the accelerator 220 of the calculation device 200 can read the lightweight weight file FILE_LWD by performing DMA operations on the memories MEM1 to MEM4 in parallel or individually, and can obtain lightweight weight data using weight fragments included in the lightweight weight file FILE_LWD. Therefore, the speed at which the accelerator 220 of the calculation device 200 accesses the lightweight weight data is improved, thereby improving the performance of the calculation device 200.

[0100] 1 and 12, in step S310, the host 100 can acquire configuration information CONFIG from the computing device 200. In step S320, the host 100 can read weight data WD from the storage device 300. In step S330, the host 100 can lighten the weight data WD and generate light weight data LWD. The operations of steps S310 to S330 are similar to the operations of steps S110 to S130 described with reference to FIG. 2, and therefore detailed description thereof will be omitted.

[0101] In step S340, the host 100 may align the lightweight weight data LWD based on the configuration information CONFIG and the size of the weight fragments to generate a lightweight weight file FILE_LWD. For example, the aligning module 120 of the host 100 may generate the lightweight weight file FILE_LWD based on the configuration information CONFIG and the size of the weight fragments so that the size of the lightweight weight file FILE_LWD is minimized. Alternatively, the aligning module 120 of the host 100 may generate the lightweight weight file FILE_LWD based on the configuration information CONFIG and the size of the weight fragments so that the lightweight weight file FILE_LWD is efficiently transferred from the memory device 210 to the accelerator 220. The aligning operation of step S340 will be described in further detail with reference to FIGS. 13A to 16.

[0102] In step S350, the host 100 may load the lightweight weight file FILE_LWD into the memory device 210.

[0103] 13A and 13B are diagrams illustrating the operation of the aligning module of the host of FIG. 1. In one embodiment, the operation of step S340 of FIG. 12 will be described with reference to FIG. 13A and 13B. For convenience of explanation, detailed description of the aforementioned components will be omitted. Referring to FIG. 1, 12, 13A, and 13B, the aligning module 120 of the host 100 may perform an aligning operation on the light weight data LWD based on the configuration information CONFIG.

[0104] For example, the light weight data LWD may be the result of 4-bit non-uniform quantization and compression of the weight data WD, i.e., the first light weight data LWD1 may include weight fragments 1a-1d W1a, W1b, W1c, W1d, the second light weight data LWD2 may include weight fragments 2a-2d W2a-W2d, the third light weight data LWD3 may include weight fragments 3a-3d W3a-W3d, and the fourth light weight data LWD4 may include weight fragments 4a-4d W4a-W4d.

[0105] The aligning module 120 can arrange the 1a-1d weight fragments W1a, W1b, W1c, and W1d of the first light weight data LWD1 into the first to fourth files FL1-2 to FL4-2, respectively. Then, the aligning module 120 can arrange the 2a-2d weight fragments W2a-W2d of the second light weight data LWD2 into the first to fourth files FL1-2 to FL4-2, respectively, based on the sizes of the 1a-1d weight fragments W1a, W1b, W1c, and W1d of the first light weight data LWD1 and the sizes of the 2a-2d weight fragments W2a-W2d of the second light weight data LWD2.

[0106] For example, of the 1a-1d weight fragments W1a, W1b, W1c, and W1d, the 1a weight fragment W1a may have the longest length or largest size. Therefore, of the 2a-2d weight fragments W2a-W2d, the 2b weight fragment W2b, which has the shortest length or smallest size, is added to the first file FL1-2 so that it is positioned after the 1a weight fragment W1a.

[0107] Of the 1a-1d weight fragments W1a, W1b, W1c, and W1d, the 1c weight fragment W1c may have the second longest length or second largest size. Therefore, of the 2a-2d weight fragments W2a-W2d, the 2c weight fragment W2c, which has the second shortest length or second smallest size, is added to the third file FL3-2 so as to be positioned after the 1c weight fragment W1c.

[0108] Of the 1a-1d weight fragments W1a, W1b, W1c, and W1d, the 1b weight fragment W1b may have the third longest length or the third largest size. Therefore, of the 2a-2d weight fragments W2a-W2d, the 2d weight fragment W2d, which has the third shortest length or the third smallest size, is added to the second file FL2-2 so that it is positioned after the 1b weight fragment W1b.

[0109] Of the 1a-1d weight fragments W1a, W1b, W1c, and W1d, the 1d weight fragment W1d may have the shortest length or smallest size. Therefore, of the 2a-2d weight fragments W2a-W2d, the 2a weight fragment W2a, which has the longest length or largest size, is added to the second file FL4-2 so that it is positioned after the 1d weight fragment W1d.

[0110] Similarly, the aligning module 120 can arrange the 3a to 3d weight fragments W3a to W3d of the third light weight data LWD3 and the 4a to 4d weight fragments W4a to W4d of the fourth light weight data LWD4 into the first to fourth files FL1-2 to FL4-2, respectively, based on the sizes of the 3a to 3d weight fragments W3a to W3d of the third light weight data LWD3 and the 4a to 4d weight fragments W4a to W4d of the fourth light weight data LWD4.

[0111] Thus, the first file FL1-1 contains 1a, 2b, 3c and 4a weight fragments W1a, W2b, W3c, W4a, the second file FL2-1 contains 1b, 2d, 3b and 4d weight fragments W1b, W2d, W3b, W4d, the third file FL3-1 contains 1c, 2c, 3d and 4c weight fragments W1c, W2c, W3d, W4c, and the first file FL1-1 contains 1a, 2b, 3c and 4a weight fragments W1a, W2b, W3c, W4a.

[0112] The first to fourth files FL1-2 to FL4-2 may be loaded into the first to fourth memories MEM1 to MEM4, respectively. The accelerator 220 can read the first to fourth files FL1-2 to FL4-2 or weight fragments from the first to fourth memories MEM1 to MEM4.

[0113] For example, as shown in FIG. 13B , the first memory controller MCT1 sequentially receives 1a, 2b, 3c, and 4a weight fragments W1a, W2b, W3c, and W4a from the first memory MEM1 through a first DMA operation; the second memory controller MCT2 sequentially receives 1b, 2d, 3b, and 4d weight fragments W1b, W2d, W3b, and W4d from the second memory MEM2 through a second DMA operation; the third memory controller MCT3 sequentially receives 1c, 2c, 3d, and 4c weight fragments W1c, W2c, W3d, and W4c from the third memory MEM3 through a third DMA operation; and the fourth memory controller MCT4 sequentially receives 1a, 2b, 3c, and 4a weight fragments W1a, W2b, W3c, and W4a from the fourth memory MEM4 through a fourth DMA operation.

[0114] At a first time t1, the accelerator 220 can generate first lightweight data LWD1, at a second time t2, the accelerator 220 can generate second lightweight data LWD2, at a third time t3, the accelerator 220 can generate third lightweight data LWD3, and at a fourth time t4, the accelerator 220 can generate fourth lightweight data LWD4.

[0115] In one embodiment, as shown in FIGS. 13A and 13B, multiple files FL1-2 to FL4-2 are generated based on the size of the weight fragments. In this case, the size of each of the multiple files FL1-2 to FL1-4 is optimized or minimized, thereby reducing the DMA time between the memory device 210 and the accelerator 220. For example, if the size of the weight fragments is not taken into consideration, relatively large weight fragments are concentrated in a specific file or a specific memory. In this case, the DMA time for the specific file or the specific memory may increase. On the other hand, according to the embodiment of FIGS. 13A and 13B, relatively large weight fragments (or relatively small weight fragments) are distributed among the multiple files FL1-2 to FL4-2, so the length of each of the multiple files FL1-2 to FL4-2 may be relatively short. In this case, the DMA time for each of the multiple memories MEM1 to MEM4 is reduced.

[0116] 14A and 14B are diagrams illustrating the operation of the aligning module of the host of FIG. 1. Referring to FIGS. 1, 14A, and 14B, the aligning module 120 of the host 100 performs an aligning operation on the light weight data (LWD or LWD1 to LWD4) based on the configuration information CONFIG. This allows multiple files FL1-3 to FL4-3 to be generated. The light weight data (LWD or LWD1 to LWD4) and the weight fragments W1a to W4d included therein have been described with reference to FIG. 4, and therefore a detailed description thereof will be omitted.

[0117] In the above-described embodiment, the aligning module 120 distributes weight fragments corresponding to one weight data to different files. Meanwhile, in the embodiment of Figures 14A and 14B, the aligning module 120 of the host 100 distributes k weight fragments (where k is a natural number) corresponding to one weight data to n files (where n is a natural number smaller than k). For example, the aligning module 120 distributes weight fragments W1a, W1b, W1c, and W1d of the first light weight data LWD1 among the first and second files FL1-3 and FL2-3, distributes weight fragments W2a, W2b, W2c, and W2d of the second light weight data LWD2 among the third and fourth files FL3-3 and FL4-3, distributes weight fragments W3a, W3b, W3c, and W3d of the third light weight data LWD3 among the first and second files FL1-3 and FL2-3, and distributes weight fragments W4a, W4b, W4c, and W4d of the fourth light weight data LWD4 among the third and fourth files FL3-3 and FL4-3.

[0118] In this case, the first file FL1-3 contains 1a, 1d, 3a and 3d weight fragments W1a, W1d, W3a and W3d, the second file FL2-3 contains 1b, 1c, 3b and 3c weight fragments W1b, W1c, W3b and W3c, the third file FL3-3 contains 2a, 2b, 4c and 4d weight fragments W2a, W2b, W4c and W4d, and the fourth file FL4-3 contains 2c, 2d, 4a and 4b weight fragments W2c, W2d, W4a and W4b.

[0119] Except for the fact that the number of weight fragments assigned to the same file among weight fragments corresponding to the same light weight data is different, the method of allocating weight fragments (e.g., a method of distributing based on the size of the weight fragment, a method of adding padding data, etc.) is similar to that described above, and therefore detailed explanation thereof will be omitted.

[0120] In one embodiment, the aligning module 120 of the host 100 can generate the first to fourth files FL1-3 to FL4-3 so that two weight fragments are read from each of the first to fourth memories MEM1 to MEM4 in one DMA interval. At this time, padding data may be added to each of the first to fourth memories MEM1 to MEM4 so that the lengths or sizes of the two weight fragments are equal.

[0121] For example, as shown in FIG. 14A, padding data may be added next to the 1b and 1c weight fragments W1b, W1c of the second file FL2-3 and next to the 2a and 2b weight fragments W2a, W2b of the third file FL3-3 so that the first size of the 1a and 1d weight fragments W1a, W1d of the first file FL1-3, the second size of the 1b and 1c weight fragments W1b, W1c of the second file FL2-3, the third size of the 2a and 2b weight fragments W2a, W2b of the third file FL3-3, and the fourth size of the 2c and 2d weight fragments W2c, W2d of the fourth file FL4_3 are all equal to each other. Padding data may be added next to the 4c ​​and 4d weight fragments W4c and W4d of the third file FL3-3 so that the fifth size of the 3a and 3d weight fragments W3a and W3d of the first file FL1-3, the sixth size of the 3b and 3c weight fragments W3b and W3c of the second file FL2-3, the seventh size of the 4c ​​and 4d weight fragments W4c and W4d of the third file FL3-3, and the eighth size of the 4a and 4b weight fragments W4a and W4b of the fourth file FL4_3 are all equal to one another. In this case, the addition of padding data may be minimized.

[0122] The generated first to fourth files FL1-3 to FL4-3 are loaded into the first to fourth memories MEM1 to MEM4, respectively, and read by the accelerator 220. For example, as shown in FIG. 14B, the first memory controller MCT1 can sequentially receive weight fragments 1a, 1d, 3a, and 3d W1a, W1d, W3a, and W3d from the first memory MEM1 through a first DMA operation. The second memory controller MCT2 can sequentially receive weight fragments 1b, 1c, 3b, and 3c W1b, W1c, W3b, and W3c from the second memory MEM2 through a second DMA operation. The third memory controller MCT3 can sequentially receive weight fragments 2a, 2b, 4c, and 4d W2a, W2b, W4c, and W4d from the third memory MEM3 through a third DMA operation. The fourth memory controller MCT4 can sequentially receive the 2c, 2d, 4a and 4b weight fragments W2c, W2d, W4a and W4b from the fourth memory MEM4 through a fourth DMA operation.

[0123] At a first time t1, the accelerator 220 can generate first light-weight weight data LWD1 using weight fragments 1a, 1b, 1c, and 1d W1a, W1b, W1c, and W1d, and can generate second light-weight weight data LWD2 using weight fragments 2a, 2b, 2c, and 2d W2a, W2b, W2c, and W2d. At a second time t2, the accelerator 220 can generate third light-weight weight data LWD3 using weight fragments 3a, 3b, 3c, and 3d W3a, W3b, W3c, and W3d, and can generate fourth light-weight weight data LWD4 using weight fragments 4a, 4b, 4c, and 4d W4a, W4b, W4c, and W4d.

[0124] Fig. 15 is a diagram for explaining the operation of the aligning module of the host in Fig. 1. Referring to Fig. 1 and Fig. 15, the aligning module 120 of the host 100 performs an aligning operation on the light weight data (LWD, or LWD1 to LWD4) based on the configuration information CONFIG, and can generate multiple files FL1-4 to FL6-4. The light weight data (LWD, or LWD1 to LWD4) and the weight fragments W1a to W4d included therein have been described with reference to Fig. 4, so detailed description thereof will be omitted.

[0125] In the embodiment described above, the memory device 210 and the accelerator 220 of the computing device 200 communicate with each other through four channels CH1 to CH4. Therefore, the aligning module 120 can generate four files FL1 to FL4, FL1-1 to FL4-1, FL1-2 to FL4-2, or FL1-3 to FL4-3 to be stored in memories MEM1 to MEM4 corresponding to the four channels CH1 to CH4, respectively. However, the scope of the present invention is not limited thereto.

[0126] As an example, the memory device 210 and the accelerator 220 of the computing device 200 may communicate with each other through six channels. In this case, the aligning module 120 may generate six files FL1-4 to FL6-4. For example, the aligning module 120 may distribute weight fragments W1a, W1b, W1c, and W1d of the first light-weight weight data LWD1 to the first, second, third, and fourth files FL1-4, FL2-4, FL3-4, and FL4-4. The aligning module 120 may distribute weight fragments W3a, W3b, W3c, and W3d of the third light-weight weight data LWD3 to the first, second, third, and fourth files FL1-4, FL2-4, FL3-4, and FL4-4. The aligning module 120 may distribute the weight fragments W4a, W4b, W4c, and W4d of the fourth light weight data LWD4 among the first, second, third, and fourth files FL1-4, FL2-4, FL3-4, and FL4-4. The aligning module 120 may distribute the weight fragments W2a, W2b, W2c, and W2d of the second light weight data LWD2 among the fifth and sixth files FL5-4 and FL6-4.

[0127] In this case, the first file FL1-4 may include 1a, 3a and 4a weight fragments W1a, W3a, W4a, the second file FL2-4 may include 1b, 3b and 4b weight fragments W1b, W3b, W4b, the third file FL3-4 may include 1c, 3c and 4c weight fragments W1c, W3c, W4c, the fourth file FL4-4 may include 1d, 3d and 4d weight fragments W1d, W3d, W4d, the fifth file FL5-4 may include 2a and 2c weight fragments W2a, W2c, and the sixth file FL6-4 may include 2b and 2d weight fragments W2b, W2d.

[0128] The aligning module 120 can add padding data so that the size of data accessed in one DMA section is equal. The method of adding padding data has been described above, so a detailed description thereof will be omitted.

[0129] The first to sixth files FL1-4 to FL6-4 may be loaded into six memories configured to communicate with the accelerator 220 via six channels, respectively. The accelerator 220 can access the first to sixth files FL1-4 to FL6-4 from the six memories via the six channels independently or in parallel to generate the first to fourth light weight data LWD1 to LWD4.

[0130] 16 is a diagram illustrating the operation of the aligning module of the host of FIG. 1. Referring to FIG. 1 and FIG. 16, the aligning module 120 of the host 100 performs an aligning operation on the light weight data (LWD, or LWD1 to LWD4) based on the configuration information CONFIG, and can generate a plurality of files FL1-5 to FL6-5. The light weight data (LWD, or LWD1 to LWD4) and the weight fragments W1a to W4d included therein have been described with reference to FIG. 4, and therefore detailed description thereof will be omitted.

[0131] In the embodiment of FIG. 16, the aligning module 120 distributes weight fragments corresponding to one lightweight weight data across two files. For example, the aligning module 120 distributes weight fragments W1a, W1b, W1c, and W1d of the first lightweight weight data LWD1 across the first and second files FL1-5 and FL2-5. The aligning module 120 distributes weight fragments W2a, W2b, W2c, and W2d of the second lightweight weight data LWD2 across the third and fourth files FL3-5 and FL4-5. The aligning module 120 distributes weight fragments W3a, W3b, W3c, and W3d of the third lightweight weight data LWD3 across the fifth and sixth files FL5-5 and FL6-5. The aligning module 120 distributes the weight fragments W4a, W4b, W4c, and W4d of the fourth light weight data LWD4 to the first, second, third, and fourth files FL1-4, FL2-4, FL3-4, and FL4-4.

[0132] In this case, the first file FL1-5 contains 1a, 1b and 4a weight fragments W1a, W1b and W4a, the second file FL2-5 contains 1c, 1d and 4b weight fragments W1c, W1d and W4b, the third file FL3-5 contains 2a, 2c and 4c weight fragments W2a, W2c and W4c, the fourth file FL4-5 contains 2b, 2d and 4a weight fragments W2b, W2d and W4d, the fifth file FL5-5 contains 3a and 3c weight fragments W3a and W3c, and the sixth file FL6-5 contains 3b and 3d weight fragments W3b and W3d.

[0133] The first to sixth files FL1-4 to FL6-4 may be loaded into six memories configured to communicate with the accelerator 220 via six channels, respectively. The accelerator 220 can access the first to sixth files FL1-4 to FL6-4 from the six memories via the six channels independently or in parallel, respectively, to generate the first to fourth light weight data LWD1 to LWD4.

[0134] As described above, the host 100 according to the present invention can generate multiple files by performing a lightening operation and an aligning operation on multiple weight data WD based on the configuration information CONFIG of the computing device 200. The host 100 can load multiple files into the memory device 210 of the computing device 200. In this case, the accelerator 220 of the computing device 200 can access the multiple files from the memory device 210 in parallel, thereby reducing the time it takes for the accelerator 220 to obtain weight data.

[0135] In one embodiment, host 100 may provide weight fragment information (e.g., starting addresses, offsets, etc.) regarding weight fragments included in multiple files to computing device 200. Computing device 200 may read the multiple files from memory device 210 based on the weight fragment information and identify weight fragments from the multiple files. For example, accelerator 220 may perform DMA configuration for multiple memory controllers MCT1-MCT4 based on the weight fragment information. Multiple memory controllers MCT1-MCT4 may perform DMA operations for multiple memories MEM1-MEM4 of memory device 210 based on the DMA configuration, thereby identifying weight fragments.

[0136] In one embodiment, the host 100 may provide the weight fragment information to the computing device 200 via a separate communication line. Alternatively, the host 100 may load the weight fragment information into a predetermined area of ​​the memory device 200. The accelerator 220 of the computing device 200 may access the predetermined area of ​​the memory device 210 to acquire the weight fragment information and access multiple files based on the acquired weight fragment information.

[0137] The above-described embodiments are merely examples for easily explaining the present invention, and the scope of the present invention is not limited thereto. For example, the configuration information CONFIG of the computing device 200 may include information regarding n-bit quantization information and k channels. In this case, each of the plurality of lightweight weight data generated by the lightweight module 110 of the host 100 may be divided into n weight fragments, and the sizes of the weight fragments may differ from one another. The aligning module 120 of the host 100 may perform an aligning operation on the weight fragments to generate a plurality of files. The plurality of files are loaded into the memory device 210 of the computing device 200, and the accelerator 220 of the computing device 200 may read the plurality of files from the memory device 210. In this case, the plurality of files may be generated based on the configuration information CONFIG of the aligning module 120 so that the accelerator 220 of the computing device 200 can efficiently read the plurality of files from the memory device 210. This improves the performance of the artificial intelligence system 10 or the computing device 200.

[0138] FIG. 17 is a block diagram illustrating an artificial intelligence system according to an embodiment of the present invention. Referring to FIG. 17, the artificial intelligence system 1000 may include a processor 1100, a memory 1200, a storage device 1300, an artificial intelligence storage device 1400, an artificial intelligence memory 1500, a first accelerator 1610, and a second accelerator 1620. The processor 1100 may control the general operation of the artificial intelligence system 1000. The processor 1100 may run or execute various operating systems or various programs. The memory 1200 is used as a system memory, buffer memory, or cache memory for the artificial intelligence system 1000. The storage device 1300 is used as a mass storage medium for the artificial intelligence system 1000.

[0139] For example, the storage device 1300 is configured to store information or data related to various operating systems and various programs run by the artificial intelligence system 1000. The various information or data included in the storage device 1300 is loaded into the memory 1200, and the processor 1100 can run or execute the various operating systems and various programs based on the information or data loaded into the memory 1200.

[0140] The artificial intelligence storage device 1400 is configured to store various information, artificial intelligence models, or weight data for artificial intelligence calculation, learning, or inference. For example, the artificial intelligence storage device 1400 may be or include the storage device 300 of FIG. 1. In one embodiment, the processor 1100 may perform a weighting and aligning operation on the various information, artificial intelligence models, or weight data for artificial intelligence calculation, learning, or inference stored in the artificial intelligence storage device 1400 to generate multiple files. The multiple files may be stored in the artificial intelligence memory 1500. In one embodiment, the processor 1100 may perform the weighting and aligning operation based on the methods described with reference to FIGS. 1 to 16. For example, the processor 1100 may be or include the host 100 of FIG. 1.

[0141] The first and second accelerators 1610 and 1620 may read multiple files from the artificial intelligence memory 1500 and generate various information, artificial intelligence models, or weight data for artificial intelligence calculation, learning, or inference based on the multiple files. The first and second accelerators 1610 and 1620 may perform artificial intelligence calculation, learning, or inference based on the generated information. For example, the first and second accelerators 1610 and 1620 may be or include the accelerator 220 of FIG. 1.

[0142] In one embodiment, the first and second accelerators 1610, 1620 may share the artificial intelligence memory 1500. For example, the first and second accelerators 1610, 1620 may access the artificial intelligence memory 1500 through the same channel. Alternatively, the first and second accelerators 1610, 1620 may access different regions of the artificial intelligence memory 1500. For example, the first accelerator 1610 may access a first region of the artificial intelligence memory 1500 through a first channel, and the second accelerator 1620 may access a second region of the artificial intelligence memory 1500 through a second channel. The artificial intelligence memory 1500 may be substantially identical to or similar to the memory device 210 of FIG. 1 .

[0143] In one embodiment, the processor 1100 can perform lightening and aligning operations based on the configuration information of each of the first and second accelerators 1610, 1620 so that each of the first and second accelerators 1610, 1620 can efficiently access the artificial intelligence memory 1500.

[0144] 17 illustrates individual components for artificial intelligence calculation, learning, or inference (e.g., artificial intelligence storage device 1400 and artificial intelligence memory 1500), but the scope of the present invention is not limited thereto. Depending on the implementation manner of the artificial intelligence system, artificial intelligence storage device 1400 and artificial intelligence memory 1500 may be omitted, and data, information, or weight data for artificial intelligence calculation, inference, or learning may be stored in storage device 1300. Multiple files generated by the lightening and aligning operations of processor 1100 may be stored in memory 1200. In this case, first and second accelerators 1610 and 1620 may access multiple files from memory 1200.

[0145] The above-described content is a specific embodiment for implementing the present invention. The present invention includes not only the above-described embodiment but also embodiments that can be simply modified or easily changed. The present invention also includes techniques that can be easily implemented by modifying the embodiment. Therefore, the scope of the present invention should not be limited to the above-described embodiment, but should be defined not only by the claims below but also by equivalents to the claims of the present invention.

Claims

1. 1. A method of operating a host configured to control a computing device that performs artificial intelligence computing, comprising: receiving configuration information from the computing device; generating a plurality of light weight data by performing a lightening process on the weight data based on the configuration information; generating a plurality of files by performing an aligning operation on the plurality of light weight data based on the configuration information; loading the plurality of files into a memory device of the computing device; the configuration information includes channel information for a plurality of channels between the memory device of the computing device and an accelerator of the computing device; the number of the plurality of files is equal to the number of available channels of the plurality of channels; How it works.

2. the configuration information includes quantization information; The step of generating the plurality of lightweight weight data includes: generating a plurality of unit data by quantizing the weight data based on the quantization information; performing a compression operation on the plurality of unit data to generate the plurality of light weight data; The method of claim 1 .

3. The step of generating the plurality of lightweight weight data includes: generating the plurality of lightweight weight data such that each of the plurality of lightweight weight data includes a plurality of weight fragments, and the number of the plurality of weight fragments included in each of the plurality of lightweight weight data corresponds to the quantization information; The method of claim 2.

4. The step of generating a plurality of files by performing an aligning operation on the light-weight data based on the configuration information includes: Distributing the weight fragments included in each of the plurality of lightweight weight data to the plurality of files based on the configuration information. The method of claim 3.

5. providing weight fragment information for the plurality of files to the computing device; the weight fragment information includes information regarding a start address and an offset of each of the plurality of weight fragments included in each of the plurality of files; 5. The method of claim 4.

6. The step of generating the plurality of files comprises: generating the plurality of files such that a first file of the plurality of files includes a first portion of a plurality of first weight fragments included in a first light weight data of the plurality of light weight data, and a second file of the plurality of files includes a second portion of the plurality of first weight fragments included in the first light weight data of the plurality of light weight data; The method of claim 3.

7. loading the plurality of files into the memory device of the computing device, loading the first file into a first memory, among a plurality of memories of the memory device, that communicates with the accelerator through a first channel; loading the second file into a second memory of the plurality of memories of the memory device, the second memory communicating with the accelerator through a second channel; 7. The method of claim 6.

8. loading the plurality of files into the memory device of the computing device, loading the plurality of files into the memory device such that a starting address of the first portion included in the first file loaded into the first memory is the same as a starting address of the second portion included in the second file loaded into the second memory, 8. The method of claim 7.

9. The step of generating the plurality of files comprises: generating the plurality of files such that the size of the first portion is greater than the size of the second portion, and the second file further includes first padding data added to the second portion.

7. The method of claim 6.

10. The step of generating the plurality of files comprises: generating the plurality of files such that each of the plurality of files has the same size as the others; The method of claim 1 .

11. A computer program causing a computer to perform the operating method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Deep convolutional neural network acceleration method and system suitable for NPU

    CN113469350A

  • Memory controller arbiter with streak and read / write transaction management

    US10402120B2

  • On-chip interconnect for memory channel controllers

    US20220309011A1

  • Hardware Acceleration

    US20220374348A1

  • Convolution-based processing

    WO2020038551A1