Electronic device, method, and non-transitory computer-readable storage medium for acquiring instance
By loading pre-trained models into volatile memory with specific configuration information, the electronic device optimizes memory usage and enhances AI model performance for natural language processing, addressing inefficiencies in existing technologies.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2026-04-02
AI Technical Summary
Existing electronic devices equipped with AI technology face challenges in efficiently utilizing pre-trained models for natural language processing due to limitations in loading and executing different configuration information, leading to suboptimal performance and resource inefficiencies.
The solution involves an electronic device with non-volatile and volatile memory, where pre-trained models are loaded into volatile memory based on specific composition information, allowing for independent configuration changes without immediate execution, enabling efficient inference operations.
This approach enhances the performance and resource utilization of AI models by optimizing memory usage and enabling flexible configuration, thereby improving the efficiency and effectiveness of natural language processing tasks.
Smart Images

Figure KR2025008207_02042026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transient computer-readable storage medium for acquiring an instance
[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable storage medium for obtaining an instance.
[0002] With the advancement of electronic devices, technological developments related to electronic devices equipped with artificial intelligence (AI) technology have recently been underway. Electronic devices equipped with AI technology can provide various services to users. For example, electronic devices equipped with AI technology can provide responses to input prompts by performing natural language processing on the input prompts.
[0003] The information described above is provided as background information to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above may be applied as prior art related to the present disclosure.
[0004] The aspects of the present disclosure are to solve at least the problems and / or disadvantages mentioned above and to provide at least the advantages described below. Accordingly, the aspects of the present disclosure provide an electronic device, a method, and a non-transient computer-readable storage medium for obtaining an instance.
[0005] Additional aspects will be described in part in the following description and may become apparent from the description or be learned by practicing the presented embodiments.
[0006] According to an aspect of the present disclosure, an electronic device is provided. The electronic device may include a non-volatile memory comprising one or more storage media for storing instructions. The electronic device may include a volatile memory comprising one or more storage media. The electronic device may include at least one processor comprising processing circuitry. The at least one processor is communicately connected to the non-volatile memory and the volatile memory. The instructions may cause the electronic device to receive input data for utilizing the function of a pre-trained model stored in the non-volatile memory when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to acquire an instance according to the loaded first composition information based on loading first composition information of the pre-trained model into the volatile memory when executed individually or collectively by the at least one processor. The above instructions may cause the electronic device to load a second configuration information of the pre-learned model, which is distinct from the first configuration information of the pre-learned model, into the volatile memory, independently of executing the acquired instance to perform inference on the input data when executed individually or collectively by the at least one processor.
[0007] According to another aspect of the present disclosure, a method is provided for performing in an electronic device having a non-volatile memory and a volatile memory. The method may include the operation of receiving input data by the electronic device for utilizing the function of a pre-trained model stored in the non-volatile memory. The method may include the operation of acquiring an instance by the electronic device according to the loaded first composition information based on loading first composition information of the pre-trained model into the volatile memory. The method may include the operation of loading second composition information of the pre-trained model, which is distinct from the first composition information of the pre-trained model, into the volatile memory by the electronic device, independently of executing the acquired instance to perform inference on the input data.
[0008] According to another aspect of the present disclosure, one or more non-transient computer-readable storage media are provided for storing one or more computer programs that, when executed individually or collectively by at least one processor of an electronic device comprising non-volatile memory and volatile memory, cause the electronic device to perform operations. The operations may include receiving input data by the electronic device for utilizing the function of a pre-trained model stored in the non-volatile memory. The operations may include acquiring an instance by the electronic device according to the loaded first composition information based on loading first composition information of the pre-trained model into the volatile memory. The operations may include loading second composition information of the pre-trained model, which is distinct from the first composition information of the pre-trained model, into the volatile memory by the electronic device, independently of executing the acquired instance to perform inference on the input data.
[0009] Other aspects, advantages, and important features of the present disclosure will become apparent to those skilled in the art from the following detailed descriptions disclosing various embodiments of the present disclosure together with the accompanying drawings.
[0010] The above and other aspects, features, and advantages of specific embodiments of the present disclosure will become more apparent from the following description together with the accompanying drawings.
[0011] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment of the present disclosure.
[0012] FIG. 2 is a diagram illustrating a neural network executed in an electronic device according to one embodiment of the present disclosure.
[0013] FIG. 3 illustrates an example of a simplified block diagram of an electronic device according to one embodiment of the present disclosure.
[0014] FIG. 4 illustrates an example of loading composition information of a model stored in a non-volatile memory according to one embodiment of the present disclosure into a volatile memory.
[0015] FIG. 5 illustrates examples of operations of an electronic device for acquiring an instance according to a model stored in a non-volatile memory according to one embodiment of the present disclosure.
[0016] FIG. 6 illustrates examples of operations of an electronic device for obtaining an instance according to one embodiment of the present disclosure.
[0017] FIG. 7 illustrates an example of a second instance for verifying inference data of a first instance according to one embodiment of the present disclosure.
[0018] FIG. 8 illustrates examples of operations of an electronic device executing one or more instances according to one embodiment of the present disclosure.
[0019] It should be noted that similar reference numbers are used throughout the drawings to describe identical or similar elements, features, and structures.
[0020] The following description, with reference to the attached drawings, is provided to aid in a comprehensive understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. While this description includes various specific details to aid understanding, they should be considered merely illustrative. Accordingly, those skilled in the art will recognize that various changes and modifications are possible with respect to the various embodiments of the present disclosure without departing from the technical spirit and scope of the rights of the present disclosure. Additionally, descriptions of known functions and configurations may be omitted for the sake of clarity and brevity.
[0021] The terms and words used in the following description and claims are not limited to their dictionary meanings and are used by the inventor to enable a clear and consistent understanding of the present disclosure. Accordingly, it will be obvious to those skilled in the art that the following description of various embodiments of the present disclosure is merely for illustrative purposes and is not intended to limit the present disclosure as defined by the appended claims and their equivalents.
[0022] Unless clearly otherwise specified in the context, the singular forms of "a," "an," and "the" should be understood to include plural objects. Thus, for example, a reference to "a component surface" is interpreted to include one or more such surfaces.
[0023] In the various embodiments of the present disclosure described below, a hardware-based approach is described as an example. However, since the various embodiments of the present disclosure include techniques using both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.
[0024] Terms used in the following description to refer to data (e.g., weight data, graph data, input data, output data, inference data, token, configuration information), terms referring to values, terms for operation states (e.g., operation, process), terms referring to objects, terms referring to network entities, terms referring to device components, etc., are provided as examples for the convenience of explanation. Accordingly, the present disclosure is not limited to the terms described below, and other terms having equivalent technical meanings may be used.
[0025] Additionally, in this disclosure, expressions of "greater than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled; however, this is merely for the purpose of expressing an example and does not exclude descriptions of "greater than" or "less than." Conditions described as "greater than" may be replaced with "greater than," conditions described as "less than" may be replaced with "less than," and conditions described as "greater than and less than" may be replaced with "greater than and less than." Furthermore, "A" to "B" below refer to at least one of elements from A (including A) to B (including B). Below, "C" and / or "D" refers to including at least one of "C" or "D," i.e., {"C", "D", "C" and "D"}.
[0026] It should be recognized that the blocks of each flowchart and combinations of flowcharts may be executed by one or more computer programs containing instructions. The whole of the one or more computer programs may be stored in a single memory device, or the one or more computer programs may be divided so that each part is stored in a plurality of different memory devices.
[0027] Any functions or operations described in this document may also be processed by a single processor or a combination of multiple processors. The single processor or combination of multiple processors is a circuit that performs processing and includes circuits such as an application processor (AP) (e.g., CPU (central processing unit)), a communication processor (CP) (e.g., modem), a graphics processing unit (GPU), a neural processing unit (NPU) (e.g., an artificial intelligence chip), a wireless fidelity (Wi-Fi) chip, a Bluetooth® chip, a global positioning system (GPS) chip, a near field communication (NFC) chip, connectivity chips, a sensor controller, a touch controller, a fingerprint sensor controller, a display driver IC (integrated circuit), an audio codec (CODEC) chip, a universal serial bus (USB) controller, a camera controller, an image processing IC, a microprocessor unit (MPU), a system on chip (SoC), an IC, etc.
[0028] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment of the present disclosure.
[0029] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0030] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0031] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0032] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, software (e.g., program (140)) and input or output data for related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0033] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0034] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0035] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0036] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0037] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0038] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0039] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0040] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0041] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0042] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0043] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0044] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0045] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0046] The wireless communication module (192) is 4G (4 thIt can support 5G networks and next-generation communication technologies following the generation) network, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave (millimeter wave) band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0047] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0048] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0049] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0050] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0051] In this disclosure, technology related to artificial intelligence (or an artificial intelligence model) may be described. Functions related to artificial intelligence are operated through a processor (e.g., processor (120)) and memory (e.g., memory (130)). The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as a CPU (central processing unit), AP (application processor), DSP (Digital Signal Processor), etc., graphics-dedicated processors such as a GPU (graphic processing unit) or VPU (Vision Processing Unit), or artificial intelligence-dedicated processors such as an NPU (neural processing unit). One or more processors process input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0052] The predefined operation rules or artificial intelligence model are characterized by being created through learning. Here, being created through learning means that a predefined operation rules or artificial intelligence model configured to perform a desired characteristic (or purpose) is created by training a basic artificial intelligence model using a number of learning data by a learning algorithm. Such learning may be performed on the device (e.g., electronic device (101)) itself where the artificial intelligence according to the present disclosure is performed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0053] An artificial intelligence model (e.g., the model (320) of FIG. 3) may be composed of multiple neural network layers (e.g., the neural network (200) of FIG. 2). Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the results of operations of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized by the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.
[0054] An electronic device according to the present disclosure (e.g., electronic device (101)) may use an artificial intelligence model to recommend, execute, and / or infer a response to input data (e.g., a prompt). A processor (e.g., processor (120)) may perform a preprocessing step on the input data to convert it into a form suitable for use as input to an artificial intelligence model. An artificial intelligence model may be created through learning. Here, being created through learning means that a basic artificial intelligence model is trained using a number of training data by a learning algorithm, thereby creating a predefined rule of operation or an artificial intelligence model configured to perform a desired characteristic (or purpose). An artificial intelligence model may be composed of a number of neural network layers. Each of the number of neural network layers has a number of weight values and performs neural network operations through operations between the result of a previous layer and the number of weights. Inference prediction is a technology that logically reasones and predicts by judging information, and includes knowledge-based reasoning, probability-based reasoning, optimization prediction, preference-based planning, and recommendation.
[0055] FIG. 2 is a diagram illustrating a neural network (200) executed in an electronic device (e.g., the electronic device (101) of FIG. 1) according to one embodiment of the present disclosure. According to one embodiment, the neural network (200) of FIG. 2 may be obtained from a set of parameters stored in a memory (e.g., the memory (130) of FIG. 1) by the electronic device (101). For example, the neural network (200) may be an example of a model stored in the memory (130). For example, the set of parameters may be included in the composition information of the model stored in the memory (130).
[0056] Referring to FIG. 2, the neural network (200) may include a plurality of layers. For example, the neural network (200) may include an input layer (210), one or more hidden layers (220), and an output layer (230). The input layer (210) may correspond to a vector and / or matrix representing input data of the neural network (200). For example, the vector representing input data may have elements corresponding to the number of nodes included in the input layer (210). For example, the elements included in the matrix representing input data may correspond to each of the nodes included in the input layer (210). Based on the input data, signals generated at each of the nodes within the input layer (210) may be transmitted from the input layer (210) to the hidden layers (220). The output layer (230) can generate output data of the neural network (200) based on one or more signals received from the hidden layers (220). For example, the output data may correspond to a vector and / or matrix having elements corresponding to the number of nodes included in the output layer (230).
[0057] According to one embodiment, first nodes included in a specific layer among a plurality of layers included in a neural network (200) may correspond to a weighted sum of at least one of the second nodes of a layer preceding the specific layer within a sequence of a plurality of layers. According to one embodiment, an electronic device (101) may identify a weight to be applied to at least one of the second nodes from a set of parameters stored in memory (130). Training the neural network (200) may include the operation of changing and / or determining one or more weights related to the weighted sum.
[0058] Referring to FIG. 2, one or more hidden layers (220) may be positioned between an input layer (210) and an output layer (230) and may convert input data transmitted through the input layer (210) into a predictable value. The input layer (210), one or more hidden layers (220), and the output layer (230) may include a plurality of nodes. One or more hidden layers (220) may be convolution filters or fully connected layers in a convolutional neural network (CNN), or various types of filters or layers grouped based on special functions or features. In one embodiment, one or more hidden layers (220) may be layers based on a recurrent neural network (RNN) in which the output value is input back into the hidden layer at the current time. According to one embodiment, a neural network (200) may include a number of hidden layers (220) to form a deep neural network. Training a deep neural network is called deep learning. Among the nodes of the neural network (200), a node included in the hidden layers (220) is referred to as a hidden node.
[0059] According to one embodiment, nodes included in the input layer (210) and one or more hidden layers (220) may be connected to each other through connecting lines having connecting weights, and nodes included in the hidden layer and output layer (230) may also be connected to each other through connecting lines having connecting weights. Tuning and / or training the neural network (200) may mean changing the connecting weights between nodes included in each of the layers included in the neural network (200) (e.g., input layer (210), one or more hidden layers (220), and output layer (230)). For example, tuning of the neural network (200) may be performed based on supervised learning and / or unsupervised learning.
[0060] According to one embodiment, in a state of acquiring a neural network (200), an electronic device (101) may identify weights corresponding to connecting lines connecting an input layer (210), one or more hidden layers (220), and / or an output layer (230) stored in memory (e.g., memory (130)). The electronic device (101) may acquire weighted sums based on connecting lines sequentially along a plurality of layers of the neural network (200) (e.g., the input layer (210), the one or more hidden layers (220), and the output layer (230)) in order to acquire output data from the neural network (200) based on the identified weights. The acquired weighted sums may be stored in at least one processor (e.g., processor (120)) and / or memory (130) of the electronic device (101). For example, the electronic device (101) can repeatedly update the weighted sum stored in memory (130) as it sequentially obtains the weighted sum along a plurality of layers.
[0061] Each of the multiple layers of the neural network (200) may have an independent data type and / or precision. For example, if the connecting lines between the first layer and the second layer among the multiple layers have weights based on a first data type for representing a floating-point number, the electronic device (101) may obtain weighted sums based on the first data type from the numerical values and weights corresponding to the nodes of the first layer. In the above example, if the connecting lines between the second layer and the third layer among the multiple layers have weights based on a second data type for representing an integer number, the electronic device (101) may obtain weighted sums based on the second data type from the obtained weighted sums and the weights based on the second data type.
[0062] According to one embodiment, when a plurality of layers have different data types, the electronic device (101) can obtain weighted sums corresponding to each of the plurality of layers based on the different data types by using at least one processor (e.g., processor (120)). As the electronic device (101) accesses the memory (130) based on the weighted sums obtained based on the different data types, the bandwidth of the memory (130) can be utilized more efficiently. For example, as the bandwidth of the memory (130) is utilized more efficiently, the electronic device (101) can obtain output data more quickly from the neural network (200) based on the plurality of layers.
[0063] According to one embodiment, the electronic device (101) may store sets of parameters representing each of a plurality of neural networks having different precisions. For example, a neural network associated with super resolution for upscaling images and / or videos may require a precision of a data type for representing floating-point numbers based on 32 bits. For example, a neural network associated with super resolution for upscaling images and / or videos may require a precision of a data type for representing floating-point numbers based on 16 bits (e.g., half-precision floating-point format defined by IEEE 754). For example, a neural network for recognizing subjects included in images and / or videos may require a precision of a data type for representing integers based on 8 bits and / or 4 bits. For example, a neural network for performing handwriting recognition may require a precision of a data type for representing an integer based on first bits and / or second bits. For example, an electronic device (101) may perform an operation to obtain a weighted sum based on different precisions corresponding to each of a plurality of neural networks.
[0064] The model described in this disclosure (e.g., the model (320) of FIG. 3) may include a large language model (LLM) (or a large multimodal model (LMM)). However, it is not limited thereto. For example, although the description of the LLM (or LMM) will be provided later, it is obvious that the artificial intelligence neural network of this disclosure may include not only a language model but also various foundation models such as a code model, an image model, and other artificial intelligence neural network models.
[0065] An artificial intelligence model according to an embodiment of the present disclosure (e.g., the model (320) of FIG. 3) may refer to an LLM, which is an artificial neural network-based language model that has learned a large amount of text data through prior training. An LLM may contain relatively more parameters (e.g., more than 10 billion) than existing general language models. An LLM may use a transformer artificial neural network structure based on an attention mechanism.
[0066] The attention mechanism is a technique that helps artificial intelligence models focus on important parts within input data. The attention mechanism can be utilized to predict output data by predicting the extent to which parts of time-series input data (e.g., input data such as voice or video, or input data for a specific layer of a neural network) contribute to the intermediate or final output of the neural network. While the recurrent neural network (RNN) structure, which processes each element of a sequence sequentially, suffers from degraded prediction performance when there is information dependency over long time-series distances, the attention mechanism can account for information dependency over long time-series distances by controlling the degree of weight concentration within the overall context (or part thereof) of the input data.
[0067] A transformer can be composed of an encoder-decoder structure. The encoder processes input data to output compressed information (e.g., contextual representation), and the decoder processes the compressed information to output data in token units. Each of the encoder and decoder may include an independent attention network and a cross-attention network connecting the encoder and the decoder.
[0068] According to one embodiment, LLM learning may include pre-training and / or fine-tuning. Pre-training is a process of enabling the LLM to acquire general linguistic knowledge using a large amount of text data, and may include, for example, self-supervised learning that predicts the next word using the previous sequence of words in a sequence of text. Fine-tuning is a process of training the LLM to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task, and the LLM may be further trained through supervised learning (or adaptive learning) using a dataset suitable for the domain purpose based on a pre-trained model. The LLM may perform tasks with text input containing natural language, referred to as a prompt.
[0069] For example, fine-tuning can be omitted during LLM training. Users can control the prompts input to the LLM to improve the performance of desired tasks. In a manner similar to in-context learning or zero-shot / few-shot learning, task examples and / or guides for performing the task can be added to the prompts. Examples of publicly available LLMs include BERT (bidirectional encoder representations from transformer) and GPT (generative pre-trained transformer).
[0070] The term 'LLM' may refer to the language neural network model itself, but it may also refer to models of LLM-based applications (e.g., chatbots, translation, summarization, text classification, sentence generation). For example, an LLM-based chatbot or translator such as ChatGPT may also be referred to as 'LLM'. 'LLM' may also include an inference engine utilizing an LLM neural network model. For example, "inputting an input prompt into the LLM" may mean "inputting an input prompt into an LLM-based inference engine." For example, "the output of the LLM for the input prompt" may refer to the output information of the last neural network layer of the LLM (or output information modified through additional processing) obtained when the input prompt is input into an LLM-based inference engine.
[0071] FIG. 3 illustrates an example of a simplified block diagram of an electronic device (301) according to one embodiment of the present disclosure. The electronic device (301) may be an example of the electronic device (101) of FIG. 1.
[0072] Referring to FIG. 3, the electronic device (301) may include at least one processor (300) and / or memory (310). For example, at least one processor (300) and / or memory (310) may be electronically and / or operably coupled with each other by a communication bus. Hereinafter, operably coupled hardware components may mean that a direct or indirect connection between hardware components is established wired or wirelessly so that a second hardware component is controlled by a first hardware component among the hardware components. Although the hardware components illustrated in FIG. 3 are illustrated based on different blocks, the present disclosure is not limited thereto.
[0073] At least one processor (300) may include a hardware component for processing data based on executing instructions. At least one processor (300) may be configured to execute instructions stored in memory (310) individually or collectively. At least one processor (300) may include a processing circuit. For example, the hardware component for processing data may include an arithmetic and logic unit (ALU), a floating point unit (FPU), and a field programmable gate array (FPGA). For example, the hardware component for processing data may include a central processing unit (CPU), a graphic processing unit (GPU), a display processing unit (DPU), a neural processing unit (NPU), a digital signal processor (DSP), an application processor (AP), and / or a microcontroller (MCU). For example, the NPU may include a hardware component dedicated to computations related to the model (320). For example, the NPU may include multiple circuits for performing operations (e.g., multiplication and / or addition) that are performed sequentially and / or in parallel based on the model (320). Multiple circuits included within the NPU may be referred to as neural engines. The NPU may perform said operations based on a specified data type (e.g., floating-point number and / or integer) associated with the model (320). For example, the GPU may include one or more pipelines that perform multiple operations for executing instructions related to computer graphics and / or parallel operations.For example, the GPU pipeline may include a graphics pipeline or a rendering pipeline for generating a 3D image and generating a 2D raster image from the generated 3D image. By using the graphics pipelines, operations related to an artificial neural network (e.g., neural network (200)) can be executed substantially simultaneously.
[0074] According to one embodiment, at least one processor (300) may include one or more cores. For example, at least one processor (300) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.
[0075] According to one embodiment, in terms of the subject performing the operations of an artificial neural network directed by the model (320) below, at least one processor (300) may be referred to as an artificial intelligence (AI) accelerator. The AI accelerator may be referred to as an accelerator.
[0076] The memory (310) may include one or more storage media. The memory (310) may include non-volatile memory (311) and / or volatile memory (312). The non-volatile memory (311) may include a hardware component for storing data and / or instructions that are input to and / or output from at least one processor (300). The non-volatile memory (311) may include, for example, at least one of ROM (read-only memory), PROM (programmable ROM), EPROM (erasable PROM), EEPROM (electrically erasable PROM), flash memory, hard disk, compact disk, or EMMC (embedded multimedia card). For example, within the non-volatile memory (311) of the electronic device (301), one or more instructions (or commands) representing operations and / or actions to be performed on data by at least one processor (300) of the electronic device (301) may be stored. A set of one or more instructions may be referred to as a program, firmware, operating system, process, routine, sub-routine and / or application. The volatile memory (312) may include at least one of RAM (random-access memory), DRAM (dynamic RAM), SRAM (static RAM), cache RAM, or PSRAM (pseudo SRAM). For example, the memory (310) may correspond to the memory (130) of FIG. 1. For example, the non-volatile memory (311) may correspond to the non-volatile memory (134) of FIG. 1. For example, the volatile memory (312) may correspond to the volatile memory (132) of FIG. 1.
[0077] According to one embodiment, the non-volatile memory (311) may include a model (320). The model (320) may be stored within the non-volatile memory (311). The model (320) may include the neural network (200) of FIG. 2. For example, the model (320) stored within the non-volatile memory (311) may be represented by a file stored within the non-volatile memory (311). For example, the electronic device (301) may execute functions similar to human cognitive processes or learning processes based on the model (320). Based on operations indicated by the model (320) and performed sequentially by a plurality of parameters, the electronic device (301) may output data containing generalized information about input data (e.g., a prompt). The non-volatile memory (311) of the electronic device (301) may store composition information of the model (320). For example, the configuration information may include graph data related to the model (320) and / or weight data related to the model (320). For example, the structure and operations of the model (320) may be described by the graph data. For example, the graph data may represent the operations, variables, and / or connection relationships of the model (320). According to one embodiment, the electronic device (301) may perform operations directed by the model (320) based on the weight data of the model (320). For example, the weight data may include weights assigned to a plurality of nodes directed by the model (320) and / or connections between the plurality of nodes. As an example, but not limited to, the configuration information may include hyperparameters related to the model (320).For example, hyperparameters may include at least one of a learning rate, a cost function, a regularization parameter, a mini-batch size, the number of training iterations, the number of hidden layers, metaparameters, or free parameters.
[0078] An electronic device (301) can load a model (320) stored in non-volatile memory (311) into volatile memory (312). For example, the electronic device (301) can load (or copy) configuration information of the model (320) stored in non-volatile memory (311) into volatile memory (312). Based on the configuration information loaded into volatile memory (312), the electronic device (301) can obtain instances (e.g., multiple instructions) for performing operations directed by the model (320). At least one processor (300) (e.g., an accelerator) can execute one or more functions related to the model (320) stored in volatile memory (312) based on the instances. One or more functions may include a function for performing inference on input data based on the model (320), a function for performing image-based object recognition, speech recognition, and / or handwriting recognition using a pre-trained model (e.g., model (320)), and at least one of a function personalized to the user of the electronic device (301) based on a neural network. However, the embodiments are not limited thereto.
[0079] According to one embodiment, an electronic device (301) can perform operations related to input data based on a model (320) using an accelerator. The electronic device (301) can obtain output data (e.g., inference data) from input data that is input to an instance according to the model (320) based on the performance of chained (or serial or consecutive) operations based on a plurality of parameters of the model (320).
[0080] According to one embodiment, the model (320) (or an instance according to the model (320)) may include a plurality of nodes. The plurality of nodes may be divided into units of layers. In one embodiment in which a plurality of parameters include weights connecting two nodes of different layers of the model (320) (or an instance according to the model (320)), the electronic device (301) may obtain values corresponding to nodes of another layer connected to the specific layer by applying weights to values corresponding to nodes of a specific layer.
[0081] According to one embodiment, the model (320) stored in non-volatile memory (311) may be a pre-trained model. For example, the pre-trained model may include an on-device model. For example, the pre-trained model may be a freeze model. For example, a freeze model may be described as a model in which at least a portion of the configuration information is fixed. For example, a freeze model may be a model in which the weight data is fixed. For example, a freeze model may be referred to as a model that has completed training and can perform inference on input data. For example, a freeze model may be referred to as an offline compiled model.
[0082] According to one embodiment, the model (320) may support an early exit inference method and / or self-speculative decoding. For example, the early exit inference method may be referred to as a method of performing inference using only some of the operations based on the model (320). The early exit inference method may have a relatively high inference speed. For example, self-speculative decoding may be referred to as a method for improving the accuracy of inference data by performing verification on the speculated inference data (e.g., inference data obtained according to the early exit inference method).
[0083] According to one embodiment, the model (320) may include a draft model and a full layer model. The draft model may be a simplified model designed to generate a fast initial result. The electronic device (301) may perform inference faster than the full layer model by using the draft model. The electronic device (301) may perform verification of inference data (e.g., tokens), which is the result of inference performed by the draft model, by using the full layer model. Although the accuracy of the inference data of the draft model may be relatively low, the accuracy of the inference data of the draft model may be guaranteed by verifying (or supplementing) the inference data of the draft model by the full layer model. The electronic device (301) may perform inference on the input data faster when using both the draft model and the full layer model than when using only the full layer model. For example, the time to first token (TTFT) when using both the draft model and the full layer model may be smaller than the TTFT when using only the full layer model. The layers in the draft model may correspond to at least some of the layers in the full layer model. For example, the number of layers in the draft model may be fewer than the number of layers in the full layer model.
[0084] According to one embodiment, when using an on-device model (e.g., model (320)), the electronic device (301) may independently store a plurality of draft models and a full layer model, respectively, in non-volatile memory (311). For example, the electronic device (301) may experience a problem in which the remaining capacity of the non-volatile memory (311) decreases as the number of draft models increases. For example, the electronic device (301) may experience a problem in which the usage of the volatile memory (312) increases as the number of draft models increases.
[0085] In the present disclosure, a technique may be described in which an electronic device (301) acquires one or more instances by reading a model (320) stored in a non-volatile memory (311). Additionally, the electronic device (301) can reduce the usage of the volatile memory (312) by loading configuration information of a model (320) into the volatile memory (312). This method will be described and illustrated with reference to FIGS. 4, FIGS. 5, FIGS. 6, FIGS. 7, and / or FIGS. 8.
[0086] FIG. 4 illustrates an example of loading composition information of a model (320) stored in a non-volatile memory (311) according to one embodiment of the present disclosure into a volatile memory (312).
[0087] Referring to FIG. 4, the model (320) may be stored in non-volatile memory (311). For example, the model (320) may be represented by a file stored in non-volatile memory (311). For example, the model (320) may be an on-device model. For example, the model (320) may be a freeze model. For the freeze model, the descriptions of the freeze model in FIG. 3 may be referenced. For example, the model (320) may be pre-compiled. For example, the model (320) may be a pre-trained model. The model (320) may be composed of multiple layers. For example, the model (320) may include first layers (401), second layers (402), and / or third layers (403).
[0088] According to one embodiment, an electronic device (e.g., electronic device (301)) may acquire or create an instance to execute the function of the model (320). The electronic device (301) may load (or copy) the model (320) into volatile memory (312) to acquire the instance. For example, the electronic device (301) may load configuration information of the model (320) into volatile memory (312). For example, the electronic device (301) may load configuration information of the model (320) into volatile memory (312) by sequentially reading the layers of the model (320). For example, the electronic device (301) may load first configuration information (410) into volatile memory (312) by reading the first layers (401). The first configuration information (410) may correspond to the first layers (401). For example, the first configuration information (410) may include weight data of nodes within the first layers (401) and / or graph data of the first layers (401). As an example without limitation, the first configuration information (410) may include early termination information. For example, the electronic device (301) may acquire or create a first instance according to the first configuration information based on identifying the early termination information.
[0089] According to one embodiment, the electronic device (301) can read the second layers (402) after reading the first layers (401). For example, the electronic device (301) can load the second configuration information (420) into the volatile memory (312) by reading the first layers (401). The second configuration information (420) may correspond to the second layers (402). For example, the second configuration information (420) may include weight data of nodes within the second layers (402) and / or graph data of the second layers (402). As an example without limitation, the second configuration information (420) may include early termination information. For example, the electronic device (301) may acquire or create a second instance according to the first configuration information (410) and the second configuration information (420) based on identifying the early termination information.
[0090] According to one embodiment, the electronic device (301) can read the third layers (403) after reading the second layers (402). For example, the electronic device (301) can load the third configuration information (430) into the volatile memory (312) by reading the third layers (403) after reading the second layers (402). The third configuration information (430) may correspond to the third layers (403). For example, the third configuration information (430) may include weight data of nodes within the third layers (403) and / or graph data of the third layers (403). As an example without limitation, the third configuration information (430) may include early termination information. For example, the electronic device (301) may acquire or create a third instance according to the first configuration information (410), the second configuration information (420), and the third configuration information (430) based on identifying early termination information.
[0091] According to one embodiment, the size (or capacity) of the configuration information (e.g., first configuration information (410), second configuration information (420), and third configuration information (430)) loaded into the volatile memory (312) may correspond to the size (or capacity) of the model (320). The first configuration information (410) loaded into the volatile memory (312) may be used to obtain a first instance, to obtain a second instance, and to obtain a third instance. The second configuration information (420) loaded into the volatile memory (312) may be used to obtain a second instance and to obtain a third instance. The electronic device (301) can efficiently utilize the usage of the volatile memory (312) and the capacity of the non-volatile memory (311) by sharing the configuration information (e.g., first configuration information (410), second configuration information (420)) loaded in the volatile memory (312) to acquire instances (e.g., second instance, third instance).
[0092] FIG. 5 illustrates examples of operations of an electronic device (e.g., electronic device (301)) that acquires an instance according to a model (e.g., model (320)) stored in a non-volatile memory (e.g., non-volatile memory (311)) according to one embodiment of the present disclosure.
[0093] Referring to FIG. 5, the model (320) may be composed of a plurality of layers. For example, the model (320) may include first layers (501) (e.g., first layers (401)), second layers (502) (e.g., second layers (402)), and / or third layers (503) (e.g., third layers (403)). The electronic device (301) may load the model (320) into volatile memory (312) by reading the plurality of layers. For example, the electronic device (301) may load configuration information of the model (320) into volatile memory (312). The descriptions of FIG. 4 may be referenced for the electronic device (301) to load configuration information of the model (320) into volatile memory (312).
[0094] According to one embodiment, the electronic device (301) can load configuration information for each of the first layers (501) into the volatile memory (312) by sequentially reading the first layers (501) (e.g., layer (501-1), layer (501-2), layer (501-3), layer (501-4), and layer (501-5)). For example, the electronic device (301) can load configuration information corresponding to layer (501-1) into the volatile memory (312) by reading layer (501-1). For example, the electronic device (301) can load configuration information corresponding to layer (501-2) into the volatile memory (312) by reading layer (501-2) after reading layer (501-1). For example, the electronic device (301) can load configuration information corresponding to layer (501-3) into the volatile memory (312) by reading layer (501-3) after reading layer (501-2). For example, the electronic device (301) can load configuration information corresponding to layer (501-4) into the volatile memory (312) by reading layer (501-4) after reading layer (501-3). For example, the electronic device (301) can load configuration information corresponding to layer (501-5) into the volatile memory (312) by reading layer (501-5) after reading layer (501-4). According to one embodiment, the configuration information corresponding to layer (501-5) may include an end point. However, it is not limited thereto. The electronic device (301) may store end point information indicating at least one end point in the non-volatile memory (311). The electronic device (301) can execute operation 510 depending on the end point.
[0095] According to one embodiment, in operation 510, the electronic device (301) may acquire or create a first instance. The electronic device (301) may acquire or create a first instance based on first configuration information (e.g., first configuration information (410)) corresponding to the first layers (501). The electronic device (301) may execute operation 511 in response to acquiring the first instance.
[0096] In operation 511, the electronic device (301) can execute the first instance in response to acquiring the first instance. By executing the first instance, the electronic device (301) can execute a function corresponding to the first layers (501). By executing the first instance, the electronic device (301) can perform inference on the input data.
[0097] According to one embodiment, the electronic device (301) can read the second layers (502) independently of operation 510 and / or operation 511. Since the execution of operation 510 and / or operation 511 of the electronic device (301) and the operation of the electronic device (301) reading the second layers (502) are independent, the execution of operation 510 and / or operation 511 of the electronic device (301) and the operation of the electronic device (301) reading the second layers (502) can be executed simultaneously. The electronic device (301) can load the configuration information of each of the second layers (502) into the volatile memory (312) by sequentially reading each of the second layers (502). For example, the electronic device (301) can identify the end point included in the configuration information of the last layer (e.g., layer 10) within the second layers (502). For example, the electronic device (301) can execute operation 520 depending on the end point.
[0098] According to one embodiment, in operation 520, the electronic device (301) may acquire or create a second instance. The electronic device (301) may acquire or create a second instance based on first configuration information (e.g., first configuration information (410)) corresponding to the first layers (501) and second configuration information (e.g., second configuration information (420)) corresponding to the second layers (502). The electronic device (301) may execute operation 521 in response to acquiring the second instance.
[0099] In operation 521, the electronic device (301) may execute the second instance in response to acquiring the second instance. By executing the second instance, the electronic device (301) may execute functions corresponding to the first layers (501) and the second layers (502). By executing the second instance, the electronic device (301) may perform inference on the input data. According to one embodiment, the second instance may be used for verification of inference data, which is the result of inference performed by executing the first instance. Verification will be described with reference to FIG. 7.
[0100] As a non-limiting example, the electronic device (301) may acquire or generate a second instance based on the second configuration information among the first configuration information corresponding to the first layers (501) and the second configuration information corresponding to the second layers (502). The electronic device (301) may execute the acquired second instance. For example, the electronic device (301) may execute a function corresponding to the second layers (502) by executing the second instance. The electronic device (301) may perform inference on the input data by executing the second instance. According to one embodiment, the electronic device (301) may perform verification of the inference data, which is the result of the inference, using a third instance to be described later.
[0101] According to one embodiment, the electronic device (301) can read the third layers (503) independently of operation 520 and / or operation 521. Since the execution of operation 520 and / or operation 521 of the electronic device (301) and the operation of the electronic device (301) reading the third layers (503) are independent, the execution of operation 520 and / or operation 521 of the electronic device (301) and the operation of the electronic device (301) reading the third layers (503) can be executed simultaneously. The electronic device (301) can load the configuration information of each of the third layers (503) into the volatile memory (312) by sequentially reading each of the third layers (503). For example, the electronic device (301) can identify the end point included in the configuration information of the last layer (e.g., layer N) within the third layers (503). For example, the electronic device (301) can execute operation 530 depending on the end point.
[0102] According to one embodiment, in operation 530, the electronic device (301) may acquire or create a third instance. The electronic device (301) may acquire or create a third instance based on first configuration information corresponding to the first layers (501) (e.g., first configuration information (410)), second configuration information corresponding to the second layers (502) (e.g., second configuration information (420)), and third configuration information corresponding to the third layers (503) (e.g., third configuration information (430)). The electronic device (301) may execute operation 531 in response to acquiring the third instance.
[0103] In operation 531, the electronic device (301) may execute the third instance in response to acquiring the third instance. By executing the third instance, the electronic device (301) may execute functions corresponding to the first layers (501), the second layers (502), and the third layers. For example, the electronic device (301) may execute functions corresponding to the model (320). By executing the third instance, the electronic device (301) may perform inference on the input data. According to one embodiment, the third instance may be used to verify the inference data, which is the result of the inference performed by executing the first instance. According to one embodiment, the third instance may be used to verify the inference data, which is the result of the inference performed by executing the second instance. Verification will be described with reference to FIG. 7.
[0104] FIG. 6 illustrates examples of operations of an electronic device (e.g., electronic device (301)) for obtaining an instance according to one embodiment of the present disclosure.
[0105] Referring to FIG. 6, in operation 601, an electronic device (301) (e.g., at least one processor (300)) may receive input data for utilizing the functions of a pre-trained model (e.g., model (320)) stored in non-volatile memory (e.g., non-volatile memory (311)). For example, the functions of the pre-trained model (320) may vary depending on the learning.
[0106] In operation 603, an electronic device (301) (e.g., at least one processor (300)) may load (or copy) first configuration information (e.g., first configuration information (410)) of a pre-trained model (320) into volatile memory (e.g., volatile memory (312)). For example, the first configuration information (410) may correspond to first layers (e.g., first layers (401, 501)). For example, the first configuration information (410) may include weight data of the first layers (401, 501) and / or graph data of the first layers (401, 501).
[0107] In operation 605, an electronic device (301) (e.g., at least one processor (300)) may acquire or create a first instance according to the first configuration information (410). By acquiring the first instance, the electronic device (301) may execute functions according to the first layers (401, 501).
[0108] In operation 607, an electronic device (301) (e.g., at least one processor (300)) can perform inference on input data based on executing a first instance. For example, the electronic device (301) can execute a first instance to predict or process inference data on input data. For example, the electronic device (301) can execute a first instance to obtain inference data on input data. For example, if the function of the first instance is translation, the electronic device (301) can obtain translation text for input text. For example, if the function of the first instance is image classification, the electronic device (301) can obtain classification data for input images.
[0109] In operation 609, an electronic device (301) (e.g., at least one processor (300)) may load (or copy) second configuration information (e.g., second configuration information (420)) of a pre-trained model (320) into volatile memory (e.g., volatile memory (312)). For example, the second configuration information (420) may correspond to second layers (e.g., second layers (402, 502)). For example, the second configuration information (420) may include weight data of the second layers (402, 502) and / or graph data of the second layers (402, 502).
[0110] According to one embodiment, the electronic device (301) may execute operation 609 independently of operation 605 and / or operation 607. For example, the electronic device (301) may execute operation 609 while operation 605 and / or operation 607 is being executed.
[0111] In operation 611, an electronic device (301) (e.g., at least one processor (300)) may acquire or create a second instance according to first configuration information (410) and second configuration information (420). By acquiring the second instance, the electronic device (301) may execute functions according to the first layers (401, 501) and the second layers (402, 502).
[0112] According to one embodiment, the number of layers constituting the first instance may be less than the number of layers constituting the second instance. For example, the layers constituting the first instance may correspond to at least some of the layers constituting the second instance. For example, the nodes of the first layer constituting the first instance may correspond to the nodes of the second layer constituting the second instance. For example, the weight assigned to the nodes of the first layer may correspond to the weight assigned to the nodes of the second layer. For example, the weight assigned to the nodes of the first layer may be the same as the weight assigned to the nodes of the second layer. For example, the nodes of the first layer constituting the first instance and / or the weight assigned to those nodes may be the same as the nodes of the second layer constituting the second instance and / or the weight assigned to those nodes, and may be used in duplicate.
[0113] In operation 613, the electronic device (301) (e.g., at least one processor (300)) can perform inference on the input data based on executing a second instance. For example, the electronic device (301) can execute the second instance to predict or process inference data on the input data. For example, the electronic device (301) can execute the second instance to obtain inference data on the input data. For example, if the function of the second instance is translation, the electronic device (301) can obtain translation text for the input text. For example, if the function of the second instance is image classification, the electronic device (301) can obtain classification data for the input image.
[0114] According to one embodiment, a second instance may be used to verify inference data inferred by a first instance. For example, an electronic device (301) may determine whether the second inference data output by the second instance corresponds to the first inference data output by the first instance. Verification will be described and illustrated in FIG. 7.
[0115] According to one embodiment, the electronic device (301) can efficiently utilize the usage of the volatile memory (312) and the capacity of the non-volatile memory (311) by using the first configuration information (410) loaded into the volatile memory (312) to obtain a second instance.
[0116] In operation 615, an electronic device (301) (e.g., at least one processor (300)) may load (or copy) third configuration information (e.g., third configuration information (430)) of a pre-trained model (320) into volatile memory (e.g., volatile memory (312)). For example, the third configuration information (430) may correspond to third layers (e.g., third layers (403, 503)). For example, the third configuration information (430) may include weight data of the third layers (403, 503) and / or graph data of the third layers (403, 503).
[0117] According to one embodiment, the electronic device (301) may execute operation 615 independently of operation 611 and / or operation 613. For example, the electronic device (301) may execute operation 615 while operation 611 and / or operation 613 is being executed.
[0118] According to one embodiment, the electronic device (301) may acquire or create a third instance according to the first configuration information (410), the second configuration information (420), and the third configuration information (430). By acquiring the third instance, the electronic device (301) may execute functions according to the first layers (401, 501), the second layers (402, 502), and the third layers (403, 503). For example, by acquiring the third instance, the electronic device (301) may execute functions according to the model (320).
[0119] According to one embodiment, a third instance may be used to verify inferred data inferred by a first instance. A third instance may be used to verify inferred data inferred by a second instance. Verification will be described and illustrated in FIG. 7.
[0120] The number of instances (e.g., first instance, second instance, and third instance) obtained by the electronic device (301) in FIG. 6 is one embodiment and is not limited thereto. For example, the number of instances generated from the model (320) may be determined according to the configuration information of the model (320). As an example not limited thereto, the number of instances generated from the model (320) may be determined according to the end point information.
[0121] FIG. 7 illustrates an example of a second instance (702) for verifying inference data of a first instance (701) according to an embodiment of the present disclosure. The first instance (701) may be an example of the first instance of FIG. 5 and / or FIG. 6. The second instance (702) may be an example of the second instance of FIG. 5 and / or FIG. 6. In FIG. 7, the first instance (701) and the second instance (702) are described as instances obtained based on an autoregressive model, but it is obvious that the embodiment is not limited thereto. The first instance (701) and the second instance (702) may include instances obtained based on various artificial intelligence models.
[0122] Referring to FIG. 7, the number of layers (710) of the first instance (701) may be less than the number of layers (720) of the second instance (702). For example, the layers (710) of the first instance (701) may correspond to at least some of the second layers (720) of the second instance (702). For example, the second layers (720) of the second instance (702) may include the layers (710) of the first instance (701).
[0123] According to one embodiment, the first instance (701) and the second instance (702) can be obtained based on a model (e.g., model (320)) stored in non-volatile memory (e.g., non-volatile memory (311)).
[0124] According to one embodiment, the input data may include a token (711) and / or a token (712). For example, an electronic device (e.g., electronic device (301)) may execute a first instance (701) to perform inference on the token (711) and the token (712). For example, the electronic device (301) may output or obtain a token (713) by the first instance (701). For example, the electronic device (301) may execute a first instance (701) to perform inference on the token (713). For example, the electronic device (301) may output or obtain a token (714) by the first instance (701). For example, the first inference data may include a token (713) and / or a token (714).
[0125] According to one embodiment, the electronic device (301) can perform verification of the first inference data obtained by the first instance (701) using the second instance (702). For example, the electronic device (301) can perform inference on the input data and the first inference data by executing the second instance (702). For example, the inference on the input data and the first inference data can be performed in parallel.
[0126] For example, the electronic device (301) can perform inference on tokens (711) and (712). The electronic device (301) can output or acquire token (723). For example, the electronic device (301) can perform inference on token (713). The electronic device (301) can output or acquire token (724). For example, the electronic device (301) can perform inference on token (714). The electronic device (301) can output or acquire token (725).
[0127] The electronic device (301) can perform verification of the token (713) by comparing the token (713) obtained by the first instance (701) with the token (723) obtained by the second instance (702). For example, the electronic device (301) can determine whether the token (713) obtained by the first instance (701) corresponds to the token (723) obtained by the second instance (702). For example, the electronic device (301) can remove the token (713) based on the determination that the token (713) does not correspond to the token (723). The electronic device (301) can store the token (723) among the token (713) and the token (723) based on the determination that the token (713) does not correspond to the token (723). For example, the electronic device (301) may decide to trust the token (713) based on the determination that the token (713) corresponds to the token (723). For example, the electronic device (301) may compare the token (714) obtained by the first instance (701) with the token (724) obtained by the second instance (702) based on the determination that the token (713) corresponds to the token (723).
[0128] According to one embodiment, the electronic device (301) can perform verification of the token (714) by comparing the token (714) obtained by the first instance (701) with the token (724) obtained by the second instance (702). For example, the electronic device (301) can determine whether the token (714) obtained by the first instance (701) corresponds to the token (724) obtained by the second instance (702). For example, the electronic device (301) can remove the token (714) based on the determination that the token (714) does not correspond to the token (724). The electronic device (301) can store the token (724) among the token (714) and the token (724) based on the determination that the token (714) does not correspond to the token (724). For example, the electronic device (301) may decide to trust the token (714) based on the determination that the token (714) corresponds to the token (724). For example, the electronic device (301) may decide to trust the token (725) obtained by the second instance (702) based on the determination that the token (714) corresponds to the token (724).
[0129] According to one embodiment, the electronic device (301) can perform verification of the first inference data obtained by the first instance (701) after executing the first instance (701) a preset number of times (e.g., 2 times).
[0130] According to one embodiment, since the number of layers (710) of the first instance (701) is less than the number of layers (720) of the second instance (702), the time taken to perform inference on a token using the first instance (701) may be less than the time taken to perform inference on a token using the second instance (702). For example, it may be assumed that the preset number of times is 2, the time taken to perform inference on a token using the first instance (701) is 2 ms, and the time taken to perform inference on a token using the second instance (702) is 10 ms. For example, the time taken for the electronic device (301) to execute the first instance 2 times to acquire the token (713) and the token (714) may be 4 ms. For example, when the electronic device (301) executes the second instance (702) once to verify the token (713) and the token (714), the time consumed may be 10 ms. In the verification, when the token (723) corresponds to the token (713) and the token (724) corresponds to the token (714), 14 ms may be consumed to obtain the token (713) determined to be trusted by the electronic device (301), the token (714) determined to be trusted by the electronic device (301), and the token (725).
[0131] According to one embodiment, the electronic device (301) may execute only the second instance (702) among the first instance (701) and the second instance (702) to obtain the token (723), token (724), and token (725), and may execute the second instance (702) three times. 30 ms may be consumed for the electronic device (301) to execute only the second instance (702) to obtain the token (723), token (724), and token (725).
[0132] In an embodiment of the present disclosure, the quality of performing inference on input data by executing the first instance (701) and the second instance (702) may be higher than the quality of performing inference on input data by executing only the second instance (702). For example, the high quality may include less time consumed to obtain output data (e.g., the final result of the inference).
[0133] FIG. 8 illustrates examples of operations of an electronic device (e.g., electronic device (301)) that executes one or more instances according to one embodiment of the present disclosure.
[0134] Referring to FIG. 8, in operation 801, an electronic device (301) (e.g., at least one processor (300)) can identify a pre-trained model (e.g., model (320)) and early termination information. The pre-trained model (320) may be stored in non-volatile memory (e.g., non-volatile memory (311)). For example, the early termination information may be stored in a file format within the non-volatile memory (311). However, it is not limited thereto. For example, the early termination information may be included within the configuration information of the pre-trained model (320).
[0135] In operation 803, an electronic device (301) (e.g., at least one processor (300)) may acquire one or more instances based on a pre-learned model (320) and early termination information. The electronic device (301) may determine the number of one or more instances based on the pre-learned model (320) and early termination information. The electronic device (301) may acquire or generate one or more instances based on reading the pre-learned model (320). For example, the electronic device (301) may stop reading the pre-learned model (320) based on the determination that the number of acquired instances is the number of one or more instances determined above.
[0136] In operation 805, an electronic device (301) (e.g., at least one processor (300)) may execute one or more instances. The electronic device (301) may execute each of the one or more instances in response to acquiring the instance. The electronic device (301) may execute one or more instances to perform inference on input data.
[0137] In an embodiment according to the present disclosure, an electronic device (e.g., electronic device (301)) may store a pre-trained model (e.g., pre-trained model (320)) in a non-volatile memory (e.g., non-volatile memory (311)). The electronic device (301) may acquire or create one or more instances based on loading configuration information of the model (320) into a volatile memory (e.g., volatile memory (312)). For example, to acquire one or more instances, the electronic device (301) may efficiently utilize the usage of the volatile memory (312) and the capacity of the non-volatile memory (311) by reusing at least a portion of the loaded configuration information (e.g., first configuration information (410), second configuration information (420)). Additionally, the electronic device (301) may acquire inference data by executing a first instance (e.g., first instance (701)) among one or more instances and performing inference on the input data. The electronic device (301) may execute a method of performing verification of inference data using a second instance (e.g., a second instance (702)) among one or more instances. The quality of performing inference on input data by executing the first instance (701) and the second instance (702) may be higher than the quality of performing inference on input data by executing only the second instance (702). For example, the higher quality may include less time consumed to obtain output data.
[0138] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.
[0139] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, an electronic device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0140] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0141] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0142] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101) of FIG. 1). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0143] According to one embodiment, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0144] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0145] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs.
[0146] As described above, an electronic device (e.g., electronic device (301)) may include a non-volatile memory (e.g., non-volatile memory (311)) comprising one or more storage media for storing instructions. The electronic device may include a volatile memory (e.g., volatile memory (312)) comprising one or more storage media. The electronic device may include at least one processor (e.g., at least one processor (300)) comprising processing circuitry. The at least one processor is communiquently coupled to the non-volatile memory and the volatile memory. The instructions may cause the electronic device to receive input data for utilizing the functions of a pre-trained model (e.g., model (320)) stored in the non-volatile memory when executed individually or collectively by the at least one processor. The above instructions may cause the electronic device to acquire an instance according to the loaded first composition information based on loading the first composition information of the pre-learned model into the volatile memory when executed individually or collectively by the at least one processor. The above instructions may cause the electronic device to load the second composition information of the pre-learned model, which is distinct from the first composition information of the pre-learned model, into the volatile memory independently of executing the acquired instance to perform inference on the input data when executed individually or collectively by the at least one processor.
[0147] According to one embodiment, the instructions may cause the electronic device to obtain a second instance distinguished from the first instance according to the first configuration information and the second configuration information, based on loading the second configuration information into the volatile memory when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to perform inference on the input data based on executing the second instance when executed individually or collectively by the at least one processor.
[0148] According to one embodiment, the instructions may cause the electronic device to obtain first inference data for the input data based on executing the first instance when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to obtain second inference data for the input data and the first inference data based on executing the second instance when executed individually or collectively by the at least one processor. The instructions may cause the electronic device to perform verification of the first inference data based on the first inference data and the second inference data when executed individually or collectively by the at least one processor.
[0149] According to one embodiment, the number of layers constituting the first instance may be less than the number of layers constituting the second instance.
[0150] According to one embodiment, the instructions may cause the electronic device to load the second configuration information into the volatile memory while executing the instance, when executed individually or collectively by the at least one processor.
[0151] According to one embodiment, the first configuration information may include first weight data of first layers in the pre-trained model. The second configuration information may include second weight data of second layers in the pre-trained model following the first layers in the pre-trained model.
[0152] According to one embodiment, the first configuration information may include first graph data of first layers in the pre-trained model. The second configuration information may include second graph data of second layers in the pre-trained model following the first layers in the pre-trained model.
[0153] According to one embodiment, the first configuration information may include first weight data of first layers within the pre-trained model. The first graph data may include operations and structures of the pre-trained model.
[0154] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor, operations to obtain the instance can be performed based on the first graph data and the first weight data.
[0155] A method performed in an electronic device (e.g., electronic device (301)) having a non-volatile memory (e.g., non-volatile memory (311)) and a volatile memory (e.g., volatile memory (312)) as described above may include the operation of receiving input data by the electronic device for utilizing the function of a pre-trained model (e.g., model (320)) stored in the non-volatile memory. The method may include the operation of acquiring an instance by the electronic device according to the loaded first composition information based on loading first composition information of the pre-trained model into the volatile memory. The method may include the operation of loading second composition information of the pre-trained model, which is distinct from the first composition information of the pre-trained model, into the volatile memory by the electronic device, independently of executing the acquired instance to perform inference on the input data.
[0156] According to one embodiment, the method may include an operation of obtaining a second instance distinguished from the first instance according to the first configuration information and the second configuration information, based on loading the second configuration information into the volatile memory. The method may include an operation of performing inference on the input data based on executing the second instance.
[0157] According to one embodiment, the method may include an operation of obtaining first inference data for the input data based on executing the first instance. The method may include an operation of obtaining second inference data for the input data and the first inference data based on executing the second instance. The method may include an operation of performing verification on the first inference data based on the first inference data and the second inference data.
[0158] According to one embodiment, the number of layers constituting the first instance may be less than the number of layers constituting the second instance.
[0159] According to one embodiment, the method may include the operation of loading the second configuration information into the volatile memory while executing the instance.
[0160] According to one embodiment, the first configuration information may include first weight data of first layers in the pre-trained model. The second configuration information may include second weight data of second layers in the pre-trained model following the first layers in the pre-trained model.
[0161] According to one embodiment, the first configuration information may include first graph data of first layers in the pre-trained model. The second configuration information may include second graph data of second layers in the pre-trained model following the first layers in the pre-trained model.
[0162] According to one embodiment, the first configuration information may include first weight data of first layers within the pre-trained model. The first graph data may include operations and structures of the pre-trained model.
[0163] According to one embodiment, the method may include operations for obtaining the instance based on the first graph data and the first weight data.
[0164] In one or more non-transient computer-readable storage media in which one or more programs are stored as described above, the one or more programs may include computer-executable instructions that cause the electronic device to perform operations when executed individually or collectively by at least one processor of an electronic device (e.g., electronic device (301)) comprising a non-volatile memory (e.g., non-volatile memory (311)) and a volatile memory (e.g., volatile memory (312)). The above operations may include receiving input data by the electronic device for utilizing the function of a pre-trained model (e.g., model (320)) stored in the non-volatile memory; acquiring an instance by the electronic device according to the loaded first composition information based on loading the first composition information of the pre-trained model into the volatile memory; and loading a second composition information of the pre-trained model, which is distinct from the first composition information of the pre-trained model, into the volatile memory by the electronic device, independently of executing the acquired instance to perform inference on the input data.
[0165] According to one embodiment, the operations may include, based on loading the second configuration information into the volatile memory, an operation of obtaining a second instance distinguished from the first instance according to the first configuration information and the second configuration information, and an operation of performing inference on the input data based on executing the second instance.
[0166] According to one embodiment, the operations may include an operation of obtaining first inference data for the input data based on executing the first instance, an operation of obtaining second inference data for the input data and the first inference data based on executing the second instance, and an operation of performing verification of the first inference data based on the first inference data and the second inference data.
[0167] According to one embodiment, the number of layers constituting the first instance may be less than the number of layers constituting the second instance.
[0168] According to one embodiment, the operations may include loading the second configuration information into the volatile memory while executing the instance.
[0169] According to one embodiment, the first configuration information may include first weight data of first layers in the pre-trained model. The second configuration information may include second weight data of second layers in the pre-trained model following the first layers in the pre-trained model.
[0170] According to one embodiment, the first configuration information may include first graph data of first layers in the pre-trained model. The second configuration information may include second graph data of second layers in the pre-trained model following the first layers in the pre-trained model.
[0171] According to one embodiment, the first configuration information may include first weight data of first layers within the pre-trained model. The first graph data may include operations and structures of the pre-trained model.
[0172] According to one embodiment, the one or more programs may include instructions that cause the head-wearing electronic device to perform operations to acquire the instance based on the first graph data and the first weight data when executed by the head-wearing electronic device.
[0173] Although the present disclosure has been illustrated and described with reference to various embodiments, it will be understood by those skilled in the art that various modifications in form or detail are possible without departing from the technical spirit and scope of the present disclosure as defined by the appended claims and their equivalents.
Claims
1. In an electronic device, Non-volatile memory comprising one or more storage media for storing instructions; Volatile memory comprising one or more storage media; and The apparatus comprises at least one processor including a processing circuit, wherein the at least one processor is communicationally coupled to the non-volatile memory and the volatile memory. When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Receiving input data to utilize the functions of a pre-trained model stored in the above non-volatile memory; Based on loading the first composition information of the above-mentioned pre-trained model into the volatile memory, an instance is obtained according to the loaded first composition information; and Independently of executing the above-mentioned acquired instance to perform inference on the above-mentioned input data, causing to load the second configuration information of the above-mentioned pre-trained model, which is distinct from the first configuration information of the above-mentioned pre-trained model, into the volatile memory, Electronic device.
2. In Claim 1, The above instance is the first instance, and When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on loading the second configuration information into the volatile memory, a second instance distinguished from the first instance is obtained according to the first configuration information and the second configuration information; and Causing to perform inference on the input data based on executing the second instance above, Electronic device.
3. In Claim 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on executing the above first instance, first inference data for the input data is obtained; Based on executing the second instance above, obtain second inference data for the input data and the first inference data; and Causing to perform verification of the first inference data based on the first inference data and the second inference data, Electronic device.
4. In Claim 2, The number of layers constituting the first instance is less than the number of layers constituting the second instance. Electronic device.
5. In Claim 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Causing the second configuration information to be loaded into the volatile memory while executing the above instance, Electronic device.
6. In Claim 1, The first configuration information above includes first weight data of first layers within the pre-trained model, and The second configuration information includes second weight data of the second layers in the pre-trained model following the first layers in the pre-trained model, Electronic device.
7. In Claim 1, The first configuration information above includes first graph data of first layers within the pre-trained model, and The second configuration information includes second graph data of second layers in the pre-trained model following the first layers in the pre-trained model, Electronic device.
8. A method performed by an electronic device having non-volatile memory and volatile memory, The operation of receiving input data by the electronic device for utilizing the function of a pre-trained model stored in the non-volatile memory; An operation of acquiring an instance by the electronic device according to the loaded first composition information based on loading the first composition information of the above-mentioned pre-trained model into the volatile memory; and Independently of executing the above-mentioned acquired instance to perform inference on the above-mentioned input data, the operation of loading the second configuration information of the above-mentioned pre-trained model, which is distinct from the first configuration information of the above-mentioned pre-trained model, into the volatile memory by the electronic device. method.
9. In Claim 8, The above instance is the first instance, and An operation to obtain a second instance distinguished from the first instance according to the first configuration information and the second configuration information, based on loading the second configuration information into the volatile memory; and Based on executing the second instance above, further comprising an operation of performing inference on the input data, method.
10. In Claim 9, An operation to obtain first inference data for the input data based on executing the first instance above; Based on executing the second instance above, an operation to obtain second inference data for the input data and the first inference data; and A method further comprising an operation of performing verification of the first inference data based on the first inference data and the second inference data. method.
11. In Claim 9, The number of layers constituting the first instance is less than the number of layers constituting the second instance. method.
12. In claim 8, While executing the above instance, the operation of loading the second configuration information into the volatile memory is included. method.
13. In Claim 8, The first configuration information above includes first weight data of first layers within the pre-trained model, and The second configuration information includes second weight data of the second layers in the pre-trained model following the first layers in the pre-trained model, method.
14. In Claim 8, The first configuration information above includes first graph data of first layers within the pre-trained model, and The second configuration information includes second graph data of second layers in the pre-trained model following the first layers in the pre-trained model, method.
15. In one or more non-transient computer-readable storage media storing one or more programs including computer-executable instructions, The above instructions cause the electronic device to perform operations when executed individually or collectively by at least one processor of an electronic device including non-volatile memory and volatile memory, and The above operations are: The operation of receiving input data by the electronic device for utilizing the function of a pre-trained model stored in the non-volatile memory; An operation of acquiring an instance by the electronic device according to the loaded first composition information based on loading the first composition information of the above-mentioned pre-trained model into the volatile memory; and Independently of executing the above-mentioned acquired instance to perform inference on the above-mentioned input data, the operation of loading the second configuration information of the above-mentioned pre-trained model, which is distinct from the first configuration information of the above-mentioned pre-trained model, into the volatile memory by the electronic device. One or more non-transient computer-readable storage media.
Citation Information
Patent Citations
Water circulation mat and control method of same
KR1020220046353A
Charger Management System
KR1020220157717A
Cartridge and inhaler comprising the same
KR1020260042814A
Apparatus for exchanging nonfreezing a hydrant
KR102280536B1
Moving Robot and controlling method
KR102500529B1