Method for performing federated learning on basis of moe and lora, electronic device supporting same, and storage medium

The federated learning method using LoRA and MoE separates training of gating networks from expert models, addressing high communication costs and inefficiencies in large-scale language model training across distributed devices, ensuring efficient and privacy-preserving updates.

WO2026034906A1PCT designated stage Publication Date: 2026-02-12SAMSUNG ELECTRONICS CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/011497
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-20
Filing Date
2025-08-01
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing federated learning methods face challenges in managing large-scale language models due to high communication costs and inefficient parameter training across distributed devices, particularly when using approaches like GLaM and MoLE, which either increase communication costs or require extensive training of multiple expert models.

Method used

Implementing a federated learning method using low-rank adaptation (LoRA) to separate the training of gating networks from expert models, reducing communication costs by transmitting only a subset of parameters through a mixture-of-experts (MoE) framework, allowing parameter-efficient training on client devices.

Benefits of technology

This approach reduces communication costs and enhances training efficiency by optimizing parameter updates across distributed devices, ensuring data privacy while maintaining model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025011497_12022026_PF_FP_ABST
    Figure KR2025011497_12022026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device, according to one embodiment, comprises: a communication circuit; at least one processor including a processing circuit; and a memory storing instructions. The instructions, when executed individually or collectively by the at least one processor, instruct the electronic device to: identify a plurality of expert models corresponding to a plurality of external electronic devices in a large language model, wherein the plurality of external electronic devices are configured to perform federated learning, and the large language model includes a gating network and the plurality of expert models; and transmit the gating network and the expert models corresponding to the respective external electronic devices to the plurality of external electronic devices via the communication circuit.
Need to check novelty before this filing date? Find Prior Art

Description

Method for performing federated learning based on MOE and LORA, electronic devices supporting the same, and storage media

[0001] The present disclosure relates to a method for performing federated learning based on mixture-of-experts (MoE) and low-rank adaptation (LoRA), an electronic device supporting the same, and a storage medium.

[0002] Portable digital communication devices have become an essential part of modern life for many people. Consumers want to access a variety of high-quality services anytime, anywhere, using these devices.

[0003] Portable digital communication devices store and process a variety of personal information. With the proliferation of portable digital communication devices, the importance of protecting collected personal information has grown. Federated learning is a method for training AI (artificial intelligence) models using distributed data while protecting data privacy.

[0004] The above information may be provided as background information to aid in understanding the present disclosure. None of the above is claimed to be prior art related to the present disclosure, nor can it be used to determine prior art.

[0005] According to one aspect of the present disclosure, an electronic device includes: a communication circuit; at least one processor including a processing circuit; and instructions that, when individually or collectively executed by the at least one processor, cause the electronic device to: identify a plurality of expert models corresponding to a plurality of external electronic devices from a large-scale language model, wherein the plurality of external electronic devices are configured to perform federated learning, and the large-scale language model includes a gating network and the plurality of expert models; and transmit the expert models corresponding to the external electronic devices corresponding to the gating network to the plurality of external electronic devices via the communication circuit.

[0006] According to one aspect of the present disclosure, a method for an electronic device to perform federated learning includes: identifying a plurality of expert models corresponding to a plurality of external electronic devices from a large-scale language model, wherein the plurality of external electronic devices are configured to perform federated learning, and the large-scale language model includes a gating network and a plurality of expert models; and transmitting the expert models corresponding to the gating network and the corresponding external electronic devices to each of the plurality of external electronic devices.

[0007] According to one aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer-executable instructions includes instructions that, when individually or collectively executed by at least one processor, cause an electronic device to identify a plurality of expert models corresponding to a plurality of external electronic devices from a large-scale language model, and transmit the expert models corresponding to a gating network and the corresponding external electronic devices to each of the plurality of external electronic devices.

[0008] According to one aspect of the present disclosure, there is provided an electronic device comprising: a communication circuit; at least one processor including a processing circuit; and a memory storing instructions, the instructions, when individually or collectively executed by the at least one processor, cause an electronic device to identify a plurality of expert models corresponding to a plurality of external electronic devices configured to perform federated learning, the plurality of expert models being included in a large-scale language model (LLM), the LLM including a gating network pre-trained to identify at least one expert model corresponding to input data among the plurality of expert models; and transmitting, through the communication circuit, the gating network and each of the plurality of expert models to the plurality of external electronic devices for training each of the gating network and the plurality of expert models.

[0009] Additional aspects will be partly set forth in the following description, and partly will become apparent from the description or may be learned by practicing the embodiments provided.

[0010] The above and other aspects, features and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:

[0011] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments of the present disclosure.

[0012] FIG. 2 is a drawing for explaining an example of a configuration of an electronic device within a network environment according to one embodiment of the present disclosure.

[0013] FIG. 3 is a diagram for explaining a communication method between an electronic device and an external electronic device within a network environment according to one embodiment of the present disclosure.

[0014] FIG. 4 is a flowchart illustrating a method by which an electronic device according to one embodiment of the present disclosure transmits an expert model corresponding to a gating network and an external electronic device.

[0015] FIG. 5 is a diagram illustrating a method for reducing communication costs with an external electronic device by an electronic device according to one embodiment of the present disclosure.

[0016] FIGS. 6A, 6B, and 6C are diagrams illustrating a method for an electronic device according to an embodiment of the present disclosure to train some parameters of a large language model.

[0017] FIG. 7 is a flowchart illustrating a method by which an electronic device according to one embodiment of the present disclosure obtains updated gating networks and updated expert models.

[0018] FIGS. 8A and 8B are diagrams illustrating a method for an electronic device according to an embodiment of the present disclosure to obtain an updated gating network and updated expert models.

[0019] FIG. 9 is a flowchart illustrating a method by which an electronic device, according to one embodiment of the present disclosure, distributes an updated gating network and updated expert models to a plurality of external electronic devices.

[0020] Fig. 10 is a diagram for explaining a generative artificial intelligence model according to one embodiment.

[0021] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0022] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.

[0023] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0024] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0025] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0026] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0027] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0028] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0029] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0030] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0031] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0032] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0033] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0034] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0035] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0036] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0037] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).

[0038] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0039] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0040] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0041] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0042] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0043] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0044] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0045] FIG. 2 is a drawing for explaining an example of a configuration of an electronic device within a network environment according to one embodiment of the present disclosure.

[0046] Referring to FIG. 2, in one embodiment, an electronic device (201) (e.g., the electronic device (101) or server (108) of FIG. 1) may include a communication circuit (210), a memory (220), and a processor (230). The electronic device (201) may be implemented as a server for performing federated learning, but is not limited thereto.

[0047] In one embodiment, the communication circuit (210) may provide a function corresponding to the communication module (190) of FIG. 1. In one embodiment, the communication circuit (210) may communicate with an external electronic device via a network (e.g., the second network (199) of FIG. 1). The external electronic device may be, for example, a client device that trains an AI model (or some parameters of an AI model stored in the external electronic device) based on collecting personalized data. The AI ​​model may include a generative AI model trained to output a response corresponding to a user utterance based on receiving a user utterance or an input prompt corresponding to the user utterance. The AI ​​model may also include a transformer-based AI model. The generative AI model may include a large language model (LLM) trained to output text information, an image generation model trained to output image information, or an LLM / RAG (retrieval-augmented generation) model trained to generate output information based on a search database. Image generation models can be implemented as generative adversarial networks (GANs), variational autoencoders (VAEs), or diffusion models.

[0048] In one embodiment, the memory (220) may provide a function corresponding to the memory (130) of FIG. 1. In one embodiment, the memory (220) may store an AI model. In one embodiment, the AI ​​model stored in the memory (220) may correspond to an AI model stored in each of a plurality of external electronic devices (or client devices) participating in federated learning from the electronic device (201) (e.g., a server). The number of parameters of the AI ​​model stored in each of the plurality of external electronic devices may be less than the number of parameters of the AI ​​model stored in the memory (220).

[0049] In one embodiment, the processor (230) may provide functions corresponding to the processor (120) of FIG. 1.

[0050] In one embodiment, the number of processors (230) may be one or more. For example, the processor (230) may have a multi-core processor structure such as a dual core, quad core, or hexa core.

[0051] In one embodiment, the processor (230) may control operations of the electronic device (101) by executing instructions stored in the memory (220). For example, the processor (230) may correspond to multiple processors that collectively perform multiple operations by dividing them among the processors.

[0052] In one embodiment, the processor (230) may control the overall operation for updating (e.g., changing parameters or weights) an AI model (or some parameters of the AI ​​model) based on parameters received from each of a plurality of external electronic devices, and transmitting (or distributing) the updated AI model to the plurality of external electronic devices. In one embodiment, the processor (230) may include one or more processors for distributing the updated AI model to the plurality of external electronic devices. The operation performed by the processor (230) to obtain an updated AI model based on AI models trained by the plurality of external electronic devices based on the initial AI model will be described below.

[0053] Although the electronic device (201) in FIG. 2 is illustrated as including a communication circuit (210), a memory (220), and / or a processor (230), it is not limited thereto. For example, the electronic device (201) may further include a configuration that provides a function corresponding to at least one configuration illustrated in FIG. 1.

[0054] FIG. 3 is a diagram for explaining a communication method between an electronic device and an external electronic device within a network environment according to one embodiment of the present disclosure.

[0055] In one embodiment, the server (301) may perform federated learning based on communicating with a plurality of external electronic devices (341_1, 341_2, to 341_M). Federated learning may be a method in which a model (e.g., a generative AI model) is trained in an electronic device (or "client") that stores local data, and the server updates the model based on the model training results (or updated parameters). Updating the model may include an operation of changing at least some of the parameters included in the model. In federated learning, data privacy may be protected because local data is not transmitted to the server and is only used for model training of the client. The server (301) may include a large language model ("LLM") (310). The large language model (310) may include a gating network (320) and a plurality of expert models (330) based on a mixture-of-experts (MoE). The large language model (310) can perform learning using some parameters without activating a large number of parameters. The gating network (320) can be trained to identify (or regulate traffic to) at least one expert model to which input data will be input among a plurality of expert models (330) in response to input data (e.g., tokens). Each of the plurality of expert models (330) can include a smaller number of parameters than the plurality of parameters corresponding to the LLM (310).

[0056] In one embodiment, the server (301) may reduce the communication cost between the server (301) and the external electronic device (or client device) by transmitting an expert model corresponding to each of the plurality of external electronic devices (341_1, 341_2, to 341_M) among the plurality of expert models (330) to each of the plurality of external electronic devices (341_1, 341_2, to 341_M). For example, the server (301) may transmit a first expert model among the plurality of expert models (330) to the first external electronic device (341_1) (or the first client). The server (301) may transmit the gating network (320) together with the first expert model to the first external electronic device (341_1). The server (301) can transmit a second expert model from among a plurality of expert models (330) to a second external electronic device (341_2) (or a second client). The server (301) can transmit a gating network (320) together with the second expert model to the second external electronic device (341_2). The server (301) can transmit an Nth expert model from among a plurality of expert models (330) to an Mth external electronic device (341_M) (or an Mth client). The server (301) can transmit a gating network (320) together with the Nth expert model to the Mth external electronic device (341_M). In one embodiment, the number of the plurality of expert models (330) (e.g., N) may be different from the number of the plurality of external electronic devices (341_1, 341_2, to 341_M) (e.g., M). In one embodiment, the number of the plurality of expert models (330) (e.g., N) may be the same as the number of the plurality of external electronic devices (341_1, 341_2, to 341_M) (e.g., M). In one embodiment, a model (or network) transmitted between the server (301) and the plurality of external electronic devices (341_1, 341_2, to 341_M) may include a plurality of parameters.

[0057] In one embodiment, each of the plurality of external electronic devices (341_1, 341_2, to 341_M) may store an LLM (350). The LLM (350) may include a gating network (351) and a plurality of expert models (353). Each of the plurality of external electronic devices (341_1, 341_2, to 341_M) may perform local learning based on the expert model and the gating network (320) received from the server (301). Each of the plurality of external electronic devices (341_1, 341_2, to 341_M) may train the gating network (351) and an expert model received from among the plurality of expert models (353) using local data (e.g., “privacy data” or “user data”) acquired by the external electronic device. For example, a first external electronic device (341_1) can train a first expert model (353_1) among a gating network (351) and a plurality of expert models (353) stored in the first external electronic device (341_1) using local data acquired by the first external electronic device (341_1). A second external electronic device (341_2) can train a second expert model (353_2) among a gating network (351) and a plurality of expert models (353) stored in the second external electronic device (341_2) using local data acquired by the second external electronic device (341_2). The M external electronic device (341_M) can train the M expert model (353_M) among the gating network (351) and the plurality of expert models (353) stored in the M external electronic device (341_M) using local data acquired by the M external electronic device (341_M).

[0058] In one embodiment, each of the plurality of external electronic devices (341_1, 341_2, to 341_M) may transmit an expert model trained using local data and a gating network trained using local data to the server (301). For example, a first external electronic device (341_1) may transmit a first expert model trained using local data acquired by the first external electronic device (341_1) and a gating network trained using local data acquired by the first external electronic device (341_1) to the server (301). A second external electronic device (341_2) may transmit a second expert model trained using local data acquired by the second external electronic device (341_2) and a gating network trained using local data acquired by the second external electronic device (341_2) to the server (301). The M external electronic device (341_M) can transmit the N expert model trained using local data acquired by the M external electronic device (341_M) and the gating network trained using local data acquired by the M external electronic device (341_M) to the server (301).

[0059] In one embodiment, the server (301) can obtain a plurality of updated expert models and an updated gating network based on a plurality of expert models and a plurality of gating networks received from a plurality of external electronic devices (341_1, 341_2, to 341_M). The server (301) can perform parameter-efficient training by separating the training of the plurality of expert models from the training of the gating network. The server (301) can train an expert model corresponding to an expert model received from an external electronic device among a plurality of expert models (330) stored in the server (301), based on the plurality of expert models received from the plurality of external electronic devices (341_1, 341_2, to 341_M). The server (301) can train a first expert model corresponding to the expert model trained based on local data of the first external electronic device (341_1), for example, based on the trained expert model received from the first external electronic device (341_1). The server (301) can obtain an updated expert model based on training the first expert model. The server (301) can train a second expert model corresponding to the expert model trained based on local data of the second external electronic device (341_2), based on the trained expert model received from the second external electronic device (341_2). The server (301) can obtain an updated expert model based on training the second expert model. The server (301) can train an Nth expert model corresponding to the expert model trained based on local data of the Mth external electronic device (341_M) based on the trained expert model received from the Mth external electronic device (341_M). The server (301) can obtain an updated expert model based on training the Nth expert model.The server (301) can train a gating network (320) stored in the server (301) based on a plurality of gating networks received from a plurality of external electronic devices (341_1, 341_2, to 341_M). The server (301) can obtain an updated gating network based on a gating network trained based on local data acquired by a first external electronic device (341_1), a gating network trained based on local data acquired by a second external electronic device (341_2), and a gating network trained based on local data acquired by an M-th external electronic device (341_M). In one embodiment, the sequential operations of the server (301) distributing a gating network (320) and an expert model corresponding to an external electronic device among a plurality of expert models (330) to a plurality of external electronic devices (341_1, 341_2, to 341_M), receiving a plurality of trained gating networks and a plurality of trained expert models from the plurality of external electronic devices (341_1, 341_2, to 341_M), and obtaining an updated gating network and a plurality of updated expert models may be referred to as a "round." The server (301) may distribute the R-th updated gating network and the R-th updated expert models to the plurality of external electronic devices (341_1, 341_2, to 341_M) based on repeating a plurality of rounds (e.g., R rounds).

[0060] In one embodiment, approaches to reducing the computational cost of large language models (e.g., generalist language model (GLaM) or mixture of language experts (MoLE)) may be unsuitable for federated learning. For example, GLaM may increase communication costs between the server and the client due to its inclusion of trillions of parameters. MoLE may be unsuitable for federated learning, which requires training of a gating network and multiple expert models, due to its assumption that multiple expert models have been trained.

[0061] In one embodiment, the server (301) can perform parameter-efficient learning based on a federated LoRA ("FLoRA") method for federated learning, thereby reducing communication costs between the server and the client.

[0062] FIG. 4 is a flowchart illustrating a method for an electronic device according to an embodiment of the present disclosure to transmit an expert model corresponding to a gating network and an external electronic device. The embodiment of FIG. 4 will be described with reference to FIGS. 5, 6A, 6B, and 6C. FIG. 5 is a diagram illustrating a method for an electronic device according to an embodiment of the present disclosure to reduce communication costs with an external electronic device. FIGS. 6A, 6B, and 6C are diagrams illustrating a method for an electronic device according to an embodiment of the present disclosure to train some parameters of a large-scale language model.

[0063] In one embodiment, the operations illustrated in FIG. 4 may be performed in various orders, not limited to the order illustrated. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. In one embodiment, more operations may be performed than those illustrated in FIG. 4, or at least one operation may be performed less than those illustrated in FIG.

[0064] Referring to FIG. 4, in operation 401, in one embodiment, the electronic device (201) (e.g., the processor (230) and / or the server (301)) may respectively (respectively) check a plurality of expert models of a large language model (LLM) corresponding to a plurality of external electronic devices (e.g., the plurality of external electronic devices (341_1, 341_2, to 341_M)) performing federated learning. The large language model (e.g., the LLM (310)) may include a gating network (e.g., the gating network (320)) and a plurality of expert models (e.g., the plurality of expert models (330)). In one embodiment, the expert-client mapping may be performed in a random manner. Based on the federated learning of a plurality of rounds being performed, the expert-client mapping may be optimized based on a linear sum assignment (LSA) algorithm.

[0065] Referring to FIG. 5, in one embodiment, each of the plurality of expert models includes a plurality of parameters of a large language model ( ) among multiple parameters obtained based on LoRA (low-rank adaptation) ) may be included. LoRA may be a technique for fine-tuning a pre-trained language model (PLM). The parameters of the PLM may be represented by a weight matrix W. The number of rows and columns of the weight matrix W may each be d. The set of matrices R of d*d is It can be illustrated as follows. Based on the fact that the gradient updated in the process of fine-tuning the PLM is sparse, the electronic device (201) can perform fine-tuning using a small number of parameters included in a specific part (attention weights) of the LLM. The electronic device (201) can obtain parameter sets (505, 507) corresponding to low-rank matrices, for example, based on performing matrix decomposition (503) on the gradient among the entire parameters (501) of the LLM. For example, the matrix set of the parameter set A (505) decomposed by rank r is It can be illustrated as follows. The matrix set of parameter set B(507) decomposed by rank r is It can be illustrated as follows. The electronic device (201) (or client) can perform parameter-efficient learning based on learning the acquired parameter sets (505, 507).

[0066] Referring to FIG. 6A, in one embodiment, the large language model (600) may include a plurality of transformer blocks (601, 602, 603, 604, 605, 606). Each of the plurality of transformer blocks (601, 602, 603, 604, 605, 606) may include a layer normalization (611) (613) (“layer norm, “LN”), a feed forward layer (612), and a masked multi self-attention layer (614). The structure of the transformer blocks is not limited to the examples described above. In one embodiment, the parameters of the large language model (600) may be in a frozen state. The electronic device (201) can fine-tune the LLM (600) using a small amount of training data, based on a parameter set for fine-tuning obtained based on matrix decomposition, without substantially changing the structure of the LLM (600). In one embodiment, the LLM (600) can include a text input prediction model trained to predict text information to be sequentially input based on input text information. The text input prediction model can be trained on the client side based on training data (e.g., local data) that includes information related to the user's typing patterns, language usage patterns, and communication preferences.

[0067] Referring to FIG. 6B, in one embodiment, the gating network (620) can be updated by learning the parameters of the masked multi-self-attention layer (614) included in each of the transformer blocks (601, 602, 603, 604, 605, 606). Each of the plurality of transformer blocks (601, 602, 603, 604, 605, 606) can include the gating network (620) and a plurality of expert models. The update of the gating network (620) can be separated from the update of the plurality of expert models.

[0068] Referring to FIG. 6C, in one embodiment, a plurality of expert models (630) can be updated by learning parameters of a masked multi-self-attention layer (614) included in each of the transformer blocks (601, 602, 603, 604, 605, 606). A small number of parameters (e.g., parameters corresponding to a gating network and the plurality of expert models (630)) decomposed by LoRA of each of the plurality of transformer blocks (601, 602, 603, 604, 605, 606) can be updated based on parameter-efficient fine-tuning. The electronic device (201) can selectively learn an expert model (631) corresponding to an expert model trained by an external electronic device among the plurality of expert models included in the transformer block. Here, the expert model (631) may be an expert model acquired (641) based on LoRA (low-rank adaptation). An operation of updating an expert model (631) corresponding to a client (or, "external electronic device") among the expert models (631) may include an operation of updating (or changing) parameters acquired based on LoRA. Since the number of parameters acquired based on LoRA is much smaller than the number of parameters of PLM, parameter-efficient fine-tuning can be performed.

[0069] In operation 403, in one embodiment, the electronic device (201) may transmit, to each of a plurality of external electronic devices, a gating network and an expert model corresponding to the external electronic device via a communication circuit (e.g., communication circuit (210)). The external electronic device may be included in the plurality of external electronic devices. The expert model may be included in the plurality of expert models. The electronic device (201) may reduce the communication cost between the electronic device (201) and the external electronic device by transmitting the expert model corresponding to the external electronic device among the gating network and the plurality of expert models without transmitting all parameters of the PLM.

[0070] FIG. 7 is a flowchart illustrating a method for an electronic device according to an embodiment of the present disclosure to obtain an updated gating network and updated expert models. The embodiment of FIG. 7 will be described with reference to FIGS. 8A and 8B. FIGS. 8A and 8B are diagrams illustrating a method for an electronic device according to an embodiment of the present disclosure to obtain an updated gating network and updated expert models.

[0071] In one embodiment, the operations illustrated in FIG. 7 may be performed in various orders, not limited to the order illustrated. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. In one embodiment, more operations may be performed than those illustrated in FIG. 7, or at least one operation may be performed less than those illustrated in FIG.

[0072] Referring to FIG. 7, in operation 701, in one embodiment, the electronic device (201) (e.g., the processor (230) and / or the server (301)) may respectively (respectively) check a plurality of expert models corresponding to a plurality of external electronic devices (e.g., a plurality of external electronic devices (341_1, 341_2, to 341_M)). In one embodiment, operation 701 may be at least partially identical to or similar to operation 401, and any description overlapping with operation 401 may not be repeated herein.

[0073] In operation 703, in one embodiment, the electronic device (201) may transmit, to each of a plurality of external electronic devices, an expert model corresponding to the gating network and the external electronic device, via a communication circuit (e.g., communication circuit (210)). In one embodiment, operation 703 may be at least partially identical to or similar to operation 403, and any description overlapping with operation 403 may not be repeated herein.

[0074] In operation 705, in one embodiment, the electronic device (201) may receive, via the communication circuit, a plurality of first trained gating networks and a plurality of first trained expert models from a plurality of external electronic devices. Each of the plurality of first trained gating networks may include a plurality of parameters of the gating network and a plurality of parameters modified based on first training data of the external electronic device. The first training data of the external electronic device may include user data acquired by the external electronic device. Each of the plurality of first trained expert models may include a plurality of parameters of an expert model corresponding to the external electronic device and a plurality of parameters modified based on the first training data of the external electronic device.

[0075] In operation 707, in one embodiment, the electronic device (201) may obtain a first updated gating network by changing a plurality of parameters of the gating network based on a plurality of first trained gating networks.

[0076] Referring to FIG. 8A, in one embodiment, the first external electronic device (341_1) may obtain a trained gating network (811_1) and a trained expert model (813_1) based on training the gating network (351) and the first expert model (353_1) using local data acquired by the first external electronic device (341_1). The gating network may include a relatively small number of parameters based on LoRA. The learning of the gating network may be performed using data , gating network parameters , FreeTrain Giant Language Model It can be performed according to mathematical formula 1. Labels required for learning the gating network can be defined by expert-client mapping.

[0077]

[0078] The above mathematical formula 1 is merely an example to aid understanding, and embodiments of the present disclosure may not be limited thereto. For example, the above mathematical formula 1 may be modified, applied, or expanded in various ways.

[0079] A first external electronic device (341_1) can transmit a trained gating network (811_1) and a trained expert model (813_1) to a server (301). A second external electronic device (341_2) can obtain a trained gating network (811_2) and a trained expert model (813_2) based on training the gating network (351) and the second expert model (353_2) using local data acquired by the second external electronic device (341_2). The second external electronic device (341_2) can transmit the trained gating network (811_2) and the trained expert model (813_2) to the server (301). The M external electronic device (341_M) can obtain a trained gating network (811_N) and a trained expert model (813_N) based on training the gating network (351) and the N expert model (353_N) using local data acquired by the M external electronic device (341_M). The M external electronic device (341_M) can transmit the trained gating network (811_N) and the trained expert model (813_N) to the server (301). The server (301) can obtain an updated gating network (821) by changing (820) multiple parameters of the gating network based on multiple trained gating networks (811_1, 811_2, to 811_N) received from multiple external electronic devices (341_1, 341_2, to 341_M).

[0080] In operation 709, in one embodiment, the electronic device (201) may obtain a plurality of first updated expert models by changing a plurality of parameters of an expert model corresponding to each of the plurality of first trained expert models based on the plurality of first trained expert models.

[0081] Referring to FIG. 8B, in one embodiment, the server (301) may obtain a plurality of updated expert models (831_1, 831_2, to 831_N) by changing a plurality of parameters of an expert model corresponding to each of the plurality of trained expert models (813_1, 813_2, to 813_N) among the plurality of expert models (330) based on a plurality of trained expert models (813_1, 813_2, to 813_N) received from a plurality of external electronic devices (341_1, 341_2, to 341_M). For local data (x, y) of client M, expert parameters can be learned according to mathematical formula 2.

[0082]

[0083] The above mathematical equation (2) is merely an example to aid understanding, and embodiments of the present disclosure may not be limited thereto. For example, the above mathematical equation (1) may be modified, applied, or expanded in various ways.

[0084] FIG. 9 is a flowchart illustrating a method by which an electronic device, according to one embodiment of the present disclosure, distributes an updated gating network and updated expert models to a plurality of external electronic devices.

[0085] In one embodiment, the operations illustrated in FIG. 9 may be performed in various orders, not limited to the order illustrated. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. In one embodiment, more operations may be performed than those illustrated in FIG. 9, or at least one operation may be performed less than those illustrated in FIG.

[0086] Referring to FIG. 9, in operation 901, in one embodiment, the electronic device (201) (e.g., the processor (230) and / or the server (301)) may respectively (respectively) check a plurality of first updated expert models of a macro language model corresponding to a plurality of external electronic devices (e.g., the plurality of external electronic devices (341_1, 341_2, to 341_M)). The macro language model may include a first updated gating network and a plurality of first updated expert models. In one embodiment, the electronic device (201) may respectively check a plurality of first updated expert models corresponding to the plurality of external electronic devices based on a linear sum assignment (LSA) method.

[0087] In operation 903, in one embodiment, the electronic device (201) may transmit, to each of a plurality of external electronic devices, a first updated gating network and a first expert model corresponding to the external electronic device via a communication circuit (e.g., communication circuit (210)). The first updated expert model may be included in the plurality of first updated expert models. The external electronic device may be included in the plurality of external electronic devices.

[0088] In operation 905, in one embodiment, the electronic device (201) may receive, via the communication circuitry, a plurality of second trained gating networks and a plurality of second trained expert models from a plurality of external electronic devices. Each of the plurality of second trained gating networks may include a plurality of parameters of the first updated gating network and a plurality of parameters modified based on second training data of the external electronic device. The second training data of the external electronic device may include user data acquired by the external electronic device. Each of the plurality of second trained expert models may include a plurality of parameters of the first updated expert model corresponding to the external electronic device and a plurality of parameters modified based on the second training data of the external electronic device.

[0089] In operation 907, in one embodiment, the electronic device (201) may obtain a second updated gating network by changing a plurality of parameters of the first updated gating network based on a plurality of second trained gating networks.

[0090] In operation 909, in one embodiment, the electronic device (201) may obtain a plurality of second updated expert models by changing a plurality of parameters of a first updated expert model corresponding to each of the plurality of second trained expert models based on the plurality of second trained expert models.

[0091] In operation 911, in one embodiment, the electronic device (201) may transmit, via the communication circuit, to each of a plurality of external electronic devices, a second updated gating network and a second updated expert model corresponding to the external electronic device. The second updated expert model may be included in the plurality of second updated expert models.

[0092] In one embodiment, the electronic device (201) can reduce the cost of communicating with multiple external electronic devices by transmitting and / or receiving a significantly smaller number of parameters than a pre-trained language model (PLM), based on the MoE (mixture-of-experts) concept and the FLoRA (federated low-rank adaptation) method. Referring to Table 1, the number of communication parameters required by each client per round can be significantly smaller than the total number of parameters of the pre-trained language model.

[0093]

[0094] Table 1 illustrates that when federated learning is performed based on the FLoRA method, a relatively smaller number of parameters (e.g., less than 200,000) are transmitted between the server and the client compared to the number of parameters of a pre-trained language model (e.g., tens of millions of parameters). Referring to Table 1, the number of parameters learned within the client (e.g., hundreds of thousands to millions) may be much smaller than the number of parameters of a pre-trained language model (e.g., tens of millions of parameters). In the present disclosure, when federated learning is performed based on the FLoRA method, the communication cost between the server and the client and the learning cost of the client can be significantly reduced.

[0095] Fig. 10 is a diagram for explaining a generative artificial intelligence model according to one embodiment.

[0096] Referring to FIG. 10, a user query / response interface (1010), an application / service component (1030), a knowledge repository (1020), an AI framework (1040), and a generative AI model (1060) may be stored in a memory (e.g., a memory (130) of FIG. 1) or stored in a separate server. At least some of the user query / response interface (1010), the application / service component (1030), the knowledge repository (1020), the AI ​​framework (1040), or the generative AI model (1060) may be implemented in software or hardware.

[0097] According to one embodiment, a user query / response interface (1010) may receive a user's input. The user's input may be in the form of natural language, images, and / or videos, but is not limited thereto. Furthermore, context information may also be transmitted when the user's input is transmitted. The context information may include various additional information at the time of the user's input. For example, the additional information may include information about the application currently being used by the user or information about the user's location. Furthermore, the user's input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Furthermore, the user's input may also be in a non-natural language form, such as selecting a menu. The user query / response interface (1010) may output the results of the generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. The user query / response interface (1010) may output the results of the generative artificial intelligence system to the user. The output can be in natural language form, in the form of specific content, or in the form of an action requested by the user.

[0098] The AI ​​framework (1040) can receive user input and coordinate and control each component necessary to perform the user's intention based on the user's query.

[0099] User input received from the user query / response interface (1010) can be transmitted to a prompt design component (1041). The prompt design component (1041) can be used to generate prompts suitable for inputting the user input into a large language model (LLM), a large vision model (LVM), or a large multimodal model (LMM). The prompt design component (1041) can be an AI component that uses a machine learning algorithm or a neural network to develop better prompts over time. The prompt design component (1041) can access a knowledge component including user preference data, a prompt library, and prompt examples based on the user input to generate prompts, and can transmit the generated prompts to the large language model (LLM) or the large multimodal model (LMM).

[0100] The API / Plug-in management component (1042) can communicate with external information when there is a request for additional information when passing user input as input to a generative model. The API / Plug-in management component (1042) can establish a channel for communicating with the outside of the AI ​​Interface through the API, and can enable access to various data sources (e.g., knowledge repositories (1020)) through the established channel. In addition, the API / Plug-in management component (1042) can request the application / service component (1030) through the API for an action that ultimately performs the user input, rather than an intermediate result, when the action needs to be performed in the application or service. Information obtained from the outside can be used to generate a prompt in the prompt design component (1041) together with the user input, or can be passed as input to the generative model.

[0101] The output modification component (also called a refiner component) (1043) can fine-tune the output from the generative model. For example, the output modification component (1043) can verify that the content generated through the LLM and / or LMM is not irrelevant, does not contain biased content, or does not contain harmful content. In addition, the output modification component (1043) can determine to what extent the content matches the result desired by the user and, if necessary, can perform additional processing. The output modification component (1043) can additionally configure and provide the user with hints to avoid unwanted output.

[0102] A generative AI model (1060) may generally refer to an artificial intelligence neural network that creates new types of data based on user input information. The generative AI model (1060) may include an image-generating model and / or a language-generating model. Representative models for generating images include a generative adversarial network (GAN) and a variational auto encoder (VAE), and examples include a VAE and a Diffusion-based generative model that uses a Transformer structure. A language-generating model is a model trained to statistically output the most appropriate output based on input values, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. In addition, there are also LMMs (large multimodal models) that can recognize various types of data input, such as text, images, and audio, and generate new data corresponding to them.

[0103] According to one embodiment of the present disclosure, an electronic device (e.g., the electronic device (201) of FIG. 2) may include a communication circuit (e.g., the communication circuit (210) of FIG. 2), at least one processor including a processing circuit (e.g., the processor (230) of FIG. 2), and a memory (e.g., the memory (220) of FIG. 2) storing instructions. The instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to respectfully identify a plurality of expert models of a large language model corresponding to a plurality of external electronic devices performing federated learning. The large language model may include a gating network and a plurality of expert models. The instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to transmit, via the communication circuit (210), to each of the plurality of external electronic devices, an expert model corresponding to the gating network and the external electronic device. The external electronic device may be included in the plurality of external electronic devices. The expert model may be included in the plurality of expert models.

[0104] In one embodiment, the expert model may include a plurality of parameters that are fewer than the plurality of parameters corresponding to the large language model. The gating network may be trained to identify at least one expert model among the plurality of expert models to which the input data will be input in response to input data. The instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to receive a plurality of first trained gating networks and a plurality of first trained expert models from a plurality of external electronic devices via the communication circuit (210). The instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to obtain a first updated gating network by changing a plurality of parameters of the gating network based on the plurality of first trained gating networks. The instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to obtain a plurality of first updated expert models by changing a plurality of parameters of an expert model corresponding to each of the plurality of first trained expert models based on the plurality of first trained expert models.

[0105] In one embodiment, each of the plurality of first trained gating networks may include a plurality of parameters of the gating network and a plurality of parameters modified based on first training data of the external electronic device. Each of the plurality of first trained expert models may include a plurality of parameters of an expert model corresponding to the external electronic device and a plurality of parameters modified based on first training data of the external electronic device. The first training data of the external electronic device may include user data acquired by the external electronic device.

[0106] In one embodiment, the instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to identify each of the plurality of first updated expert models of the macro language model corresponding to the plurality of external electronic devices. The macro language model may include the first updated gating network and the plurality of first updated expert models. The instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to transmit, via the communication circuit (210), to each of the plurality of external electronic devices, the first updated gating network and the first updated expert model corresponding to the external electronic device. The first updated expert model may be included in the plurality of first updated expert models. The external electronic device may be included in the plurality of external electronic devices.

[0107] In one embodiment, the instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to identify each of the plurality of first updated expert models of the large language model corresponding to the plurality of external electronic devices based on a linear sum assignment (LSA) method.

[0108] In one embodiment, the instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to receive, via the communication circuit (210), a plurality of second trained gating networks and a plurality of second trained expert models from a plurality of external electronic devices. The instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to obtain a second updated gating network by modifying a plurality of parameters of the first updated gating network based on the plurality of second trained gating networks. The instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to obtain a plurality of second updated expert models by changing a plurality of parameters of a first updated expert model corresponding to each of the plurality of second trained expert models based on the plurality of second trained expert models. The instructions, when individually or collectively executed by the at least one processor (230), may cause the electronic device (201) to transmit, via the communication circuit (210), to each of the plurality of external electronic devices, the second updated gating network and the second updated expert model corresponding to the external electronic device. The second updated expert model may be included in the plurality of second updated expert models.

[0109] In one embodiment, each of the plurality of second trained gating networks may include a plurality of parameters of the first updated gating network and a plurality of parameters modified based on second training data of the external electronic device. Each of the plurality of second trained expert models may include a plurality of parameters of the first updated expert model corresponding to the external electronic device and a plurality of parameters modified based on second training data of the external electronic device. The second training data of the external electronic device may include user data acquired by the external electronic device.

[0110] In one embodiment, the large language model may include a plurality of transformer blocks. Each of the plurality of transformer blocks may include the gating network and the plurality of expert models. Each of the plurality of expert models may include a plurality of parameters obtained based on low-rank adaptation (LoRA) among the plurality of parameters of the large language model.

[0111] In one embodiment, the large-scale language model may include a text input prediction model trained to predict text information to be sequentially input based on input text information. The text input prediction model may be trained based on training data including information related to the user's typing patterns, language usage patterns, and communication preferences.

[0112] According to one embodiment of the present disclosure, a method may include an operation of verifying a plurality of expert models of a large language model corresponding to a plurality of external electronic devices performing federated learning, respectively. The large language model may include a gating network and a plurality of expert models. The method may include an operation of transmitting, to each of the plurality of external electronic devices, the gating network and the expert model corresponding to the external electronic device via a communication circuit of the electronic device. The external electronic device may be included in the plurality of external electronic devices. The expert model may be included in the plurality of expert models.

[0113] In one embodiment, the expert model may include a plurality of parameters that are fewer than the plurality of parameters corresponding to the large language model. The gating network may be trained to identify at least one expert model to which the input data will be input among the plurality of expert models in response to input data. The method may further include receiving, from a plurality of external electronic devices via the communication circuit (210), a plurality of first trained gating networks and a plurality of first trained expert models. The method may further include obtaining a first updated gating network by changing a plurality of parameters of the gating network based on the plurality of first trained gating networks. The method may further include obtaining a plurality of first updated expert models by changing a plurality of parameters of an expert model corresponding to each of the plurality of first trained expert models based on the plurality of first trained expert models.

[0114] In one embodiment, each of the plurality of first trained gating networks may include a plurality of parameters of the gating network and a plurality of parameters modified based on first training data of the external electronic device. Each of the plurality of first trained expert models may include a plurality of parameters of an expert model corresponding to the external electronic device and a plurality of parameters modified based on first training data of the external electronic device. The first training data of the external electronic device may include user data acquired by the external electronic device.

[0115] In one embodiment, the method may further include an operation of checking each of the plurality of first updated expert models of the large language model corresponding to the plurality of external electronic devices. The large language model may include the first updated gating network and the plurality of first updated expert models. The method may further include an operation of transmitting, to each of the plurality of external electronic devices, the first updated gating network and the first updated expert model corresponding to the external electronic device via the communication circuit (210). The first updated expert model may be included in the plurality of first updated expert models. The external electronic device may be included in the plurality of external electronic devices.

[0116] In one embodiment, the operation of checking each of the plurality of first updated expert models of the giant language model corresponding to the plurality of external electronic devices may include an operation of checking each of the plurality of first updated expert models of the giant language model corresponding to the plurality of external electronic devices based on a linear sum assignment (LSA) method.

[0117] In one embodiment, the method may further include receiving, via the communication circuit (210), a plurality of second trained gating networks and a plurality of second trained expert models from a plurality of external electronic devices. The method may further include obtaining a second updated gating network by changing a plurality of parameters of the first updated gating network based on the plurality of second trained gating networks. The method may further include obtaining a plurality of second updated expert models by changing a plurality of parameters of a first updated expert model corresponding to each of the plurality of second trained expert models based on the plurality of second trained expert models. The method may further include transmitting, via the communication circuit (210), to each of the plurality of external electronic devices, the second updated gating network and the second updated expert model corresponding to the external electronic device. The second updated expert model may be included in the plurality of second updated expert models.

[0118] In one embodiment, each of the plurality of second trained gating networks may include a plurality of parameters of the first updated gating network and a plurality of parameters modified based on second training data of the external electronic device. Each of the plurality of second trained expert models may include a plurality of parameters of the first updated expert model corresponding to the external electronic device and a plurality of parameters modified based on second training data of the external electronic device. The second training data of the external electronic device may include user data acquired by the external electronic device.

[0119] In one embodiment, the large language model may include a plurality of transformer blocks. Each of the plurality of transformer blocks may include the gating network and the plurality of expert models. Each of the plurality of expert models may include a plurality of parameters obtained based on low-rank adaptation (LoRA) among the plurality of parameters of the large language model.

[0120] In one embodiment, the large-scale language model may include a text input prediction model trained to predict text information to be sequentially input based on input text information. The text input prediction model may be trained based on training data including information related to the user's typing patterns, language usage patterns, and communication preferences.

[0121] According to one embodiment of the present disclosure, a non-transitory computer-readable storage medium having computer-executable instructions recorded thereon may cause an electronic device (201) to, when individually or collectively executed by at least one processor (230), identify a plurality of expert models of a large-scale language model corresponding to a plurality of external electronic devices performing federated learning, respectively. The large-scale language model may include a gating network and a plurality of expert models. The computer-executable instructions, when individually or collectively executed by at least one processor (230), may cause the electronic device (201) to transmit, via a communication circuit (210) of the electronic device (201), to each of the plurality of external electronic devices, the gating network and the expert model corresponding to the external electronic device. The external electronic device may be included in the plurality of external electronic devices. The expert model may be included in the plurality of expert models.

[0122] In one embodiment, the expert model may include a plurality of parameters that are fewer than the plurality of parameters corresponding to the large language model. The gating network may be trained to identify at least one expert model among the plurality of expert models in response to input data to which the input data will be input. The computer-executable instructions, when individually or collectively executed by at least one processor (230), may cause the electronic device (201) to receive, from a plurality of external electronic devices via the communication circuit (210), a plurality of first trained gating networks and a plurality of first trained expert models. The computer-executable instructions, when individually or collectively executed by at least one processor (230), may cause the electronic device (201) to obtain a first updated gating network by changing a plurality of parameters of the gating network based on the plurality of first trained gating networks. The computer-executable instructions, when executed individually or collectively by at least one processor (230), may cause the electronic device (201) to obtain a plurality of first updated expert models by changing a plurality of parameters of an expert model corresponding to each of the plurality of first trained expert models based on the plurality of first trained expert models.

[0123] An electronic device according to an embodiment of the present disclosure may take various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a home appliance. The electronic device according to an embodiment of the present disclosure is not limited to the aforementioned devices.

[0124] It should be understood that the embodiments of the present disclosure and the terminology used herein are not intended to limit the technical features described in the present disclosure to specific embodiments, but include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In the present disclosure, each of the phrases "A or B," "at least one of A and B," "at least one of A or B," "A, B, or C," "at least one of A, B, and C," and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among the phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0125] The term "module" used in one embodiment of the present disclosure may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0126] An embodiment of the present disclosure may be implemented as software (e.g., a program (440)) including one or more instructions stored in a storage medium (e.g., an internal memory (436) or an external memory (438)) readable by a machine (e.g., an electronic device (411)). For example, a processor (e.g., a processor (420)) of the machine (e.g., an electronic device (411)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0127] According to one embodiment, the method according to one embodiment disclosed in the present disclosure may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0128] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0129] Additionally, the structure of the data used in the embodiments of the present invention described above can be recorded on a computer-readable recording medium through various means. The computer-readable recording medium includes storage media such as magnetic storage media (e.g., ROM, floppy disk, hard disk, etc.) and optical reading media (e.g., CD-ROM, DVD, etc.).

[0130] The present invention has been described above, focusing on preferred embodiments thereof. Those skilled in the art will appreciate that the present invention can be implemented in modified forms without departing from its essential characteristics. Therefore, the disclosed embodiments should be considered illustrative rather than limiting. The scope of the present invention is set forth in the claims, not the foregoing description, and all differences within the scope equivalent thereto should be construed as being encompassed by the present invention.

Claims

1. In the electronic device (201), Communication circuit (210); At least one processor (230) comprising a processing circuit; and A memory (220) for storing instructions, said instructions, when individually or collectively executed by said at least one processor (230), cause said electronic device (201) to: Identifying a plurality of expert models corresponding to a plurality of external electronic devices configured to perform federated learning, wherein the plurality of expert models are included in a large-scale language model (LLM), and the LLM includes a pre-trained gating network to identify at least one expert model corresponding to input data among the plurality of expert models, An electronic device (201) comprising instructions for causing the gating network and each of the plurality of expert models to be transmitted to the plurality of external electronic devices through the communication circuit, such that each of the gating network and each of the plurality of expert models is trained.

2. In paragraph 1, The above expert model includes a plurality of parameters that are smaller than the plurality of parameters corresponding to the giant language model, The above gating network is trained to identify at least one expert model among the plurality of expert models to which input data will be input in response to input data, The above instructions, when individually or collectively executed by the at least one processor (230), cause the electronic device (201) to: Through the above communication circuit (210), a plurality of first trained gating networks and a plurality of first trained expert models are received from a plurality of external electronic devices, Based on the plurality of first trained gating networks, a first updated gating network is obtained by changing a plurality of parameters of the gating network, An electronic device (201) that causes a plurality of first updated expert models to be obtained by changing a plurality of parameters of an expert model corresponding to each of the plurality of first trained expert models based on the plurality of first trained expert models.

3. In paragraph 1 or 2, Each of the plurality of first trained gating networks includes a plurality of parameters of the gating network and a plurality of parameters changed based on first training data of the external electronic device, Each of the plurality of first trained expert models includes a plurality of parameters of an expert model corresponding to the external electronic device and a plurality of parameters changed based on first training data of the external electronic device, The first training data of the external electronic device includes user data obtained by the external electronic device, an electronic device (201).

4. In any one of paragraphs 1 to 3, The above instructions, when individually or collectively executed by the at least one processor (230), cause the electronic device (201) to: Identifying the plurality of first updated expert models of the giant language model corresponding to the plurality of external electronic devices, respectively, wherein the giant language model includes the first updated gating network and the plurality of first updated expert models, An electronic device (201) that causes the first updated gating network and the first updated expert model corresponding to the external electronic device to be transmitted to each of the plurality of external electronic devices through the communication circuit (210), wherein the first updated expert model is included in the plurality of first updated expert models, and the external electronic device is included in the plurality of external electronic devices.

5. In any one of paragraphs 1 to 4, The electronic device (201), wherein the instructions, when individually or collectively executed by the at least one processor (230), cause the electronic device (201) to respectively identify the plurality of first updated expert models of the giant language model corresponding to the plurality of external electronic devices based on a linear sum assignment (LSA) method.

6. In any one of paragraphs 1 to 5, The above instructions, when individually or collectively executed by the at least one processor (230), cause the electronic device (201) to: Through the above communication circuit (210), a plurality of second trained gating networks and a plurality of second trained expert models are received from a plurality of external electronic devices, Based on the plurality of second trained gating networks, a second updated gating network is obtained by changing a plurality of parameters of the first updated gating network, Based on the plurality of second trained expert models, a plurality of second updated expert models are obtained by changing a plurality of parameters of a first updated expert model corresponding to each of the plurality of second trained expert models, An electronic device (201) that causes the second updated gating network and the second updated expert model corresponding to the external electronic device to be transmitted to each of the plurality of external electronic devices through the communication circuit (210), wherein the second updated expert model is included in the plurality of second updated expert models.

7. In any one of paragraphs 1 to 6, Each of the plurality of second trained gating networks includes a plurality of parameters of the first updated gating network and a plurality of parameters changed based on second training data of the external electronic device, Each of the plurality of second trained expert models includes a plurality of parameters of a first updated expert model corresponding to the external electronic device and a plurality of parameters changed based on second training data of the external electronic device, The second training data of the external electronic device includes user data obtained by the external electronic device, an electronic device (201).

8. In any one of paragraphs 1 to 7, The above giant language model includes multiple transformer blocks, Each of the plurality of transformer blocks includes the gating network and the plurality of expert models, An electronic device (201), wherein each of the plurality of expert models includes a plurality of parameters obtained based on LoRA (low-rank adaptation) among the plurality of parameters of the large language model.

9. In any one of paragraphs 1 to 8, The above large language model includes a text input prediction model trained to predict text information to be sequentially input based on input text information, An electronic device (201) wherein the text input prediction model is trained based on training data including information related to the user's typing patterns, language usage patterns, and communication preferences.

10. In the method, An operation of identifying a plurality of expert models corresponding to a plurality of external electronic devices configured to perform federated learning, wherein the plurality of expert models are included in a large-scale language model (LLM), and the LLM includes a gating network pre-trained to identify at least one expert model corresponding to input data among the plurality of expert models, A method comprising the operation of transmitting the gating network and each of the plurality of expert models to the plurality of external electronic devices through the communication circuit so that each of the gating network and each of the plurality of expert models is trained.

11. In paragraph 10, The above expert model includes a plurality of parameters that are smaller than the plurality of parameters corresponding to the giant language model, The above gating network is trained to identify at least one expert model among the plurality of expert models to which input data will be input in response to input data, An operation of receiving a plurality of first trained gating networks and a plurality of first trained expert models from a plurality of external electronic devices through the above communication circuit (210); An operation of obtaining a first updated gating network by changing a plurality of parameters of the gating network based on the plurality of first trained gating networks; and A method further comprising an operation of obtaining a plurality of first updated expert models by changing a plurality of parameters of an expert model corresponding to each of the plurality of first trained expert models based on the plurality of first trained expert models.

12. In paragraph 10 or 11, Each of the plurality of first trained gating networks includes a plurality of parameters of the gating network and a plurality of parameters changed based on first training data of the external electronic device, Each of the plurality of first trained expert models includes a plurality of parameters of an expert model corresponding to the external electronic device and a plurality of parameters changed based on first training data of the external electronic device, A method wherein the first training data of the external electronic device includes user data acquired by the external electronic device.

13. In any one of paragraphs 10 to 12, An operation of verifying each of the plurality of first updated expert models of the giant language model corresponding to the plurality of external electronic devices, wherein the giant language model includes the first updated gating network and the plurality of first updated expert models; A method further comprising: transmitting, through the communication circuit (210), to each of the plurality of external electronic devices, the first updated gating network and the first updated expert model corresponding to the external electronic device, wherein the first updated expert model is included in the plurality of first updated expert models, and the external electronic device is included in the plurality of external electronic devices.

14. In any one of paragraphs 10 to 13, A method wherein the operation of checking each of the plurality of first updated expert models of the giant language model corresponding to the plurality of external electronic devices includes an operation of checking each of the plurality of first updated expert models of the giant language model corresponding to the plurality of external electronic devices based on a linear sum assignment (LSA) method.

15. In a non-transitory computer-readable storage medium having recorded thereon computer-executable instructions, the computer-executable instructions, when individually or collectively executed by at least one processor (230), cause an electronic device (201) to: Identifying a plurality of expert models corresponding to a plurality of external electronic devices configured to perform federated learning, wherein the plurality of expert models are included in a large-scale language model (LLM), and the LLM includes a pre-trained gating network to identify at least one expert model corresponding to input data among the plurality of expert models, A non-transitory computer-readable storage medium comprising instructions for causing the gating network and each of the plurality of expert models to be transmitted to the plurality of external electronic devices through the communication circuit, such that each of the gating network and each of the plurality of expert models is trained.

Citation Information

Patent Citations

  • Big data transaction and quality evaluation method and system taking big language model as medium

    CN117332247A

  • A federated pre-training method for large language models based on a trusted execution environment

    CN117648998B

  • A large model knowledge distillation low-rank adaptive federated learning method, electronic device and readable storage medium

    CN118070876B