Federated learning apparatus and method

By learning the local adapter of the client and using it to inform the server's global adapter, while distributing a second global adapter to the client, the system efficiently learns neural network models for both servers and clients, improving task performance and reducing data leakage risks.

WO2025095267A1PCT designated stage expired Publication Date: 2025-05-08LG ELECTRONICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/008751
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-02
Filing Date
2024-06-25
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing combined learning systems face challenges in efficiently learning neural network models for both servers and clients, leading to suboptimal task performance due to the dual-dual learning network structure and the risk of personal data leakage during centralized data processing.

Method used

The proposed solution involves learning the local adapter of the client, using this data to learn a global adapter on the server, and distributing a second global adapter to the client, thereby efficiently reflecting both common and individual characteristics in various environments.

Benefits of technology

This approach enhances task performance by distributing the second global adapter and learning resources to the client, maintaining alignment between the global neural network models of both the client and server, and reducing the risk of personal data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024008751_08052025_PF_FP_ABST
    Figure KR2024008751_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a federated learning apparatus and method capable of efficiently training a neural network model of a client and a server on the basis of a local adapter, and the apparatus may comprise: at least one client for generating a multi-modal feature on the basis of a first neural network model including a first global adapter and a local adapter; and a server for generating a task action corresponding to the multi-modal feature on the basis of a second neural network model including a second global adapter, and providing the task action to the client, wherein the server collects a learned local adapter from the client, retrains the second global adapter of the second neural network model on the basis of the collected local adapter, and distributes the retrained second global adapter to the client so as to update the first global adapter of the client with the retrained second global adapter.
Need to check novelty before this filing date? Find Prior Art

Description

Federated learning device and method

[0001] The present disclosure relates to a federated learning device and method capable of efficiently training neural network models of a server and a client based on a local adapter.

[0002] In general, artificial intelligence is a field of computer engineering and information technology that studies ways to enable computers to think, learn, and develop themselves in ways that human intelligence can do. It means enabling computers to imitate human intelligent behavior.

[0003] Furthermore, artificial intelligence does not exist in isolation; rather, it is closely related, both directly and indirectly, to other fields of computer science. In particular, in modern times, there are active attempts to introduce AI elements into various fields of information technology and utilize them to solve problems in those fields.

[0004] Training AI models requires massive computer resources to perform large-scale calculations. Cloud computing services are the best solution for easily providing computing infrastructure to train AI models without the need for complex hardware and software installation.

[0005] Cloud computing relies on the centralization of resources, requiring all necessary data to be stored in cloud memory and utilized for model training. While data centralization offers numerous benefits in terms of maximizing efficiency, it also carries the risk of personal data leakage, a risk that is becoming increasingly critical for business as data transfer increases.

[0006] Recently, to overcome these problems, many learning algorithms have been introduced to support federated learning systems and federated learning architectures.

[0007] Federated learning is a learning method that centrally collects models learned from users' personal data on the client, rather than centrally collecting users' personal data on the server and learning from there.

[0008] However, since this federated learning has a dual model structure while the server and client have deep learning networks with the same structure, problems may occur in the overall task performance as the downstream task input of the server changes when the model is retrained on the client.

[0009] Therefore, in the future, it is necessary to develop a federated learning device that can efficiently train neural network models of servers and clients to provide various services.

[0010] The present disclosure aims to solve the above-mentioned problems and other problems.

[0011] The present disclosure provides a federated learning device and method capable of efficiently training neural network models of a server and a client to reflect both common characteristics and individual unique characteristics in various real environments by training a local adapter of a client, retraining a global adapter of a server based on the trained local data, and distributing the retrained second global adapter to the client.

[0012] A federated learning device according to one embodiment of the present disclosure includes at least one client that generates a multi-modal feature based on a first neural network model including a first global adapter and a local adapter, and a server that generates a task action corresponding to the multi-modal feature based on a second neural network model including a second global adapter and provides the task action to the client, wherein the server collects local adapters learned from the client, retrains a second global adapter of a second neural network model based on the collected local adapters, and distributes the retrained second global adapter to the client so as to update the first global adapter of the client with the retrained second global adapter.

[0013] A federated learning method according to one embodiment of the present disclosure is a federated learning method of a federated learning device including at least one client that generates multi-modal features based on a first neural network model including a first global adapter and a local adapter, and a server that generates a task action corresponding to the multi-modal features based on a second neural network model including a second global adapter and provides the task action to the client, wherein the method may include a step of the client retraining a local adapter of the first neural network model based on at least one of local data and global data, a step of the client transmitting the retrained local adapter to the server, a step of the server collecting the retrained local adapter from the client, a step of the server retraining a second global adapter of the second neural network model based on the local adapter collected by the server, a step of the server distributing the retrained second global adapter to the client, a step of the client receiving the retrained second global adapter from the server, and a step of the client updating a first global adapter of the first neural network model based on the retrained second global adapter.

[0014] According to one embodiment of the present disclosure, a federated learning device can efficiently train neural network models of a server and a client to reflect both common characteristics and individual unique characteristics in various real environments by training a local adapter of a client, retraining a global adapter of a server based on the trained local data, and distributing the retrained second global adapter to the client.

[0015] In addition, the present disclosure can improve task performance by training each client's neural network model to be aligned with the server's global neural network model by having the server distribute a retrained second global adapter and retraining resources to the clients.

[0016] FIG. 1 illustrates an artificial intelligence device according to one embodiment of the present disclosure.

[0017] FIG. 2 illustrates an artificial intelligence server according to one embodiment of the present disclosure.

[0018] FIG. 3 illustrates an artificial intelligence system according to one embodiment of the present disclosure.

[0019] FIG. 4 is a diagram for explaining the operation of a federated learning device according to one embodiment of the present disclosure.

[0020] FIG. 5 is a diagram illustrating a neural network model of a client of a federated learning device according to an embodiment of the present disclosure.

[0021] FIG. 6 is a diagram for explaining a neural network model of a server of a federated learning device according to an embodiment of the present disclosure.

[0022] FIG. 7 is a diagram for explaining the client operation of a federated learning device according to one embodiment of the present disclosure.

[0023] FIG. 8 is a diagram for explaining the server operation of a federated learning device according to one embodiment of the present disclosure.

[0024] FIG. 9 is a drawing for explaining the overall structure of a federated learning device according to one embodiment of the present disclosure.

[0025] FIG. 10 is a diagram illustrating a global update process of a federated learning device according to an embodiment of the present disclosure.

[0026] FIG. 11 is a diagram for explaining a local update process of a federated learning device according to an embodiment of the present disclosure.

[0027] FIG. 12 is a diagram for explaining a multi-modal based service provision process of a federated learning device according to an embodiment of the present disclosure.

[0028] FIGS. 13 to 15 are diagrams for explaining an example of a multi-modal-based service application of a federated learning device according to an embodiment of the present disclosure.

[0029] FIG. 16 and FIG. 17 are drawings for explaining a server product and a client product of a federated learning device according to one embodiment of the present disclosure.

[0030] FIGS. 18 to 20 are diagrams for explaining a federated learning method of a federated learning device according to an embodiment of the present disclosure.

[0031] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be given the same reference numbers and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably only for the convenience of writing the specification, and do not in themselves have distinct meanings or roles. In addition, when describing the embodiments disclosed in this specification, if it is determined that a specific description of a related known technology may obscure the gist of the embodiments disclosed in this specification, a detailed description thereof will be omitted. In addition, the attached drawings are only intended to facilitate easy understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, and substitutes included in the spirit and technical scope of the present disclosure.

[0032] Terms that include ordinal numbers, such as first, second, etc., may be used to describe various components, but the components are not limited by these terms. These terms are used solely to distinguish one component from another.

[0033] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0034] The neural network, artificial neural network, and network function of the present disclosure may often be used interchangeably.

[0035] Additionally, in the present disclosure, the terms "neural network," "neural network," and "network function" may be used interchangeably. A neural network may be comprised of a set of interconnected computational units, generally referred to as "nodes." These "nodes" may also be referred to as "neurons." A neural network comprises at least two or more nodes. The nodes (or neurons) constituting the neural networks may be interconnected by one or more "links."

[0036] Artificial Intelligence (AI)

[0037] Artificial intelligence (AI) is the study of artificial intelligence or the methodologies for creating it, while machine learning (ML) defines various problems in the field of AI and studies the methodologies for solving them. Machine learning is also defined as an algorithm that improves performance on a task through consistent experience.

[0038] An artificial neural network (ANN) is a model used in machine learning. It can refer to a model with problem-solving capabilities, comprised of artificial neurons (nodes) formed by the connection of synapses. An ANN can be defined by the connection patterns between neurons in different layers, the learning process that updates model parameters, and the activation function that generates output values.

[0039] An artificial neural network may include an input layer, an output layer, and optionally one or more hidden layers. Each layer may contain one or more neurons, and the artificial neural network may include synapses connecting neurons. In an artificial neural network, each neuron can output a function value of an activation function based on input signals, weights, and biases received through the synapses.

[0040] Model parameters are parameters determined through learning, including synaptic connection weights and neuron biases. Hyperparameters are parameters that must be set before learning in machine learning algorithms, including the learning rate, number of iterations, mini-batch size, and initialization function.

[0041] The goal of artificial neural network training can be seen as determining model parameters that minimize a loss function. The loss function can be used as an indicator for determining optimal model parameters during the artificial neural network training process.

[0042] Machine learning can be classified into supervised learning, unsupervised learning, and reinforcement learning depending on the learning method.

[0043] Supervised learning refers to a method for training an artificial neural network given labels for training data. Labels can refer to the correct answer (or output) that the artificial neural network must infer when the training data is input to the artificial neural network. Unsupervised learning refers to a method for training an artificial neural network without given labels for the training data. Reinforcement learning refers to a learning method in which an agent defined within a given environment is trained to select actions or action sequences that maximize cumulative rewards in each state.

[0044] Among artificial neural networks, machine learning implemented with a deep neural network (DNN) containing multiple hidden layers is sometimes called deep learning, and deep learning is a subset of machine learning. Hereinafter, "machine learning" is used to encompass deep learning.

[0045] Robot

[0046] A robot can be defined as a machine that automatically performs or operates a given task based on its own capabilities. Specifically, a robot capable of perceiving its environment, making independent judgments, and performing actions can be called an intelligent robot.

[0047] Robots can be classified into industrial, medical, household, and military types depending on their purpose or field of use.

[0048] Robots are equipped with actuators or motors, enabling them to perform various physical actions, such as moving robot joints. Furthermore, mobile robots include wheels, brakes, propellers, and other actuators within their actuators, enabling them to move on the ground or fly in the air.

[0049] Self-Driving

[0050] Autonomous driving refers to the technology of driving on its own, and an autonomous vehicle refers to a vehicle that drives without user intervention or with minimal user intervention.

[0051] For example, autonomous driving can include technologies that maintain the driving lane, technologies that automatically adjust speed such as adaptive cruise control, technologies that automatically drive along a set route, and technologies that automatically set a route and drive when a destination is set.

[0052] Vehicles include vehicles equipped only with internal combustion engines, hybrid vehicles equipped with both internal combustion engines and electric motors, and electric vehicles equipped only with electric motors, and may include not only automobiles but also trains, motorcycles, etc.

[0053] At this time, autonomous vehicles can be viewed as robots with autonomous driving functions.

[0054] Extended Reality (XR)

[0055] Extended reality is a general term for virtual reality (VR), augmented reality (AR), and mixed reality (MR). VR technology presents real-world objects and backgrounds as CG images only, AR technology presents virtual CG images over images of real objects, and MR technology is a computer graphics technology that blends and combines virtual objects with the real world.

[0056] MR technology is similar to AR in that it presents both real and virtual objects simultaneously. However, while AR uses virtual objects to complement real objects, MR uses virtual and real objects on an equal footing.

[0057] XR technology can be applied to HMD (Head-Mount Display), HUD (Head-Up Display), mobile phones, tablet PCs, laptops, desktops, TVs, digital signage, etc., and devices to which XR technology is applied can be called XR devices.

[0058] Figure 1 illustrates an AI device (100) according to one embodiment of the present disclosure.

[0059] The AI ​​device (100) can be implemented as a fixed device or a movable device, such as a TV, a projector, a mobile phone, a smart phone, a desktop computer, a laptop, a digital broadcasting terminal, a PDA (personal digital assistant), a PMP (portable multimedia player), a navigation device, a tablet PC, a wearable device, a set-top box (STB), a DMB receiver, a radio, a washing machine, a refrigerator, a desktop computer, digital signage, a robot, a vehicle, etc.

[0060] Referring to FIG. 1, the AI ​​device (100) may include a communication unit (110), an input unit (120), a learning processor (130), a sensing unit (140), an output unit (150), a memory (170), and a processor (180).

[0061] The communication unit (110) can transmit and receive data with external devices such as other AI devices (100a to 100e) or AI servers (200) using wired or wireless communication technology. For example, the communication unit (110) can transmit and receive sensor information, user input, learning models, control signals, etc. with external devices.

[0062] At this time, the communication technologies used by the communication unit (110) include GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), LTE (Long Term Evolution), 5G, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Bluetooth (Bluetooth), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), ZigBee, NFC (Near Field Communication), etc.

[0063] The input unit (120) can obtain various types of data.

[0064] At this time, the input unit (120) may include a camera for inputting a video signal, a microphone for receiving an audio signal, a user input unit for receiving information from a user, etc. Here, the camera or microphone may be treated as a sensor, and a signal obtained from the camera or microphone may be referred to as sensing data or sensor information.

[0065] The input unit (120) can obtain input data to be used when obtaining output using learning data and learning models for model learning. The input unit (120) can also obtain unprocessed input data, in which case the processor (180) or learning processor (130) can extract input features as preprocessing for the input data.

[0066] The learning processor (130) can train a model composed of an artificial neural network using learning data. Here, the trained artificial neural network may be referred to as a learning model. The learning model can be used to infer result values ​​for new input data other than the learning data, and the inferred values ​​can be used as a basis for making decisions regarding certain actions.

[0067] At this time, the running processor (130) can perform AI processing together with the running processor (240) of the AI ​​server (200) of FIG. 2.

[0068] At this time, the running processor (130) may include a memory integrated or implemented in the AI ​​device (100). Alternatively, the running processor (130) may be implemented using a memory (170), an external memory directly coupled to the AI ​​device (100), or a memory maintained in an external device.

[0069] The sensing unit (140) can obtain at least one of internal information of the AI ​​device (100), information about the surrounding environment of the AI ​​device (100), and user information using various sensors.

[0070] At this time, the sensors included in the sensing unit (140) include a proximity sensor, a light sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, a lidar, a radar, etc.

[0071] The output unit (150) can generate output related to vision, hearing, or touch.

[0072] At this time, the output unit (150) may include a display unit that outputs visual information, a speaker that outputs auditory information, a haptic module that outputs tactile information, etc.

[0073] The memory (170) can store data that supports various functions of the AI ​​device (100). For example, the memory (170) can store input data, learning data, learning models, learning history, etc. obtained from the input unit (120).

[0074] The processor (180) may determine at least one executable operation of the AI ​​device (100) based on information determined or generated using a data analysis algorithm or a machine learning algorithm. Then, the processor (180) may control components of the AI ​​device (100) to perform the determined operation.

[0075] To this end, the processor (180) may request, search, receive, or utilize data from the running processor (130) or memory (170), and control components of the AI ​​device (100) to execute at least one of the executable operations, a predicted operation, or an operation determined to be desirable.

[0076] At this time, if connection of an external device is required to perform a determined operation, the processor (180) can generate a control signal for controlling the external device and transmit the generated control signal to the external device.

[0077] The processor (180) can obtain intent information for user input and determine the user's requirement based on the obtained intent information.

[0078] At this time, the processor (180) can obtain intent information corresponding to the user input by using at least one of an STT (Speech To Text) engine for converting voice input into a string or a natural language processing (NLP) engine for obtaining intent information of natural language.

[0079] At this time, at least one of the STT engine or the NLP engine may be configured with an artificial neural network, at least in part, trained according to a machine learning algorithm. Furthermore, at least one of the STT engine or the NLP engine may be trained by a learning processor (130), a learning processor (240) of an AI server (200), or a distributed processing thereof.

[0080] The processor (180) can collect history information including the operation details of the AI ​​device (100) or the user's feedback on the operation, and store the information in the memory (170) or the learning processor (130), or transmit the information to an external device such as an AI server (200). The collected history information can be used to update the learning model.

[0081] The processor (180) can control at least some of the components of the AI ​​device (100) to drive an application program stored in the memory (170). Furthermore, the processor (180) can operate two or more of the components included in the AI ​​device (100) in combination to drive the application program.

[0082] Figure 2 illustrates an AI server (200) according to one embodiment of the present disclosure.

[0083] Referring to FIG. 2, the AI ​​server (200) may refer to a device that trains an artificial neural network using a machine learning algorithm or utilizes a trained artificial neural network. Here, the AI ​​server (200) may be composed of multiple servers to perform distributed processing, and may be defined as a 5G network. In this case, the AI ​​server (200) may be included as part of the AI ​​device (100) and may perform at least a portion of the AI ​​processing.

[0084] The AI ​​server (200) may include a communication unit (210), memory (230), a learning processor (240), and a processor (260).

[0085] The communication unit (210) can transmit and receive data with an external device such as an AI device (100).

[0086] The memory (230) may include a model storage unit (231). The model storage unit (231) may store a model (or artificial neural network, 231a) that is being learned or has been learned through the learning processor (240).

[0087] The learning processor (240) can train an artificial neural network (231a) using learning data. The learning model can be used while mounted on the AI ​​server (200) of the artificial neural network, or can be mounted on an external device such as an AI device (100).

[0088] The learning model may be implemented in hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented in software, one or more instructions constituting the learning model may be stored in memory (230).

[0089] The processor (260) can infer a result value for new input data using a learning model and generate a response or control command based on the inferred result value.

[0090] Figure 3 shows an AI system (1) according to one embodiment of the present invention.

[0091] Referring to FIG. 3, an AI system (1) is connected to a cloud network (10) by at least one of an AI server (200), a robot (100a), an autonomous vehicle (100b), an XR device (100c), a smartphone (100d), or an appliance (100e). Here, a robot (100a), an autonomous vehicle (100b), an XR device (100c), a smartphone (100d), or an appliance (100e) to which AI technology is applied may be referred to as an AI device (100a to 100e).

[0092] A cloud network (10) may refer to a network that constitutes part of a cloud computing infrastructure or exists within a cloud computing infrastructure. Here, the cloud network (10) may be configured using a 3G network, a 4G or LTE (Long Term Evolution) network, a 5G network, etc.

[0093] That is, each device (100a to 100e, 200) constituting the AI ​​system (1) can be connected to each other via a cloud network (10). In particular, each device (100a to 100e, 200) can communicate with each other via a base station, but can also communicate with each other directly without going through a base station.

[0094] The AI ​​server (200) may include a server that performs AI processing and a server that performs operations on big data.

[0095] The AI ​​server (200) is connected to at least one of the AI ​​devices constituting the AI ​​system (1), such as a robot (100a), an autonomous vehicle (100b), an XR device (100c), a smartphone (100d), or a home appliance (100e), through a cloud network (10), and can assist at least part of the AI ​​processing of the connected AI devices (100a to 100e).

[0096] At this time, the AI ​​server (200) can train an artificial neural network according to a machine learning algorithm on behalf of the AI ​​devices (100a to 100e), and can directly store the learning model or transmit it to the AI ​​devices (100a to 100e).

[0097] At this time, the AI ​​server (200) can receive input data from the AI ​​devices (100a to 100e), infer a result value for the received input data using a learning model, and generate a response or control command based on the inferred result value and transmit it to the AI ​​devices (100a to 100e).

[0098] Alternatively, the AI ​​device (100a to 100e) may infer a result value for input data using a direct learning model and generate a response or control command based on the inferred result value.

[0099] Below, various embodiments of AI devices (100a to 100e) to which the above-described technology is applied are described. Here, the AI ​​devices (100a to 100e) illustrated in FIG. 3 can be viewed as specific embodiments of the AI ​​device (100) illustrated in FIG. 1.

[0100] <AI+로봇>

[0101] The robot (100a) can be implemented as a guide robot, transport robot, cleaning robot, wearable robot, entertainment robot, pet robot, unmanned flying robot, etc. by applying AI technology.

[0102] The robot (100a) may include a robot control module for controlling movement, and the robot control module may mean a software module or a chip that implements the same in hardware.

[0103] The robot (100a) can obtain status information of the robot (100a), detect (recognize) the surrounding environment and objects, generate map data, determine a movement path and driving plan, determine a response to user interaction, or determine an action using sensor information obtained from various types of sensors.

[0104] Here, the robot (100a) can use sensor information acquired from at least one sensor among lidar, radar, and camera to determine a movement path and driving plan.

[0105] The robot (100a) can perform the above-described operations using a learning model comprised of at least one artificial neural network. For example, the robot (100a) can recognize its surroundings and objects using the learning model, and determine operations using the recognized surrounding environment information or object information. Here, the learning model may be learned directly by the robot (100a) or by an external device such as an AI server (200).

[0106] At this time, the robot (100a) may perform an action by generating a result using a direct learning model, but may also perform an action by transmitting sensor information to an external device such as an AI server (200) and receiving the result generated accordingly.

[0107] The robot (100a) can determine a movement path and a driving plan using at least one of map data, object information detected from sensor information, or object information acquired from an external device, and control a driving unit to drive the robot (100a) according to the determined movement path and driving plan.

[0108] Map data may include object identification information for various objects positioned in the space where the robot (100a) moves. For example, map data may include object identification information for fixed objects such as walls and doors, as well as movable objects such as flower pots and desks. Furthermore, object identification information may include name, type, distance, location, etc.

[0109] Additionally, the robot (100a) can perform actions or drive by controlling the driving unit based on the user's control / interaction. At this time, the robot (100a) can acquire intention information regarding the interaction based on the user's actions or voice utterances, and determine a response based on the acquired intention information to perform the action.

[0110] <AI+자율주행>

[0111] An autonomous vehicle (100b) can be implemented as a mobile robot, vehicle, or unmanned aerial vehicle by applying AI technology.

[0112] The autonomous vehicle (100b) may include an autonomous driving control module for controlling autonomous driving functions. The autonomous driving control module may refer to a software module or a chip implementing the same as hardware. The autonomous driving control module may be included internally as a component of the autonomous vehicle (100b), but may also be configured as separate hardware and connected to the exterior of the autonomous vehicle (100b).

[0113] An autonomous vehicle (100b) can obtain status information of the autonomous vehicle (100b), detect (recognize) the surrounding environment and objects, generate map data, determine a movement path and driving plan, or determine an action using sensor information obtained from various types of sensors.

[0114] Here, the autonomous vehicle (100b) can use sensor information acquired from at least one sensor among lidar, radar, and camera, similar to the robot (100a), to determine a movement path and driving plan.

[0115] In particular, the autonomous vehicle (100b) can recognize the environment or objects in an area where the field of view is obstructed or an area beyond a certain distance by receiving sensor information from external devices, or can receive information recognized directly from external devices.

[0116] The autonomous vehicle (100b) can perform the above-described operations using a learning model comprised of at least one artificial neural network. For example, the autonomous vehicle (100b) can recognize its surroundings and objects using the learning model, and determine a driving route using the recognized surrounding environment information or object information. Here, the learning model may be learned directly by the autonomous vehicle (100b) or by an external device such as an AI server (200).

[0117] At this time, the autonomous vehicle (100b) may perform an action by generating a result using a direct learning model, but may also perform an action by transmitting sensor information to an external device such as an AI server (200) and receiving the result generated accordingly.

[0118] An autonomous vehicle (100b) can determine a movement path and a driving plan using at least one of map data, object information detected from sensor information, or object information acquired from an external device, and control a driving unit to drive the autonomous vehicle (100b) according to the determined movement path and driving plan.

[0119] Map data may include object identification information for various objects located in the space (e.g., a road) where the autonomous vehicle (100b) travels. For example, map data may include object identification information for fixed objects such as streetlights, rocks, and buildings, as well as movable objects such as vehicles and pedestrians. Furthermore, object identification information may include name, type, distance, location, and the like.

[0120] Additionally, the autonomous vehicle (100b) can perform actions or drive by controlling the driving unit based on the user's control / interaction. At this time, the autonomous vehicle (100b) can acquire intention information regarding the interaction based on the user's actions or voice utterances, and determine a response based on the acquired intention information to perform the action.

[0121] <AI+XR>

[0122] The XR device (100c) can be implemented as an HMD (Head-Mount Display), a HUD (Head-Up Display) installed in a vehicle, a television, a mobile phone, a smart phone, a computer, a wearable device, a home appliance, digital signage, a vehicle, a fixed robot, or a mobile robot by applying AI technology.

[0123] The XR device (100c) can obtain information about surrounding space or real objects by analyzing 3D point cloud data or image data acquired through various sensors or from an external device to generate location data and attribute data for 3D points, and can render and output an XR object to be output. For example, the XR device (100c) can output an XR object including additional information about a recognized object in correspondence with the recognized object.

[0124] The XR device (100c) can perform the above-described operations using a learning model composed of at least one artificial neural network. For example, the XR device (100c) can recognize a real-world object from 3D point cloud data or image data using the learning model, and provide information corresponding to the recognized real-world object. Here, the learning model may be learned directly in the XR device (100c) or learned from an external device such as an AI server (200).

[0125] At this time, the XR device (100c) may perform an operation by generating a result using a direct learning model, but may also perform an operation by transmitting sensor information to an external device such as an AI server (200) and receiving the result generated accordingly.

[0126] <AI+로봇+자율주행>

[0127] The robot (100a) can be implemented as a guide robot, transport robot, cleaning robot, wearable robot, entertainment robot, pet robot, unmanned flying robot, etc. by applying AI technology and autonomous driving technology.

[0128] A robot (100a) to which AI technology and autonomous driving technology are applied may refer to a robot itself with autonomous driving function, or a robot (100a) that interacts with an autonomous vehicle (100b).

[0129] A robot (100a) with an autonomous driving function can be a general term for devices that move on their own along a given path without user control or move by determining the path on their own.

[0130] A robot (100a) with autonomous driving capabilities and an autonomous vehicle (100b) may use a common sensing method to determine one or more of a movement path or a driving plan. For example, a robot (100a) with autonomous driving capabilities and an autonomous vehicle (100b) may use information sensed via lidar, radar, and cameras to determine one or more of a movement path or a driving plan.

[0131] A robot (100a) interacting with an autonomous vehicle (100b) may exist separately from the autonomous vehicle (100b), and may be linked to autonomous driving functions within the autonomous vehicle (100b) or perform actions linked to a user riding in the autonomous vehicle (100b).

[0132] At this time, the robot (100a) interacting with the autonomous vehicle (100b) can control or assist the autonomous driving function of the autonomous vehicle (100b) by acquiring sensor information on behalf of the autonomous vehicle (100b) and providing it to the autonomous vehicle (100b), or by acquiring sensor information and generating surrounding environment information or object information and providing it to the autonomous vehicle (100b).

[0133] Alternatively, a robot (100a) interacting with an autonomous vehicle (100b) may monitor a user riding in the autonomous vehicle (100b) or control functions of the autonomous vehicle (100b) through interaction with the user. For example, if the robot (100a) determines that the driver is drowsy, it may activate the autonomous driving function of the autonomous vehicle (100b) or assist in controlling the driving unit of the autonomous vehicle (100b). Here, the functions of the autonomous vehicle (100b) controlled by the robot (100a) may include not only the autonomous driving function, but also functions provided by a navigation system or audio system installed inside the autonomous vehicle (100b).

[0134] Alternatively, a robot (100a) interacting with an autonomous vehicle (100b) may provide information to the autonomous vehicle (100b) or assist functions from outside the autonomous vehicle (100b). For example, the robot (100a) may provide traffic information, including signal information, to the autonomous vehicle (100b), such as a smart traffic light, or may interact with the autonomous vehicle (100b) to automatically connect an electric charger to a charging port, such as an automatic electric charger for an electric vehicle.

[0135] <AI+로봇+XR>

[0136] The robot (100a) can be implemented as a guide robot, transport robot, cleaning robot, wearable robot, entertainment robot, pet robot, unmanned flying robot, drone, etc. by applying AI technology and XR technology.

[0137] A robot (100a) to which XR technology is applied may refer to a robot that is the subject of control / interaction within an XR image. In this case, the robot (100a) is distinct from the XR device (100c) and can be linked with each other.

[0138] When a robot (100a) that is the target of control / interaction within an XR image obtains sensor information from sensors including a camera, the robot (100a) or the XR device (100c) can generate an XR image based on the sensor information, and the XR device (100c) can output the generated XR image. In addition, the robot (100a) can operate based on a control signal input through the XR device (100c) or a user's interaction.

[0139] For example, a user can check an XR image corresponding to the viewpoint of a remotely connected robot (100a) through an external device such as an XR device (100c), and through interaction, adjust the autonomous driving path of the robot (100a), control the operation or driving, or check information on surrounding objects.

[0140] <AI+자율주행+XR>

[0141] Autonomous vehicles (100b) can be implemented as mobile robots, vehicles, unmanned aerial vehicles, etc. by applying AI technology and XR technology.

[0142] An autonomous vehicle (100b) to which XR technology is applied may refer to an autonomous vehicle equipped with a means for providing XR images, an autonomous vehicle that is the subject of control / interaction within an XR image, etc. In particular, an autonomous vehicle (100b) that is the subject of control / interaction within an XR image is distinct from an XR device (100c) and can be linked with each other.

[0143] An autonomous vehicle (100b) equipped with a means for providing XR images can acquire sensor information from sensors including cameras and output XR images generated based on the acquired sensor information. For example, the autonomous vehicle (100b) can be equipped with a HUD to output XR images, thereby providing passengers with XR objects corresponding to real objects or objects on the screen.

[0144] At this time, when the XR object is output to the HUD, at least a part of the XR object may be output so as to overlap with an actual object toward which the passenger's gaze is directed. On the other hand, when the XR object is output to a display provided inside the autonomous vehicle (100b), at least a part of the XR object may be output so as to overlap with an object on the screen. For example, the autonomous vehicle (100b) may output XR objects corresponding to objects such as a road, another vehicle, a traffic light, a traffic sign, a two-wheeled vehicle, a pedestrian, a building, etc.

[0145] When an autonomous vehicle (100b) that is the target of control / interaction within an XR image obtains sensor information from sensors including a camera, the autonomous vehicle (100b) or the XR device (100c) can generate an XR image based on the sensor information, and the XR device (100c) can output the generated XR image. In addition, the autonomous vehicle (100b) can operate based on a control signal input through an external device such as the XR device (100c) or a user's interaction.

[0146] FIG. 4 is a diagram for explaining the operation of a federated learning device according to one embodiment of the present disclosure.

[0147] As illustrated in FIG. 4, the federated learning device of the present disclosure may include a server (400) and one or more clients (500) that are communicatively connected to the server (400).

[0148] Here, the client (500) may be an individual device, an individual hub that is connected to multiple devices, or an individual cloud that is connected to multiple hubs, which can be adjusted according to the physical range.

[0149] The client (500) can generate multi-modal features based on a first neural network model (510) including a first global adapter (512) and a local adapter (514).

[0150] For example, the first neural network model (510) may include a pre-trained first foundation model, a first global adapter (512) that reflects common characteristics of the client, and a local adapter (514) that reflects unique characteristics of the client.

[0151] Here, the first global adapter (512) is updated with the second global adapter (410) distributed from the server (400), and the local adapter (514) can be retrained based on at least one of the local data of the client (500) and the global data of the server (400).

[0152] The client (500) can classify the result values ​​output from the pre-learned first foundation model into a first global adapter (512) reflecting common characteristics and a local adapter (514) reflecting unique characteristics, and generate multi-modal features based on the first global adapter (512) and the local adapter (514) and provide them to the server (400).

[0153] Here, the client (500) collects multi-modal data when classifying the first global adapter (512) or local adapter (514), inputs the collected multi-modal data into a pre-learned first foundation model to output an inferred result value, and can classify the result value as the first global adapter (512) or local adapter (514) based on the characteristic weights assigned to the result value.

[0154] For example, when classifying a result value, the client (500) may assign weights to the result value for global characteristics and local characteristics according to the characteristics of the result value, and if the weight assigned to the result value is greater for the local characteristic than for the global characteristic, the result value may be classified as a local adapter (514), and if the weight assigned to the result value is less for the local characteristic than for the global characteristic, the result value may be classified as a first global adapter (512).

[0155] In addition, when providing multi-modal features to the server (400), the client (500) can average the weights of the features corresponding to the first global adapter (512) and the weights of the features corresponding to the local adapter (514), fuse the features corresponding to the first global adapter (512) and the features corresponding to the local adapter (514), and generate a multi-modal feature including the fused features and provide the result to the server (400).

[0156] As an example, if the client (500) is unable to perform self-learning on the local adapter (514), the client (500) may store local data corresponding to the operation of the first neural network model (510) and transmit the local data to the server (400) if it determines that self-learning on the local adapter (514) is impossible.

[0157] Here, when transmitting local data, the client (500) can check whether the stored local data satisfies preset transmission conditions, and if it satisfies the preset transmission conditions, transmit the local data to the server (400).

[0158] For example, when checking whether a preset transmission condition is satisfied, the client (500) can check whether the preset transmission condition satisfies at least one of a condition for transmitting by a set period, a condition for transmitting by a set amount of data, a condition for transmitting by data having a set characteristic, and a condition for transmitting by a set communication status. This is only one example and is not limited thereto.

[0159] In addition, the client (500) can check whether the relearned second global adapter is received from the server (400), and if the relearned second global adapter is received, can update the local adapter (514) based on the second global adapter.

[0160] Additionally, the client (500) may request a second global adapter from the server (400) before operating the first neural network model (510), and upon receiving the second global adapter from the server (400), may initialize a local adapter (514) based on the second global adapter.

[0161] In another embodiment, if the client (500) is capable of self-learning for the local adapter (514), the client (500), if it determines that self-learning for the local adapter (514) is possible, determines whether re-learning for the local adapter (514) of the first neural network model (510) is necessary, and if it determines that re-learning for the local adapter (514) is necessary, re-learns the local adapter (514) based on at least one of the local data of the client (500) and the global data of the server (400), and transmits the re-learned local adapter (514) to the server (400).

[0162] Here, the client (500) can check whether the relearned second global adapter is received from the server (400), and if the relearned second global adapter is received, can update the first global adapter (512) based on the relearned second global adapter.

[0163] In addition, when the client (500) receives the relearned second global adapter and the relearning resource from the server (400) at the same time, the client (500) can classify the relearned second global adapter and the relearning resource, store the classified relearning resource, and update the first global adapter (512) based on the relearned second global adapter.

[0164] In some cases, the client (500) may sequentially perform a process of storing the relearning resources classified in the order of reception time when the relearned second global adapter and the relearning resources are received from the server (400) at different times, and a process of updating the first global adapter (512) based on the relearned second global adapter.

[0165] In another case, the client (500) may request a relearning resource from the server (400), and when the relearning resource is received from the server (400), the client (500) may store the relearning resource.

[0166] Here, the client (500) may request relearning resources from the server (400) whenever a second global adapter that has been relearned is received from the server (400), or may request relearning resources from the server (400) at preset intervals.

[0167] Additionally, the client (500) may request a second global adapter from the server (400) before operating the first neural network model (510), and upon receiving the second global adapter from the server (400), initialize a local adapter (514) based on the second global adapter.

[0168] Next, when the client (500) receives a task action from the server (400), it can execute a service corresponding to the task action and provide it to the user.

[0169] Meanwhile, the server (400) can be a hub that unites multiple devices when the client (500) is an individual device, a cloud that unites multiple hubs when the client (500) is an individual hub, or a large cloud that unites multiple clouds when the client (500) is an individual cloud, which can be adjusted according to the physical range.

[0170] The server (400) can generate a task action corresponding to a multi-modal feature based on a second neural network model (410) including a second global adapter and provide it to the client (500).

[0171] Here, the server (400) collects local adapters (514) learned from the client (500), retrains the second global adapter of the second neural network model (410) based on the collected local adapters (514), and distributes the retrained second global adapter to the client (500) so that the first global adapter (512) of the client (500) can be updated with the retrained second global adapter.

[0172] For example, the second neural network model (410) may include a pre-trained second foundation model and a second global adapter that reflects common characteristics of the federated clients (500).

[0173] Here, the second global adapter can be relearned and updated based on multiple local adapters (514) received from the federated clients (500).

[0174] Additionally, the server (400) can initialize a second neural network model (410) based on global data and distribute a second global adapter and learning resources to the federated clients (500) before acquiring multi-modal features from the federated clients (500).

[0175] Next, the server (400) collects local adapters and local data from the federated clients (500), classifies the collected local data and local data by client, checks whether local adapters and local data have been collected from all the federated clients, and when local adapters and local data have been collected from all the federated clients, the server can retrain the second global adapter based on the local adapters collected from the federated clients.

[0176] Here, when retraining the second global adapter, the server (400) checks whether a change in resources for retraining is necessary when the local adapters (514) of all the federated clients (500) are collected, and if a change in resources for retraining is necessary, the server updates the second neural network model (410) based on the local adapters and local data of the clients, and then retrains the second global adapter based on the collected local adapters (514).

[0177] At this time, if a change in the resource for relearning is unnecessary, the server (400) can relearn only the second global adapter of the second neural network model (410) based on the collected local adapter (514) and distribute the relearned second global adapter to the associated clients (500).

[0178] In some cases, the server (400) may check whether the second neural network model (410) excluding the retrained second global adapter needs to be lightened, and if the second neural network model (410) needs to be lightened, the server may generate a lightweight model for distribution corresponding to the retrained second global adapter and update the retraining resource.

[0179] In another case, the server (400) may check whether the pre-stored global data has changed after the second global adapter has been retrained, and if the global data has changed, select the source data for distribution corresponding to the second global adapter retrained from the global data and update the retraining resource.

[0180] Here, the server (400) can check whether global data has changed when the retraining resource is updated, resample the training data for distribution when the global data has changed, retrain the second global adapter based on the resampled training data, and distribute the second global adapter and the retraining resource including at least one of the training data and the lightweight model to the associated client.

[0181] In addition, when generating a task action, the server (400) may select a task corresponding to the multi-modal features when multi-modal features are received from the clients (500) and generate a task action corresponding to the selected task.

[0182] In this way, the present disclosure can efficiently train neural network models of servers and clients to reflect both common characteristics and individual unique characteristics in various real environments by training a local adapter of a client, retraining a global adapter of a server based on the trained local data, and distributing the retrained second global adapter to the client.

[0183] In addition, the present disclosure can improve task performance by training each client's neural network model to be aligned with the server's global neural network model by having the server distribute a retrained second global adapter and retraining resources to the clients.

[0184] FIG. 5 is a diagram illustrating a neural network model of a client of a federated learning device according to an embodiment of the present disclosure.

[0185] As illustrated in FIG. 5, when multi-modal data is received, the client of the present disclosure can input the received multi-modal data into a first neural network model (510) to output a multi-modal feature.

[0186] The first neural network model (510) may include a pre-trained first foundation model (516), a first global adapter (512) that reflects common characteristics of the client, and a local adapter (514) that reflects unique characteristics of the client.

[0187] Here, the first global adapter (512) can be updated with the second global adapter distributed from the server, and the local adapter (514) can be retrained based on at least one of the client's local data and the server's global data.

[0188] The first neural network model (510) classifies the output values ​​from the pre-trained first foundation model (516) into a first global adapter (512) reflecting common characteristics and a local adapter (514) reflecting unique characteristics, and can output multi-modal features based on the first global adapter (512) and the local adapter (514).

[0189] The first neural network model (510) assigns weights to the global and local features to the result value according to the characteristics of the result value, and if the weight assigned to the result value is greater for the local feature than for the global feature, the result value is classified as a local adapter (514), and if the weight assigned to the result value is less for the local feature than for the global feature, the result value can be classified as a first global adapter (512). This is only one example and is not limited thereto.

[0190] Here, the multi-modal features may include fused features in which the weights of the features corresponding to the first global adapter (512) and the weights of the features corresponding to the local adapter (514) are averaged and the features corresponding to the first global adapter (512) and the features corresponding to the local adapter (514) are fused.

[0191] As an example, if the client is unable to self-learn the local adapter (514), the first neural network model (510) can transmit multi-modal features including local data to the server, and when the re-learned second global adapter is received from the server, the local adapter (514) can be updated based on the re-learned second global adapter.

[0192] In another embodiment, if the client is capable of self-learning the local adapter (514), the first neural network model (510) can send the retrained local adapter (514) to the server, and when the retrained second global adapter is received from the server, the first global adapter (512) can be updated based on the retrained second global adapter.

[0193] In some cases, the first neural network model (510) may be updated based on the retrained second global adapter and the retraining resource when the retrained second global adapter and the retraining resource are received from the server.

[0194] FIG. 6 is a diagram for explaining a neural network model of a server of a federated learning device according to an embodiment of the present disclosure.

[0195] As illustrated in FIG. 6, the server of the present disclosure can input multi-modal features into a second neural network model (410) when multi-modal features are received from a plurality of clients to be connected, and output a task action corresponding to the multi-modal features.

[0196] The second neural network model (410) may include a pre-trained second foundation model (414) and a second global adapter (412) that reflects common characteristics of the federated clients.

[0197] Here, the second global adapter (412) can be retrained and updated based on multiple local adapters when local adapters are extracted from multi-modal features received from the federated clients.

[0198] The second neural network model (410) is initialized based on global data before acquiring multi-modal features from the federated clients, and when local adapters and local data are collected from the federated clients, the second global adapter (412) can be retrained based on the collected local adapters.

[0199] Additionally, the second neural network model (410) can retrain the second global adapter (412) based on the collected local adapter after updating it based on the client's local adapter and local data if a change in retraining resources is required.

[0200] At this time, the second neural network model (410) can retrain only the second global adapter (412) based on the collected local adapter if there is no need to change the retraining resource.

[0201] In some cases, if lightweighting is required in areas other than the second global adapter (412), the second neural network model (410) may be transformed into a lightweight model for distribution corresponding to the retrained second global adapter, thereby updating the resources for retraining.

[0202] In another case, the second neural network model (410) can resample the training data for distribution when the global data changes after the second global adapter is retrained, and retrain the second global adapter (412) based on the resampled training data.

[0203] FIG. 7 is a diagram for explaining the client operation of a federated learning device according to one embodiment of the present disclosure.

[0204] As illustrated in FIG. 7, the client (500) of the present disclosure may include a memory (570) in which a first neural network model (510) including a first global adapter and a local adapter is stored, and a processor (580) that generates multi-modal features based on the output values ​​of the first neural network model (510).

[0205] Here, the processor (580) updates the first global adapter based on the relearned second global adapter when the relearned second global adapter is received from the server, relearns the local adapter based on at least one of the client's local data and the server's global data, and generates multi-modal features based on the first global adapter and the local adapter and provides the multi-modal features to the server.

[0206] The processor (580) can classify the result values ​​output from the pre-learned first foundation model into a first global adapter reflecting common characteristics and a local adapter reflecting unique characteristics, and generate multi-modal features based on the first global adapter and the local adapter and provide them to the server.

[0207] The processor (580) collects multi-modal data, inputs the collected multi-modal data into a first foundation model that has been pre-learned, outputs an inferred result value, and can classify the result value into a first global adapter or a local adapter based on the characteristic weights assigned to the result value.

[0208] For example, the processor (580) may assign weights to the global and local characteristics to the result value according to the characteristics of the result value, and if the weight assigned to the result value is greater for the local characteristic than for the global characteristic, the result value may be classified as a local adapter, and if the weight assigned to the result value is less for the local characteristic than for the global characteristic, the result value may be classified as a first global adapter.

[0209] Additionally, the processor (580) can average the weights of the features corresponding to the first global adapter and the weights of the features corresponding to the local adapter, fuse the features corresponding to the first global adapter and the features corresponding to the local adapter, and generate a multi-modal feature including the fused features and provide the multi-modal feature to the server.

[0210] And, if the processor (580) determines that self-learning for the local adapter is impossible, it can store local data corresponding to the operation of the first neural network model (510), transmit the local data to the server, check whether a re-learned second global adapter is received from the server, and, if the re-learned second global adapter is received, update the local adapter based on the second global adapter.

[0211] Here, the processor (580) may request a second global adapter from the server before operating the first neural network model (510), and upon receiving the second global adapter from the server, initialize a local adapter based on the second global adapter.

[0212] Additionally, the processor (580) can check whether the stored local data satisfies preset transmission conditions, and if the preset transmission conditions are satisfied, transmit the local data to the server.

[0213] For example, when checking whether a preset transmission condition is satisfied, the processor (580) may check whether the preset transmission condition satisfies at least one of a condition for transmitting by a set period, a condition for transmitting by a set amount of data, a condition for transmitting by data having a set characteristic, and a condition for transmitting by a set communication status. However, this is only one example and is not limited thereto.

[0214] Next, the processor (580), if it determines that self-learning for the local adapter is possible, determines whether re-learning for the local adapter of the first neural network model (510) is necessary, and if it determines that re-learning for the local adapter is necessary, re-learns the local adapter based on at least one of the client's local data and the server's global data, transmits the re-learned local adapter to the server, determines whether a second re-learned global adapter is received from the server, and if the second re-learned global adapter is received, updates the first global adapter based on the second re-learned global adapter.

[0215] Here, the processor (580) may request a second global adapter from the server before operating the first neural network model (510), and upon receiving the second global adapter from the server, initialize a local adapter based on the second global adapter.

[0216] In addition, when a relearned second global adapter and a resource for relearning are simultaneously received from the server, the processor (580) can classify the relearned second global adapter and the resource for relearning, store the classified relearning resource, and update the first global adapter based on the relearned second global adapter.

[0217] In some cases, the processor (580) may sequentially perform a process of storing the relearning resources classified in the order of reception times when the relearned second global adapter and the relearning resources are received from the server at different times, and a process of updating the first global adapter based on the relearned second global adapter.

[0218] Additionally, the processor (580) can request resources for relearning from the server and store the resources for relearning when the resources for relearning are received from the server.

[0219] Here, the processor (580) may request resources for retraining from the server each time a second global adapter that has been retrained is received from the server, or may request resources for retraining from the server at preset intervals.

[0220] FIG. 8 is a diagram for explaining the server operation of a federated learning device according to one embodiment of the present disclosure.

[0221] As illustrated in FIG. 8, the server (400) may include a memory (470) in which a second neural network model (410) including a second global adapter is stored, and a processor (480) that generates a task action corresponding to a multi-modal feature received from a client.

[0222] Here, the processor (480) collects local adapters and local data from the federated clients, retrains a second global adapter based on the local adapters collected from the federated clients, checks whether a retraining resource corresponding to the retrained second global adapter is updated, and when the retraining resource is updated, distributes the retrained second global adapter and the updated retraining resource to the federated clients.

[0223] Here, the processor (480) can initialize a second neural network model (410) based on global data before acquiring multi-modal features from the federated clients, and distribute the second global adapter and learning resources to the federated clients.

[0224] And, the processor (480) collects local adapters and local data from the federated clients, classifies the collected local data and local data by client, checks whether local adapters and local data have been collected from all the federated clients, and when local adapters and local data have been collected from all the federated clients, the processor (480) can retrain the second global adapter based on the local adapters collected from the federated clients.

[0225] Here, the processor (480) checks whether a change in resources for retraining is necessary when the local adapters of all the federated clients are collected, and if a change in resources for retraining is necessary, the processor updates the second neural network model (410) based on the local adapters and local data of the clients, and then retrains the second global adapter based on the collected local adapters.

[0226] At this time, if a change in resources for relearning is unnecessary, the processor (480) can relearn only the second global adapter of the second neural network model based on the collected local adapter and distribute the relearned second global adapter to the associated clients.

[0227] In some cases, the processor (480) may check whether the second neural network model (410) excluding the retrained second global adapter needs to be lightened, and if the second neural network model (410) needs to be lightened, the processor may generate a lightweight model for distribution corresponding to the retrained second global adapter and update the retraining resource.

[0228] In another case, the processor (480) may check whether the pre-stored global data has changed after the second global adapter has been retrained, and if the global data has changed, select the source data for distribution corresponding to the second global adapter retrained from the global data and update the retraining resource.

[0229] Here, the processor (480) can check whether global data has changed when the retraining resource is updated, resample the training data for distribution if the global data has changed, retrain the second global adapter based on the resampled training data, and distribute the second global adapter and the retraining resource including at least one of the training data and the lightweight model to the associated client.

[0230] FIG. 9 is a drawing for explaining the overall structure of a federated learning device according to one embodiment of the present disclosure.

[0231] As illustrated in FIG. 9, the client (500) of the present disclosure can generate multi-modal features based on a first neural network model including a first global adapter, a local adapter, and a foundation model core module, and store local data corresponding to the operation of the first neural network model in a database.

[0232] In addition, the server (400) of the present disclosure selects a task corresponding to the multi-modal characteristics of the client (500) based on a second neural network model including a second global adapter and a foundation model core module, generates a task action corresponding to the selected task and provides it to the client (500), and stores global data corresponding to the operation of the second neural network model in a database.

[0233] In addition, the server (400) may check whether the second neural network model, excluding the retrained second global adapter, needs to be lightened, and if the second neural network model needs to be lightened, generate a lightweight model for distribution corresponding to the retrained second global adapter and update the retraining resource.

[0234] In addition, the server (400) can check whether the global data stored in the database has changed after the second global adapter has been retrained, and if the global data has changed, select the source data for distribution corresponding to the retrained second global adapter from the global data stored in the database and update the retraining resource.

[0235] Here, the server (400) can resample training data for distribution from the database when global data stored in the database is changed, retrain a second global adapter based on the resampled training data, and distribute a retraining resource including the second global adapter and at least one of the training data and the lightweight model to the client (500).

[0236] FIG. 10 is a diagram illustrating a global update process of a federated learning device according to an embodiment of the present disclosure.

[0237] As illustrated in FIG. 10, the server of the present disclosure can receive a local adapter containing unique characteristics of each client from a plurality of federated clients.

[0238] The server can then retrain the global adapter based on the local adapter containing the unique characteristics of each client.

[0239] Here, the server's global adapter is retrained based on the local adapter containing the unique characteristics of each client, so that it can learn various contexts experienced by multiple clients.

[0240] Therefore, the server's global adapter can produce results that reflect both the common characteristics of multiple clients and the unique characteristics of each individual client during inference.

[0241] Next, the server can check whether the neural network model excluding the retrained global adapter needs to be lightweight, and if the neural network model needs to be lightweight, it can update the retraining resource by generating a lightweight model for distribution corresponding to the retrained global adapter.

[0242] Additionally, the server can check whether the global data has changed after the global adapter has been retrained, and if the global data has changed, select the source data for distribution corresponding to the retrained global adapter from the global data and update the retraining resource.

[0243] In addition, the server can check whether the global data has changed when the retraining resource is updated, resample the training data for distribution if the global data has changed, retrain the global adapter based on the resampled training data, and distribute the retrained global adapter to multiple federated clients.

[0244] Additionally, the server can distribute updated retraining resources for training local models to multiple federated clients.

[0245] Here, the distributed retraining resource may include at least one of training data for distribution and a lightweight model for distribution.

[0246] In this way, the present disclosure can improve task performance by training each client's neural network model to be aligned with the server's global neural network model by having the server distribute a retrained second global adapter and retraining resources to the clients.

[0247] FIG. 11 is a diagram for explaining a local update process of a federated learning device according to an embodiment of the present disclosure.

[0248] As illustrated in FIG. 11, the client of the present disclosure may request a relearning resource from the server, and store the relearning resource when the relearning resource is received from the server.

[0249] Here, the client can request retraining resources from the server whenever a retrained global adapter is received from the server, or can request retraining resources from the server at preset intervals.

[0250] When a client receives retraining resources from the server, the client can store the source data and lightweight model for model training among the retraining resources in a database.

[0251] Next, the client can retrain the local adapter based on at least one of the global data and the local data, which includes the retraining resource.

[0252] Here, the client retrains the local adapter based on the global data containing the retraining resources received from the server and the local data stored in the database, so that the client's neural network model and the server's global neural network model can be aligned with each other.

[0253] Next, the client can send the relearned local adapter to the server.

[0254] Here, the client can send multi-modal features including relearned local adapters to the server.

[0255] Additionally, when the client receives a retrained global adapter from the server, the client can update its global adapter based on the retrained server global adapter.

[0256] Here, when the client simultaneously receives the retrained global adapter and the retraining resource from the server, the client can classify the retrained global adapter and the retraining resource, store the classified retraining resource, and update the client's global adapter based on the retrained server's global adapter.

[0257] In some cases, the client may sequentially perform a process of storing the retraining resources sorted by the reception time order when the retrained global adapter and the retraining resources are received from the server at different times, and a process of updating the client's global adapter based on the retrained server's global adapter.

[0258] FIG. 12 is a diagram for explaining a multi-modal based service provision process of a federated learning device according to an embodiment of the present disclosure.

[0259] As illustrated in FIG. 12, when the client (500) of the present disclosure receives multi-modal data from a user (600) or the like, it can input the received multi-modal data into a neural network model to generate multi-modal features and transmit the same to the server (400).

[0260] The neural network model of the client (500) may include a pre-learned foundation model, a global adapter reflecting common characteristics of the client, and a local adapter reflecting unique characteristics of the client.

[0261] Here, the global adapter is updated with a global adapter distributed from the server, and the local adapter can be retrained based on at least one of the client's local data and the server's global data.

[0262] Next, when multi-modal features are received from multiple clients that are connected, the server (400) can input the multi-modal features into a neural network model to generate a task action corresponding to the multi-modal features and transmit it to the client.

[0263] Here, the neural network model of the server may include a pre-trained foundation model and a global adapter that reflects common characteristics of the federated clients.

[0264] Here, the global adapter can be retrained and updated based on multiple local adapters as local adapters are extracted from multi-modal features received from federated clients.

[0265] Next, when the client (500) receives a task action from the server (400), it can execute a service corresponding to the task action through a task action executor and provide it to the user.

[0266] Here, if the client (500) is unable to perform self-learning on the local adapter, the client (500) can update the local adapter of the client (500) based on the relearned global adapter when the relearned global adapter is received from the server (400).

[0267] In some cases, if the client (500) is capable of self-learning for the local adapter, the client (500) may transmit the relearned local adapter to the server (400), and when the relearned global adapter is received from the server (400), the client (500) may update its global adapter based on the relearned global adapter of the server (400).

[0268] Additionally, when the client (500) receives the retrained global adapter and the retraining resource from the server (400), the client (500) can update the model based on the retrained global adapter and the retraining resource.

[0269] FIGS. 13 to 15 are diagrams for explaining an example of a multi-modal-based service application of a federated learning device according to an embodiment of the present disclosure.

[0270] As illustrated in FIG. 13, the federated learning device of the present disclosure can provide a multi-modal based smart device control service when the client is a home hub and the server is a cloud that is connected to multiple home hubs.

[0271] For example, the client of the present disclosure can recognize that a user (600) is going outside through image data acquired from a camera.

[0272] Next, the client of the present disclosure can recognize the sound of a door closing through audio data acquired from a microphone.

[0273] Next, the client of the present disclosure can recognize that a vehicle is leaving the parking lot through log data.

[0274] In addition, the client of the present disclosure can input multimodal data including image data, audio data, and log data into a neural network model to generate multimodal features and transmit them to a server.

[0275] Next, the server of the present disclosure can determine the user's status based on multi-modal features collected from the client.

[0276] Next, when the server of the present disclosure recognizes the user's status as being away from home, it can identify the status of a home appliance within the client, select a task to determine whether to execute away mode for a specific home appliance, and generate a task action corresponding to the selected task and transmit it to the client.

[0277] And, when a task action is received, the client of the present disclosure can execute an away mode for the electronic device (610) corresponding to the task action.

[0278] For example, if the electronic devices (610) corresponding to the task action are a robot vacuum cleaner and an air purifier, the client of the present disclosure can provide a multi-modal based smart device control service that turns on the robot vacuum cleaner and turns off the air purifier.

[0279] As illustrated in FIG. 14, the federated learning device of the present disclosure can provide a multi-modal based interactive service when the client is a home hub and the server is a cloud that is connected to multiple home hubs.

[0280] For example, the client of the present disclosure can recognize that a user (600) is angry through vision data acquired from a camera.

[0281] Next, the client of the present disclosure can recognize that the user (600) is angry by the raised voice tone through sound data acquired from the microphone.

[0282] Next, the client of the present disclosure can recognize that an error is continuously occurring in the washing machine through log data.

[0283] In addition, the client of the present disclosure can input multimodal data including vision data, sound data, and log data into a neural network model to generate multimodal features and transmit them to a server.

[0284] Next, the server of the present disclosure can determine the status of the user and the status of the washing machine based on multi-modal features collected from the client.

[0285] Next, the server of the present disclosure can generate a consultation sentence based on the user's status and the washing machine's status, and transmit the generated consultation sentence to the client's local device.

[0286] And, the client of the present disclosure can conduct a conversation with a user (600) through a chatbot (620) of a local device.

[0287] For example, a chatbot (620) on a local device can provide a multi-modal based conversational service that soothes an angry user (600) or adds empathetic small talk through a conversation engine into which context is input.

[0288] As illustrated in FIG. 15, the federated learning device of the present disclosure can provide a multi-modal based smart car service when the client is a smart car and the server is a smart car agent that is connected to multiple smart cars.

[0289] For example, the client of the present disclosure can obtain a user command from a user (600) driving a smart car (650) to check the status of a child sitting in the back seat of the smart car (650), and transmit the obtained user command to the server.

[0290] Next, the server of the present disclosure can analyze a user command and request a rear seat camera (652) image of a smart car (650) to the client.

[0291] Next, the client of the present disclosure can activate the rear seat camera (652) of the smart car (650) to obtain an image of a child sitting in the rear seat and transmit it to the server.

[0292] In addition, the server of the present disclosure can check the current status of the child based on the video of the child sitting in the back seat, and convert the current status of the child into a notification message in the form of text and sound and transmit it to the client.

[0293] Next, the server of the present disclosure can select a task based on the child's current status, generate a task action corresponding to the selected task, and transmit it to the client.

[0294] For example, the server of the present disclosure may select a task as a rear-seat sleep environment mode when the child's current state is a sleep state, and generate a task action for controlling electronic devices of a smart car (650) in sleep mode corresponding to the selected task and transmit the task action to the client.

[0295] Next, the client of the present disclosure can execute a sleep mode for an electronic device corresponding to a task action when a task action is received.

[0296] For example, a multi-modal based smart car service can be provided in which the electronic device corresponding to the task action adjusts the incline of the rear seat to sleep mode if it is a rear seat, adjusts the temperature and humidity device to sleep mode if it is a temperature and humidity device, and turns off the audio device if it is an audio device.

[0297] FIG. 16 and FIG. 17 are drawings for explaining a server product and a client product of a federated learning device according to one embodiment of the present disclosure.

[0298] As illustrated in FIG. 16, the federated learning device of the present disclosure can set the physical configuration and logical configuration between the server and the client in various ranges.

[0299] Here, the client can be an individual device, an individual hub that communicates with multiple devices, or an individual cloud that communicates with multiple hubs, which can be adjusted according to the physical scope.

[0300] Additionally, the server can be a hub that federates multiple devices when the clients are individual devices, a cloud that federates multiple hubs when the clients are individual hubs, or a large cloud that federates multiple clouds when the clients are individual clouds, which can be adjusted according to the physical scope.

[0301] As shown in Fig. 16, the server can be a single cloud that federates multiple home hubs, and the clients can be individual home hubs.

[0302] Additionally, the server can be a single home hub that connects multiple devices, and the clients can be individual devices.

[0303] Meanwhile, as illustrated in FIG. 17, in the federated learning device of the present disclosure, clients can be classified into a first client type that is incapable of self-learning and a second client type that is capable of self-learning.

[0304] The clients of the present disclosure have differences in the learning process depending on the first client type that is incapable of self-learning and the second client type that is capable of self-learning.

[0305] The present disclosure determines that self-learning of a client's local adapter is impossible, stores local data corresponding to the operation of a neural network model, transmits the local data to a server, checks whether a re-learned global adapter is received from the server, and, when the re-learned global adapter of the server is received, updates the local adapter based on the global adapter of the server.

[0306] And, the present disclosure determines whether re-learning of the local adapter of a neural network model is necessary when it is determined that self-learning of the local adapter of the client is possible, and if it is determined that re-learning of the local adapter is necessary, the local adapter is re-learned based on at least one of the local data of the client and the global data of the server, the re-learned local adapter is transmitted to the server, and whether the re-learned global adapter is received from the server, and when the re-learned global adapter of the server is received, the global adapter of the client can be updated based on the re-learned global adapter of the server.

[0307] FIGS. 18 to 20 are diagrams for explaining a federated learning method of a federated learning device according to an embodiment of the present disclosure.

[0308] Figure 18 is a flowchart for explaining the federated learning process of the server in the federated learning device of the present disclosure.

[0309] As illustrated in Fig. 18, the server can initialize the server's neural network model based on global data (S110).

[0310] And, the server can distribute global adapters and learning resources to clients (S120).

[0311] Next, the server can collect local adapters and local data from the federated clients (S130).

[0312] Next, the server can check whether local adapters and local data have been collected from all federated clients (S140).

[0313] And, the server can check whether a change in resources for relearning is necessary when local adapters and local data are collected from all federated clients (S150).

[0314] Next, if a change in resources for retraining is required, the server can retrain the global adapter based on the collected local adapter after updating the neural network model of the server based on the client's local adapter and local data (S160).

[0315] Here, the server can retrain only the second global adapter of the neural network model of the server based on the collected local adapter if a change in the retraining resource is unnecessary (S220), and distribute the retrained global adapter to the federated clients (S230).

[0316] And, the server can check whether the neural network model excluding the retrained global adapter needs to be lightened (S170).

[0317] Next, if the lightweighting of the neural network model is required, the server can update the retraining resources by creating a lightweight model for distribution corresponding to the retrained global adapter (S200).

[0318] Next, the server can check whether the pre-stored global data has changed after the global adapter has been relearned (S180).

[0319] And, when the global data changes, the server can select the source data for distribution corresponding to the global adapter retrained from the global data and update the retraining resource.

[0320] Here, the server can resample training data for distribution when global data changes (S210), retrain the global adapter based on the resampled training data (S190), and distribute the global adapter and a retraining resource including at least one of the training data and the lightweight model to the associated client (S120).

[0321] In this way, the server of the present disclosure can select a task corresponding to multi-modal features received from clients based on a neural network model including a retrained global adapter, and generate a task action corresponding to the selected task.

[0322] Figure 19 is a flowchart for explaining the federated learning process of a client that is unable to learn on its own in the federated learning device of the present disclosure.

[0323] As illustrated in Fig. 19, the client can request a global adapter from the server, and upon receiving the global adapter from the server, the client can initialize the local adapter of the client based on the global adapter of the server (S310).

[0324] And, the client can operate a neural network model and store local data corresponding to the operation of the neural network model (S320).

[0325] Next, the client can check whether transmission of local data is possible (S330).

[0326] Here, the client can check whether the stored local data satisfies the preset transmission conditions, and if the preset transmission conditions are satisfied, the local data can be transmitted to the server (S350).

[0327] For example, when checking whether preset transmission conditions are satisfied, the client may check whether the preset transmission conditions satisfy at least one of the conditions for transmitting at a set interval, the conditions for transmitting at a set amount of data, the conditions for transmitting at a set characteristic rate, and the conditions for transmitting at a set communication status. This is only one example and is not limited thereto.

[0328] Then, the client can check whether a relearned global adapter is received from the server (S340), and if the relearned global adapter of the server is received, the client's local adapter can be updated based on the global adapter of the server (S360).

[0329] In this way, the client can receive a task action from the server based on the neural network model including the updated local adapter and execute a service corresponding to the task action to provide it to the user.

[0330] Figure 20 is a flowchart for explaining the federated learning process of a client capable of self-learning in the federated learning device of the present disclosure.

[0331] As illustrated in Figure 20, the client can request a global adapter from the server, and upon receiving the global adapter from the server, the client can initialize the local adapter of the client based on the global adapter of the server (S410).

[0332] Next, the client can operate the neural network model (S420) and determine whether retraining is necessary for the local adapter of the neural network model (S430).

[0333] And, if the client determines that re-learning of the local adapter is necessary, the client can re-learn the local adapter based on at least one of the client's local data and the server's global data (S440) and transmit the re-learned local adapter to the server (S450).

[0334] Next, the client can check whether the global adapter of the retrained server is received from the server (S460), and if the global adapter of the retrained server is received, the client's global adapter can be updated based on the global adapter of the retrained server (S470).

[0335] Here, when the client simultaneously receives the retrained global adapter and the retraining resource from the server, the client can classify the retrained global adapter and the retraining resource, store the classified retraining resource, and update the client's global adapter based on the retrained server's global adapter.

[0336] In some cases, if the client receives the retrained global adapter and the retraining resources from the server at different times, the client may sequentially perform the process of storing the retraining resources sorted by the reception time order and the process of updating the client's global adapter based on the retrained server's global adapter.

[0337] In another case, the client may request retraining resources from the server and store the retraining resources when they are received from the server.

[0338] Here, the client can request retraining resources from the server whenever a retrained global adapter is received from the server, or can request retraining resources from the server at preset intervals.

[0339] In this way, the client can receive a task action from the server based on a neural network model including a retrained local adapter and execute a service corresponding to the task action to provide it to the user.

[0340] The present disclosure efficiently trains neural network models of servers and clients to reflect both common characteristics and individual unique characteristics in various real environments by training a local adapter of a client, retraining a global adapter of a server based on the trained local data, and distributing the retrained second global adapter to the client.

[0341] In addition, the present disclosure can improve task performance by training each client's neural network model to be aligned with the server's global neural network model by having the server distribute a retrained second global adapter and retraining resources to the clients.

[0342] The above-described present disclosure can be implemented as computer-readable code on a program-recorded medium. The computer-readable medium includes all types of recording devices that store data that can be read by a computer system. Examples of computer-readable media include hard disk drives (HDDs), solid-state disk drives (SSDs), silicon disk drives (SDDs), read-only memory (ROM), random access memory (RAM), CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, the computer may include a processor (180) of an artificial intelligence device.

[0343] The federated learning device according to the present disclosure has remarkable industrial applicability because it has the effect of efficiently training the neural network models of the server and client to reflect both common characteristics and individual unique characteristics in various real environments.

Claims

1. At least one client generating multimodal features based on a first neural network model including a first global adapter and a local adapter; and A server that generates a task action corresponding to the multi-modal feature based on a second neural network model including a second global adapter and provides the task action to the client, The above server, A federated learning device characterized in that it collects local adapters learned from the client, retrains a second global adapter of the second neural network model based on the collected local adapters, and distributes the retrained second global adapter to the client to update the first global adapter of the client with the retrained second global adapter.

2. In paragraph 1, The above first neural network model is, Pre-trained first foundation model; A first global adapter reflecting the common characteristics of the above clients; and Includes a local adapter that reflects the unique characteristics of the client, The above first global adapter, Updated with the second global adapter distributed from the above server, The above local adapter is, A federated learning device characterized in that relearning is performed based on at least one of the local data of the client and the global data of the server.

3. In paragraph 1, The above client, A federated learning device characterized in that, when it is determined that self-learning for the local adapter is impossible, local data corresponding to the operation of the first neural network model is stored and the local data is transmitted to the server.

4. In paragraph 3, The above client, A federated learning device characterized in that it checks whether the relearned second global adapter is received from the server, and when the relearned second global adapter is received, it updates the local adapter based on the second global adapter.

5. In paragraph 1, The above client, A federated learning device characterized in that, if it is determined that self-learning for the local adapter is possible, it determines whether re-learning for the local adapter of the first neural network model is necessary, and if it is determined that re-learning for the local adapter is necessary, it re-learns the local adapter based on at least one of the local data of the client and the global data of the server, and transmits the re-learned local adapter to the server.

6. In paragraph 5, The above client, A federated learning device characterized in that it checks whether the relearned second global adapter is received from the server, and when the relearned second global adapter is received, it updates the first global adapter based on the relearned second global adapter.

7. In paragraph 1, The above second neural network model is, A pre-trained second foundation model; and, Includes a second global adapter that reflects the common characteristics of the federated clients, The above second global adapter, A federated learning device characterized in that it is relearned and updated based on a plurality of local adapters received from the federated clients.

8. In paragraph 1, The above server, A federated learning device characterized in that, before acquiring multi-modal features from federated clients, the second neural network model is initialized based on global data, and the second global adapter and learning resources are distributed to the federated clients.

9. In paragraph 8, The above server, A federated learning device characterized in that it collects local adapters and local data from the federated clients, classifies the collected local data and local data by client, checks whether local adapters and local data have been collected from all the federated clients, and retrains the second global adapter based on the local adapters collected from the federated clients when the local adapters and local data have been collected from all the federated clients.

10. In paragraph 9, The above server, A federated learning device characterized in that when retraining the second global adapter, when the local adapters of all the federated clients are collected, it is checked whether a change in the retraining resource is necessary, and if a change in the retraining resource is necessary, the second neural network model is updated based on the local adapter and local data of the client, and then the second global adapter is retrained based on the collected local adapter.

11. In paragraph 10, The above server, A federated learning device characterized in that, if the change of the relearning resource is unnecessary, only the second global adapter of the second neural network model is relearned based on the collected local adapter, and the relearned second global adapter is distributed to the federated clients.

12. In paragraph 10, The above server, A federated learning device characterized in that it checks whether a second neural network model excluding the retrained second global adapter needs to be lightweight, and if the second neural network model needs to be lightweight, it creates a lightweight model for distribution corresponding to the retrained second global adapter and updates the retraining resource.

13. In paragraph 10, The above server, A federated learning device characterized in that after the second global adapter is retrained, it checks whether pre-stored global data has changed, and if the global data has changed, it selects source data for distribution corresponding to the retrained second global adapter from the global data and updates the retraining resource.

14. In paragraph 13, The above server, A federated learning device characterized in that when the above retraining resource is updated, it checks whether global data has changed, and if the global data has changed, it resamples learning data for distribution, it retrains the second global adapter based on the resampled learning data, and it distributes the second global adapter and the retraining resource including at least one of the learning data and the lightweight model to the federated client.

15. A federated learning method of a federated learning device including at least one client generating multi-modal features based on a first neural network model including a first global adapter and a local adapter, and a server generating a task action corresponding to the multi-modal features based on a second neural network model including a second global adapter and providing the same to the client, A step in which the client retrains a local adapter of the first neural network model based on at least one of local data and global data; The step of the client transmitting the relearned local adapter to the server; A step in which the server collects relearned local adapters from the client; A step in which the server retrains the second global adapter of the second neural network model based on the collected local adapter; The step of the server distributing the relearned second global adapter to the client; The step of the client receiving the second global adapter relearned from the server; and A federated learning method, characterized in that the client includes a step of updating the first global adapter of the first neural network model based on the retrained second global adapter.

Citation Information

Patent Citations

  • Apparatus and system for managing federated learning resource, and resource efficiency method thereof

    KR102567565B1

  • Framework for rapidly prototyping federated learning algorithms

    US20220129786A1

  • KR20230062553A