Artificial intelligence server and operation method thereof

By transmitting a fine-tuned multi-modal model layer from a cloud server to a home server and updating it with encrypted user data, the AI server achieves personalized and accurate context inference and actions, addressing the limitations of non-personalized global models.

WO2025105524A1PCT designated stage expired Publication Date: 2025-05-22LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2023/018329
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing artificial intelligence servers provide only non-personalized global models, failing to account for user-specific input information, and lack parameter settings for encoders to encrypt input data effectively across different modalities.

Method used

A personalized multi-modal model is achieved by transmitting a fine-tuned layer of a multi-modal model learned by a cloud server to a home server, where the model is updated using encrypted input information collected by the home server, with modality-specific encryption modules.

Benefits of technology

This approach enables user-tailored context inference and actions, ensuring accurate situation inference and appropriate responses using various types of data, while maintaining privacy through encryption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023018329_22052025_PF_FP_ABST
    Figure KR2023018329_22052025_PF_FP_ABST
Patent Text Reader

Abstract

An artificial intelligence server according to an embodiment of the present disclosure may comprise: a memory for storing a multi-modal model for inferring a situation by using a plurality of modalities of different types; a communication circuit for communicating with a plurality of edge apparatuses or a cloud server; and a processor. The processor: receives parameter information about a finely tuned layer of the multi-modal model from the cloud server; receives a plurality of encrypted modalities from the plurality of edge apparatuses; and updates the parameter information about the finely tuned layer of the multi-modal model by using the received plurality of encrypted modalities.
Need to check novelty before this filing date? Find Prior Art

Description

Artificial intelligence server and its operation method

[0001] The present invention relates to an artificial intelligence server, and more particularly, to an artificial intelligence server that infers a situation using a multi-modal model.

[0002] A multimodal model represents a model that processes multiple types of input data and produces multiple types of output.

[0003] Multimodal models can primarily handle multiple data types simultaneously, including text, images, and audio. For example, they can simultaneously understand images and captions to generate descriptions, or process text and speech together to improve natural language understanding and speech recognition.

[0004] Previously, there was a technology that encrypted input information on edge devices and sent it to an artificial intelligence server for processing to protect privacy.

[0005] However, previously, only non-personalized global artificial intelligence models were provided, even though input information differed for each user.

[0006] Additionally, for each modality (each data type), there was no setting of encoder parameters for encrypting input information.

[0007] The purpose of the present disclosure is to provide a personalized multi-modal model by transmitting a fine-tuned layer of a multi-modal model learned by a cloud server to a home server and updating the fine-tuned layer using encrypted input information collected from the home server.

[0008] The purpose of the present disclosure is to learn a multi-modal model by encrypting modality data using an encryption module (encoder) according to the type of modality.

[0009] An artificial intelligence server according to one embodiment of the present disclosure may include a memory for storing a multi-modal model for inferring a situation using a plurality of modalities having different types, a communication circuit for communicating with a plurality of edge devices or a cloud server, and a processor for receiving parameter information for a fine-tuned layer of the multi-modal model from the cloud server, receiving a plurality of encrypted modalities from the plurality of edge devices, and updating the parameter information of the fine-tuned layer of the multi-modal model using the received plurality of encrypted modalities.

[0010] According to one embodiment of the present disclosure, an artificial intelligence server may include a method of operating an artificial intelligence server, the method including: receiving parameter information for a fine-tuned layer of a multi-modal model that infers a situation using a plurality of modalities having different types from a cloud server; receiving a plurality of encrypted modalities from the plurality of edge devices; and updating the parameter information of the fine-tuned layer of the multi-modal model using the received plurality of encrypted modalities.

[0011] According to embodiments of the present disclosure, user-tailored context inference and actions for context inference can be effectively performed based on a personalized multi-modal model.

[0012] According to an embodiment of the present disclosure, by using various types of data, an accurate situation can be inferred and an appropriate action can be taken for the inferred situation.

[0013] Figure 1 illustrates an AI device according to one embodiment of the present disclosure.

[0014] FIG. 2 illustrates an AI server according to one embodiment of the present disclosure.

[0015] FIG. 3 is a diagram for explaining the configuration of an artificial intelligence system according to one embodiment of the present disclosure.

[0016] FIG. 4 is a ladder diagram for explaining an operation method of an artificial intelligence system according to an embodiment of the present disclosure.

[0017] FIG. 5a is a diagram illustrating a process in which a cloud server collects modalities according to an embodiment of the present disclosure, and FIG. 5b is a diagram illustrating a process in which a cloud server learns a multi-modal model using the collected modalities.

[0018] FIG. 6 is a diagram illustrating a fine tuning process of a multi-modal model according to an embodiment of the present disclosure.

[0019] FIG. 7 is a ladder diagram for explaining an operation method of an artificial intelligence system according to an embodiment of the present disclosure.

[0020] FIGS. 8A to 8C are diagrams illustrating a process of inferring a situation from multiple modalities using a multi-modal model learned according to embodiments of the present disclosure.

[0021] FIG. 9 is a diagram illustrating an example in which each of a plurality of home servers has a personalized multi-modal model according to one embodiment of the present disclosure.

[0022] Artificial intelligence (AI) is the study of artificial intelligence or the methodologies for creating it, while machine learning (ML) defines various problems in the field of AI and studies the methodologies for solving them. Machine learning is also defined as an algorithm that improves performance on a task through consistent experience.

[0023] An artificial neural network (ANN) is a model used in machine learning. It can refer to a model with problem-solving capabilities, comprised of artificial neurons (nodes) formed by the connection of synapses. An ANN can be defined by the connection patterns between neurons in different layers, the learning process that updates model parameters, and the activation function that generates output values.

[0024] An artificial neural network may include an input layer, an output layer, and optionally one or more hidden layers. Each layer contains one or more neurons, and the artificial neural network may include synapses connecting neurons. In an artificial neural network, each neuron can output a function value of an activation function based on input signals, weights, and biases received through the synapses.

[0025] Model parameters are parameters determined through learning, including synaptic connection weights and neuron biases. Hyperparameters are parameters that must be set before learning in machine learning algorithms, including the learning rate, number of iterations, mini-batch size, and initialization function.

[0026] The goal of artificial neural network training can be seen as determining model parameters that minimize a loss function. The loss function can be used as an indicator for determining optimal model parameters during the artificial neural network training process.

[0027] Machine learning can be classified into supervised learning, unsupervised learning, and reinforcement learning depending on the learning method.

[0028] Supervised learning refers to a method for training an artificial neural network when given labels for the training data. The labels can refer to the correct answer (or output value) that the artificial neural network must infer when the training data is input to the artificial neural network. Unsupervised learning can refer to a method for training an artificial neural network when the training data is not given labels. Reinforcement learning can refer to a learning method in which an agent defined within a given environment is trained to select actions or action sequences that maximize the cumulative reward in each state.

[0029] Machine learning implemented with a deep neural network (DNN) containing multiple hidden layers among artificial neural networks is also called deep learning, and deep learning is a subset of machine learning. Hereinafter, the term "machine learning" is used to encompass deep learning.

[0030] FIG. 1 illustrates an artificial intelligence (AI) device according to one embodiment of the present disclosure.

[0031] The AI ​​device (100) can be implemented as a fixed device or a movable device, such as a TV, a projector, a mobile phone, a smart phone, a desktop computer, a laptop, a digital broadcasting terminal, a PDA (personal digital assistant), a PMP (portable multimedia player), a navigation device, a tablet PC, a wearable device, a set-top box (STB), a DMB receiver, a radio, a washing machine, a refrigerator, a desktop computer, digital signage, a robot, a vehicle, etc.

[0032] Referring to FIG. 1, the terminal (100) may include a communication circuit (110), an input interface (120), a running processor (130), a sensor (140), an output interface (150), a memory (170), and a processor (180).

[0033] The communication circuit (110) can transmit and receive data with external devices such as other AI devices (100a to 100e) or AI servers (200) using wired or wireless communication technology.

[0034] The communication circuit (110) can transmit and receive sensor information, user input, learning models, control signals, etc. with external devices.

[0035] Communication technologies used by the communication circuit (110) include GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), LTE (Long Term Evolution), 5G, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Bluetooth™, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), ZigBee, NFC (Near Field Communication), etc.

[0036] The input interface (120) can obtain various types of data.

[0037] The input interface (120) may include a camera for inputting a video signal, a microphone for receiving an audio signal, a user input interface for receiving information from a user, etc.

[0038] Cameras and microphones can be treated as sensors, and signals obtained from the camera or microphone can be called sensing data or sensor information.

[0039] The input interface (120) can obtain input data to be used when obtaining output using learning data and learning models for model learning. The input interface (120) can also obtain unprocessed input data, in which case the processor (180) or learning processor (130) can extract input features as preprocessing for the input data.

[0040] The learning processor (130) can train a model composed of an artificial neural network using learning data. Here, the trained artificial neural network may be referred to as a learning model. The learning model can be used to infer result values ​​for new input data other than the learning data, and the inferred values ​​can be used as a basis for making decisions regarding certain actions.

[0041] The running processor (130) can perform AI processing together with the running processor (240) of the AI ​​server (200).

[0042] The running processor (130) may include memory integrated or implemented in the AI ​​device (100). Alternatively, the running processor (130) may be implemented using memory (170), external memory directly coupled to the AI ​​device (100), or memory maintained in an external device.

[0043] The sensor (140) can obtain at least one of internal information of the AI ​​device (100), information about the surrounding environment of the AI ​​device (100), and user information using a plurality of sensors.

[0044] The sensor (140) may include one or more of a proximity sensor, a light sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, a lidar, and a radar.

[0045] The output interface (150) can generate output related to visual, auditory, or tactile sensations.

[0046] The output interface (150) may include a display unit that outputs visual information, a speaker that outputs auditory information, a haptic module that outputs tactile information, etc.

[0047] The memory (170) can store data that supports various functions of the AI ​​device (100). The memory (170) can store input data, learning data, learning models, learning history, etc. obtained from the input interface (120).

[0048] The processor (180) can determine at least one executable operation of the AI ​​device (100) based on information determined or generated using a data analysis algorithm or a machine learning algorithm.

[0049] The processor (180) can control components of the AI ​​device (100) to perform determined operations.

[0050] To this end, the processor (180) can request, retrieve, receive or utilize data from the running processor (130) or memory (170), and control components of the AI ​​device (100) to execute at least one of the executable operations, a predicted operation or an operation determined to be desirable.

[0051] When the processor (180) requires connection to an external device to perform a determined operation, it can generate a control signal for controlling the external device and transmit the generated control signal to the external device.

[0052] The processor (180) can obtain intent information for user input and determine the user's requirements based on the obtained intent information.

[0053] The processor (180) can obtain intent information corresponding to the user input by using at least one of a STT (Speech To Text) engine for converting voice input into a string or a natural language processing (NLP) engine for obtaining intent information of natural language.

[0054] At least one of the STT engine or the NLP engine may be configured with an artificial neural network, at least in part, trained according to a machine learning algorithm. Furthermore, at least one of the STT engine or the NLP engine may be trained by the learning processor (130), the learning processor (240) of the AI ​​server (200), or through distributed processing thereof.

[0055] The processor (180) can collect history information including the operation details of the AI ​​device (100) or the user's feedback on the operation, and store the information in the memory (170) or the learning processor (130), or transmit the information to an external device such as an AI server (200). The collected history information can be used to update the learning model.

[0056] The processor (180) can control at least some of the components of the AI ​​device (100) to drive an application program stored in the memory (170). Furthermore, the processor (180) can operate two or more of the components included in the AI ​​device (100) in combination to drive the application program.

[0057] FIG. 2 illustrates an AI server (200) according to one embodiment of the present disclosure.

[0058] Referring to FIG. 2, the AI ​​server (200) may refer to a device that trains an artificial neural network using a machine learning algorithm or utilizes a trained artificial neural network. Here, the AI ​​server (200) may be composed of multiple servers to perform distributed processing, and may be defined as a 5G network.

[0059] The AI ​​server (200) may be included as part of the configuration of the AI ​​device (100) and may perform at least part of the AI ​​processing together.

[0060] The AI ​​server (200) may include a communication circuit (210), a memory (230), a learning processor (240), and a processor (260).

[0061] The communication circuit (210) can transmit and receive data with an external device such as an AI device (100).

[0062] The memory (230) may include a model storage unit (231). The model storage unit (231) may store a model (or artificial neural network, 231a) being learned or learned through the learning processor (240).

[0063] A learning processor (240) can train an artificial neural network (231a) using learning data. The learning model can be used while mounted on the AI ​​server (200) of the artificial neural network, or can be mounted on an external device such as an AI device (100).

[0064] The learning model may be implemented in hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented in software, one or more instructions constituting the learning model may be stored in memory (230).

[0065] The processor (260) can use a learning model to infer a result value for new input data and generate a response or control command based on the inferred result value.

[0066] FIG. 3 is a diagram for explaining the configuration of an artificial intelligence system according to one embodiment of the present disclosure.

[0067] The artificial intelligence system (30) may include a cloud server (200-1), a home server (200-2), a first edge device (100-1), and a second edge device (100-2).

[0068] Although Figure 3 illustrates that there are two edge devices, this is only an example and there may be more edge devices.

[0069] Each of the cloud server (200-1) and the home server (200-2) may include the components of FIG. 2. That is, each of the cloud server (200-1) and the home server (200-2) may be an example of the AI ​​server (200) of FIG. 2. Each of the cloud server (200-1) and the home server (200-2) may be referred to as an artificial intelligence server.

[0070] Each of the first edge device (100-1) and the second edge device (100-2) may include all of the components illustrated in FIG. 1. That is, each of the first edge device (100-1) and the second edge device (100-2) may be an example of the AI ​​device (100) illustrated in FIG. 1.

[0071] As another example, each of the first edge device (100-1) and the second edge device (100-2) may be a component of the AI ​​device (100).

[0072] Each of the first edge device (100-1) and the second edge device (100-2) may be an electronic device capable of collecting modality data, such as a camera, a sensor, a home appliance, a TV, or a microphone.

[0073] The first edge device (100-1) and the second edge device (100-2) may be electronic devices used by the same user. The first edge device (100-1) and the second edge device (100-2) may be installed in a specific space, for example, within a home.

[0074] FIG. 4 is a ladder diagram for explaining an operation method of an artificial intelligence system according to an embodiment of the present disclosure.

[0075] The input information below may be modality information collected by the edge device.

[0076] Referring to FIG. 4, the cloud server (200-1) can learn a multi-model model (S401), and the cloud server (200-1) can set different encoders for each modality (S403).

[0077] The cloud server (200-1) may be equipped with multiple encoders corresponding to each of the multiple modalities. The multiple encoders may be included in the processor (260) or may be equipped separately from the processor (260).

[0078] Each of the plurality of encoders may be an encryption module that encrypts the modality received from the edge device (100-1).

[0079] The running processor (240) or processor (260) of the cloud server (200-1) can set the parameters of the encoder differently for each modality.

[0080] The learning processor (240) or processor (260) of the cloud server (200-1) can learn a multi-model model.

[0081] A multimodal model can be an artificial neural network-based model that infers situations by taking multiple modalities as input.

[0082] A multimodal model may be a model trained using contrastive learning.

[0083] Contrastive learning can be a method to learn the distances between similar sample pairs in the embedding space to be close to each other and the distances between dissimilar sample pairs to be farther away.

[0084] The processor (260) of the cloud server (200-1) can learn a multi-modal model so that the distance between feature vectors of the same type of modality becomes close, and the distance between feature vectors of different types of modality becomes far from each other.

[0085] Modality can be named as multi-modal data as input data of a multi-modal model.

[0086] Multimodal data may include one or more of image data, text data, sound data, sensor data, and log data (table data or knowledge data) of an edge device (100-1).

[0087] In one embodiment, the cloud server (200-1) can collect multiple modalities from multiple edge devices.

[0088] In another embodiment, the multiple modalities may be data generated from a generative AI model of a generative AI server rather than actual data. This is because privacy issues may arise in the case of actual data.

[0089] FIG. 5a is a diagram illustrating a process in which a cloud server collects modalities according to an embodiment of the present disclosure, and FIG. 5b is a diagram illustrating a process in which a cloud server learns a multi-modal model using the collected modalities.

[0090] The cloud server (200-1) can collect image data, sensing data, sound data, and text data.

[0091] The cloud server (200-1) can receive modalities from each of a plurality of edge devices.

[0092] The plurality of encoders may include an image encoder (501) that encrypts image data, a sensor encoder (503) that encrypts sensing data, a sound encoder (505) that encrypts sound data, and a text encoder (507) that encrypts text data.

[0093] The learning processor (240) or processor (260) of the cloud server (200-1) can train an image encoder (501) corresponding to image data through a convolutional neural network (CNN).

[0094] The learning processor (240) or processor (260) of the cloud server (200-1) can train a sound encoder (505) corresponding to sound data through a recurrent neural network (RNN).

[0095] The learning processor (240) or processor (260) of the cloud server (200-1) can train a text encoder (507) corresponding to text data through the BERT (Bidirectional Encoder Representations from Transformers) algorithm.

[0096] The image encoder (501) can output a first vector set (511) that is input to a multi-modal model using image data.

[0097] The sensor encoder (503) can output a second vector set (513) that is input to a multi-modal model using sensing data. The sensing data may be a millimeter wave signal related to the presence or absence of a person or the person's posture.

[0098] The sound encoder (505) can output a third vector set (515) that is input to the multi-modal model using sound data.

[0099] The text encoder (507) can output a set of fourth vectors (517) that are input to the multi-modal model using text data.

[0100] The first to fourth vector sets (511 to 517) can be used later for a multi-modal model to infer situations. Each of the first to fourth feature vector sets (511 to 517) may include 512 vectors, but this is merely an example.

[0101] In addition, the cloud server (100-1) can collect additional log information from edge devices (100-1) or home appliances. The log information from home appliances may include information about the operating status and operating history of the home appliances.

[0102] Referring to FIG. 5b, the cloud server (200-1) can train a multi-modal model using the first to fourth vector sets (511 to 517).

[0103] The cloud server (200-1) can train a multi-modal model through contrastive learning. The multi-modal model can output an inferred situation using the first to fourth vector sets (511 to 517).

[0104] The multi-modal model can be trained to cluster the first to fourth vector sets (511 to 517) to specific data points in the embedding space.

[0105] For example, a multi-modal model can infer a cooking situation (Context A) or a TV watching situation (Context B) using the first to fourth vector sets (511 to 517) as input.

[0106] Again, Figure 4 is explained.

[0107] The processor (260) of the cloud server (200-1) can transmit first encoder parameter information according to the settings to the first edge device (100-1) through the communication circuit (210) (S405), and can transmit second encoder parameter information to the second edge device (100-2) through the communication circuit (210) (S407).

[0108] The first encoder parameter information may include information about parameters of the first encoder.

[0109] The second encoder parameter information may include information about parameters of the second encoder.

[0110] The parameter information may include parameters that can be used to encrypt input information input to the encoder according to the type of modality.

[0111] The processor (260) of the cloud server (200-1) can transmit first encoder parameter information corresponding to a modality type that the first edge device (100-1) can collect to the first edge device (100-1).

[0112] The processor (260) of the cloud server (200-1) can transmit second encoder parameter information corresponding to a modality type that the second edge device (100-2) can collect to the second edge device (100-2).

[0113] The first edge device (100-1) can encrypt the first input information based on the received first encoder parameter information (S409) and transmit the encrypted first input information to the cloud server (200-1) (S411).

[0114] The first edge device (100-1) can collect one or more modalities. The collected modalities may be data with privacy issues. Specifically, the modalities may be data with privacy issues, such as image data captured within a home or audio data of a speaker conversing within a home.

[0115] The first edge device (100-1) may be equipped with a first encoder corresponding to the type of modality it can collect.

[0116] For example, if the first edge device (100-1) is capable of collecting image data, the first edge device (100-1) may be equipped with an image encoder capable of encrypting the image data.

[0117] The first edge device (100-1) can update the parameters of the first encoder based on the first encoder parameter information.

[0118] The first edge device (100-1) can encrypt the first input information using the parameters of the updated first encoder.

[0119] The second edge device (100-2) can encrypt the second input information based on the received second encoder parameter information (S413) and transmit the encrypted second input information to the cloud server (200-1) (S415).

[0120] The second edge device (100-2) can collect one or more modalities. The collected modalities may be data with privacy issues. Specifically, the modalities may be data with privacy issues, such as image data captured within a home or audio data of a speaker conversing within a home.

[0121] The second edge device (100-2) may be equipped with a first encoder corresponding to the type of modality it can collect.

[0122] For example, if the second edge device (100-2) is capable of collecting sensing data, the second edge device (100-2) may be equipped with a sensing encoder capable of encrypting the sensing data.

[0123] The second edge device (100-2) can update the parameters of the second encoder based on the second encoder parameter information.

[0124] The second edge device (100-2) can encrypt the second input information using the parameters of the updated second encoder and transmit the encrypted second input information to the cloud server (200-1).

[0125] The cloud server (200-1) can fine-tune the layers of the multi-modal model based on the encrypted first and second input information (S417).

[0126] The encrypted input information may be encrypted in one or more of multiple modalities.

[0127] A multimodal model can include multiple layers. The multiple layers can include an input layer that is input to the multimodal model, a hidden layer that extracts features from the data input to the input layer, and an output layer that outputs an inferred situation using the extracted features.

[0128] Each of the input layer, hidden layer, and output layer can consist of one or more layers.

[0129] The running processor (240) or processor (260) of the cloud server (200-1) can fine-tune the parameters of one or more layers among the plurality of layers constituting the multi-modal model based on encrypted input information.

[0130] The parameters of a layer can be weights.

[0131] Fine tuning can be the process of adjusting some of the parameters of the layers that make up a pre-trained multimodal model to suit a new purpose (or task).

[0132] Fine tuning can be a process of improving the situational inference performance of a multi-modal model by adjusting various parameters such as the amount of data used for learning, learning rate, and epoch.

[0133] The cloud server (200-1) can generate a personalized multi-modal model using first input information received from the first edge device (100-1) and second input information received from the second edge device (100-2). The personalized multi-modal model may be a model in which parameters of some layers are adjusted based on encrypted first and second input information.

[0134] FIG. 6 is a diagram illustrating a fine tuning process of a multi-modal model according to an embodiment of the present disclosure.

[0135] Referring to FIG. 6, the first edge device (100-1) may be a camera that collects image data. The first edge device (100-1) may include an image encoder (610) that encrypts the image data.

[0136] The first edge device (100-1) can collect first input information and generate encrypted first input information through an image encoder (610). The encrypted first input information may be a first vector set including 512 vectors. The first edge device (100-1) can transmit the encrypted first input information to a cloud server (200-1).

[0137] The second edge device (100-2) may be a sensor that collects sensing data. The second edge device (100-2) may include a sensor encoder (630) that encrypts the sensing data.

[0138] The second edge device (100-2) can collect second input information and generate encrypted second input information through the sensor encoder (630). The encrypted second input information may be a second vector set including 512 vectors. The second edge device (100-2) can transmit the encrypted second input information to the cloud server (200-1).

[0139] The cloud server (200-1) can fine-tune some layers of the multi-modal model based on first input information received from the first edge device (100-1) and second input information received from the second edge device (100-2).

[0140] The cloud server (200-1) can fine-tune some layers of the multi-modal model so that the multi-modal model, which has received the encrypted first and second input information, can infer that the multi-modal model is in a specific context (Context A). The specific context may be a situation where the user is cooking.

[0141] Again, Figure 4 is explained.

[0142] The cloud server (200-1) can transmit parameter information of the fine-tuned layer to the home server (200-2) through a communication circuit (210) (S419).

[0143] The processor (260) of the cloud server (200-1) can transmit information about updated parameters according to fine tuning of the layers of the multi-modal model to the home server (200-2) through the communication circuit (210).

[0144] The parameter information of a fine-tuned layer may include the weights of the parameters that constitute the layer. The weights may have values ​​between 0 and 1, and the number of weights may be 262144 (512 x 512).

[0145] The cloud server (200-1) can transmit parameter information for layers that have undergone different fine tuning to each home server. That is, the cloud server (200-1) can perform fine tuning of a multi-modal model through multiple modalities received from edge devices (100-1, 100-2) connected to the home server (200-2). Accordingly, a personalized multi-modal model can be provided to each home server.

[0146] In another embodiment, the processor (260) of the cloud server (200-1) may transmit a multi-modal model to the home server (200-2). The processor (260) of the cloud server (200-1) may transmit parameter information for a plurality of layers constituting the multi-modal model to the home server (200-2).

[0147] The home server (200-2) can store parameter information of a fine-tuned layer received from the cloud server (200-1) or parameter information for a plurality of layers constituting a multi-modal model in the memory (230).

[0148] The home server (200-2) can set parameters for layers of a pre-stored multi-modal model using parameter information of a fine-tuned layer received from the cloud server (200-1).

[0149] After that, the first edge device (100-1) encrypts the third input information (S421) and transmits the encrypted third input information to the home server (200-2) (S423).

[0150] The second edge device (100-2) encrypts the fourth input information (S425) and transmits the encrypted fourth input information to the home server (200-2) (S427).

[0151] The home server (200-2) can independently update the layers of the multi-modal model based on the encrypted third and fourth input information (S429).

[0152] The running processor (240) or processor (260) of the home server (200-2) can fine-tune one or more layers among the plurality of layers constituting the multi-modal model based on the encrypted third and fourth input information.

[0153] The running processor (240) or processor (260) of the home server (200-2) can obtain a third vector set using a first encoder corresponding to the third input information, and can obtain a fourth vector set using a second encoder corresponding to the fourth input information.

[0154] The running processor (240) or processor (260) of the home server (200-2) can adjust the weights of the parameters of one or more layers among the plurality of layers constituting the multi-modal model using the third and fourth vector sets.

[0155] FIG. 7 is a ladder diagram for explaining an operation method of an artificial intelligence system according to an embodiment of the present disclosure.

[0156] Figure 4 is a diagram explaining the learning and fine tuning process of a multi-modal model, and Figure 7 is a diagram explaining the process of outputting situation inference using a fine-tuned multi-modal model.

[0157] The embodiment of Fig. 7 may occur after the embodiment of Fig. 4.

[0158] Referring to FIG. 7, the first edge device (100-1) can transmit an encrypted first modality to the home server (200-2) (S701), and the second edge device (100-2) can transmit an encrypted second modality to the home server (200-2) (S703).

[0159] The first modality may be data collected by the first edge device (100-1), and the second modality may be data collected by the second edge device (100-2). Each of the first and second modalities may be any one of image data, sensing data, sound data, or text data.

[0160] The processor (260) of the home server (200-2) can obtain the output value of the multi-modal model using the encrypted first modality and the encrypted second modality (S705).

[0161] The multi-modal model may be stored in the memory (230) of the home server (200-2).

[0162] The processor (260) of the home server (200-2) can extract a first feature vector set from the encrypted first modality and can extract a second feature vector set from the encrypted second modality.

[0163] Each of the first feature vector set and the second feature vector set can include 512 vectors.

[0164] The processor (260) of the home server (200-2) can input the first feature vector set and the second feature vector set into a multi-modal model learned through contrastive learning, thereby obtaining an output value of the multi-modal model.

[0165] The output values ​​of a multimodal model can be expressed as a set of 512 output vectors.

[0166] The processor (260) of the home server (200-2) can obtain a situation inference result matching the output value of the acquired multi-modal model (S707).

[0167] The output values ​​of a multimodal model may be matched to specific situations.

[0168] Each of the multiple output values ​​may be matched with a corresponding set of situational inference results. The situational inference results may be any of the following: the user's emotional state, the presence of a person in the home, an intrusion, or an emergency situation in the vehicle.

[0169] The processor (260) of the home server (200-2) can determine whether a response to the acquired situation inference result is possible (S709).

[0170] The processor (260) of the home server (200-2) can determine that a response is possible if action is possible for the situation inference result, and can determine that a response is impossible if action is impossible for the situation inference result.

[0171] For example, if the situation inference result is that the home is not present, control may be required to minimize the power consumption of home appliances installed in the home. The processor (260) of the home server (200-2) may determine that a response to the situation inference result is impossible if the home server (200-2) does not have the authority to control the home appliances, and may determine that a response is possible if the home server (200-2) has the authority to control the home appliances.

[0172] If the processor (260) of the home server (200-2) determines that a response to the acquired situation inference result is possible, it can transmit the response result to the situation inference result to the first edge device (100-1) or the second edge device (100-2) through the communication circuit (210) (S711, 713).

[0173] If the situation inference result indicates that the user is not home, the response may be a control command to minimize power consumption of home appliances. The control command may be to turn off the appliance or place it in standby mode.

[0174] The processor (260) of the home server (200-2) may transmit the response result for the situation inference result to a specific home appliance other than the first and second edge devices (100-1, 100-2).

[0175] If the processor (260) of the home server (200-2) determines that a response to the acquired situation inference result is impossible, the processor (260) can transmit the situation inference result to the cloud server (200-1) via the communication circuit (210) (S713).

[0176] The processor (260) of the home server (200-2) can transmit the situation inference result and the output value of the multi-modal model to the cloud server (200-1) through the communication circuit (210).

[0177] The processor (260) of the home server (200-2) may transmit the situation inference result and the encrypted first and second modalities to the cloud server (200-1) through the communication circuit (210).

[0178] The processor (260) of the cloud server (200-1) can transmit a response result to the situation inference result to the first edge device (100-1) or the second edge device (100-2) through the communication circuit (210) (S715, 717).

[0179] The processor (260) of the cloud server (200-1) can obtain a response result for the situation inference result.

[0180] The processor (260) of the cloud server (200-1) may also transmit the acquired response result to the home server (200-2). In this case, the home server (200-2) may transmit the response result received from the cloud server (200-1) to the first edge device (100-1), the second edge device (100-2), or a specific home appliance.

[0181] Meanwhile, according to another embodiment of the present disclosure, the first situation inference result of the multi-modal model of the home server (200-2) and the second situation inference result of the multi-modal model of the cloud server (200-1) may be different from each other.

[0182] The cloud server (200-1) can receive the encrypted first and second modalities from the home server (200-2) and the first situation inference result obtained by the home server (200-2).

[0183] The cloud server (200-1) can obtain the second situation inference result from the encrypted first and second modalities using the multi-modal model it stores.

[0184] If the first situation inference result and the second situation inference result are different, the cloud server (200-1) can transmit a notification indicating that the inference results are different to the home server (200-2) or the first and second edge devices (100-1, 100-2).

[0185] For example, the home server (200-2) may recognize the current situation as a normal situation, while the cloud server (200-2) may recognize the current situation as an emergency or critical situation. The cloud server (200-2) may transmit a notification indicating an emergency or critical situation to the home server (200-2) or the first and second edge devices (100-1, 100-2). Accordingly, even if the situation determined by the multi-modal model provided in the home server (200-2) is incorrect, notification of an emergency / critical situation can be effectively provided.

[0186] FIGS. 8A to 8C are diagrams illustrating a process of inferring a situation from multiple modalities using a multi-modal model learned according to embodiments of the present disclosure.

[0187] In FIGS. 8A to 8C, the multi-modal model (800) may be a global model stored in a cloud server (200-1) or a personalized model stored in a home server (200-2).

[0188] First, Figure 8a is described.

[0189] The first, second, and third edge devices (100-1, 100-2, and 100-3) may be provided in a single device. For example, a home robot may include the first, second, and third edge devices (100-1, 100-2, and 100-3).

[0190] The multi-modal model (800) can receive encrypted image data from a first edge device (100-1), encrypted sound data from a second edge device (100-2), and text data from a third edge device (100-3).

[0191] The first edge device (100-1) can obtain image data including a user's facial expression and encrypt the image data through an image encoder.

[0192] The second edge device (100-2) can obtain sound data representing the voice spoken by the user and encrypt the sound data through a sound encoder.

[0193] The third edge device (100-3) can obtain text data representing the content of a conversation spoken by a user. The text data is data that does not require privacy, and the third edge device (100-3) may not be equipped with an encoder for encrypting the text data.

[0194] Text data can be converted into a vector set through a text encoder provided in a cloud server (200-1) or home server (200-2).

[0195] The multi-modal model (800) can output a situation inference result from a first vector set that is encrypted image data, a second vector set that is encrypted voice data, and a third vector set corresponding to text data.

[0196] For example, the situational inference result may indicate the user's emotional state. Emotional states may include anger and happiness.

[0197] The multi-modal model (800) can place a first vector set, which is encrypted image data, a second vector set, which is encrypted voice data, and a third vector set, which is encrypted voice data, in an embedding space including a plurality of data points.

[0198] Each of the multiple data points can correspond to a specific emotional state.

[0199] The multi-modal model (800) can obtain the emotional state corresponding to the data point closest to the first, second, and third vector sets as a situation inference result.

[0200] The multi-modal model (800) can obtain the angry state as a situation inference result, as shown in Fig. 8a.

[0201] The home server (200-2) or cloud server (200-1) may, in response to an angry state, transmit a control command to the home robot requesting a soothing or empathetic conversation with the user. The home robot may output a soothing or empathetic voice according to the control command.

[0202] Next, Figure 8b is described.

[0203] The multi-modal model (800) can receive encrypted sound data from a first edge device (100-1), encrypted sensing data from a second edge device (100-2), and log data from a third edge device (100-3).

[0204] The first edge device (100-1) can acquire sound data generated within a home and encrypt the sound data through a sound encoder.

[0205] The second edge device (100-2) can obtain sensing data representing military waves and encrypt the sensing data through a sensor encoder.

[0206] The third edge device (100-3) is a home appliance that can collect log data from the appliance. The log data may include information about the appliance's operating time, operating status, and power status. The log data may be data that does not require privacy.

[0207] Log data can be converted into a vector set through an encoder provided in a cloud server (200-1) or home server (200-2).

[0208] The multi-modal model (800) can output a situation inference result from a first vector set that is encrypted sound data, a second vector set that is encrypted sensing data, and a third vector set corresponding to log data.

[0209] For example, the situational inference result may indicate whether the home is occupied, absent, or intruded by an outsider.

[0210] The multi-modal model (800) can place a first vector set, which is encrypted sound data, a second vector set, and a third vector set, which are encrypted sensing data, in an embedding space including a plurality of data points.

[0211] Each of the multiple data points can correspond to a presence state, an absence state, or an outsider intrusion state.

[0212] The multi-modal model (800) can obtain a state corresponding to the data point closest to the first, second, and third vector sets as a situation inference result.

[0213] The multi-modal model (800) can obtain the absence status as a situation inference result, as illustrated in FIG. 8b.

[0214] The home server (200-2) or cloud server (200-1) can, in response to the absence status, transmit control commands to home appliances to minimize power consumption of the home appliances.

[0215] If the multi-modal model (800) acquires an external intrusion status as a result of situation inference, the home server (200-2) or cloud server (200-1) can transmit a security notification to the user's terminal.

[0216] Next, Figure 8c is described.

[0217] The first, second, and third edge devices (100-1, 100-2, and 100-3) may be provided in a single device. For example, a vehicle may include the first, second, and third edge devices (100-1, 100-2, and 100-3).

[0218] The multi-modal model (800) can receive encrypted image data from a first edge device (100-1), encrypted voice data from a second edge device (100-2), and log data from a third edge device (100-3).

[0219] The first edge device (100-1) can obtain image data captured from inside a vehicle and encrypt the image data through an image encoder.

[0220] The second edge device (100-2) can obtain sound data within the vehicle and encrypt the sound data through a sound encoder.

[0221] The third edge device (100-3) can collect log data using the vehicle's processor. The log data may include the vehicle's door open / close status, vehicle speed, whether the vehicle is stationary, and vehicle operating status. The log data may be data that does not require privacy.

[0222] Log data can be converted into a vector set through an encoder provided in a cloud server (200-1) or home server (200-2).

[0223] The multi-modal model (800) can output a situation inference result from a first vector set that is encrypted image data, a second vector set that is encrypted sound data, and a third vector set corresponding to log data.

[0224] For example, the situation inference result may indicate an emergency condition in a locked vehicle or a normal condition in a locked vehicle.

[0225] The multi-modal model (800) can place a first vector set, which is encrypted image data, a second vector set, which is encrypted sound data, and a third vector set, which is encrypted sound data, in an embedding space including a plurality of data points.

[0226] Each of the multiple data points may represent a critical condition within a locked vehicle or a normal condition within a locked vehicle.

[0227] The multi-modal model (800) can obtain a state corresponding to the data point closest to the first, second, and third vector sets as a situation inference result.

[0228] The multi-modal model (800) can obtain an emergency situation inside a locked vehicle as a situation inference result, as illustrated in FIG. 8c.

[0229] The home server (200-2) or cloud server (200-1) can respond to an emergency situation in a locked vehicle by sending an emergency notification to the vehicle owner's terminal or unlocking the vehicle.

[0230] In this way, according to an embodiment of the present disclosure, an accurate situation can be inferred using various types of data, and actions appropriate to the inferred situation can be taken.

[0231] FIG. 9 is a diagram illustrating an example in which each of a plurality of home servers has a personalized multi-modal model according to one embodiment of the present disclosure.

[0232] Referring to FIG. 9, the cloud server (200-1) can store a global multi-modal model (900).

[0233] The global multi-modal model (900) can be transmitted to multiple home servers (200-2A, 200-2B).

[0234] The first home server (200-2A) can perform fine tuning on some layers of the global multi-modal model (900) and can be equipped with a fine-tuned first personalized model (900-1). This process may correspond to step S429 of FIG. 4.

[0235] The second home server (200-2B) can perform fine tuning on some layers of the global multi-modal model (900) and can be equipped with a fine-tuned second personalized model (900-2). This process may correspond to step S429 of FIG. 4.

[0236] That is, each home server can perform fine tuning on a global multi-model (900) using multiple modalities received from edge devices, and obtain a personalized multi-modal model according to the fine tuning performed.

[0237] Based on a personalized multi-modal model, user-tailored contextual inference and actions based on contextual inference can be effectively performed.

[0238] The above-described present disclosure can be implemented as computer-readable code on a program-recorded medium. The computer-readable medium includes all types of recording devices that store data that can be read by a computer system. Examples of computer-readable media include hard disk drives (HDDs), solid-state disk drives (SSDs), silicon disk drives (SDDs), read-only memory (ROM), random-access memory (RAM), CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices.

[0239] Additionally, the computer may include a processor (180) of an artificial intelligence device.

Claims

1. In the artificial intelligence server, A memory that stores a multi-modal model that infers situations using multiple modalities with different types; A communication circuit that communicates with multiple edge devices or cloud servers; and A processor including: receiving parameter information for a fine-tuned layer of the multi-modal model from the cloud server; receiving a plurality of encrypted modalities from the plurality of edge devices; and updating the parameter information of the fine-tuned layer of the multi-modal model using the received plurality of encrypted modalities. Artificial intelligence server.

2. In paragraph 1, The above processor Receive additional encrypted modalities, and obtain a first situation inference result from the additionally received encrypted modalities using the updated multi-modal model. Artificial intelligence server.

3. In paragraph 2, The above processor Transmitting a plurality of encrypted modalities and the first situation inference result to the cloud server through the above communication circuit. Artificial intelligence server.

4. In paragraph 3, The above processor If the second situation inference result obtained by the cloud server is different from the first situation inference result, a notification is received from the cloud server. Artificial intelligence server.

5. In paragraph 2, The above processor It is determined whether a response to the above first situation inference result is possible, and based on determining that a response is possible, the response result is transmitted to the plurality of edge devices through the communication circuit, and based on determining that a response is impossible, the situation inference result is transmitted to the cloud server. Artificial intelligence server.

6. In paragraph 5, The above processor In response to transmitting the situation inference result from the above cloud server to the above cloud server, a response result is received. Artificial intelligence server.

7. In paragraph 1, Each of the above multiple modalities Any one of image data, sound data, sensing data, text data, or log data. Artificial intelligence server.

8. In paragraph 1, The above multi-modal model A model based on an artificial neural network trained through contrastive learning. Artificial intelligence server.

9. In the method of operation of the artificial intelligence server, A step of receiving parameter information for a fine-tuned layer of a multi-modal model that infers a situation by using multiple modalities having different types from a cloud server; A step of receiving a plurality of encrypted modalities from the plurality of edge devices; and A step of updating the parameter information of the fine-tuned layer of the multi-modal model using the received multiple encrypted modalities. How the AI ​​server works.

10. In paragraph 9, Further comprising a step of additionally receiving a plurality of encrypted modalities and obtaining a first situation inference result from the additionally received plurality of encrypted modalities using the updated multi-modal model. How the AI ​​server works.

11. In paragraph 10, Further comprising a step of transmitting a plurality of encrypted modalities and the first situation inference result to the cloud server through the communication circuit. How the AI ​​server works.

12. In paragraph 11, If the second situation inference result obtained by the cloud server is different from the first situation inference result, the method further includes receiving a notification from the cloud server. How the AI ​​server works.

13. In paragraph 10, A step of determining whether a response to the above first situation inference result is possible; A step of transmitting the response result to the plurality of edge devices through the communication circuit based on determining that a response is possible; and Further comprising a step of transmitting the situation inference result to the cloud server based on determining that the response is impossible. How the AI ​​server works.

14. In paragraph 13, Further comprising a step of receiving a response result in response to transmitting the situation inference result from the cloud server to the cloud server. How the AI ​​server works.

15. In paragraph 9, Each of the above multiple modalities Any one of image data, sound data, sensing data, text data, or log data. How the AI ​​server works.

Citation Information

Patent Citations

  • Federal learning method and system for sample sparsity

    CN113128701A

  • Notification service module method to easily access the management know-how of camper vehicle

    KR1020220107640A

  • Fish-shaped bread vending machine

    KR1020220114831A

  • Method and apparatus for split rendering architecture to support real-time media delivery in a mobile communication system

    KR1020230155740A

  • Odor control system of drying machines

    KR102524193B1