Electronic device for training artificial intelligence model, and control method therefor
By using a loss function to minimize vector differences during encoder retraining, the method addresses the inefficiencies in updating AI models, reducing time and resources while maintaining decoder performance.
Patent Information
- Application Number
- PCT/KR2024/020220
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-04
- Filing Date
- 2024-12-10
- Publication Date
- 2025-07-10
AI Technical Summary
Training an artificial intelligence model requires significant time and resources when updating learning data, particularly due to the need to retrain both the encoder and decoder components, which can disrupt the performance of the decoder and increase the time and resources required.
The method involves retraining the encoder using a loss function that minimizes the difference between pre- and post-retraining feature vectors and decoder outputs, allowing the encoder to be updated independently while maintaining decoder performance.
This approach significantly reduces the time and resources needed for retraining, ensuring efficient adaptation of the AI model to new learning data without compromising decoder performance.
Smart Images

Figure KR2024020220_10072025_PF_FP_ABST
Abstract
Description
Electronic device for training artificial intelligence model and control method thereof
[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more particularly, to an electronic device for training an artificial intelligence model and a method for controlling the same.
[0002] An artificial intelligence (AI) system is a system in which machines learn (train) on their own, make judgments, and produce desired results or perform desired actions.
[0003] AI technology consists of machine learning, such as deep learning, and elements that utilize machine learning. AI technology is widely used in fields such as linguistic understanding, visual understanding, inference / prediction, knowledge representation, and motion control.
[0004] For example, AI technology can be utilized in the fields of visual understanding and inference / prediction. Specifically, AI technology can be used to implement technologies for analyzing and classifying input data. In other words, methods and devices can be implemented to analyze and / or classify input data to obtain desired results.
[0005] Here, when the AI model generates output data corresponding to the input data, the accuracy of the output data may vary depending on the learning data.
[0006] At this time, there is a problem that it takes a lot of time and resources to train the entire neural network of the AI model as the learning data or learning algorithm is updated.
[0007] A method for training an artificial intelligence model including an encoder trained based on first training data according to one or more embodiments of the present disclosure includes the steps of obtaining second training data and retraining the encoder based on the second training data, wherein the step of retraining the encoder retrains the encoder such that, in an embedding space, a distance between a first feature vector when the second training data is input to the encoder before the retraining and a second feature vector when the second training data is input to the encoder during the retraining becomes less than a preset distance.
[0008] The step of retraining the encoder may retrain the encoder using a loss function including a first difference value between a first feature vector when the second learning data is input to the encoder before the retraining and a second feature vector when the second learning data is input to the encoder during the retraining.
[0009] The above artificial intelligence model further includes a decoder, and the step of retraining the encoder may retrain the encoder using a loss function including the first difference value and a second difference value between an output when the second learning data is input to the decoder and a target output of the decoder.
[0010] The method may further include a step of normalizing the first difference value and the second difference value to obtain the loss function.
[0011] The method may further include a step of obtaining the loss function by adding a product of a first weight value and the first feature vector and a product of a second weight value and the second feature vector.
[0012] The sum of the first weight value and the second weight value may be a preset value.
[0013] The above control method may include a step of obtaining input data, a step of inputting the input data to the relearned encoder, a step of inputting an output of the relearned encoder to the decoder, and a step of providing a service corresponding to the input data using the output of the decoder.
[0014] The above second learning data may be updated data of the first learning data.
[0015] An electronic device for training an artificial intelligence model including an encoder trained based on first learning data according to one or more embodiments of the present disclosure includes a memory and a processor, wherein the processor obtains second learning data, retrains the encoder based on the second learning data, and the processor retrains the encoder such that, in an embedding space, a distance between a first feature vector when the second learning data is input to the encoder before the retraining and a second feature vector when the second learning data is input to the encoder during the retraining becomes less than a preset distance.
[0016] The processor may retrain the encoder using a loss function including a first difference value between a first feature vector when the second learning data is input to the encoder before the retraining and a second feature vector when the second learning data is input to the encoder during the retraining.
[0017] The artificial intelligence model further includes a decoder, and the processor can retrain the encoder using a loss function including the first difference value and a second difference value between an output of the decoder when the second learning data is input and a target output of the decoder.
[0018] The above processor can obtain the loss function by normalizing the first difference value and the second difference value.
[0019] The processor can obtain the loss function by adding the product of the first weight value and the first feature vector and the product of the second weight value and the second feature vector.
[0020] The sum of the first weight value and the second weight value may be a preset value.
[0021] The processor may obtain input data, input the input data to the retrained encoder, input the output of the retrained encoder to the decoder, and provide a service corresponding to the input data using the output of the decoder.
[0022] The above second learning data may be updated data of the first learning data.
[0023] A non-transitory computer-readable recording medium including a program for executing a control method of an electronic device for training an artificial intelligence model including an encoder trained based on first training data according to one or more embodiments of the present disclosure, the control method including the steps of obtaining second training data and retraining the encoder based on the second training data, wherein the step of retraining the encoder retrains the encoder such that, in an embedding space, a distance between a first feature vector when the second training data is input to the encoder before the retraining and a second feature vector when the second training data is input to the encoder during the retraining becomes less than a preset distance.
[0024] The step of retraining the encoder may retrain the encoder using a loss function including a first difference value between a first feature vector when the second learning data is input to the encoder before the retraining and a second feature vector when the second learning data is input to the encoder during the retraining.
[0025] The above artificial intelligence model further includes a decoder, and the step of retraining the encoder may retrain the encoder using a loss function including the first difference value and a second difference value between an output when the second learning data is input to the decoder and a target output of the decoder.
[0026] The method may include a step of obtaining the loss function by normalizing the first difference value and the second difference value.
[0027] FIG. 1 is a block diagram illustrating a configuration of an electronic device according to one or more embodiments of the present disclosure.
[0028] FIG. 2 is a diagram illustrating the operation of a plurality of modules and an artificial intelligence model according to one or more embodiments of the present disclosure.
[0029] FIG. 3 is a diagram illustrating the operation of a plurality of modules and an artificial intelligence model according to one or more embodiments of the present disclosure.
[0030] FIG. 4 is a flowchart illustrating a method for an electronic device to retrain an encoder according to one or more embodiments of the present disclosure.
[0031] FIG. 5 is a diagram illustrating a feature vector output by an encoder according to one or more embodiments of the present disclosure in an embedding space.
[0032] FIG. 6 is a flowchart illustrating a method for an electronic device to obtain a loss function according to one or more embodiments of the present disclosure.
[0033] FIG. 7 is a flowchart illustrating a method for an electronic device according to one or more embodiments of the present disclosure to acquire a learning algorithm and retrain an encoder.
[0034] The present embodiments may be modified and have various embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope to specific embodiments, but should be understood to encompass various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0035] In describing the present disclosure, if it is determined that a specific description of a related known function or configuration may unnecessarily obscure the gist of the present disclosure, a detailed description thereof will be omitted.
[0036] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concepts of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to further faithfully and completely convey the technical concepts of the present disclosure to those skilled in the art.
[0037] The terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the scope of the rights. Singular expressions include plural expressions unless the context clearly dictates otherwise.
[0038] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.
[0039] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0040] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0041] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that said component may be directly coupled to said other component, or may be coupled via another component (e.g., a third component).
[0042] On the other hand, when it is said that a component (e.g., a first component) is "directly connected" or "directly connected" to another component (e.g., a second component), it can be understood that no other component (e.g., a third component) exists between said component and said other component.
[0043] The expression "configured to" used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.
[0044] Instead, in some contexts, the phrase "a device configured to" may mean that the device is "capable of" doing something in conjunction with other devices or components. For example, the phrase "a processor configured (or set) to perform A, B, and C" may mean a dedicated processor (130) for performing the actions, or a general-purpose processor (130) (e.g., a CPU or application processor) that can perform the actions by executing one or more software programs stored in a memory device.
[0045] In the embodiments, a 'module' or 'part' performs at least one function or operation, and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, a plurality of 'modules' or 'parts' may be integrated into at least one module and implemented as at least one processor, except for a 'module' or 'part' that needs to be implemented as a specific hardware.
[0046] Meanwhile, the various elements and areas in the drawings are schematically drawn. Therefore, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0047]
[0048] *41 Below, with reference to the attached drawings, embodiments according to the present disclosure are described in detail so that a person having ordinary knowledge in the technical field to which the present disclosure pertains can easily carry out the present disclosure.
[0049] FIG. 1 is a block diagram illustrating a configuration of an electronic device according to one or more embodiments of the present disclosure.
[0050] The electronic device (100) may include at least one of a memory (110), a communication interface (120), and a processor (130). The electronic device (100) may further include other components in addition to the above components.
[0051] The electronic device (100) may be implemented as a server, but this is only one or more embodiments, and the electronic device (100) may be implemented in various forms, such as a smartphone, a TV, a smart TV, a set-top box, a mobile phone, a PDA (personal digital assistant), a laptop, a media player, an e-book reader, a digital broadcasting terminal, a navigation device, a kiosk, an MP3 player, a wearable device, a home appliance, and other mobile or non-mobile computing devices.
[0052] The memory (110) can store at least one instruction regarding the electronic device (100). The memory (110) can store an operating system (O / S) for driving the electronic device (100). In addition, the memory (110) can store various software programs or applications for operating the electronic device (100) according to various embodiments of the present disclosure. In addition, the memory (110) can include a semiconductor memory such as a flash memory (110) or a magnetic storage medium such as a hard disk.
[0053] Specifically, the memory (110) can store various software modules for operating the electronic device (100) according to various embodiments of the present disclosure, and the processor (130) can control the operation of the electronic device (100) by executing various software modules stored in the memory (110). That is, the memory (110) is accessed by the processor (130), and data reading / recording / modifying / deleting / updating, etc. can be performed by the processor (130).
[0054] Meanwhile, in the present disclosure, the term memory (110) may be used to mean a memory (110), a ROM (not shown), a RAM (not shown) in a processor (130), or a memory card (not shown) (e.g., a micro SD card, a memory stick) mounted on an electronic device (100).
[0055] And, the communication interface (120) includes a circuitry and is a configuration capable of communicating with external devices and servers. The communication interface (120) can communicate with external devices or servers based on a wired or wireless communication method. The communication interface (120) may include a Bluetooth module (not shown), a Wi-Fi module (not shown), an IR (infrared) module, a LAN (Local Area Network) module, an Ethernet module, etc. Here, each communication module may be implemented in the form of at least one hardware chip. In addition to the above-described communication method, the wireless communication module may include at least one communication chip that performs communication according to various wireless communication standards such as Zigbee, USB (Universal Serial Bus), MIPI CSI (Mobile Industry Processor Interface Camera Serial Interface), 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), etc. However, this is only one embodiment, and the communication interface (120) can utilize at least one communication module among various communication modules.
[0056] The processor (130) can control the overall operation and function of the electronic device (100). Specifically, the processor (130) is connected to the configuration of the electronic device (100) including the memory (110), and can control the overall operation of the electronic device (100) by executing at least one command stored in the memory (110) as described above.
[0057] The processor (130) may be implemented in various ways. For example, the processor (130) may be implemented as at least one of an application specific integrated circuit (ASIC), a logic integrated circuit, an embedded processor, a microcomputer (Micom), a microprocessor, hardware control logic, a hardware finite state machine (FSM), and a digital signal processor (130).
[0058] In particular, the processor (130) may include one or more processors. Specifically, the one or more processors may include one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), a Neural Processing Unit (NPU), a Main Processing Unit (MPU), a hardware accelerator, or a machine learning accelerator. The one or more processors may control one or any combination of other components of the electronic device, and may perform operations related to communication or data processing. The one or more processors may execute one or more programs or instructions stored in a memory. For example, the one or more processors may perform a method according to an embodiment of the present disclosure by executing one or more instructions stored in a memory.
[0059] When a method according to one or more embodiments of the present disclosure includes multiple operations, the multiple operations may be performed by one processor or by multiple processors. That is, when a first operation, a second operation, and a third operation are performed by a method according to one or more embodiments, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (130) and the third operation may be performed by the second processor (130).
[0060] One or more processors may be implemented as a single-core processor (130) including one core, or may be implemented as one or more multi-core processors (130) including multiple cores (e.g., homogeneous multi-cores or heterogeneous multi-cores). When one or more processors are implemented as a multi-core processor, each of the multiple cores included in the multi-core processor may include internal processor memory, such as cache memory or on-chip memory, and a common cache shared by the multiple cores may be included in the multi-core processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multi-core processor may independently read and execute a program instruction for implementing a method according to one or more embodiments of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to one or more embodiments of the present disclosure.
[0061] When a method according to one or more embodiments of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one core among the plurality of cores included in a multi-core processor, or may be performed by the plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to one or more embodiments, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.
[0062] In embodiments of the present disclosure, the processor (130) may mean a system on a chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but embodiments of the present disclosure are not limited thereto.
[0063] The operation of the processor (130) for implementing various embodiments of the present disclosure can be implemented through an artificial intelligence model and multiple modules.
[0064] Specifically, data for an artificial intelligence model and a plurality of modules according to the present disclosure can be stored in a memory (110), and a processor (130) can access the memory (110) to load data for a plurality of modules into a memory or buffer within the processor (130), and then implement various embodiments according to the present disclosure using the plurality of modules.
[0065] However, at least one of the plurality of modules according to the present disclosure may be implemented in hardware and included in the processor (130) in the form of a system on chip.
[0066] Alternatively, at least one of the plurality of modules according to the present disclosure may be implemented as a separate external device, and the electronic device (100) and each module may communicate and perform operations according to the present disclosure.
[0067] Hereinafter, with reference to the attached drawings, the operation of a processor (130) using multiple modules according to the present disclosure will be described in detail.
[0068] FIG. 2 is a diagram illustrating the operation of a plurality of modules and an artificial intelligence model according to one or more embodiments of the present disclosure.
[0069] Referring to FIG. 2, the data acquisition module (111) can acquire data. The data acquisition module (111) can receive data from an external device or load data stored in the memory (110).
[0070] At this time, the acquired data may be input data for inputting into an artificial intelligence model (112). Alternatively, the acquired data may be learning data for (re)training the artificial intelligence model (112).
[0071] Here, the data may include at least one of user profile information (10), user service usage information (20), and service item information (30), but is not limited thereto.
[0072] Specifically, the data acquisition module (111) can receive user profile information from a user terminal device through a communication interface (120).
[0073] Here, user profile information may include at least one of the following: user personal information, user preference information, and user log records. For example, user profile information may include at least one of the following: user gender, user age, user's residential area, user's occupation, user's preferred music, user's preferred food, user's movement history, user's visited locations, and user terminal autonomy log records.
[0074] Alternatively, the data acquisition module (111) can receive the user's service usage information (20) from the service providing server through the communication interface (120).
[0075] Here, the user's service usage information (20) may include the user's service usage history. For example, the user's service usage history may include at least one of the following: the website the user accessed, the search terms the user searched, the number of times the user used the service, the time the user used the service, the location the user used the service, and the log record of the user's service usage.
[0076] Alternatively, the data acquisition module (111) can receive service information (30) from the service providing server through the communication interface (120).
[0077] Here, service item information (30) may include information about items provided by the service. In the present disclosure, "item" may be replaced with expressions representing identical / similar concepts, such as "content," "product," and "place."
[0078] For example, information about an item provided in a service may include at least one of identification information about an item provided in a service, information about a list of items provided in a service, information about a description of an item provided in a service, information about a review of an item provided in a service, and information about a star rating of an item provided in a service. According to one or more embodiments of the present disclosure, the electronic device (100) may also load user profile information or service usage records stored in the memory (110).
[0079] When data is acquired, the data acquisition module (111) can store the received data in the data DB (40) of the memory (110).
[0080] Meanwhile, the data acquisition module (111) may receive data in the form of raw data or in the form of preprocessed data (e.g., data embedding vector). Here, data preprocessing may mean data encryption or data feature extraction.
[0081] According to one or more embodiments of the present disclosure, when the data acquisition module (111) receives the first embedding vector, the electronic device (100) can input the acquired first embedding vector into the encoder (112a) to acquire a second embedding vector corresponding to the features of the first embedding vector.
[0082] When data is received in the form of raw data, the data acquisition module (111) can preprocess the received data and store the preprocessed data in the data DB (40).
[0083] The artificial intelligence model (112) of the present disclosure may include an encoder (112a) and a decoder (112b).
[0084] The encoder (112a) is a configuration for extracting a feature vector of data input to the encoder. Specifically, the encoder (112a) may refer to a neural network trained to output a feature vector of input data when data is input.
[0085] In the present disclosure, “encoder” may be replaced with expressions representing the same / similar concepts, such as “embedding model,” “embedding layer,” “embedding neural network,” “feature output model,” “feature output layer,” or “feature output neural network.”
[0086] In the present disclosure, the “feature vector” of the input data may be replaced with an expression representing the same / similar concept, such as “feature”, “embedding”, or “embedding vector”.
[0087] In the present disclosure, the decoder (112b) is configured to output data of a predetermined type when input data features are input. Specifically, the decoder (112b) may refer to a neural network trained to output data of a predetermined type when a feature vector output from an encoder is input. In other words, the decoder can reconstruct the feature vector output by the encoder into data of a desired type.
[0088] Here, the preset type of data may be, but is not limited to, data for classifying a class of input data, data for identifying recommended content corresponding to the input data, or data for identifying an advertisement corresponding to the input data.
[0089] For example, when a feature vector of user profile information (10) is input to a decoder (112b), the decoder (112b) can output data on recommended movies optimized for the user.
[0090] Referring again to FIG. 2, the processor (130) can input data acquired by the data acquisition module (111) or data stored in the data DB into the encoder (112a).
[0091] Accordingly, the encoder (112a) can output a feature vector (50) of the input data.
[0092] At this time, the processor (130) can store the output feature vector (50) in the feature vector DB (60) stored in the memory (110).
[0093] Meanwhile, among the above-described contents, the operation of storing data in a data DB or storing a feature vector (50) in a feature vector DB may be omitted.
[0094] And, the processor (130) can input the feature vector (50) output by the encoder (112a) or the feature vector stored in the feature vector DB to the decoder (112b).
[0095] Accordingly, the decoder (112b) can output data of a preset type.
[0096] In addition, the service provision module (113) can provide a preset service using data output by the decoder (112b).
[0097] For example, when the decoder (112b) outputs data about personalized recommended content to the user, the service providing module (113) can provide information about the recommended content to the user using the data output from the decoder (112b).
[0098] FIG. 3 is a diagram illustrating the operation of multiple modules and an artificial intelligence model according to one or more embodiments of the present disclosure. Since the components, excluding the decoder and service provision module, have been described with reference to FIG. 2, any redundant description will be omitted.
[0099] Referring to FIG. 3, the artificial intelligence model (112) may include multiple decoders (310). At this time, the multiple decoders (310) may share the feature vector output by the encoder (112a). At this time, each of the multiple decoders (310) may output data of a preset type.
[0100] In addition, the memory (110) may include a plurality of service providing modules (320) corresponding to each of a plurality of decoders (310).
[0101] At this time, the first service providing module can provide the first service using data output from the first decoder. Furthermore, the second service providing module can provide a second service different from the first service using data output from the second decoder.
[0102] For example, the first decoder may output data regarding recommended content corresponding to the user's profile information and service usage information. Furthermore, the second decoder may output data regarding recommended advertisements corresponding to the user's profile information and service usage information.
[0103] At this time, the first service provision module can provide the user with information about recommended content using the data output from the first decoder. Furthermore, the second service provision module can provide the user with information about recommended advertisements using the data output from the second decoder.
[0104] Meanwhile, each component of the above-described plurality of modules and artificial intelligence models (112) may be stored in the electronic device (100), but this is only one embodiment, and at least one of each component of the plurality of modules and artificial intelligence models may be implemented as a separate external device, and the electronic device (100) and the external device may transmit and receive data, perform communication, and perform operations according to the present disclosure.
[0105] According to one or more embodiments of the present disclosure, the electronic device (100) may include a data acquisition module and an encoder. In addition, the decoder and the service provision module may be stored in the user terminal device.
[0106] In this case, the electronic device (100) transmits the feature vector output through the encoder to the user terminal device, and the user terminal device inputs the received feature vector into the decoder and inputs data output from the decoder into the service providing module to provide a service to the user.
[0107] According to one or more embodiments of the present disclosure, the electronic device (100) may include a data acquisition module, an encoder, and a decoder, and the service provision module may be stored in the user terminal device.
[0108] In this case, the electronic device (100) transmits data output through the decoder to the user terminal device, and the user terminal device inputs the received data into the service provision module to provide a service to the user.
[0109] According to one or more embodiments of the present disclosure, the artificial intelligence model of the present disclosure may be a model learned based on first learning data. Specifically, the encoder (112a) and decoder (112b) included in the artificial intelligence model (112) may be models learned based on first learning data.
[0110] Meanwhile, the encoder needs to be retrained periodically due to changes in training data or changes in the encoder learning algorithm.
[0111] When an encoder is retrained, the problem of the data output by the encoder changing significantly for identical / similar input data may arise.
[0112] For example, for the profile information of user A, the encoder before retraining can output a feature vector such as "(1, 0, 1, 1, 0, 0, 1, 0)". And, for the profile information of user A, the encoder after retraining can output a feature vector such as (0, 1, 1, 1, 0, 1, 0, 0).
[0113] Therefore, when retraining only the encoder excluding the decoder, the decoder that receives the feature vector output by the encoder may have a problem in that it cannot maintain the performance of the decoder before the encoder was retrained.
[0114] To avoid these problems, the decoder must be retrained along with the encoder, but this increases the time and resources required for retraining.
[0115] Therefore, when training an artificial intelligence model (112) based on new training data, there is an inconvenience of having to train the decoder together with the encoder.
[0116] That is, the meaning of the feature vector output by the retrained encoder is different from the meaning of the feature vector output by the encoder before retraining, so there is an inconvenience of having to retrain the decoder to match the meaning of the changed feature vector.
[0117] In particular, as described above with reference to FIG. 3, the AI model of the present disclosure may include multiple decoders that share the output of an encoder. In this case, since each of the multiple decoders must be retrained along with the encoder, the time and resources required for retraining increase further.
[0118] The electronic device (100) of the present disclosure can retrain only the encoder while maintaining the performance of the decoder. Accordingly, there is a technical effect of significantly reducing the time and resources required for retraining.
[0119] Specifically, the electronic device (100) of the present disclosure can obtain a loss function including a value indicating compatibility between an encoder before retraining and an encoder during retraining, and retrain the encoder so that the value of the obtained loss function is minimized. In the present disclosure, "compatibility" can be replaced with expressions indicating identical / similar concepts, such as "fitness" and "similarity."
[0120] The detailed operation of the electronic device (100) is described with reference to the drawings below.
[0121] FIG. 4 is a flowchart illustrating a method for an electronic device to retrain an encoder according to one or more embodiments of the present disclosure.
[0122] Referring to FIG. 4, the electronic device (100) can obtain second learning data (S410).
[0123] In the present disclosure, learning data may be, but is not limited to, user profile information containing personalized information for the user. In this case, the electronic device (100) may receive the user profile information from the user terminal device via the communication interface (120).
[0124] In the present disclosure, the second learning data may be updated data from the first learning data. The second learning data may be data in which some attributes of the first learning data have been transformed. Alternatively, the second learning data may be data in which additional data has been combined with the first learning data. Alternatively, the second learning data may be data acquired later than the first learning data.
[0125] For example, if the learning data is user profile information, the first learning data may include data about the user's gender and age. The second learning data may include data about the user's gender, age, and region of residence.
[0126] For example, if the learning data is a user's service usage log record, the first learning data may be a user's log record acquired during a first period, and the second learning data may be a user's log record acquired during a second period. In this case, the second period may be a period after the first period. Alternatively, the second period may be a period that includes the first period and the period after the first period.
[0127] In the present disclosure, the first learning data may refer to learning data used to train an artificial intelligence model (112). That is, the encoder and decoder constituting the artificial intelligence model in the present disclosure may be models trained based on the first learning data.
[0128] Based on the second learning data, the electronic device (100) can retrain the encoder.
[0129] Specifically, when the second learning data is acquired, the electronic device (100) can acquire a loss function for retraining the encoder (S320).
[0130] In this disclosure, "retraining" may refer to an operation of retraining an encoder, previously trained based on first training data, based on second training data. Retraining may be performed to improve the performance of the encoder or optimize the encoder for the second training data. In this disclosure, "retraining" may be replaced with an expression of the same or similar concept, such as "fine-tuning."
[0131] The electronic device (100) of the present disclosure can retrain the encoder so that the distance between the first feature vector when the second learning data is input to the encoder before retraining and the second feature vector when the second learning data is input to the encoder during retraining is less than a preset distance in the embedding space.
[0132] Specifically, the encoder can output an embedding vector that can be mapped to the embedding space for each training data. At this time, the output embedding vector can correspond to the input training data. In addition, the electronic device (100) can retrain the encoder so that the distribution of the embedding vector output by the encoder before retraining and the distribution of the embedding vector output by the retrained encoder are similar in the embedding space.
[0133] FIG. 5 is a diagram illustrating a feature vector output by an encoder according to one or more embodiments of the present disclosure in an embedding space.
[0134] Referring to FIG. 5, for specific input data, the distance (531, 532, 533) between the feature vector (511, 512, 513) output by the encoder before retraining and the feature vector (521, 522, 523) output by the retrained encoder may be less than a preset distance.
[0135] Accordingly, the output of the encoder before retraining and the output of the retrained encoder can maintain the semantic structure.
[0136] Specifically, the electronic device (100) of the present disclosure can obtain a loss function including a value indicating compatibility between an encoder before retraining and an encoder during retraining, and retrain the encoder so that the value of the obtained loss function is minimized. In the present disclosure, "compatibility" can be replaced with expressions indicating identical / similar concepts, such as "fitness" and "similarity."
[0137] In the present disclosure, an encoder undergoing retraining may mean an encoder that has progressed through n epochs while retraining the encoder. Here, n may be any natural number.
[0138] According to one or more embodiments of the present disclosure, the electronic device (100) can obtain a loss function including a first difference value between a first feature vector when second learning data is input to an encoder before relearning and a second feature vector when second learning data is input to an encoder during relearning.
[0139] At this time, the loss function can be expressed as in mathematical equation 1 below.
[0140] Mathematical formula 1
[0141] Loss function = a (V old - V new ) = a (first difference value)
[0142] Here, V old can mean the first feature vector when the second training data is input to the encoder before retraining. And, V new may refer to the second feature vector when the second learning data is input to the encoder being retrained.
[0143] Here, a may represent the first weight. a may be any real number. In this case, a may be a weight for normalizing the first difference value.
[0144] According to one or more embodiments of the present disclosure, the electronic device (100) may obtain a loss function including a first difference value and a second difference value between an output value of an encoder before retraining when the output of the encoder is input to the decoder and a target output value of the decoder. Here, the output of the encoder before retraining may be an output when second training data is input to the encoder before retraining. And, the target output value of the decoder may be a target value labeled with the second training data.
[0145] In this case, the loss function can be expressed as in mathematical equation 2 below.
[0146] Mathematical formula 2
[0147] Loss function = a (V old - V new ) + b (Y - Y') = a (first difference value) + b (second difference value)
[0148] Here, Y may refer to the target output value of the decoder. And, Y' may refer to the output value of the decoder when the output of the encoder being retrained is input.
[0149] Here, a may represent the first weight, and b may represent the second weight. At this time, a and b may be any real numbers.
[0150] At this time, a and b may be weights for normalizing the first difference value and the second difference value.
[0151] That is, the electronic device (100) can obtain a loss function by normalizing the first difference value and the second difference value.
[0152] Specifically, that is, each of the first weight and the second weight may be a weight for normalizing the first difference value and the second difference value.
[0153] Referring to the mathematical expression 2 above, the electronic device (100) can obtain a loss function by adding the product of the first weight value and the first feature vector and the product of the second weight value and the second feature vector.
[0154] Also, the sum of the first weight value and the second weight value may be a preset value. For example, the first weight value and the weight value b may be 1. In this case, the second weight value b may be expressed in the form (1-a).
[0155] According to one or more embodiments of the present disclosure, the artificial intelligence model may include a plurality of decoders, wherein the output of the encoder may be input to each of the plurality of decoders.
[0156] At this time, the electronic device can obtain a loss function including a difference value between the first difference value and the result value when the output of the encoder before relearning is input to each of the plurality of decoders and the target output value of each of the plurality of decoders.
[0157] In this case, the loss function can be expressed as in mathematical equation 3 below.
[0158] Mathematical formula 3
[0159] Loss function = a (V old - V new ) +
[0160] The mathematical expression 3 above can represent the loss function when there are k decoders.
[0161] Here, Y k can mean the target output value of the kth decoder. And, Y' k can mean the output value of the decoder when the output of the encoder being retrained is input. W k may be a weight to normalize the kth decoder.
[0162] Specifically, the electronic device (100) can obtain a loss function by normalizing each difference value that constitutes the loss function.
[0163] At this time, the sum of multiple weights may be a preset value.
[0164] Then, the electronic device (100) can retrain the encoder using the acquired loss function (S330). At this time, the electronic device (100) can retrain the encoder so that the value of the acquired loss function approaches 0 (converges). Alternatively, the electronic device (100) can retrain the encoder so that the value of the acquired loss function approaches a preset value.
[0165] Meanwhile, according to one or more embodiments of the present disclosure, the electronic device may obtain a loss function differently depending on whether a pre-learned encoder exists.
[0166] FIG. 6 is a flowchart illustrating a method for an electronic device to obtain a loss function according to one or more embodiments of the present disclosure.
[0167] Referring to FIG. 6, the electronic device (100) can acquire second learning data (S610). At this time, the operation of S410 may be the same as the operation of S310 described above, and any duplicate description will be omitted.
[0168] When the second learning data is acquired, the electronic device (100) can identify whether a pre-learned encoder exists (S620).
[0169] If a pre-trained encoder exists (S620-Y), the electronic device (100) can obtain a first loss function for retraining the encoder (S630). At this time, the operation of S630 may be the same as the operation of S420 described above, and a duplicate description will be omitted. That is, the first loss function in step S630 may be one of the loss functions described above in operation S420.
[0170] If a pre-trained encoder does not exist (S620-N), the electronic device (100) can obtain a second loss function for training the encoder (S640). At this time, the second loss function may include the difference between the output value of the decoder when the output of the encoder into which the second training data is input is input to the decoder and the target output value of the decoder.
[0171] At this time, the second loss function can be expressed according to the mathematical formula 4 below.
[0172] Mathematical formula 4
[0173] Loss function = c (Y - Y')
[0174] Here, Y can mean the target output value of the decoder. And, Y' can mean the output value of the decoder when the output of the encoder being trained is input.
[0175] Here, c can be a weight to normalize the difference between the target output value of the decoder and the output value of the decoder.
[0176] Meanwhile, according to the above, the operation of S620 can be replaced with an operation of identifying whether the learned feature vector is stored in the feature vector DB (60). In this case, if the feature vector is stored in the feature vector DB (60), the electronic device (100) can perform the operation of S630. In addition, if the feature vector is not stored in the feature vector DB (60), the electronic device (100) can perform the operation of S640.
[0177] Once the retraining of the encoder is completed, the electronic device (100) can perform a preset operation using an artificial intelligence model including the retrained encoder.
[0178] According to one or more embodiments of the present disclosure, an electronic device (100) may acquire input data, input the input data to a retrained encoder, input the output of the retrained encoder to a decoder, and provide a service corresponding to the input data using the output of the decoder. Here, the input data may be user profile information or a user's service usage history.
[0179] As described above, the electronic device (100) can obtain second learning data and retrain the encoder based on the second learning data, but is not limited thereto.
[0180] FIG. 7 is a flowchart illustrating a method for an electronic device according to one or more embodiments of the present disclosure to acquire a learning algorithm and retrain an encoder.
[0181] According to one or more embodiments of the present disclosure, the electronic device (100) can obtain a learning algorithm for retraining an encoder, and retrain the encoder based on the obtained learning algorithm.
[0182] Specifically, referring to FIG. 7, the electronic device (100) can obtain a second learning algorithm for retraining the encoder (S710).
[0183] *175 In the present disclosure, the artificial intelligence model (112) may be a model learned based on the first learning algorithm. That is, the encoder or decoder constituting the artificial intelligence model (112) may be a neural network acquired based on the first learning algorithm. In this case, the learning data learned by the encoder before learning and the learning data for retraining the encoder may be the same, but are not limited thereto.
[0184] The electronic device (100) can obtain a loss function based on the acquired second learning algorithm (S720).
[0185] At this time, the electronic device (100) can obtain a loss function including a difference value between the feature vector output by the encoder being trained with the second learning algorithm obtained when the feature learning data was input and the feature vector output by the encoder before re-learning when the learning data was input.
[0186] At this time, the loss function can be obtained in the form described above through Fig. 4, and redundant descriptions are omitted.
[0187] In addition, the electronic device (100) can retrain the encoder using the acquired loss function (S730). At this time, the operation of S730 may be the same as the operation of S430 or S650.
[0188] Meanwhile, in the process of acquiring the loss function, the operations of S720 and S420 can be combined.
[0189] Specifically, the electronic device (100) can obtain second learning data and a second learning algorithm.
[0190] At this time, the artificial intelligence model (112) may be a pre-learned model based on the first learning data and the first learning algorithm.
[0191] At this time, the electronic device (100) can obtain a loss function including a first difference value between a first feature vector output when second learning data is input to the encoder before relearning and a second feature vector output when second learning data is input to the encoder during relearning based on the second learning algorithm. At this time, the encoder before relearning may be a neural network trained based on the first learning data and the first learning algorithm.
[0192] At this time, the electronic device (100) can obtain a loss function including a second difference value and a first difference value between the output of the encoder into which the second learning data is input and the target output of the decoder when the output is input to the decoder.
[0193] Once the loss function is obtained, the electronic device (100) can retrain the encoder so that the loss function approaches 0 (S630).
[0194] At this time, the method by which the electronic device (100) retrains the encoder may be the same as the method described above, and redundant descriptions are omitted.
[0195] Specifically, the artificial intelligence model may be a model learned based on first learning data and a first learning algorithm.
[0196] At this time, the electronic device (100) can obtain information about a second algorithm that is different from the first learning algorithm.
[0197] Additionally, the electronic device (100) can retrain the encoder using a loss function that includes the difference between the output of the encoder before learning and the output of the encoder after learning.
[0198] At this time, the method by which the electronic device (100) retrains the encoder is as described above, so redundant description is omitted.
[0199] Although various embodiments have been described above, each embodiment is not necessarily implemented individually, and may be implemented together in a single product by being combined in whole or in part with at least one other embodiment.
[0200] Meanwhile, the terms "part" or "module" used in the present disclosure include units composed of hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A "part" or "module" may be an integrally composed component, a minimum unit performing one or more functions, or a portion thereof. For example, a module may be composed of an application-specific integrated circuit (ASIC).
[0201] Various embodiments of the present disclosure may be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device, which is a device capable of calling instructions stored in the storage medium and operating according to the called instructions, may include an electronic device (100) according to the disclosed embodiments. When the instructions are executed by a processor, the processor may directly or under the control of the processor perform a function corresponding to the instructions using other components. The instructions may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" means that the storage medium does not contain signals and is tangible, but does not distinguish between data being stored semi-permanently or temporarily in the storage medium.
[0202] According to one or more embodiments, the methods according to the various embodiments disclosed herein may be provided as a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0203] Each component (e.g., a module or a program) according to various embodiments may be composed of one or more entities, and some of the aforementioned sub-components may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., a module or a program) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the respective components prior to integration. Operations performed by a module, program, or other component according to various embodiments may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
Claims
1. A method for training an artificial intelligence model including an encoder trained based on first training data, Step of acquiring second learning data; and A step of retraining the encoder based on the second learning data; The step of retraining the above encoder is: A method for retraining an encoder such that the distance between a first feature vector when the second learning data is input to the encoder before retraining and a second feature vector when the second learning data is input to the encoder during retraining is less than a preset distance in the embedding space.
2. In paragraph 1, The step of retraining the above encoder is: A method for retraining the encoder using a loss function including a first difference value between a first feature vector when the second learning data is input to the encoder before the retraining and a second feature vector when the second learning data is input to the encoder during the retraining.
3. In paragraph 2, The above artificial intelligence model further includes a decoder, The step of retraining the above encoder is: A method for retraining the encoder using a loss function including the first difference value and the second difference value between the output of the decoder when the second training data is input and the target output of the decoder.
4. In paragraph 3, The above method, A method further comprising: a step of obtaining the loss function by normalizing the first difference value and the second difference value.
5. In paragraph 3, The above method, A control method further comprising: a step of obtaining the loss function by adding a product of a first weight value and the first feature vector and a product of a second weight value and the second feature vector.
6. In paragraph 5, A control method, wherein the sum of the first weighted value and the second weighted value is a preset value.
7. In paragraph 3, The above control method is, Step of obtaining input data; A step of inputting input data into the above retrained encoder; a step of inputting the output of the above relearned encoder to the decoder; and A method comprising: providing a service corresponding to the input data using the output of the decoder.
8. In paragraph 1, A method wherein the second learning data is updated data of the first learning data.
9. An electronic device for training an artificial intelligence model including an encoder trained based on first training data, memory; and a processor; including; The above processor, Obtain the second learning data, Retraining the encoder based on the second learning data, The above processor, An electronic device that retrains the encoder so that the distance between the first feature vector when the second learning data is input to the encoder before the retraining and the second feature vector when the second learning data is input to the encoder during the retraining becomes less than a preset distance in the embedding space.
10. In paragraph 9, The above processor, An electronic device that retrains the encoder using a loss function including a first difference value between a first feature vector when the second learning data is input to the encoder before the retraining and a second feature vector when the second learning data is input to the encoder during the retraining.
11. In paragraph 10, The above artificial intelligence model further includes a decoder, The above processor, An electronic device that retrains the encoder using a loss function including the first difference value and the second difference value between the output of the decoder when the second learning data is input and the target output of the decoder.
12. In paragraph 11, The above processor, An electronic device that obtains the loss function by normalizing the first difference value and the second difference value.
13. In paragraph 11, The above processor, A control electronic device that obtains the loss function by adding the product of the first weighted value and the first feature vector and the product of the second weighted value and the second feature vector.
14. In paragraph 13, A control electronic device, wherein the sum of the first weighted value and the second weighted value is a preset value.
15. In paragraph 11, The above processor, Obtain input data, Input data is input to the above retrained encoder, The output of the above retrained encoder is input to the decoder, An electronic device that provides a service corresponding to the input data by using the output of the decoder.
Citation Information
Patent Citations
Equipment fault prediction method and device, electronic equipment and storage medium
CN111383220A
Sample recognition model generation method and device, computer equipment and storage medium
CN111444952A
Method and appartus for providing docent services with art work
KR1020240078547A
Learning method using artificial neural network
KR102205430B1
Ammonia Fuel Supply System For Ship
KR102705014B1