Method and apparatus for artificial intelligence model personalization

CN113811895BActive Publication Date: 2026-09-18SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080034927.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-28
Filing Date
2020-07-08
Publication Date
2026-09-18
Estimated Expiration
2040-07-08

AI Technical Summary

Technical Problem

然而,这样的方法具有存储容量需要根据训练数据的增加而持续地增加、用于学习的计算机资源增加,并且模型的更新所需的时间增加的问题

Benefits of technology

[0034] As described above, according to various embodiments of this disclosure, one or more general training data and personal training data are generated on an electronic device through a training data generation model, thus eliminating the need to store large amounts of training data (e.g., 1TB to 2TB) for training the artificial intelligence model. Therefore, even on electronic devices with small memory capacities, according to this embodiment, artificial intelligence models can be trained efficiently by using a training data generation model (e.g., a model size of 10MB to 20MB).

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113811895B_ABST
    Figure CN113811895B_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. The electronic device may include: a memory configured to store one or more training data generation models and an artificial intelligence model; and a processor configured to use the one or more training data generation models to generate personal training data reflecting the characteristics of a user, use the personal learning data as training data to train the artificial intelligence model, and store the trained artificial intelligence model in the memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to electronic devices and control methods thereof that are updated based on training data, and more specifically, to methods and apparatus for personalizing artificial intelligence models through incremental learning. Background Technology

[0002] Artificial neural networks can be designed and trained to perform a wide range of functions, and their applications can be used in speech recognition, object recognition, and more. Artificial neural networks can exhibit improved performance by using large amounts of training data from large databases. Specifically, in cases where AI models need to identify elements (such as speech) that are different for each user, a large amount of data is required, as it may be necessary to train the AI ​​model using both individual training data that includes the characteristics of the user of the electronic device and general training data that includes the characteristics of general users.

[0003] Therefore, one could consider, for example Figure 1 The method shown continuously accumulates training data for updating the artificial intelligence model and trains the model based on the accumulated training data. However, such a method has the problems of requiring storage capacity to continuously increase with the increase of training data, increasing computing resources used for learning, and increasing the time required for model updates.

[0004] Alternatively, one could consider training an AI model by sampling only a portion of the training data, but this approach suffers from the problem of inefficiently performing the sampling.

[0005] Furthermore, methods that do not maintain or only partially maintain existing training data and train AI models based on new training data have the catastrophic forgetting problem that causes AI models to forget previously learned knowledge. Summary of the Invention

[0006] Technical issues

[0007] The embodiments of this disclosure overcome the above-described disadvantages and other disadvantages not described above. Furthermore, this disclosure does not need to overcome the above-described disadvantages, and the embodiments of this disclosure may not overcome any of the above-described problems.

[0008] Technical solutions to the problem

[0009] According to one aspect of this disclosure, an electronic device may include: a memory configured to store one or more training data generation models and an artificial intelligence model; and a processor configured to use one or more training data generation models to generate personal training data reflecting the characteristics of a user, use the personal learning data as training data to train the artificial intelligence model, and store the trained artificial intelligence model in the memory.

[0010] One or more training data generation models may include: a personal training data generation model trained to generate personal training data that reflects the characteristics of a user; and a general training data generation model trained to generate general training data corresponding to the usage data of multiple users.

[0011] Artificial intelligence models can be updated based on at least one of personal training data, general training data, or actual user data obtained from users.

[0012] The model generated from personal training data can be updated based on at least one of user data or personal training data.

[0013] Artificial intelligence models can be speech recognition models, handwriting recognition models, object recognition models, speaker recognition models, word recommendation models, or translation models.

[0014] General training data may include first input data, personal training data may include second input data, and the artificial intelligence model performs unsupervised learning based on user data, first input data, and second input data.

[0015] General training data may include first input data, personal training data may include second input data, and the artificial intelligence model may generate first output data corresponding to the first input data based on the first input data, generate second output data corresponding to the second input data based on the second input data, and be trained based on user data, first input data, first output data, second input data and second output data.

[0016] The general training data generation model can be downloaded from the server and stored in memory.

[0017] The processor can upload artificial intelligence models to the server.

[0018] The processor can train an artificial intelligence model based on the electronic device being in a charging state, the occurrence of a predetermined time, or the absence of user manipulation of the electronic device within a predetermined time.

[0019] According to one aspect of this disclosure, a method for controlling an electronic device comprising one or more training data generation models and an artificial intelligence model may include: using one or more training data generation models to generate personal training data reflecting the characteristics of a user; using the personal training data as training data to train an artificial intelligence model; and storing the trained artificial intelligence model.

[0020] One or more training data generation models may include: a personal training data generation model trained to generate personal training data that reflects the characteristics of a user; and a general training data generation model trained to generate general training data corresponding to the usage data of multiple users.

[0021] Artificial intelligence models can be updated based on at least one of personal training data, general training data, or actual user data obtained from users.

[0022] The model generated from personal training data can be updated based on at least one of user data or personal training data.

[0023] Artificial intelligence models can be speech recognition models, handwriting recognition models, object recognition models, speaker recognition models, word recommendation models, or translation models.

[0024] General training data may include first input data, personal training data may include second input data, and the artificial intelligence model may perform unsupervised learning based on user data, first input data, and second input data.

[0025] General training data may include first input data, personal training data may include second input data, and the artificial intelligence model may generate first output data corresponding to the first input data based on the first input data, generate second output data corresponding to the second input data based on the second input data, and be trained based on user data, first input data, first output data, second input data and second output data.

[0026] The general training data can be used to generate models that can be downloaded from the server and stored in the memory of an electronic device.

[0027] Artificial intelligence models can be uploaded to a server.

[0028] The method may include training an artificial intelligence model based on the fact that the electronic device is in a charging state, a predetermined time has occurred, or no user manipulation of the electronic device has been detected within a predetermined time.

[0029] According to one aspect of this disclosure, a non-transitory computer-readable medium may store a program for performing a personalized method of an artificial intelligence model in an electronic device, the method comprising: generating personal training data reflecting characteristics of a user; and using the personal training data as training data to train the artificial intelligence model.

[0030] This method may include generating general training data that reflects the characteristics of multiple users.

[0031] This method may include collecting and storing actual usage data of users obtained through devices in electronic devices.

[0032] This method may include training an artificial intelligence model using at least one of personal training data, general training data, or actual user usage data obtained from the user.

[0033] Beneficial effects of the invention

[0034] As described above, according to various embodiments of this disclosure, one or more general training data and personal training data are generated on an electronic device through a training data generation model, thus eliminating the need to store large amounts of training data (e.g., 1TB to 2TB) for training the artificial intelligence model. Therefore, even on electronic devices with small memory capacities, according to this embodiment, artificial intelligence models can be trained efficiently by using a training data generation model (e.g., a model size of 10MB to 20MB).

[0035] Furthermore, because it does not require downloading training data from an external server, artificial intelligence models can be trained freely even without a network connection.

[0036] Furthermore, since AI models are trained on personal training data that reflects the characteristics of users of electronic devices, they can achieve improved accuracy in recognizing user data such as user voice.

[0037] Additional and / or other aspects and advantages of this disclosure will be set forth in part in the description which follows, and will also be apparent in part from the description or may be learned by practice of this disclosure. Attached Figure Description

[0038] The above and other aspects, features, and advantages of certain embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:

[0039] Figure 1 It is a diagram used to describe related technologies;

[0040] Figure 2 It is a diagram used to schematically illustrate the configuration of an electronic system according to an embodiment;

[0041] Figure 3 It is a block diagram used to describe the operation of an electronic device according to an embodiment;

[0042] Figure 4 This is a flowchart illustrating an illustrative process of training an artificial intelligence model and performing identification according to an embodiment;

[0043] Figure 5 This is a diagram used to describe the process of generating a model from individual training data according to an embodiment;

[0044] Figure 6 This is a diagram used to describe the process of training an artificial intelligence model according to an embodiment;

[0045] Figure 7 It is a diagram used to describe specific operations performed between the processor and the memory according to an embodiment;

[0046] Figure 8 This is a diagram used to describe a case where an artificial intelligence model is implemented using an object recognition model according to another embodiment; and

[0047] Figure 9 This is a diagram used to describe the process of training an artificial intelligence model on an electronic device according to an embodiment. Detailed Implementation

[0048] The present disclosure will be described in detail below with reference to the accompanying drawings.

[0049] After a brief description of the terminology used in the specification, this disclosure will be described in detail.

[0050] In consideration of the functions described herein, currently widely used general terms have been selected as the terms used in the embodiments of this disclosure, but may be changed based on the intent of those skilled in the art, judicial precedent, the emergence of new technologies, etc. Furthermore, in certain circumstances, terms may be arbitrarily chosen by the applicant. In such cases, the meanings of these terms will be mentioned in detail in the corresponding descriptive sections of this disclosure. Therefore, the terms used in the embodiments of this disclosure should be defined based on their meaning and throughout the content of this disclosure, rather than on their simple names.

[0051] Because this disclosure can be modified in various ways and has several embodiments, specific embodiments of this disclosure will be shown in the accompanying drawings and described in detail in the specific embodiments. However, it should be understood that this disclosure is not limited to the specific embodiments, but includes all modifications, equivalents, and substitutions without departing from the scope and spirit of this disclosure. Detailed descriptions of known technologies related to this disclosure will be omitted where it is determined that such detailed descriptions may obscure the gist of this disclosure.

[0052] The singular form of a term is intended to include the plural form of the term unless the context clearly indicates otherwise. It should be understood that terms such as “comprising” or “including” as used in this specification specify the presence of a feature, number, step, operation, component, part, or combination thereof mentioned in this specification, but do not exclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0053] The statement “at least one of A and / or B” should be understood as indicating “A only”, “B only”, or “A and B”.

[0054] The use of terms such as "first" and "second" in the specification may refer to various components, regardless of the order and / or importance of the components, and will only be used to distinguish one component from other components, and does not limit the corresponding components.

[0055] When any component (e.g., the first component) is (operably or communicatively) coupled to or connected to another component (e.g., the second component), it should be understood that any component is directly coupled to another component or can be coupled to another component through another component (e.g., the third component).

[0056] In this disclosure, a "module" or "device" may perform at least one function or operation and may be implemented by hardware or software, or by a combination of hardware and software. Furthermore, multiple "modules" or multiple "devices" may be integrated into at least one module, and in addition to "modules" or "devices" implemented by specific hardware, they may be implemented by at least one processor (not shown). In this disclosure, the term "user" may refer to a person using an electronic device or a device using an electronic device (e.g., an artificial intelligence (AI) electronic device).

[0057] In the following description, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings, enabling those skilled in the art to readily practice the disclosure. However, the present disclosure can be modified in various different forms and is not limited to the embodiments described herein. Furthermore, in the drawings, portions unrelated to the description will be omitted so as not to obscure the present disclosure, and similar reference numerals will be used throughout the specification to describe similar portions.

[0058] In the following, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0059] Figure 2 This is a diagram used to schematically depict the configuration of an electronic system according to an embodiment.

[0060] refer to Figure 2 The electronic system 10 according to the embodiment includes an electronic device 100 and a server 200.

[0061] The electronic device 100 according to an embodiment can be configured to perform specific operations to identify user data by using an artificial intelligence model (or alternatively, a neural network model or a learning network model). Here, user data is data that reflects the unique characteristics of a user, such as the user's voice, handwriting, images captured by the user, user input character data, translation data, etc. Furthermore, user data can also refer to actual usage data obtained from the user and may be called actual user data.

[0062] The AI-related functions according to this disclosure are executed by a processor and memory. The processor can be implemented by one or more processors. Here, the one or more processors can be general-purpose processors (such as central processing units (CPUs), application processors (APs), or digital signal processors (DSPs)), graphics-specific processors (such as graphics processing units (GPUs) or vision processing units (VPUs)), or AI-specific processors (such as neural processing units (NPUs)). The one or more processors can be configured to perform control to process input data according to predefined operating rules or AI models stored in memory. Alternatively, in the case where one or more processors are AI-specific processors, the AI-specific processors can be designed with hardware architectures dedicated to processing specific AI models.

[0063] Predefined operating rules or artificial intelligence models are obtained through training. Here, obtaining predefined operating rules or artificial intelligence models through training means using training data and training algorithms to train a basic artificial intelligence model to obtain a predefined set of operating rules or artificial intelligence models, thereby achieving the desired characteristics (or purpose). Training can be performed by a device in which artificial intelligence is performed according to this disclosure, or it can be performed by a separate server and / or system. Examples of training algorithms may include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning.

[0064] Artificial intelligence models can include multiple neural network layers. Each of these layers has multiple weight values, and neural network computations are performed using the results of the previous layer's computations and by using these multiple weight values. The weight values ​​of the multiple neural network layers can be optimized (or improved) through the training results of the artificial intelligence model. For example, multiple weight values ​​can be updated during the training process to reduce or minimize the loss or cost values ​​acquired by the artificial intelligence model. Artificial neural networks can include deep neural networks (DNNs). For example, an artificial neural network can be a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or a deep Q-network, but is not limited to these.

[0065] In cases where the artificial intelligence model is a speech recognition model according to an embodiment, electronic device 100 may include a speech recognition model that recognizes user speech and thus can provide virtual assistant functionality. For example, electronic device 100 may be implemented in various forms, such as smartphones, tablet PCs, mobile phones, video phones, e-book readers, desktop PCs, laptop PCs, netbooks, workstations, servers, personal digital assistants (PDAs), portable multimedia players (PMPs), MP3 players, medical devices, cameras, and wearable devices.

[0066] Artificial intelligence models can include multiple neural network layers and can be trained to increase recognition efficiency.

[0067] The electronic device 100 according to an embodiment may include a training data generation model that generates training data itself without retaining a large amount of training data for training an artificial intelligence model, thus allowing the artificial intelligence model to be updated in real time as needed.

[0068] The training data to generate the model can be pre-installed on the electronic device, or it can be downloaded from server 200 and installed on the electronic device.

[0069] As mentioned above, artificial intelligence models can be trained based on training data generated by training data generation models installed in electronic devices.

[0070] Here, the training data generation model may include at least one personal training data generation model and a general training data generation model. The personal training data generation model may be a model that generates personal training data to reflect the characteristics of the user of the electronic device 100, and the general training data generation model may be a model that generates general training data to reflect the characteristics of a general user.

[0071] Artificial intelligence models can be trained using personal training data generated from personal training data, general training data generated from general training data, or both personal and general training data.

[0072] An AI model can be trained using both personal training data and general training data. This is because training an AI model using only general training data may not reflect the characteristics of the user of the electronic device 100, while training an AI model using only personal training data may result in an over-biased focus on the characteristics of a specific user during the training process.

[0073] Server 200 is a device configured to manage at least one of a personal training data generation model, a general training data generation model, or an artificial intelligence model, and can be implemented through a central server, cloud server, etc. Server 200 can send at least one of the personal training data generation model, general training data generation model, or artificial intelligence model to electronic device 100 based on requests from electronic device 100. Specifically, server 200 can send a general training data generation model pre-trained based on general user data to electronic device 100. Furthermore, server 200 can send update information (e.g., weight values ​​or bias information for each layer) or the updated general training data generation model itself to electronic device 100 as needed.

[0074] The personal training data generation model and artificial intelligence model sent to electronic device 100 can be untrained models. These models can be pre-installed on the electronic device. Considering user privacy, the personal training data generation model and artificial intelligence model can be trained on electronic device 100 without the need to send user data related to training to a server. In some cases, user data can be sent to server 200, and the personal training data generation model and artificial intelligence model can be trained on server 200.

[0075] The training data generation model can directly generate training data within the electronic device 100. Therefore, the electronic device 100 may not need to store large amounts of actual training data or receive training data separately from the server 200. Furthermore, in addition to general training data, the artificial intelligence model can also be trained using personal training data. Thus, the artificial intelligence model can be trained as a personalized AI model that reflects the user's characteristics and can be incrementally updated.

[0076] Besides speech recognition models, artificial intelligence models can be implemented using various other models, such as handwriting recognition models, visual object recognition models, speaker recognition models, word recommendation models, and translation models. However, for ease of explanation, the following text will primarily describe speech recognition models.

[0077] Figure 3 This is a block diagram used to describe the operation of an electronic device according to embodiments of the present disclosure.

[0078] refer to Figure 3 The electronic device 100 includes a memory 110 and a processor 120.

[0079] The memory 110 is electrically connected to the processor 120 and can store data for implementing various embodiments of the present disclosure.

[0080] Depending on the purpose of data storage, memory 110 may be implemented as a memory embedded in electronic device 100 or as a memory that can be attached to and detached from electronic device 100. For example, data for driving electronic device 100 may be stored in memory embedded in electronic device 100, and data for extended functions of electronic device 100 may be stored in memory that can be attached to and detached from electronic device 100. The memory embedded in the electronic device 100 may be implemented by at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM) or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one-time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard disk drive or solid-state drive (SSD)), and the memory that can be attached to and removed from the electronic device 100 may be implemented by memory cards (e.g., compact flash (CF), secure digital (SD), micro-secure digital (Micro-SD), mini-secure digital (Mini-SD), high-speed digital (xD), multimedia card (MMC)), external memory that can be connected to a USB port (e.g., USB memory), etc.

[0081] According to an embodiment, the memory 110 may store one or more training data generation models and artificial intelligence models. Here, the one or more training data generation models may include general training data generation models and personal training data generation models.

[0082] Here, the general training data generation model can be a model trained to generate general training data corresponding to the usage data of general users. In other words, the general training data can be data that reflects the characteristics of general users. According to an embodiment, the general training data generation model includes multiple neural network layers, each of which includes multiple parameters, and each layer can perform neural network computation by using the computation results of the previous layer and by using the computation of multiple parameters. For example, computation can be sequentially performed in multiple pre-trained neural network layers included in the general training data generation model based on random input values, so that general training data can be generated.

[0083] The personal training data generation model is a model trained to generate personal training data that reflects the characteristics of the user of the electronic device 100. For example, computations are sequentially performed in multiple pre-trained neural network layers included in the personal training data generation model based on random input values, making it possible to generate personal training data.

[0084] According to embodiments, the artificial intelligence model can be a model that recognizes data obtained from user actions (such as speech or handwriting). The artificial intelligence model can be a speech recognition model, a handwriting recognition model, an object recognition model, a speaker recognition model, a word recommendation model, a translation model, etc. As an example, where the artificial intelligence model is implemented by a speech recognition model, the artificial intelligence model can be implemented by an automatic voice recognition (ASR) model that recognizes the user's speech and outputs text corresponding to the recognized speech. However, according to another embodiment, the artificial intelligence model can be a model that generates various outputs based on the user's speech (e.g., outputting transformed speech corresponding to the user's speech). In addition to the models described above, the artificial intelligence model can be implemented by various types of models, as long as it can be trained using personal training data to reflect individual characteristics.

[0085] Personal training data can be used to generate models and artificial intelligence models, which can be received from server 200 and stored in memory 110. However, personal training data can also be used to generate models and artificial intelligence models, which can be stored in memory 110 during the manufacture of electronic device 100, or received from external devices other than the server and stored in memory 110.

[0086] Furthermore, the memory 110 can store user data. Here, user data refers to the user's actual user data. For example, user data can be speech data directly spoken by the user, handwriting actually written by the user, images directly captured by the user, etc. User data has a concept different from that of personal training data or general training data. For example, user data can be speech data directly spoken by the user, and personal training data can be speech data similar to the user's directly spoken speech, which is artificially generated by a personal training data generation model used to train an artificial intelligence model.

[0087] As an example, when the user data is user voice data, the user voice data may be user voice data received through a microphone (not shown) provided in the electronic device 100 or user voice data received from an external device. For example, the user voice data may be data such as a WAV file or an MP3 file, but is not limited to these.

[0088] As another example, in the case where the user data is user handwriting data, the user handwriting data may be handwriting data input by a user through a display (not shown) provided in the electronic device 100, by touch, or by a stylus.

[0089] User data can be stored in a different memory than the memory used to store general training data generation models, personal training data generation models, and artificial intelligence models.

[0090] Processor 120 is electrically connected to memory 110 and controls the overall operation of electronic device 100. Processor 120 controls the overall operation of electronic device 100 using various instructions or programs stored in memory 110. Specifically, according to an embodiment, the main central processing unit (CPU) can copy a program to RAM according to instructions stored in ROM and access RAM to execute the program. Here, the program may include artificial intelligence models, etc.

[0091] According to an embodiment, the processor 120 can train an artificial intelligence model based on personal training data that reflects user characteristics and is generated by a personal training data generation model, and user data that corresponds to the user's actual usage data. However, when training an artificial intelligence model solely using personal training data, the training of the artificial intelligence model may be overly biased towards the characteristics of a specific user.

[0092] Therefore, the processor 120 can train an artificial intelligence model based on general training data, personal training data, and user data generated by a general training data generation model that reflects the characteristics of a general user. Using both general training data and personal training data, the artificial intelligence model can be trained to be personalized for the user, and training of the artificial intelligence model can be performed based on the general training data without excessively biasing towards a specific user. Embodiments of training an artificial intelligence model based on general training data and personal training data will be described below.

[0093] Figure 4 This is a flowchart illustrating an illustrative process of training an artificial intelligence model and performing identification according to an embodiment of the present disclosure.

[0094] The electronic device 100 can store a general training data generation model, a personal training data generation model, and an artificial intelligence model in the memory 110.

[0095] The personal training data generation model can be trained based on user data (operation S410). According to an embodiment, the user data can be data obtained in the training data acquisition mode of the electronic device 100, thereby obtaining actual user usage data as training data. For example, once the training data acquisition mode is activated, predetermined text can be displayed on the screen, and a message requesting the displayed text can be provided. For example, the processor 120 can control a speaker (not shown) to output audio output such as "Please read the displayed text aloud," or can control the screen to display a user interface (UI) window showing the message "Please read the displayed text aloud." Then, user speech corresponding to the displayed text (such as "How's the weather today?") is input through a microphone (not shown) provided in the electronic device 100, and the input speech can be used as user data to train the personal training data generation model and the artificial intelligence model. Alternatively, in the training data acquisition mode, the user can input handwriting corresponding to a specific text based on a predetermined message. The aforementioned handwriting input can be used as user data.

[0096] However, this disclosure is not limited to this, and even when the training data acquisition mode is not activated, the processor 120 can acquire user voice through the microphone and use the acquired user voice as training data. For example, user voice inquiries, user voice commands, etc., entered during normal use of the electronic device can be acquired during a predetermined period and used as training data.

[0097] As described in more detail below, after generating personal training data, a personal training data generation model can be trained based on at least one of user voice data or personal training data. (See reference...) Figure 5 Describe operation S410 in detail.

[0098] A general training data generation model is pre-trained on server 200 and then sent to electronic device 100, so it may not be necessary to train the general training data generation model separately on electronic device 100. However, the general training data generation model can be updated on server 200, and the updated general training data generation model can be periodically sent to electronic device 100 to replace or update the existing general training data generation model.

[0099] A general training data generation model can generate general training data that reflects the characteristics of a typical user. A personal training data generation model can generate personal training data that reflects the characteristics of the user of electronic device 100 (operation S420). Personal training data includes training data that incorporates the user's characteristics and can be used to train an artificial intelligence model into a personalized model.

[0100] According to one embodiment, the artificial intelligence model can be trained and updated based on user data obtained within a predetermined time period, generated general training data, and generated personal training data (operation S430). However, according to another embodiment, the artificial intelligence model can also be trained based on user data and generated personal training data. The artificial intelligence model can be trained and updated based on at least one or any combination of user data, generated general training data, or generated personal training data.

[0101] Unlike the generated general training data and the generated personal training data, user data is actual usage data as raw, unprocessed data. Therefore, AI models can be trained using user data alongside general and personal training data to improve their recognition accuracy. User data can be used to train models generated from personal training data and also to train AI models. (See reference...) Figure 6 Describe operation S430 in detail.

[0102] Then, user data identification (operation S440) can be performed based on the trained (personalized and updated) artificial intelligence model. Here, user data is the data to be identified by the artificial intelligence model, and can be, for example, voice data used for voice queries. The user data identified by the artificial intelligence can also be used to update the personally trained data generation model and the artificial intelligence model.

[0103] Specifically, when a user's voice is input into an artificial intelligence model, the model can recognize the user's voice, convert the user's voice into text, and output the text.

[0104] Figure 5 This is a diagram illustrating the process of training a model from personal training data according to embodiments of the present disclosure.

[0105] Processor 120 can load user data stored in memory 110 (operation S510). For example, processor 120 can load user data stored in memory 110 outside processor 120 into internal memory of processor 120 (not shown).

[0106] Here, user data can be data input according to a request from processor 120. For example, if the user data is user voice data, processor 120 can request the user to speak text including predetermined words or sentences to obtain user voice data as training data. As an example, processor 120 can display predetermined words or sentences on a display (not shown) to guide the user to speak the predetermined words or sentences, or it can output speech through a speaker (not shown) to request the user to speak the predetermined words or sentences. Processor 120 can store the user voice data input as described above in memory 110, and then load the stored user voice data to train a personal training data generation model. Therefore, user voice data stored in memory 110 can be loaded into processor 120. However, this disclosure is not limited to this, and if the user data is user handwriting data, characters or numbers handwritten by the user according to a request from processor 120 can be input into electronic device 100.

[0107] Furthermore, the processor 120 can load the personal training data generated model stored in the memory 110 (operation S520). The order of operations S510 and S520 can be changed.

[0108] Then, a personal training data generation model can be trained based on user data (operation S530). Specifically, the personal training data generation model can be trained using the distribution of user data. For example, the frequency distribution in user speech data can be used to update the personal training data generation model. The frequency distribution in speech is different for each person. Therefore, by training the personal training data generation model based on the frequency distribution characteristics, the speech characteristics of the user of the electronic device 100 can be accurately reflected.

[0109] Therefore, a model generated from personal training data can be trained to be personalized for the user.

[0110] Then, the trained personal training data generation model can generate personal training data (operation S540). The personal training data generation model is trained to reflect the characteristics of the user of the electronic device 100; therefore, the personal training data generated by the personal training data generation model can be similar to the user data. As an example, personal training data can be generated as speech very similar to the user's directly spoken speech. As another example, personal training data can be generated as text with a form very similar to the text directly written by the user. The similarity between the user data and the personal training data can be determined based on the updated results of the personal training data generation model.

[0111] The generated personal training data can be used together with general training data and user data to train artificial intelligence models (operation S550).

[0112] In addition to user data, personal training data generation models can also be trained and updated based on the generated personal training data.

[0113] Then, the processor 120 can store the trained individual training data generation model in the memory 110 (operation S560).

[0114] The personal training data generation model can be updated by continuously repeating the above operations S510 to S560.

[0115] Figure 6 This is a diagram used to describe the process of training an artificial intelligence model according to an embodiment.

[0116] The processor 120 can load training data stored in the memory 110 (operation S610). Specifically, the processor 120 can load user data obtained from the user, as well as general training data and personal training data generated by the training data generation model.

[0117] Processor 120 can identify whether the artificial intelligence model is in a trainable state (operation S620). Processor 120 can identify whether sufficient computing resources have been ensured for training the artificial intelligence model. For example, because a large amount of computing resources have been used, the number of operations performed other than training the artificial intelligence model on electronic device 100 is relatively low, which can identify that the artificial intelligence model is in a trainable state.

[0118] Specifically, the trainable states of the artificial intelligence model can include situations where the electronic device 100 is charging, a predetermined time has occurred, or no user interaction occurs within a predetermined time. Because the electronic device 100 is powered while charging, it can remain powered on during the training of the artificial intelligence model. Furthermore, the predetermined time could be, for example, the time when the user begins to sleep. As an example, if the predetermined time is 1:00 AM, the artificial intelligence model can be trained based on training data from 1:00 AM. The predetermined time can be obtained by monitoring usage patterns of the electronic device 100 or can be a time set by the user. Moreover, if no user interaction occurs within the predetermined time, it is recognized that there is likely to be little user interaction even afterward, thus allowing the artificial intelligence model to be trained. As an example, if no user interaction occurs with the electronic device 100 for one hour, the artificial intelligence model can be trained.

[0119] Furthermore, according to embodiments, training of the artificial intelligence model can be performed when additional learning start conditions, besides the trainable state of the artificial intelligence model, are met. Learning start conditions may include at least one of the following: commands for training the artificial intelligence model have been input; the artificial intelligence model has made a predetermined number or more identification errors; or a predetermined amount or more of training data has been accumulated. For example, learning start conditions may be met when identification errors occur five or more times within a predetermined time period, or when 10MB or more of user voice data has been accumulated.

[0120] If it is determined that the artificial intelligence model is not in a trainable state (operation S620 - No), the processor 120 may store the loaded training data (operation S625). For example, the processor 120 may store the loaded user voice data again in the memory 110.

[0121] If the artificial intelligence model is identified as being in a trainable state (operation S620 - Yes), the processor 120 can load the artificial intelligence model stored in the memory 110 (operation S630).

[0122] The loaded artificial intelligence model can be trained based on user voice data, general training data, and personal training data (operation S640). Specifically, the artificial intelligence model can be trained sequentially using user voice data, general training data, and personal training data.

[0123] The processor 120 can store the trained artificial intelligence model in the memory 110 (operation S650).

[0124] Then, according to the example, processor 120 can delete training data from memory 110 (operation S660). For example, processor 120 can delete user voice data stored in memory 110. In other words, user voice data used to train personal training data generation models and artificial intelligence models can be deleted.

[0125] According to another example, if the amount of data stored in memory 110 exceeds a predetermined amount, training data can be deleted sequentially starting from the oldest training data stored in memory 110.

[0126] The order in which S650 and S660 are operated can be changed.

[0127] In the above embodiments, the use of pre-generated and stored personal (general) training data has been described when the learning start conditions for the artificial intelligence model are met. However, personal (general) training data can also be generated and used by operating the training data generation module based on the met conditions.

[0128] Figure 7 It is a diagram used to describe specific operations performed between the processor and the memory according to an embodiment.

[0129] Omitted references Figures 4 to 6 A detailed description of the overlapping parts described.

[0130] Under the control of the processor 120, the general training data generation model and the personal training data generation model stored in the memory 110 can be loaded into the processor 120 (①).

[0131] Then, the general training data generation model and the personal training data generation model loaded into the processor 120 can generate general training data and personal training data respectively (②).

[0132] Furthermore, under the control of the processor 120, an artificial intelligence model stored in the memory 110 can be loaded into the processor 120 (③). For example, the processor 120 can load an artificial intelligence model stored in the memory 110 outside the processor 120 into the internal memory (not shown) of the processor 120. Then, the processor 120 can access the artificial intelligence model loaded into the internal memory.

[0133] The artificial intelligence model loaded onto the processor 120 can be trained based on user data, general training data, and personal training data (④). Here, user data can be data loaded from the memory 110.

[0134] The processor 120 can store the trained artificial intelligence model in the memory 110 (⑤). In addition, the processor 120 can upload the trained artificial intelligence model to the server 200.

[0135] According to an embodiment, general training data may include first input data, personal training data may include second input data, and the artificial intelligence model may perform unsupervised learning based on user data, the first input data, and the second input data.

[0136] For example, suppose the artificial intelligence model is implemented using a speech recognition model. General training data may include first speech data (first input data), and personal training data may include second speech data (second input data). In this case, the artificial intelligence model can perform unsupervised learning based on the user's speech data, the first speech data, and the second speech data.

[0137] According to another embodiment, general training data may include first input data, and personal training data may include second input data. In this case, the artificial intelligence model can generate first output data corresponding to the first input data based on the first input data, and the artificial intelligence model can generate second output data corresponding to the second input data based on the second input data. The artificial intelligence model can then be trained based on user data, the first input data, the first output data, the second input data, and the second output data.

[0138] For example, suppose the artificial intelligence model is implemented by a speech recognition model. The artificial intelligence model can generate first text data (first output data) corresponding to the first speech data (first input data) based on the first input speech data, and can generate second text data (second output data) corresponding to the second speech data (second input data) based on the second input speech data.

[0139] Here, the first and second text data can be generated in the form of probability distributions. For example, if the first speech data includes the speech "meeting room" or similar speech, the first text data can be generated as a probability vector (y1, y2, and y3). For example, y1 could be the probability that the text corresponding to the first speech data is "meeting room," y2 could be the probability that the text corresponding to the first speech data is "lounge room," and y3 could be the probability that the text corresponding to the first speech data is neither "meeting room" nor "lounge room."

[0140] Artificial intelligence models can also be trained based on user voice data, first voice data, and second voice data, as well as first and second text data generated in the form of probability distributions. Furthermore, personal training data generation models can be trained based on second text data generated by the artificial intelligence model.

[0141] According to another embodiment, an artificial intelligence model can be trained based on first input data, first label data corresponding to the first input data, second input data, and second label data corresponding to the second input data. Here, label data can refer to explicit correct answer data of the input data. For example, in the case where the artificial intelligence model is a speech recognition model, the first label data can refer to correct answer text data corresponding to the first speech data. As an example, when a user is asked to speak the displayed text, the spoken speech can be user speech data, and the displayed text can be text data corresponding to that user speech data. Alternatively, the text corresponding to the user's speech can be displayed as output on a display (not shown), and the electronic device 100 can receive feedback from the user regarding that text. For example, if the output text corresponding to the user's speech is "nail this to me" and the text received as feedback from the user is "mail this to me", "mail this to me" can be text data corresponding to the user's speech data as a label.

[0142] Similarly, general training data may include first speech data and first label data corresponding to the first speech data.

[0143] In other words, general training data may include first speech data and first text data corresponding to the first speech data, and personal training data may include second speech data and second text data corresponding to the second speech data. In the following text, for ease of explanation, when training data is configured as pairs, general training data is described as a general training data pair, and personal training data is described as a personal training data pair.

[0144] Artificial intelligence models can be trained based on general training data pairs, personal training data pairs, and user voice data pairs. Each of the general training data pairs and personal training data pairs includes voice data and corresponding text data.

[0145] Specifically, the AI ​​model can perform supervised learning, where first speech data is input into the AI ​​model and the output text data is compared with the first text data. Similarly, the AI ​​model can perform supervised learning, where second speech data is input into the AI ​​model and the output text data is compared with the second text data.

[0146] Although user data and user voice data have been described, this disclosure is not limited thereto, provided that the training data can be configured accordingly. For example, personal training data may include handwriting data generated similar to the user's handwriting and corresponding text data. Here, the text data may be data input as labels for the handwriting data.

[0147] Figure 8 This is a diagram used to describe a case where an artificial intelligence model is implemented using an object recognition model according to another embodiment.

[0148] Figure 8 It is an image showing digital handwriting. For example, Figure 8 The handwriting shown can be either directly written by the user or generated by a model based on personally trained data. Similar to speech, handwriting is different for each user; therefore, by training the AI ​​model according to user characteristics to personalize it, the recognition accuracy of the AI ​​model can be improved. The process of training the object recognition model is the same as described above. Figure 4 The process shown is the same.

[0149] Specifically, a personal training data generation model can be trained based on user handwriting data. Here, user handwriting data can include data that includes handwriting input according to a request from processor 120. For example, processor 120 can request the user to write numbers from 0 to 9. In this case, digital handwriting can be input to a touch display using a stylus, or an image including digital handwriting written with a pen can be input to electronic device 100.

[0150] Then, a personalized training data generation model can be trained based on user handwriting data. Therefore, the personalized training data generation model can be trained to be a personalized model for each user. The trained general training data generation model and the trained personalized training data generation model can generate general training data and personalized training data, respectively. Here, the general training data can be training data reflecting the handwriting characteristics of a typical user, and the personalized training data can be training data reflecting the handwriting characteristics of the user of the electronic device 100. An artificial intelligence model can be trained based on the general training data, personalized training data, and user handwriting data. Therefore, the trained artificial intelligence model reflects the user's handwriting characteristics, thereby improving the accuracy of user handwriting recognition.

[0151] However, this disclosure is not limited to this, and artificial intelligence models can be trained based on handwriting data and personal training data.

[0152] Although already Figure 8The example described uses digital handwriting, but the examples can also be applied to handwriting recognition models that identify characters such as the alphabet, as well as various other object recognition models.

[0153] Figure 9 This is a diagram used to describe the process of training an artificial intelligence model on an electronic device according to an embodiment.

[0154] Omitted references Figures 4 to 6 A detailed description of the overlapping parts described.

[0155] A personal training data generation model can be trained based on user data stored in the user data storage. The general training data generation model downloaded from the server and the trained personal training data generation model can be stored in the training data generation model storage. Here, the training data generation model storage can be, but is not limited to, the same physical storage as the user data storage, and can be different from the user data storage. The general training data generation model and the personal training data generation model can generate general training data and personal training data respectively. Here, the general training data generation model and the personal training data generation model can each be implemented by a variational autoencoder (VAE), a generative adversarial network (GAN), etc., but are not limited to these.

[0156] Then, the artificial intelligence model can be updated based on user data, general training data, and personal training data stored in the user data store.

[0157] The updated AI model can be uploaded to server 200 periodically. For example, the AI ​​model can be uploaded to server 200 based on at least one of the following: a predetermined time period has elapsed or the amount of data learned by the AI ​​model is a predetermined amount.

[0158] A new (or updated) version of the artificial intelligence model can be sent from server 200 to electronic device 100. In this case, the existing trained artificial intelligence model can be replaced (or updated) by the new (or updated) version of the artificial intelligence model sent from server 200. The new (or updated) version of the artificial intelligence model can be trained to reflect the user's personal characteristics, and even if the user data (e.g., user voice data stored as WAV or MP3 files) is not stored in memory 110, the new (or updated) version of the artificial intelligence model can be trained relatively quickly based on personal training data generated by a pre-trained personal training data generation model. In other words, because the personal training data generation model is trained based on user data, similar effects can be obtained even when the artificial intelligence model is trained based on personal training data generated by the personal training data generation model, as in the case where the artificial intelligence model is trained based on user data. Therefore, even when the artificial intelligence model is replaced (or updated) by a new version of the artificial intelligence model, electronic device 100 can be trained to reflect personal characteristics based on the generated personal training data without storing large amounts of user data.

[0159] In addition to the memory 110 and the processor 120, the electronic device 100 may also include a communication interface (not shown), a display (not shown), a microphone (not shown), and a speaker (not shown).

[0160] The communication interface includes a circuit system and can perform communication with server 200.

[0161] The communication interface may include a Wi-Fi module (not shown), a Bluetooth module (not shown), an infrared (IR) module, a local area network (LAN) module, an Ethernet module, etc. Each communication module may be implemented as at least one hardware chip. In addition to the communication methods described above, the wireless communication module may include at least one communication chip that performs communication according to various wireless communication protocols, such as Zigbee, Universal Serial Bus (USB), Mobile Industrial Processor Interface Camera Serial Interface (MIPI CSI), 3G, 3GPP, LTE, LTE-A Advanced, 4G, and 5G. However, this is only an example, and the communication interface may use at least one of various communication modules. Furthermore, the communication interface may also communicate with a server via a wired connection.

[0162] The communication interface can receive general training data generation models, personally trained data generation models, and artificial intelligence models from server 200 via wired or wireless communication. Furthermore, the communication interface can send trained artificial intelligence models to server 200.

[0163] The display can be implemented using a touchscreen in which the display and touchpad are configured in a mutual layer structure. Here, the touchscreen can be configured to detect the location and area of ​​a touch, as well as the pressure applied during the touch. For example, the display can detect handwriting input from a stylus.

[0164] A microphone is a component configured to receive user speech. The received user speech can be used as training data, or it can be used to output text corresponding to the user's speech based on a speech recognition model.

[0165] The methods according to the various embodiments of the present disclosure described above can be implemented in the form of applications that can be installed in existing electronic devices.

[0166] Furthermore, the methods according to the various embodiments of this disclosure described above can be implemented through software or hardware upgrades of existing electronic devices.

[0167] Furthermore, the various embodiments described above can also be implemented by an embedded server disposed in an electronic device or by at least one external server of the electronic device.

[0168] According to embodiments of this disclosure, the various embodiments described above can be implemented by software including instructions stored in a machine-readable storage medium (e.g., a non-transitory computer-readable storage medium). A machine is an apparatus that can invoke stored instructions from a storage medium and operate according to the invoked instructions. A machine may include an electronic device according to the disclosed embodiments. When an instruction is executed by a processor, the processor may directly execute the function corresponding to the instruction, or other components may execute the function corresponding to the instruction under the control of the processor. Instructions may include code created or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory" means that the storage medium is tangible and does not distinguish whether data is stored semi-permanently or temporarily on the storage medium. As an example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0169] Furthermore, according to embodiments of this disclosure, the methods according to the various embodiments described above can be included in and provided in a computer program product. The computer program product can be traded as a product between a seller and a buyer. The computer program product can be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)) or through an app store (e.g., the Play Store). TM Online distribution is permitted. In the case of online distribution, at least a portion of the computer program product may be stored at least temporarily in a storage medium (such as the storage of the manufacturer's server, the application store's server, or a relay server) or may be temporarily created.

[0170] Furthermore, according to embodiments of this disclosure, the various embodiments described above can be implemented using software, hardware, or a combination of software and hardware, on a computer or a computer-readable recording medium. In some cases, the embodiments described in this disclosure can be implemented by the processor itself. Depending on the software implementation, the embodiments such as processes and functions described in this disclosure can be implemented by separate software modules. Each of these software modules can perform one or more functions and operations described in this disclosure.

[0171] Computer instructions for performing processing operations of a machine according to the various embodiments of the present disclosure described above may be stored in a non-transitory computer-readable medium. According to the various embodiments described above, when the computer instructions stored in the non-transitory computer-readable medium are executed by a processor of a particular machine, they allow the particular machine to perform processing operations within the machine.

[0172] Non-transitory computer-readable media can be media in which data is stored semi-permanently and can be read by a machine. Specific examples of non-transitory computer-readable media can include compact discs (CDs), digital versatile discs (DVDs), hard disks, Blu-ray discs, universal serial bus (USB), memory cards, ROMs, etc.

[0173] Furthermore, each of the components (e.g., modules or programs) according to the various embodiments described above may include a single entity or multiple entities, and some of the corresponding sub-components may be omitted in the various embodiments, or other sub-components may be further included. Alternatively or additionally, some of the components (e.g., modules or programs) may be integrated into one entity, and may perform the functions performed by the respective corresponding components prior to integration in the same or similar manner. Operations performed by modules, programs, or other components according to the various embodiments may be performed sequentially, in parallel, iteratively, or heuristically, and may be performed in different orders or at least some of the operations may be omitted, or other operations may be added.

[0174] Although embodiments of the present disclosure have been shown and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by those skilled in the art to which this disclosure pertains without departing from the spirit of the present disclosure as disclosed in the appended claims. These modifications should also be understood to fall within the scope and spirit of this disclosure.

Claims

1. An electronic device comprising: microphone; The memory is configured to store one or more training data generation models and artificial intelligence models; The one or more training data generation models include: a personal training data generation model trained to generate personal training data reflecting the characteristics of users of the electronic device; and a general training data generation model trained to generate general training data corresponding to usage data of multiple users; and The processor is configured as follows: The general training data is generated using the general training data generation model; The personal training data generation model is used to generate personal training data that reflects the characteristics of the user of the electronic device; The learning start condition is determined based on the predetermined number of identification errors occurring in the artificial intelligence model. Based on meeting the learning initiation conditions, the artificial intelligence model is trained using the personal training data, the general training data, and actual user data obtained from the user of the electronic device as training data; and The artificial intelligence model trained using the individual training data and the general training data is stored in the memory. The artificial intelligence model, trained using the personal training data and the general training data stored in the memory, performs recognition on the input data received through the microphone. The artificial intelligence model is trained to be a personalized AI model that reflects the characteristics of the user of the electronic device. The artificial intelligence model is updated based on the individual training data, the general training data, and the actual user data obtained from the users of the electronic device. The personal training data generation model is updated based on at least one of the user data or the personal training data. The user data refers to the raw data to be identified by the artificial intelligence model, and this user data differs from the individual training data and the general training data. The general training data includes the first input data. The personal training data includes the second input data, and The artificial intelligence model generates first output data corresponding to the first input data based on the first input data, and generates second output data corresponding to the second input data based on the second input data. It is trained using user data, the first input data, the first output data, the second input data, and the second output data. The artificial intelligence model is a speech recognition model, and the personal training data generation model is a model trained using the user's speech data from the electronic device. Both the first input data and the second input data include speech data. Both the first output data and the second output data include text data corresponding to the voice data. 2.The electronic device of claim 1, wherein, The general training data includes the first input data. The personal training data includes the second input data, and The artificial intelligence model performs unsupervised learning based on user data, first input data, and second input data. 3.The electronic device of claim 1, wherein The general training data generation model is downloaded from the server and stored in the memory. 4.The electronic device of claim 3, wherein, The processor is configured to upload artificial intelligence models to the server. 5.The electronic device of claim 1, wherein The processor is configured to train an artificial intelligence model based on the electronic device being in a charging state, the occurrence of a predetermined time, or the absence of user manipulation of the electronic device within a predetermined time. 6.A control method performed by an electronic device, wherein The electronic device includes a processor and a memory configured to store one or more training data generation models and artificial intelligence models as neural networks. The one or more training data generation models include: a personal training data generation model trained to generate personal training data reflecting the characteristics of the user of the electronic device, and a general training data generation model trained to generate general training data corresponding to the usage data of multiple users. The control method includes: The general training data is generated using the general training data generation model; The personal training data generation model is used to generate personal training data that reflects the characteristics of the user of the electronic device. The learning start condition is determined based on the predetermined number of identification errors occurring in the artificial intelligence model. Based on meeting the learning initiation conditions, the artificial intelligence model is trained using the personal training data, the general training data, and actual user data obtained from the user of the electronic device as training data; and The artificial intelligence model trained using the individual training data and the general training data is stored in the memory. The artificial intelligence model, trained using the personal training data and the general training data stored in the memory, performs recognition on input data received via a microphone, wherein the artificial intelligence model is trained to be a personalized artificial intelligence model reflecting the characteristics of the user of the electronic device. The artificial intelligence model is updated based on the individual training data, the general training data, and the actual user data obtained from the users of the electronic device. The personal training data generation model is updated based on at least one of the user data or the personal training data. The user data refers to the raw data to be identified by the artificial intelligence model, and this user data differs from the individual training data and the general training data. The general training data includes the first input data. The personal training data includes the second input data, and The artificial intelligence model generates first output data corresponding to the first input data based on the first input data, and generates second output data corresponding to the second input data based on the second input data. It is trained using user data, the first input data, the first output data, the second input data, and the second output data. The artificial intelligence model is a speech recognition model, and the personal training data generation model is a model trained using the user's speech data from the electronic device. Both the first input data and the second input data include speech data. Both the first output data and the second output data include text data corresponding to the voice data.

Citation Information

Patent Citations

  • Speaker verification and identification using artificial neural network-based sub-phonetic unit discrimination

    US20140195236A1

  • Usage modeling

    US20150371023A1