Information processing method, information processing device, and information processing program
By clustering and selecting training data based on proximity to cluster representatives, the method addresses biased data selection in replay buffers, enhancing AI model generalization through continuous learning.
Patent Information
- Application Number
- PCT/JP2025/004075
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-18
- Filing Date
- 2025-02-07
- Publication Date
- 2025-09-18
AI Technical Summary
Conventional continual learning methods suffer from biased training data selection in replay buffers, leading to degraded generalization performance of AI models due to catastrophic forgetting.
Classify learning data into clusters using unsupervised clustering methods and select data closest to cluster representatives for storage in a buffer, combining with new task data for continuous learning.
Prevents bias in the buffer data and improves the generalization performance of AI models by ensuring diverse and important training data is used for continuous learning.
Smart Images

Figure JP2025004075_18092025_PF_FP_ABST
Abstract
Description
Information processing method, information processing device, and information processing program
[0001] The present disclosure relates to a technology for continuously training an AI model that has trained a first task to perform a second task.
[0002] Continual learning is a well-known technique that allows a person to learn a series of tasks without forgetting previously acquired knowledge.
[0003] For example, Non-Patent Document 1 discloses Lifelong Unsupervised Mixup (LUMP), a continuous learning method that uses unlabeled data. LUMP uses a replay buffer that stores part of the training data used in learning past tasks to suppress catastrophic forgetting.
[0004] However, in the above-mentioned conventional technology, the training data to be stored in the replay buffer is randomly selected, which may result in bias in the training data stored in the replay buffer, and may degrade the generalization performance of the model.
[0005] Divyam Madaan, 4 others, “Representational Continuity For Unsupervised Continual Learning”, ICLR 2022, January 29, 2022
[0006] The present disclosure has been made to solve the above problems, and aims to provide a technology that can suppress bias in the training data stored in the buffer and improve the generalization performance of the AI model.
[0007] An information processing method according to one aspect of the present disclosure is an information processing method executed by an information processing device that causes an artificial intelligence (AI) model that has been trained on a first task to continuously learn a second task, and includes: classifying a plurality of first learning data used in learning the first task into a predetermined number of clusters using an unsupervised clustering method; selecting, from the plurality of first learning data, the predetermined number of first learning data that are closest to representative points of each of the predetermined number of clusters; storing the selected predetermined number of first learning data in a buffer; and causing the AI model to continuously learn the second task using the predetermined number of first learning data stored in the buffer and a plurality of second learning data used in learning the second task.
[0008] According to the present disclosure, it is possible to prevent bias in the learning data stored in the buffer and improve the generalization performance of the AI model.
[0009] FIG. 1 is a diagram illustrating a configuration of an information processing device according to the present embodiment. FIG. 2 is a diagram illustrating learning of a first task according to the present embodiment. FIG. 3 is a diagram illustrating processing for storing a predetermined number of first learning data out of a plurality of first learning data in a replay buffer according to the present embodiment. FIG. 4 is a diagram illustrating learning of a second task according to the present embodiment. FIG. 5 is a flowchart illustrating learning processing of a first task by an information processing device according to an embodiment of the present disclosure. FIG. 6 is a flowchart illustrating learning processing of a second task by an information processing device according to an embodiment of the present disclosure.
[0010] (Findings underlying the present disclosure) Conventional AI learning processes discard learning data used in past learning and use learning data appropriate for a new task for learning. In this way, in conventional learning processes, the learning data changes every time the task changes, which can lead to catastrophic forgetting, where the learning data adapts to the new task data and becomes unable to solve previously learned tasks. LUMP, as described above, combines the learning data for a new task with the learning data for previously learned tasks stored in a replay buffer, and uses the combined learning data to learn the new task. This suppresses catastrophic forgetting in LUMP.
[0011] However, because LUMP randomly selects the training data to be stored in the replay buffer, there is a risk that the training data stored in the replay buffer will be biased, which could result in a deterioration in the generalization performance of the model.
[0012] In order to solve the above problems, the following techniques are disclosed.
[0013] (1) An information processing method according to one aspect of the present disclosure is an information processing method executed by an information processing device that causes an artificial intelligence (AI) model that has been trained on a first task to continuously learn a second task, and includes: classifying a plurality of first learning data used in learning the first task into a predetermined number of clusters using an unsupervised clustering method; selecting, from the plurality of first learning data, the predetermined number of first learning data that are closest to representative points of each of the predetermined number of clusters; storing the selected predetermined number of first learning data in a buffer; and causing the AI model to continuously learn the second task using the predetermined number of first learning data stored in the buffer and a plurality of second learning data used in learning the second task.
[0014] According to this configuration, the plurality of first learning data used in learning the first task are classified into a predetermined number of clusters by an unsupervised clustering method, and a predetermined number of the plurality of first learning data that are closest to the representative points of each of the predetermined number of clusters are stored in the buffer, thereby preventing bias in the learning data stored in the buffer. Furthermore, the second task is continuously learned for the AI model using the predetermined number of first learning data stored in the buffer and classified into the predetermined number of clusters, and the plurality of second learning data used in learning the second task, thereby improving the generalization performance of the AI model.
[0015] (2) The information processing method described in (1) above may further include calculating gradient features of each of the plurality of first learning data based on a function of the AI model, wherein classifying the predetermined number of clusters includes classifying the plurality of gradient features into the predetermined number of clusters by the unsupervised clustering method, and selecting the predetermined number of first learning data may include selecting, from the plurality of first learning data, the predetermined number of first learning data corresponding to the predetermined number of gradient features that are closest to representative points of each of the predetermined number of clusters.
[0016] According to this configuration, the gradient features of each of the multiple first training data are calculated, and the multiple gradient features are classified into a predetermined number of clusters in the gradient space by an unsupervised clustering method, making it possible to sample training data that is highly important and diverse and necessary for training an AI model.
[0017] (3) In the information processing method described in (2) above, the AI model may include an input layer, a plurality of intermediate layers, and an output layer, and the calculation of the gradient features may include calculating the plurality of gradient features based on the function possessed by an intermediate layer among the plurality of intermediate layers that is close to the output layer.
[0018] According to this configuration, multiple gradient features are calculated based on the function of the intermediate layer closest to the output layer among multiple intermediate layers, so that continuous learning can be performed taking into account the variability in the output data of the AI model.
[0019] (4) In the information processing method described in (2) above, the AI model may include an input layer, a plurality of intermediate layers, and an output layer, and the calculation of the gradient features may include calculating the plurality of gradient features based on the function possessed by an intermediate layer among the plurality of intermediate layers that is close to the input layer.
[0020] According to this configuration, multiple gradient features are calculated based on the function of the intermediate layer closest to the input layer among multiple intermediate layers, so that continuous learning can be performed taking into account the variability in the input data of the AI model.
[0021] (5) The information processing method according to any one of (1) to (4) above, further comprising: reducing the predetermined number of first learning data stored in the buffer to a first predetermined number of first learning data; classifying the plurality of second learning data used in learning the second task into a second predetermined number of clusters by the unsupervised clustering method; selecting, from the plurality of second learning data, the second predetermined number of second learning data that are closest to representative points of each of the second predetermined number of clusters; and storing the selected second predetermined number of second learning data in the buffer, wherein the sum of the first predetermined number and the second predetermined number may be equal to or less than a maximum number of learning data that can be stored in the buffer.
[0022] According to this configuration, a first predetermined number of first learning data and a second predetermined number of second learning data are stored in the buffer, so that new tasks from the second task onwards can be continuously learned using the first learning data and second learning data used in learning the past first task and second task.
[0023] (6) In the information processing method described in (5) above, the first predetermined number and the second predetermined number may be the same number.
[0024] According to this configuration, the first predetermined number and the second predetermined number are the same number, so that the learning data for each task can be stored in the buffer without variation.
[0025] (7) In the information processing method described in (5) above, the method may further include inputting a plurality of first evaluation data used to evaluate the first task into the AI model that has undergone continuous training, and obtaining a first output result from the AI model for each of the plurality of first evaluation data; inputting a plurality of second evaluation data used to evaluate the second task into the AI model that has undergone continuous training, and obtaining a second output result from the AI model for each of the plurality of second evaluation data; and determining the first predetermined number and the second predetermined number according to the ratio between the number of errors in the first output result and the number of errors in the second output result.
[0026] According to this configuration, learning data with a large number of errors is stored in the buffer in larger amounts than learning data with a small number of errors, thereby making it possible to prevent past knowledge from being forgotten.
[0027] (8) In the information processing method described in (1) above, the first learning data and the second learning data may be image data, and the first task and the second task may recognize which of a plurality of classes the input image data belongs to.
[0028] This configuration enables continuous learning of the second task of recognizing which of a plurality of classes input image data belongs to.
[0029] (9) In the information processing method described in (1) above, the continuous learning of the second task may be self-supervised learning.
[0030] According to this configuration, the continuous learning of the second task is self-supervised learning, so there is no need to assign labels to the learning data, and the cost of generating the learning data can be reduced.
[0031] Furthermore, the present disclosure can be realized not only as an information processing method that executes the characteristic processes described above, but also as an information processing device having a characteristic configuration corresponding to the characteristic processes executed by the information processing method. Furthermore, the present disclosure can also be realized as a computer program that causes a computer to execute the characteristic processes included in such an information processing method. Therefore, the same effects as those of the above information processing method can also be achieved in the following other aspects.
[0032] (10) Another aspect of the present disclosure provides an information processing device that continuously trains an artificial intelligence (AI) model that has been trained on a first task to perform a second task, the information processing device comprising a processor and a buffer, wherein the processor classifies a plurality of first learning data used in training the first task into a predetermined number of clusters using an unsupervised clustering method, selects the predetermined number of first learning data from the plurality of first learning data that are closest to representative points of each of the predetermined number of clusters, stores the selected predetermined number of first learning data in the buffer, and causes the AI model to continuously train the second task using the predetermined number of first learning data stored in the buffer and a plurality of second learning data used in training the second task.
[0033] (11) Another aspect of the present disclosure provides an information processing program for causing an artificial intelligence (AI) model that has been trained on a first task to continuously learn a second task, the information processing program including: classifying a plurality of first learning data used in learning the first task into a predetermined number of clusters using an unsupervised clustering method; selecting, from the plurality of first learning data, a predetermined number of first learning data that are closest to representative points of each of the predetermined number of clusters; storing the selected predetermined number of first learning data in a buffer; and causing a computer to function such that the AI model continuously learns the second task using the predetermined number of first learning data stored in the buffer and a plurality of second learning data used in learning the second task.
[0034] (12) A non-transitory computer-readable recording medium according to another aspect of the present disclosure records the information processing program described in (11) above.
[0035] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept are described as optional components. Furthermore, in all embodiments, the respective contents can be combined.
[0036] (Embodiment) FIG. 1 is a diagram showing the configuration of an information processing device 1 according to this embodiment.
[0037] The information processing device 1 causes an artificial intelligence (AI) model that has learned a first task to continuously learn a second task. First, the information processing device 1 causes the AI model to learn the first task. Then, the information processing device 1 causes the AI model to continuously learn the second task. The information processing device 1 includes a processor 11 and a memory 12.
[0038] The information processing device 1 includes at least a computer system including, for example, a control program, a processing circuit such as a processor or logic circuit that executes the control program, and a recording device such as an internal memory or an accessible external memory that stores the control program. The information processing device 1 may be implemented, for example, by hardware implementation using the processing circuit, or by execution of a software program stored in the memory by the processing circuit or distributed from an external server, or by a combination of these hardware and software implementations. The information processing device 1 may also be a server, a terminal, or a system including a server and a terminal.
[0039] The processor 11 is, for example, a CPU (Central Processing Unit). The processor 11 includes a learning data acquisition unit 101, a synthesis unit 102, a learning unit 103, a gradient calculation unit 104, a clustering unit 105, a selection unit 106, and a storage processing unit 107. The learning data acquisition unit 101, the synthesis unit 102, the learning unit 103, the gradient calculation unit 104, the clustering unit 105, the selection unit 106, and the storage processing unit 107 may be realized by the processor 11 executing an information processing program, or may be configured by a dedicated hardware circuit such as an ASIC. The information processing program may be recorded on a non-transitory computer-readable recording medium.
[0040] The memory 12 is configured as a non-volatile rewritable storage device such as a hard disk drive or a solid state drive. The memory 12 stores various information. The memory 12 includes a replay buffer 201. The replay buffer 201 is an example of a buffer.
[0041] In learning the first task, the learning data acquisition unit 101 acquires a plurality of first learning data to be used for learning the first task. Note that the learning data acquisition unit 101 may read, from the memory 12, the plurality of first learning data pre-stored in the memory 12. Alternatively, the learning data acquisition unit 101 may acquire the plurality of first learning data received from an external device via a communication unit (not shown). Alternatively, the learning data acquisition unit 101 may read, from the auxiliary storage device, the plurality of first learning data stored in the auxiliary storage device, such as a USB (Universal Serial Bus) memory.
[0042] The first learning data is image data. The first task recognizes which of a plurality of classes the input image data belongs to. For example, if the first task recognizes which of two classes the input image data belongs to, the first learning data includes two image data items that are recognized as each of the two classes.
[0043] FIG. 2 is a diagram for explaining the learning of the first task in this embodiment.
[0044] In learning the first task, the learning unit 103 causes the AI model to learn the first task using a plurality of first learning data acquired by the learning data acquisition unit 101. The learning of the first task is, for example, unsupervised learning using unlabeled first learning data. The unsupervised learning is, for example, self-supervised learning. The self-supervised learning is, for example, Simsiam. The learning unit 103 creates two pieces of data by applying different data augmentations to the input first learning data, and inputs each of the created two pieces of data to two encoders. Then, the learning unit 103 inputs the output of one of the encoders to a predictor and calculates the cosine similarity between the output of the predictor and the output of the other encoder as a loss. Then, the learning unit 103 backpropagates the gradient to only one of the encoders and updates the encoder.
[0045] The learning process performed by the learning unit 103 may be unsupervised learning using unlabeled learning data, supervised learning using labeled learning data, or reinforcement learning.
[0046] FIG. 3 is a diagram for explaining a process of storing a predetermined number of first learning data out of a plurality of first learning data in the replay buffer 201 in this embodiment.
[0047] In learning the first task, the gradient calculation unit 104 calculates gradient features of each of the multiple first learning data acquired by the learning data acquisition unit 101 based on a function of the AI model. The gradient features are gradients calculated using stochastic gradient descent. The gradient calculation unit 104 calculates, as the gradient feature, a derivative f′(x) obtained by differentiating a function f(x) that inputs the first learning data as a variable x. In the gradient space 301 of FIG. 3 , in learning the first task of classifying image data into either a first class or a second class, the gradient features of the first learning data classified into the first class are represented by white circles, and the gradient features of the first learning data classified into the second class are represented by black circles.
[0048] The AI model is a multi-layered neural network that includes an input layer, multiple intermediate layers, and an output layer. The gradient calculation unit 104 may calculate multiple gradient features based on a function of an intermediate layer that is closest to the output layer among the multiple intermediate layers. The gradient calculation unit 104 may calculate multiple gradient features based on a function of an intermediate layer that is closest to the output layer among the multiple intermediate layers. This allows continuous learning to be performed while taking into account variations in the output data of the AI model.
[0049] Furthermore, the gradient calculation unit 104 may calculate the plurality of gradient features based on a function of an intermediate layer that is closest to the input layer among the plurality of intermediate layers. The gradient calculation unit 104 may calculate the plurality of gradient features based on a function of an intermediate layer that is closest to the input layer among the plurality of intermediate layers. This allows continuous learning to be performed while taking into account variations in the input data of the AI model.
[0050] In learning the first task, the clustering unit 105 classifies the plurality of first learning data used in learning the first task into a predetermined number of clusters using an unsupervised clustering method. That is, the clustering unit 105 classifies the plurality of gradient features calculated by the gradient calculation unit 104 into a predetermined number of clusters using an unsupervised clustering method. The unsupervised clustering method is, for example, the k-means method. The predetermined number is the maximum number of data that can be stored in the replay buffer 201. In the gradient space 302 of FIG. 3, the plurality of gradient features are classified into five clusters, and the center of gravity of each cluster is represented by an X mark.
[0051] In this embodiment, the clustering unit 105 may calculate the feature amounts of each of the plurality of first learning data, and classify the calculated feature amounts into a predetermined number of clusters by an unsupervised clustering method.
[0052] In learning the first task, the selection unit 106 selects, from the plurality of first learning data, a predetermined number of first learning data that are closest to the representative points of the predetermined number of clusters classified by the clustering unit 105. The representative points are, for example, centroids. That is, from the plurality of first learning data, the selection unit 106 selects a predetermined number of first learning data that correspond to a predetermined number of gradient features that are closest to the representative points of the predetermined number of clusters classified by the clustering unit 105. In the gradient space 303 of FIG. 3 , the five gradient features closest to the centroids of the five clusters are represented by black circles.
[0053] In learning the first task, the storage processing unit 107 stores a predetermined number of first learning data selected by the selection unit 106 in the replay buffer 201. In FIG. 3 , five first learning data corresponding to five gradient features are stored in the replay buffer 201.
[0054] The replay buffer 201 stores a predetermined number of first learning data.
[0055] In learning the second task, the learning data acquisition unit 101 acquires a plurality of second learning data to be used for learning the second task. Note that the learning data acquisition unit 101 may read, from the memory 12, the plurality of second learning data pre-stored in the memory 12. Alternatively, the learning data acquisition unit 101 may acquire the plurality of second learning data received from an external device via a communication unit (not shown). Alternatively, the learning data acquisition unit 101 may read, from the auxiliary storage device, the plurality of second learning data stored in the auxiliary storage device, such as a USB (Universal Serial Bus) memory.
[0056] The second learning data is image data. The second task recognizes which of multiple classes the input image data belongs to. For example, if the second task recognizes which of two classes the input image data belongs to, the second learning data includes two image data that are recognized as each of the two classes. The multiple classes classified in the second task are different from the multiple classes classified in the first task. For example, if the AI model recognizes N classes in the first task and M classes in the second task, and continuous learning of the second task is performed after learning of the first task, the AI model will be able to recognize N+M classes.
[0057] Furthermore, the first task and the second task may be any of object recognition, which recognizes objects appearing in the image data, object detection, which identifies objects appearing in the image data and locates the objects, and area classification, which identifies objects for each pixel in the image data.
[0058] The first learning data and the second learning data may be voice data. The first learning data and the second learning data may be sensing data such as acceleration data measured by an acceleration sensor worn by the user or heart rate data measured by a heart rate monitor worn by the user. In this case, the first task and the second task may be a task for recognizing the user's behavior. The first learning data and the second learning data may be continuous values. The first learning data and the second learning data may be text data. In this case, the first task and the second task may be a natural language processing task for processing text data.
[0059] In learning the second task, the synthesis unit 102 creates multiple pieces of synthetic learning data by synthesizing each of a predetermined number of first learning data stored in the replay buffer 201 with each of the multiple pieces of second learning data used in learning the second task. Here, the synthesis unit 102 sequentially extracts one piece of second learning data from the multiple pieces of second learning data, and randomly extracts one piece of first learning data from the predetermined number of pieces of first learning data stored in the replay buffer 201. Then, the synthesis unit 102 creates synthetic learning data by linearly interpolating the extracted one piece of second learning data and the extracted one piece of first learning data. The synthesis unit 102 creates as many pieces of synthetic learning data as there are multiple pieces of second learning data.
[0060] FIG. 4 is a diagram for explaining the learning of the second task in this embodiment.
[0061] In learning the second task, the learning unit 103 continuously trains the AI model on the second task using a predetermined number of first learning data stored in the replay buffer 201 and multiple second learning data used in learning the second task. More specifically, the learning unit 103 continuously trains the AI model on the second task using the composite learning data created by the synthesis unit 102. The learning of the second task is, for example, unsupervised learning using unlabeled composite learning data. The unsupervised learning is, for example, self-supervised learning. The self-supervised learning is, for example, Simsiam. The learning unit 103 performs self-supervised learning by using the first learning data as input data for the first task and composite learning data obtained by combining the second learning data and the first learning data stored in the replay buffer 201 as input data for the second task. The learning of the second task is the same as the learning of the first task.
[0062] In learning the second task, the storage processing unit 107 reduces a predetermined number of first learning data stored in the replay buffer 201 to a first predetermined number of first learning data. Here, the gradient calculation unit 104 calculates gradient features for each of the predetermined number of first learning data stored in the replay buffer 201 based on a function of the AI model. The clustering unit 105 classifies the predetermined number of gradient features calculated by the gradient calculation unit 104 into a first predetermined number of clusters using an unsupervised clustering method. The selection unit 106 selects, from the predetermined number of first learning data, a first predetermined number of first learning data corresponding to a first predetermined number of gradient features that are closest to the representative points of each of the first predetermined number of clusters classified by the clustering unit 105. The storage processing unit 107 stores the first predetermined number of first learning data selected by the selection unit 106 in the replay buffer 201. As a result, the predetermined number of first learning data stored in the replay buffer 201 is reduced to the first predetermined number of first learning data.
[0063] In learning the second task, the gradient calculation unit 104 calculates, based on a function of the AI model, gradient features of each of the plurality of second learning data acquired by the learning data acquisition unit 101. The gradient calculation unit 104 calculates, as the gradient feature, a derivative f′(x) obtained by differentiating a function f(x) that inputs the second learning data as a variable x.
[0064] In learning the second task, the clustering unit 105 classifies the plurality of second learning data used in learning the second task into a second predetermined number of clusters using an unsupervised clustering method. That is, the clustering unit 105 classifies the plurality of gradient features calculated by the gradient calculation unit 104 into the second predetermined number of clusters using an unsupervised clustering method. The unsupervised clustering method is, for example, the k-means method. The sum of the first predetermined number and the second predetermined number is less than or equal to the maximum number of learning data that can be stored in the replay buffer 201. The first predetermined number and the second predetermined number are the same number. The first predetermined number and the second predetermined number are half the maximum number of learning data that can be stored in the replay buffer 201.
[0065] In this embodiment, the clustering unit 105 may calculate the feature amounts of each of the plurality of second learning data, and classify the calculated feature amounts into a predetermined number of clusters by an unsupervised clustering method.
[0066] In learning the second task, the selection unit 106 selects, from the plurality of second learning data, a second predetermined number of second learning data that are closest to the representative points of the second predetermined number of clusters classified by the clustering unit 105. The representative points are, for example, center of gravity points. That is, the selection unit 106 selects, from the plurality of second learning data, a second predetermined number of second learning data that correspond to the second predetermined number of gradient features that are closest to the representative points of the second predetermined number of clusters classified by the clustering unit 105.
[0067] In learning the second task, the storage processing unit 107 stores the second predetermined number of second learning data selected by the selection unit 106 in the replay buffer 201. The replay buffer 201 stores the first predetermined number of first learning data and the second predetermined number of second learning data.
[0068] For example, if the maximum number of data items that can be stored in the replay buffer 201 is 128, the replay buffer 201 stores 128 pieces of first learning data in learning the first task. Then, in learning the second task, the storage processing unit 107 reduces the 128 pieces of first learning data stored in the replay buffer 201 to 64 pieces of first learning data. Furthermore, the storage processing unit 107 stores 64 pieces of second learning data in the replay buffer 201. As a result, the replay buffer 201 stores 64 pieces of first learning data and 64 pieces of second learning data.
[0069] Note that the third and subsequent continuous learning may be performed after the second continuous learning. The third and subsequent continuous learning are performed in the same manner as the second continuous learning, but the number of learning data for each task stored in the replay buffer 201 is the same. If the maximum number of learning data that can be stored in the replay buffer 201 is X, in the nth continuous learning of the nth task, the storage processing unit 107 reduces the X / (n-1) pieces of learning data for each task from the first task to the (n-1)th task stored in the replay buffer 201 to X / n pieces of learning data. The clustering unit 105 then classifies the nth learning data used in learning the nth task into X / n clusters. The selection unit 106 selects, from the nth learning data used in learning the nth task, X / n pieces of nth learning data that are closest to the representative points of each of the X / n clusters. The storage processing unit 107 stores the X / n pieces of nth learning data in the replay buffer 201.
[0070] Next, the learning process of the first task performed by the information processing device 1 according to the embodiment of the present disclosure will be described.
[0071] FIG. 5 is a flowchart illustrating the learning process of the first task performed by the information processing device 1 according to the embodiment of the present disclosure.
[0072] First, in step S1, the learning data acquisition unit 101 acquires a plurality of first learning data to be used for learning the first task.
[0073] Next, in step S2, the learning unit 103 causes the AI model to learn the first task using the plurality of first learning data acquired by the learning data acquisition unit 101. At this time, the learning unit 103 extracts one piece of first learning data from the plurality of first learning data acquired by the learning data acquisition unit 101. The learning unit 103 causes the AI model to learn the first task using the extracted one piece of first learning data. Then, the learning unit 103 repeats the extraction of one piece of first learning data and the learning of the first task using the extracted one piece of first learning data until learning of the first task using all of the first learning data is completed.
[0074] Next, in step S3, the gradient calculation unit 104 calculates the gradient feature of each of the multiple first learning data acquired by the learning data acquisition unit 101 based on the function of the AI model.
[0075] Next, in step S4, the clustering unit 105 classifies the plurality of gradient features calculated by the gradient calculation unit 104 into a predetermined number of clusters using an unsupervised clustering method.
[0076] Next, in step S5, the selection unit 106 selects, from the plurality of first learning data, a predetermined number of first learning data corresponding to a predetermined number of gradient features that are closest to the representative points of each of the predetermined number of clusters classified by the clustering unit 105.
[0077] Next, in step S6 , the storage processing unit 107 stores the predetermined number of first learning data selected by the selection unit 106 in the replay buffer 201 .
[0078] Next, the learning process for the second task performed by the information processing device 1 according to the embodiment of the present disclosure will be described.
[0079] FIG. 6 is a flowchart illustrating the learning process of the second task performed by the information processing device 1 according to the embodiment of the present disclosure.
[0080] First, in step S11, the learning data acquisition unit 101 acquires a plurality of second learning data to be used for learning the second task.
[0081] Next, in step S12 , the synthesis unit 102 extracts one piece of second learning data from the plurality of pieces of second learning data acquired by the learning data acquisition unit 101 .
[0082] Next, in step S13 , the synthesis unit 102 randomly extracts one piece of first learning data from a predetermined number of pieces of first learning data stored in the replay buffer 201 .
[0083] Next, in step S14, the combining unit 102 combines the extracted one piece of second learning data with the extracted one piece of first learning data.
[0084] Next, in step S15, the learning unit 103 causes the AI model to continue learning the second task using the synthesized learning data synthesized by the synthesis unit 102.
[0085] Next, in step S16, the synthesis unit 102 determines whether or not all of the second learning data have been extracted, acquired by the learning data acquisition unit 101. If it is determined that all of the second learning data have not been extracted (NO in step S16), the process returns to step S12, where the synthesis unit 102 extracts one piece of second learning data that has not yet been extracted from the remaining second learning data.
[0086] On the other hand, if it is determined that all of the second learning data have been extracted (YES in step S16), in step S17, the gradient calculation unit 104 calculates the gradient features of each of the predetermined number of first learning data stored in the replay buffer 201 based on the function of the AI model.
[0087] Next, in step S18, the clustering unit 105 classifies the predetermined number of gradient features calculated by the gradient calculation unit 104 into a first predetermined number of clusters by the k-means method.
[0088] Next, in step S19, the selection unit 106 selects, from the predetermined number of first learning data, a first predetermined number of first learning data corresponding to a first predetermined number of gradient features that are closest to the representative points of each of the first predetermined number of clusters classified by the clustering unit 105.
[0089] Next, in step S20 , the storage processing unit 107 stores the first predetermined number of first learning data selected by the selection unit 106 in the replay buffer 201 .
[0090] Next, in step S21, the gradient calculation unit 104 calculates the gradient feature of each of the plurality of second learning data acquired by the learning data acquisition unit 101 based on the function of the AI model.
[0091] Next, in step S22, the clustering unit 105 classifies the plurality of gradient features calculated by the gradient calculation unit 104 into a second predetermined number of clusters using the k-means method.
[0092] Next, in step S23, the selection unit 106 selects, from the plurality of second learning data, a second predetermined number of second learning data corresponding to a second predetermined number of gradient features that are closest to the representative points of each of the second predetermined number of clusters classified by the clustering unit 105.
[0093] Next, in step S24, the storage processing unit 107 stores the second predetermined number of second learning data selected by the selection unit 106 in the replay buffer 201. As a result, the first predetermined number of first learning data and the second predetermined number of second learning data are stored in the replay buffer 201. When learning of the third task is performed, the synthesis unit 102 randomly extracts one piece of first learning data or one piece of second learning data from the first predetermined number of first learning data and the second predetermined number of second learning data stored in the replay buffer 201. The synthesis unit 102 synthesizes the extracted piece of third learning data with the extracted piece of first learning data or one piece of second learning data.
[0094] In this way, the plurality of first learning data used in learning the first task are classified into a predetermined number of clusters by the unsupervised clustering method, and a predetermined number of the plurality of first learning data that are closest to the representative points of each of the predetermined number of clusters are stored in the replay buffer 201, thereby preventing bias from occurring in the learning data stored in the replay buffer 201. Furthermore, the second task is continuously learned for the AI model using the predetermined number of first learning data stored in the replay buffer 201 and classified into the predetermined number of clusters, and the plurality of second learning data used in learning the second task, thereby improving the generalization performance of the AI model.
[0095] In the present embodiment, the first predetermined number and the second predetermined number are the same, but the present disclosure is not limited thereto, and the first predetermined number and the second predetermined number may be different numbers. The processor 11 may further include a data number determination unit. After the continuous learning of the second task is completed, the data number determination unit inputs a plurality of first evaluation data used to evaluate the first task to the continuously trained AI model and obtains a first output result from the AI model for each of the plurality of first evaluation data. The data number determination unit then inputs a plurality of second evaluation data used to evaluate the second task to the continuously trained AI model and obtains a second output result from the AI model for each of the plurality of second evaluation data. The data number determination unit then determines the first predetermined number and the second predetermined number based on the ratio between the number of errors in the first output result and the number of errors in the second output result. In the object identification task, the number of errors represents the number of times an object of class A is mistakenly recognized as a different class B. In an object detection task, the number of errors represents the number of areas other than the target object that were detected or the number of times an object to be detected was not detected. For example, if the number of errors in the first output result is 3 and the number of errors in the second output result is 2, the data number determination unit determines the first and second predetermined numbers so that the ratio of the first and second predetermined numbers is 3:2. In other words, the more errors a task has, the more learning data it may store in the replay buffer 201.
[0096] Note that some or all of the functions of the device according to the embodiment of the present disclosure may be realized by a processor such as a CPU executing a program.
[0097] Furthermore, all the numbers used above are merely examples to specifically explain the present disclosure, and the present disclosure is not limited to the numbers used as examples.
[0098] The order in which the steps are executed in the above flowchart is merely an example for specifically explaining the present disclosure, and other orders may be used as long as similar effects are obtained. Also, some of the steps may be executed simultaneously (in parallel) with other steps.
[0099] The technology disclosed herein can suppress bias in the learning data stored in the buffer and improve the generalization performance of the AI model, and is therefore useful as a technology for continuously learning a second task for an AI model that has learned a first task.
Claims
1. An information processing method executed by an information processing device that causes an artificial intelligence (AI) model that has been trained on a first task to continuously learn a second task, the information processing method including: classifying a plurality of first learning data used in learning the first task into a predetermined number of clusters using an unsupervised clustering method; selecting, from the plurality of first learning data, the predetermined number of first learning data that are closest to representative points of each of the predetermined number of clusters; storing the selected predetermined number of first learning data in a buffer; and causing the AI model to continuously learn the second task using the predetermined number of first learning data stored in the buffer and a plurality of second learning data used in learning the second task.
2. The information processing method according to claim 1, further comprising calculating gradient features of each of the plurality of first learning data based on a function of the AI model, wherein classifying the predetermined number of clusters comprises classifying the plurality of gradient features into the predetermined number of clusters using the unsupervised clustering method, and selecting the predetermined number of first learning data comprises selecting, from the plurality of first learning data, the predetermined number of first learning data corresponding to the predetermined number of gradient features that are closest to representative points of each of the predetermined number of clusters.
3. The information processing method according to claim 2, wherein the AI model includes an input layer, multiple intermediate layers, and an output layer, and the calculation of the gradient features includes calculating the multiple gradient features based on the function of an intermediate layer that is closest to the output layer among the multiple intermediate layers.
4. The information processing method according to claim 2, wherein the AI model includes an input layer, multiple intermediate layers, and an output layer, and the calculation of the gradient features includes calculating the multiple gradient features based on the function possessed by an intermediate layer, among the multiple intermediate layers, that is closest to the input layer.
5. The information processing method according to any one of claims 1 to 4, further comprising: reducing the predetermined number of first learning data stored in the buffer to a first predetermined number of first learning data; classifying the plurality of second learning data used in learning the second task into a second predetermined number of clusters by the unsupervised clustering method; selecting, from the plurality of second learning data, the second predetermined number of second learning data that are closest to representative points of each of the second predetermined number of clusters; and storing the selected second predetermined number of second learning data in the buffer, wherein the sum of the first predetermined number and the second predetermined number is equal to or less than the maximum number of learning data that can be stored in the buffer.
6. The information processing method according to claim 5, wherein the first predetermined number and the second predetermined number are the same number.
7. The information processing method of claim 5, further comprising: inputting a plurality of first evaluation data used to evaluate the first task into the AI model that has undergone continuous training, and obtaining a first output result from the AI model for each of the plurality of first evaluation data; inputting a plurality of second evaluation data used to evaluate the second task into the AI model that has undergone continuous training, and obtaining a second output result from the AI model for each of the plurality of second evaluation data; and determining the first predetermined number and the second predetermined number according to the ratio between the number of errors in the first output result and the number of errors in the second output result.
8. An information processing method according to any one of claims 1 to 4, wherein the first learning data and the second learning data are image data, and the first task and the second task recognize which of a plurality of classes the input image data belongs to.
9. The information processing method according to any one of claims 1 to 4, wherein the continuous learning of the second task is self-supervised learning.
10. An information processing device that causes an artificial intelligence (AI) model that has been trained on a first task to continue learning a second task, comprising: a processor; and a buffer, wherein the processor classifies a plurality of first learning data used in learning the first task into a predetermined number of clusters using an unsupervised clustering method; selects, from the plurality of first learning data, the predetermined number of first learning data that are closest to representative points of each of the predetermined number of clusters; stores the selected predetermined number of first learning data in the buffer; and causes the AI model to continue learning the second task using the predetermined number of first learning data stored in the buffer and a plurality of second learning data used in learning the second task.
11. An information processing program for causing an artificial intelligence (AI) model that has been trained on a first task to continuously learn a second task, the information processing program comprising the steps of: classifying a plurality of first learning data used in learning the first task into a predetermined number of clusters using an unsupervised clustering method; selecting, from the plurality of first learning data, the predetermined number of first learning data that are closest to representative points of each of the predetermined number of clusters; storing the selected predetermined number of first learning data in a buffer; and causing a computer to function so as to cause the AI model to continuously learn the second task using the predetermined number of first learning data stored in the buffer and a plurality of second learning data used in learning the second task.
Citation Information
Patent Citations
Continuous learning device, continuous learning method, and continuous learning program
JP2023160198A