Information processing method, information processing apparatus, and information processing program
The information processing method addresses the limitations of existing model generation techniques by optimizing the learning process and model architecture, resulting in more flexible and accurate foundation models.
Patent Information
- Application Number
- JP2024019684
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-27
- Filing Date
- 2024-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-02-13
AI Technical Summary
Existing techniques for generating models, such as language models, face limitations in flexibility and effectiveness, particularly in learning industry-specific expressions and fine-tuning models for specific tasks.
An information processing method that acquires training data to learn a foundation model through multiple stages of processing, optimizing the learning method and model architecture to improve model generation flexibility and accuracy.
The method enables the proper generation of foundation models, enhancing model learning and flexibility, and improving the accuracy of models for various tasks.
Smart Images

Figure 2025073953000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing method, an information processing device, and an information processing program. [Background technology]
[0002] In recent years, a technology has been proposed for generating models by having various models, such as language models, learn features of training data. Models, such as language models, trained in this way are used for various inference processes, such as various predictions and classifications. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2023-072863 A Summary of the Invention [Problem to be solved by the invention]
[0004] In addition, the above-mentioned technology has room for improvement in model generation. For example, in the above-mentioned example, a model that can accurately interpret an expression using a sentence containing an industry-specific expression as input can be generated, but there is room for improvement in the learning of the model, and it is desired to generate a model more flexibly. For example, there is room for improvement in the learning of a base model that can be used as part of a model for various applications, and a fine-tuned model that is fine-tuned to be applied to a specific task. Therefore, for example, it is desired to appropriately generate a base model. [Means for solving the problem]
[0005] The information processing method of the present application is an information processing method executed by a computer, and is characterized in that it includes an acquisition process for acquiring learning data used to train a base model using text as input, and a generation process for generating the base model through a multi-stage learning process using the learning data. Effect of the Invention
[0006] According to one aspect of the embodiment, the foundation model can be appropriately generated. [Brief description of the drawings]
[0007] [Figure 1] FIG. 1 is a diagram illustrating an example of an information processing system according to an embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of a flow of model generation using an information processing device in an embodiment. [Diagram 3] FIG. 1 is a diagram illustrating an example of a configuration of an information processing device according to an embodiment. [Figure 4] FIG. 11 is a diagram showing an example of information registered in a learning data database according to the embodiment. [Diagram 5] 10 is a flowchart illustrating an example of a flow of information processing according to the embodiment. [Figure 6] 10 is a flowchart illustrating an example of a flow of information processing according to the embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of a structure of a model according to an embodiment. [Figure 8] FIG. 13 is a diagram illustrating an example of designation of an input of a model according to the embodiment. [Figure 9] FIG. 13 is a diagram illustrating an example of designation of an input of a model according to the embodiment. [Figure 10] FIG. 4 is a diagram illustrating an example of an input type according to the embodiment. [Figure 11] FIG. 13 is a diagram illustrating an example of an input of a model according to the embodiment. [Figure 12] FIG. 13 is a diagram illustrating another example of the structure of the model according to the embodiment. [Figure 13] FIG. 13 is a diagram illustrating an example of an input of a model according to the embodiment. [Figure 14] FIG. 13 is a diagram showing an example of an experimental result. [Figure 15] FIG. 13 is a diagram showing an example of an experimental result. [Figure 16] FIG. 13 is a diagram showing an example of an experimental result. [Figure 17] FIG. 13 is a diagram showing an example of an experimental result. [Figure 18] FIG. 13 is a diagram showing an example of an experimental result. [Figure 19] FIG. 11 is a diagram illustrating an example of a model learning process according to the embodiment. [Figure 20] FIG. 11 is a diagram illustrating an example of a model learning process according to the embodiment. [Figure 21] FIG. 11 is a diagram illustrating an example of a model learning process according to the embodiment. [Figure 22] FIG. 13 is a diagram showing an example of an experimental result. [Figure 23] FIG. 13 is a diagram showing an example of an experimental result. [Figure 24] FIG. 13 is a diagram showing an example of an experimental result. [Diagram 25] FIG. 13 is a diagram showing an example of an experimental result. [Figure 26] FIG. 2 illustrates an example of a hardware configuration. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] Hereinafter, the information processing method, information processing device, and information processing program according to the present application will be described in detail with reference to the drawings. Note that the information processing method, information processing device, and information processing program according to the present application are not limited to these embodiments. In addition, each embodiment can be appropriately combined as long as the processing contents are not contradictory. In addition, the same parts in each of the following embodiments are given the same reference numerals, and duplicated explanations will be omitted.
[0009] [Embodiment] In the following embodiment, first, the premise of the system configuration and the like will be described, and then the process of generating a model using text as input will be described. In this embodiment, before showing the generation of the above-mentioned model, experimental results, and the like, first, the configuration of the information processing system 1 that generates the model will be described.
[0010] [1. Configuration of information processing system] First, a configuration of an information processing system having an information processing device 10, which is an example of an information processing device, will be described with reference to FIG. 1. FIG. 1 is a diagram showing an example of an information processing system according to an embodiment. As shown in FIG. 1, the information processing system 1 has an information processing device 10, a model generation server 2, and a terminal device 3. Note that the information processing system 1 may have a plurality of model generation servers 2 and a plurality of terminal devices 3. The information processing device 10 and the model generation server 2 may be realized by the same server device, a cloud system, or the like. Here, the information processing device 10, the model generation server 2, and the terminal device 3 are connected to each other via a network N (see, for example, FIG. 3) so as to be able to communicate with each other by wire or wirelessly.
[0011] The information processing device 10 is an information processing device that executes an index generation process that generates generation indicators, which are indicators for model generation (i.e., a recipe for a model), and a model generation process that generates a model according to the generation indicators, and provides the generated generation indicators and models, and is realized, for example, by a server device, a cloud system, etc.
[0012] The model generation server 2 is an information processing device that generates a model that has learned the characteristics of the learning data, and is realized by, for example, a server device or a cloud system. For example, when the model generation server 2 receives a configuration file that indicates the type and behavior of the model to be generated and how to learn the characteristics of the learning data as a model generation index, the model generation server 2 automatically generates a model according to the received configuration file. Note that the model generation server 2 may learn the model using any model learning method. Also, for example, the model generation server 2 may be any of various existing services such as AutoML (Automated Machine Learning).
[0013] The terminal device 3 is a terminal device used by a user U, and is realized by, for example, a PC (Personal Computer), a server device, etc. For example, the terminal device 3 generates a generation index of a model through an exchange with the information processing device 10, and acquires a model generated by the model generation server 2 according to the generated generation index.
[0014] 2. Overview of Processing Executed by Information Processing Device 10 First, an overview of the process executed by the information processing device 10 will be described. First, the information processing device 10 receives, from the terminal device 3, an indication of learning data for causing the model to learn features (step S1). For example, the information processing device 10 stores various learning data used for learning in a predetermined storage device, and receives an indication of learning data designated by the user U as the learning data. Note that the information processing device 10 may acquire the learning data used for learning, for example, from the terminal device 3 or various external servers.
[0015] Here, any data can be used as the learning data. For example, the information processing device 10 may use various information related to users, such as the location history of each user, the history of web content viewed by each user, the purchase history and search query history of each user, as the learning data. The information processing device 10 may also use demographic attributes and psychographic attributes of users as the learning data. The information processing device 10 may also use metadata such as the type, content, and creator of various web contents to be distributed as the learning data.
[0016] In such a case, the information processing device 10 generates candidates for generation indicators based on statistical information of the learning data used for learning (step S2). For example, the information processing device 10 generates candidates for generation indicators indicating what model should be used for learning and what learning method should be used for learning, based on the characteristics of the values included in the learning data. In other words, the information processing device 10 generates, as generation indicators, models capable of accurately learning the characteristics of the learning data and learning methods for making the models accurately learn the characteristics. That is, the information processing device 10 optimizes the learning method. Note that the content of the generation indicators to be generated when what type of learning data is selected will be described later.
[0017] Next, the information processing device 10 provides candidates for the generation index to the terminal device 3 (step S3). In such a case, the user U modifies the candidates for the generation index according to preferences, rules of thumb, etc. (step S4). Then, the information processing device 10 provides each candidate for the generation index and learning data to the model generation server 2 (step S5).
[0018] On the other hand, the model generation server 2 generates a model for each generation index (step S6). For example, the model generation server 2 trains the model having the structure indicated by the generation index to learn the features of the training data by the learning method indicated by the generation index. Then, the model generation server 2 provides the generated model to the information processing device 10 (step S7).
[0019] Here, it is considered that the accuracy of each model generated by the model generation server 2 differs due to the difference in the generation index. Therefore, the information processing device 10 generates a new generation index by a genetic algorithm based on the accuracy of each model (step S8), and repeatedly generates a model using the newly generated generation index (step S9).
[0020] For example, the information processing device 10 divides the learning data into evaluation data and learning data, and acquires a plurality of models that are models that have been trained based on features of the learning data and that are generated according to different generation indices. For example, the information processing device 10 generates 10 generation indices, and generates 10 models using the generated 10 generation indices and the learning data. In such a case, the information processing device 10 measures the accuracy of each of the 10 models using the evaluation data.
[0021] Next, the information processing device 10 selects a predetermined number of models (for example, five) from the ten models in order of increasing accuracy. Then, the information processing device 10 generates new generation indices from the generation indices adopted when generating the five selected models. For example, the information processing device 10 regards each generation index as an individual of a genetic algorithm, and regards the type of model, the structure of the model, and various learning methods (i.e., various indices indicated by the generation index) indicated by each generation index as genes in the genetic algorithm. Then, the information processing device 10 generates ten new generation indices for the next generation by selecting individuals to perform gene crossover and performing gene crossover. Note that the information processing device 10 may take into consideration mutation when performing gene crossover. Also, the information processing device 10 may perform two-point crossover, multi-point crossover, uniform crossover, or random selection of genes to be crossover targets. Also, the information processing device 10 may adjust the crossover rate when performing crossover so that, for example, the higher the accuracy of the model of an individual, the more genes are passed on to the next generation of individuals.
[0022] Furthermore, the information processing device 10 again generates new 10 models using the generation index of the next generation. Then, the information processing device 10 generates new generation indexes by the above-mentioned genetic algorithm based on the accuracy of the new 10 models. By repeatedly executing such processing, the information processing device 10 can bring the generation indexes closer to the generation indexes according to the characteristics of the learning data, that is, the optimized generation indexes.
[0023] Furthermore, when a predetermined condition is met, such as when a new generation index is generated a predetermined number of times or when the maximum, average, or minimum value of the accuracy of the model exceeds a predetermined threshold, the information processing device 10 selects the model with the highest accuracy as the model to be provided. Then, the information processing device 10 provides the selected model and the corresponding generation index to the terminal device 3 (step S10). As a result of such processing, the information processing device 10 can generate a generation index of an appropriate model and provide a model according to the generated generation index, simply by selecting learning data from the user.
[0024] In the above example, the information processing device 10 realizes stepwise optimization of the generation index using a genetic algorithm, but the embodiment is not limited to this. As will be clear from the following description, the accuracy of a model varies greatly depending on the index when generating a model (i.e., when learning the characteristics of the learning data), such as not only the characteristics of the model itself, such as the type and structure of the model, but also what learning data is input to the model and how, and what hyperparameters are used to learn the model.
[0025] Therefore, the information processing device 10 may not perform optimization using a genetic algorithm as long as it generates a generation index estimated to be optimal according to the learning data. For example, the information processing device 10 may present to the user a generation index generated according to whether or not the learning data satisfies various conditions generated according to empirical rules, and generate a model according to the presented generation index. Furthermore, when the information processing device 10 receives a correction to the presented generation index, it may generate a model according to the received corrected generation index, present the accuracy of the generated model to the user, and receive a correction to the generation index again. That is, the information processing device 10 may cause the user U to find an optimal generation index by trial and error.
[0026] [3. About the generation of generated indicators] An example of what kind of generation index is generated for what kind of learning data will be described below. Note that the following example is merely an example, and any process can be adopted as long as it generates a generation index according to the characteristics of the learning data.
[0027] [3-1. About generation indicators] First, an example of information indicated by a generation index will be described. For example, when a model is made to learn features of training data, it is considered that the manner in which the training data is input to the model, the manner of the model, and the learning manner of the model (i.e., the features indicated by the hyperparameters) contribute to the accuracy of the model finally obtained. Therefore, the information processing device 10 improves the accuracy of the model by generating a generation index in which each manner is optimized according to the features of the training data.
[0028] For example, it is considered that the learning data includes data to which various labels are attached, that is, data showing various features. However, if data showing features that are not useful when classifying data is used as the learning data, the accuracy of the finally obtained model may deteriorate. Therefore, the information processing device 10 determines the features of the learning data to be input as a mode when inputting the learning data to the model. For example, the information processing device 10 determines which label-attached data (i.e., data showing which feature) is to be input from among the learning data. In other words, the information processing device 10 optimizes the combination of features to be input.
[0029] In addition, it is considered that the learning data includes columns of various formats, such as data containing only numerical values and data containing character strings. When such learning data is input to a model, it is considered that the accuracy of the model changes depending on whether the learning data is input as is or converted into data of another format. For example, when multiple types of learning data (learning data showing different characteristics) are input, and learning data of character strings and learning data of numerical values are input, it is considered that the accuracy of the model changes depending on whether the character strings and numerical values are input as is, whether the character strings are converted into numerical values and only numerical values are input, and whether the numerical values are regarded as character strings and input. Therefore, the information processing device 10 determines the format of the learning data to be input to the model. For example, the information processing device 10 determines whether the learning data to be input to the model is to be a numerical value or a character string. In other words, the information processing device 10 optimizes the column type of the input feature.
[0030] In addition, when there are learning data showing different features, it is considered that the accuracy of the model changes depending on which combination of features is input simultaneously. That is, when there are learning data showing different features, it is considered that the accuracy of the model changes depending on which combination of features (i.e., the relationship of a combination of multiple features) is learned. For example, when there is learning data showing a first feature (e.g., gender), learning data showing a second feature (e.g., address), and learning data showing a third feature (e.g., purchase history), it is considered that the accuracy of the model changes when the learning data showing the first feature and the learning data showing the second feature are input simultaneously and when the learning data showing the first feature and the learning data showing the third feature are input simultaneously. Therefore, the information processing device 10 optimizes the combination of features (cross future) that allows the model to learn the relationship.
[0031] Here, various models project input data into a space of a predetermined dimension divided by a predetermined hyperplane, and classify the input data according to which of the divided spaces the projected position belongs to. Therefore, if the number of dimensions of the space to which the input data is projected is lower than the optimal number of dimensions, the classification ability of the input data deteriorates, and the accuracy of the model deteriorates. In addition, if the number of dimensions of the space to which the input data is projected is higher than the optimal number of dimensions, the inner product value with the hyperplane changes, and as a result, data different from the data used during learning may not be properly classified. Therefore, the information processing device 10 optimizes the number of dimensions of the input data input to the model. For example, the information processing device 10 optimizes the number of dimensions of the input data by controlling the number of nodes in the input layer of the model. In other words, the information processing device 10 optimizes the number of dimensions of the space to which the input data is embedded.
[0032] In addition to SVM, there are also models such as neural networks with multiple intermediate layers (hidden layers). Various types of neural networks are known, such as feedforward DNNs in which information is transmitted in one direction from the input layer to the output layer, convolutional neural networks (CNNs) that convolve information in intermediate layers, recurrent neural networks (RNNs) with directed loops, and Boltzmann machines. These various neural networks include long short-term memory (LSTM) and various other neural networks.
[0033] In this way, when the type of model that learns various features of the learning data is different, the accuracy of the model is considered to change. Therefore, the information processing device 10 selects a type of model that is estimated to learn the features of the learning data with high accuracy. For example, the information processing device 10 selects a type of model depending on what label is attached as the label of the learning data. To give a more specific example, when data with a term related to "history" attached as a label exists, the information processing device 10 selects an RNN that is considered to be able to learn the features of history better, and when data with a term related to "image" attached as a label exists, the information processing device 10 selects a CNN that is considered to be able to learn the features of the image better. In addition to these, the information processing device 10 may determine whether the label is a term specified in advance or a term similar to a term, and select a model of a type that is previously associated with a term determined to be the same or similar.
[0034] In addition, when the number of intermediate layers of the model or the number of nodes included in one intermediate layer changes, the learning accuracy of the model is considered to change. For example, when the number of intermediate layers of the model is large (when the model is deep), it is considered that classification according to more abstract features can be realized, but as a result, local errors in backpropagation are less likely to propagate to the input layer, and learning may not be performed appropriately. In addition, when the number of nodes included in the intermediate layers is small, more advanced abstraction can be performed, but when the number of nodes is too small, there is a high possibility that information required for classification will be lost. Therefore, the information processing device 10 optimizes the number of intermediate layers and the number of nodes included in the intermediate layers. That is, the information processing device 10 optimizes the architecture of the model.
[0035] In addition, it is considered that the accuracy of the nodes changes depending on whether attention is present or absent, whether the nodes included in the model have autoregression or not, and which nodes are connected. Therefore, the information processing device 10 optimizes the network by determining whether or not the network has autoregression and which nodes are connected.
[0036] Furthermore, when learning a model, the optimization method for the model (algorithm used during learning), the dropout rate, the activation function of the node, the number of units, etc. are set as hyperparameters. It is considered that the accuracy of the model changes when such hyperparameters change. Therefore, the information processing device 10 optimizes the learning mode when learning the model, that is, the hyperparameters.
[0037] Furthermore, when the size of the model (the number of input layers, intermediate layers, and output layers, and the number of nodes) changes, the accuracy of the model also changes. Therefore, the information processing device 10 also optimizes the size of the model.
[0038] In this way, the information processing device 10 optimizes the indices used when generating the various models described above. For example, the information processing device 10 holds in advance conditions corresponding to each index. Note that such conditions are set, for example, based on empirical rules such as the accuracy of various models generated from past learning models. Then, the information processing device 10 determines whether the learning data satisfies each condition, and adopts an index that is previously associated with a condition that the learning data satisfies or does not satisfy as a generation index (or a candidate thereof). As a result, the information processing device 10 can generate a generation index that can accurately learn the features of the learning data.
[0039] As described above, when the generation index is automatically generated from the training data and the process of creating a model according to the generation index is automatically performed, the user does not need to refer to the inside of the training data and determine what kind of distribution data is present. As a result, the information processing device 10 can reduce the effort required for a data scientist or the like to recognize the training data when creating a model, and can prevent the loss of privacy that accompanies the recognition of the training data.
[0040] [3-2. Generated indicators according to data type] An example of a condition for generating a generation index will be described below. First, an example of a condition according to what kind of data is adopted as the learning data will be described.
[0041] For example, the learning data used for learning includes integers, floating points, character strings, etc. As a result, it is estimated that if an appropriate model is selected for the format of input data, the learning accuracy of the model will be higher. Therefore, the information processing device 10 generates a generation index based on whether the learning data is an integer, a floating point, or a character string.
[0042] For example, when the training data is an integer, the information processing device 10 generates a generation index based on the continuity of the training data. For example, when the density of the training data exceeds a predetermined first threshold, the information processing device 10 regards the training data as data having continuity, and generates a generation index based on whether or not the maximum value of the training data exceeds a predetermined second threshold. Also, when the density of the training data is below the predetermined first threshold, the information processing device 10 regards the training data as sparse training data, and generates a generation index based on whether or not the number of unique values included in the training data exceeds a predetermined third threshold.
[0043] A more specific example will be described. In the following example, an example of a process of selecting a feature function as a generation index from a configuration file to be transmitted to a model generation server 2 that automatically generates a model by AutoML will be described. For example, when the learning data is an integer, the information processing device 10 determines whether or not the density exceeds a predetermined first threshold. For example, the information processing device 10 calculates the density by dividing the number of unique values included in the learning data by the maximum value of the learning data plus 1.
[0044] Next, when the density exceeds a predetermined first threshold, the information processing device 10 determines that the learning data is learning data having continuity, and determines whether or not a value obtained by adding 1 to the maximum value of the learning data exceeds a second threshold. Then, when a value obtained by adding 1 to the maximum value of the learning data exceeds the second threshold, the information processing device 10 selects "Categorical_column_with_identity & embedding_column" as the feature function. On the other hand, when a value obtained by adding 1 to the maximum value of the learning data falls below the second threshold, the information processing device 10 selects "Categorical_column_with_identity" as the feature function.
[0045] On the other hand, when the density is below a predetermined first threshold, the information processing device 10 determines that the training data is sparse, and determines whether the number of unique values included in the training data exceeds a predetermined third threshold. When the number of unique values included in the training data exceeds the predetermined third threshold, the information processing device 10 selects "Categorical_column_with_hash_bucket & embedding_column" as the feature function, and when the number of unique values included in the training data is below the predetermined third threshold, the information processing device 10 selects "Categorical_column_with_hash_bucket" as the feature function.
[0046] Furthermore, when the learning data is a character string, the information processing device 10 generates a generation index based on the number of types of character strings included in the learning data. For example, the information processing device 10 counts the number of unique character strings (the number of unique data) included in the learning data, and when the counted number is below a predetermined fourth threshold, selects "categorical_column_with_vocabulary_list" and / or "categorical_column_with_vocabulary_file" as the feature function. When the counted number is below a fifth threshold that is greater than the predetermined fourth threshold, the information processing device 10 selects "categorical_column_with_vocabulary_file & embedding_column" as the feature function. When the counted number is greater than a fifth threshold that is greater than the predetermined fourth threshold, the information processing device 10 selects "categorical_column_with_hash_bucket & embedding_column" as the feature function.
[0047] Furthermore, when the learning data is a floating point, the information processing device 10 generates a conversion index to input data for inputting the learning data to the model as a generation index of the model. For example, the information processing device 10 selects "bucketized_column" or "numeric_column" as a feature function. That is, the information processing device 10 bucketizes (groups) the learning data and selects whether to input the bucket number or to input the numerical value as it is. Note that the information processing device 10 may bucketize the learning data so that the range of numerical values associated with each bucket is approximately the same, and may associate a numerical range with each bucket so that the number of learning data classified into each bucket is approximately the same. Furthermore, the information processing device 10 may select the number of buckets or the range of numerical values associated with the bucket as a generation index.
[0048] Furthermore, the information processing device 10 acquires learning data showing multiple features, and generates, as a generation index for the model, a generation index showing a feature to be learned by the model among the features possessed by the learning data. For example, the information processing device 10 determines which label of the learning data is to be input to the model, and generates a generation index showing the determined label. Furthermore, the information processing device 10 generates, as a generation index for the model, a generation index showing multiple types of the learning data types for which correlation is to be learned by the model. For example, the information processing device 10 determines a combination of labels to be simultaneously input to the model, and generates a generation index showing the determined combination.
[0049] Furthermore, the information processing device 10 generates a generation index indicating the number of dimensions of the learning data input to the model as a generation index of the model. For example, the information processing device 10 may determine the number of nodes in the input layer of the model according to the number of unique data included in the learning data, the number of labels input to the model, a combination of the numbers of labels input to the model, the number of buckets, etc.
[0050] Furthermore, the information processing device 10 generates a generation index indicating the type of model that learns the features of the learning data as a generation index of the model. For example, the information processing device 10 determines the type of model to be generated according to the density and sparseness of the learning data previously used as the learning target, the label contents, the number of labels, the number of label combinations, and the like, and generates a generation index indicating the determined type. For example, the information processing device 10 generates a generation index indicating "BaselineClassifier", "LinearClassifier", "DNNClassifier", "DNNLinearCombinedClassifier", "BoostedTreesClassifier", "AdaNetClassifier", "RNNClassifier", "DNNResNetClassifier", "AutoIntClassifier", and the like as the class of the model in AutoML.
[0051] The information processing device 10 may generate generation indices indicating various independent variables of the models of each class. For example, the information processing device 10 may generate, as the generation indices of the model, a generation indices indicating the number of intermediate layers the model has or the number of nodes included in each layer. Furthermore, the information processing device 10 may generate, as the generation indices of the model, a generation indices indicating the connection state between the nodes in the model or a generation indices indicating the size of the model. These independent variables are appropriately selected depending on whether various statistical features of the learning data satisfy predetermined conditions.
[0052] Furthermore, the information processing device 10 may generate, as a generation index of the model, a generation index indicating a learning mode when the model learns the features of the learning data, that is, a hyperparameter. For example, the information processing device 10 may generate a generation index indicating "stop_if_no_decrease_hook", "stop_if_no_increase_hook", "stop_if_higher_hook", or "stop_if_lower_hook" in the setting of the learning mode in AutoML.
[0053] That is, the information processing device 10 generates a generation index indicating the characteristics of the learning data to be learned by the model, the mode of the model to be generated, and the learning mode when the model is trained on the characteristics of the learning data, based on the labels of the learning data used for learning and the characteristics of the data itself. More specifically, the information processing device 10 generates a configuration file for controlling the generation of a model in AutoML.
[0054] [3-3. Order of determining generation indicators] Here, the information processing device 10 may optimize the various indices described above simultaneously or in an appropriate order. The information processing device 10 may also be able to change the order in which the indices are optimized. That is, the information processing device 10 may receive from the user a designation of the order for determining the features of the learning data to be learned by the model, the mode of the model to be generated, and the learning mode when the features of the learning data are learned by the model, and may determine the indices in the order received.
[0055] For example, when the information processing device 10 starts generating generated indices, it optimizes input features such as the features of the input learning data and the manner in which the learning data is to be input, and then optimizes input cross features to determine which combination of features is to be learned. Next, the information processing device 10 selects a model and optimizes the model structure. After that, the information processing device 10 optimizes hyperparameters and ends the generation of generated indices.
[0056] Here, in the input feature optimization, the information processing device 10 may iteratively optimize the input features by selecting or modifying various input features such as the characteristics and input mode of the learning data to be input, and selecting new input features using a genetic algorithm. Similarly, in the input cross feature optimization, the information processing device 10 may iteratively optimize the input cross features, and may iteratively execute model selection and model structure optimization. Furthermore, the information processing device 10 may iteratively execute hyperparameter optimization. Furthermore, the information processing device 10 may iteratively execute a series of processes, including input feature optimization, input cross feature optimization, model selection, model structure optimization, and hyperparameter optimization, to optimize each index.
[0057] Furthermore, the information processing device 10 may, for example, perform hyperparameter optimization before performing model selection or model structure optimization, or may perform input feature optimization or input cross feature optimization after model selection or model structure optimization. Furthermore, the information processing device 10 may, for example, repeatedly perform input feature optimization and then repeatedly perform input cross feature optimization. Thereafter, the information processing device 10 may repeatedly perform input feature optimization and input cross feature optimization. In this way, any setting can be adopted for which indexes are optimized in what order, and which optimization processes are repeatedly executed in the optimization.
[0058] [3-4. Flow of model generation realized by information processing device] Next, an example of a flow of model generation using the information processing device 10 will be described with reference to Fig. 2. Fig. 2 is a diagram for explaining an example of a flow of model generation using the information processing device in the embodiment. For example, the information processing device 10 accepts learning data and a label of each learning data. Note that the information processing device 10 may accept a label together with the designation of the learning data.
[0059] In such a case, the information processing device 10 analyzes the data and divides the data according to the analysis result. For example, the information processing device 10 divides the learning data into training data used for learning the model and evaluation data used for evaluating the model (i.e., measuring the accuracy). The information processing device 10 may further divide data for various tests. The process of dividing such learning data into training data and evaluation data can employ any of various known techniques.
[0060] Furthermore, the information processing device 10 generates the various generation indices described above using the learning data. For example, the information processing device 10 generates a configuration file that defines a model generated in AutoML and learning of the model. In such a configuration file, various functions used in AutoML are stored as they are as information indicating the generation indices. Then, the information processing device 10 generates a model by providing the training data and the generation indices to the model generation server 2.
[0061] Here, the information processing device 10 may optimize the generation index and thus the model by repeatedly performing model evaluation by the user and automatic generation of the model. For example, the information processing device 10 optimizes input features (optimization of input features and input cross features), optimizes hyperparameters, and optimizes the model to be generated, and automatically generates a model according to the optimized generation index. Then, the information processing device 10 provides the generated model to the user.
[0062] Meanwhile, the user trains, evaluates, and tests the automatically generated model, and analyzes and provides the model. The user then corrects the generated generation index to automatically generate a new model again, and evaluates and tests it. By repeatedly executing such processing, it is possible to realize processing for improving the accuracy of the model through trial and error, without executing complex processing.
[0063] 4. Configuration of Information Processing Device Next, an example of a functional configuration of the information processing device 10 according to the embodiment will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of a configuration of the information processing device according to the embodiment. As shown in Fig. 3, the information processing device 10 has a communication unit 20, a storage unit 30, and a control unit 40.
[0064] The communication unit 20 is realized by, for example, a network interface card (NIC) etc. The communication unit 20 is connected to the network N by wire or wirelessly, and transmits and receives information to and from the model generation server 2 and the terminal device 3.
[0065] The storage unit 30 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 30 also has a training data database 31 and a model generation database 32.
[0066] The training data database 31 stores various information related to data used for training. The training data database 31 stores a data set of training data used for model training. FIG. 4 is a diagram showing an example of information registered in the training data database according to the embodiment. In the example of FIG. 4, the training data database 31 includes items such as "data set ID", "data ID", and "data".
[0067] "Dataset ID" indicates identification information for identifying a dataset. "Data ID" indicates identification information for identifying each piece of data. Also, "Data" indicates data identified by the data ID. For example, in the example of FIG. 4, corresponding data (learning data) is registered in association with the data ID that identifies each piece of learning data.
[0068] In the example of FIG. 4, a data set (data set DS1) identified by a data set ID "DS1" includes a plurality of data "DT1", "DT2", "DT3", etc., identified by data IDs "DID1", "DID2", "DID3", etc. In FIG. 4, the data is shown as abstract character strings such as "DT1", "DT2", "DT3", etc., but the data may be registered in any format such as various integers, real numbers, character strings, sentences, etc., or information in any format such as various integers, real numbers, character strings, sentences, etc., converted into text may be registered. For example, the learning data database 31 may store data as shown in FIG. 10.
[0069] Although not shown in the figures, the learning data database 31 may store a label (correct answer information) corresponding to each piece of data in association with each piece of data. Also, for example, a data group including a plurality of pieces of data may be stored in association with one label. In this case, the data group including a plurality of pieces of data corresponds to data (input data) input to the model. For example, information in any format, such as a numerical value or a character string, may be used as the label.
[0070] The learning data database 31 may store various information according to the purpose, without being limited to the above. For example, the learning data database 31 stores information indicating which of a plurality of learning stages each piece of data is used in, in association with each piece of data. For example, the learning data database 31 stores information indicating whether each piece of data is a first data group or a second data group in association with each piece of data. For example, the learning data database 31 may store data in a manner that allows each piece of data to be identified as data used in a learning process (training data) or data used for evaluation (evaluation data), etc. For example, the learning data database 31 may store information (such as a flag) that identifies whether each piece of data is training data or evaluation data in association with each piece of data.
[0071] The model generation database 32 stores various information used to generate a model other than the learning data. The model generation database 32 stores various information related to the model to be generated. For example, the model generation database 32 stores information used to generate a model based on a genetic algorithm. For example, the model generation database 32 stores information specifying the number of combinations of types to be inherited in subsequent processing based on the genetic algorithm.
[0072] For example, the model generation database 32 stores setting values such as various parameters related to the model to be generated. The model generation database 32 stores an upper limit value of the size of the model (also called "upper size limit value"). The model generation database 32 stores information indicating the structure of the model, such as the number of partial models (blocks) included in the model to be generated and information on each partial model. The model generation database 32 stores information on modules used as components of the partial models. Note that the partial models (blocks) may, for example, constitute a part of the model, or may function as a model by themselves. Also, a module is a functional unit element for realizing a function realized by, for example, a partial model (block).
[0073] The model generation database 32 stores information indicating what kind of processing each module performs, information on the elements that constitute each module, etc. The model generation database 32 stores various information on the processing that constitutes each module. The model generation database 32 stores information on the processing that constitutes each module, such as normalization, dropout, etc.
[0074] For example, the model generation database 32 stores information about each partial model. The model generation database 32 stores information indicating what modules each partial model is made up of. For example, the model generation database 32 stores information indicating the number of modules each partial model has. The model generation database 32 stores information indicating the modules included in each partial model.
[0075] Information indicating the type of data that each partial model uses as an input is stored in the model generation database 32. For example, information indicating a combination of types of data that each partial model uses as an input is stored in the model generation database 32.
[0076] The model generation database 32 is not limited to the above, and may store various types of information as long as the information is used for generating a model.
[0077] Returning to FIG. 3, the description will be continued. The control unit 40 is realized by, for example, a central processing unit (CPU) or a micro processing unit (MPU) executing various programs (for example, a generation program for executing a process for generating a model, an information processing program, etc.) stored in a storage device inside the information processing device 10 using a RAM as a working area. The information processing program is used to operate a computer as a model having at least one partial model (block). For example, the information processing program operates a computer (for example, the information processing device 10) as a model that has been learned using learning data. In addition, the control unit 40 is realized by, for example, an integrated circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). As shown in FIG. 3, the control unit 40 has an acquisition unit 41, a determination unit 42, a reception unit 43, a generation unit 44, a processing unit 45, and a provision unit 46.
[0078] The acquisition unit 41 acquires information from the storage unit 30. The acquisition unit 41 acquires a data set of learning data used for model training. The acquisition unit 41 acquires the learning data used for model training. For example, when the acquisition unit 41 receives various data to be used as learning data and labels to be assigned to the various data from the terminal device 3, the acquisition unit 41 registers the received data and labels as learning data in the learning data database 31. The acquisition unit 41 may also receive designation of the learning data ID and label of the learning data to be used for model training from among the data registered in advance in the learning data database 31.
[0079] The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of pieces of input information including converted information obtained by converting information other than sentence text, which is a sentence. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including text obtained by converting information other than sentence text. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including text obtained by converting information other than sentence text, which is a sentence (also referred to as "sentence"). The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including text obtained by converting information other than sentence text, which is a sentence.
[0080] The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including text corresponding to a sentence. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including sentence text. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including text corresponding to a sentence posted on the Internet. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including sentence text posted on the Internet. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including text corresponding to a question posted on the Internet. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including text corresponding to an answer posted on the Internet.
[0081] The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including texts into which numerical values have been converted. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including texts into which integers have been converted. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including texts into which real numbers have been converted. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including texts into which numerical values indicating the date and time at which a sentence was posted have been converted.
[0082] The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including texts into which character strings other than sentences have been converted. The acquisition unit 41 acquires learning data used for training a model that receives as input a plurality of texts including texts into which character strings indicating the day of the week on which a sentence was posted have been converted.
[0083] The acquisition unit 41 acquires learning data used for training a base model that uses text as input. The base model here may be a model (partial model) that has been trained to be able to execute a wide variety of tasks, for example, so that it can be used as a part of models for various purposes. For example, the base model may be a model that can be applied as a base (partial model) for a wide variety of applications. The acquisition unit 41 acquires learning data including a first data group used for training general-purpose language ability. The acquisition unit 41 acquires a first data group including natural language data used for training a large-scale language model.
[0084] The acquisition unit 41 acquires learning data including a second data group used for learning language ability for a specific process. The acquisition unit 41 acquires learning data including a second data group that is data related to posts on the Internet. The acquisition unit 41 acquires the second data group including text corresponding to a sentence posted on the Internet. The acquisition unit 41 acquires the second data group including text corresponding to a question posted on the Internet. The acquisition unit 41 acquires the second data group including text corresponding to an answer posted on the Internet.
[0085] The determination unit 42 determines various information related to the learning process. The determination unit 42 determines the learning mode. The determination unit 42 determines initial values, etc., in the learning process by the generation unit 44. The determination unit 42 determines the initial value of each parameter. The determination unit 42 refers to a setting file indicating the initial setting value of each parameter, and determines the initial value of each parameter. For example, the determination unit 42 determines the maximum number of partial models (blocks) to be included in the model. The determination unit 42 determines the maximum number of modules to be included in a partial model (block). The determination unit 42 determines the dropout rate. The determination unit 42 determines the dropout rate of each partial model (block). The determination unit 42 determines the size of the model. The determination unit 42 determines the number of modules to be included in each partial model (block).
[0086] The receiving unit 43 receives a correction to the generation index presented to the user. The receiving unit 43 also receives from the user a designation of the order for determining the features of the learning data to be learned by the model, the mode of the model to be generated, and the learning mode when the features of the learning data are learned by the model.
[0087] The generating unit 44 generates various information in response to the determination by the determining unit 42. In addition, the generating unit 44 generates various information in response to an instruction received by the receiving unit 43. For example, the generating unit 44 may generate a generation index of the model.
[0088] The generation unit 44 generates a model that can input information other than sentences as text by learning using the learning data. The generation unit 44 generates a model that receives as input a plurality of texts based on data in a tabular format. The generation unit 44 generates a model that is a language model that can input information other than sentences as text by learning using the learning data.
[0089] The generation unit 44 generates a model having as input a plurality of texts including text corresponding to a sentence. The generation unit 44 generates a model having as input a plurality of texts including sentence text. The generation unit 44 generates a model having as input a plurality of texts including text corresponding to a question. The generation unit 44 generates a model having as input a plurality of texts including text corresponding to an answer.
[0090] The generation unit 44 generates a model in which a numerical value can be input as text. The generation unit 44 generates a model in which an integer can be input as text. The generation unit 44 generates a model in which a real number can be input as text. The generation unit 44 generates a model in which the date and time when a sentence is posted can be input as text. The generation unit 44 generates a model in which a character string other than a sentence can be input as text. The generation unit 44 generates a model in which the day of the week when a sentence is posted can be input as text.
[0091] The generation unit 44 generates a base model by a multi-stage learning process using the learning data. The generation unit 44 generates a base model by a multi-stage learning process including a first-stage learning process using a first data group. The generation unit 44 generates a base model by a multi-stage learning process including a second-stage learning process using a second data group.
[0092] The generation unit 44 generates the base model through a multiple-stage learning process including a third-stage learning process in which layers of the base model are increased and learning is performed. The generation unit 44 generates the base model through the third-stage learning process using data used to learn language ability for a specific process. The generation unit 44 generates the base model through the third-stage learning process using data related to posts on the Internet.
[0093] The generation unit 44 generates a base model by a third-stage learning process using text corresponding to a sentence posted on the Internet. The generation unit 44 repeatedly executes the third-stage learning process to increase layers of the base model and to perform a learning process on the increased base model, thereby generating a base model.
[0094] The generation unit 44 generates a model capable of converting and inputting information other than sentence text by learning using the learning data. The generation unit 44 generates a base model and a fine-tuning model that fine-tunes the base model to apply it to a specified task by learning using information other than sentences. The generation unit 44 generates a base model and a fine-tuning model in which input information is input differently from the base model. The generation unit 44 generates a base model and a fine-tuning model in which the input order of input information is the same as that of the base model. The generation unit 44 uses label information as input information in learning the base model.
[0095] The generating unit 44 may generate a model based on a genetic algorithm. For example, the generating unit 44 generates a plurality of models by targeting a plurality of combination candidates each having a different type combination. The generating unit 44 may further generate a model using combination candidates (also called "inheritance candidates") corresponding to a predetermined number (e.g., two) of models with high accuracy among the generated plurality of models. For example, the generating unit 44 may inherit some type combinations from each of the inheritance candidates, and generate a model using a type candidate to which the type combination of the inheritance candidate is copied. The generating unit 44 may generate a model to be finally used by repeating the process of generating a model by inheriting the type combination of the inheritance candidate described above.
[0096] The generation unit 44 requests the model generation server 2 to learn the model by sending data to be used for generating the model to the external model generation server 2, and generates the model by receiving from the model generation server 2 the model that the model generation server 2 has learned.
[0097] For example, the generation unit 44 generates a model using data registered in the learning data database 31. The generation unit 44 generates a model based on each piece of data used as training data and a label. The generation unit 44 generates a model by performing learning so that an output result output by the model when training data is input matches the label. For example, the generation unit 44 generates a model by transmitting each piece of data used as training data and a label to the model generation server 2 and having the model generation server 2 learn the model.
[0098] For example, the generation unit 44 measures the accuracy of the model using data registered in the learning data database 31. The generation unit 44 measures the accuracy of the model based on each piece of data used as evaluation data and the label. The generation unit 44 measures the accuracy of the model by collecting the result of comparing the output result output by the model when the evaluation data is input with the label.
[0099] The processing unit 45 performs various processes. The processing unit 45 functions as an inference unit that performs inference processing. The processing unit 45 performs inference processing using a model (e.g., model M1) stored in the storage unit 30. The processing unit 45 performs inference processing using a model acquired by the acquisition unit 41. The processing unit 45 performs inference processing using a model generated by the generation unit 44. The processing unit 45 performs inference processing using a model learned using the model generation server 2. The processing unit 45 performs inference processing to input data into a model and generate an inference result corresponding to the data.
[0100] The processing unit 45 executes an inference process using the model generated by the generation unit 44. The processing unit 45 inputs a plurality of texts including text in which information other than sentences has been converted to the model, and executes an inference process based on output data output by the model. The processing unit 45 inputs a plurality of texts including text in which information other than sentence text, which is a sentence, has been converted to the model, and executes an inference process based on output data output by the model. The processing unit 45 inputs a plurality of texts in which at least one of an integer, a real number, and a character string has been converted to text to the model, and executes an inference process based on output data output by the model.
[0101] The processing unit 45 may execute the inference process by using an external device (an inference server) having a model. For example, the processing unit 45 may transmit input data to the inference server having the model, receive information (inference information) generated by the external device using the input data and the model received, and perform the inference process by using the received inference information.
[0102] The providing unit 46 provides the generated model to the user. The providing unit 46 transmits to the user's terminal device 3 an information processing program that causes the user's terminal device 3 to operate as a model (e.g., model M1) used in inference processing. For example, when the accuracy of the model generated by the generating unit 44 exceeds a predetermined threshold, the providing unit 46 transmits the model together with a generation index corresponding to the model to the terminal device 3. As a result, the user can evaluate and try out the model, and also modify the generation index.
[0103] The providing unit 46 presents the index generated by the generating unit 44 to the user. For example, the providing unit 46 transmits the AutoML configuration file generated as the generated index to the terminal device 3. Furthermore, the providing unit 46 may present the generated index to the user every time the generated index is generated, or may present to the user only the generated index corresponding to a model whose accuracy exceeds a predetermined threshold, for example.
[0104] [5. Processing flow of information processing system] Next, the procedure of the process executed by the information processing device 10 will be described with reference to Fig. 5 and Fig. 6. Fig. 5 and Fig. 6 are flowcharts showing an example of the flow of information processing according to the embodiment. In addition, the following will be described as an example in which the information processing system 1 performs the process, but the process shown below may be performed by any device included in the information processing system 1, such as the information processing device 10, the model generation server 2, the terminal device 3, etc., included in the information processing system 1.
[0105] First, a processing example shown in Fig. 5 will be described. In Fig. 5, the information processing system 1 acquires learning data used for learning a model that receives as input a plurality of pieces of input information including converted information obtained by converting information other than sentence text (step S101). Then, the information processing system 1 generates a model that can be input by converting information other than sentence text through learning using the learning data (step S102).
[0106] Next, a processing example shown in Fig. 6 will be described. In Fig. 6, the information processing system 1 acquires learning data used for learning a base model with text as input (step S201). Then, the information processing system 1 generates a base model through a multi-stage learning process using the learning data (step S202).
[0107] [6. Examples of processing in information processing systems] Here, an example in which the information processing system 1 performs the processes in Figs. 5 and 6 described above will be described. The information processing device 10 acquires learning data. The information processing device 10 acquires information such as parameters used to generate a model. For example, the information processing device 10 acquires information indicating various upper limit values for the model to be generated. For example, the information processing device 10 acquires information indicating an upper limit value for the size of the model to be generated. In addition, the information processing device 10 acquires various setting values in the genetic algorithm. For example, the information processing device 10 acquires information indicating the number of inheritance candidates in the genetic algorithm.
[0108] The information processing device 10 generates a model based on learning data, information indicating the structure of the model, various upper limits such as an upper size limit, and information indicating settings in a genetic algorithm. The information processing device 10 generates the model capable of inputting information other than sentence text by converting it. The information processing device 10 learns a model that receives as input a plurality of pieces of input information including converted information obtained by converting information other than sentence text, which is sentence text. For example, the information processing device 10 generates a model capable of inputting information other than sentence text as text. The information processing device 10 generates a base model through a multi-stage learning process using the learning data.
[0109] The information processing device 10 transmits information used for generating a model to a model generation server 2 that learns the model. For example, the information processing device 10 transmits to the model generation server 2 information indicating learning data, information indicating the structure of the model, various upper limits such as an upper size limit value, setting values in a genetic algorithm, and the like.
[0110] The model generation server 2 that receives information from the information processing device 10 generates a model through a learning process. Then, the model generation server 2 transmits the generated model to the information processing device 10. In this way, the concept of "generating a model" in this application includes not only learning a model in one's own device, but also generating and instructing a model to another device by providing information necessary for generating the model to the other device, and receiving the model learned by the other device. In the information processing system 1, the information processing device 10 generates a model by transmitting information used for generating the model to the model generation server 2 that learns the model, and acquiring the model generated by the model generation server 2. In this way, the information processing device 10 generates a model by requesting the generation of a model by transmitting information used for generating the model to the other device, and having the other device that receives the request generate the model.
[0111] [7. Model] From here, the model will be described. Below, various points related to the model, such as the structure and learning mode of the model generated in the information processing system 1, will be described. In the example shown below, text into which information other than sentence text, which is a sentence, is converted will be described as an example of converted information. Note that the format of the converted information is not limited to text, and the input to the model is not limited to text, but this point will be described later.
[0112] [7-1. Example of model structure] First, an example of the structure of a model to be generated will be described with reference to FIG. 7. The information processing system 1 generates a model M1 as shown in FIG. 7. FIG. 7 is a diagram showing an example of the structure of a model according to an embodiment. In FIG. 7, the information processing system 1 generates a model M1 having various configurations such as a partial model PM1, which is an example of a base model, and a plurality of partial models such as a partial model PM2 in which the output of the partial model PM1 is used as an input. When describing the partial models PM1, PM2, etc. without making a particular distinction, they may be described as "partial models PM" or simply "partial models". Note that FIG. 7 shows an example in which the model M1 has two partial models PM, but the model M1 may have three or more partial models PM, or may have only one partial model PM.
[0113] For example, the partial model PM1 is a model (language model) based on Transformer. The Transformer (model) is the same as the conventional Transformer, and detailed description will be omitted. The partial model PM1 may be any model as long as it can input information such as integers and real numbers in a format (text) similar to a sentence (also simply called a "sentence"). For example, the partial model PM1 may be a model configured based on any natural language processing model such as BERT (Bidirectional Encoder Representations from Transformers), RoBERTa (A Robustly Optimized BERT Pretraining Approach), or DeBERTa (Decoding-enhanced BERT with disentangled attention). In addition, the partial model PM1 may be a model of any configuration, not limited to a model configured based on BERT, as long as it can input information such as integers and real numbers in a format (text) similar to a sentence.
[0114] In the example shown in FIG. 7, the partial model PM1 has a plurality of layers (module layers) such as layers EL10, EL11, EL12, EL15, etc. In the following, when layers EL10, EL11, EL12, EL15, etc. are not particularly distinguished from one another, they may be referred to as "layer EL" or simply "layer". FIG. 7 shows an example of a configuration in which each layer EL includes one Transformer. Note that the number of layers shown in FIG. 7 is merely an example, and the partial model PM1 may include any number of layers EL. Furthermore, the number of layers EL in the partial model PM1 may be changed (increased) depending on learning, but this point will be described later.
[0115] In FIG. 7, layer EL10 is the layer located closest to the input side of partial model PM1. For example, layer EL10 may be a layer (input layer) to which input data of partial model PM1 is input. FIG. 7 shows, as an example, a state in which a token "CLS" indicating the beginning of a text (sentence), the text "This is a pen", a token "SEP" indicating a break in a sentence, and the like are input to (layer EL10 of) partial model PM1. For example, after the token "SEP", information in which a numerical value (integer) indicating time is made into text (for example, the character "7"), information in which a character string indicating a day of the week is made into text (for example, the character "Sat"), and the like are input in succession.
[0116] In the partial model PM1, a layer EL11 is placed after the layer EL10. That is, the layer EL11 is a layer EL to which the output of the layer EL10 is input. In the partial model PM1, a layer EL12 is placed after the layer EL11. That is, the layer EL12 is a layer EL to which the output of the layer EL11 is input.
[0117] In Fig. 7, layer EL15 is the layer located closest to the output side of partial model PM1. For example, the output of layer EL15 is used as the output of partial model PM1. Note that Fig. 7 is merely an example, and the partial model PM1 can have any configuration.
[0118] The partial model PM2, denoted as "DNN Sparse" in FIG. 7, is a partial model PM in which the output of the partial model PM1 is used as an input. For example, the partial model PM2 is a sparse DNN (deep neural network) configured using any technique such as dropout. Note that the partial model PM2 may be any model that uses the output from the partial model PM1 as an input and outputs a desired inference result. For example, the partial model PM2 is not limited to a sparse DNN and may be any DNN, or may be any model in addition to a DNN.
[0119] [7-2. Model input example] Moreover, Fig. 7 shows an example of input to model M1. Fig. 7 shows an example in which two sentences, Sentence#1 and Sentence#2, are input to model M1. For example, in Fig. 7, a text in which information indicating the question category has been converted into text, a token "SEP", and text corresponding to the question (question sentence) are arranged in this order is input to model M1 as Sentence#1. That is, in Fig. 7, a single text in which information indicating the question category has been converted into text and text corresponding to the question are linked by the token "SEP" is used for Sentence#1.
[0120] Also, in FIG. 7, text corresponding to the answer (answer sentence) is input as Sentence#2 to model M1. That is, in FIG. 7, text corresponding to the answer is used for Sentence#2. An example of input designation shown in FIG. 7 is shown in FIG. 8. FIG. 8 is a diagram showing an example of input designation for a model according to an embodiment. As shown in FIG. 8, input to model M1 is performed by designating column (item) names corresponding to various types of information.
[0121] For example, when specifying a category and question as Sentence#1, which is an input for model M1, a string in which strings indicating the category are separated by a delimiter (a comma in the case of Figure 8) and enclosed in square brackets is used. Specifically, a string such as "[category,question]" is used to specify Sentence#1.
[0122] In addition, when specifying an answer as Sentence#2, which is an input to model M1, a character string indicating the type of answer is enclosed in square brackets. Specifically, a character string such as "[answer]" is used to specify Sentence#2.
[0123] In the example of Fig. 8, the entire input of model M1 is a character string in which information about each type of Sentence#1 and Sentence#2 is separated by a delimiter (a comma in Fig. 8). Specifically, the input of model M1 is specified as a character string such as "tokenizerColumns: [[category,question], [answer]]".
[0124] Note that FIG. 8 is merely an example, and as long as the input of the model M1 can be specified, it may be specified in any manner. Moreover, the input shown in FIG. 7 is merely an example, and any combination of text may be used as the input of the model M1. For example, when a question and an answer are used as the input of the model M1, they may be specified as shown in FIG. 9. FIG. 9 is a diagram showing an example of the specification of the input of the model according to the embodiment. As shown in FIG. 9, when a question is specified as Sentence#1, which is an input of the model M1, and an answer is specified as Sentence#2, a character string such as "tokenizerColumns: [[question], [answer]]" is used.
[0125] As described above, information other than sentences is also converted into text and input to the model M1. An example of this will be described with reference to FIG. 10. FIG. 10 is a diagram showing an example of the type of input according to the embodiment. Each row in FIG. 10 indicates the type of each piece of information contained in the data. Note that the information corresponding to "label" in FIG. 10 does not have to be the input of the model. For example, the information corresponding to "label" in FIG. 10 may be a label (correct answer information) indicating whether or not the questions, answers, etc. in each column correspond to a violation.
[0126] For example, the information corresponding to "hour" in Fig. 10 is an integer (numeric value) indicating the date and time when the question or answer in the corresponding column was posted. When the information corresponding to "hour" in Fig. 10 is used as input to model M1, the integer (numeric value) is converted to text and then input.
[0127] Also, for example, the information corresponding to "day_week" in FIG. 10 is a character string indicating the day of the week on which the question or answer in the corresponding column was posted. When the information corresponding to "day_week" in FIG. 10 is used as an input to the model M1, the character string is converted to text and input. This enables the model M1 to input integers, real numbers, and character strings as text. Note that when the character string indicating the day of the week can be used as text as is, the information corresponding to "day_week" in FIG. 10 may be used as input to the model M1 as is.
[0128] Also, for example, the information corresponding to "question" in Fig. 10 is a sentence indicating a question in the corresponding column. When the information corresponding to "question" in Fig. 10 is used as an input to model M1, the sentence (text) is used as it is as an input to model M1.
[0129] Also, for example, the information corresponding to "answer" in FIG. 10 is a sentence indicating the answer in the corresponding column. When the information corresponding to "answer" in FIG. 10 is used as an input for model M1, the sentence (text) is used as is as an input for model M1. Note that the type of information shown in FIG. 10 is merely an example, and the type of information used as an input for model M1 is not limited to that shown in FIG. 10. For example, the type of information used as an input for model M1 may include categories, etc., as described above.
[0130] For example, the model M1 receives as input tabular data including multiple types of data as shown in Fig. 10. For example, the model M1 receives as input a combination of information in which each of the tabular data including multiple types of data as shown in Fig. 10 has been converted into text.
[0131] Note that a combination of the various types of information described above may be used for inputting the model M1. An example of this point will be described. For example, the input to the model M1 may be an input as shown in FIG. 11. FIG. 11 is a diagram showing an example of an input of the model according to the embodiment. Note that the description of the same points as those described in FIG. 7 etc. will be omitted as appropriate.
[0132] For example, in Fig. 11, model M1 is input with the following text, in this order: text information indicating the date and time the question was posted, the token "SEP", text information indicating the category of the question, the token "SEP", and text corresponding to the question (question sentence), as Sentence#1. That is, in Fig. 11, Sentence#1 uses a single text in which the text information indicating the date and time the question was posted, the text information indicating the category of the question, and the text corresponding to the question are linked by the token "SEP".
[0133] 11, model M1 is input with information indicating the day of the week when the answer was posted converted into text, the token "SEP", and text corresponding to the answer (answer sentence) as Sentence#2. That is, in Fig. 11, Sentence#2 uses a single text in which the information indicating the day of the week when the answer was posted converted into text and the text corresponding to the answer are linked by the token "SEP".
[0134] Moreover, the input information CM1 shows an example of the input designation shown in FIG.
[0135] For example, when specifying the date, time, category, and question as Sentence#1, which is the input of model M1, a string in which the strings indicating the types are separated by a delimiter (a comma in the case of Figure 11) and enclosed in square brackets is used. Specifically, a string such as "[hour,category,question]" is used to specify Sentence#1.
[0136] In addition, when specifying the day of the week and the answer as Sentence#2, which is an input for model M1, a string in which the string indicating the type is separated by a delimiter (a comma in the case of Figure 11) and enclosed in square brackets is used. Specifically, a string such as "[day_week,answer]" is used to specify Sentence#2.
[0137] In the example of Fig. 11, the entire input of model M1 is a character string in which information about each type of Sentence #1 and Sentence #2 is separated by a delimiter (a comma in the case of Fig. 11). Specifically, the input of model M1 is specified as a character string such as "tokenizerColumns: [[hour,category,question], [day_week,answer]]".
[0138] [7-3. Examples of other model structures] The above-mentioned structure of the model is merely an example, and the model can adopt any configuration. An example of this point will be described with reference to FIG. 12. FIG. 12 is a diagram showing another example of the structure of the model according to the embodiment. Note that the same points as those in FIG. 7 will be appropriately omitted from the description by assigning the same reference numerals, etc.
[0139] 12, the information processing system 1 generates a model M11 having various configurations such as partial models PM1 and PM11, which are examples of base models, and a plurality of partial models such as a partial model PM12 in which the outputs of the partial models PM1 and PM11 are used as inputs. That is, the model M11 differs from the model M1 in that it includes a partial model P11 and a partial model PM12 in which the output of the partial model P11 and the output of the partial model PM1 are used as inputs.
[0140] For example, the partial model PM11 is a model (language model) based on Transformer, similar to the partial model PM1. Note that the internal configuration of the partial model PM11 is similar to that of the partial model PM1, and therefore a description thereof will be omitted.
[0141] The partial model PM12, which is indicated as "DNN Sparse" in Fig. 12, is a partial model PM in which the output of the partial model PM1 and the output of the partial model PM11 are used as inputs. Note that the partial model PM12 is similar to the partial model PM2 except that the output of the partial model PM11 is used as inputs, and therefore a description thereof will be omitted.
[0142] Fig. 12 shows an example of input to the model M11. Fig. 12 shows a case where Sentence#1 and Sentence#2 are input to the partial model PM1, and Sentence#3 and Sentence#4 are input to the partial model PM11. Note that Sentence#1 and Sentence#2 input to the partial model PM1 are the same as in Fig. 7, and therefore description thereof will be omitted.
[0143] FIG. 12 shows an example in which two sentences, Sentence#3 and Sentence#4, are input to model M11 (partial model PM11). For example, in FIG. 12, text corresponding to an answer (answer sentence) is input to partial model PM11 as Sentence#3. Text in which text corresponding to a question (question sentence), the token "SEP", and information indicating the question category in text form are arranged in that order is input as Sentence#4. An example of the input designation shown in FIG. 12 is shown in FIG. 13. FIG. 13 is a diagram showing an example of an input to a model according to an embodiment.
[0144] For example, when specifying an answer as Sentence#3, which is an input for model M11, a character string indicating the type of answer enclosed in square brackets is used. Specifically, a character string such as "[answer]" is used to specify Sentence#3.
[0145] Furthermore, when specifying a question and category as Sentence#4, which is an input to model M11, a string in which character strings indicating the type are arranged, separated by a delimiter (a comma in the case of Fig. 13), and enclosed in square brackets is used. Specifically, a string such as "[question,category]" is used to specify Sentence#4.
[0146] In the example of Fig. 13, the entire input of the model M11 is a character string in which information for each type of Sentence#1, Sentence#2, Sentence#3, and Sentence#4 is separated by a delimiter (a comma in the case of Fig. 13). Specifically, the input of the model M11 is specified as a character string such as "tokenizerColumns: [[category,question], [answer], [answer], [question,category]]". In this way, the model M11 is a model that allows duplicate text input.
[0147] 7-4. Experimental Results Next, an example of the results of an experiment performed using the model generated by the above-mentioned processing will be described with reference to Figures 14 to 18. Figures 14 to 18 are diagrams showing an example of the results of the experiment.
[0148] First, an example shown in Fig. 14 will be described. An example of an experiment result when determining whether an answer posting violates rules in a knowledge sharing service where questions, answers, and the like are posted will be shown.
[0149] Result RS11 shown on the left side of FIG. 14 indicates the experimental results when the conventional model was used. The vertical axis of result RS11 indicates precision, and the horizontal axis of result RS11 indicates recall. The waveform in the graph of result RS11 indicates the PR curve. For example, the conventional model is a model that uses the above-mentioned BERT, and is a model that accepts answers that are sentences as input.
[0150] Meanwhile, the result RS12 shown on the right side of FIG. 14 indicates the experimental result when the model of this method is used. The vertical axis of the result RS12 indicates precision, and the horizontal axis of the result RS12 indicates recall. The waveform in the graph of the result RS12 indicates the PR curve. For example, the model of this method is a model including a base model (partial model PM1, etc.) such as the above-mentioned model M1, and a DNNSparse model (partial model PM2, etc.). In addition, the model of this method is a model that accepts as input information in which categories are converted into text, a combination of questions that are text, and answers that are text.
[0151] As shown in Results RS11 and RS12 in Figure 14, the model of this method achieved an 80% improvement in accuracy over the conventional model. In this way, it was shown that by using a model that can input types of information that could not be accepted as text in the past, it is possible to improve the accuracy of determining whether an answer is a violation.
[0152] Next, an example shown in Fig. 15 will be described. An example of an experiment result when determining whether an answer posting is in violation in a knowledge sharing service by posting questions, answers, etc. is shown. Note that the same points as in Fig. 14 will not be described as appropriate.
[0153] Result RS21 shown on the left side of Fig. 15 shows the experimental result when the conventional model was used, similar to result RS11 shown on the left side of Fig. 14. The waveform in the graph of result RS21 shows the PR curve. Thus, result RS21 shown on the left side of Fig. 15 is similar to result RS11 shown on the left side of Fig. 14, and therefore a description thereof will be omitted.
[0154] Meanwhile, result RS22 shown on the right side of FIG. 15 indicates the experimental results when the model of this method is used. The vertical axis of result RS22 indicates precision (precision rate), and the horizontal axis of result RS22 indicates recall (recall rate). The waveform in the graph of result RS22 indicates the PR curve. For example, the model of this method is a model obtained by learning the above-mentioned base model (partial model PM1, etc.) using data from a knowledge sharing service to learn a base model suitable for a knowledge sharing service. In addition, the model of this method is a model that accepts as input information in which categories are converted into text, a combination of questions that are text, and answers that are text.
[0155] As shown in Results RS21 and RS22 in Figure 15, the model of this method achieved a 106.6% improvement in accuracy over the conventional model. In this way, it was shown that it is possible to further improve the accuracy of determining whether an answer is a violation by learning the base model of a model that can input types of information that could not be accepted as text before using data from the corresponding service.
[0156] Next, an example shown in Fig. 16 will be described. An example of an experiment result when determining whether a question is posted in violation in a knowledge sharing service that posts questions and answers, etc. will be shown. Note that the same points as those in Fig. 14, Fig. 15, etc. will not be described as appropriate.
[0157] Result RS31 shown on the left side of FIG. 16 indicates the experimental results when the conventional model was used. The vertical axis of result RS31 indicates precision, and the horizontal axis of result RS31 indicates recall. The waveform in the graph of result RS31 indicates the PR curve. For example, the conventional model is a model that uses the above-mentioned BERT, and is a model that accepts a question, which is a sentence, as input.
[0158] Meanwhile, the result RS32 shown on the right side of FIG. 16 indicates the experimental result when the model of this method is used. The vertical axis of the result RS32 indicates precision (precision rate), and the horizontal axis of the result RS32 indicates recall (recall rate). The waveform in the graph of the result RS32 indicates the PR curve. For example, the model of this method is a model that includes a base model (partial model PM1, etc.) such as the above-mentioned model M1, and a DNNSparse model (partial model PM2, etc.). In addition, the model of this method is a model that accepts as input a combination of information in which categories are converted into text and questions, which are sentences.
[0159] As shown in Results RS31 and RS32 in Figure 16, the model of this method achieved a 31% improvement in accuracy over the conventional model. In this way, it was shown that by using a model that can input types of information that could not be accepted as text before, it is possible to improve the accuracy of determining whether a question is a violation.
[0160] Next, an example shown in Fig. 17 will be described. An example of an experiment result when determining whether a question is posted in violation in a knowledge sharing service by posting questions, answers, etc. will be shown. Note that the description of the same points as Figs. 14 to 16 will be omitted as appropriate.
[0161] Result RS41 shown on the left side of Fig. 17 shows the experimental result when the conventional model was used, similar to result RS31 shown on the left side of Fig. 16. The waveform in the graph of result RS41 shows the PR curve. Thus, result RS41 shown on the left side of Fig. 17 is the same as result RS31 shown on the left side of Fig. 16, and therefore a description thereof will be omitted.
[0162] Meanwhile, result RS42 shown on the right side of FIG. 17 indicates the experimental results when the model of this method is used. The vertical axis of result RS42 indicates precision (precision rate), and the horizontal axis of result RS42 indicates recall (recall rate). The waveform in the graph of result RS42 indicates the PR curve. For example, the model of this method is a model obtained by learning the above-mentioned base model (partial model PM1, etc.) using data from a knowledge sharing service to learn a base model suitable for a knowledge sharing service. In addition, the model of this method is a model that accepts as input a combination of information in which categories are converted into text and questions, which are sentences.
[0163] As shown in Results RS41 and RS42 in Figure 17, the model of this method achieved a 45.6% improvement in accuracy over the conventional model. In this way, it was shown that it is possible to further improve the accuracy of determining whether a question is a violation by training the base model of a model that can input types of information that could not be accepted as text before using data from a corresponding service.
[0164] Next, an example shown in Fig. 18 will be described. An example of an experiment result when determining whether a comment posted in violation in a comment service for a news article is shown. Note that the description of the same points as Figs. 14 to 17 etc. will be omitted as appropriate.
[0165] Result RS51 shown on the left side of Figure 18 shows the experimental results when the conventional model was used. The vertical axis of result RS51 indicates precision, and the horizontal axis of result RS51 indicates recall. The waveform in the graph of result RS51 indicates PR AUC. For example, the conventional model is a model that uses the above-mentioned BERT, and is a model that accepts comments, which are sentences, as input.
[0166] Meanwhile, the result RS52 shown on the right side of FIG. 18 indicates the experimental result when the model of this method is used. The vertical axis of the result RS52 indicates precision (the precision rate), and the horizontal axis of the result RS52 indicates recall (the recall rate). The waveform in the graph of the result RS52 indicates the PR AUC. For example, the model of this method is a model obtained by learning the above-mentioned base model (partial model PM1, etc.) using data from a comment service to learn a base model suitable for a comment service. In addition, the model of this method is a model that accepts as input a combination of information in which the headlines of news articles are converted into text and comments, which are sentences.
[0167] As shown in Results RS51 and RS52 in Figure 18, the model of this method achieved a 78.3% improvement in accuracy over the conventional model. In this way, it was shown that it is possible to further improve the accuracy of determining whether a comment is a violation by learning the base model of a model that can input types of information that could not be accepted as text before using data from a corresponding service.
[0168] [8. Learning process example] The information processing system 1 may learn various models such as the above-mentioned model M1 and model M11 by various learning methods. For example, the information processing system 1 may learn various partial models PM such as the partial model PM1, which is a base model of the model M1, by any learning method.
[0169] For example, the information processing device 10 may generate a partial model PM1 that is a base model of the model M1 through a multi-stage learning process. For example, the information processing device 10 may generate the partial model PM1 that is a base model of the model M1 through a multi-stage learning process including a first-stage learning process using a first data group used for learning general language ability among the learning data.
[0170] Furthermore, the information processing device 10 may generate the partial model PM1, which is a base model of the model M1, through a multi-stage learning process including a second-stage learning process using a second data group used for learning language ability for a specific process. For example, the second data group is data posted on the Internet used for learning a model used for processing related to posts on the Internet.
[0171] Furthermore, the information processing device 10 may generate a partial model PM1 that is a base model of the model M1 through a multiple-stage learning process including a third-stage learning process in which layers of the base model are increased and learning is performed. The data group used in the third-stage learning process may be the second data group. For example, the information processing device 10 may repeatedly execute the third-stage learning process to increase the layers of the base model and perform a learning process on the base model after the increase, thereby generating a partial model PM1 that is a base model of the model M1.
[0172] An example of the learning process of the partial model PM1, which is a base model of the above-mentioned model M1, will be described with reference to Figs. 19 to 21. Figs. 19 to 21 are diagrams showing an example of the learning process of a model according to the embodiment. For example, Fig. 19 is a diagram showing an example of a first-stage learning process targeting the partial model PM1. Fig. 20 is a diagram showing an example of a second-stage learning process targeting the partial model PM1. Fig. 21 is a diagram showing an example of a third-stage learning process targeting the partial model PM1. Note that explanations of points similar to those described above will be omitted as appropriate.
[0173] First, the information processing device 10 learns the partial model PM1, which is a base model of the model M1, by a first-stage learning process as shown in FIG. 19. The first-stage learning process shown in FIG. 19 is performed in a state where the output of the partial model PM1 is input to a partial model PM3 denoted as "Language Masked Prediction Layer". For example, the partial model PM3 is a model that predicts a masked character string in an input sentence (text). In this way, when learning the partial model PM1, which is a base model, the information processing device 10 may replace another partial model (for example, the partial model PM2) included in a model (for example, the model M1) used for a service after learning with a partial model (for example, the partial model PM3) having a different function and perform learning.
[0174] For example, in the first-stage learning process shown in FIG. 19, natural language data or the like used for learning a large-scale language model is used as the first data group FD1. Any data set such as Wikipedia, CC-100, or OSCAR Data may be used as the first data group FD1. For example, the information processing device 10 executes the first-stage learning process using documents (also referred to as "first sentences") included in the first data group FD1. For example, the information processing device 10 learns the partial model PM1 by learning the partial model PM3 to accurately predict character strings in the masked parts of the first sentences using documents in which parts of the first sentences included in the first data group FD1 are masked as input. This allows the information processing device 10 to learn (generate) the partial model PM1 that has acquired general-purpose language ability.
[0175] Next, the information processing device 10 learns a partial model PM1, which is a base model of the model M1, through a second-stage learning process as shown in Fig. 20. Note that in Fig. 20, explanations of points similar to those in Fig. 19 will be omitted as appropriate. For example, the information processing device 10 further learns the partial model PM1 through a second-stage learning process as shown in Fig. 20, using the partial model PM1 learned through the first-stage learning process shown in Fig. 19.
[0176] 20, the information processing device 10 further learns the partial model PM1 by additional learning using a second data group SD1 in a model configuration similar to that of the first-stage learning process. For example, the information processing device 10 learns the partial model PM1 by the second-stage learning process using data for fine tuning as the second data group SD1. For example, the information processing device 10 learns the partial model PM1 by the second-stage learning process using the second data group SD1 used for fine-tuning a model (e.g., model M1, etc.) used for determining whether a post is in violation.
[0177] Fig. 20 shows an example in which two sentences, Sentence#1 and Sentence#2, are input to partial model PM1. For example, in Fig. 20, a text in which the information indicating the question category has been converted into text, the token "SEP", and the text corresponding to the question (question sentence) are arranged in this order is input as Sentence#1 to partial model PM1. That is, in Fig. 20, a single text in which the information indicating the question category has been converted into text and the text corresponding to the question are linked by the token "SEP" is used for Sentence#1.
[0178] Furthermore, in Fig. 20, text corresponding to the answer (answer sentence) is input as Sentence#2 to the partial model PM1. That is, in Fig. 20, text corresponding to the answer is used for Sentence#2.
[0179] In this way, the information processing device 10 executes the second stage of learning processing using the second data group SD1, which is data for fine tuning to learn a model for inference processing related to posts such as questions and answers. For example, the information processing device 10 executes the second stage of learning processing using documents (also referred to as "second sentences") included in the second data group SD1. For example, the information processing device 10 learns the partial model PM1 by using documents in which a part of each second sentence included in the second data group SD1 is masked as input, and learning so that the partial model PM3 accurately predicts character strings in the masked parts of the second sentences. This allows the information processing device 10 to learn (generate) the partial model PM1 that has acquired language ability for specific processing.
[0180] Next, the information processing device 10 learns a partial model PM1, which is a base model of the model M1, through a third-stage learning process as shown in Fig. 21. Note that in Fig. 21, explanations of points similar to those in Fig. 19 and Fig. 20 will be omitted as appropriate. For example, the information processing device 10 further learns the partial model PM1 through a third-stage learning process as shown in Fig. 21, using the partial model PM1 learned through the second-stage learning process as shown in Fig. 20.
[0181] In Fig. 21, the information processing device 10 further learns the partial model PM1 through a third-stage learning process in which a layer is added to the partial model PM1 learned through the first-stage and second-stage learning processes and learning is performed. In Fig. 21, as shown in the partial model PM1-1, the information processing device 10 further learns the partial model PM1 through a third-stage learning process in which learning is performed by adding an additional layer such as layer AL1. For example, the information processing device 10 learns the partial model PM1 by adding a layer to the partial model PM1 through the third-stage learning process using data for fine tuning as the second data group SD1.
[0182] For example, the information processing device 10 may increase data in the third stage learning process by adding new records or duplicate copies. In addition, the information processing device 10 may perform the third stage learning process by optimizing the learning rate or setting the scheduler to constant. For example, the information processing device 10 may perform the third stage learning process by setting the batch size to an arbitrary value (e.g., 34,560, etc.). For example, the information processing device 10 can stabilize learning and perform optimization to shorten the learning time by setting the batch size to 6,900 or more. In addition, for example, the MLM (Masked Language Model) probability is an optimization target, and may be in the range of 0.15 to 0.45.
[0183] 21 is merely one example of the configuration of the partial model PM1, and two or more layers may be added to the partial model PM1. For example, the information processing device 10 may repeatedly execute the learning process of the third stage to increase the layers of the base model and to perform a learning process on the increased base model, thereby adding further layers to the partial model PM1 to generate the partial model PM1.
[0184] 8-1. Experimental Results Next, an example of the results of an experiment performed using a model generated by the above-described multi-stage learning process will be described with reference to Figures 22 to 25. Figures 22 to 25 are diagrams showing an example of the results of the experiment.
[0185] First, an example shown in Fig. 22 will be described. Fig. 22 shows an example of an experimental result for a base model (e.g., partial model PM1) used in a knowledge sharing service through posting of questions, answers, etc. In the result RS61 in Fig. 22, the horizontal axis indicates the number of steps related to the learning process, and the vertical axis indicates the accuracy of the MLM Task using the base model.
[0186] Line LN11 in Fig. 22 indicates the accuracy of the MLM Task using the foundation model in the first stage of learning processing (Step 1 in Fig. 22). Fig. 22 shows that the number of layers of the foundation model trained in the first stage of learning processing is 26, and the accuracy of the MLM Task using that foundation model is "0.7280".
[0187] Line LN12 in Fig. 22 shows the accuracy of the MLM Task using the foundation model in the second stage of learning processing (Step 2 in Fig. 22). Fig. 22 shows that the number of layers of the foundation model trained in the second stage of learning processing is 26, and the accuracy of the MLM Task using that foundation model has increased to "0.7724".
[0188] Line LN13 in Fig. 22 shows the accuracy of the MLM Task using the foundation model in the third stage of learning processing (Step 3 in Fig. 22). Fig. 22 shows that the number of layers of the foundation model trained in the third stage of learning processing is 28, and the accuracy of the MLM Task using that foundation model has increased to "0.7821".
[0189] Thus, in the result RS61 of Fig. 22, as shown by lines LN11 to LN13, for the base model optimized for the knowledge sharing service, the second-stage learning process and the third-stage learning process achieved a 7.4% improvement in accuracy. In this way, it was shown that the information processing device 10 is capable of improving accuracy by multiple stages of learning processes.
[0190] Next, an example shown in Fig. 23 will be described. An example of an experiment result when determining whether an answer posting is in violation in a knowledge sharing service that posts questions, answers, etc. will be shown. Note that a description of points similar to those described above will be omitted as appropriate.
[0191] Result RS71 shown on the left side of FIG. 23 indicates the experimental result when a conventional model is used. The vertical axis of result RS71 indicates precision, and the horizontal axis of result RS71 indicates recall. The waveform in the graph of result RS71 indicates a PR curve. For example, the conventional model is a model that uses the above-mentioned DeBERTa, and is a model that accepts answers that are text as input. Result RS71 is a model that has undergone learning processing corresponding to the first stage learning processing, and indicates the experimental result when a model with 26 layers is used.
[0192] Meanwhile, the result RS72 shown on the right side of FIG. 23 indicates the experimental result when the model of this method is used. The vertical axis of the result RS72 indicates precision (fit rate), and the horizontal axis of the result RS72 indicates recall (recall rate). The waveform in the graph of the result RS72 indicates the PR curve. For example, the model of this method is a model in which a base model (partial model PM1, etc.) trained by the above-mentioned multiple-stage learning process is subjected to second and third-stage learning process using data from a knowledge sharing service, and the base model is optimized for the knowledge sharing service, and is a model that accepts answers that are text as input. The result RS72 is a model that has undergone learning process including all of the first to third stages of learning process, and indicates the experimental result when a model with 28 layers is used.
[0193] As shown in Results RS71 and RS72 in Figure 23, the model of this method achieved a 26.6% improvement in accuracy over the conventional model. In this way, it was shown that it is possible to further improve the accuracy of judging whether an answer is a violation by learning through a multi-stage learning process.
[0194] First, the example shown in Fig. 24 will be described. Fig. 24 shows an example of an experimental result for a base model (e.g., partial model PM1) used for a comment service for news articles. Note that explanations of points similar to those described above will be omitted as appropriate. In the result RS81 in Fig. 24, the horizontal axis indicates the number of steps related to the learning process, and the vertical axis indicates the accuracy of the MLM Task using the base model.
[0195] Line LN20 in Fig. 24 shows the accuracy of the MLM Task using the conventional model. In Fig. 24, the conventional model has 12 layers of the base model trained based on DeBERTa, and the accuracy of the MLM Task using the base model is "0.6790".
[0196] Line LN21 in Fig. 24 indicates the accuracy of the MLM Task using the foundation model in the first stage of learning processing (Step 1 in Fig. 24). Fig. 24 shows that the number of layers of the foundation model trained in the first stage of learning processing is 26, and the accuracy of the MLM Task using that foundation model is "0.7109".
[0197] Line LN22 in Fig. 24 indicates the accuracy of the MLM Task using the foundation model in the second stage of learning processing (Step 2 in Fig. 24). Fig. 24 shows that the number of layers of the foundation model trained in the second stage of learning processing is 26, and the accuracy of the MLM Task using that foundation model has increased to "0.7441".
[0198] Line LN23 in Fig. 24 indicates the accuracy of the MLM Task using the foundation model in the third stage of learning processing (Step 3 in Fig. 24). Fig. 24 shows that the number of layers of the foundation model trained in the third stage of learning processing is 28, and the accuracy of the MLM Task using that foundation model has increased to "0.7741".
[0199] Thus, in the result RS81 of Fig. 24, as shown by lines LN21 to LN23, for the base model optimized for the comment service on news articles, the second-stage learning process and the third-stage learning process achieved an accuracy improvement of 8.9%. In this way, it was shown that the information processing device 10 is capable of improving accuracy by multiple stages of learning processes.
[0200] Next, an example shown in Fig. 25 will be described. An example of an experiment result when determining whether a comment posted in violation of the rules in a comment service for a news article is shown. Note that a description of the same points as those described above will be omitted as appropriate.
[0201] Result RS91 shown on the left side of FIG. 25 indicates the experimental result when a conventional model is used. The vertical axis of result RS91 indicates precision, and the horizontal axis of result RS91 indicates recall. The waveform in the graph of result RS91 indicates PR AUC. For example, the conventional model is a model that uses the above-mentioned DeBERTa, and is a model that accepts comments, which are text, as input. Result RS91 is a model that has undergone learning processing corresponding to the first stage learning processing, and indicates the experimental result when a model with 26 layers is used.
[0202] On the other hand, the result RS92 shown on the right side of FIG. 25 indicates the experimental result when the model of this method is used. The vertical axis of the result RS92 indicates precision (the accuracy rate), and the horizontal axis of the result RS92 indicates recall (the recall rate). The waveform in the graph of the result RS92 indicates the PR AUC. For example, the model of this method is a model in which a base model (such as the partial model PM1) trained by the above-mentioned multiple-stage learning process is subjected to second and third-stage learning process using data of a comment service for news articles, and the base model is optimized for the comment service for news articles, and is a model that accepts comments, which are text, as input. The result RS92 is a model that has undergone learning process including all of the first to third stages of learning process, and indicates the experimental result when a model with 28 layers is used.
[0203] As shown in Results RS91 and RS92 in Figure 25, the model of this method achieved a 29.7% improvement in accuracy over the conventional model. In this way, it was shown that by learning through a multi-stage learning process, it is possible to further improve the accuracy of judging violations posted by comment services on news articles.
[0204] [8-2. Other processing examples] From here, other processing examples will be described based on the above content. For example, in the above example, text into which information other than sentence text, which is a sentence, is converted is described as an example of converted information, but the input to the model is not limited to text.
[0205] For example, the information processing system 1 may generate a base model and a model (also called a "fine-tuning model") that is fine-tuned so that the base model is applied to a specified task by learning using information other than text. For example, the information processing system 1 learns a partial model PM1 or the like that is a base model by learning using information other than text. For example, the information processing system 1 learns a fine-tuning model (e.g., model M1, etc.) that includes the base model (e.g., partial model PM1, etc.) and is fine-tuned so that it is applied to a specified task by learning using information other than text. For example, the information processing system 1 learns the fine-tuning model by using data for fine tuning.
[0206] This allows the information processing system 1 to improve the accuracy of both the base model and the fine-tuning model. In addition, the information processing system 1 can apply various language models, including not only the BERT-based model related to the BERT described above but also the GPT-based model related to GPT (Generative Pretrained Transformer), to tabular data (tabular format data) other than text data.
[0207] For example, the information processing system 1 may generate a base model and a fine-tuning model that has input information different from that of the base model. For example, the information processing system 1 learns the fine-tuning model (e.g., model M1, etc.) by using an input of the fine-tuning model (e.g., model M1, etc.) including the base model (e.g., partial model PM1, etc.) as an input different from that of the base model (e.g., partial model PM1, etc.).
[0208] This allows the information processing system 1 to freely set the input order of text and non-text information for both the base model and the fine-tuning model, including duplication and deletion of input, which enables the information processing system 1 to automatically optimize the input order and improve the accuracy of the base model and the fine-tuning model.
[0209] For example, the information processing system 1 may generate a base model and a fine-tuning model in which the input order of input information is the same as that of the base model. For example, the information processing system 1 trains the fine-tuning model (e.g., model M1, etc.) by setting the input of the fine-tuning model (e.g., model M1, etc.) including the base model (e.g., partial model PM1, etc.) to the same input order as that of the base model (e.g., partial model PM1, etc.).
[0210] This allows the information processing system 1 to make the input order of text and non-text information the same for the base model and the fine-tuning model, which enables the information processing system 1 to use features related to the positions (token positions) of the input information learned by the base model, thereby improving the accuracy of the fine-tuning model.
[0211] For example, the information processing system 1 may convert non-text information such as integers, real numbers, and character strings into non-text information. For example, the information processing system 1 may convert into any format such as "identity", "vocabulary", "numeric", "bucketize", "identity + embedding", "vocabulary + embedding", and "bucketize + embedding". For example, when inputting non-text information such as integers, real numbers, and character strings, the information processing system 1 considers that the accuracy of the model changes depending on whether the non-text information such as integers, real numbers, and character strings is input as is or whether the non-text information such as integers, real numbers, and character strings is converted and input.
[0212] Therefore, the information processing device 10 converts (determines) the format of information other than text, such as integers, real numbers, and character strings, to be input to the model as described above. For example, the information processing device 10 determines whether the format of information other than text, such as integers, real numbers, and character strings, to be input to the model should be the same format, a text format, or another format. For example, when inputting an integer or real number, the information processing system 1 determines whether to input it as a numerical value, convert it into text, or convert it into another format such as a character string. For example, when inputting a character string, the information processing system 1 determines whether to input it as a string, convert it into text, or convert it into another format such as a numerical value. For example, the information processing system 1 determines to use the input format when the most accurate model is generated.
[0213] Through the above-mentioned processing, the information processing device 10 can optimize the format of the information (features) to be input. This allows the information processing system 1 to automatically optimize the above-mentioned arbitrary formats and their combinations using a feature optimization algorithm. This allows the information processing system 1 to improve the accuracy of the fine-tuning model.
[0214] For example, the information processing system 1 may use label information as input information in training the base model. This allows the information processing system 1 to input label information to the base model as input information (input feature) during pre-training of the base model. In addition, it becomes possible to mask label information during pre-training, and the information processing system 1 can train the base model so that label information can also be predicted during pre-training. This allows the information processing system 1 to improve the accuracy of the fine-tuning model.
[0215] 9. Modifications An example of the information processing has been described above. However, the embodiment is not limited to this. Below, a modified example of the information processing will be described.
[0216] [9-1. Equipment configuration] In the above embodiment, an example has been described in which the information processing system 1 includes the information processing device 10 that generates generation indicators and the model generation server 2 that generates a model according to the generation indicators, but the embodiment is not limited to this. For example, the information processing device 10 may have the functions of the model generation server 2. Furthermore, the functions performed by the information processing device 10 may be included in the terminal device 3. In such a case, the terminal device 3 will automatically generate generation indicators and automatically generate a model using the model generation server 2.
[0217] [9-2.Other] In addition, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by a known method. In addition, the information including the processing procedures, specific names, various data and parameters shown in the above documents and drawings can be changed arbitrarily unless otherwise specified. For example, the various information shown in each drawing is not limited to the illustrated information.
[0218] In addition, each component of each device shown in the figure is a functional concept, and does not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads, usage conditions, etc.
[0219] Furthermore, the above-described embodiments can be appropriately combined as long as the processing contents are not contradictory.
[0220] [9-3. Program] The information processing device 10 according to the embodiment described above is realized by a computer 1000 having a configuration as shown in Fig. 26, for example. Fig. 26 is a diagram showing an example of a hardware configuration. The computer 1000 is connected to an output device 1010 and an input device 1020, and has a configuration in which a calculation device 1030, a primary storage device 1040, a secondary storage device 1050, an output IF (Interface) 1060, an input IF 1070, and a network IF 1080 are connected by a bus 1090.
[0221] The arithmetic device 1030 operates based on programs stored in the primary storage device 1040 and the secondary storage device 1050 and programs read from the input device 1020, and executes various processes. The primary storage device 1040 is a memory device, such as a RAM, that temporarily stores data used by the arithmetic device 1030 for various calculations. The secondary storage device 1050 is a storage device in which data used by the arithmetic device 1030 for various calculations and various databases are registered, and is realized by a ROM (Read Only Memory), HDD, flash memory, or the like.
[0222] The output IF 1060 is an interface for transmitting information to be output to an output device 1010 that outputs various types of information, such as a monitor or a printer, and is realized by a connector conforming to a standard such as USB (Universal Serial Bus), DVI (Digital Visual Interface), or HDMI (registered trademark) (High Definition Multimedia Interface). The input IF 1070 is an interface for receiving information from various input devices 1020, such as a mouse, keyboard, and scanner, and is realized by a USB, for example.
[0223] The input device 1020 may be a device that reads information from, for example, an optical recording medium such as a CD (Compact Disc), a DVD (Digital Versatile Disc), or a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory. The input device 1020 may also be an external storage medium such as a USB memory.
[0224] The network IF 1080 receives data from other devices via the network N and sends it to the arithmetic device 1030, and also transmits data generated by the arithmetic device 1030 to other devices via the network N.
[0225] The arithmetic unit 1030 controls the output device 1010 and the input device 1020 via the output IF 1060 and the input IF 1070. For example, the arithmetic unit 1030 loads a program from the input device 1020 or the secondary storage device 1050 onto the primary storage device 1040 and executes the loaded program.
[0226] For example, when the computer 1000 functions as the information processing device 10 , the arithmetic unit 1030 of the computer 1000 executes a program loaded onto the primary storage device 1040 to realize the functions of the control unit 40 .
[0227] [10. Effects] As described above, the information processing device 10 has an acquisition unit (acquisition unit 41 in the embodiment) that acquires learning data used for learning a base model (for example, partial model PM1 in the embodiment) that uses text as input, and a generation unit (generation unit 44 in the embodiment) that generates the base model through a multi-stage learning process using the learning data. This allows the information processing device 10 to appropriately generate the base model.
[0228] The acquisition unit acquires learning data including a first data group used for learning the general language ability. The generation unit generates a base model through a multi-stage learning process including a first-stage learning process using the first data group. This allows the information processing device 10 to generate a model through a multi-stage learning process including a first-stage learning process for learning the general language ability, and therefore the base model can be appropriately generated.
[0229] The acquisition unit also acquires a first data group including natural language data used to train the large-scale language model. This allows the information processing device 10 to generate a model that has acquired general-purpose language capabilities through multiple stages of learning processing including a first stage of learning processing using the natural language data used to train the large-scale language model, thereby making it possible to appropriately generate a base model.
[0230] The acquisition unit also acquires learning data including a second data group used to learn language skills for a specific process. The generation unit generates a base model through a multi-stage learning process including a second-stage learning process using the second data group. This allows the information processing device 10 to generate a model through a multi-stage learning process including a second-stage learning process for learning language skills for a specific process, and therefore allows the base model to be generated appropriately.
[0231] The acquisition unit also acquires learning data including a second data group that is data related to posts on the Internet. This allows the information processing device 10 to generate a model suitable for a specific process through a multiple-stage learning process including a second-stage learning process using data related to posts on the Internet, thereby enabling the information processing device 10 to appropriately generate a base model.
[0232] The acquisition unit also acquires a second data group including text corresponding to the sentence posted on the Internet. This allows the information processing device 10 to generate a model suitable for processing the posted sentence through a multi-stage learning process including a second-stage learning process using the text corresponding to the sentence posted on the Internet, thereby enabling the information processing device 10 to appropriately generate a base model.
[0233] The acquisition unit also acquires a second data group including text corresponding to a question posted on the Internet. This allows the information processing device 10 to generate a model suitable for processing the posted question through a multiple-stage learning process including a second-stage learning process using the text corresponding to the question posted on the Internet, thereby enabling the information processing device 10 to appropriately generate a base model.
[0234] The acquisition unit also acquires a second data group including text corresponding to the answers posted on the Internet. This allows the information processing device 10 to generate a model suitable for processing the posted answers through a multi-stage learning process including a second-stage learning process using the text corresponding to the answers posted on the Internet, thereby enabling the information processing device 10 to appropriately generate a base model.
[0235] The generation unit generates the base model through a multi-stage learning process including a third-stage learning process in which the layers of the base model are increased and learning is performed. This allows the information processing device 10 to generate a model through a multi-stage learning process including a third-stage learning process in which the layers of the base model are increased and learning is performed, and therefore the base model can be appropriately generated.
[0236] The generation unit generates the base model through a third-stage learning process using data used to learn language skills for a specific process. This allows the information processing device 10 to generate a model suitable for a specific process through multiple stages of learning processes including the third-stage learning process using data used to learn language skills for a specific process, and therefore allows the information processing device 10 to appropriately generate the base model.
[0237] The generating unit generates the base model through a third-stage learning process using data related to posts on the Internet. This allows the information processing device 10 to generate a model suitable for processing posted text through multiple stages of learning processes including the third-stage learning process using data related to posts on the Internet, and therefore allows the information processing device 10 to appropriately generate the base model.
[0238] The generation unit also generates the base model through a third-stage learning process using text corresponding to the sentence posted on the Internet. This allows the information processing device 10 to generate a model suitable for processing the posted sentence through multiple stages of learning processes including the third-stage learning process using text corresponding to the sentence posted on the Internet, and therefore allows the information processing device 10 to appropriately generate the base model.
[0239] In addition, the generation unit repeatedly executes the third stage learning process to increase the layers of the base model and to perform a learning process on the increased base model, thereby generating a base model. This allows the information processing device 10 to appropriately generate the base model by increasing the layers of the base model and to perform a learning process on the increased base model.
[0240] In addition, the generation unit generates a base model and a fine-tuning model that is fine-tuned so as to apply the base model to a predetermined task by learning using information other than text. This allows the information processing device 10 to properly learn both the base model and the fine-tuning model, and therefore to generate a model that can properly input information other than text.
[0241] The generation unit also generates a base model and a fine-tuning model that receives input information different from that of the base model. This allows the information processing device 10 to properly learn both the base model and the fine-tuning model that receives input information different from that of the base model, and therefore allows the information processing device 10 to generate a model that can properly input information other than text.
[0242] The generation unit also generates a base model and a fine-tuning model in which the input order of the input information is the same as that of the base model. This allows the information processing device 10 to properly learn both the base model and the fine-tuning model in which the input order is the same as that of the base model, and therefore allows the information processing device 10 to generate a model in which information other than text can be properly input.
[0243] The generation unit generates the base model and the fine-tuning model using a plurality of pieces of input information including converted information obtained by converting information other than sentence text, which is a sentence, as input. This allows the information processing device 10 to properly learn both the base model and the fine-tuning model, which can input information other than sentences as text, and therefore allows the information processing device 10 to generate a model that can properly input information other than sentences.
[0244] In addition, the generation unit uses the label information as input information in learning the base model. This allows the information processing device 10 to appropriately learn the base model using the label information, and therefore, it is possible to generate a model that can appropriately input information other than text.
[0245] Although some of the embodiments of the present application have been described in detail above with reference to the drawings, these are merely examples, and the present invention can be embodied in other forms that incorporate various modifications and improvements based on the knowledge of those skilled in the art, including the forms described in the Disclosure of the Invention section.
[0246] Moreover, the above-mentioned "section, module, unit" can be read as "means" or "circuit", etc. For example, an acquisition section can be read as an acquisition means or an acquisition circuit. [Explanation of symbols]
[0247] 1. Information Processing Systems 2. Model Generation Server 3 Terminal Equipment 10. Information processing device 20 Communications Department 30 Storage section 40 Control section 41 Acquisition Department 42 Decision Section 43 Reception 44 Generation part 45 Processing section 46 Providing Department
Claims
1. 1. A computer-implemented information processing method, comprising: An acquisition step of acquiring learning data used for training a foundation model using text as input; A generation process of generating the base model through a multiple-stage learning process using the learning data; 13. An information processing method comprising:
2. The obtaining step includes: acquiring the learning data including a first data group used for learning general language skills; The generating step includes: The base model is generated by the multiple-stage learning process including the first-stage learning process using the first data group.
2. The information processing method according to claim 1,
3. The obtaining step includes: Obtaining the first set of data including natural language data used to train a large-scale language model.
3. The information processing method according to claim 2.
4. The obtaining step includes: acquiring the training data including a second data group used for training language skills for a specific process; The generating step includes: The base model is generated by the multiple-stage learning process including the second-stage learning process using the second data group.
2. The information processing method according to claim 1,
5. The obtaining step includes: The learning data includes the second data group, which is data related to posts on the Internet.
5. The information processing method according to claim 4.
6. The obtaining step includes: The second data set includes text corresponding to the article posted on the Internet.
6. The information processing method according to claim 5,
7. The obtaining step includes: and acquiring the second set of data including text corresponding to a question posted on the Internet.
7. The information processing method according to claim 6,
8. The obtaining step includes: and obtaining the second set of data including text corresponding to the answers posted on the Internet.
7. The information processing method according to claim 6,
9. The generating step includes: The base model is generated by the multiple-stage learning process including a third-stage learning process in which layers of the base model are increased and learning is performed.
2. The information processing method according to claim 1,
10. The generating step includes: The base model is generated by the third stage learning process using data used for learning language skills for a specific process.
10. The information processing method according to claim 9.
11. The generating step includes: The base model is generated by the third stage learning process using data related to posts on the Internet.
11. The information processing method according to claim 10.
12. The generating step includes: The base model is generated by the third stage learning process using text corresponding to sentences posted on the Internet.
12. The information processing method according to claim 11.
13. The generating step includes: By repeatedly executing the third stage learning process, a process of increasing layers of the base model and a learning process for the increased base model are performed to generate the base model.
10. The information processing method according to claim 9.
14. The generating step includes: By learning using information other than text, the base model and a fine-tuned model that is fine-tuned to apply the base model to a predetermined task are generated.
2. The information processing method according to claim 1,
15. The generating step includes: The base model and the fine-tuning model having input information different from that of the base model are generated.
15. The information processing method according to claim 14.
16. The generating step includes: Generate the base model and the fine-tuning model in which the input information is input in the same order as that of the base model.
15. The information processing method according to claim 14.
17. The generating step includes: The base model and the fine-tuning model are generated by inputting a plurality of pieces of input information including converted information in which information other than the text of the sentence is converted.
15. The information processing method according to claim 14.
18. The generating step includes: In learning the base model, label information is used as input information.
2. The information processing method according to claim 1,
19. An acquisition unit that acquires learning data used for training a foundation model using text as input; A generation unit that generates the base model through a multiple-stage learning process using the learning data; An information processing device having the above configuration.
20. An acquisition step of acquiring training data used to train a foundation model using text as input; A generation step of generating the base model by a multiple-stage learning process using the learning data; An information processing program for causing a computer to execute the above.
Citation Information
Patent Citations
Neural network structure extension method, dimension reduction method, and device using method
JP2016103262A
Information processing apparatus, and information processing method
JP2021051589A
Leaning device, learning method, and learning program
JP2022047529A
Learning device for document classification, document classifier, and program
JP2023136771A
Model generation device and model generation method
WO2022180989A1