Machine learning program, machine learning method, and information processing apparatus

By identifying and vectorizing proper names and verbs in sentences, the method improves data selection for domain adaptation, enhancing model accuracy and reducing adaptation time in machine learning models.

JP7707638B2Active Publication Date: 2025-07-15FUJITSU LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021080360
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-11
Publication Date
2025-07-15
Estimated Expiration
2041-05-11

AI Technical Summary

Technical Problem

Inaccurate selection of training data during domain adaptation in machine learning models leads to deterioration in model accuracy, particularly when applying a pre-trained model to a target domain with multiple subdomains.

Method used

The method involves identifying proper names and verbs with dependency relationships in sentences, vectorizing these elements, and selecting similar sentences based on threshold values to train the model, thereby improving data selection for domain adaptation.

Benefits of technology

This approach enhances the accuracy of machine learning models by ensuring appropriate training data is used, reducing model deterioration and shortening adaptation time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007707638000001
    Figure 0007707638000001
  • Figure 0007707638000002
    Figure 0007707638000002
  • Figure 0007707638000003
    Figure 0007707638000003
Patent Text Reader

Abstract

To provide a machine learning program, a machine learning method, and an information processing apparatus that can prevent a deterioration in the accuracy of a machine learning model.SOLUTION: An information processing apparatus has: a specification unit that specifies, from each of a plurality of sentences, named entries and verbs having dependency with the named entries; a vectorization processing unit that, based on the named entries and the verbs having dependency with the named entries, vectorizes each of the plurality of sentences; a selection unit that, based on a plurality of vectors generated by the vectorization processing unit, specifies, from the plurality of sentences, one or more sentences having similarity equal to or more than a threshold to a specific sentence; and a training unit that, based on the one or more sentences selected by the selection unit, executes training of the machine learning model, generates a language model adapted to domains, and stores the language model in a storage unit.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the generation of machine learning models.

Background Art

[0002] In many fields that utilize machine learning models, techniques related to domain adaptation are used, where a machine learning model generated using training data from one domain is applied to another domain. Domain adaptation applies knowledge obtained from a source domain with sufficient training data to a target domain, thereby generating an identifier or the like that operates with high accuracy in the target domain. Here, a domain refers to, for example, a collection of data.

[0003] For example, in the field of natural language processing, when applying a pre-trained language model generated using a source domain to the target domain side, the pre-trained language model is retrained using the training data on the target domain side.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, when retraining a machine learning model using training data on the domain side, inappropriate training data may be included, and the accuracy of the machine learning model after retraining may deteriorate. For example, the training data within the target domain includes training data belonging to various subdomains. When performing retraining of a machine learning model applicable to a specific subdomain, the corresponding training data is selected from the target domain. However, if this selection is not accurate, training data from various subdomains will be included, deteriorating the accuracy of the machine learning model.

[0006] In one aspect, an object is to provide a machine learning program, a machine learning method, and an information processing apparatus capable of suppressing deterioration in the accuracy of a machine learning model.

Means for Solving the Problem

[0007] In the first aspect, the machine learning program causes a computer to execute a process of identifying, from each of a plurality of sentences, a proper name and a verb having a dependency relationship with the proper name, vectorizing each of the plurality of sentences based on the proper name and the verb having the dependency relationship with the proper name, identifying one or more sentences among the plurality of sentences that are similar to a specific sentence by a threshold or more based on the plurality of vectors generated by the vectorization process, and performing training of a machine learning model based on the one or more sentences.

Effects of the Invention

[0008] According to one embodiment, deterioration in the accuracy of a machine learning model can be suppressed.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, examples of the machine learning program, machine learning method, and information processing apparatus disclosed in the present application will be described in detail with reference to the drawings. Note that the present invention is not limited to these examples. Also, the respective examples can be appropriately combined within a non - contradictory range.

[0011] FIG. 1 is a diagram for explaining the information processing apparatus 10 according to Embodiment 1. The information processing apparatus 10 shown in FIG. 1 generates a machine learning model applicable to a certain task by extracting appropriate data from the data included in the corpus data and using the extracted data for machine learning as training data to generate a machine learning model.

[0012] Here, in this embodiment, as an example, the domain adaptation of the machine learning model will be described. However, it can also be applied to other situations such as the generation of the machine learning model. Specifically, an example will be described in which the information processing apparatus 10 domain-adapts the machine learning model generated using the data of the source domain as training data by retraining it using the data of the appropriate sub-domain 3 from the target domain (Target Domain) including a plurality of sub-domains 1, 2, and 3.

[0013] Here, as domain adaptation, a method of selecting training data used for domain adaptation based on the similarity between two sentences based on Bag-of-Words (BoW) is often used. However, in this method, when calculating the similarity, the named entity of the sentence and the syntactic information of the verb are not considered, so the data selection is not sufficient, and the accuracy of the machine learning model after domain adaptation may not be good.

[0014] For example, consider an example of domain-adapting a machine learning model to a biomedical sub-domain. That is, consider an example of only performing named entity recognition (NER) of the biomedical sub-domain for a downstream task. Since words such as "Lactococcus lactis" are used in both the biomedical sub-domain and the news sub-domain, both sub-domains are selected as corpus data (training data) for domain adaptation. As a result, since the machine learning model is trained to be applicable to both the biomedical sub-domain and the news sub-domain, the accuracy for the data (downstream task) of the biomedical sub-domain decreases.

[0015] Therefore, the information processing apparatus 10 according to the first embodiment suppresses the deterioration of the accuracy of the machine learning model by selecting training data for domain adaptation using syntactic information based on the combination of named entities and verbs appearing in the sentence.

[0016] Specifically, the information processing apparatus 10 identifies, from each of a plurality of sentences included in the target domain, an entity expression and a verb having a dependency relationship with the entity expression. Subsequently, the information processing apparatus 10 vectorizes each of the plurality of sentences based on the entity expression and the verb having a dependency relationship with the entity expression. Then, the information processing apparatus 10 identifies one or more sentences that are similar to or more similar than a threshold value to a specific sentence corresponding to a downstream task among the plurality of sentences based on the plurality of vectors generated by the vectorization process. Thereafter, the information processing apparatus 10 executes domain adaptation of the machine learning model by training the machine learning model based on the one or more sentences.

[0017] For example, the information processing apparatus 10 generates a vector (vector data) obtained by combining an entity expression and a verb as a comparison target for each sentence. Then, the information processing apparatus 10 selects, as training data for domain adaptation, a document that is similar to the vector of the sentence of the downstream task by comparing the vectors of each sentence. Thereafter, the information processing apparatus 10 executes retraining of the machine learning model using the selected training data (document).

[0018] As described above, since the information processing apparatus 10 vectorizes the feature amounts of the sentences of the downstream task, selects them as training data for domain adaptation by similarity determination using vector values, and executes retraining, it is possible to suppress deterioration in the accuracy of the machine learning model after domain adaptation.

[0019] FIG. 2 is a functional block diagram showing the functional configuration of the information processing apparatus 10 according to the first embodiment. As shown in FIG. 2, the information processing apparatus 10 includes a communication unit 11, a storage unit 12, and a control unit 20.

[0020] The communication unit 11 controls communication with other devices. For example, the communication unit 11 acquires a machine learning model generated using a source domain from an administrator terminal or the like, and transmits a processing result to the administrator terminal or the like by the control unit 20.

[0021] The memory unit 12 stores various data and programs executed by the control unit 20. This memory unit 12 stores the pre-trained language model 13, the task DB 14, the corpus data DB 15, and the language model 16.

[0022] The pre-trained language model 13 is a machine learning model generated using training data belonging to the source domain. For example, the pre-trained language model 13 is a machine learning model for domain adaptation target, and is an example of a machine learning model that performs named entity recognition extraction, and for example, converts a sentence into a vector representation.

[0023] The task DB 14 is a database that stores at least one sentence corresponding to a task to be determined by the machine learning model after domain adaptation. That is, the sentences stored in the task DB 14 correspond to the downstream tasks or specific sentences. For example, the task DB 14 stores sentences belonging to the biomedical subdomain.

[0024] The corpus data DB 15 is a database that stores sentences used for domain adaptation of the pre-trained language model 13. This corpus data DB 15 stores sentences classified into a plurality of subdomains corresponding to the target domain. FIG. 3 is a diagram showing an example of information stored in the corpus data DB 15. As shown in FIG. 3, the corpus data DB 15 stores sentences belonging to the news subdomain, sentences belonging to the biomedical subdomain, sentences belonging to the sports subdomain, and the like.

[0025] The language model 16 is a language model after domain application. That is, the language model 16 is a machine learning model for NER finally generated by the information processing apparatus 10. In the above example, the language model 16 is a machine learning model obtained by adapting the pre-trained language model 13 to the downstream task.

[0026] The control unit 20 is a processing unit that controls the entire information processing apparatus 10, and includes a specifying unit 21, a vectorization processing unit 22, a selection unit 23, and a training unit 24.

[0027] The specific part 21 identifies, from each of a plurality of sentences, an idiomatic expression and a verb having a dependency relationship with the idiomatic expression. For example, the specific part 21 identifies, for each sentence corresponding to the downstream task stored in the task DB 14 and each sentence belonging to the target domain stored in the corpus data DB 15, an idiomatic expression and a verb having a dependency relationship with the idiomatic expression. Here, as the dependency relationship, for example, distance or a preset combination can be adopted. For example, the specific part 21 identifies the verb that appears at the position closest to the idiomatic expression and generates a combination of the idiomatic expression and the verb.

[0028] The vectorization processing unit 22 vectorizes each of the plurality of sentences based on the idiomatic expression, the idiomatic expression, and the verb having a dependency relationship. Specifically, the vectorization processing unit vectorizes each sentence by vectorizing each combination identified by the specific part 21 for each sentence corresponding to the downstream task. Also, the vectorization processing unit vectorizes each sentence by vectorizing each combination identified by the specific part 21 for each sentence belonging to the target domain. An example of vectorization will be described later.

[0029] The selection unit 23 identifies one or more sentences among the sentences belonging to the target domain that are similar to the downstream task by a threshold or more based on the plurality of vectors generated by the vectorization processing by the vectorization processing unit 22. That is, the selection unit 23 selects a sentence suitable for domain adaptation.

[0030] The training unit 24 executes machine learning of the pre-trained language model 13 based on the one or more sentences selected by the selection unit 23. That is, the training unit 24 executes machine learning of the pre-trained language model 13 using the document of the target domain selected by the selection unit 23 to generate a domain-adapted language model 16. Then, the training unit 24 stores the generated language model 16 in the storage unit 12.

[0031] Here, the above-described domain adaptation process will be specifically described. FIG. 4 is a diagram for explaining a specific example of a set of proper expressions and verbs in a sentence. As an example, it will be described using a document belonging to a downstream task, but the same process is executed for each document belonging to the target domain.

[0032] As shown in FIG. 4, the specifying unit 21 performs morphological analysis and the like on the sentence 1 "the force-distance curves were analyzed to determine the physical and nanomechanical properties of L. lactis pili.". Then, the specifying unit 21 extracts "the force-distance", "L. lactis pili.", and "the physical and nanomechanical properties" as proper expressions. Similarly, the specifying unit 21 specifies "curves", "analyzed", and "determine" as verbs.

[0033] Next, the vectorization processing unit 22 vectorizes the document using the set of proper expressions and verbs. Specifically, the vectorization processing unit 22 vectorizes the set of the proper expressions specified by the specifying unit 21 and the verb closest to the proper expression to calculate "syntactic representation".

[0034] FIG. 5 is a diagram for explaining an example of calculating the syntactic representation of the verb set in each sentence, and FIG. 6 is a diagram for explaining an example of calculating the syntactic representation of a sentence.

[0035] As shown in FIG. 5, the vectorization processing unit 22 identifies, as the nearest verb sets (combinations), combination 1 "the force - distance, curves", combination 2 "L. lactis pili., determine", and combination 3 "the physical and nanomechanical properties, determine" according to the occurrence positions of each specific expression and each verb. Then, the vectorization processing unit 22 inputs each of combinations 1 to 3 into "word embedding architecture", which is an example of a generated machine learning model, and generates vector representations (vector data) emb(combination 1), emb(combination 2), and emb(combination 3).

[0036] In this way, the vectorization processing unit 22 generates vector representations "emb(combination 1), emb(combination 2), emb(combination 3)" for sentence 1 "the force - distance curves were analyzed to determine the physical and nanomechanical properties of L. lactis pili.".

[0037] After that, the vectorization processing unit 22 generates an integrated vector representation of the entire sentence 1. As shown in FIG. 6, for example, the vectorization processing unit 22 calculates the similarity of each of emb(combination 1), emb(combination 2), and emb(combination 3), and calculates the average value of the similarities as "syntactic representation". Note that known calculation methods such as cosine similarity and Euclidean distance can be adopted for calculating the similarity. Also, not limited to the average value of the similarities, the average value (average vector) or the total value of the vector representations may be used.

[0038] Next, the selection unit 23 selects corpus data for domain adaptation according to the similarity of the "syntactic representation" of each document generated by the vectorization processing unit 22.

[0039] FIG. 7 is a diagram for explaining an example of selecting corpus data. As shown in FIG. 7, for each of "Document 1, Document 2, Document 3" belonging to the downstream task, the vectorization processing unit 22 calculates the above "syntactic representation". Similarly, for each of "Document A, Document B, Document C..." belonging to the target domain, the vectorization processing unit 22 calculates the above "syntactic representation".

[0040] Then, the selection unit 23 calculates the similarity between the "syntactic representation" of each of "Document 1, Document 2, Document 3" belonging to the downstream task and the "syntactic representation" of each document belonging to the target domain. Note that for calculating the similarity, known calculation methods such as cosine similarity and Euclidean distance can be adopted.

[0041] Subsequently, the selection unit 23 calculates the average value of the similarities of each document (Document 1, Document 2, Document 3) belonging to the downstream task with respect to Document A in the target domain. That is, the selection unit 23 calculates the similarity between Document A in the target domain and Document 1, the similarity between Document A and Document 2, and the similarity between Document A and Document 3. Then, the selection unit 23 calculates the average value of each similarity with respect to Document A.

[0042] Similarly, the selection unit 23 calculates the average value of the similarities of each document (Document 1, Document 2, Document 3) belonging to the downstream task with respect to Document B in the target domain, and calculates the average value of the similarities of each document (Document 1, Document 2, Document 3) belonging to the downstream task with respect to Document C in the target domain. Then, the selection unit 23 selects the top k documents (Document A... Document L) with high average values among the documents in the target domain to generate new corpus data.

[0043] Next, the training unit 24 uses the document selected by the selection unit 23 to execute the training of the machine learning model. FIG. 8 is a diagram for explaining the training using corpus data. As shown in FIG. 8, the training unit 24 uses the top k documents, which are new corpus data, to execute the retraining of the pre-trained language model 13 and generate the language model 16 after domain adaptation.

[0044] Note that as the training method, a known training method for the machine learning model used in NER can be adopted. For example, when the downstream task is the "biomedical domain", the training unit 24 extracts and vectorizes the unique expressions of each selected document, and assigns the label "biomedical domain" to each vector expression obtained from the document. Then, the training unit 24 inputs each vector into the pre-trained language model 13 and executes the training of the pre-trained language model 13 so that the pre-trained language model 13 recognizes each unique expression as a unique expression in the "biomedical domain", and generates the language model 16 adapted to the domain of the downstream task.

[0045] Next, the above-described processing flow will be described. FIG. 9 is a flowchart showing the flow of the training process of the machine learning model. As shown in FIG. 9, the specifying unit 21 selects a downstream task (S101). For example, the specifying unit 21 selects one or more sentences of the downstream task according to an administrator's instruction, schedule, or the like.

[0046] Then, the vectorization processing unit 22 calculates the "syntactic representation" for each sentence of the downstream task (S102). For example, the vectorization processing unit 22 vectorizes the set of the unique expression specified by the specifying unit 21 and the verb closest to the unique expression to calculate the "syntactic representation".

[0047] Also, the specifying unit 21 selects each sentence of the target domain (S103). For example, the specifying unit 21 selects each sentence belonging to the target domain regardless of each subdomain in the target domain.

[0048] Then, for each sentence in the target domain, the vectorization processing unit 22 calculates a "syntactic representation" (S104). For example, the vectorization processing unit 22 vectorizes a set of the specific expressions identified by the specifying unit 21 and the verb closest to the specific expression to calculate the "syntactic representation".

[0049] After that, the selection unit 23 calculates the average value of the similarity with each document of the downstream task for each document belonging to the target domain (S105). For example, the selection unit 23 calculates the similarity between the "syntactic representation" of each document belonging to the target domain and the "syntactic representation" of each document belonging to the downstream task. Then, the selection unit 23 calculates the average value of the similarity for each document belonging to the target domain.

[0050] Then, the selection unit 23 selects the top k sentences with high similarity from each document belonging to the target domain (S106). After that, the training unit 24 generates a language model using the above k sentences as training data (S107).

[0051] As described above, since the information processing apparatus 10 can select an appropriate sentence from the target domain and generate a machine learning model by domain adaptation using the sentence, by using the machine learning model, the downstream task can be determined more accurately. Also, since the information processing apparatus 10 can suppress training using unnecessary training data, the time required for domain adaptation can be shortened.

[0052] In addition, the information processing apparatus 10 can generate a machine learning model adapted to a downstream task by executing each step (process) of vectorizing a sentence using an idiomatic expression, extracting a feature amount of the sentence, and selecting a domain adaptation sentence based on the feature amount. As a result, the downstream task can be determined more accurately.

[0053] In addition, since the information processing apparatus 10 can identify the verb closest to the idiomatic expression and generate a vector based on the set of the idiomatic expression and the verb, the accuracy of the vector expression representing the characteristics of the document can be improved. As a result, since the information processing apparatus 10 can select a similar document using an accurate vector expression, a high-precision machine learning model can be generated.

[0054] In addition, the information processing apparatus 10 can provide an application that executes each step (process) of vectorizing a sentence using an idiomatic expression, extracting a feature amount of the sentence, and selecting a domain adaptation sentence based on the feature amount. Further, the information processing apparatus 10 can also provide an application including up to generating a machine learning model adapted to a downstream task in the above steps.

[0055] The data examples, the above k (k is an arbitrary integer), numerical examples, number of domains, domain examples, sentences, specific examples, etc. used in the above embodiments are merely examples and can be arbitrarily changed.

[0056] Regarding the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above document and drawings, they can be arbitrarily changed unless otherwise specified.

[0057] In addition, each component of each illustrated apparatus is conceptually functional and does not necessarily need to be physically configured as shown in the figure. That is, the specific form of the dispersion and integration of each apparatus is not limited to that shown in the figure. That is, all or part of it can be functionally or physically dispersed and integrated in any unit according to various loads, usage situations, etc.

[0058] Furthermore, all or any part of each processing function performed by each device can be implemented by a CPU and a program analyzed and executed by the CPU, or can be implemented as hardware by wired logic.

[0059] FIG. 10 is a diagram for explaining a hardware configuration example. As shown in FIG. 10, the information processing apparatus 10 includes a communication device 10a, a HDD (Hard Disk Drive) 10b, a memory 10c, and a processor 10d. Further, each part shown in FIG. 10 is interconnected by a bus or the like.

[0060] The communication device 10a is a network interface card or the like and communicates with other devices. The HDD 10b stores a program and a DB for operating the functions shown in FIG. 5.

[0061] The processor 10d reads out a program for executing the same processing as each processing unit shown in FIG. 2 from the HDD 10b or the like and expands it in the memory 10c, thereby operating a process for executing each function described in FIG. 2 or the like. For example, this process executes the same functions as each processing unit included in the information processing apparatus 10. Specifically, the processor 10d reads out a program having the same functions as the specifying unit 21, the vectorization processing unit 22, the selection unit 23, the training unit 24, etc. from the HDD 10b or the like. Then, the processor 10d executes a process for executing the same processing as the specifying unit 21, the vectorization processing unit 22, the selection unit 23, the training unit 24, etc.

[0062] In this way, the information processing apparatus 10 operates as an information processing apparatus that executes a machine learning method by reading and executing a program. Further, the information processing apparatus 10 can also realize the same functions as those of the above-described embodiments by reading the program from the recording medium by the medium reading apparatus and executing the read program. Note that the program in this other embodiment is not limited to being executed by the information processing apparatus 10. For example, the present invention can be similarly applied when another computer or server executes the program, or when these cooperate to execute the program.

[0063] This program can be distributed via a network such as the Internet. Further, this program is recorded on a computer-readable recording medium such as a hard disk, a flexible disk (FD), a CD-ROM, a MO (Magneto-Optical disk), a DVD (Digital Versatile Disc), etc., and can be executed by being read from the recording medium by a computer.

Explanation of Signs

[0064] 10 Information processing apparatus 11 Communication unit 12 Storage unit 13 Pre-trained language model 14 Task DB 15 Corpus data DB 16 Language model 20 Control unit 21 Identification unit 22 Vectorization processing unit 23 Selection unit 24 Training unit

Claims

Claims 1. Identify a plurality of first combinations, each being a combination of an idiomatic expression and a verb having a dependency relationship with the idiomatic expression, from a specific document, and identify a plurality of second combinations, each being a combination of an idiomatic expression and a verb having a dependency relationship with the idiomatic expression, from each of a plurality of sentences. Generate a plurality of first vector values obtained by vectorizing each of the plurality of first combinations, and a plurality of second vector values obtained by vectorizing each of the plurality of second combinations. Calculate a first average value, which is the average value of the similarities of each of the plurality of first vector values corresponding to the specific document, and a plurality of second average values, which are the average values of the similarities of each of the plurality of second vector values corresponding to each of the plurality of sentences. Among the plurality of second average values corresponding to each of the plurality of sentences, identify the second average value whose similarity to the first average value of the specific document is equal to or greater than a threshold value. Identify one or more sentences corresponding to the identified second average value among the plurality of sentences. Execute training of a machine learning model based on the one or more sentences. A machine learning program characterized by causing a computer to execute the processing. Claims 2. The identifying process is Among the verbs included in the specific sentence, identify the verb having the closest distance from the idiomatic expression as the verb having a dependency relationship with the idiomatic expression, and identify the combination of the idiomatic expression and the verb having the closest distance as the first combination. Among the verbs included in the plurality of sentences, identify the verb having the closest distance from the idiomatic expression as the verb having a dependency relationship with the idiomatic expression, and identify the combination of the idiomatic expression and the verb having the closest distance as the second combination. The machine learning program according to claim 1, characterized in that. Claims 3. Identify a plurality of first combinations, each being a combination of an idiomatic expression and a verb having a dependency relationship with the idiomatic expression, from a specific document, and identify a plurality of second combinations, each being a combination of an idiomatic expression and a verb having a dependency relationship with the idiomatic expression, from each of a plurality of sentences. Generate a plurality of first vector values obtained by vectorizing each of the plurality of first combinations, and a plurality of second vector values obtained by vectorizing each of the plurality of second combinations. Calculate a first average value, which is the average value of the similarities of the plurality of first vector values corresponding to the specific document, and a plurality of second average values, which are the average values of the similarities of the plurality of second vector values corresponding to the plurality of sentences respectively. Among the plurality of second average values corresponding to the plurality of sentences respectively, identify the second average values whose similarity to the first average value of the specific document is equal to or greater than a threshold value. Among the plurality of sentences, identify one or more sentences corresponding to the identified second average values. Execute training of a machine learning model based on the one or more sentences. A machine learning method characterized in that a computer executes the processing.

4. Identify a plurality of first combinations, which are combinations of an idiomatic expression and a verb having a dependency relationship with the idiomatic expression, from a specific document, and identify a plurality of second combinations, which are combinations of an idiomatic expression and a verb having a dependency relationship with the idiomatic expression, from each of the plurality of sentences. Generate a plurality of first vector values obtained by vectorizing each of the plurality of first combinations, and a plurality of second vector values obtained by vectorizing each of the plurality of second combinations. Calculate a first average value, which is the average value of the similarities of the plurality of first vector values corresponding to the specific document, and a plurality of second average values, which are the average values of the similarities of the plurality of second vector values corresponding to the plurality of sentences respectively. Among the plurality of second average values corresponding to the plurality of sentences respectively, identify the second average values whose similarity to the first average value of the specific document is equal to or greater than a threshold value. Among the plurality of sentences, identify one or more sentences corresponding to the identified second average values. Execute training of a machine learning model based on the one or more sentences. An information processing apparatus characterized by having a control unit.

Citation Information

Patent Citations

  • Session response scheme determination method and device, equipment and medium

    CN111753062A

  • Information processor and processing method, and program

    JP2008226104A

  • Method to select learning text for language model, method to learn language model by using the same learning text, and computer and computer program for executing the methods

    JP2016024759A

  • Natural language analysis device, method and program

    JP2016162308A

  • Passage type questioning and answering device, method, and program

    JP2018124914A