Information processing device, information processing method, and program

A dual-dataset training approach for a general-purpose language model improves name matching accuracy and versatility by concurrently learning name matching and title inference tasks.

WO2025141700A1PCT designated stage expired Publication Date: 2025-07-03FAST ACCOUNTING INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/046681
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing machine learning models for name matching require task-specific training for each target, leading to increased training costs and reduced accuracy when applied to different tasks.

Method used

A general-purpose language model is fine-tuned using two datasets: one for name matching and another for title inference, allowing simultaneous learning of both tasks to improve determination accuracy without task specialization.

Benefits of technology

Enhances the accuracy of name matching while maintaining versatility across different tasks by leveraging a single model trained on diverse datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023046681_03072025_PF_FP_ABST
    Figure JP2023046681_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention is an information processing device 1 comprising: a storage unit 12 which stores a trained model M that, upon input of inference data which is a data set for inference and in which the name of a first merge target, the name of a second merge target, an explanation of the first merge target, and an explanation of the second merge target are associated with each other, outputs information indicating whether the first merge target and the second merge target match and that has learned a merge task which is for outputting whether the first merge target and the second merge target match and a title inference task which is for outputting the name of an estimation target using an explanation of the estimation target as input; an acquisition unit 131 which acquires the inference data; a determination unit 132 which inputs the acquired inference data into the trained model M and which determines whether the first merge target in the inference data and the second merge target in the inference data match; and an output unit 133 which outputs the result of determination by the determination unit 132.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present invention relates to an information processing device, an information processing method, and a program.

[0002] A technique for determining name matching using a machine learning model that is fine-tuned from a pre-trained language model is known (for example, Non-Patent Document 1).

[0003] Yuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan, and Wang-Chiew Tan. 2020. Deep entity matching with pre-trained language models. Proceedings of the VLDB Endowment, 14(1):50-60. Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958.

[0004] With existing methods, in order to improve accuracy, it is necessary to train a machine learning model for each task or target of name matching. This means that re-training is required when performing name matching for other tasks or target of name matching, resulting in high training costs.

[0005] Therefore, the present invention has been made in consideration of these points, and aims to improve the judgment accuracy of the name identification task while maintaining versatility.

[0006] In a first aspect of the information processing device of the present invention, a trained model outputs information indicating whether or not the first name matching target in the input inference data matches the second name matching target in the inference data, when the trained model receives inference data that associates a name of a first name matching target, a name of a second name matching target, a description of the first name matching target, and a description of the second name matching target, and the first name matching target matches the second name matching target. The trained model outputs information indicating whether or not the first name matching target in the input inference data matches the second name matching target in the inference data, when the trained model receives inference data that associates a name of a first name matching target, a name of a second name matching target, a description of the first name matching target, and a description of the second name matching target, the first name matching target matches the second name matching target. The training model includes: (1) a name matching task that, when a description of a first name matching target and a description of a second name matching target are input, outputs information indicating whether the first name matching target and the second name matching target match; and (2) a title inference task that, based on a second learning dataset in which the name of an inference target and the description of the inference target are associated, inputs a description of the inference target and outputs the name of the inference target; a memory unit that stores the trained model that has learned the above; an acquisition unit that acquires the inference data; a determination unit that inputs the acquired inference data into the trained model and determines whether the first name matching target in the inference data and the second name matching target in the inference data match; and an output unit that outputs the result determined by the determination unit.

[0007] The description of the first matching target associated in the inference data and the first training dataset may include information indicating details of the first matching target, which is a product, and the price of the first matching target; the description of the second matching target associated in the inference data and the first training dataset may include information indicating details of the second matching target, which is a product, and the price of the second matching target; and the description of the estimated target associated in the second training dataset may include information indicating details of the estimated target.

[0008] The trained model may have trained a general-purpose language model to perform the name matching task based on the first training dataset and the title inference task based on the second training dataset.

[0009] The acquisition unit may acquire the first training data set and the second training data set, and the information processing device may further include a learning unit that generates the trained model trained on (1) the name matching task based on the first training data set acquired by the acquisition unit and (2) the title inference task based on the second training data set, and stores the trained trained model in the memory unit.

[0010] The learning unit may perform parallel learning of the name matching task based on the first learning data set and the title inference task based on the second learning data set in a single learning process.

[0011] An information processing method according to a second aspect of the present invention includes: an acquisition unit that acquires inference data sets, the inference data being associated with a name of a first target, a name of a second target, a description of the first target, and a description of the second target; and a trained model stored in a storage unit, the trained model outputting, when the inference data is input, information indicating whether or not the first target in the inference data matches the second target in the inference data, the trained model including: (1) a first training data set associated with a name of the first target, a name of the second target, a description of the first target, a description of the second target, and a label indicating whether or not the first target matches the second target; The method includes: (1) a name matching task that, when a name of a first name matching target, a name of a second name matching target, a description of the first name matching target, and a description of the second name matching target are input based on a data set, outputs information indicating whether the first name matching target and the second name matching target match; and (2) a title inference task that, based on a second learning data set in which the name of an inference target and a description of the inference target are associated, outputs the name of the inference target using the description of the inference target as input. The method includes a step of inputting the inference data into the trained model that has trained the above-mentioned name matching task, and determining whether the first name matching target in the inference data matches the second name matching target in the inference data; and a step of outputting the result determined in the determining step.

[0012] In a third aspect of the program of the present invention, a computer is provided with an acquisition unit that acquires inference data sets, the inference data sets being associated with a name of a first target, a name of a second target, a description of the first target, and a description of the second target; and a trained model stored in a storage unit, the trained model outputting, when the inference data is input, information indicating whether or not the first target in the inference data matches the second target in the inference data, the trained model including: (1) a first training data set in which the name of the first target, the name of the second target, the description of the first target, the description of the second target, and a label indicating whether or not the first target matches the second target are associated with each other; and (2) a title inference task that, based on a second learning dataset in which the name of an inference target and the description of the inference target are associated, outputs the name of the inference target, taking as input the description of the inference target. The trained model is trained to execute the following steps: inputting the inference data into the trained model; determining whether the first name matching target in the inference data and the second name matching target in the inference data match; and outputting the result determined in the determining step.

[0013] According to the present invention, it is possible to improve the determination accuracy of a name identification task while maintaining versatility.

[0014] FIG. 1 is a diagram for explaining an overview of an information processing system S. FIG. 2 is a diagram showing an example of a learning dataset. FIG. 3 is a block diagram showing the configuration of an information processing device 1. FIG. 4 is a flowchart for explaining a learning flow in a learning unit 134. FIG. 5 is a diagram showing an example of a prompt acquired by an acquisition unit 131. FIG. 6 is a flowchart showing the flow of processing in the information processing device 1.

[0015] [Outline of Information Processing System S] Fig. 1 is a diagram for explaining an outline of the information processing system S. Fig. 1(a) shows the configuration of the information processing system S. The information processing system S is a system for performing name matching. Name matching is a task performed by a machine learning model, and is a task of determining whether multiple given targets match.

[0016] The objects of name matching performed by the information processing device system S are, for example, names of products or services, but are not limited to these. The information processing system S may also perform name matching on corporate names, personal names, or other names. The information processing system S has an information processing device 1 and an information terminal 2. The information processing device 1 and the information terminal 2 are connected to each other so as to be able to communicate with each other via a network.

[0017] The information processing device 1 is a device for performing name matching. The information processing device 1 is, for example, a server. The information processing device 1 trains a machine learning model, and when data to be matched is given, the information processing device 1 uses the machine learning model to determine whether the target in the given data matches.

[0018] The information terminal 2 is a terminal used by a user of the information processing system S. As an example, the information terminal 2 transmits a data set to be used for learning or inference to the information processing device 1, instructs the information processing device 1 to execute learning or inference, receives the inference results from the information processing device 1, and displays them on a display unit. Note that the information processing device 1 and the information terminal 2 may be configured as an integrated unit. That is, the information processing device 1 has an input / output interface, accepts operations from a user, and displays the inference results.

[0019] The processing in the information processing system S will be described with reference to FIG. 1( b). The information processing device 1 stores a pre-trained model M1. The pre-trained model M1 is a general-purpose language model, and is a trained model that has been trained to be able to execute natural language processing tasks based on a large amount of data set. The information processing device 1 trains the pre-trained model M1 on a name matching task and a title inference task, and generates a trained model M2.

[0020] The name matching task is a task in which the names of multiple objects to be matched and text describing each object are given, and the task determines whether the multiple objects match. The name of the object indicates the name of the product, natural person, corporation, etc. that is the object. The description of the object indicates the characteristics of the object. For example, if the object is a product, the description of the object includes the product's size, color, function, place of manufacture, manufacturer, seller, model number, operating environment, raw materials, selling points, price, etc.

[0021] If the target is a natural person, the description of the target includes information such as the date of birth, birthplace, alma mater, occupation, achievements, etc. If the target is a corporation, the description of the target includes information such as the corporation's address, number of employees, year of establishment, history, composition of officers, sales, etc.

[0022] Specifically, the information processing device 1 trains the pre-trained model M1 to perform a name matching task based on a first training dataset. An example of the first training dataset is shown in FIG. 2( a). In the first training dataset, the name of the first target, the name of the second target, a description of the first target, a description of the second target, and a label indicating whether the first target matches the second target are associated with each other.

[0023] The title inference task is a task in which, given text indicating a description of an object whose title is to be inferred, the title of the object indicated by the given text is generated. Specifically, the information processing device 1 trains the pre-trained model M1 on the title inference task based on a second training dataset. An example of the second training dataset is shown in FIG. 2(b). In the second training dataset, the name of the inference object and a description of the inference object are associated with each other.

[0024] It is particularly suitable to train the second training dataset based on a dataset containing named entities such as various products, model numbers, brands, place names, or corporate names as descriptions or titles. By training the title inference task, the trained model M2 can learn named entities used in descriptions that may affect the title. This enables the model to recognize important expressions in the text that affect the name matching results. As a result, it is expected that the accuracy of the name matching task will be improved without compromising the versatility of the model.

[0025] The trained model M2 is trained to output a determination result D2 corresponding to the input inference data D1 when the inference data D1 is input. The inference data D1 associates the name of a first name matching target, the name of a second name matching target, a description of the first name matching target, and a description of the second name matching target. The determination result D2 indicates whether the first name matching target in the inference data matches the second name matching target in the inference data.

[0026] The information processing device 1 inputs the inference data D1 into the trained model M2 and outputs the judgment result D2.

[0027] By configuring the information processing system S in this manner, it is possible to achieve the effect of improving the judgment accuracy of the name matching task while maintaining versatility.

[0028] 3 is a block diagram showing the configuration of the information processing device 1. The information processing device 1 has a communication unit 11, a storage unit 12, and a control unit 13. The control unit 13 has an acquisition unit 131, a determination unit 132, an output unit 133, and a learning unit 134.

[0029] The communication unit 11 is a communication interface for transmitting and receiving data to and from other devices via a network. The storage unit 12 is a storage medium including a read-only memory (ROM), a random access memory (RAM), a solid-state drive (SSD), a hard disk drive, etc. The storage unit 12 pre-stores a program to be executed by the control unit 13. The storage unit 12 stores a pre-trained model M1 and a trained model M2.

[0030] The control unit 13 is a processor such as a CPU (Central Processing Unit), and functions as an acquisition unit 131, a determination unit 132, an output unit 133, and a learning unit 134 by executing a program stored in the storage unit 12.

[0031] The acquisition unit 131 acquires the inference data D1. For example, the acquisition unit 131 acquires the inference data D1 from the information terminal 2. The acquisition unit 131 may acquire the inference data D1 from the storage unit 12 or from an external device (not shown). The acquisition unit 131 may acquire the first training data set and the second training data set, and output them to the learning unit 134.

[0032] The determination unit 132 inputs the acquired inference data into the trained model M2 and determines whether the first name matching target in the inference data matches the second name matching target in the inference data. The output unit 133 outputs the determination result D2 determined by the determination unit 132. As an example, the output unit 133 displays the determination result D2 on the display unit of the information terminal 2.

[0033] When the object of name matching performed by the information processing device 1 is a product, the name matching may be performed based on a data set including the price of the product.

[0034] A description of a first object associated in the inference data and the first training dataset includes information indicating details of the first object, which is a commodity, and the price of the first object. A description of a second object associated in the inference data and the first training dataset includes information indicating details of the second object, which is a commodity, and the price of the second object.

[0035] Note that even when a dataset including product prices is used in the name matching task, the product prices do not need to be used in learning in the title inference task. That is, the description of the associated estimation target in the second learning dataset includes information indicating details of the estimation target. This is because the product prices have a large impact on the results of the name matching task, while the product prices have a relatively small impact on the results of the title inference task.

[0036] In this way, by performing name matching based on information including the price of the product to be matched, the accuracy of name matching can be improved.

[0037] The learning unit 134 trains the pre-trained model M1 based on the first training data set and the second training data set acquired by the acquisition unit 131, generates a trained model M2 by updating parameters of the pre-trained model M1, and stores the generated trained model M2 in the storage unit 12. Note that the learning unit 134 may additionally train the trained model M2 on a name matching task or a title inference task.

[0038] Specifically, the training unit 134 trains the pre-trained model M1 based on the first training dataset. Specifically, the training unit 134 inputs the name of the first matching target, the name of the second matching target, a description of the first matching target, and a description of the second matching target into the pre-trained model M1, and outputs a determination result indicating whether the first matching target and the second matching target match. The training unit 134 calculates a loss based on the determination result output by the pre-trained model M1 and the labels included in the first training dataset, updates the parameters of the pre-trained model M1 based on the calculated loss, and trains the pre-trained model M1.

[0039] The learning unit 134 inputs a description of the estimation target included in the second training dataset to the pre-trained model M1, and outputs a name of the estimation target corresponding to the input description. The learning unit 134 calculates a loss based on the name of the estimation target output by the pre-trained model M1 and the name of the estimation target as training data included in the second training dataset, updates the parameters of the pre-trained model M1 based on the calculated loss, and trains the pre-trained model M1.

[0040] The learning of the name matching task and the learning of the title inference task may be performed simultaneously. The learning unit 134 may perform parallel learning of the name matching task based on the first learning data set and the title inference task based on the second learning data set in a single learning process. Fig. 4 is a flowchart for explaining the learning flow in this case. The flowchart shown in Fig. 4 starts when the information processing device 1 receives an instruction to start learning from the information terminal 2.

[0041] The acquisition unit 131 acquires a first training data set (S01). The acquisition unit 131 acquires a second training data set (S02). The learning unit 134 determines a termination condition (S03). The termination condition may be, for example, that learning has been performed a predetermined number of times.

[0042] If the termination condition is not satisfied (NO in S03), the learning unit 134 causes the pre-trained model M1 to execute a name matching task based on the first training data set and output the results (S04).The learning unit 134 causes the pre-trained model M1 to execute a title inference task based on the second training data set and output the results (S05).

[0043] The learning unit 134 calculates a loss based on the result output by the pre-trained model M1 in the name matching task and the associated label in the first training dataset (S06).The learning unit 134 also calculates a loss based on the result output by the pre-trained model M1 in the title inference task and the name of the estimation target as training data associated in the second training dataset (S06).

[0044] The learning unit 134 updates the parameters of the pre-trained model M1 based on the calculated loss (S07). As an example, the learning unit 134 calculates the gradient in the title inference task and the gradient in the name matching task based on the calculated loss, and updates the parameters based on the average value of the gradient in the title inference task and the gradient in the name matching task.

[0045] The extent to which the parameters are updated per step may differ between the title inference task and the name identification task. That is, the parameter update amount may be calculated by multiplying each gradient by a different predetermined coefficient, or different learning rates may be set for the title inference task and the name identification task. The information processing device 1 proceeds to step S03.

[0046] If the termination condition is satisfied (YES in S03), the information processing device 1 stores the trained model M2, which is the pre-trained model M1 whose parameters have been updated, in the storage unit 12 (S08). Then, the information processing device 1 terminates the process.

[0047] The acquiring unit 131 may be configured to acquire inference data included in a prompt (command) written in a natural language. As an example, the acquiring unit 131 displays a screen for accepting a prompt on the information terminal 2 and acquires the prompt written in a natural language from the information terminal 2. FIG. 5 shows an example of a prompt acquired by the acquiring unit 131. As shown in FIG. 5 , the prompt includes an instruction (P1) on the content of a task to be executed, a first object to be identified (P2), and a second object to be identified (P3). The first object to be identified (P2) and the second object to be identified (P3) each include a name (P21, P31) and a description (P22, P32).

[0048] In this case, the trained model M2 is trained to use a prompt as input, identify the task content to be executed contained in the prompt, and if the content of the identified task is a name matching task, execute the name matching task based on the inference data contained in the prompt.

[0049] 6 is a flowchart showing the flow of processing in the information processing device 1. The flowchart shown in FIG. 6 starts from the point in time when an instruction to perform inference is received from the information terminal 2.

[0050] The acquisition unit 131 acquires inference data (S11). The acquisition unit 131 inputs the inference data to the trained model M2 and determines whether the first name matching target in the inference data matches the second name matching target in the inference data (S12). The output unit 133 outputs the determination result output by the trained model M2 (S13). As an example, the output unit 133 displays the determination result output by the trained model M2 on the information terminal 2. Then, the information processing device 1 ends the processing.

[0051] [Effects of this embodiment] As described above, the information processing device 1 can improve the accuracy of judgment in the name matching task while maintaining versatility, by having the information processing device 1 learn the title inference task and the name matching task, without specializing in a specific task.

[0052] The present invention has been described above using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by functionally or physically distributing or integrating in any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments.

[0053] REFERENCE SIGNS LIST 1 Information processing device 2 Information terminal 11 Communication unit 12 Storage unit 13 Control unit 131 Acquisition unit 132 Determination unit 133 Output unit 134 Learning unit

Claims

1. A data set for inference, which, when inputting inference data in which the name of the first naming target, the name of the second naming target, the description of the first naming target, and the description of the second naming target are associated, outputs information indicating whether the first naming target in the input inference data matches the second naming target in the inference data. A learned model, (1) Based on a first training data set in which the name of the first naming target, the name of the second naming target, the description of the first naming target, the description of the second naming target, and a label indicating whether the first naming target and the second naming target match are associated, when inputting the name of the first naming target, the name of the second naming target, the description of the first naming target, and the description of the second naming target, it is a task of a naming task that outputs information indicating whether the first naming target and the second naming target match, (2) Based on a second training data set in which the name of the estimation target and the description of the estimation target are associated, a title inference task that outputs the name of the estimation target with the description of the estimation target as input, A storage unit that stores the learned model that has learned the above, An acquisition unit that acquires the inference data, An input the acquired inference data into the learned model, and a determination unit that determines whether the first naming target in the inference data matches the second naming target in the inference data, An output unit that outputs the result determined by the determination unit, An information processing apparatus having the above.

2. In the description of the first naming target associated in the inference data and the first training data set, it includes information indicating the details of the first naming target that is a product and the amount of the first naming target, In the description of the second naming target associated in the inference data and the first training data set, it includes information indicating the details of the second naming target that is a product and the amount of the second naming target, In the description of the estimation target associated in the second training data set, it includes information indicating the details of the estimation target, The information processing apparatus according to claim 1.

3. The learned model according to claim 1 is a learned model that has learned the clustering task based on the first learning dataset for a general-purpose language model and the title inference task based on the second learning dataset. The information processing apparatus according to claim 1.

4. The acquisition unit acquires the first learning dataset and the second learning dataset. The information processing apparatus further includes a learning unit that generates the learned model that has learned: (1) the clustering task based on the first learning dataset acquired by the acquisition unit; and (2) the title inference task based on the second learning dataset, and stores the learned model in the storage unit. The information processing apparatus according to claim 1.

5. The learning unit according to claim 4 learns the clustering task based on the first learning dataset and the title inference task based on the second learning dataset in parallel in a single learning process. The information processing apparatus according to claim 4.

6. An inference dataset executed by a computer, comprising: an acquisition unit that acquires inference data in which the name of a first naming target, the name of a second naming target, the description of the first naming target, and the description of the second naming target are associated; a learned model stored in a storage unit that outputs information indicating whether the first naming target in the inference data matches the second naming target in the inference data when the inference data is input; (1) A naming task that, based on a first learning dataset in which the name of a first naming target, the name of a second naming target, the description of the first naming target, the description of the second naming target, and a label indicating whether the first naming target matches the second naming target are associated, outputs information indicating whether the first naming target matches the second naming target when the name of the first naming target, the name of the second naming target, the description of the first naming target, and the description of the second naming target are input; (2) A title inference task that outputs the name of an estimation target using the description of the estimation target as an input based on a second learning dataset in which the name of the estimation target and the description of the estimation target are associated; A step of inputting the inference data into the learned model that has learned the naming task and the title inference task, and determining whether the first naming target in the inference data matches the second naming target in the inference data; A step of outputting the result determined in the determining step; An information processing method having the above steps.

7. A computer is caused to perform: an acquisition unit that acquires inference data in which the name of a first naming target, the name of a second naming target, the description of the first naming target, and the description of the second naming target are associated; a learned model stored in a storage unit that, when the inference data is input, outputs information indicating whether or not the first naming target in the inference data matches the second naming target in the inference data, where the learned model is learned based on: (1) a first learning dataset in which the name of the first naming target, the name of the second naming target, the description of the first naming target, the description of the second naming target, and a label indicating whether or not the first naming target matches the second naming target are associated, and which, when the name of the first naming target, the name of the second naming target, the description of the first naming target, and the description of the second naming target are input, outputs information indicating whether or not the first naming target matches the second naming target, which is a naming task; and (2) a second learning dataset in which the name of an estimation target and the description of the estimation target are associated, and which, when the description of the estimation target is input, outputs the name of the estimation target, which is a title inference task; a step of inputting the inference data into the learned model and determining whether or not the first naming target in the inference data matches the second naming target in the inference data; and a step of outputting the result determined in the determining step.

Citation Information

Patent Citations

  • Learning program and learning method

    JP2019185244A

  • Information processing device, information processing method, and program

    WO2023132029A1

  • Information processing device, information processing method, and information processing program

    WO2023162206A1