Information Processing Apparatus, Information Processing Method, and Information Processing Program

The information processing apparatus addresses the challenges of record matching by converting record pairs into a format suitable for model input, allowing for similarity calculation without training data and handling heterogeneous data, thus enhancing the efficiency and accuracy of record matching.

JP7686923B2Active Publication Date: 2025-06-03NEC CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024502726
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-06-03
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

Existing record matching techniques require large amounts of training data and struggle with handling heterogeneous data, where records have different data formats, and unsupervised methods face challenges in aligning attributes across different datasets.

Method used

An information processing apparatus and method that acquires a record pair, converts it into a format suitable for input into a model, calculates the similarity of the converted record pair using the model, and outputs the calculated similarity, without requiring training data specific to the record pair and capable of handling heterogeneous data.

Benefits of technology

Enables efficient calculation of record pair similarity without needing training data and effectively handles heterogeneous data, improving the scalability and accuracy of record matching processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007686923000001
    Figure 0007686923000001
  • Figure 0007686923000002
    Figure 0007686923000002
  • Figure 0007686923000003
    Figure 0007686923000003
Patent Text Reader

Abstract

The present invention provides, as a technology for calculating the degree of similarity between a record pair, a technology that does not require training data pertaining to a record pair and that is capable of handling data of differing types. To this end, an information processing device (1) comprises: an acquisition unit (11) that acquires a record pair; a conversion unit (12) that converts the record pair so as to generate a converted record pair; a similarity degree calculation unit (13) that inputs the converted record pair into a model so as to calculate a similarity degree pertaining to the converted record pair; and an output unit (14) that outputs the similarity degree which has been calculated by the similarity degree calculation unit (13).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] Processing is performed to identify and associate combinations of identical or similar records from records stored in different datasets. Such processing is also called record matching. Record matching enables unified management of tables and data expansion. As a technique for performing record matching, there is a technique for performing matching by machine learning. For example, Patent Document 1 describes an apparatus that calculates the similarity of record pairs using a plurality of similarity functions for calculating the similarity of record pairs and learns the weights of the similarities by supervised machine learning using training data. Here, the training data is a dataset to which a label indicating whether combinations of records are identical is attached. Non-Patent Document 1 also describes a technique called DITTO that performs record matching by supervised machine learning. Non-Patent Document 2 also describes a technique called ZeroER that collates records by unsupervised machine learning without using training data.

[0003] In recent years, language models (for example, Non-Patent Documents 3 to 5) and image classification models (for example, Non-Patent Document 6) have been proposed as models generated by machine learning.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Non-Patent Documents

[0005]

Non-Patent Document 1

[0006] However, since supervised machine learning requires a large amount of training data, in the technologies described in Patent Document 1 and Non-Patent Document 1, the cost and time for collecting training data are significant, and there is also a problem that heterogeneous data cannot be handled. Here, heterogeneous data refers to a combination of records where the data formats are not the same. Also, in ZeroER, an unsupervised machine learning method described in Non-Patent Document 2, although training data is not required, there is a problem that it cannot be applied to heterogeneous data with different attributes because the attributes need to be aligned.

[0007] One aspect of the present invention has been made in view of the above problems, and an example of its object is to provide a technique for calculating the similarity of a record pair that does not require training data regarding the record pair and can also handle heterogeneous data.

Means for Solving the Problems

[0008] An information processing apparatus according to one aspect of the present invention includes an acquisition unit that acquires a record pair, a conversion unit that generates a converted record pair by converting the record pair, a similarity calculation unit that calculates a similarity regarding the converted record pair by inputting the converted record pair into a model, and an output unit that outputs the similarity calculated by the similarity calculation unit.

[0009] An information processing method according to one aspect of the present invention includes at least one processor acquiring a record pair, generating a converted record pair by converting the record pair, calculating a similarity regarding the converted record pair by inputting the converted record pair into a model, and outputting the calculated similarity.

[0010] An information processing program according to one aspect of the present invention causes a computer to execute an acquisition process of acquiring a record pair, a conversion process of generating a converted record pair by converting the record pair, a similarity calculation process of calculating a similarity regarding the converted record pair by inputting the converted record pair into a model, and an output process of outputting the similarity calculated in the similarity calculation process.

Effect of the Invention

[0011] According to one aspect of the present invention, as a technique for calculating the similarity of a record pair, it is possible to provide a technique that does not require training data regarding the record pair and can also handle heterogeneous data.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Embodiments for Carrying Out the Invention

[0013] 〔First Exemplary Embodiment〕 The first exemplary embodiment of the present invention will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for the exemplary embodiments described later.

[0014] <Configuration of Information Processing Apparatus 1> The configuration of the information processing apparatus 1 according to the present exemplary embodiment will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the information processing apparatus 1. The information processing apparatus 1 is an apparatus that calculates the similarity between records. Here, a record is a unit of data for which similarity is calculated. Examples of the data included in a record include structured data such as table data, semi-structured data described in a data description language such as JSON (JavaScript Object Notation: registered trademark) or XML (Extensible Markup Language), and unstructured data representing a document written in a natural language. As an example, a record is a row of a table and includes a set of one or more attribute names and attribute values corresponding to the columns of the table. Also, a record may be graph data.

[0015] The information processing apparatus 1 includes an acquisition unit 11, a conversion unit 12, a similarity calculation unit 13, and an output unit 14.

[0016] (Acquisition Unit 11) The acquisition unit 11 acquires a pair of records. A pair of records is a set of records, and as an example, it is a set of a record included in a first table and a record included in a second table. The first table and the second table are, as an example, tables storing a business operator's customer information or tables storing product information. However, the first table and the second table are not limited to the above-described examples and may be other tables. Also, the first table and the second table may be the same or different.

[0017] The plurality of records included in the pair of records may each have a different data format. More specifically, for example, when a record is a row of a table, some of the attribute names included in the record may be different, or all of the attribute names included in the record may be different.

[0018] The acquisition unit 11 may acquire a record pair by reading the record pair from a storage device, or may acquire the record pair by receiving the record pair from another device connected via a communication interface. Further, the acquisition unit 11 may acquire the record pair input from an input device via an input / output interface.

[0019] (Conversion unit 12) The conversion unit 12 generates a converted record pair by converting the above-mentioned record pair. As an example, the conversion unit 12 converts the records included in the record pair into data representing a document, an image, audio, or a graph. More specifically, as an example, the conversion unit 12 converts a record into an affirmative sentence or an interrogative sentence. However, the method by which the conversion unit 12 converts the record pair is not limited to the above-mentioned examples, and the conversion unit 12 may convert the record pair by other methods.

[0020] (Similarity calculation unit 13) The similarity calculation unit 13 calculates the similarity regarding the converted record pair by inputting the converted record pair into a model. Here, the model is a model for calculating similarity, and as an example, it is a model that is generally publicly available and can be used by any user. The model may be a model generated by machine learning or a rule-based model created by a human. Specifically, as an example, the model is a document classification model, an image classification model, an audio classification model, or a graph classification model. The document classification model is a model for classifying document data. The image classification model is a model for classifying image data. The audio classification model is a model for classifying audio data. The graph classification model is a model for classifying graph data.

[0021] Examples of document classification models include, for example, document embedding models, entailment recognition models, paraphrase prediction models, question answering models, and masked language models. A document embedding model is a model that embeds a document or a word into a vector space. An entailment recognition model is a model that predicts the entailment relationship between multiple documents. A paraphrase prediction model is a model that predicts whether two documents are paraphrased expressions. A question answering model is a model that extracts and outputs an answer from a given document for a question. A masked language model is a model for predicting words that match the masks in a document.

[0022] Examples of image classification models include, for example, image embedding models. An image embedding model is a model that embeds image data into a vector space. Examples of speech classification models include, for example, speech embedding models. A speech embedding model is a model that embeds speech data into a vector space.

[0023] The input to the above models includes, for example, at least any one of text data, image data, speech data, graphs, and vectors. The output of the above models includes, for example, a vector or a score indicating a confidence level. The score is, for example, a score indicating the confidence level regarding the inclusion relationship of a document or a score indicating the confidence level of whether it is a paraphrased expression. However, the input and output of the model are not limited to the above examples and may include other information.

[0024] When the above model is generated by machine learning, examples of the model include, but are not limited to, the language models described in Non-Patent Documents 3 to 5, the image classification model described in Non-Patent Document 6, or a speech classification model that classifies speech data. Also, the above model may be stored in the memory of the information processing device 1 or may be stored in another device that can communicate with the information processing device 1.

[0025] Similarity is information regarding the degree of similarity between records included in a record pair, and as an example, it is the cosine similarity of a vector pair. Also, the similarity may be a value calculated from a score output by the model.

[0026] (Output unit 14) The output unit 14 outputs the similarity calculated by the similarity calculation unit 13. As an example, the output unit 14 may output the similarity by writing it to a storage device, or may output the similarity by transmitting the similarity to another device via a communication interface. Also, the output unit 14 may output the similarity to an output device (not shown) connected via an input / output interface. The output device is, as an example, a display, a printer, a projector, or a speaker.

[0027] The similarity output by the output unit 14 is used, for example, in table integration processing or information search processing. In the case of table integration processing, multiple tables can be integrated and data can be managed centrally by associating records predicted to be the same based on the similarity calculated by the similarity calculation unit 13. Also, in information search, the similarity calculation unit 13 may calculate the similarity for a record pair between a record serving as a search key (for example, a record specified by a user) and any other record registered in a predetermined table. In this case, the information processing apparatus 1 may output, as search results, the records included in the record pair predicted to be the same based on the similarity calculated by the similarity calculation unit 13. Thereby, search processing using the search key becomes possible even in a table not associated with the record serving as the search key.

[0028] <Effect of the information processing apparatus 1> As described above, in the information processing apparatus 1 according to the present exemplary embodiment, there are provided an acquisition unit 11 that acquires a record pair, a conversion unit 12 that generates a converted record pair by converting the record pair, a similarity calculation unit 13 that calculates a similarity regarding the converted record pair by inputting the converted record pair into a model, and an output unit 14 that outputs the similarity calculated by the similarity calculation unit 13. Therefore, according to the information processing apparatus 1 according to the present exemplary embodiment, as a technique for calculating the similarity of a record pair, an effect can be obtained that it is possible to provide a technique that does not require training data regarding the record pair and can also handle heterogeneous data.

[0029] <Information processing program> The functions of the above-described information processing apparatus 1 can also be realized by a program. The information processing program according to the present exemplary embodiment causes a computer to execute an acquisition process of acquiring a record pair, a conversion process of generating a converted record pair by converting the record pair, a similarity calculation process of calculating a similarity regarding the converted record pair by inputting the converted record pair into a model, and an output process of outputting the similarity calculated in the similarity calculation process.

[0030] <Flow of information processing method S1> The flow of the information processing method S1 according to the present exemplary embodiment will be described with reference to FIG. 2. FIG. 2 is a flowchart showing the flow of the information processing method S1. The execution subject of each step in the information processing method S1 may be a processor included in the information processing apparatus 1, may be a processor included in another apparatus, or may be processors provided in different apparatuses for each step.

[0031] In step S11, at least one processor acquires a record pair. In step S12, at least one processor generates a converted record pair by converting the record pair. In step S13, at least one processor calculates a similarity regarding the converted record pair by inputting the converted record pair into a model. In step S14, at least one processor outputs the calculated similarity.

[0032] <Effect of Information Processing Method S1> As described above, in the information processing method S1 according to this exemplary embodiment, a configuration is adopted in which at least one processor acquires a record pair, generates a converted record pair by converting the record pair, calculates a similarity regarding the converted record pair by inputting the converted record pair into a model, and outputs the calculated similarity. Therefore, according to the information processing method S1 according to this exemplary embodiment, as a technique for calculating the similarity of record pairs, an effect can be obtained in that a technique that does not require training data regarding the record pairs and can also handle heterogeneous data can be provided.

[0033] 〔Exemplary Embodiment 2〕 The second exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in Exemplary Embodiment 1 are denoted by the same reference numerals, and the description thereof will not be repeated.

[0034] <Outline of Information Processing Apparatus 1A> FIG. 3 is a block diagram showing the configuration of an information processing apparatus 1A according to this exemplary embodiment. The information processing apparatus 1A has a function of determining the identity between records. The data containing records is, as an example, structured data such as table data, semi-structured data described in a data description language such as JSON or XML, or unstructured data representing a document written in a natural language.

[0035] FIG. 4 is a diagram showing a specific example of data including records. In FIG. 4, data D1 is a table. In this case, a record is each row of the table. Also, in FIG. 4, data D2 is semi-structured data described in a data description language such as a markup language. In this case, as an example, a record is a web page. Data D3 is unstructured data representing a document written in a natural language. In this case, as an example, a record is a file generated in a predetermined file format.

[0036] Here, an overview of the processing performed by the information processing apparatus 1A according to this exemplary embodiment will be described with reference to FIG. 5. FIG. 5 is a diagram showing an overview of the processing flow performed by the information processing apparatus 1A. The information processing apparatus 1A roughly performs (i) a record pair generation process, (ii) a similarity calculation process, and (iii) an identity determination process.

[0037] (i) In the record pair generation process, the information processing apparatus 1A generates record pairs from first data x including a plurality of records e and second data x' including a plurality of records e'. As an example, the information processing apparatus 1A generates all combinations of the records e included in the first data x and the records e' included in the second data x'. Also, in generating record pairs, the information processing apparatus 1A may narrow down candidates for identity determination of the second data x' for the records e of the first data x by a technique called blocking.

[0038] (ii) In the similarity calculation process, the information processing apparatus 1A calculates the similarity between the records included in the record pairs. In this exemplary embodiment, the information processing apparatus 1A calculates the similarity by inputting a converted record pair obtained by converting the records into a model. Details of the similarity calculation process will be described later.

[0039] (iii) In the identity determination process, the information processing apparatus 1A determines the identity of the records included in the record pair based on the calculated similarity. As an example, when the similarity is equal to or greater than a threshold value, the information processing apparatus 1A determines that the records are identical. However, the method for determining identity is not limited to the method described above, and the information processing apparatus 1A may determine the identity of the records by other methods.

[0040] FIG. 6 is a diagram showing a specific example of the identity determination result. In FIG. 6, table TBL1 is an example of the first data x and includes a plurality of rows and a plurality of columns. Also, table TBL2 is an example of the second data x' and includes a plurality of rows and a plurality of columns. In table TBL1 and table TBL2, a record is a row of the table. Table TBL1 includes records l1, l2, l3, and l4, and table TBL2 includes records r1, r2, and r3.

[0041] In the example of FIG. 6, the information processing apparatus 1A determines that record l1 and record r2 are identical, record l2 and record r3 are identical, and record l3 and record r1 are identical by the above processes (i) to (iii).

[0042] <Configuration of Information Processing Apparatus 1A> As shown in FIG. 3, the information processing apparatus 1A includes a control unit 10A, a storage unit 20A, a communication unit 30A, and an input / output unit 40A.

[0043] (Communication Unit 30A) The communication unit 30A communicates with a device external to the information processing apparatus 1A via a communication line. The specific configuration of the communication line does not limit this exemplary embodiment, but as an example, the communication line is a wireless LAN (Local Area Network), a wired LAN, a WAN (Wide Area Network), a public switched telephone network, a mobile data communication network, or a combination thereof. The communication unit 30A transmits data supplied from the control unit 10A to other devices or supplies data received from other devices to the control unit 10A.

[0044] (Input / Output Unit 40A) Input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel are connected to the input / output unit 40A. The input / output unit 40A receives input of various information to the information processing apparatus 1A from the connected input devices. Also, the input / output unit 40A outputs various information to the connected output devices under the control of the control unit 10A. Examples of the input / output unit 40A include an interface such as USB (Universal Serial Bus).

[0045] (Control Unit 10A) As shown in FIG. 3, the control unit 10A includes an acquisition unit 11, a conversion unit 12, a similarity calculation unit 13, an output unit 14, an identity determination unit 15A, and an integration unit 16A.

[0046] (Acquisition Unit 11) In this exemplary embodiment, the acquisition unit 11 generates a record pair including the record e and the record e' from the first data x including the record e and the second data x' including the record e'. However, the acquisition unit 11 may not perform the process of generating the record pair. As an example, the acquisition unit 11 may acquire the record pair by reading it from the storage unit 20A or another external storage device, or may acquire the record pair received from another device via the communication unit 30A. Also, the acquisition unit 11 may acquire the record pair input from the input device connected to the input / output unit 40A.

[0047] (Conversion Unit 12) The conversion unit 12 generates a converted record pair by converting the record pair. As an example, the conversion unit 12 converts the records included in the record pair into document data, image data, audio data, or a graph. The conversion process executed by the conversion unit 12 will be described later.

[0048] (Similarity Calculation Unit 13) The similarity calculation unit 13 calculates the similarity s for the pair of converted records by inputting the pair of converted records into the model MA. The similarity s is information regarding the degree of similarity between the records included in the record pair, and as an example, it is a value calculated based on the cosine similarity of the vector pair or the score output by the model MA. Details of the process in which the similarity calculation unit 13 calculates the similarity s will be described later.

[0049] (Output unit 14) The output unit 14 outputs the similarity s calculated by the similarity calculation unit 13. As an example, the output unit 14 outputs by writing the similarity to the storage unit 20A. However, the method by which the output unit 14 outputs the similarity is not limited to the example described above, and the similarity s may be output by other methods. As an example, the output unit 14 may transmit the similarity to another device connected via the communication unit 30A, or may output the similarity to an output device connected via the input / output unit 40A.

[0050] (Identity determination unit 15A) The identity determination unit 15A determines the identity between the records included in the record pair based on the similarity s. As an example, the identity determination unit 15A determines that the records are identical when the similarity s is equal to or greater than a threshold value. Also, the identity determination unit 15A may determine the identity based on the rank when the record pairs are sorted in descending order of similarity, such as determining that the top x record pairs with the highest similarity are identical. Also, as an example, the identity determination unit 15A may apply a matching algorithm such as the stable marriage problem algorithm to determine the identity.

[0051] In addition, the method by which the identity determination unit 15A determines identity is not limited to the example described above, and the identity determination unit 15A may determine identity by other methods. As an example, the identity determination unit 15A may predict the identity of a record pair by inputting the record pair and the similarity into a prediction model generated by machine learning. In this case, the input to the prediction model includes, as an example, the record pair and the similarity. Also, the output of the prediction model includes, as an example, the prediction result of identity. In this case, the machine learning method of the prediction model is not limited, and as an example, a decision tree-based, linear regression, or neural network method may be used, or two or more of these methods may be used.

[0052] (Integration unit 16A) Based on the determination result of the identity determination unit 15A, the integration unit 16A integrates the first data x and the second data x'. As an example, the integration unit 16A integrates the first data x and the second data x' by increasing the number of records and / or increasing the number of data attributes. Examples of data integration performed by the integration unit 16A include (i) entity integration, (ii) data cleansing, and (iii) schema matching. (i) Entity integration refers to unifying the notations of different attributes and their values when the same set of records is given. (ii) Data cleansing refers to unifying differences in the description formats of company names, addresses, area codes, etc. (such as "(Co., Ltd.)" and "Limited Company"). (iii) Schema matching refers to taking an alignment (matching) of a plurality of attributes with different notations.

[0053] (Storage unit 20A) The storage unit 20A stores the first data x and the second data x', and also stores the similarity s calculated by the similarity calculation unit 13. The storage unit 20A also stores the model MA. Note that storing the model MA in the storage unit 20A means that the parameters defining the model MA are stored in the storage unit 20A.

[0054] (Model MA) Model MA is a model for calculating similarity and, as an example, is a model that is generally publicly available and can be used by any user. Model MA may be a model generated by machine learning or may also be a rule-based model created by humans.

[0055] As an example, model MA includes at least one of a document classification model, an image classification model, an audio classification model, and a graph classification model. Examples of the document classification model include, for example, a document embedding model, an entailment recognition model, a paraphrase prediction model, and a masked language model. Examples of the image classification model include, for example, an image embedding model that embeds image data into a vector space. Examples of the audio classification model include, for example, an audio embedding model that embeds audio data into a vector space.

[0056] The input of model MA includes, as an example, at least one of document data, image data, audio data, a graph, and a vector. The output of model MA includes, as an example, at least one of a vector and a score.

[0057] (Example 1 of document classification model: Document embedding model) The document embedding model is a model that embeds a document or a word into a vector space. The document embedding model is generated, as an example, by RoBERTa described in Non-Patent Document 3. The input of the document embedding model is, as an example, a document or a word (for example, the sentence "An old man is walking in the park."). The output is, as an example, a vector.

[0058] (Example 2 of document classification model: Entailment recognition model) The implicature recognition model is a model that predicts whether there is an implicature relationship such as "if Document 1, then Document 2". The implicature recognition model is generated by the technology described in Non-Patent Document 4 as an example. The input of the implicature recognition model is, for example, two documents. The two documents are, for example, Document 1: "An elderly man is walking in the park", and Document 2: "A man is in the park". In this case, Document 2 implicates Document 1. Figure 8 is a diagram schematically showing the implicature relationship between documents. In the example of Figure 8, Document 2 implicates Document 1.

[0059] Also, the output of the implicature recognition model is, for example, an implicature score. The implicature score is a numerical value indicating the confidence level of the implicature relationship, and is, for example, a real number between 0 and 1. The implicature score indicates that, for example, the higher the value, the higher the confidence level of the implicature relationship.

[0060] (Example 3 of Document Classification Model: Paraphrase Prediction Model) The paraphrase prediction model is a model that predicts whether two documents are paraphrases. The paraphrase prediction model is generated by RoBERTa described in Non-Patent Document 3 as an example. The input of the paraphrase prediction model is, for example, two documents. The two documents are, for example, Document 1: "NEC is an IT company.", and Document 2: "Nippon Electric Co., Ltd. conducts IT business.". The output of the model includes, for example, a paraphrase score. The paraphrase score is a score indicating the confidence level that two documents are paraphrases, and is, for example, a real number between 0 and 1. The paraphrase score indicates that, for example, the higher the value, the higher the confidence level that two documents are paraphrases.

[0061] (Example 4 of Document Classification Model: Masked Language Model) A masked language model is a model that predicts words that fit the masks in a document. The masked language model is, for example, a model generated by RoBERTa described in Non-Patent Document 3. The input of the document classification model is, for example, a document (e.g., a sentence such as "This pizza is very delicious. I [mask] this pizza."). The output of the model includes a word (e.g., "like") and a score. The score is a value indicating the confidence that the word fits the mask, and is, for example, a real number between 0 and 1.

[0062] (Image Classification Model) An image classification model is a model that classifies images. The image classification model is, for example, a model generated by the technique described in Non-Patent Document 6. The input of the image classification model is, for example, an image (e.g., an image of a dog). Also, the intermediate output of the model is, for example, a vector representation of the image, and the output of the model is, for example, a label (e.g., a label indicating "dog") and a score. The score is a value indicating the confidence of the label, and is, for example, a real number between 0 and 1. In this exemplary embodiment, the similarity calculation unit 13 calculates the similarity s using the vector representation of the image.

[0063] (Voice Classification Model) A voice classification model is a model that classifies voices. The input of the voice classification model is, for example, voice data (e.g., the barking of a dog). The intermediate output of the model is, for example, a vector representation of the voice, and the output of the model is, for example, a label (e.g., a label indicating that the voice is the barking of a dog) and a score. The score is a value indicating the confidence of the label, and is, for example, a real number between 0 and 1. In this exemplary embodiment, the similarity calculation unit 13 calculates the similarity s using the vector representation of the voice, which is the intermediate output.

[0064] (Graph Classification Model) The graph classification model is a model for classifying graphs. The input of the graph classification model is, for example, graph data (such as a graph representing facial features). The intermediate output of the model is, for example, the vector representation of the graph, and the output of the model is, for example, a label (such as a label indicating a person) and a score. The score is a value indicating the confidence level of the label and is, for example, a real number between 0 and 1. In this exemplary embodiment, the similarity calculation unit 13 calculates the similarity s using the vector representation of the graph, which is the intermediate output.

[0065] <Flow of the information processing method S100A> FIG. 7 is a flowchart showing the flow of the information processing method S100A executed by the information processing apparatus 1A. Note that, among the steps included in the information processing method S100A, some steps may be executed in parallel or in a different order. Also, the description of the content already explained will not be repeated.

[0066] (Step S101) In step S101, the acquisition unit 11 reads the model MA. In this exemplary embodiment, the model MA is selected from a plurality of model candidates. The plurality of model candidates include, for example, at least any one of a document classification model such as a document embedding model and an entailment recognition model, an image classification model, and an audio classification model. The selection of the model MA may be performed, for example, based on a user operation, or may be performed according to a predetermined algorithm. The model MA may be a single model or a set of a plurality of models.

[0067] (Step S102) In step S102, the acquisition unit 11 reads a record pair. The record e and the record e' included in the record pair are, for example, e = ((a_j, v_j))_{j = 1,…,d}, e' = ((a'_j, v'_j))_{j = 1,…,d'} It is represented as follows. Here, the attribute name a_j ∈ A_j, and A_j is, for example, a string space. The attribute value v_j ∈ V_j, and V_j is, for example, a string space or a real number space. In this example, the record l1 in FIG. 6 is l1 = ((title, sims 2 glamour life stuff pack), (manufacturer, aspyr media), (price, 24.99)), and the record r2 is r2 = ((title, aspyr media inc sims 2 glamour life stuff pack), (price, NaN)) is as follows.

[0068] In this exemplary embodiment, the acquisition unit 11 generates a record pair from the first data x and the second data x'. The acquisition unit 11, for example, generates all combinations of the record e included in the first data x and the record e' included in the second data x'. Further, in generating the record pair, the acquisition unit 11 may narrow down candidates for identity determination of the second data x' for the record e of the first data x by a technique called blocking.

[0069] (Step S103) In step S103, the conversion unit 12 generates a converted record pair by converting the record pair (e, e') into a format corresponding to the input of the model MA. As an example, the model MA includes a document classification model, and the conversion process by the conversion unit 12 includes a process of generating a converted record pair by converting the record pair (e, e') into a document. Also, as an example, the model MA includes an image classification model, and the conversion process by the conversion unit 12 includes a process of generating a converted record pair by converting the record pair (e, e') into an image. Also, as an example, the model MA includes an audio classification model, and the conversion process by the conversion unit 12 includes a process of generating a converted record pair by converting the record pair into audio. Also, as an example, the model MA includes a graph classification model, and the conversion process by the conversion unit 12 includes a process of generating a converted record pair by converting the record pair into a graph.

[0070] (Step S104) In step S104, the similarity calculation unit 13 calculates the similarity s regarding the converted record pair by inputting the converted record pair into the model MA.

[0071] (Processing examples 1 to 5 of steps S101 to S104) Here, as processing examples of steps S101 to S104, processing examples 1 to 5 will be described. Processing example 1 is a processing example when using a document embedding model. Processing example 2 is a processing example when using an image classification model. Processing example 3 is a processing example when using an audio classification model. Processing example 4 is a processing example when using an implication recognition model. Processing example 5 is a processing example when using a paraphrase prediction model.

[0072] (Processing example 1: Document embedding model) In this example, in step S101, the acquisition unit 11 reads a document embedding model. Also, in step S103, the conversion unit 12 converts the record pair (e, e') into a document pair (t, t'). The conversion unit 12, as an example, e = ((title, sims 2 glamour life stuff pack), (manufacturer, aspyr media), (price, 24.99)) e´ = ((title, aspyr media inc sims 2 glamour life stuff pack), (price, NaN)) A record pair (e, e´) including the record e and the record e´ is t = “Title is sims 2 glamour life stuff pack. Manufacturer is aspyr media. Price is 24.99.” t´ = “Title is aspyr media inc sims 2 glamour life stuff pack.” It is converted into a document pair (t, t´) including the document t and the document t´.

[0073] Also, in step S104, the similarity calculation unit 13 converts the document pair (t, t´) into a vector pair (v, v´) by means of a document embedding model. Here, v = M(t) and v´ = M(t´). Also, the similarity calculation unit 13 calculates a similarity s from the vector pair (v, v´). The similarity s is, as an example, s = exp(-||v - v´|| / c), where c > 0. Also, the similarity s may be the cosine similarity s = v^Tv´ / (||v|| ||v´||). Here, ^T is a symbol representing transposition.

[0074] (Processing Example 2: Image Embedding Model) In this example, in step S101, the acquisition unit 11 reads an image embedding model. In step S103, the conversion unit 12 converts the record pair (e, e´) into an image pair (i, i´).

[0075] FIG. 9 is a diagram showing an example of the image converted by the conversion unit 12. The conversion unit 12, as an example, e = ((title, sims 2 glamour life stuff pack), (manufacturer, aspyr media), (price, 24.99)) e' = ((title, aspyr media inc sims 2 glamour life stuff pack), (price, NaN)) Convert the record pair (e, e') including the record e and the record e' into the images i, i' shown in FIG. 9.

[0076] Also, in step S104, the similarity calculation unit 13 converts the image pair (i, i') into a vector pair (v, v') by using the image embedding model. Here, v = M(i) and v' = M(i'). Further, the similarity calculation unit 13 calculates the similarity s from the vector pair (v, v'). The similarity s is, as an example, = exp(-||v - v'|| / c), where c > 0.

[0077] When using the image embedding model, the conversion unit 12 may convert one record into one image, or may perform image conversion for each element (for example, word) included in the record. When the conversion unit 12 performs image conversion for each element, the similarity calculation unit 13 calculates the similarity s by using the set of images for each element. Also, when performing image conversion for each element, the conversion unit 12 may not perform image conversion for the missing values of the record.

[0078] (Processing Example 3: Voice Embedding Model) In this example, in step S101, the acquisition unit 11 reads the voice embedding model. Also, in step S103, the conversion unit 12 converts the record pair (e, e') into a voice pair (i, i'). The conversion unit 12, as an example, e = ((title, sims 2 glamour life stuff pack), (manufacturer, aspyr media), (price, 24.99)) e´ = ((title, aspyr media inc sims 2 glamour life stuff pack),(price, NaN)) Convert the record pair (e, e´) including the record e and the record e´ into voice data i representing the voice of reading the record e and voice data i´ representing the voice of reading the record e´.

[0079] Also, in step S104, the similarity calculation unit 13 converts the voice data pair (i, i´) into a vector pair (v, v´) by the voice embedding model. Here, v = M(i) and v´ = M(i´). Further, the similarity calculation unit 13 calculates the similarity s from the vector pair (v, v´). The similarity s is, as an example, = exp(-||v - v´|| / c), where c > 0.

[0080] (Processing Example 4: Implication Recognition Model) In this example, in step S101, the acquisition unit 11 loads the implication recognition model. Also, in step S103, the conversion unit 12 converts the record pair (e, e´) into a document pair (t, t´).

[0081] The conversion unit 12, as an example, Converts the record e = ((a_j, v_j))_{j = 1,…,d} into The document t = "a_1 is v_1.... a_d is v_d." However, the conversion unit 12 does not include the missing value in the document t.

[0082] The conversion unit 12, as an example, e = ((title, sims 2 glamour life stuff pack),(manufacturer, aspyr media),(price, 24.99)) e´ = ((title, aspyr media inc sims 2 glamour life stuff pack),(price, NaN)) The record pair (e, e´) including the record e and the record e´ where t = "Title is sims 2 glamour life stuff pack. Manufacturer is aspyr media. Price is 24.99." t´ = "Title is aspyr media inc sims 2 glamour life stuff pack." Convert to a document pair (t, t´) including the document t and the document t´.

[0083] Also, in step S104, the similarity calculation unit 13 calculates the implicature score of the document pair (t, t´) using the implicature recognition model. Further, the similarity calculation unit 13 calculates the similarity s using the implicature score. As an example, the similarity s is s = M(t, t´) × M(t´, t). In other words, the similarity s is the product value of the implicature score M(t, t´) of the implicature relationship "if it is document t, then it is document t´" and the implicature score M(t´, t) of the implicature relationship "if it is document t´, then it is document t". However, the similarity s is not limited to the above example and may be other values. The similarity s may be, for example, the maximum value of the implicature score M(t, t´) and the implicature score M(t´, t), or the sum of the implicature score M(t, t´) and the implicature score M(t´, t).

[0084] If the record e and the record e´ included in the record pair (e, e´) are the same, then e ⊂ e´ and e´ ⊂ e. In this processing example, the similarity calculation unit 13 uses this relationship to calculate the similarity.

[0085] (Processing Example 5: Paraphrase Prediction Model) In this example, in step S101, the acquisition unit 11 reads the paraphrase prediction model. Also, in step S103, the conversion unit 12 converts the record pair (e, e´) into a document pair (t, t´).

[0086] The conversion unit 12, as an example, The record e = ((a_j, v_j))_{j = 1,…,d} is Convert the document t = "v_1 … v_d". However, the conversion unit 12 does not include missing values in the document.

[0087] As an example, the conversion unit 12 e = ((title, sims 2 glamour life stuff pack),(manufacturer, aspyr media),(price, 24.99)) e´ = ((title, aspyr media inc sims 2 glamour life stuff pack),(price, NaN)) Convert the record pair (e, e´) including the record e and the record e´ which are t = “sims 2 glamour life stuff pack aspyr media 24.99” t´ = “aspyr media inc sims 2 glamour life stuff pack” into the document pair (t, t´) including the document t and the document t´ which are

[0088] Also, in step S104, the similarity calculation unit 13 calculates the paraphrase score of the document pair (t, t´) using the paraphrase prediction model, and sets the calculated paraphrase score as the similarity s of the record pair. That is, in this processing example, the similarity calculation unit 13 calculates the similarity s by reducing the record pair to a form that asks whether it is a paraphrased expression.

[0089] (Steps S105·S106) In step S105, the output unit 14 outputs the similarity s calculated by the similarity calculation unit 13. As an example, the output unit 14 outputs by writing the similarity s to the storage unit 20A. In step S106, the identity determination unit 15A determines the identity between the records included in the record pair based on the similarity s.

[0090] (Step S107) In step S106, the integration unit 16A generates integrated data from the first data x and the second data x' with reference to the determination result of the identity determination unit 15A. As an example, the integrated data includes a record obtained by integrating the records included in the record pair determined to be identical by the identity determination unit 15A.

[0091] <Effect of the information processing apparatus 1A> By the way, in recent years, high-precision models (inference models) trained with huge datasets have been publicly available and applied to various tasks (such as text classification, question answering, and image classification) in the fields of natural language processing and image processing. For example, DITTO, which is a clustering model of supervised machine learning (see Non-Patent Document 1), partially applies the pre-trained language model of BERT (Bidirectional Encoder Representations from Transformers). However, since the data formats vary depending on the task, it has not been possible to perform clustering only with existing inference models. In particular, it has not been possible to perform clustering on records containing unexpected attributes with existing inference models.

[0092] On the other hand, according to this exemplary embodiment, the information processing apparatus 1A calculates the similarity s regarding the record pair by converting the record pair into a format corresponding to the input of the model MA and inputting it to the model MA. By converting the records and reducing the clustering task to a task in the field of natural language processing or image processing, etc., it becomes possible to utilize inference models trained with a large amount of data. That is, according to this exemplary embodiment, it is possible to determine the identity of records having various attributes without requiring the training data of the model MA.

[0093] In the information processing apparatus 1A according to this exemplary embodiment, the model MA is selected from a plurality of model candidates, and the conversion unit 12 is configured to generate a converted record pair by converting the record pair (e, e') into a format corresponding to the input of the model MA. By converting the record pair by the conversion unit 12, the record pair becomes data in a format that can be input to the model MA. That is, regardless of what attributes the records to be calculated for similarity include, the similarity calculation unit 13 can calculate the similarity s by inputting the converted record pair to the model MA. Thus, according to the information processing apparatus 1A according to this exemplary embodiment, the effect that the similarity can be calculated for records having various attributes using the model MA without requiring the training data of the model MA is obtained.

[0094] In the information processing apparatus 1A according to this exemplary embodiment, the model MA includes a document classification model, and the conversion process by the conversion unit 12 includes a process of generating a converted record pair by converting the record pair into a document. For this reason, according to the information processing apparatus 1A according to this exemplary embodiment, the effect that the similarity of records having various attributes can be calculated using the model MA without training the model MA, which is a document classification model, is obtained.

[0095] Also, in the information processing apparatus 1A according to this exemplary embodiment, the model MA includes an image classification model, and the conversion process by the conversion unit 12 includes a process of generating a converted record pair by converting the record pair into an image. With the information processing apparatus 1A converting the record pair into an image, the similarity between records that are different as characters but have similar notations can be calculated more suitably. For example, in the case of a record including the word "glamour" and a record including the word "glqmour", although the character strings included in the records are different, since the shapes of the characters are similar, a high similarity is calculated. Thus, according to the information processing apparatus 1A according to this exemplary embodiment, in addition to the effect exhibited by the information processing apparatus 1 according to Exemplary Embodiment 1, an effect of being able to calculate a similarity that reflects the degree of similarity of the character shapes can be obtained.

[0096] Also, in the information processing apparatus 1A according to this exemplary embodiment, the model MA includes an audio classification model, and the conversion process by the conversion unit 12 includes a process of generating a converted record pair by converting the record pair into audio. With the information processing apparatus 1A converting the record pair into audio, the similarity between records that are different as characters but have similar phonemes can be calculated more suitably. For example, in the case of a record including the word "glamour" and a record including the word "glamar", although the character strings included in the records are different, since the pronunciations of the words are similar, a high similarity is calculated. Thus, according to the information processing apparatus 1A according to this exemplary embodiment, in addition to the effect exhibited by the information processing apparatus 1 according to Exemplary Embodiment 1, an effect of being able to calculate a similarity that reflects the degree of similarity of the sounds can be obtained.

[0097] In the information processing apparatus 1A according to this exemplary embodiment, the model MA includes a graph classification model, and the conversion process by the conversion unit 12 includes a process of generating a converted record pair by converting a record pair into a graph. By converting the record pair into a graph by the information processing apparatus 1A, the effect that the similarity of records having various attributes can be calculated using the model MA without training the model MA, which is a graph classification model, can be obtained.

[0098] 〔Exemplary Embodiment 3〕 The third exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the first and second exemplary embodiments are denoted by the same reference numerals, and the description thereof will not be repeated.

[0099] <Configuration of Information Processing Apparatus 1B> FIG. 10 is a block diagram showing the configuration of the information processing apparatus 1B according to this exemplary embodiment. The control unit 10B of the information processing apparatus 1B includes an acquisition unit 11B, a conversion unit 12B, a similarity calculation unit 13B, an output unit 14, an identity determination unit 15A, and an integration unit 16A. The storage unit 20B stores the model MB in addition to the first data x, the second data x', and the similarity s.

[0100] In addition to the record pair (e, e'), the acquisition unit 11B further acquires an auxiliary record. The auxiliary record is an auxiliary record used for calculating the similarity of the record pair (e, e'). As an example, the auxiliary record is a record included in the first data x and other than the record e included in the record pair (e, e'). Also, as an example, the auxiliary record is a record included in the second data x' and other than the record e' included in the record pair (e, e').

[0101] The conversion unit 12B converts the record pair acquired by the acquisition unit 11B to generate a converted record pair. Further, the conversion unit 12B generates a converted auxiliary record by converting the auxiliary record. As an example, the conversion unit 12B converts the auxiliary record into data representing a document, an image, audio, or a graph.

[0102] More specifically, as an example, the conversion unit 12B generates the above-mentioned converted record pair by converting one record e included in the record pair (e, e') into a question sentence and converting each of the other record e included in the record pair (e, e') and the auxiliary record into a response sentence.

[0103] The similarity calculation unit 13B calculates the similarity regarding the converted record pair by inputting the converted record pair and the converted auxiliary record into the model MB. The model MB includes, as an example, a question-and-answer model that takes a question sentence and a response sentence as inputs.

[0104] (Question-and-answer model) The question-and-answer model is a model that extracts and outputs an answer sentence from the documents given for the question sentence. The question-and-answer model is, as an example, a model generated by a technique called TANDA described in Non-Patent Document 5. The input of the question-and-answer model includes, as an example, a question sentence and a document. The question sentence is, as an example, a sentence such as "Where is the head office of NEC?" The document is, as an example, a document such as "Nippon Electric Co., Ltd. (NEC Corporation) is an electric appliance manufacturer of the Sumitomo Group with its head office located at Shiba 5-chome, Minato-ku, Tokyo. It is one of the constituent stocks of the Nikkei Stock Average."

[0105] The output of the model includes, as an example, an answer sentence and a score. The answer sentence is, as an example, a sentence such as "Shiba 5-chome, Minato-ku, Tokyo." The score is, as an example, a real number from 0 to 1. Here, when obtaining the output of the model, the score of each word may be calculated. For example, by the question-and-answer model, the score of "Nippon Electric Co., Ltd." is "0.1", the score of "Sumitomo Group" is "0.02", and the score of "Nikkei Stock Average" is "0.08" are calculated.

[0106] <Flow of Information Processing Method S100B> FIG. 11 is a flowchart showing the flow of information processing method S100B executed by information processing apparatus 1B. Note that some steps may be executed in parallel or in a different order. Also, descriptions of content already explained will not be repeated.

[0107] Information processing method S100B includes step S101B, step S102, step S102B, step S103B, step S104B, step S105, step S106, and step S107.

[0108] In step S101B, acquisition unit 11B reads model MB. In step S102B, acquisition unit 11B reads the auxiliary record.

[0109] In step S103B, conversion unit 12B generates a converted record pair by converting the record pair and generates a converted auxiliary record by converting the auxiliary record. In step S104B, similarity calculation unit 13B calculates the similarity regarding the converted record pair by inputting the converted record pair and the converted auxiliary record into model MB.

[0110] (Processing Example 6 of Steps S101 to S104B: Question-Answer Model) Here, as an example of the processing from step S101 to S104B, an example of the processing when using a question-and-answer model will be described. In this example, in step S101, the similarity calculation unit 13B reads the question-and-answer model. Also, in steps S102 and S102B, the acquisition unit 11B reads the record pair (e, e') and the auxiliary record R = {e_1,..., e_k} (k is a natural number). The auxiliary record R is, as an example, the set of all records included in the second data x'. However, the auxiliary record R is not limited to the above example and may be a set of other records. For example, the auxiliary record R may be a set of records selected from the second data x' by a random selection algorithm. Also, the auxiliary record R may be a blocked record set such as a record set obtained by extracting records including words common to the record e from the second data x'.

[0111] In this processing example, the auxiliary record R includes the record e' included in the record pair (e, e'). The auxiliary record is, as an example, the set of records r1 to r3 in the table TBL2 of FIG. 6, that is, R = {((title, adobe photoshop elements 4.0 photo-editing software for mac), (price, 85.95)), ((title, aspyr media inc sims 2 glamour life stuff pack), (price, NaN)), ((title, final-draft final draft av 2.5 screenwriting software mac / win screen writing software), (price, 199.95))} is as follows.

[0112] In this processing example, in step S103B, the conversion unit 12B converts the record e and the auxiliary record R (e' ∈ R). More specifically, the conversion unit 12B converts the record e into the question sentence q = T1(e). Here, the question sentence q is preferably of the so-called 5W1H open question type. Also, the conversion unit 12B converts the auxiliary record R into a document including a plurality of answer sentences.

[0113] As an example, the conversion unit 12B converts the record e = ((a_1, v_1),…,(a_d, v_d)) into T1(e) = “What is characterized as v_1 of a_1, … and v_d of a_d?’’ Also, the conversion unit 12B converts the auxiliary record R = {e_1,…,e_k} into the document T2(R) = “T3(e_1). T3(e_2). … T3(e_k).’’ Here, the answer sentence T3(e_j) (1 ≤ j ≤ k) included in the document T2(R) is T3(e_j) = ``{ID of e_j} is characterized as v_1 of a_1, …, and v_d of a_d’’ where {ID of e_j} is a unique ID assigned to the record e_j ∈ R. However, the conversion unit 12B does not include missing values in the document during conversion.

[0114] As an example, the conversion unit 12B e = ((title, sims 2 glamour life stuff pack),(manufacturer, aspyr media),(price, 24.99)) into q = “What is characterized as title of sims 2 glamour life stuff pack, manufacturer of aspyr media, and price of 24.99?” Convert it to. Also, auxiliary record R = {((title, adobe photoshop elements 4.0 photo - editing software for mac), (price, 85.95)), ((title, aspyr media inc sims 2 glamour life stuff pack), (price, NaN)), ((title, final - draft final draft av 2.5 screenwriting software mac / win screen writing software), (price, 199.95))} Convert it to document c = “r1 is characterized as title of adobe photoshop elements 4.0 photo - editing software for mac and price of 85.95. r2 is characterized as title of aspyr media inc sims 2 glamour life stuff pack. r3 is characterized as title of final - draft final draft av 2.5 screenwriting software mac / win screen writing software and price of 199.95.” Convert it to.

[0115] Also, in step S104B, the similarity calculation unit 13B inputs the question sentence q and the document c into the question-answering model. The question-answering model outputs a score indicating the confidence that the answer to the input question sentence q is the answer sentence T3(e_j) (1 ≦ j ≦ k) extracted from the document c. The similarity calculation unit 13B calculates the similarity s based on the score output by the question-answering model. The similarity s is, as an example, MB(q, c, {ID of e´}), that is, the confidence that the record e´ included in the record pair (e, e´) is the answer sentence. However, the similarity s is not limited to this example, and the similarity calculation unit 13B may calculate the similarity s by other methods. The similarity calculation unit 13B may, as an example, use the sum of the score when record e is used as the question sentence and the score when record e´ is used as the question sentence as the similarity.

[0116] FIG. 12 is a conceptual diagram of the similarity calculation process using the question-answering model. In the example of FIG. 12, the conversion unit 12B converts the record e and the auxiliary record R into a question sentence and a document, and the similarity calculation unit 13B inputs the question sentence and the document into the model MB, which is the question-answering model, to calculate the similarity s. Thus, in this processing example, the information processing apparatus 1B calculates the similarity of the records by converting the records into the question-answering format.

[0117] <Effect of the information processing apparatus 1B> As described above, in the information processing apparatus 1B according to this exemplary embodiment, the acquisition unit 11B further acquires the auxiliary record, the conversion unit 12B generates the converted auxiliary record by converting the auxiliary record, and the similarity calculation unit 13B inputs the converted record pair and the converted auxiliary record into the model MB to calculate the similarity regarding the converted record pair. Therefore, according to the information processing apparatus 1B according to this exemplary embodiment, the effect that the similarity can be calculated for records having various attributes using the model MB without training the model MB is obtained.

[0118] Also, in the information processing apparatus 1B according to this exemplary embodiment, the model MB includes a question-and-answer model that takes a question sentence and an answer sentence as inputs. The conversion unit 12B converts one record included in the record pair into a question sentence, and generates the converted record pair by converting each of the other record included in the record pair and the auxiliary record into an answer sentence. Therefore, according to the information processing apparatus 1B according to this exemplary embodiment, the effect that the similarity can be calculated using the question-and-answer model for records having various attributes without training the question-and-answer model is obtained.

[0119] 〔Exemplary Embodiment 4〕 A fourth exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the first to third exemplary embodiments are denoted by the same reference numerals, and the description thereof will not be repeated.

[0120] <Configuration of Information Processing Apparatus 1C> FIG. 13 is a block diagram showing the configuration of the information processing apparatus 1C according to this exemplary embodiment. The control unit 10C of the information processing apparatus 1C includes an acquisition unit 11, a conversion unit 12, a similarity calculation unit 13C, a similarity integration unit 17C, an output unit 14C, an identity determination unit 15A, and an integration unit 16A. The storage unit 20C stores the model MC in addition to the first data x, the second data x', and the similarity s.

[0121] (Similarity Calculation Unit 13C) The similarity calculation unit 13C calculates a plurality of similarities si for one record pair (e, e'). As an example, the similarity calculation unit 13C calculates a first similarity s1 by inputting the two records included in the record pair (e, e') to the model MC without swapping them. Further, the similarity calculation unit 13C calculates a second similarity s2 by inputting the two records included in the record pair (e, e') to the model MC after swapping them.

[0122] In Processing Examples 4 to 5 described in the above Exemplary Embodiment 2 and Processing Example 6 described in Exemplary Embodiment 3, the similarity of the record pair (e, e') is different from the similarity of the record pair (e', e). Therefore, in this exemplary embodiment, the similarity calculation unit 13C calculates the similarity of the record pair (e, e') and the similarity of the record pair (e', e), respectively, and determines identity with reference to these similarities.

[0123] However, the method by which the similarity calculation unit 13C calculates a plurality of similarities si is not limited to the above-described example, and the similarity calculation unit 13C may calculate a plurality of similarities si by other methods. For example, the similarity calculation unit 13C may calculate a plurality of similarities si using a plurality of models. In this case, for example, the conversion unit 12 performs a plurality of conversions on one record pair, and the similarity calculation unit 13C inputs the converted record pairs into respective models (document classification model, image classification model, ···) to calculate a plurality of similarities si.

[0124] Alternatively, the similarity calculation unit 13C may generate a plurality of converted record pairs by converting one record pair by a plurality of conversion methods, and calculate a plurality of similarities si by inputting the plurality of converted record pairs into one model.

[0125] (Similarity integration unit 17C) The similarity integration unit 17C integrates a plurality of similarities si into an integrated similarity s. As an example, the similarity integration unit 17C calculates the integrated similarity s by taking the average or weighted average of the plurality of similarities si. However, the method by which the similarity integration unit 17C integrates a plurality of similarities si is not limited to the above-described example, and the similarity integration unit 17C may calculate the integrated similarity s by other methods. For example, the similarity integration unit 17C may use the sum or integrated value of the plurality of similarities si as the integrated similarity s.

[0126] In this specification, it can also be said that the similarity integration unit 17C is configured to determine the identity of the target record pair based on a plurality of similarities si regarding the target record pair.

[0127] (Output unit 14C) The output unit 14C outputs the integrated similarity s obtained by integrating the plurality of similarities si. As an example, the output unit 14C outputs by writing the similarity s to the storage unit 20C.

[0128] (Model MC) The model MC is a model for calculating similarity. As an example, the model MC is a model having asymmetry with respect to the mutual interchange of two elements input to the model. The model MC includes, as an example, at least any one of an implication recognition model, a paraphrase prediction model, and a question-and-answer model.

[0129] FIG. 14 is a diagram showing a specific example of the similarity si calculated by the similarity calculation unit 13C. In FIG. 14, the first similarity s1 calculated by the similarity calculation unit 13C for the record pair (L1, R1) is "9", and the second similarity s2 calculated by the similarity calculation unit 13C for the record pair (R1, L1) obtained by interchanging the two records is "10". Thus, the similarity calculation unit 13C calculates the first similarity s1 and the second similarity s2 for one record pair, and the identity determination unit 15A determines that the records are identical if both the first similarity s1 and the second similarity s2 are the highest compared to other record pairs. In the example of FIG. 14, the identity determination unit 15A determines that the record L1 and the record R1 are identical, and also determines that the record L2 and the record R3 are identical.

[0130] FIG. 15 is a diagram showing another example of the similarity si calculated by the similarity calculation unit 13C. In FIG. 15, the similarity integration unit 17C aggregates the bidirectional similarities. As an example, the similarity integration unit 17C sets the similarity s as the sum of the similarity s1 of the record pair (L1, R1) and the similarity s2 of the record pair (R1, L1). In the example of FIG. 15, the similarity s of the record pair (L1, R1) is the sum of "10" and "9", that is, "19", and the similarity s of the record pair (L1, R2) is the sum of "9" and "7", that is, "16". Also, the similarity s of the record pair (L2, R2) is the sum of "9" and "4", that is, "13", and the similarity s of the record pair (L2, R3) is the sum of "8" and "8", that is, "16".

[0131] In the example of FIG. 15, the identity determination unit 15A determines that the record L1 and the record R1 are identical, and that the record L2 and the record R3 are identical, in the same manner as in the example of FIG. 14. In this example, furthermore, the identity determination unit 15A also determines that record pairs for which the similarity s is equal to or greater than a predetermined threshold among the record pairs determined to be identical are identical. Here, the threshold value is, for example, the minimum value of the similarity s of the record pairs determined to be identical (in the example of FIG. 15, "13"). The threshold value may be determined based on that ratio if the ratio of identical to non-identical is known. In the case where the threshold value is "13" in the example of FIG. 15, the identity determination unit 15A also determines that the records are identical for the record pair for which the similarity s is "13" or more, that is, the record pair (L1, R2) for which the similarity s is "16".

[0132] <Effect of the information processing apparatus 1C> As described above, in the information processing apparatus 1C according to this exemplary embodiment, the similarity calculation unit 13C calculates a plurality of similarities si for the above record pairs, and the output unit 14C outputs the integrated similarity s obtained by integrating the plurality of similarities si. Therefore, according to the information processing apparatus 1C according to this exemplary embodiment, an effect is obtained in that the similarity s of the record pairs can be calculated more accurately.

[0133] Further, in the information processing apparatus 1C according to the present exemplary embodiment, the model MC is a model having asymmetry with respect to the mutual interchange of two elements input to the model, and the similarity calculation unit 13C calculates the first similarity s1 by inputting the two records included in the record pair (e, e') to the model MC without interchanging them with each other, and calculates the second similarity s2 by inputting the two records included in the record pair (e, e') to the model MC after interchanging them with each other. Therefore, according to the information processing apparatus 1C according to the present exemplary embodiment, by integrating the first similarity s1 and the second similarity s2, an effect that the similarity s of the records can be calculated more accurately can be obtained.

[0134] 〔Exemplary Embodiment 5〕 The fifth exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the exemplary embodiments 1 to 3 are denoted by the same reference numerals, and the description thereof will not be repeated.

[0135] <Configuration of Information Processing Apparatus 1D> FIG. 16 is a block diagram showing the configuration of an information processing apparatus 1D according to the present exemplary embodiment. The control unit 10D of the information processing apparatus 1D includes an acquisition unit 11, a conversion unit 12, a similarity calculation unit 13, an output unit 14, an identity determination unit 15A, and a search result output unit 18D.

[0136] The acquisition unit 11 according to the present exemplary embodiment acquires input data from the user as the first record e included in the record pair (e, e'). The input data from the user is input, for example, by an input device (such as a keyboard, a mouse, etc.) connected to the input / output unit 40A.

[0137] Further, the acquisition unit 11 acquires one of the plurality of records included in the target data as the second record e' included in the record pair (e, e'). The target data is data to be searched, and includes, for example, one or more tables.

[0138] The identity determination unit 15A performs identity prediction for each record pair of the first record e and each of the plurality of records included in the target data.

[0139] Based on the similarity s calculated by the similarity calculation unit 13, the search result output unit 18D outputs a search result based on the input data and targeted at the target data. As an example, the search result output unit 18D refers to the determination result of the identity determination unit 15A and outputs a search result based on the input data and targeted at the target data. As an example, the search result output unit 18D outputs the search result to an output device (display, printer, etc.) connected to the input / output unit 40A. Also, the search result output unit 18D may output the search result by transmitting the search result to another device connected via the communication unit 30A. Further, the search result output unit 18D may output the search result by storing the search result in the storage unit 20A or an external storage device.

[0140] FIG. 17 is a diagram showing a specific example of the screen display output by the search result output unit 18D. In the example of FIG. 17, the input data is a character string input by the user into the text box 51, and the target data is the table T1 and the table T2 having a plurality of records. The identity determination unit 15A determines the identity for each record pair of the first record e, which is the input data of the user, and each of the records included in the table T1 and the record e' included in the table T2.

[0141] In the example of FIG. 17, the search result output unit 18D refers to the determination result of the identity determination unit 15A and outputs the search result 53 and the search result 54 based on the input data. The search result 53 is a search result retrieved from the table T1 with the character string "potato" as the input data. The search result 54 is a search result retrieved from the table T2 with the character string "potato" as the input data.

[0142] <Effect of the information processing apparatus 1D> As described above, in the information processing apparatus 1D according to the present exemplary embodiment, a configuration is adopted in which, with reference to the determination result of the identity determination unit 15A, a search result based on input data and including a search result for target data as the search target is output. Therefore, according to the information processing apparatus 1D according to the present exemplary embodiment, in addition to the effects achieved by the information processing apparatus 1 according to Exemplary Embodiment 1, an effect can be obtained that the search for target data based on input data can be performed more suitably.

[0143] The information processing apparatus 1D can also be described as follows. Acquisition means for acquiring, as a record pair, input data from a user and one of a plurality of records included in target data; Conversion means for generating a converted record pair by converting the record pair; Similarity calculation means for calculating a similarity regarding the converted record pair by inputting the converted record pair into a model; Output means for outputting, with reference to the similarity calculated by the similarity calculation means, a search result based on the input data and including the target data as the search target; An information processing apparatus comprising:

[0144] 〔Modification Example〕 <Modification Example 1> In each of the above-described exemplary embodiments, the information processing apparatuses 1, 1A, 1B, 1C, 1D (hereinafter referred to as "information processing apparatuses 1 etc.") determined the identity between the record e included in the first data x and the record e' included in the second data x'. The plurality of records to be determined by the information processing apparatuses 1 etc. may be records included in different data, or may be records included in common data. In other words, the information processing apparatuses 1 etc. may execute a process of searching for the same record from one database. Further, in the above-described exemplary embodiments, the case where the first data x and the second data x' are integrated has been described, but the information processing apparatuses 1 etc. may integrate three or more pieces of data.

[0145] <Modification Example 2> In each of the above exemplary embodiments, the information processing apparatus 1 or the like may select models MA, MB, and MC (hereinafter referred to as "model M") from a plurality of model candidates, or the user may select model M from a plurality of model candidates. The algorithm for the information processing apparatus 1 or the like to select model M is not limited. As an example, the information processing apparatus 1 or the like may select model M based on rules. For example, the information processing apparatus 1 or the like may select model M according to the characteristics of the record pair. Here, the characteristics of the record pair include, as an example, the attributes of the records included in the record pair, the data size of the records, the type of the database to which the records belong, and the attributes of the database.

[0146] <Modification Example 3> In each of the above exemplary embodiments, the data including records e and e' may be semi-structured data such as JSON or XML. By applying the information processing apparatus according to the above exemplary embodiments to the semi-structured data, it is possible to determine the identity of document data or web pages. For example, in a housing information site that provides housing information, there may be a plurality of web pages created for the same property. In this case, by performing identity determination on the web pages, the web pages can be grouped by property.

[0147] In this example, as an example, the record is a web page included in the target site. For example, record e = {id1: value1, id2: {id2-1: value2-1, id2-2: value2-1}, id3: value3} If so, the converted document is, as an example, “id1 is value1. id2-1 of id2 is value2-1. id2-2 of id2 is value2-1. id3 is value3.” That's it.

[0148] <Modification Example 4> Also, the record according to this specification may be graph data as shown in, for example, FIG. 18. FIG. 18 is a diagram showing an example of graph data. By applying the information processing apparatus 1 and the like according to this specification to the graph data, for example, face verification can be performed. For example, when the record is the graph shown in FIG. 18, the converted document is, as an example, “1 and 2 are linked. 1 and 4 are linked. 2 and 3 are linked. 2 and 4 are linked.” as follows.

[0149] <Modification Example 5> Also, the data included in the record may be a graph database as shown in, for example, FIG. 19. By applying the information processing apparatus 1 and the like according to this specification to the graph database, for example, the identity of communities of different social networking services (SNS) can be determined, and it can be applied to, for example, the investigation of criminal organizations. In this example, when the graph database is the one shown in FIG. 19, the converted document is, as an example, “Taro of age 23 follows Sakura of age 26. Taro of age 23 follows Emi of age 25. Sakura of age 26 follows Emi of age 25. Sakura of age 26 wrote via smartphone tweet of text “I’m sleepy.” date 20XX / YY / ZZ. Emi of age 25 follows Sakura of age 26. Emi of age 25 follows Taro of age 23.” as follows.

[0150] <Modification Example 6> In each of the above exemplary embodiments, a configuration may be adopted in which the information processing apparatus 1 or the like executes a learning phase for learning the model M. The machine learning method of the model M is not limited. As an example, a decision tree-based, linear regression, or neural network method may be used, or two or more of these methods may be used.

[0151] <Modification Example 7> In each of the above exemplary embodiments, a configuration may be adopted in which a converter having learnable parameters is added before and after the output of the model M. FIG. 20 is a diagram schematically showing a configuration in which the learned converters 121 and 122 having learnable parameters are provided before and after the output of the model M. The learned converters 121 and 122 have learnable parameters, and the learning unit (not shown) is a model that optimizes the way of record conversion (such as the way of writing sentences or the number of auxiliary records) and / or the parameters of the conversion using the training data. By providing the learned converters 121 and 122, the similarity of the records can be calculated with higher accuracy.

[0152] The machine learning method of the learned converters 121 and 122 is not limited. As an example, a decision tree-based, linear regression, or neural network method may be used, or two or more of these methods may be used. Further, the learned converters 121 and 122 may be models generated by active learning.

[0153] 〔Example of Realization by Software〕 Some or all of the functions of the information processing apparatuses 1, 1A, 1B, 1C, and 1D may be realized by hardware such as an integrated circuit (IC chip), or may be realized by software.

[0154] In the latter case, the information processing apparatuses 1, 1A, 1B, 1C, and 1D are realized by, for example, a computer that executes instructions of a program, which is software for realizing each function. An example of such a computer (hereinafter referred to as computer C) is shown in FIG. 18. Computer C includes at least one processor C1 and at least one memory C2. A program P for operating computer C as information processing apparatuses 1, 1A, 1B, 1C, and 1D is recorded in memory C2. In computer C, each function of information processing apparatuses 1, 1A, 1B, 1C, and 1D is realized by processor C1 reading program P from memory C2 and executing it.

[0155] As processor C1, for example, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), a DSP (Digital Signal Processor), an MPU (Micro Processing Unit), an FPU (Floating point number Processing Unit), a PPU (Physics Processing Unit), a microcontroller, or a combination thereof can be used. As memory C2, for example, a flash memory, an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a combination thereof can be used.

[0156] Note that computer C may further include a RAM (Random Access Memory) for expanding program P during execution or temporarily storing various data. Also, computer C may further include a communication interface for transmitting and receiving data to and from other devices. Also, computer C may further include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.

[0157] Also, the program P can be recorded on a non-transitory tangible recording medium M that is readable by the computer C. As such a recording medium M, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit can be used. The computer C can obtain the program P via such a recording medium M. Also, the program P can be transmitted via a transmission medium. As such a transmission medium, for example, a communication network or a broadcast wave can be used. The computer C can also obtain the program P via such a transmission medium.

[0158] [Supplementary Note 1] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope indicated in the claims. For example, embodiments obtained by appropriately combining the technical means disclosed in the above-described embodiments are also included in the technical scope of the present invention.

[0159] [Supplementary Note 2] Some or all of the above-described embodiments may also be described as follows. However, the present invention is not limited to the aspects described below. (Supplementary Note 1) An acquisition means for acquiring a record pair, A conversion means for generating a converted record pair by converting the record pair, A similarity calculation means for calculating a similarity regarding the converted record pair by inputting the converted record pair into a model, An output means for outputting the similarity calculated by the similarity calculation means, An information processing apparatus comprising the above.

[0160] (Supplementary Note 2) The model is selected from a plurality of model candidates, The conversion means generates the converted record pair by converting the record pair into a format corresponding to the input of the model, The information processing apparatus according to Supplementary Note 1.

[0161] (Appendix 3) The model includes a document classification model, The conversion process by the conversion means includes a process of generating the converted record pair by converting the record pair into a document, The information processing apparatus according to any one of Appendix 1 or 2.

[0162] (Appendix 4) The model includes an image classification model, The conversion process by the conversion means includes a process of generating the converted record pair by converting the record pair into an image, The information processing apparatus according to any one of Appendices 1 to 3.

[0163] (Appendix 5) The model includes an audio classification model, The conversion process by the conversion means includes a process of generating the converted record pair by converting the record pair into audio, The information processing apparatus according to any one of Appendices 1 to 4.

[0164] (Appendix 6) The model includes a graph classification model, The conversion process by the conversion means includes a process of generating the converted record pair by converting the record pair into a graph, The information processing apparatus according to any one of Appendices 1 to 5.

[0165] (Appendix 7) The acquisition means further acquires an auxiliary record, The conversion means generates a converted auxiliary record by converting the auxiliary record, The similarity calculation means calculates the similarity regarding the converted record pair by inputting the converted record pair and the converted auxiliary record into the model. The information processing apparatus according to any one of Appendices 1 to 6.

[0166] (Appendix 8) The model includes a question-and-answer model that takes a question sentence and an answer sentence as inputs, The conversion means generates the converted record pair by converting one record included in the record pair into a question sentence and converting each of the other record included in the record pair and the auxiliary record into an answer sentence. The information processing apparatus according to Appendix 7.

[0167] (Appendix 9) The similarity calculation means calculates a plurality of similarities for the record pair, The output means outputs the integrated similarity obtained by integrating the plurality of similarities. The information processing apparatus according to any one of Appendices 1 to 8.

[0168] (Appendix 10) The model is a model having asymmetry with respect to the mutual interchange of two elements input to the model, The similarity calculation means, calculates a first similarity by inputting the two records included in the record pair to the model without interchanging them with each other, calculates a second similarity by inputting the two records included in the record pair to the model after interchanging them with each other. The information processing apparatus according to Appendix 9.

[0169] (Appendix 11) At least one processor, acquires a record pair, generates a converted record pair by converting the record pair, calculates a similarity regarding the converted record pair by inputting the converted record pair to a model, outputs the calculated similarity, and includes an information processing method.

[0170] (Appendix 12) Cause the computer to perform an acquisition process of acquiring a record pair, a conversion process of generating a converted record pair by converting the record pair, a similarity calculation process of calculating a similarity regarding the converted record pair by inputting the converted record pair into a model, and an output process of outputting the similarity calculated in the similarity calculation process. An information processing program for causing the above to be executed.

[0171] (Appendix 13) An acquisition means for acquiring, as a record pair, input data from a user and one of a plurality of records included in target data, a conversion means for generating a converted record pair by converting the record pair, a similarity calculation means for calculating a similarity regarding the converted record pair by inputting the converted record pair into a model, and an output means for outputting a search result based on the input data and targeting the target data with reference to the similarity calculated by the similarity calculation means. An information processing apparatus comprising the above.

[0172] [Supplementary Note 3] Some or all of the above-described embodiments can also be expressed as follows.

[0173] An information processing apparatus comprising at least one processor, wherein the processor performs an acquisition process of acquiring a record pair, a conversion process of generating a converted record pair by converting the record pair, a similarity calculation process of calculating a similarity regarding the converted record pair by inputting the converted record pair into a model, and an output process of outputting the similarity calculated in the similarity calculation process.

[0174] Incidentally, this information processing apparatus may further include a memory, and a program for causing the processor to execute the acquisition process, the conversion process, the similarity calculation process, and the output process may be stored in this memory. Further, this program may be recorded on a non-transitory tangible computer-readable recording medium.

Explanation of Signs

[0175] 1, 1A, 1B, 1C, 1D Information processing apparatus 11, 11B Acquisition unit 12, 12B Conversion unit 13, 13B, 13C Similarity calculation unit 14, 14C Output unit 16A Integration unit

Claims

1. an acquisition means for acquiring a record pair; a conversion means for generating a converted record pair by converting the record pair; a similarity calculation means for calculating a similarity regarding the converted record pair by inputting the converted record pair into a model; an output means for outputting the similarity calculated by the similarity calculation means ; and the conversion means generates the converted record pair by selecting a model from a plurality of model candidates according to the characteristics of the record pair and converting the record pair into a format corresponding to the input of the selected model, an information processing apparatus.

2. an acquisition means for acquiring a record pair; a conversion means for generating a converted record pair by converting the record pair; a similarity calculation means for calculating a similarity regarding the converted record pair by inputting the converted record pair into a model; an output means for outputting the similarity calculated by the similarity calculation means; and the model includes an image classification model, and the conversion process by the conversion means includes a process of generating the converted record pair by converting the record pair into an image an information processing apparatus.

3. an acquisition means for acquiring a record pair; a conversion means for generating a converted record pair by converting the record pair; a similarity calculation means for calculating a similarity regarding the converted record pair by inputting the converted record pair into a model; an output means for outputting the similarity calculated by the similarity calculation means; and the model includes a voice classification model, and the conversion process by the conversion means includes a process of generating the converted record pair by converting the record pair into voice an information processing apparatus.

4. an acquisition means for acquiring a record pair; a conversion means for generating a converted record pair by converting the record pair; a similarity calculation means for calculating a similarity regarding the converted record pair by inputting the converted record pair into a model; an output means for outputting the similarity calculated by the similarity calculation means; and the model includes a graph classification model, and the conversion process by the conversion means includes a process of generating the converted record pair by converting the record pair into a graph, an information processing apparatus.

5. an acquisition means for acquiring a record pair; Conversion means for generating a converted record pair by converting the record pair; Similarity calculation means for calculating the similarity regarding the converted record pair by inputting the converted record pair into a model; Output means for outputting the similarity calculated by the similarity calculation means; comprising; the acquisition means further acquires an auxiliary record; the conversion means generates a converted auxiliary record by converting the auxiliary record; the similarity calculation means calculates the similarity regarding the converted record pair by inputting the converted record pair and the converted auxiliary record into the model An information processing apparatus.

6. The model includes a question-and-answer model that takes a question sentence and an answer sentence as inputs; the conversion means generates the converted record pair by converting one record included in the record pair into a question sentence and converting each of the other record included in the record pair and the auxiliary record into an answer sentence The information processing apparatus according to claim 5.

7. Acquisition means for acquiring a record pair; Conversion means for generating a converted record pair by converting the record pair; Similarity calculation means for calculating the similarity regarding the converted record pair by inputting the converted record pair into a model; Output means for outputting the similarity calculated by the similarity calculation means; comprising; the similarity calculation means calculates a plurality of similarities regarding the record pair; the output means outputs an integrated similarity obtained by integrating the plurality of similarities; the model is a model having asymmetry with respect to the mutual interchange of two elements input to the model; the similarity calculation means, calculates a first similarity by inputting the two records included in the record pair into the model without interchanging them with each other; calculates a second similarity by inputting the two records included in the record pair into the model after interchanging them with each other An information processing apparatus.

8. At least one processor acquires a record pair; generates a converted record pair by converting the record pair; calculates the similarity regarding the converted record pair by inputting the converted record pair into a model; outputs the calculated similarity including In the step of generating the converted record pair, the at least one processor selects a model from a plurality of model candidates according to the characteristics of the record pair, and generates the converted record pair by converting the record pair into a format corresponding to the input of the selected model. Information processing method. **Claim 9** A computer is caused to perform an acquisition process of acquiring a record pair, a conversion process of generating a converted record pair by converting the record pair, a similarity calculation process of calculating a similarity regarding the converted record pair by inputting the converted record pair into a model, and an output process of outputting the similarity calculated in the similarity calculation process. In the conversion process, a model is selected from a plurality of model candidates according to the characteristics of the record pair, and the converted record pair is generated by converting the record pair into a format corresponding to the input of the selected model. Information processing program.

Citation Information

Patent Citations

  • Learning program and learning method

    JP2019185244A

  • Information processing device, information processing method and information processing program

    JP2021174300A

  • Transliteration of data records for improved data matching

    JP2022510818A

  • Organizational data enrichment

    US20170091274A1

  • Automated and dynamic method and system for clustering data records

    US20210374164A1