Method and system for processing data records

By using a trained data representation learning model to convert data records into feature vectors, the method addresses the challenges of duplicate removal and record matching in data management systems, achieving efficient storage and real-time processing.

JP7691185B2Active Publication Date: 2025-06-11INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022569281
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-05-28
Filing Date
2021-04-16
Publication Date
2025-06-11
Estimated Expiration
2041-04-16

Smart Images

  • Figure 0007691185000012
    Figure 0007691185000012
  • Figure 0007691185000013
    Figure 0007691185000013
  • Figure 0007691185000014
    Figure 0007691185000014
Patent Text Reader

Abstract

The present disclosure relates to a method that includes providing a set of one or more records, each record in the set of records including a set of one or more attributes; inputting values ​​of the set of attributes of the set of records into a trained data representation learning model, thereby receiving, as an output of the trained data representation model, a set of feature vectors that each represent the set of records; and storing the set of feature vectors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital computer systems, and more particularly, to a method for processing data records.

Background Art

[0002] Storing and processing data can be an essential requirement for a data management system to function properly. For example, removing duplicate records or detecting matches within a database is a very important step in the data cleansing process because duplicates can potentially have a significant impact on the results of any subsequent data processing or data mining. The matching process for record linkage becomes even more complex as the complexity increases for a number of different attributes required across different geographical regions, countries, etc., and constitutes one of the major challenges regarding record linkage algorithms.

Summary of the Invention

[0003] Various embodiments provide a method, a computer system, and a computer program product for processing data records, as described by the subject matter of the independent claims. Advantageous embodiments are described in the dependent claims. Embodiments of the present invention can be freely combined with each other if they are not mutually exclusive.

[0004] In one aspect, the present invention is to provide a set of one or more records, wherein each record of the set of records includes a set of one or more attributes, and to input the values of the set of attributes of the set of records into a trained data representation learning model, whereby, as an output of the trained data representation model, receive a set of feature vectors respectively representing the set of records, and store the set of feature vectors.

[0005] In another aspect, the present invention relates to a computer program product comprising a computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being configured to perform all of the steps of the method according to the foregoing embodiments.

[0006] In another aspect, the present invention relates to a computer system configured to input a set of values of a set of attributes of a set of one or more records into a trained data representation learning model, thereby receiving, as an output of the trained data representation model, a set of feature vectors respectively representing the set of records, and storing the set of feature vectors.

[0007] This subject matter may enable efficient storage and representation of data records in a database. The generated set of feature vectors may have the advantages of saving storage resources and providing a compact or alternative storage solution for storing data records. The generated set of feature vectors may further have the advantage of optimizing the processing of data records by using them instead of the records. For example, it may be fast, efficient, and reliable to match feature vectors instead of records. A trained data representation learning model may be used to provide an accurate and controllable definition of a vector space for representing records. The vector space may be defined such that, for example, duplicate feature vectors can be identified by simplified vector operations in the vector space. An exemplary vector operation may be a distance function.

[0008] In the following, embodiments of the present invention will be described in more detail with reference to the following drawings by way of example only.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 7A

Figure 7B

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0010] The description of various embodiments of the present invention is presented for illustrative purposes and is not intended to be exhaustive or to limit the invention to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, actual application, or technical improvements over technologies found in the marketplace, or to enable other skilled artisans to understand the embodiments disclosed herein.

[0011] A data record or record is a collection of related data items such as the name, date of birth, and class of a particular user. A record represents an entity, where an entity refers to a user, object, or concept related to the information stored in the record. The terms "data record" and "record" are used interchangeably.

[0012] A dataset may include a set of records processed by the present subject matter. For example, a dataset may be provided in the form of a collection of related records contained in a file, for example, a dataset may be a file containing the records of all students in a class. A dataset may be, for example, a database table or a file in a Hadoop(R) file system. In another example, a dataset may include one or more documents such as an HTML page or other document types.

[0013] The dataset may be stored, for example, in a central repository. The computer system may be configured to access the central repository. For example, the central repository may or may not be part of the computer system. The central repository may be a data store or storage that stores data received from a plurality of client systems. The dataset may include a subset of existing records in the central repository that are identified or selected for processing by the present method. The subset of records may be selected, for example, based on the values of one or more attributes of the records. For example, the subset of records may represent entities belonging to a particular country or region. The records of the dataset may be pre-processed, for example, before being processed by the present method. This pre-processing may include, for example, converting the format of the attribute values of the records of the dataset. For example, the attribute values may be displayed in uppercase, noise characters (such as characters like -,., / ) of the attribute values may be removed, and anonymous attribute values (such as attribute values like city = nowhere or first name = test) may be removed.

[0014] The data representation learning model is configured to output a feature vector of a record, and the distance between two feature vectors generated based on two records using the data representation learning model is made to be a measure of the similarity of the attributes of the two records. For example, the data representation learning model may receive a record as its input and may be trained to generate, typically, a d-dimensional vector space, and each unique record is assigned a corresponding vector in the space. The feature vectors may be arranged in the vector space such that records sharing similar records or attribute values are located close to each other in the space. The feature vector may be a single ordered array. The feature vector may be an example of a first-order tensor.

[0015] In one example, the trained data representation learning model may be a single similarity encoder configured to receive a set of one or more attributes of a record and generate a feature vector representing the record. The feature vector may be obtained, for example, by collectively processing the values of the set of attributes, or may be obtained as a concatenation of individual feature vectors associated with the set of attributes, and the similarity encoder is configured to receive the values of the set of attributes sequentially and generate the associated individual feature vectors.

[0016] In another example, the trained data representation learning model may include one similarity encoder for each attribute of a set of attributes of a record, and each similarity encoder of the plurality of similarity encoders is configured to receive each attribute of the set of attributes and generate a corresponding individual feature vector. Individual feature vectors may be concatenated to obtain a feature vector representing the record.

[0017] In another example, the trained data representation learning model may be a neural network configured to receive a set of one or more attributes of a record and generate a feature vector representing the record. In another example, the trained data representation learning model may include one neural network for each attribute of a set of attributes of a record, and each neural network of the plurality of neural networks is configured to receive each attribute of the set of attributes and generate a corresponding individual feature vector. The individual feature vectors may be concatenated to obtain a feature vector representing the record.

[0018] The concatenation of the individual feature vectors may be a concatenation or addition of the elements of the individual feature vectors to the feature vector such that the feature vector contains the elements of the individual feature vectors.

[0019] This method may be deployed and used, for example, in big data solutions (such as IBM BigMatch technology running on IBM BigInsights(R), Cloudera(R), and Hortonworks(R)), together with information integration software (such as Informatica(R) Power Center(R), IBM Information Server). The computer system may be, for example, a master data management (MDM) system.

[0020] According to one embodiment, this method may further include receiving a record that includes a set of attributes. To obtain a feature vector of the received record from a trained data representation learning model, the values of the set of attributes of the received record may be input into the trained data representation learning model. The obtained feature vector may be compared with at least a portion of the set of feature vectors to determine a matching level with the set of feature vectors. The obtained feature vector or the received record or both may be stored based on the matching level. In one example, the obtained feature vector may be compared with each feature vector of the set of feature vectors. In another example, the obtained feature vector may be compared with each feature vector of a selected subset of the set of feature vectors.

[0021] Conventional matching algorithms may use bucketing of features. However, if the buckets are too large (more than 1000 - 1500 records), the matching performance may degrade to the point where the matching cannot be used in real time. Additionally, creating an appropriate bucketing strategy to avoid such large buckets may require deep human subject matter expertise, which is difficult and expensive to obtain. Also, rebucketing may require human expertise and multiple iterations of adjustment (e.g., 4 - 6 weeks). This embodiment can solve these problems by using a vector space that enables easy matching of feature vectors within large buckets. The matching according to this embodiment is not limited to a small number of threshold values for attributes because distance calculation at the vector level enables matching of a large number of attributes. Another advantage of using feature vectors may be as follows. Comparison functions such as edit distance and phonetic distance may work well for simple attributes such as first name attributes. However, these comparison functions may not work for free text fields such as product descriptions of 200 words. This can be solved by using feature vectors. For this purpose, the content of the free text description is encoded as a vector, and nearby vectors refer to similar free text descriptions.

[0022] According to one embodiment, storing a set of feature vectors includes clustering the set of feature vectors into clusters and associating each of the stored feature vectors with cluster information indicating the corresponding cluster. Storing feature vectors as clusters may be advantageous because it may enable optimal access to the stored data based on the clusters. For example, queries for accessing the stored data can be improved using criteria regarding the clusters. This can improve the speed of access to the data.

[0023] According to one embodiment, storing a set of feature vectors includes clustering the set of feature vectors into clusters. This method further includes determining the distance between the acquired feature vector and each of the plurality of clusters, wherein at least a portion of the set of feature vectors includes the cluster of feature vectors having the closest distance to the acquired feature vector. That is, a selected subset of the set of feature vectors includes the feature vectors among the clusters of feature vectors having the closest distance to the acquired feature vector. The distance between a cluster and the acquired feature vector may be the distance between the acquired feature vector and a vector representing the cluster of feature vectors. The vector representing the cluster may be, for example, the centroid of the cluster of feature vectors. In another example, any feature vector of the cluster may be the vector representing the cluster.

[0024] This embodiment can improve the speed of the matching process, and thus can save the processing resources required to perform the matching with all the stored feature vectors in other ways. Therefore, the matching process enabled by this embodiment may be used in real time, for example, during the operation of an MDM system.

[0025] According to one embodiment, this method is executed in real time and a record is received as part of a creation operation or an update operation. For example, this method may be executed during the execution of an MDM system. For example, in response to receiving a data record, this method may be automatically executed for that record. For example, in order to execute the record matching process in real time during a creation / update operation, the matching process may need to be completed within 100 milliseconds. Using feature vectors according to this subject matter may make it possible to shorten the time of the matching process to the level required by real-time processing.

[0026] According to one embodiment, the trained data representation learning model outputs a feature vector among a set of feature vectors by generating an individual feature vector for each attribute of a set of attributes and combining the individual feature vectors to obtain the feature vector. By processing attributes at the individual level, this embodiment can enable more detailed access to the features of records, and in this way, can provide a reliable representation of the records.

[0027] The data representation learning model may include, for example, a similarity encoder for each attribute of a set of attributes, and the similarity encoder is configured to output an individual feature vector of a related attribute among the set of attributes such that the distance between two individual feature vectors generated based on two attributes using the similarity encoder is a measure of the similarity of the two attributes.

[0028] According to one embodiment, the trained data representation learning model is configured to receive an input value and output a set of feature vectors in parallel. That is, the data representation learning model is configured to process the values of a set of attributes in parallel (or simultaneously) and generate outputs in parallel (or simultaneously). Thereby, according to the present subject matter, for example, the speed of the matching process can be further improved.

[0029] According to one embodiment, the trained data representation learning model includes a set of trained data representation models at the attribute level each associated with a set of attributes, and the output of each feature vector of the set of feature vectors includes inputting the value of each attribute of the set of attributes to the trained data representation model at the associated attribute level, receiving an individual feature vector from each of the trained data representation models at the attribute level in response to this input, and combining the individual feature vectors to obtain the feature vector. The combination of the individual feature vectors may be a concatenation or addition of the elements of the individual feature vectors to the feature vector such that the feature vector includes the elements of the individual feature vectors.

[0030] This embodiment may enable structuring a data representation learning model into a pipeline. Each attribute to be compared belongs to one pipeline. Each pipeline predicts the similarity of its attribute. The pipeline may be advantageous because it can provide an extensible structure. This may enable dealing with an increase in the number of attributes by adding new pipelines. For example, the pipeline approach may enable a customer to add a new pipeline to predict regarding attributes specific to that customer (e.g., a unique id of a specific customer).

[0031] According to one embodiment, the trained data representation learning model further includes a set of trained weights respectively associated with a set of attributes, and combining includes weighting each of the individual feature vectors using each of the trained weights of the set of trained weights and combining, for example, by concatenating the weighted individual feature vectors. This may enable controlling the prediction process at the attribute level by changing the set of weights according to their corresponding attributes. For example, during training, the weights associated with a specific attribute may be changed differently compared to other attributes. This weighting process occurs as part of the learning step in the data representation learning model.

[0032] This embodiment may further enable a customer to weight the impact on different comparisons. For example, a customer may have high-quality address data. Thus, the comparison regarding the address field may have a large impact on the overall prediction. Therefore, the customer may change the trained weights associated with the address data.

[0033] According to one embodiment, each of the attribute-level trained data representation models is a neural network.

[0034] According to one embodiment, a trained data representation learning model is trained to optimize a loss function. In one example, to optimize the loss function, the trained data representation learning model is trained using backpropagation. The loss function may indicate (or be a function of) a measure of similarity between the feature vectors of pairs of records. The measure of similarity may be a combination of individual similarities. Each individual similarity of the plurality of individual similarities indicates the similarity of two individual feature vectors generated for the same attribute within a pair of records. The measure of similarity between two feature vectors may be, for example, the Euclidean distance, the cosine distance, or the difference of pairs of elements of the two feature vectors.

[0035] According to one embodiment, a trained data representation learning model is trained to optimize a loss function, and the loss function is a measure of similarity between the feature vectors of pairs of records. In one example, to optimize the loss function, the trained data representation learning model is trained using backpropagation.

[0036] According to one embodiment, the trained data representation learning model includes at least one neural network trained according to a siamese neural network architecture. For example, the neural network of the trained data representation learning model may be one of two networks of a trained siamese neural network. This may be advantageous in order to achieve a semantically significant space where related patterns (e.g., records of the same entity) are close to each other while avoiding the proximity of semantically unrelated patterns (e.g., records of different entities), realizing a non-linear embedding of the data. The comparison between feature vectors may be simplified using the distance. In another example, the trained data representation learning model includes an autoencoder.

[0037] According to one embodiment, the trained data representation learning model includes one trained neural network for each attribute of a set of attributes, and the output of each feature vector of the set of feature vectors is input to the trained neural network that associates the value of each attribute of the set of attributes, and in response to this input, receiving an individual feature vector from each of the trained neural networks, and combining the individual feature vectors to obtain the feature vector.

[0038] According to one embodiment, the method further includes receiving a training data set including pairs of records including a set of attributes, each pair being labeled, and training a data representation learning model using the training data set to generate a trained data representation learning model. The data representation learning model may be trained in a supervised manner using pairs of records, and each pair of records is labeled with a pre-defined label. The label of the pair of records may indicate "same" or "different". Here, "same" means that the pair of records represents the same entity, and "different" means that these two records belong to different physical entities (e.g., people).

[0039] According to one embodiment, the data representation learning model includes a set of attribute-level data representation models each associated with a set of attributes. Training of the data representation learning model includes, for each pair of records and for each attribute in the set of attributes, inputting a pair of values of the attribute (within the pair of records) into the corresponding attribute-level data representation model, thereby obtaining a pair of individual feature vectors, calculating an individual similarity level between the pair of individual feature vectors, and weighting the individual similarity levels using trainable weights of the attributes. Training of the data representation learning model further includes, for each pair of records, determining a measure of similarity between the feature vectors of the pair of records as a combination of the weighted individual similarity levels, and evaluating a loss function using this measure. The evaluated loss function may be used in a minimization process during training of the data representation learning model. The combination of the weighted individual similarity levels may be the sum of the weighted individual similarity levels.

[0040] The trainable parameters of the set of attribute-level data representation models and the set of weights may be changed during training to achieve an optimal value of the loss function. For example, the set of weights may increase the importance of one attribute over another. To that end, during training, the set of weights may be changed differently with respect to the set of attributes. In a first weight change configuration, the set of attributes may be ranked according to a priority defined by the user, and the weights may be changed using an amount (difference value) corresponding to their ranking. In another example, the user may change the trained weights during the test / inference phase, e.g., the user may increase or decrease one or more of the trained weights. This may provide an opportunity to adjust the weights according to a priority defined by the user.

[0041] According to one embodiment, a set of attributes includes a first subset of attributes and a second subset of attributes. The method further includes receiving a first trained data representation learning model including a first subset of attribute-level trained data representation learning models respectively associated with the first subset of attributes, the first trained data representation learning model being configured to receive values of the first subset of attributes of a record, input a feature vector of the record into the associated attribute-level trained data representation models of the first subset of the attribute-level trained data representation learning models for each attribute of the first subset of attributes, receive individual feature vectors from each of the attribute-level trained data representation models of the first subset of the attribute-level trained data representation learning models, and output by combining the individual feature vectors to obtain the feature vector. A second subset of attribute-level data representation learning models may be provided for the second subset of attributes, and each attribute-level data representation learning model of the second subset is configured to generate a feature vector for each attribute of the second subset of attributes. The data representation learning model may be created to include the first trained data representation learning model and the second subset of attribute-level data representation learning models. The created data representation learning model may be trained to generate a trained data representation learning model.

[0042] The first trained data representation learning model may be centrally generated, for example, by a service provider. The first trained data representation learning model may be provided to a plurality of users or clients. Each user of the plurality of users may use the first trained data representation learning model according to the present subject matter to adapt to the user's needs in an efficient and controlled manner. For example, a user may add one or more new pipelines associated with new attributes specific to the user.

[0043] According to one embodiment, the first trained data representation learning model further includes a first subset of trained weights associated with a first subset of attributes, and combining is performed using the first subset of trained weights. For example, individual feature vectors are weighted by each weight, and the weighted individual feature vectors are combined. The created data representation learning model further includes a second subset of trainable weights associated with a second subset of attributes, and the trained data representation learning model inputs each feature vector of the set of feature vectors into the associated attribute-level trained data representation models of the first and second subsets of the set of attributes, and receives individual feature vectors from each of the attribute-level trained data representation models of the first and second subsets of the attribute-level trained data representation learning model, and combines the individual feature vectors using each weight of the first and second subsets of weights and outputs by obtaining the feature vectors. During training, the first subset of weights and the second subset of weights may be changed differently.

[0044] FIG. 1 shows an exemplary computer system 100. The computer system 100 may be configured to perform, for example, master data management or data warehousing or both. For example, the computer system 100 may enable a deduplication system. The computer system 100 includes a data integration system 101 and one or more client systems or data sources 105. The client system 105 may include a computer system (as described with reference to FIG. 6, for example). The client system 105 may communicate with the data integration system 101 via a network connection, which may include, for example, a wireless local area network (WLAN) connection, a WAN (Wide Area Network) connection, a LAN (Local Area Network) connection, the Internet, or a combination thereof. The data integration system 101 may control access rights (such as read access rights and write access rights) to the central repository 103.

[0045] The data set 107 of records stored in the central repository 103 may include a set a of attributes such as a first name attribute 1 ...a N (N≧1) values. This example is described with respect to several attributes, although more or fewer attributes may be used. The data set 107 used in accordance with this subject matter may include at least a portion of the records in the central repository 103.

[0046] The data records stored in the central repository 103 may be received from the client system 105 and processed by the data integration system 101 (e.g., to convert to a unified structure) before being stored in the central repository 103. For example, the records received from the client system 105 may have a structure different from the structure of the records stored in the central repository 103. For example, the client system 105 may be configured to provide the records in XML format, JSON format, or other formats that enable associating attributes and corresponding attribute values.

[0047] In another example, the data integration system 101 may import the data records of the central repository 103 from the client system 105 using one or more extract-transform-load (ETL) batch processes, or via hypertext transfer protocol (HTTP) communication, or via other types of data exchange.

[0048] The data integration system 101 may be configured to process the received records using one or more algorithms. The data integration system 101 may include, for example, a trained data representation learning model 120. The trained data representation learning model 120 may be received, for example, from a service provider. In another example, the trained data representation learning model 120 may be created in the data integration system 101. The trained data representation learning model 120 is configured to generate a feature vector representing a specific record. The specific record includes, for example, a set of attributes a 1 ...a N For this purpose, the trained data representation learning model 120 may be configured to receive all or some of the values of the attributes a 1 ...a N of the record and generate a feature vector.

[0049] In one example, the trained data representation learning model 120 may include a plurality of attribute-level trained data representation learning models 121.1 to 121.N. Each of the attribute-level trained data representation learning models 121.1 to 121.N may be associated with each attribute of the set of attributes a 1 ...a N Each of the attribute-level trained data representation learning models 121.1 to 121.N may be configured to receive the values of each attribute a 1 ...a N and generate corresponding individual feature vectors. The individual feature vectors may be combined to obtain a single feature vector representing the record. This combination may be performed, for example, using N trained weights α 1 ...α N associated with the set of N attributes a 1 ...α N .

[0050] FIG. 2 is a flowchart of a method for storing data according to an example of the present subject matter. For the purpose of illustration, the method described in FIG. 2 may be implemented in the system shown in FIG. 1, but is not limited to this implementation. The method of FIG. 2 may be executed, for example, by the data integration system 101.

[0051] A set R of K records 1 ...R K may be provided in step 201, where K ≧ 1. For example, the data set may include a set R of K records 1 ...R K Each record in the set of K records has a set of N attributes a 1 ...a Nmay include and N ≧ 1. In one example, each record in a set of K records may include one attribute (i.e., N = 1). In accordance with this subject matter, generating a feature vector using one attribute may be advantageous because the feature vector can be used to represent specific features of the entity being investigated. For example, this may enable clustering or matching records representing students of the same age or the same region.

[0052] In another example, each record in a set of records may include multiple attributes. The set of records may share one or more subsets of attributes from the set of attributes a 1 ...a N and thus may or may not include the same complete set of attributes a 1 ...a N Using all attributes can enable representing the entire record. For example, a feature vector can provide features representing the whole. This may be advantageous for detecting duplicate records.

[0053] In one example, the set of records may include all records of an existing database such as repository 103. This may enable providing feature vectors for all existing records. In another example, the set of records may include one record. This one record may be, for example, a newly received record. This may be advantageous, for example, when constructing a new database.

[0054] In step 203, the values of the set of attributes of the set of records may be input into the trained data representation learning model 120. In step 205, the output of the trained data representation learning model 120 may be received. This output is a set of K feature vectors F 1 ...F KIt includes. Steps 203 and 205 may enable the inference of the trained data representation learning model 120. Therefore, the two steps 203 and 205 may be collectively referred to as the inference step.

[0055] In the example of the first inference, the trained data representation learning model 120 may be configured to process each record R of the set of records as follows i (i = 1, 2,... or K). The values of the set of attributes of record R i may be input to the trained data representation learning model 120 simultaneously (e.g., in parallel). The trained data representation learning model 120 may be configured to generate a feature vector of the received values of record R i . The generated feature vector may represent record R i . In this example, the trained data representation learning model 120 may include, for example, one neural network.

[0056] In the example of the second inference, the trained data representation learning model 120 may be configured to process each record R of the set of records as follows i . The values of the set of attributes of record R i may be input to the trained data representation learning model 120 simultaneously. The trained data representation learning model 120 may be configured to generate individual feature vectors for each received value of the set of attributes. The individual feature vectors may be combined by the trained data representation learning model 120 to generate a feature vector representing record R i . The individual feature vectors may be generated by the attribute-level trained data representation models of the trained data representation learning model 120 respectively associated with the set of attributes.

[0057] In the third inference example, the trained data representation learning model 120 may be configured to process each record R of the set of records as follows i . The values of the set of attributes of record Ri Values of the set of attributes may be successively input to the trained data representation learning model 120. The trained data representation learning model 120 may be configured to generate an individual feature vector for each received value. The individual feature vectors may be combined by the trained data representation learning model 120 to generate a feature vector representing record R i This example may be particularly advantageous when the set of attributes is of the same type. That is, a single trained data representation learning model (e.g., a single neural network) may effectively generate feature vectors of different attributes of the same type. For example, if the set of attributes includes a person's business and private phone numbers, the trained data representation learning model 120 may be used to generate (successively) the feature vectors of these two attributes. For example, if the entity is a product and the set of attributes is height and width, the trained data representation learning model 120 may be used to generate (successively) the feature vectors of these two attributes.

[0058] The combination of the individual feature vectors may be the concatenation or addition of the elements of the individual feature vectors to the feature vector such that the feature vector contains the elements of the individual feature vectors.

[0059] The inference of the trained data representation learning model 120 may result in a set F of K feature vectors representing the set of records provided in step 201 1 ...F K Each of the sets of feature vectors may be a representation within a d-dimensional mathematical space. In step 207, the set of feature vectors may be stored. This may enable an improvement in storage utilization.

[0060] In the example of the first storage, instead of storing a set of records, a set of feature vectors may be stored. This may provide a lightweight version of the record database. This may be particularly advantageous when the database is used for specific applications that can be satisfied by the feature vectors and do not require the entire record.

[0061] In the example of the second storage, the set of feature vectors may be stored in relation to each set of records. This may be particularly advantageous because the set of feature vectors may not require a large amount of storage resources. This may make it possible to provide the set of feature vectors as metadata for the set of records.

[0062] In the example of the third storage, the set of feature vectors may be clustered into clusters. Cluster information may be determined for each of the plurality of clusters. The cluster information may be, for example, a cluster ID or the centroid of the cluster, and the set of feature vectors may be stored in relation to the cluster information of the clusters to which they belong.

[0063] FIG. 3 is a flowchart of a method for collating records according to an example of the present subject matter. For purposes of explanation, the method described in FIG. 3 may be implemented in the system shown in FIG. 1, but is not limited to this implementation. The method of FIG. 3 may be executed, for example, by the data integration system 101.

[0064] In step 301, record R K+1 may be received. The received record has a set of attributes a 1 ... a N and matches each of the N sets of attributes A 1 ... A N . The number of attributes in the set of attributes A 1 ... A N is equal to the number of attributes in the N sets of attributes used in FIG. 2. However, the N sets of attributes A 1 ... A Nis a set a of N attributes 1 ...a N and may or may not be exactly the same. For example, each attribute A i may be equivalent to a i or exactly the same. For example, a 1 may be the "first name" attribute, and A 1 may be the "name" attribute, or a 1 may be the "private phone number" attribute, and A 1 may be the "business phone number" attribute.

[0065] Record R K+1 may be received, for example, in a data request, for example, from client system 105. The data request may be, for example, an update operation request or a create operation request. The received record may be a structured record or an unstructured record. In the case of an unstructured record received (e.g., an article), step 301 may further include processing the unstructured record to identify the attribute values of the set of attributes encoded within the received record. In another example, the request may be a matching request for matching a record with a set of records. In either case, matching of the record with the set of records may be required. For example, the received record may be matched with existing records to prevent storing duplicate records before storing the received record.

[0066] In step 303, the set A of N attributes of the received record R K+1 ...A 1 ...A N may be input to the trained data representation learning model 120. In step 305, the output of the trained data representation learning model 120 may be received. This output includes the feature vector F K+1 representing the received record R K+1 .

[0067] The set of feature vectors F 1 ...FK with the feature vector F K+1 In step 307, to determine the level of match with the received record R K+1 the feature vector F representing it K+1 may be compared with at least a part of the set of feature vectors F 1 ...F K The comparison may be performed based on the vector space defined by the trained data representation learning model. For example, if the training of the data representation learning model tries to find a semantically important space where related patterns (e.g., records of the same entity) are close to each other, the feature vector F K+1 and the set of feature vectors F 1 ...F K The comparison may be performed by calculating the distance between the feature vectors. This distance may be, for example, the Euclidean distance or the cosine distance.

[0068] In a first example of comparison, the feature vector F K+1 may be compared with each of the set of feature vectors F 1 ...F K This may result in K similarity levels. The highest similarity level among the K similarity levels may be provided as the match level in step 307. This may enable accurate results.

[0069] In a second example of comparison, a cluster of feature vectors closest to the feature vector F K+1 may be identified. The feature vector F K+1 may be compared with each feature vector of the identified cluster that results in multiple similarity levels, and the match level is the highest similarity level among those multiple similarity levels. This may save resources while still providing reliable comparison results.

[0070] In step 309, the feature vector F K+1 or the received record R K+1Alternatively, both of them may be stored based on a matching level. For example, if the matching level is less than a predefined threshold value, the feature vector F K+1 or the received record R K+1 or both of them may be stored, and if not, the feature vector F K+1 or the received record R K+1 need not be stored.

[0071] The method of FIG. 3 may be automatically executed upon receipt of the record R K+1 In one example, the method of FIG. 3 may be executed in real time. For example, the record R K+1 may be received as part of a creation operation or an update operation.

[0072] FIG. 4 is a flowchart of an inference method according to an example of the present subject matter. For the purpose of explanation, the method described in FIG. 4 may be implemented in the system shown in FIG. 1, but is not limited to this implementation. The method of FIG. 4 may be executed, for example, by the data integration system 101.

[0073] The method of FIG. 4 provides an exemplary implementation of steps 203 and 205 of FIG. 2. In particular, the method of FIG. 4 enables the generation of a feature vector F 1 ...R K for each record R i (i = 1, 2,... or K) of the set of records R i to be generated.

[0074] In step 401, the values of each of the N attributes a i ...a 1 ...a N of the record R j (j = 1, 2,... or N) may be input into the relevant one of the attribute-level trained data representation models 121.1 to 121.N of the attribute-level trained data representation models.

[0075] In step 403, an individual feature vector v may be received from each of the attribute-level trained data representation models 121.1 to 121.N. As a result, N individual feature vectors may be obtained for each record R. ij i

[0076] In step 405, to obtain a single feature vector F representing record R, the N individual feature vectors v of record R may be combined. This combination may be performed, for example, as follows: i i i ij where α is the trained weight associated with attribute a, and

Number

Number

[0077] FIG. 5 is a flowchart of a method for training a data representation learning model according to an example of the present subject matter.

[0078] In step 501, a training dataset may be received. The training dataset includes pairs of records (e.g., pairs of similar records). Each pair of records in the training dataset may be associated with a label indicating whether the pair of records represents the same entity or different entities. Each record in the training dataset has a set of attributes a 1 ...a N ​​​​​​​​It includes. For example, the training set may be obtained from one or more sources (e.g., 105).

[0079] In step 503, the data representation learning model may be trained using the training dataset. Thereby, a trained data representation learning model may be obtained. The data representation learning model may be, for example, an autoencoder or a deep neural network.

[0080] In the first training example, training may be performed to find a semantically important vector space in which the feature vectors of related records of the same entity are close to each other. Thereby, it may be possible to maintain the similarity of the original pair between two records within the vector space. Thereby, it may be possible to measure the similarity between feature vectors using distance. In this case, a sham neural network architecture may be advantageously used to train a data representation learning model that is a deep neural network. For example, the deep neural network may be one of the two networks of SiNN.

[0081] In the second training example, training may be performed to find a semantically important vector space in which the feature vectors of related records of the same entity can be identified by using the difference between individual elements of two feature vectors within a predefined range. For example, if this difference is outside this range, it may indicate that the record is not a record of the same entity.

[0082] FIG. 6A is a flowchart of a method for training a data representation learning model according to an example of the present subject matter. For simplicity of explanation, the method of FIG. 6A may be described with reference to the example of FIG. 6B.

[0083] The data representation learning model is a set a of N attributes 1 ...a NIt may include a set of data representation models of N attribute levels respectively associated therewith. Each of the data representation models of the attribute levels may be the neural network systems 611.1 to 611.N shown in FIG. 6B. Each of the neural network systems 611.1 to 611.N may include two neural networks respectively associated with pairs of values of the same attribute. The two networks may share the same weights in the same way as SiNN. In one example, each of the neural network systems 611.1 to 611.N may be a sham neural network. As shown in FIG. 6B, the data representation learning model is structured in a pipeline. Each attribute to be compared (e.g., first name, last name,...) belongs to one pipeline. Each pipeline predicts the similarity of that attribute. Next, these pipelines are weighted.

[0084] J labeled pairs of records [Number] Training of the data representation learning model may be performed using a training data set including m, where m varies between 1 and J. FIG. 6B shows an example of a pair of records [Number] and [Number] Each pair of records includes N pairs of attribute values of a set of N attributes a 1 ...a N Following the example of FIG. 6B, the pair of records [Number] includes N pairs of attribute values, ("Tony", "Tony"), ("Stark", "Starc")... ("NY", "Ohio"), respectively referenced by reference numbers 610.1 to 610.N.

[0085] In step 601, for each of the N pairs 610.1 to 610.N of the attribute values of the current pair of records (e.g.,

Number

[0086] In step 602, in response to receiving the input, each of the data representation models 611.1 to 611.N of the attribute level may output each pair 612.1 to 612.N of individual feature vectors. For example, since the input values are the same, the data representation model 611.1 of the attribute level may output the pair 612.1 of the same individual feature vectors. Since the input values are not the same, the data representation model 611.2 of the attribute level may output the pair 612.2 of different individual feature vectors.

[0087] In step 603, each of the pairs 612.1 to 612.N of individual feature vectors may be weighted by each weight α 1 ... α N . As a result, within each pipeline, v 1 and v 2Different pairs of weighted individual feature vectors with the name may be obtained. These weights increase the importance of one pipeline over another. Depending on the customer's configuration (number and selection of pipelines), the weights may be adjusted for a particular customer.

[0088] In step 604, individual similarity levels 613.1 to 613.N between each pair of weighted individual feature vectors may be calculated. This calculation may be performed, for example, by calculating the distance between two weighted individual feature vectors of each pair. This is shown in FIG. 6B, where the output of each pipeline is the distance ||v 1 -v 2 || 2 which is the individual similarity level quantified by.

[0089] In step 605, an overall measure of the similarity between the current pair of feature vectors of the record may be determined. This determination may be performed, for example, by a combination of weighted individual similarity levels 613.1 to 613.N.

[0090] The overall measure may be determined, for example, using two methods. In the first method, individual vector v1 may be concatenated for all attributes, and individual vector v2 may be concatenated for all attributes. This may result in concatenated vectors V1 and V2. The Euclidean distance (scalar) between the concatenated vectors V1 and V2 may indicate the overall measure. In the second method, individual distances (scalars) between individual vector v1 and individual vector v2 may be calculated for each attribute (as described, for example, in step 604). The sum of the individual distances may indicate the overall measure.

[0091] Vector concatenation may include adding the elements of a vector to a concatenated vector such that the concatenated vector contains the elements of the vector.

[0092] Using the overall measure determined in step 605, the loss function 616 may be evaluated in step 606. Steps 601-606 may be repeated for each pair of at least some records of the training dataset using backpropagation until an optimum value of the loss function is achieved. The term "e" shown in FIG. 6B x may be used to convert a distance to a probability and may be used as part of the loss function. During the training process, in addition to the weights α 1 ... α N of individual vectors, trainable parameters of the attribute-level data representation learning models 611.1 - 611.N are learned. If each of the attribute-level data representation learning models 611.1 - 611.N is a neural network, the neural network may include a group of network weights (e.g., network weights from the input layer to the first hidden layer, network weights from the first hidden layer to the second hidden layer, etc.). Prior to training the attribute-level data representation learning models 611.1 - 611.N, the network weights may be initialized with random numbers or values. Training may be performed to search for optimization parameters (e.g., network weights and biases) of the attribute-level data representation learning models 611.1 - 611.N and minimize the classification error or residual. For example, a training set may be used as input to feed forward through each of the attribute-level data representation learning models 611.1 - 611.N. This may enable the calculation of data loss by the loss function. The data loss may measure the compatibility between the predicted task and the ground truth label. After obtaining the data loss, the data loss may be minimized by changing the weights and biases of each network of the attribute-level data representation learning models 611.1 - 611.N. This minimization may be performed, for example, by backpropagating the loss to all layers and neurons using gradient descent.

[0093] FIG. 7A is a flowchart of a method for training a data representation learning model according to an example of the present subject matter. For simplicity of explanation, the method of FIG. 7A may be described with reference to the example of FIG. 7B.

[0094] In step 701, a trained data representation learning model may be received. For example, the trained data representation learning model obtained from the method of FIG. 6A may be received in step 701. The trained data representation learning model of FIG. 6A may be centrally generated by a service provider and used by various clients. In one example, the trained data representation learning model may be used without modification in a client system or may be updated as described using FIGS. 7A-7B. For example, a user may need to use one or more additional attributes that were not used to generate the trained data representation learning model. As shown in FIG. 7B, the user may need to add an additional attribute a N+1 which is an employee ID attribute. To that end, in step 703, a user-specific data representation learning model may be created. The user-specific data representation learning model may include the trained attribute-level trained data representation models 611.1-611.N and one additional attribute-level trained data representation model 611.N+1 associated with the additional attribute a N+1 . For example, a user who wants to add their unique employee ID to the matching process may add a new pipeline to the structure of the received data representation model. The trained parameters of the trained attribute-level trained data representation models 611.1-611.N are frozen during the training of the user-specific data representation learning model. This freezing may be advantageous because the client can add a new pipeline for one additional attribute and only need to train the network for that one attribute instead of retraining the entire system. The received trained weights α 1 ...αN is not frozen and may be retrained. Additionally, an additional weight α N+1 is associated with an additional pipeline. Thus, a user-specific data representation learning model includes the trainable parameters of the additional attribute-level trained data representation model 611.N+1 and the N+1 trainable weights α 1 ...α N+1 . The trained parameters of the trained attribute-level trained data representation models 611.1 to 611.N are frozen during the training of the user-specific data representation learning model and need not be changed.

[0095] In step 705, a user-specific data representation learning model may be trained using a training set. This training set includes pairs of records and associated labels, and each record of the training set has N+1 attributes a 1 ...a N+1 . The training of the user-specific data representation learning model may be performed as described with reference to FIG. 6A. The parameters of the trained attribute-level trained data representation models 611.1 to 611.N are fixed during training and not changed.

[0096] FIG. 8 is a diagram illustrating a method for storing feature vectors according to an example of the present subject matter.

[0097] A set of records 801.1 to 801.4 may be provided. Each of the records 801.1 to 801.4 has a set of attributes a 1 ...a N for generating a trained data representation learning model 803, for example, as described with reference to FIG. 6A. To generate a feature vector 804 (named SimVec or similarity vector) representing the records, the set of attributes a 1 ...a NThe value of 1 ...a N may be input into the trained data representation learning model 803. For each set of attributes a 1 ...α N The value of is input into the trained machine learning model for each attribute level of the trained data representation learning model 803. For each input record, an individual feature vector is generated using an individual pipeline. The trained weights α

[0098] Figure 9 is a diagram illustrating a prediction process according to an example of the present subject matter. For this purpose, a trained data representation learning model 903 may be provided as described with reference to FIG. 6A, for example. A set of attributes a 1 ...a N including may be provided in a record 901. To generate a feature vector 904 representing the record 901, the set of attributes a of the record 901 1 ...a NThe value of may be input into the trained data representation learning model 903. A set of attributes a 1 ...a N The values of are input into the trained machine learning models at each attribute level of the trained data representation learning model 903. Individual feature vectors are generated using individual pipelines. Trained weights α 1 ...α NIndividual feature vectors are weighted using [weighting method]. To create the feature vector 904 (or similarity vector), the weighted individual feature vectors are concatenated or combined. To find the nearest cluster, the feature vector 904 is compared to the centroid 908 of the cluster 907 created, for example, with reference to FIG. 8. To do so, the distance 910 between each of the feature vector 904 and the centroid 908 may be calculated, and the cluster associated with the minimum calculated distance may be the cluster associated with the feature vector 904. After identifying the cluster associated with the feature vector 904, a potential match between the feature vector 904 and each feature vector 912 of the cluster associated with the feature vector 904 may be determined. To do so, a match metric 914 between each of the feature vector 904 and the feature vector 912 may be calculated. Thereby, a match level between the feature vector and the stored feature vectors may be obtained. This match level may be, for example, the minimum calculated value of the metric 914. If the match level is higher than a threshold, this may indicate that the record 901 has a matching or duplicate stored record. Next, a deduplication system constructed according to the present invention may merge these records because the records represent the same entity. Merging of records is an operation that can be performed in various ways. For example, merging two records may include creating a golden record as a replacement for the records that appear to be similar and have been detected as overlapping with each other. This merge is known as data fusion, or physical folding with any of the records, or survivor rights at the attribute level. If the match level is below the threshold, this indicates that the record 901 does not match any record in the cluster and may therefore be stored.

[0099] FIG. 10 is a flowchart of a method 1000 for collating records according to an example of the present subject matter. The method 1000 includes a machine learning stage 1001 and an application stage 1002.

[0100] A neural network system may be trained (1003) such that a trained neural network system can generate a feature vector (SimVec) representing a data record. The trained neural network system 1006 may be used to generate feature vectors for all records of the MDM database (1004). A clustering algorithm may be trained on the generated feature vectors (1005). Thereby, a cluster centroid 1007 of the clusters of the generated feature vectors may be obtained. Steps 1003, 1004, 1005, and 1007 are part of the machine learning stage 1001.

[0101] During the application stage 1002, after feature vectors are generated for all records of the MDM database, a request to add a new record may be received (1008). The trained neural network system 1006 may be used to generate a feature vector for the received record (1009). The cluster centroid 1007 may be used to determine the centroid closest to the generated feature vector of the received record (1010). The MDM database may be queried regarding the generated feature vectors belonging to the cluster having the closest centroid (1011). The queried feature vectors may be compared with the generated feature vector of the received record (1012). This comparison may be performed by calculating the distance between the compared feature vectors. If the distance between two compared feature vectors is less than a threshold value (1013), this indicates that these feature vectors are duplicates and the corresponding records may be merged (1014). If the distance between all pairs of compared feature vectors is greater than or equal to the threshold value (1013), this indicates that these feature vectors are not duplicates and the received record may be stored in the MDM database.

[0102] Figure 11 is a diagram of a system 1100 for matching records according to an example of the present subject matter. An external system (e.g., a customer database) 1101 may provide an MDM including new record entries (e.g., people, organizations). These entries are sent to the MDM backend 1102, and the MDM backend 1102 sends a request to the ML service 1104 to generate a similarity vector 1111 from the new entry 1110. Subsequently, this newly created vector 1111 may be used to find its corresponding cluster using the already trained clusters 1108 encrypted and stored inside the database 1103. The center 1107 of the clusters stored in the database 1103 may be used to find the corresponding clusters. The backend 1102 that has obtained the clusters may then query all the existing vectors inside that cluster to find possible matches with the received entry. However, the feature vectors stored in the database 1103 may be encrypted using homomorphic encryption such as Paillier encryption. Therefore, in order to find a match between the feature vector 1111 of the received entry 1110 and the encrypted feature vector 1112 of the cluster, the feature vector 1111 of the received entry 1110 is not encrypted and can be compared with the encrypted feature vector 1112 in an unencrypted form. For this purpose, the Euclidean distance may be reconstructed so as to be able to benefit from homomorphic encryption. The reconstructed Euclidean distance may be as follows.

[0103]

Number

[0104] This distance consists of three terms

Number

Number

[0105] The reconstructed distance between the feature vector 1111 and each encrypted feature vector 1112 of the cluster may be calculated. Thereby, the encrypted similarity 1113 may be obtained. To obtain each decrypted similarity 1114, the encrypted similarity may be decrypted. Each of the decrypted similarities 1114 may be compared to a threshold to determine whether there is a match between the new entry 1110 and any of the stored entries in the database 1103. If no match is detected, the generated feature vector 1111 of the new entry may first be encrypted and then stored in the database 1103.

[0106] FIG. 12 depicts a general computerized system 1600 (e.g., a data integration system) suitable for implementing at least some of the steps of the method involved in the present disclosure.

[0107] It will be understood that the methods described herein are at least partially non-interactive and are automated by a computerized system such as a server or an embedded system. However, in exemplary embodiments, the methods described herein can be implemented in a (partially) interactive system. These methods can be further implemented in software 1612, 1622 (including firmware 1622), hardware (processor) 605, or combinations thereof. In exemplary embodiments, the methods described herein are implemented in software as an executable program and are executed by a dedicated or general-purpose digital computer such as a personal computer, a workstation, a minicomputer, or a mainframe computer. Accordingly, the most common system 600 includes a general-purpose computer 601.

[0108] In exemplary embodiments, with respect to the hardware architecture, as shown in FIG. 6, the computer 601 includes a processor 605, a memory (main memory) 1610 coupled to a memory controller 1615, and one or more input and / or output (I / O) devices (or peripherals) 10, 1645 communicatively coupled via a local input / output controller 1635. The input / output controller 1635 can be one or more buses or other wired or wireless connections known in the art, but is not limited thereto. The input / output controller 1635 may include additional elements such as controllers, buffers (caches), drivers, repeaters, and receivers for enabling communication, which are omitted for simplicity. Further, the local interface may include address, control, or data connections, or combinations thereof, to enable proper communication between the aforementioned components. As described herein, the I / O devices 10, 1645 may typically include any generalized cryptographic card or smart card known in the art.

[0109] Processor 1605 is a hardware device for executing software, particularly the software stored in memory 1610. Processor 1605 can be a custom-made or commercially available processor, a central processing unit (CPU), an auxiliary processor among a plurality of processors associated with computer 1601, a semiconductor-based microprocessor (in the form of a microchip or chip set), a macroprocessor, or generally any device for executing software instructions.

[0110] Memory 1610 can include any one or a combination of volatile memory elements (e.g., random access memory (RAM) such as DRAM, SRAM, SDRAM, etc.) and non-volatile memory elements (e.g., ROM, erasable programmable read only memory (EPROM), electronically erasable programmable read only memory (EEPROM), programmable read only memory (PROM)). Note that memory 1610 can include a distributed architecture where various components are located far from each other but can be accessed by processor 1605.

[0111] The software in memory 1610 may include one or more separate programs, each of which includes an ordered list of executable instructions for implementing a logical function, particularly a function involved in an embodiment of the present invention. In the example of FIG. 6, the software in memory 1610 includes instruction 1612 (e.g., an instruction for managing a database such as a database management system).

[0112] The software in the memory 1610 should usually also include a suitable operating system (OS). The OS 1611 basically controls the execution of other computer programs, such as the software 1612 for implementing the method as described herein in some cases.

[0113] The method described herein may be in the form of any other entity including the source program 1612, the executable program 1612 (object code), the script, or a set of instructions to be executed 1612. In the case of the source program, the program may need to be converted via a compiler, an assembler, an interpreter, etc., which may or may not be included in the memory 1610 to operate properly in connection with the OS 1611. Further, the method can be described as an object-oriented programming language including classes of data and methods, or a procedural programming language including routines, subroutines, or functions, or a combination thereof.

[0114] In an example embodiment, a conventional keyboard 1650 and mouse 1655 can be coupled to an input / output controller 1635. Other output devices such as I / O device 1645 may include, for example, input devices such as printers, scanners, microphones, etc., but are not limited thereto. Finally, I / O devices 10, 1645 may further include devices that communicate with both input and output, such as, for example, a network interface card (NIC), a modem / demodulator (for accessing other files, devices, systems, or networks), a radio frequency (RF) or other transceiver, a telephone interface, a bridge, a router, etc., but are not limited thereto. I / O devices 10, 1645 can be any generalized cryptographic card or smart card known in the art. System 1600 can further include a display controller 1625 coupled to a display 1630. In an example embodiment, system 1600 can further include a network interface for coupling to network 1665. Network 1665 can be an IP-based network for communication via a broadband connection between computer 1601 and any external server, client, etc. Network 1665 can transmit and receive data between computer 1601, which can be involved in performing some or all of the steps of the methods described herein, and external system 30. In an example embodiment, network 1665 can be an administrative IP network managed by a service provider. Network 1665 may be implemented wirelessly, for example, using wireless protocols and wireless technologies such as WiFi(R), WiMax(R), etc. Network 1665 can also be a packet-switched network, such as a local area network, a wide area network, a metropolitan area network, an Internet network, or other similar types of network environments.Network 1665 may be a fixed wireless network, a wireless local area network (LAN), a wireless wide area network (WAN), a personal area network (PAN), a virtual private network (VPN), an intranet, or other suitable network system, and includes devices for transmitting and receiving signals.

[0115] When computer 1601 is a PC, a workstation, an intelligent device, etc., the software in memory 1610 may further include a basic input output system (BIOS) 1622. The BIOS is a set of basic software routines that initialize and test the hardware at startup, start up OS 1611, and support data transfer between hardware devices. The BIOS is stored in the ROM so that it can be executed when computer 1601 is started.

[0116] When computer 1601 is operating, processor 1605 is configured to execute software 1612 stored in memory 1610, communicate data with memory 1610, and overall control the operation of computer 1601 according to the software. The methods and OS 1611 described herein are read, in whole or in part (usually the latter), by processor 1605 and, in some cases, buffered within processor 1605 before being executed.

[0117] The systems and methods described herein are implemented in software 1612 as shown in FIG. 6, and the methods can be stored on any computer-readable medium, such as storage 1620, for use by or in connection with any computer-related system or method. Storage 1620 may include disk storage, such as HDD storage.

[0118] The subject matter includes the following clauses.

[0119] Clause 1: A computer-implemented method comprising: providing a set of one or more records, each record of the set of records including a set of one or more attributes; inputting the values of the set of attributes of the set of records into a trained data representation learning model, thereby receiving, as an output of the trained data representation model, a set of feature vectors representing each of the set of records; storing the set of feature vectors; and a computer-implemented method.

[0120] Clause 2: receiving a further record including a set of attributes; inputting the values of the set of attributes of the received further record into a trained data representation learning model, thereby obtaining, from the trained data representation learning model, a feature vector of the received further record; comparing the obtained feature vector with at least a portion of the set of feature vectors to determine a level of match of the obtained feature vector with the set of feature vectors; storing the obtained feature vector or the received further record or both based on the level of match; and further comprising the method according to clause 1.

[0121] Item 3: The method according to claim 1 or 2, wherein storing the set of feature vectors includes clustering the set of feature vectors into clusters and associating each of the stored feature vectors with cluster information indicating the corresponding cluster.

[0122] Item 4: The method according to claim 1 or 2, wherein storing the set of feature vectors includes clustering the set of feature vectors into clusters of feature vectors, and the method further includes determining the distance between the acquired feature vectors and the vectors representing each of the clusters, and at least a part of the set of feature vectors includes the clusters represented by the vectors having the closest distance to the acquired feature vectors.

[0123] Item 5: The method according to any one of claims 2 to 4, which is executed in real time and a record is received as part of a creation operation or an update operation.

[0124] Item 6: The method according to any one of claims 1 to 5, wherein outputting each feature vector of the set of feature vectors includes generating individual feature vectors for each attribute of the set of attributes and combining the individual feature vectors to obtain the feature vectors.

[0125] Item 7: The method according to any one of claims 1 to 6, wherein the trained data representation learning model is configured to process the input values in parallel.

[0126] Item 8: The trained data representation learning model includes a set of attribute-level trained data representation models, each of the set of attribute-level trained data representation models is associated with each attribute of the set of attributes, and the output of each feature vector of the set of feature vectors is inputting the value of each attribute of the set of attributes into the associated attribute-level trained data representation model; In response to this input, receiving an individual feature vector from each of the attribute-level trained data representation models, and combining the individual feature vectors to obtain the feature vector, and The method according to any one of claims 1 to 7, comprising:

[0127] Claim 9: The trained data representation learning model further includes a set of trained weights, each weight of the set of weights being associated with each attribute of the set of attributes, and combining includes weighting each of the individual feature vectors using each of the trained weights of the set of trained weights. The method according to claim 8.

[0128] Claim 10: The method according to claim 8 or 9, wherein each of the attribute-level trained data representation models is a neural network.

[0129] Claim 11: The trained data representation learning model is obtained from training to optimize a loss function, the loss function being a measure of similarity between the feature vectors of pairs of records, the measure of similarity being a combination of individual similarities, and each individual similarity of the individual similarities indicating the similarity between two individual feature vectors generated for the same attribute within the pair of records. The method according to any one of claims 8 to 10.

[0130] Claim 12: The method according to any one of claims 1 to 11, wherein the trained data representation learning model is trained to optimize a loss function, the loss function being a measure of similarity between the feature vectors of pairs of records.

[0131] Claim 13: The method according to any one of claims 1 to 12, wherein the trained data representation learning model includes at least one neural network trained according to a Sham neural network architecture.

[0132] Claim 14: The trained data representation learning model includes one trained neural network for each attribute of the set of attributes, and the output of each feature vector of the set of feature vectors is inputting the value of each attribute of the set of attributes into the trained neural network associated with the attribute, in response to this input, receiving an individual feature vector from each of the trained neural networks, combining the individual feature vectors to obtain the feature vector The method according to claim 13, comprising:

[0133] Claim 15: Receiving a training data set including pairs of similar records including a set of attributes, training a data representation learning model using the training data set, thereby generating a trained data representation learning model The method according to any one of claims 1 to 14, further comprising:

[0134] Claim 16: The data representation learning model includes a set of attribute-level trained data representation models respectively associated with a set of attributes, and the training of the data representation learning model is performed for each pair of similar records for each attribute of the set of attributes inputting the pair of values of the attributes in the pair of records into the corresponding attribute-level trained data representation model, thereby obtaining a pair of individual feature vectors, calculating an individual similarity level between the pair of individual feature vectors, weighting the individual similarity levels using trainable weights of the attributes, determining a measure of similarity between the feature vectors of the pair of records as a combination of the weighted individual similarity levels, evaluating a loss function using this measure for training The method according to claim 15, comprising:

[0135] Item 17: The set of attributes includes a first subset of attributes and a second subset of attributes, and the method receiving a first trained data representation learning model that includes a first subset of trained data representation learning models of attribute levels respectively associated with the first subset of attributes, wherein the first trained data representation learning model is configured to receive values of the first subset of attributes of a record, input the feature vector of the record into the associated attribute level trained data representation models of the first subset of the first subset of attribute level trained data representation learning models, receive individual feature vectors from each of the attribute level trained data representation models of the first subset of the first subset of attribute level trained data representation learning models, and output by combining the individual feature vectors to obtain the feature vector, and providing, for the second subset of attributes, a second subset of attribute level data representation learning models, wherein each attribute level data representation learning model of the second subset is configured to generate a feature vector for each attribute of the second subset of attributes, and creating a data representation learning model that includes the first trained data representation learning model and the second subset of attribute level data representation learning models, and training the data representation learning model to thereby generate a trained data representation learning model The method according to any one of Items 1 to 16, further comprising.

[0136] Claim 18: The first trained data representation learning model further includes a first subset of trained weights associated with a first subset of attributes, and combining is performed using the first subset of trained weights, and the created data representation learning model further includes a second subset of trainable weights associated with a second subset of attributes, The trained data representation learning model, for a feature vector of a set of feature vectors, inputting the value of each attribute of the set of attributes into the relevant attribute-level trained data representation models of the first and second subsets of the attribute-level trained data representation learning model; in response to this input, receiving individual feature vectors from each of the attribute-level trained data representation models of the first and second subsets of the attribute-level trained data representation learning model; combining the individual feature vectors using each weight of the first and second subsets of weights to obtain the feature vectors The method according to claim 17, configured to output by.

[0137] The present invention may be a system, method, or computer program product, or a combination thereof, at any possible integrated technical detail level. The computer program product may include a computer-readable storage medium containing computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0138] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy (R) disk, mechanically encoded devices such as punch cards or raised structures in grooves in which instructions are recorded, and any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted via a wire.

[0139] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). This network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers them for storage on a computer-readable storage medium within each computing / processing device.

[0140] Computer-readable program instructions for executing the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk(R), C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially executed on the user's computer as a stand-alone software package, partially executed on the user's computer and a remote computer respectively, or executed entirely on the remote computer or a server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to execute aspects of the present invention, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions for personalizing the electronic circuit by utilizing the state information of the computer-readable program instructions.

[0141] Aspects of the invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0142] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may be stored in a computer-readable storage medium that includes instructions that when executed cause a computer, programmable data processing apparatus, or other device to function in a particular manner, such that the computer-readable storage medium forms a product including instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0143] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0144] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be implemented as one step according to the functionality involved, may be executed at the same time, substantially simultaneously, in a partially or fully temporally overlapping manner, or, in some cases, in the reverse order. It should also be noted that each block of the block diagrams or flowchart diagrams, or both, and combinations of blocks in the block diagrams or flowchart diagrams, or both, can be implemented by a dedicated hardware-based system that performs the specified functions or operations or a combination of dedicated hardware and computer instructions.

Claims

1. A computer-implemented method comprising: providing a set of one or more records, each record of the set of records including a set of one or more attributes, the set of attributes including a first subset of attributes and a second subset of attributes; receiving a first trained data representation learning model including a first subset of attribute-level trained data representation learning models respectively associated with the first subset of attributes, the first trained data representation learning model being configured to receive values of the first subset of attributes of a record, input the values of each attribute of the first subset of attributes into the associated attribute-level trained data representation learning model of the first subset of the attribute-level trained data representation learning models to obtain a feature vector of the record, receive individual feature vectors from each of the attribute-level trained data representation learning models of the first subset of the attribute-level trained data representation learning models, and output the feature vector by combining the individual feature vectors; providing a second subset of attribute-level data representation learning models for the second subset of attributes, each attribute-level data representation learning model of the second subset being configured to generate a feature vector for each attribute of the second subset of attributes; creating a data representation learning model including the first trained data representation learning model and the second subset of attribute-level data representation learning models; training the data representation learning model to generate a trained data representation learning model; inputting values of a set of one or more attributes in a record into the generated trained data representation learning model to receive, as an output of the trained data representation learning model, a set of feature vectors respectively representing the set of records; storing the received set of feature vectors in a storage; and a computer-implemented method.

2. Receiving additional records that include the set of attributes; Inputting the values of the set of attributes of the received additional records into the trained data representation learning model, thereby obtaining, from the trained data representation learning model, a feature vector of the received additional records; Comparing the obtained feature vector with at least a part of the set of feature vectors to determine a matching level of the obtained feature vector with the set of feature vectors; Storing the obtained feature vector, or the received additional record, or both, when the determined matching level is less than a predefined threshold; The method according to claim 1, further comprising.

3. The storing of the set of feature vectors includes clustering the set of feature vectors into clusters and associating each of the stored feature vectors with cluster information indicating a corresponding cluster. The method according to claim 1 or 2.

4. The storing of the set of feature vectors includes clustering the set of feature vectors into clusters of feature vectors. The method further includes determining a distance between the obtained feature vector and a vector representing each cluster of the clusters. At least a part of the set of feature vectors includes the cluster represented by the vector having the closest distance to the obtained feature vector. The method according to any one of claim 2 or claim 3 when dependent on claim 2.

5. In response to receiving additional records that include the set of attributes, the obtaining, the determining, and the storing according to claim 2 are performed in real time. The method according to claim 2.

6. Outputting each feature vector of the set of feature vectors by the trained data representation learning model includes generating individual feature vectors for each attribute of the set of attributes and combining the individual feature vectors to obtain the feature vector. The method according to any one of claims 1 to 5.

7. The method according to any one of claims 1 to 6, wherein the trained data representation learning model is configured to process the input values in parallel.

8. The trained data representation learning model includes a set of attribute-level trained data representation learning models, each of the set of attribute-level trained data representation learning models being associated with each attribute of the set of attributes, and the output of each feature vector of the set of feature vectors being inputting the values of each attribute of the set of attributes into the associated attribute-level trained data representation learning model; receiving, in response to the input, an individual feature vector from each of the attribute-level trained data representation learning models; combining the individual feature vectors to obtain the feature vector The method according to any one of claims 1 to 7, comprising:

9. The trained data representation learning model further includes a set of trained weights, each weight of the set of weights being associated with each attribute of the set of attributes, and the combining includes weighting each of the individual feature vectors using each of the trained weights of the set of trained weights. The method according to any one of claims 1 to 8.

10. The method according to any one of claims 1 to 9, wherein each of the attribute-level trained data representation learning models is a neural network.

11. The trained data representation learning model is trained to optimize a loss function, the loss function being a measure of similarity between feature vectors of pairs of records, the measure of similarity being a combination of individual similarities, and each individual similarity of the individual similarities indicating a similarity between two of the individual feature vectors generated for the same attribute within the pair of records. The method according to any one of claims 1 to 10.

12. To optimize the loss function, the trained data representation learning model is trained using backpropagation, the loss function being a measure of similarity between feature vectors of pairs of records, the measure of similarity being a combination of individual similarities, each individual similarity of the plurality of individual similarities indicating the similarity of two individual feature vectors generated with respect to the same attribute within a pair of records. The method according to any one of claims 1 to 11.

13. The method according to any one of claims 1 to 12, wherein the trained data representation learning model includes at least one neural network trained according to a sham neural network architecture.

14. The trained data representation learning model includes one trained neural network for each attribute of the set of attributes, and the output of each feature vector of the set of feature vectors is inputting the value of each attribute of the set of attributes into the associated trained neural network; receiving an individual feature vector from each of the trained neural networks in response to the input; combining the individual feature vectors to obtain the feature vector The method according to any one of claims 1 to 13.

15. receiving a training dataset including pairs of records including the set of attributes; training a data representation learning model using the received training dataset, thereby generating the trained data representation learning model; further comprising The method according to claim 12, wherein the generated trained data representation learning model is used in the training using the backpropagation.

16. The data representation learning model includes a set of attribute-level data representation learning models respectively associated with the set of attributes, and the training of the data representation learning model is, for each pair of records, for each attribute of the set of attributes, inputting the pair of values of the attributes within the pair of records into the corresponding attribute-level data representation learning model, thereby obtaining a pair of individual feature vectors; calculating an individual similarity level between the pair of individual feature vectors; weighting the individual similarity levels using the trainable weights of the attributes; determining a measure of similarity between the feature vectors of the pair of records as a combination of the weighted individual similarity levels; evaluating a loss function using the measure for the training; The method according to any one of claims 1 to 15, comprising: **Claim 17** The first trained data representation learning model further includes a first subset of trained weights associated with a first subset of the attributes, the combining is performed using the first subset of trained weights, and the created data representation learning model further includes a second subset of trainable weights associated with a second subset of the attributes; The trained data representation learning model transforms the feature vectors of the set of feature vectors; inputting the values of each of the attributes of the set of attributes into the associated attribute-level trained data representation learning models of the first and second subsets of the attribute-level trained data representation learning models; in response to the input, receiving individual feature vectors from each of the first and second subsets of the attribute-level trained data representation learning models of the attribute-level trained data representation learning models; combining the individual feature vectors using each of the first and second subsets of the weights to obtain the feature vectors; The method according to any one of claims 1 to 16, configured to output by: **Claim 18** The method according to any one of claims 1 to 17, wherein the received set of feature vectors is used instead of the set of one or more records. **Claim 19** A computer program comprising instructions for performing the method according to any one of claims 1 to 18. **Claim 20** A computer system, comprising: a memory; a central processing unit connected to the memory; and is provided with: causing the central processing unit to execute a computer program including instructions for performing the method according to any one of claims 1 to 18; A computer system.

Citation Information

Patent Citations

  • Generating vector from data

    CN110993091A

  • Determination device, determination method, and determination program

    JP2018045505A

  • Learning program and learning method

    JP2019185244A