A method for entity alignment of industrial product master data based on BERT model

By adopting the sequence labeling method of BERT model and BiLSTM-CRF in the master data entity alignment of industrial products, combined with the BERT-CNN model, the problem of insufficient entity alignment accuracy and efficiency in the prior art is solved, and a higher accuracy entity alignment effect is achieved.

CN118839247BActive Publication Date: 2025-06-06NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410884041.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-03
Publication Date
2025-06-06
Estimated Expiration
2044-07-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture semantic information when realizing the alignment of multimodal master data entities of industrial products, especially when facing heterogeneous data sources and cross-language scenarios, accuracy and efficiency are insufficient.

Method used

The BERT model is used to combine BiLSTM-CRF with sequence labeling method to perform entity naming and recognition, and the entity relationship triplet is extracted through BERT. Further, the BERT-CNN model is proposed. By combining the 1D convolution layer on the BERT model fine-tuned by triple data, it accurately captures local features in the text sequence and improves the accuracy of entity alignment.

Benefits of technology

Through the combination of BiLSTM-CRF and BERT models, the automated recognition of entity naming and accurate capture of semantic relationships are realized, and the accuracy and efficiency of entity alignment are improved. The BERT-CNN model can capture local features in text sequences more finely, further improving the effect of entity alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118839247B_ABST
    Figure CN118839247B_ABST
Patent Text Reader

Abstract

The present invention discloses an entity alignment method for industrial product master data based on a BERT model, including: encoding, feature extraction and decoding text data of industrial product master data to be aligned based on BERT-BiLSTM-CRF to obtain entity recognition results; splicing two types of entities and texts that may be related and inputting them into a BERT model for classification to obtain triple extraction results; combining the triple extraction results with the label set of the original entity text to obtain a labeled triple data set, and fine-tuning the BERT model; building a BERT-CNN entity alignment model for triple data; inputting triple data into the BERT-CNN entity alignment model to obtain entity alignment probability, and realizing prediction of entity alignment results. The present invention can realize entity alignment with higher accuracy for triple data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence entity alignment, and specifically relates to an industrial product master data entity alignment method based on a BERT model. Background Art

[0002] The systems within modern enterprises are diverse and large in scale. In the production process of industrial products, each system usually processes different data. At the same time, due to problems such as data silos, access rights, technology and platform heterogeneity, there are obstacles to the flow of data between different systems. Master data refers to the most important and stable data required for business operations and business decisions. It is the benchmark data in the business system. Master data management aims to maintain the consistency, accuracy, completeness and controllability of master data to support the business processes and decisions of the enterprise. Through reasonable master data management, enterprises can better respond to business challenges, ensure data quality, and provide reliable benchmark data for business systems.

[0003] Traditional master data focuses on text data, but due to the business needs of industrial production, data management, and other situations, other types of data, such as images (workpiece drawings), time series data, and videos, are also considered as master data. Compared with traditional single-text master data, managing multimodal master data of industrial products is more challenging. One possible solution is to fuse master data from different systems, and the key step to achieve this solution is to build a knowledge graph for the multimodal industrial product master data of the enterprise. In the process of building a knowledge graph, the core task is entity alignment. The knowledge graph is a structured data model that represents knowledge in the form of a graph. Its core is to represent things in the real world and their relationships through nodes (entities) and edges (relationships). The basic unit of the knowledge graph is triple data used to represent entities and relationships.

[0004] Entity alignment faces many challenges, including differences in heterogeneous data sources, various changes in entity names, incomplete and inconsistent entity attributes, etc. Traditional entity alignment technology mainly focuses on the syntactic and structural levels, especially early entity alignment and mapping technology, the core of which is to calculate the similarity between labels and characters between entities. This method has certain limitations. For example, it is unable to capture information at the semantic level, and it is also powerless when faced with entity alignment problems in literal heterogeneity or cross-language scenarios. For the scenario of master data fusion of system data within modern enterprises, achieving efficient and concise entity alignment is challenging for traditional rules or feature methods.

[0005] As a pre-trained model with strong semantic understanding capabilities, the BERT model understands the semantics of words and phrases through bidirectional context information, so it can more accurately capture its semantic information when processing entity names. At the same time, through supervised fine-tuning, the BERT model can better adapt to the data characteristics of triple data and improve the accuracy of alignment. Summary of the invention

[0006] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned prior art and provide an industrial product master data entity alignment method based on the BERT model. The BiLSTM-CRF sequence labeling method is used to realize entity naming recognition, and entity relationship triples are extracted through BERT. In addition, a BERT-CNN model is proposed. The BERT model fine-tuned on the triple data is combined with a 1D convolutional layer to accurately capture the local features in the text sequence, thereby achieving entity alignment of the triple data with higher accuracy.

[0007] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:

[0008] A method for aligning industrial product master data entities based on a BERT model, comprising:

[0009] Step 1. Encode, extract features and decode the text data of the industrial product master data to be aligned based on BERT-BiLSTM-CRF to obtain entity recognition results;

[0010] Step 2: Combine entity recognition results with relationship labels, concatenate two types of entities and text that may be related, and input them into the BERT model for classification to obtain triple extraction results;

[0011] Step 3. Combine the triple extraction results with the label set of the original entity text to obtain a labeled triple dataset, and use this dataset to fine-tune the BERT model;

[0012] Step 4. Based on the fine-tuned BERT model and CNN network, build a BERT-CNN entity alignment model for triple data;

[0013] Step 5. Input the triplet data into the BERT-CNN entity alignment model to obtain the entity alignment probability and predict the entity alignment result.

[0014] To optimize the above technical solutions, the specific measures taken also include:

[0015] In step 1 above, the pre-trained BERT model is first used to convert the input sequence into a semantically rich word vector. Secondly, BiLSTM is used to extract features from the vector representation obtained by the BERT model. Finally, based on the output of BiLSTM, CRF is used for decoding to determine the final label of each word.

[0016] In step 3 above, the fine-tuning training process of the BERT model is as follows:

[0017] Step 3.1. Tokenize the attribute pairs in the triple data set, then concatenate them with "-", set a [CLS] symbol at the beginning of the sentence, and set a [SEP] symbol at the end of the sentence. Convert all data samples into a token index array with a fixed length of 128. The [CLS] and [SEP] indexes are set to 101 and 102. If the sample is less than 128, fill it with 0;

[0018] Step 3.2. Map each token in the token index array to the pre-trained word vector space to generate word vector embeddings to capture the semantic and contextual information between tokens; assign segmentation embeddings to each token to identify the tag of the sentence to which the token belongs; assign position embeddings to each token to indicate the position information of the token in the sentence; concatenate the word vector embeddings, segmentation embeddings, and position embeddings as input data for the BERT model;

[0019] Step 3.3. Input the input data into the BERT model. Through its internal 12-layer Transformer structure, it deeply learns the semantic expression of the data structure text and encodes the contextual semantic information to obtain the predicted label, which contains the output vector of the tag [CLS] and the output vector of other characters. The output vectors all carry their own unique semantic information.

[0020] Step 3.4. Use the cross entropy loss function to calculate the difference between the predicted label and the actual label; through iterative fine-tuning, continuously improve the model's ability to extract abstract semantic features.

[0021] In step 3.4 above, the cross entropy loss function is:

[0022]

[0023] Where N is the number of samples, y i represents the true label of sample i, It represents the probability that the model's predicted label for sample i belongs to the positive class.

[0024] In step 4 above, the process of building a BERT-CNN entity alignment model for triple data is as follows:

[0025] Step 4.1. For the triple data structure, distinguish the entity name and attribute pairs therein;

[0026] Step 4.2. For entity names, use the fine-tuned BERT model to generate word vector embeddings for the entity names; calculate the cosine similarity between word vector embeddings to measure the similarity between entities; adjust the binary output layer matrix of the BERT model according to the set threshold;

[0027] Step 4.3. For attribute pairs, based on the fine-tuned BERT model, a Dropout layer is introduced to concatenate the output of the BERT model, and L2 regularization is added to alleviate potential overfitting problems;

[0028] Step 4.4. Introduce a 1D convolutional layer, and use the output in step 4.3 to learn n-gram-level features at different positions through the 1D convolutional layer; then use AdaptiveMaxPool1d for pooling, and map the features learned by the convolutional layer to an output of a fixed size as the output result of the convolutional layer;

[0029] Step 4.5. Introduce the fully connected layer and softmax function to further process the output results of the convolutional layer: first introduce the Dropout layer to connect the output results of the convolutional layer, then input the output results into the fully connected layer for linear transformation, and finally use the softmax function for normalization to obtain the probability of each category, thereby realizing the prediction of the semantic relationship between the two text pairs.

[0030] In step 4.3 above, the loss function of L2 regularization is:

[0031]

[0032] Where L is the total loss function including L2 regularization, N is the number of samples, and x i is the input feature of the i-th sample, y i is the true label of the i-th sample, θ is the set of weight parameters of the model, θ j is the jth weight parameter in the model; L data (y i ,f(x i ; θ)) is the data loss function, which is used to measure the true value y i and the model prediction value f(x i ; θ); λ is a hyperparameter of the regularization strength.

[0033] In step 4 above, the total loss function of the BERT-CNN entity alignment model is:

[0034]

[0035] Among them, L 1 is the cross entropy loss, which ensures that the model can correctly predict the labels in the binary classification problem. λ is the weight decay coefficient of L2 regularization. A larger λ will impose a greater penalty on the model parameters, making the parameters smaller and more conducive to preventing overfitting; a smaller λ will focus more on data fitting and may lead to overfitting.

[0036] A computer device includes a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, the steps of the industrial product master data entity alignment method based on the BERT model are implemented.

[0037] A computer-readable storage medium stores a program, which, when executed by a processor, implements the steps of the industrial product master data entity alignment method based on the BERT model.

[0038] The present invention has the following beneficial effects:

[0039] The present invention realizes automatic recognition of entity naming through the sequence annotation method of BiLSTM-CRF; fine-tunes the pre-trained BERT model to better adapt to the data of triple structure, so as to accurately capture the semantic relationship between entities and improve the accuracy and efficiency of entity alignment; introduces CNN (convolutional neural network) and proposes BERT-CNN model, which can learn n-gram level features at different positions, so as to more finely capture local features in text sequences. The present invention integrates various semantic information through the BERT model and improves the accuracy of entity alignment tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 The figure is a schematic diagram of the overall process of the industrial product master data entity alignment method based on the BERT model according to an embodiment of the present invention.

[0041] Figure 2 This is a schematic diagram of the industrial product master data entity alignment model framework based on the BERT model built in an embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0043] Although the steps in the present invention are arranged with numbers, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It is understood that the term "and / or" used in this article involves and covers any and all possible combinations of one or more of the associated listed items.

[0044] Example 1

[0045] The method for aligning industrial product master data entities based on the BERT model of the present invention comprises the following steps:

[0046] Step 1. Encode, extract features and decode the text data of the industrial product master data to be aligned based on BERT-BiLSTM-CRF to obtain entity recognition results;

[0047] In this step, the BERT model is used to extract features from the text description to be aligned to obtain vector representation; then the BiLSTM module is trained to perform bidirectional encoding on the text to extract features, and the CRF module is used to decode the entity recognition results, thus realizing entity naming recognition based on the BiLSTM-CRF sequence annotation method, including:

[0048] 1. Use the pre-trained BERT model to obtain the vector representation of each word. BERT can be regarded as a word embedding generator, which can generate a high-dimensional vector representation for each word.

[0049] 2. Use BiLSTM to extract features from the word embedding generated by BERT. BiLSTM (Bidirectional Long Short-Term Memory Network) can capture the contextual information of words in a sentence.

[0050] 3. Based on the output of BiLSTM, the labeling logic and sequence decoding mechanism of CRF (conditional random field) are used to decode to determine the final label of each word and obtain the sequence labeling result with the highest probability, that is, the entity recognition result. CRF can consider the dependency between labels and perform global optimization.

[0051] Step 2. Entity relationship triple extraction: The entity recognition results are combined with the relationship labels. According to the relationship between the entity pairs, the two types of entities and texts that may have a relationship are spliced ​​and input into the BERT model for classification to obtain the triple extraction results.

[0052] Step 3. Combine the triple extraction results with the label set of the original entity text to obtain a labeled triple entity alignment dataset, and use this dataset to fine-tune the BERT model;

[0053] This step fine-tunes the BERT model: the triple dataset obtained in step 2 is combined with the tag set of the original entity alignment text to form a labeled entity alignment triple dataset; using the triple dataset, BERT is fine-tuned to adapt to the data in triple form. The fine-tuning training process of the BERT model is as follows:

[0054] Step 3.1. Token processing. From the triple data set, two sets of attribute pairs are concatenated with “-”, first tokenized, and then concatenated. Specifically, a [CLS] symbol is set at the beginning of the sentence, and a [SEP] symbol is set at the end of the sentence. For example, “[CLS] Variety-Fruit Color-Red Origin-Orchard [SEP] Brand-Mobile Phone Color-Black Founder-Jobs [SEP]”; convert all samples into a token index array with a fixed length of 128, and set the [CLS] and [SEP] indexes to 101 and 102. If the number of samples is less than 128, fill them with 0;

[0055] Step 3.2. Convert the token sequence to meet the input requirements of the BERT model. Map each token to the pre-trained word vector space to generate word vector embeddings to capture the semantic and contextual information between tokens; assign segmentation embeddings to each token to identify the tag of the sentence to which the token belongs; assign position embeddings to each token to indicate the position information of the token in the sentence; concatenate the word vector embeddings, segmentation embeddings, and position embeddings and use them as the input of the BERT model;

[0056] Step 3.3. Input the processed data into BERT. Through its internal 12-layer Transformer structure, it deeply learns the semantic expression of the data structure text and successfully encodes the contextual semantic information, thereby deriving the predicted label; which includes the output vector of the specific tag [CLS] and the output vector of other characters, each of which carries its own unique semantic information;

[0057] Step 3.4. Through iterative fine-tuning, the model's ability to extract abstract semantic features is continuously improved. In each iteration, the cross entropy loss function is used to calculate the difference between the predicted label and the actual label, and the gradient is calculated through back propagation. Finally, the optimizer is used to update the model's learning rate and weight parameters; the cross entropy loss function of the BERT model fine-tuning is shown in formula (1):

[0058]

[0059] Where N is the number of samples, y i represents the true label of sample i, It indicates the probability that the model predicts that sample i belongs to the positive class. As the probability distribution of the binary classification results tends to be consistent, the cross entropy loss value will decrease accordingly, reflecting that the accuracy of the model prediction results is increasing.

[0060] Step 4. Based on the fine-tuned BERT model and CNN network, build a BERT-CNN entity alignment model for triple data;

[0061] This step is based on the BERT model and CNN network. A BERT-CNN model is built to achieve triple data and entity alignment. The process of building a BERT-CNN entity alignment model for triple data is as follows:

[0062] Step 4.1. For the triple data structure, distinguish the entity name and attribute pairs therein;

[0063] Step 4.2. For entity names, use the BERT model to generate word vector embeddings for entity names; calculate the cosine similarity between these embeddings to measure the similarity between entities; adjust the binary classification output layer matrix according to the set threshold;

[0064] Step 4.3. For the attribute pair part, based on the fine-tuned BERT model, a Dropout layer is introduced to connect the output of the BERT model, and L2 regularization is added to alleviate the potential overfitting problem;

[0065] The loss function of the L2 regularization of the BERT-CNN model is shown in formula (2):

[0066]

[0067] Where L is the total loss function including L2 regularization, N is the number of samples, and x i is the input feature of the i-th sample, y i is the true label of the i-th sample, θ is the set of weight parameters of the model, θ j is the jth weight parameter in the model. data (y i ,f(x i ; θ)) is the data loss function, which is used to measure the true value y i and the model prediction value f(x i ; θ). λ is a hyperparameter of regularization strength, which determines the influence weight of the regularization term in the total loss function. Increasing the λ value will enhance the model's preference for smaller weight parameters, thereby effectively alleviating the overfitting problem.

[0068] Step 4.4. Introduce a 1D convolutional layer (conv1d), and use the output in step 4.3 to learn n-gram-level features at different positions through the convolutional layer; use AdaptiveMaxPool1d for pooling, and map the convolutional features to an output of a fixed size as the output result of the convolutional layer;

[0069] Step 4.5. Use the fully connected layer and softmax function to further process the output of the convolution layer. Introduce the Dropout layer to connect the output of the convolution layer, then input the output into the fully connected layer for linear transformation, and finally use the softmax function to normalize to get the probability of each category, so as to predict the semantic relationship between the two text pairs;

[0070] The present invention proposes a BERT-CNN model to realize the task of aligning the master data entities of industrial products. The specific form of the task is to perform binary classification on the matching results of the master data entities. The BERT-CNN model in the present invention adopts the cross entropy function (Formula 1) combined with the L2 regularized loss function (Formula 2). The total loss function is shown in Formula (3):

[0071]

[0072] Among them, L 1 is the cross entropy loss, which ensures that the model can correctly predict the labels in the binary classification problem. λ is the weight decay coefficient of L2 regularization. A larger λ will impose a greater penalty on the model parameters, making the parameters smaller and more conducive to preventing overfitting; a smaller λ will focus more on data fitting and may lead to overfitting.

[0073] Step 5. The triple data is input into the BERT-CNN entity alignment model, which calculates and outputs the entity alignment probability, predicts the entity alignment result, and completes the entity matching of the triple data.

[0074] Example 2

[0075] A computer device of the present invention includes a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, the steps of the above-mentioned industrial product master data entity alignment method based on the BERT model are implemented. In this embodiment, the computer device is any device or apparatus with data processing capability, which will not be described in detail here.

[0076] Example 3

[0077] A computer-readable storage medium of the present invention stores a program thereon, and when the program is executed by a processor, the steps of the above-mentioned method for aligning the master data entity of industrial products based on the BERT model are implemented. The computer-readable storage medium can be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card (Flash Card), etc. equipped on the device.

[0078] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0079] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A method for aligning industrial product master data entities based on the BERT model, characterized in that: include: Step 1. Encode, extract features and decode the text data of the industrial product master data to be aligned based on BERT-BiLSTM-CRF to obtain entity recognition results; Step 2: Combine entity recognition results with relationship labels, concatenate two types of entities and text that may be related, and input them into the BERT model for classification to obtain triple extraction results; Step 3. Combine the triple extraction results with the label set of the original entity text to obtain a labeled triple dataset, and use this dataset to fine-tune the BERT model; Step 4. Based on the fine-tuned BERT model and CNN network, build a BERT-CNN entity alignment model for triple data; Step 5. Input the triplet data into the BERT-CNN entity alignment model to obtain the entity alignment probability and predict the entity alignment result; In step 4, the process of building a BERT-CNN entity alignment model for triple data is as follows: Step 4.

1. For the triple data structure, distinguish the entity name and attribute pairs therein; Step 4.

2. For entity names, use the fine-tuned BERT model to generate word vector embeddings for the entity names; calculate the cosine similarity between word vector embeddings to measure the similarity between entities; adjust the binary output layer matrix of the BERT model according to the set threshold; Step 4.

3. For attribute pairs, based on the fine-tuned BERT model, a Dropout layer is introduced to concatenate the output of the BERT model, and L2 regularization is added to alleviate potential overfitting problems; Step 4.

4. Introduce a 1D convolutional layer, and use the output in step 4.3 to learn n-gram-level features at different positions through the 1D convolutional layer; then use AdaptiveMaxPool1d for pooling, and map the features learned by the convolutional layer to an output of a fixed size as the output result of the convolutional layer; Step 4.

5. Introduce the fully connected layer and softmax function to further process the output results of the convolutional layer: first introduce the Dropout layer to connect the output results of the convolutional layer, then input the output results into the fully connected layer for linear transformation, and finally use the softmax function for normalization to obtain the probability of each category, thereby realizing the prediction of the semantic relationship between the two text pairs.

2. According to claim 1, a method for aligning industrial product master data entities based on a BERT model is characterized in that: The step 1 first uses the pre-trained BERT model to obtain the vector representation of each word; secondly, BiLSTM is used to extract features of the vector representation obtained by the BERT model; finally, based on the output of BiLSTM, CRF is used for decoding to determine the final label of each word, that is, the entity recognition result.

3. According to claim 1, a method for aligning industrial product master data entities based on a BERT model is characterized in that: In step 3, the fine-tuning training process of the BERT model is as follows: Step 3.

1. Tokenize the attribute pairs in the triple data set, then concatenate them with "-", set a [CLS] symbol at the beginning of the sentence, and set a [SEP] symbol at the end of the sentence. Convert all data samples into a token index array with a fixed length of 128. The [CLS] and [SEP] indexes are set to 101 and 102. If the sample is less than 128, fill it with 0; Step 3.

2. Map each token in the token index array to the pre-trained word vector space to generate word vector embeddings to capture the semantic and contextual information between tokens; assign segmentation embeddings to each token to identify the tag of the sentence to which the token belongs; Assign position embedding to each token to indicate the position information of the token in the sentence; concatenate word vector embedding, segmentation embedding and position embedding as input data of the BERT model; Step 3.

3. Input the input data into the BERT model. Through its internal 12-layer Transformer structure, it deeply learns the semantic expression of the data structure text and encodes the contextual semantic information to obtain the predicted label, which contains the output vector of the tag [CLS] and the output vector of other characters. The output vectors all carry their own unique semantic information. Step 3.

4. Use the cross entropy loss function to calculate the difference between the predicted label and the actual label; through iterative fine-tuning, continuously improve the model's ability to extract abstract semantic features.

4. According to claim 3, a method for aligning industrial product master data entities based on a BERT model is characterized in that: In step 3.4, the cross entropy loss function is: Where N is the number of samples, y i represents the true label of sample i, It represents the probability that the model's predicted label for sample i belongs to the positive class.

5. The method for aligning industrial product master data entities based on the BERT model according to claim 1, characterized in that: In step 4.3, the loss function of L2 regularization is: Where L is the total loss function including L2 regularization, N is the number of samples, and x i is the input feature of the i-th sample, y i is the true label of the i-th sample, θ is the set of weight parameters of the model, θ j is the jth weight parameter in the model; L data (y i ,f(x i ; θ)) is the data loss function, which is used to measure the true value y i and the model prediction value f(x i ; θ); λ is a hyperparameter of the regularization strength.

6. The method for aligning industrial product master data entities based on the BERT model according to claim 1, characterized in that: In step 4, the total loss function of the BERT-CNN entity alignment model is: Among them, L1 is the cross entropy loss and λ is the weight decay coefficient of L2 regularization.

7. A computer device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, the steps of the industrial product master data entity alignment method based on the BERT model as described in any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps of the industrial product master data entity alignment method based on the BERT model as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Knowledge graph construction method based on improved BERT model

    CN110390023A

  • Relation extraction and knowledge graph construction method based on deep learning model

    CN110598000A