Name matching method, training method, device and storage medium
The problem of name matching cost and efficiency in the prior art is solved through a pre-trained neural network, and efficient name matching is achieved.
Patent Information
- Application Number
- CN202210153195.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-02-18
AI Technical Summary
In the prior art, as the list increases, the maintenance cost, storage space and matching calculation time required in the name matching method will greatly increase, making it difficult to effectively process various variants of the name.
Through a pre-trained neural network, strings of different variants of the same name are converted into the same representation vector, and then these representation vectors are used for similarity measurement in the name matching process to determine whether the name to be matched matches the reference name.
This method reduces the amount of data to be stored, reduces maintenance costs, and improves the accuracy and efficiency of name matching, and can handle multiple variants in one match, reducing matching time.
Smart Images

Figure CN114510944B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of computer software technology, and in particular to a name matching method, a training method, a device, and a storage medium. Background Art
[0002] Name matching refers to the process of comparing a pair of names to determine whether they refer to the same entity. The entity here can be a person or a thing (such as an object, a group, or a company, etc.). For example, name matching, as a basic means of identity recognition, plays an important role in the fields of financial compliance, administrative law enforcement, and homeland security. There are naturally many variants of names, such as abbreviations, spelling mistakes, aliases, nicknames, transliterations, and translations (multilingual), which increase the difficulty of name matching.
[0003] One name matching method in the related art is list-based multilingual name matching, that is, collecting and sorting out the written forms of the same name in multiple languages and various spelling variants to form a list database, and directly searching and judging whether they match when in use. Its limitation is that as the list grows, the required maintenance cost, storage space, and matching calculation time will increase significantly. Summary of the Invention
[0004] In view of this, one or more embodiments of this specification provide a name matching method, a device training method, a device, and a storage medium.
[0005] To achieve the above object, one or more embodiments of this specification provide the following technical solutions:
[0006] According to the first aspect of one or more embodiments of this specification, a name matching method is proposed, including:
[0007] Obtain the name string of the name to be matched;
[0008] Convert the name string of the name to be matched into a representation vector according to a pre-trained neural network; wherein, the neural network is used to convert the strings of different variants of the same name into the same representation vector;
[0009] Determine the similarity between the representation vector of the name to be matched and the representation vectors of a number of pre-stored reference names; wherein, the representation vector of the reference name is obtained by inputting any variant string of the reference name into the neural network;
[0010] Determine whether the name to be matched matches the reference name according to the similarity.
[0011] According to the second aspect of one or more embodiments of this specification, a training method for a neural network for name matching is proposed, including:
[0012] Obtain a number of triple samples, where the triple samples include two positive samples and one negative sample. The two positive samples include strings of different variants of the same named sample, and the negative sample includes a string belonging to a different named sample from the positive samples;
[0013] Input the triple samples into a preset neural network with three branches, and process one of the samples in the triple samples by each branch to obtain three feature vectors; wherein, the weights of the three branches are shared;
[0014] Adjust the parameters of the preset neural network according to the similarity degree between the feature vectors corresponding to the two positive samples and / or the difference degree between the feature vector of one positive sample and the feature vector of the negative sample, and obtain a trained neural network; wherein, the trained neural network includes at least one of the branches; the trained neural network is used to convert strings of different variants of the same name into the same feature vector.
[0015] According to the third aspect of one or more embodiments of this specification, a training method for a neural network for name matching is proposed, including:
[0016] Obtain a number of binary samples, where a part of the binary samples include two positive samples, and another part of the binary samples include one positive sample and one negative sample; the two positive samples include strings of different variants of the same named sample, and the negative sample includes a string belonging to a different named sample from the positive samples;
[0017] Input the binary samples into a preset neural network with two branches, and process one of the samples in the binary samples by each branch to obtain two feature vectors; wherein, the weights of the two branches are shared;
[0018] Adjust the parameters of the preset neural network according to the similarity degree between the feature vectors corresponding to the two positive samples and / or the difference degree between the feature vector of the positive sample and the feature vector of the negative sample, and obtain a trained neural network; wherein, the trained neural network has at least one of the branches; the trained neural network is used to convert strings of different variants of the same name into the same feature vector.
[0019] According to the fourth aspect of one or more embodiments of this specification, an electronic device is proposed, including:
[0020] A processor;
[0021] A memory for storing instructions executable by the processor;
[0022] Wherein, the processor runs the executable instructions to implement the method described in any one of the first aspect, the second aspect, or the third aspect.
[0023] According to a fifth aspect of one or more embodiments of the present specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in any one of the first aspect, the second aspect, or the third aspect are implemented.
[0024] The name matching method, device training method, device, and storage medium provided by one or more embodiments of the present specification. By using a pre-trained neural network to obtain the representation vectors of several reference names, the neural network can convert the strings of different variants of the same name into the same representation vector, so that it is not necessary to store the strings of different variants of the reference names, but only one representation vector corresponding to the reference name needs to be pre-stored, reducing the amount of data to be stored and also reducing the maintenance cost.
[0025] Furthermore, in the name matching process, since the strings of different variants of the same name can be represented by the same representation vector, the name string of the to-be-matched name can be converted into a representation vector according to the pre-trained neural network, and then by measuring the similarity between the representation vector of the to-be-matched name and the representation vectors of several pre-stored reference names, it is determined whether the to-be-matched name matches the reference name, improving the matching accuracy, and realizing a one-time matching process between the to-be-matched name and each reference name based on the representation vector, without separately matching the to-be-matched name with multiple variants of the reference name, which is beneficial to reducing the matching time and improving the matching efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a flowchart of a training method for a neural network for name matching provided by an exemplary embodiment.
[0027] Figure 2 is a schematic structural diagram of a neural network provided by an exemplary embodiment.
[0028] Figure 3 is a schematic structural diagram of another neural network provided by an exemplary embodiment.
[0029] Figure 4 is a schematic diagram of the change in the distance between the representation vectors of two samples before and after the training of a neural network provided by an exemplary embodiment.
[0030] Figure 5 is a flowchart of another training method for a neural network for name matching provided by an exemplary embodiment.
[0031] Figure 6It is a flowchart of a name matching method provided by an exemplary embodiment.
[0032] Figure 7 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment. Detailed implementation manners
[0033] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0034] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0035] Name matching refers to the process of comparing whether a pair of names refer to the same entity. The entity mentioned here includes but is not limited to humans, which can be a person or a thing (such as an object, a group, a commodity, a company, equipment materials, or a disease, etc.).
[0036] Name matching can be applied to different scenarios. For example, in the scenario of person name matching, person name matching is a very important technology in the field of risk control. For example, the risk control system records the names of each determined illegal user in the blacklist. Then, when performing risk control, for each user currently conducting business, by scanning, the name of each user is matched with the names in the blacklist. If the match is successful, it can be considered that the user is an illegal user and their business is rejected to prevent risks.
[0037] In another example, such as in a shopping scenario, when a user browses a shopping website, they will search for the desired commodity, for example, by entering the commodity name. Then, the shopping platform needs to match the commodity name entered by the user with the commodity names of each commodity in the database and push the commodity information corresponding to the successful match to the user.
[0038] For example, in a medical scenario, a matching process for disease names may be required; in a material procurement scenario, a matching process for equipment material names may be required; in a financial scenario, matching for company names, unit names, etc. may be required, and so on.
[0039] Names naturally have many variants, such as abbreviations, spelling mistakes, aliases, nicknames, pinyin, transliteration, and translations (multilingual), etc. These situations increase the difficulty of name matching. For example, the two different strings "Xiaoming" and the pinyin "xiaoming" are different variants of the same name, referring to the same person. Another example is that the two different strings "XX Technology Co., Ltd." and the abbreviation "X Tech" are different variants of the same company name, referring to the same company.
[0040] A name matching method in the related technology is multilingual name matching based on a list, that is, collecting and organizing the writings in multiple languages of the same name and various spelling variants to form a list database, and directly searching and judging whether there is a match when using it. Its limitation is that as the list increases, the required maintenance cost, storage space, and matching calculation time will increase significantly.
[0041] Based on this, this specification implements a pre-trained neural network that can convert strings of different variants of the same name into the same representation vector, so that it is not necessary to store strings of different variants of the same name, but only need to pre-store a representation vector corresponding to the name, reducing the amount of data to be stored and also reducing the maintenance cost; and using the neural network.
[0042] In the process of name matching, since strings of different variants of the same name can be represented by the same representation vector, after obtaining the name string of the name to be matched, the name string of the name to be matched can be converted into a representation vector according to the pre-trained neural network, and then the similarity between the representation vector of the name to be matched and the representation vectors of several pre-stored reference names is determined; furthermore, according to the similarity, it is determined whether the name to be matched matches the reference name, which is beneficial to improving the matching accuracy, and this embodiment realizes a one-time matching process of the name to be matched and each reference name based on the similarity judgment of the representation vector, without separately matching the name to be matched with multiple variants of the reference name, which is beneficial to reducing the matching time and improving the matching efficiency.
[0043] Among them, the training process of the neural network and the process of using the neural network for name matching can be executed by an electronic device, and the electronic device includes, but is not limited to, devices with computing capabilities such as servers, computer room devices, computers, tablets, or mobile terminals. Exemplarily, this specification provides a program product integrated in an electronic device, so that when the electronic device runs the program product, it can execute the name matching method provided in this specification or the training method of the neural network for name matching. Exemplarily, the electronic device includes a processor and a memory, and the processor realizes the name matching method provided in this specification or the training method of the neural network for name matching by running the executable instructions stored in the memory.
[0044] It can be understood that the training process of the neural network and the process of using the neural network for name matching can be executed by the same electronic device, or can be executed by different electronic devices, and can be specifically set according to the actual application scenario. This embodiment does not make any restrictions on this. In one example, in order to improve the accuracy of the neural network, the training process of the neural network can be carried out using a large number of training samples in an electronic device with higher computing power. After the neural network is trained, the trained neural network can be transplanted into the electronic device that needs to perform name matching.
[0045] Here, the training process of the neural network will be described first. The neural network is obtained by performing contrastive learning and representation learning based on name samples with multiple variants.
[0046] Contrastive Learning belongs to a type of self-supervised learning. Contrastive Learning learns the feature representation of samples by comparing data with positive and negative samples in the feature space respectively. Contrastive Learning focuses on learning the common features among similar instances and distinguishing the differences among non-similar instances. Compared with Generative Learning, Contrastive Learning does not need to pay attention to the cumbersome details on the instances, and only needs to learn to distinguish data in the feature space at the abstract semantic level. Therefore, the model and its optimization become simpler, and the generalization ability is stronger.
[0047] Representation learning is a collection of techniques for learning a feature, which transforms the original data into a form that can be effectively exploited by machine learning. It avoids the trouble of manually extracting features, allows the computer to learn to use features while also learning how to extract features: learning how to learn.
[0048] In some embodiments, during the neural network training process, the neural network can determine the similarity of strings of different variants of name samples to learn the representation vectors of names. Such similarity includes semantic similarity and / or phonetic similarity. Exemplarily, the name sample with multiple variants includes strings of different variants that are phonetically similar, and these strings of different variants that are phonetically similar point to the same entity, so that the neural network can learn phonetic similarity features during the training process. Exemplarily, when the strings of different variants that are phonetically similar meet the condition of pointing to the same entity, they can be strings in different languages or strings in the same language; such as phonetic transcriptions of personal names in different languages, phonetic transcriptions of company names, etc. After the neural network is trained, for different strings obtained by phonetic transcription of the same entity in different languages, the neural network that has learned phonetic similarity features can convert the different strings indicating the same entity into the same representation vector.
[0049] Exemplarily, the name sample with multiple variants includes strings of different variants that are semantically similar, and these strings of different variants that are semantically similar point to the same entity, so that the neural network can learn semantic similarity features during the training process. Exemplarily, when the strings of different variants that are semantically similar meet the condition of pointing to the same entity, they can be strings in different languages or strings in the same language; such as translation variants or abbreviations of company names, aliases or nicknames of personal names, etc. After the neural network is trained, for different strings that are semantically the same or similar for the same entity, the neural network that has learned semantic similarity features can convert the different strings indicating the same entity into the same representation vector.
[0050] Exemplarily, the name sample with multiple variants includes strings of different variants that are phonetically similar and strings of different variants that are semantically similar, and multiple variants of the name sample all point to the same entity, so that the neural network can learn phonetic similarity features and semantic similarity features during the training process. After the neural network is trained, for different variants of the same name, such as abbreviations, spelling mistakes, aliases, nicknames, pinyin, phonetic transcriptions, and translations (multilingual), etc., the neural network can convert strings of different variants of the same name (there is phonetic similarity and / or semantic similarity between these different variants) into the same representation vector.
[0051] After learning phonetic similarity features and / or semantic similarity features, the neural network provided by the embodiments of this specification can achieve cross - language comparison without translating strings in other languages, and has wide applicability.
[0052] In some embodiments, the optimization objective of the neural network includes: minimizing the distance between the representation vectors corresponding to the strings of different variants of samples with the same name, and / or maximizing the distance between the representation vectors corresponding to at least two strings of samples with different names. In other words, during the training process of the neural network, based on name samples with multiple variants, the neural network learns a function F that can encode the input data into a representation vector, such that the representation vectors corresponding to the strings of different variants of samples with the same name are as similar as possible, while the representation vectors corresponding to at least two strings of samples with different names are as different as possible, improving the accuracy of subsequent name matching using the neural network.
[0053] In some embodiments, the neural network training can be performed through a Triplet Network structure or a Siamese Network structure to learn the process of encoding strings of different variants of the same name into the same representation vector and encoding strings of different names into different representation vectors.
[0054] In an exemplary embodiment, taking the neural network training example with a Triplet Network structure as an illustration, please refer to Figure 1 , the embodiments of the present specification provide a training method for a neural network for name matching. The method can be executed by an electronic device. The method includes:
[0055] In step S101, a number of triplet samples are obtained. The triplet samples include two positive samples and one negative sample. The two positive samples include strings of different variants of the same name sample, and the negative sample includes a string that belongs to a different name sample from the positive samples.
[0056] In step S102, the triplet samples are input into a preset neural network with three branches. Each branch processes one of the samples in the triplet samples to obtain three representation vectors; wherein, the weights of the three branches are shared.
[0057] In step S103, according to the similarity degree between the representation vectors corresponding to the two positive samples and / or the difference degree between the representation vector of one positive sample and the representation vector of the negative sample, the parameters of the preset neural network are adjusted to obtain a trained neural network; wherein, the trained neural network includes at least one of the branches; the trained neural network is used to convert strings of different variants of the same name into the same representation vector.
[0058] Exemplarily, in the process of obtaining triple samples, for each string x in the dataset, x can be used as one of the positive samples, and any string selected from its matching list can be used as another positive sample x+, and any non-matching string selected from the same training batch can be used as the negative sample x-, so as to construct a triple sample (x+, x, x-); based on the above method, a number of triple samples (x+, x, x-) can be constructed. Among them, the string x and the strings in its matching list all point to the same entity, and the matching list of the string x includes semantically similar strings and / or phonetically similar strings, so that the neural network can learn semantically similar features and / or phonetically similar features.
[0059] Among them, the neural network has three branches, and the weights of the three branches are shared. After the electronic device obtains a number of triple samples (x+, x, x-), it inputs the triple samples into a preset neural network with three branches, and each branch processes one of the samples in the triple samples to obtain three representation vectors.
[0060] Exemplarily, please refer to Figure 2 , each branch in the neural network 100 includes at least an Embedding layer and an encoder 20; after the electronic device obtains a number of triple samples (x+, x, x-), it tokenizes each sample (i.e., string) in the triple samples to obtain a character set, and then inputs the three obtained character sets into the Embedding layers 10 of the three branches respectively. It can be understood that in this embodiment, no specific limitation is imposed on the specific type of the tokenizer used by the electronic device, and it can be specifically set according to the actual application scenario. For example, a character-level tokenizer can be used for tokenization.
[0061] In each branch, considering the discreteness of the character set, after the Embedding layer 10 obtains the input character set, it can perform conversion processing on the discrete character set to obtain continuous embedding vectors. Exemplarily, the Embedding layer 10 includes an embedding matrix with learnable weights, and the embedding matrix can be used to perform conversion processing on the input character set to obtain continuous embedding vectors. Through the embedding operation of the Embedding layer 10 in this embodiment, the discrete character set can be reduced to low-dimensional dense features, reducing the dimension of the vector, which is beneficial to improving the encoding efficiency of the subsequent encoder 20.
[0062] Considering that the embedding vectors obtained by the embedding layer 10 belong to the feature data in the character vector space, it is difficult to judge the semantic similarity and / or phonetic similarity between strings through the embedding vectors. Therefore, in this embodiment, an encoder 20 is provided. The encoder 20 is used to map the embedding vectors from the character vector space to the numerical vector space to obtain the representation vectors, which are continuous numerical representation vectors, so as to realize the similarity learning between strings through the representation vectors in the numerical vector space. It can be understood that the specific type of the encoder 20 in the embodiments of this specification is not limited in any way and can be specifically set according to the actual application scenario. For example, the encoder 20 can be a bidirectional long short-term memory (bi-LSTM), or it can also be the encoder 20 part of the Transformer model. Among them, the input of the encoder 20 at each time step is a character.
[0063] In order for the neural network 100 to learn sufficient semantic similarity features and / or phonetic similarity features, the dimension of the representation vectors obtained by the encoder 20 is usually relatively large. If the electronic device has sufficient computing resources, the representation vectors output by the encoder 20 can be directly supplied to the loss function for evaluation; if the computing resources of the electronic device are insufficient or there are efficiency requirements, the representation vectors output by the encoder 20 can be dimensionally reduced, and the dimensionally reduced representation vectors can be supplied to the loss function for evaluation.
[0064] Exemplarily, please refer to Figure 3 , each branch in the neural network 100 further includes a fully connected layer 30. The fully connected layer 30 is used to dimensionally reduce the representation vectors output by the encoder 20, so as to project the representation vectors output by the encoder 20 into a vector space with a smaller dimension, which is beneficial to reducing the amount of data to be processed in subsequent steps and improving the training efficiency. Exemplarily, the fully connected layer 30 can be composed of a multi-layer perceptron (MLP).
[0065] After obtaining the three representation vectors, the electronic device can adjust the parameters of the preset neural network 100 according to the similarity degree between the representation vectors corresponding to the two positive samples and / or the difference degree between the representation vector of the positive sample and the representation vector of the negative sample, so as to obtain the trained neural network 100. Exemplarily, please refer to Figure 2, a loss function is linked after the fully connected layer 30. The loss function can be used to measure the similarity degree between the representation vectors corresponding to two positive samples, and / or the difference degree between the representation vector of the positive sample and the representation vector of the negative sample. Furthermore, the parameters of the neural network 100 are adjusted according to the loss value of the loss function, so as to achieve reducing the distance between the representation vectors of the same-class samples (two positive samples), and at the same time increasing the distance between the representation vectors of the non-same-class samples (positive sample and negative sample), until the optimization goal of the neural network 100 is achieved (that is, minimizing the distance between the representation vectors corresponding to the strings of different variants of the samples belonging to the same name, and maximizing the distance between the representation vectors corresponding to at least two strings of the samples belonging to different names).
[0066] Exemplarily, the loss function includes a triplet loss function. The mathematical expression of the triplet loss function is: where ε is a hyperparameter that balances the similarity measure and the dissimilarity measure. The triplet loss function can be used to measure the similarity degree between the representation vectors corresponding to two positive samples (x and x_+), and the difference degree between the representation vector of the positive sample (x) and the representation vector of the negative sample (x_-). The parameters of the neural network can be adjusted according to the loss value of the triplet loss function, so that the representation vector of the positive sample x and the representation vector of the positive sample x_+ that matches it are as similar as possible, while being as different as possible from the representation vector of the unmatched sample x_-. This is beneficial to improving the accuracy of subsequent name matching using the neural network, and the precision-recall rate is higher.
[0067] In one example, please refer to Figure 4 , the circle "●" represents the representation vector of each sample. d1 represents the distance between the representation vectors corresponding to the positive sample x and the positive sample x_+ indicating the same entity in the vector space, and d2 represents the distance between the representation vectors corresponding to the positive sample x and the negative sample x_- indicating different entities in the vector space is d2. From Figure 3 it can be seen that before the neural network training, d1 > d2, which does not meet the requirements; after the neural network undergoes contrastive learning, during the training process, the parameters of the neural network are adjusted based on the loss value of the triplet loss function, so as to achieve reducing the distance between the representation vectors of the positive sample x and the positive sample x_+ that matches it, and increasing the distance between the representation vector of the positive sample x and the representation vector of the unmatched sample x_-, making d1 < d2. In other words, making the representation vector of the positive sample x and the representation vector of the positive sample x_+ that matches it as similar as possible, while being as different as possible from the representation vector of the unmatched sample x_-. This is beneficial to improving the accuracy of subsequent name matching using the neural network, and the precision-recall rate is higher.
[0068] Exemplarily, the loss function can also be a contrastive loss function, which can be used to measure the similarity between the representation vectors corresponding to two positive samples (x and x_+), or the difference between the representation vector of the positive sample (x) and the representation vector of the negative sample (x_-). The parameters of the neural network can be adjusted according to the loss value of the contrastive loss function, so as to reduce the distance between the representation vectors of the positive sample x and its matching positive sample x_+, making them as similar as possible, while increasing the distance between the representation vector of the positive sample x and the representation vector of the non-matching sample x_-, making them as different as possible. This is conducive to improving the accuracy of subsequent name matching using the neural network, with a higher precision-recall rate.
[0069] In another exemplary embodiment, an example of neural network training using a Siamese Network structure is described. Please refer to Figure 5 , an embodiment of this specification provides a training method for a neural network for name matching. The method can be executed by an electronic device, and the method includes:
[0070] In step S201, a number of binary samples are obtained, where some of the binary samples include two positive samples, and some of the binary samples include one positive sample and one negative sample; the two positive samples include strings of different variants of the same name sample, and the negative sample includes a string belonging to a different name sample from the positive sample.
[0071] In step S202, the binary samples are input into a preset neural network with two branches, and each branch processes one of the samples in the binary samples to obtain two representation vectors; wherein, the weights of the two branches are shared.
[0072] In step S203, according to the similarity between the representation vectors corresponding to the two positive samples and / or the difference between the representation vector of the positive sample and the representation vector of the negative sample, the parameters of the preset neural network are adjusted to obtain a trained neural network; wherein, the trained neural network has at least one of the branches; the trained neural network is used to convert strings of different variants of the same name into the same representation vector.
[0073] Exemplarily, in the process of obtaining the binary tuple samples, for each string x in the dataset, x can be used as one of the positive samples, and any string selected from its matching list can be used as another positive sample x_+. Then, a binary tuple sample (x_+, x) can be constructed. Or, any non-matching string selected from the same training batch can be used as the negative sample x_-. Then, a binary tuple sample (x, x_-) can be constructed. Based on the above method, a part of the binary tuple samples including two positive samples (x_+, x) and a part of the binary tuple samples including one positive sample and one negative sample (x, x_-) can be constructed.
[0074] Among them, the string x and the strings in its matching list all point to the same entity. The matching list of the string x includes semantically similar strings and / or phonetically similar strings, so that the neural network can learn semantically similar features and / or phonetically similar features.
[0075] When training the neural network in the structure of a Siamese Network, the neural network has two branches with shared weights. After the electronic device obtains a number of binary tuple samples, the binary tuple samples are input into a preset neural network with two branches. Each branch processes one of the samples in the binary tuple samples to obtain two feature vectors.
[0076] The structure of each branch in the neural network is similar to the structure of each branch in the triplet network structure. Each branch includes at least an embedding layer and an encoder. Or, each branch includes an embedding layer, an encoder, and a fully connected layer. For the relevant parts, please refer to the description of the embodiments shown in Figure 2 and Figure 3 the description of the embodiments shown.
[0077] After obtaining two representation vectors, the electronic device may adjust the parameters of the preset neural network according to the similarity degree between the representation vectors corresponding to the two positive samples and / or the difference degree between the representation vector of the positive sample and the representation vector of the negative sample, so as to obtain a trained neural network. Exemplarily, after obtaining two representation vectors, the two representation vectors may be input into a loss function, and the loss function includes a contrastive loss function. The contrastive loss function may be used to measure the similarity degree between the representation vectors corresponding to the two positive samples (x and x_+), or the difference degree between the representation vector of the positive sample (x) and the representation vector of the negative sample (x_-). The parameters of the neural network may be adjusted according to the loss value of the contrastive loss function, so that the representation vector of the positive sample x and the representation vector of the positive sample x_+ that matches it are as similar as possible, while being as different as possible from the representation vector of the unmatched sample x_-, thereby facilitating improving the accuracy of subsequent name matching using the neural network and having a higher precision-recall rate.
[0078] It can be understood that during the neural network training process, in addition to using two branches or three branches with weight sharing for model training, under the condition of following contrastive learning and representation learning, more branches can also be used to participate in the training process, and this embodiment does not impose any restrictions on this. For example, four branches with weight sharing are used for model training, and the training samples are quadruple samples, including at least one positive sample and at least one negative sample.
[0079] After training the neural network based on the above-described process, the trained neural network can convert strings of different variants that are phonetically and / or semantically similar in the same name (indicating the same entity) into the same representation vector, and convert strings of different names (indicating different entities) into different representation vectors. Moreover, the neural network provided in the embodiments of this specification can achieve cross-language comparison after learning phonetic similarity features and / or semantic similarity features, and has wide applicability. After obtaining the trained neural network, the electronic device can use the trained neural network to convert all reference names in the database into representation vectors. For example, a string of any variant of the reference name can be input into the neural network to obtain the representation vector of the reference name. After obtaining the representation vectors of all reference names in the database, the representation vectors of the reference names can be pre-stored for subsequent name matching processes, and all reference names and all their variants in the database are no longer needed during the name matching process. In this embodiment, only the representation vectors of the reference names need to be pre-stored for subsequent name matching processes, which helps to reduce the amount of data to be stored.
[0080] Next, the process of using the trained neural network for name matching will be described: Please refer toFigure 6 , Figure 6 It is a schematic flowchart of a name matching method provided by an embodiment of this specification. The method can be executed by an electronic device, and the method includes:
[0081] In step S301, obtain the name string of the name to be matched.
[0082] In step S302, convert the name string of the name to be matched into a representation vector according to a pre-trained neural network; wherein, the neural network is used to convert strings of different variants of the same name into the same representation vector.
[0083] In step S303, determine the similarity between the representation vector of the name to be matched and the representation vectors of a plurality of pre-stored reference names respectively; wherein, the representation vector of the reference name is obtained by inputting the string of any variant of the reference name into the neural network.
[0084] In step S304, determine whether the name to be matched matches the reference name according to the similarity.
[0085] In this embodiment, since strings of different variants of the same name can be represented by the same representation vector, the name string of the name to be matched can be converted into a representation vector by using a pre-trained neural network. Furthermore, by measuring the similarity between the representation vector of the name to be matched and the representation vectors of a plurality of pre-stored reference names respectively, it is determined whether the name to be matched matches the reference name. While improving the matching accuracy, a one-time matching process between the name to be matched and each reference name is realized based on the representation vector, without separately matching the name to be matched with multiple variants of the reference name, which is beneficial to reducing the matching time and improving the matching efficiency.
[0086] In some embodiments, after obtaining the name string of the name to be matched, the electronic device performs word segmentation processing on the name string of the name to be matched to obtain a character set of the name to be matched; then inputs the character set into a pre-trained neural network, and converts the character set into a representation vector through the neural network.
[0087] Exemplarily, the neural network includes an embedding layer and an encoder. Through the embedding layer, the character set obtained after word segmentation of the name string is subjected to conversion processing to obtain an embedding vector, which is beneficial to reducing the dimension of the input data; through the encoder, the embedding vector is mapped from the character vector space to the numerical vector space to obtain the representation vector.
[0088] Exemplarily, to improve the processing efficiency, the neural network further includes a fully connected layer, through which the feature vector output by the encoder can be dimensionally reduced. With the reduction of the dimension of the feature vector, it is beneficial to reduce the amount of calculation in the similarity calculation process and improve the calculation efficiency.
[0089] Among them, the electronic device pre-stores the feature vectors of several reference names. Exemplarily, the feature vector of the reference name is obtained by inputting the string of any variant of the reference name into the neural network. When obtaining the feature vector of the name to be matched, the electronic device can determine the similarity between the feature vector of the name to be matched and the pre-stored feature vectors of several reference names, and then determine whether the name to be matched matches the reference name according to the similarity.
[0090] It can be understood that the embodiments of this specification do not impose any restrictions on the similarity algorithm used by the electronic device, and can be specifically set according to the actual application scenario. For example, the Euclidean distance, cosine distance, etc. between the feature vector of the name to be matched and the pre-stored feature vectors of several reference names can be calculated to determine the similarity between the feature vector of the name to be matched and the pre-stored feature vectors of several reference names.
[0091] Exemplarily, a preset threshold for similarity can be set according to the actual application scenario. If the similarity is greater than the preset threshold, it is determined that the name to be matched matches the reference name, and the name to be matched and the reference name point to the same entity; otherwise, it is determined that the name to be matched does not match the reference name. Furthermore, the electronic device can execute relevant service processing procedures according to the matching result.
[0092] Correspondingly, the embodiments of this specification also provide an electronic device, including:
[0093] A processor;
[0094] A memory for storing executable instructions of the processor;
[0095] Among them, the processor realizes the method described in any one of the above by running the executable instructions.
[0096] The processor executes the executable instructions included in the memory. The processor can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0097] The memory stores the executable instructions of the above method. The memory can include at least one type of storage medium, and the storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), Random Access Memory (RAM), Static Random Access Memory (SRAM), Read Only Memory (ROM), Electrically Erasable Programmable Read Only Memory (EEPROM), Programmable Read Only Memory (PROM), magnetic memory, magnetic disk, optical disk, etc. Moreover, the electronic device can cooperate with a network storage device that performs the storage function of the memory through a network connection. The memory can be an internal storage unit of the electronic device, such as the hard disk or memory of the electronic device. The memory can also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory can also include both the internal storage unit of the electronic device and the external storage device. The memory is used to store executable instructions and other programs and data required by the device. The memory can also be used to temporarily store data that has been output or will be output.
[0098] Exemplarily, Figure 7 is a schematic structural diagram of an electronic device 700 provided by an exemplary embodiment. Please refer to Figure 7, at the hardware level, the device 700 includes a processor 702, an internal bus 704, a network interface 706, a memory 708, and a non-volatile memory 710. Of course, it may also include other hardware required for other services. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 702 reads the corresponding computer program from the non-volatile memory 710 into the memory 708 and then runs it. Of course, in addition to the software implementation manner, one or more embodiments of this specification do not exclude other implementation manners, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or logical devices.
[0099] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided. For example, a memory including instructions, and the above instructions can be executed by the processor of the device to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0100] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the terminal, enables the terminal to execute the above method.
[0101] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.
[0102] In a typical configuration, a computer includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0103] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0104] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0105] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0106] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0107] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0108] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0109] The above are only the preferred embodiments of one or more embodiments of this specification, and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope protected by one or more embodiments of this specification.
Claims
1. A name matching method, comprising: Obtain the name string of the name to be matched; Convert the name string of the name to be matched into a representation vector according to a pre-trained neural network; wherein, the neural network is used to convert strings of different variants of the same name into the same representation vector; Determine the similarity between the representation vector of the name to be matched and the representation vectors of a number of pre-stored reference names; wherein, the representation vector of the reference name is obtained by inputting the string of any variant of the reference name into the neural network; Determine whether the name to be matched matches the reference name according to the similarity.
2. The method according to claim 1, wherein the neural network is used to convert strings of different variants that are phonetically and / or semantically similar in the same name into the same representation vector.
3. The method according to claim 1, wherein converting the name string of the name to be matched into a representation vector according to the pre-trained neural network comprises: Perform word segmentation processing on the name string of the name to be matched to obtain the character set of the name to be matched; Input the character set into a pre-trained neural network, and convert the character set into a representation vector through the neural network.
4. The method according to claim 1 or 3, wherein the neural network at least comprises an embedding layer and an encoder; The embedding layer is used to perform conversion processing on the character set obtained after word segmentation of the name string to obtain an embedding vector; The encoder is used to map the embedding vector from the character vector space to the numerical vector space to obtain the representation vector.
5. The method according to claim 4, wherein the neural network further comprises a fully connected layer; The fully connected layer is used to perform dimensionality reduction processing on the representation vector output by the encoder.
6. The method according to claim 1, wherein during the training process, the neural network is obtained by performing contrast learning and representation learning according to name samples with multiple variants; Wherein, The name sample with multiple variants includes strings of different variants that are phonetically similar and / or semantically similar.
7. The method according to claim 6, wherein during the training process, the optimization objective of the neural network comprises: Minimize the distance between the representation vectors corresponding to the strings of different variants belonging to the same name sample and / or maximize the distance between the representation vectors corresponding to at least two strings belonging to different name samples.
8. The method according to claim 7, further comprising: During the training process, obtain a number of triple samples, where the triple samples include two positive samples and one negative sample, the two positive samples include strings of different variants of the same name sample, and the negative sample includes a string belonging to a different name sample from the positive sample; Input the triple samples into a preset neural network with three branches, and each branch processes one of the samples in the triple samples to obtain three representation vectors; wherein, the weights of the three branches are shared; Adjust the parameters of the preset neural network according to the similarity degree between the representation vectors corresponding to the two positive samples and / or the difference degree between the representation vector of one positive sample and the representation vector of the negative sample to obtain the trained neural network; wherein, the trained neural network includes at least one of the branches.
9. The method according to claim 7, further comprising: During the training process, obtain a number of pair samples, where a part of the pair samples include two positive samples, and another part of the pair samples include one positive sample and one negative sample; the two positive samples include strings of different variants of the same name sample, and the negative sample includes a string belonging to a different name sample from the positive sample; Input the pair samples into a preset neural network with two branches, and each branch processes one of the samples in the pair samples to obtain two representation vectors; wherein, the weights of the two branches are shared; Adjust the parameters of the preset neural network according to the similarity degree between the representation vectors corresponding to the two positive samples and / or the difference degree between the representation vector of the positive sample and the representation vector of the negative sample to obtain the trained neural network; wherein, the trained neural network has at least one of the branches.
10. The method according to claim 8 or 9, wherein during the training process, the loss function of the neural network comprises a triplet loss function and / or a contrast loss function; The ternary loss function is used to measure the similarity degree between the representation vectors corresponding to two positive samples respectively, and the difference degree between the representation vector of one positive sample and the representation vector of the negative sample; The contrastive loss function is used to measure the similarity degree between the representation vectors corresponding to two positive samples respectively, or the difference degree between the representation vector of the positive sample and the representation vector of the negative sample.
11. The method according to claim 1, wherein determining whether the name to be matched matches the reference name according to the similarity degree comprises: If the similarity is greater than a preset threshold, it is determined that the name to be matched matches the reference name, and the name to be matched and the reference name point to the same entity; otherwise, it is determined that the name to be matched does not match the reference name.
12. A method for training a neural network for name matching, comprising: Obtain a number of triple samples, where the triple samples include two positive samples and one negative sample. The two positive samples include strings of different variants of the same name sample, and the negative sample includes a string that belongs to a different name sample from the positive samples. Input the triple samples into a preset neural network with three branches, and each branch processes one of the samples in the triple samples to obtain three feature vectors; among them, the weights of the three branches are shared. Adjust the parameters of the preset neural network according to the similarity degree between the feature vectors corresponding to the two positive samples and / or the difference degree between the feature vector of one positive sample and the feature vector of the negative sample, and obtain a trained neural network; where the trained neural network includes at least one of the branches; the trained neural network is used to convert strings of different variants of the same name into the same feature vector.
13. A method for training a neural network for name matching, comprising: Obtain a number of binary samples, where some of the binary samples include two positive samples, and some of the binary samples include one positive sample and one negative sample; the two positive samples include strings of different variants of the same name sample, and the negative sample includes a string that belongs to a different name sample from the positive samples. Input the binary samples into a preset neural network with two branches, and each branch processes one of the samples in the binary samples to obtain two feature vectors; among them, the weights of the two branches are shared. Adjust the parameters of the preset neural network according to the similarity degree between the feature vectors corresponding to the two positive samples and / or the difference degree between the feature vector of the positive sample and the feature vector of the negative sample, and obtain a trained neural network; where the trained neural network has at least one of the branches; the trained neural network is used to convert strings of different variants of the same name into the same feature vector.
14. An electronic device, comprising: Processor; A memory for storing instructions executable by the processor; Wherein, the processor realizes the method according to any one of claims 1 to 13 by running the executable instructions.
15. A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Zero-sample learning image classification method and device and electronic equipment
CN111738316A
Text recognition method and device, computer equipment and storage medium
CN113822264A