Name Disambiguation Method Based on Multi-Relation Deep Retrieval Text Matching

By integrating cross-modal data and graph neural networks to dynamically update enterprise-person relationship graphs, the method addresses the challenge of accurately disambiguating person names, enhancing precision and robustness in identifying individuals within enterprises.

CN119862878BActive Publication Date: 2025-07-15ANHUI CNBI SOFTWARE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510338638.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-15
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The existing disambiguation methods of person names are difficult to accurately identify the relationship between enterprises and personnel, and cannot effectively build a human-enterprise relationship database. The traditional methods lack sufficient information for precise disambiguation, so they cannot adapt to dynamic changes and real-time updates of data.

Method used

By obtaining cross-modal data related to enterprises and people, using entity alignment algorithms to form a structured data set of people-enterprises, generating semantic vectors based on pre-trained language models, combining co-occurrence frequency and spatiotemporal correlation characteristics, using adversarial neural networks and graph attention networks for real-time updates, and dynamically adjusting weights to improve disambiguation accuracy.

Benefits of technology

It significantly improves the accuracy and robustness of name disambiguation, can adapt to dynamically changing corporate and character data, solves the limitations of traditional methods in cross-modal data fusion and multi-dimensional feature mining, and improves the system's adaptability in dynamic data environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862878B_ABST
    Figure CN119862878B_ABST
Patent Text Reader

Abstract

The present invention provides a method for disambiguating personal names based on multi-relationship deep retrieval text matching, which relates to the technical field of personal name disambiguation. The present invention obtains cross-modal data related to enterprises and people, performs data alignment and data fusion on the cross-modal data through an entity alignment algorithm to form a person-enterprise structured data set, generates semantic vectors by establishing a multi-relationship deep retrieval model based on a pre-trained language model, calculates semantic similarity according to the semantic vectors, and generates a personal embedding vector according to the co-occurrence frequency and spatio-temporal correlation features calculated from the person-enterprise structured data set. An anti-disambiguation recognition model is established based on an adversarial neural network, personal name disambiguation is performed according to the personal embedding vector, a structure update model is established through a graph attention network to update the graph structure in real time, the confidence levels of semantic similarity and co-occurrence frequency are calculated according to Bayes' theorem, and the weights of semantic similarity and co-occurrence frequency are updated according to the confidence levels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of personal name disambiguation, and specifically provides a personal name disambiguation method based on multi-relationship deep retrieval text matching. Background Technique

[0002] In the past few decades, personal name disambiguation aims to eliminate the ambiguity of personal names in different environments, and the main task is to identify the real person corresponding to the personal name based on the personal name and other information. Existing enterprise and industrial and commercial publicity data often only involve personal names for personnel, without the identity information of the corresponding personnel, which makes it impossible for us to accurately identify the relationship between individuals and enterprises, thus hindering the construction of the person-enterprise relationship database and further posing challenges to commercial operations such as supply chain management and risk access screening. Due to the huge number of enterprises and personnel involved, existing personal name disambiguation methods are difficult to efficiently and accurately complete the task of personnel identification. Most traditional personal name disambiguation methods are based on the enterprise's own information and enterprise-related information, rarely considering personnel information, and thus it is impossible to identify the situation where there are actually two different people with the same name in the same company. Further, there will be problems of misidentification.

[0003] In the prior art, the publication number CN 114611516 A discloses a personal name disambiguation method, device, storage medium, and electronic device, which generate a first relationship graph with enterprises and natural persons as nodes based on the obtained enterprise relationship data; split the first relationship graph to obtain several second relationship graphs; train a preset graph network model based on the several second relationship graphs to obtain a vector representation model; determine whether the natural person to be identified needs to be merged with the homonymous node in the first relationship graph according to the associated nodes and the vector representation model or according to the attribute information of the natural person to be identified and the vector representation model. However, this solution mainly focuses on the relationship between enterprises and natural persons, but in personal name disambiguation, enterprise relationship data cannot cover all context information and lacks sufficient information for accurate disambiguation. At the same time, training a preset graph network model based on the split second relationship graphs, this method is a static model and cannot adapt to the dynamic changes and real-time updates of data.

[0004] The above information disclosed in the background art section is only used to strengthen the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of the present invention is to provide a personal name disambiguation method based on multi-relationship deep retrieval text matching to solve the problems raised in the above background art.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A method for person name disambiguation based on multi-relation deep retrieval text matching, the specific steps include:

[0008] Step 1: Obtain cross-modal data related to enterprises and people, form an initial enterprise-person dataset based on person names and enterprise names, perform data alignment on the cross-modal data through an entity alignment algorithm, and perform data fusion on the aligned data to form an enterprise-person structured dataset;

[0009] Step 2: According to the enterprise-person structured dataset, generate semantic vectors through a multi-relation deep retrieval model established based on a pre-trained language model, and calculate semantic similarity based on the semantic vectors;

[0010] Step 3: Calculate the co-occurrence frequency and spatio-temporal correlation features according to the enterprise-person structured dataset, generate person embedding vectors based on semantic similarity, co-occurrence frequency, and spatio-temporal correlation features, establish a disambiguation recognition model based on an adversarial neural network, and perform person name disambiguation based on the person embedding vectors;

[0011] Step 4: Through the obtained real-time updated enterprise-person structured dataset and person embedding vectors, establish a structure update model through a graph attention network, and capture the high-order relationship changes between each node and node relationships with person names as nodes to real-time update the graph structure;

[0012] Step 5: Through the analysis of semantic similarity and co-occurrence frequency, calculate the confidence degrees of semantic similarity and co-occurrence frequency according to Bayes' theorem, and update the weights of semantic similarity and co-occurrence frequency according to the confidence degrees.

[0013] Furthermore, the ways to obtain the cross-modal data are respectively: through the enterprise publicity system, through social software, and news aggregation platforms;

[0014] The cross-modal data includes: text data, structured data, image data, and geographical information data;

[0015] The method for establishing the initial enterprise-person dataset is:

[0016] The initial enterprise-person dataset is expressed as , where each element represents data related to a person name and an enterprise, including a person name and multiple related modal information, and the mathematical expression form is:

[0017]

[0018] Among them, is the initial enterprise-person dataset, is the number of data related to person names and enterprise names in the initial enterprise-person dataset;

[0019] Among them, each element The expression form is as follows:

[0020]

[0021] Among them, is the data related to the names of individuals and enterprise names in the initial dataset of the personal enterprise, which are the names of individuals and enterprise names, text features, structured data, image data, and geographic information data respectively.

[0022] Furthermore, the calculation method for data alignment of cross-modal data by the entity alignment algorithm is as follows:

[0023]

[0024] Among them, is the internal data type of the person-enterprise name data, , , is the cosine similarity;

[0025] If is true, it is determined that the two data types are similar due to the internal factors of the same person-enterprise name data;

[0026] Among them, is the similarity judgment threshold.

[0027] Furthermore, the calculation formula for data fusion of the aligned data is as follows:

[0028]

[0029] Among them, is the data related to the names of individuals and enterprise names in the personal enterprise initial dataset after fusion, are the weights of text features, structured data, image data, and geographic information data respectively, .

[0030] Furthermore, the specific steps for generating semantic vectors are as follows:

[0031] Use [CLS] and [SEP] to mark the text boundaries in the person-enterprise structured dataset, separate fields with commas, and use RoBERTa to extract the semantic vectors marked by [CLS]:

[0032]

[0033] Among them, , The person names and enterprise names in the person-enterprise structured dataset respectively, is the language training model, is the sentence-level representation vector, , are the semantic vectors of the person name and the enterprise name respectively.

[0034] Furthermore, the calculation formula for the semantic similarity is:

[0035]

[0036] where, is the semantic similarity between the person name and the enterprise name.

[0037] Furthermore, the calculation formula for the co-occurrence frequency is:

[0038]

[0039] where, is the co-occurrence frequency, is the number of times the person name and the enterprise name appear in the person-enterprise structured dataset, is the number of times the person name appears in the person-enterprise structured dataset, is the number of times the enterprise name appears in the person-enterprise structured dataset;

[0040] The spatio-temporal correlation feature:

[0041]

[0042] where, is the time correlation feature, , are the start time and end time of the association between the person name and the enterprise name respectively, , are the start time and end time of the association between the enterprise name and the person name respectively;

[0043] where, is the spatial correlation feature, are the latitude and longitude of the person name respectively, are the latitude and longitude of the enterprise name respectively, is the radius of the earth, is the spherical distance between the person name and the enterprise name;

[0044] The calculation formula for generating the person embedding vector is:

[0045] where, is the combined feature vector, is the person embedding vector, is the projection matrix, is the bias term.

[0046] Furthermore, the disambiguation recognition model is established based on the adversarial neural network, and name disambiguation is performed according to the person embedding vector:

[0047] The disambiguation recognition model is established based on the adversarial neural network. By inputting the person embedding vectors of different historical names into the recognition model, whether different names refer to the same person is used as the result of the recognition model to train the recognition model. The person embedding vector of the name to be recognized is used as the input of the trained recognition model to perform name recognition and disambiguation:

[0048] The disambiguation recognition model includes a generator and a discriminator:

[0049] The generator maps the input to the output of the name disambiguation result through a neural network, and the calculation formula is:

[0050]

[0051] where, is the name disambiguation result, are the person embedding vectors of different names, are the parameters of the generator, is the mapping function of the generator;

[0052] The goal of the discriminator is to judge whether the name disambiguation result is generated by the generator, and based on this judgment, it gives a probability value:

[0053]

[0054] where, represents the output name disambiguation result, are the weights of the discriminator, is the bias of the discriminator.

[0055] Furthermore, the calculation formula for establishing the structure update model through the graph attention network is:

[0056]

[0057] where, is the feature vector of the node after update, is the set of neighbor nodes of the node, is the attention weight between the neighbor nodes of the node, is the self-feature of the node, is the activation function;

[0058] The calculation method of the attention weight between the neighbor nodes of the node is:

[0059]

[0060] Among them, is the attention score between nodes, is a learnable parameter vector, is the transpose operation, are the node feature vectors respectively;

[0061] The calculation formula for the attention score between nodes is:

[0062]

[0063] Among them, represents the th neighbor node of the node, represents the th attention score of the neighbor node.

[0064] Furthermore, the specific steps for calculating the confidence of semantic similarity and co-occurrence frequency according to Bayes' theorem and updating the weights of semantic similarity and co-occurrence frequency according to the confidence are:

[0065] The calculation formula for the prior probability is:

[0066]

[0067] Among them, in is the prior probability of semantic similarity when in is the prior probability of co-occurrence frequency when is the number of correct disambiguation times of semantic similarity and co-occurrence frequency, is the total number of disambiguation times of semantic similarity and co-occurrence frequency;

[0068] The calculation formula for the likelihood probability is:

[0069]

[0070] Among them, is the probability of correct disambiguation through semantic similarity and co-occurrence frequency, is the semantic similarity and co-occurrence probability of person names and enterprise names during the current disambiguation, is the historical semantic similarity and co-occurrence probability of person names and enterprise names during the disambiguation, is the variance of the semantic similarity and co-occurrence probability of person names and enterprise names during the disambiguation;

[0071] The posterior probability is:

[0072]

[0073] Among them, is the confidence level, is a normalization constant;

[0074] The calculation formula for updating the weights is:

[0075]

[0076] Among them, is the updated weight for semantic matching and co-occurrence analysis. When is true, is indicating the probability of correct disambiguation of semantic similarity. When is true, is indicating the probability of correct disambiguation of co-occurrence frequency.

[0077] Compared with the prior art, the beneficial effects of the present invention are:

[0078] By obtaining cross-modal data related to enterprises and people, aligning and fusing the cross-modal data through an entity alignment algorithm to form a person-enterprise structured data set, generating semantic vectors through a multi-relationship deep retrieval model based on a pre-trained language model, calculating semantic similarity according to the semantic vectors, and generating person embedding vectors according to the co-occurrence frequency and spatio-temporal correlation features calculated from the person-enterprise structured data set, establishing a disambiguation recognition model based on an adversarial neural network, performing name disambiguation according to the person embedding vectors, establishing a structure update model through a graph attention network to update the graph structure in real time, calculating the confidence levels of semantic similarity and co-occurrence frequency according to Bayes' theorem, and updating the weights of semantic similarity and co-occurrence frequency according to the confidence levels;

[0079] Based on the fusion of cross-modal data from multiple channels, the multi-relationship deep retrieval technology based on a pre-trained language model, and the structure update ability of a graph neural network, the present invention can significantly improve the accuracy and robustness of name disambiguation. First, the solution uses an entity alignment algorithm and cross-modal data fusion to form a high-quality person-enterprise structured data set, providing a more comprehensive information basis for subsequent disambiguation tasks. Second, through deep learning and semantic vector generation, the method can capture the deep semantic relationships between people, and combined with the features of spatio-temporal correlation and co-occurrence frequency, further optimize the matching and disambiguation effects of names. In particular, the disambiguation model based on an adversarial neural network makes the system more accurate in dealing with complex and ambiguous names;

[0080] The present invention also dynamically updates the structured data through a graph attention network, which can capture the changes in the relationships between enterprises and individuals in real time, ensuring that the disambiguation model remains efficient and accurate in an ever-changing information environment, better adapting to the dynamically changing enterprise and individual data, and solving the limitations of the static model in Solution 1. Generally speaking, through the combination of various technical means, this solution not only solves the limitations of traditional methods in cross-modal data fusion and multi-dimensional feature mining, but also improves the adaptability of the system in a dynamic data environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 It is a schematic diagram of the overall method flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0082] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments.

[0083] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs. The "first", "second", and similar terms used in the present invention do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0084] Embodiment:

[0085] Please refer to Figure 1 , the present invention provides a technical solution:

[0086] A method for disambiguating personal names based on multi-relationship deep retrieval text matching, the specific steps include:

[0087] Step 1: Obtain cross-modal data related to enterprises and individuals, form an initial enterprise-person dataset based on personal names and enterprise names, perform data alignment on the cross-modal data through an entity alignment algorithm, and perform data fusion on the aligned data to form an enterprise-person structured dataset.

[0088] Cross-modal data is obtained through multiple channels such as enterprise publicity systems, social software, and news aggregation platforms, covering a variety of data forms, including text, structured data, image data, and geographic information data. Compared with existing technologies, the advantage of this method is that it can comprehensively and systematically collect and combine information from multiple platforms, thus providing more dimensional information for subsequent disambiguation tasks. The data provided by these channels not only has rich sources but also can show different contexts and associations in different backgrounds, making the data set more comprehensive and capable of accurately describing the relationship between person names and enterprises. This solution can supplement and verify information between different modalities, thereby improving the integrity of the data and the accuracy of disambiguation. In particular, image data and geographic information data can provide additional context information for the disambiguation task, further reducing the probability of mis-matching.

[0089] The cross-modal data is aligned through an entity alignment algorithm, ensuring that data from different sources and modalities can be correctly matched and unified, reducing information redundancy and ensuring information consistency in the disambiguation task. Alignment is performed through the cosine distance calculation method, which helps to accurately evaluate the similarity between different modalities. Especially in the relationship matching between text data and structured data, the cosine distance, as a common similarity metric method, can effectively measure semantic similarity and has high computational efficiency. Compared with traditional alignment methods based on simple rules or static matching, the cosine distance calculation method is more efficient and accurate when dealing with cross-modal data, avoiding information loss or mis-matching.

[0090] In this embodiment, the methods for obtaining the cross-modal data are respectively: through enterprise publicity systems, through social software, and news aggregation platforms;

[0091] The cross-modal data includes: text data, structured data, image data, and geographic information data;

[0092] The method for establishing the initial person-enterprise data set is:

[0093] The initial person-enterprise data set is represented as , where each element represents data related to a person name and an enterprise, containing a person name and multiple related modal information, and its mathematical expression form is:

[0094]

[0095] Among them, is the initial person-enterprise data set, is the number of data related to person names and enterprise names in the initial person-enterprise data set;

[0096] Among them, each element represents the expression form of:

[0097]

[0098] Among them, is the data related to the name of a person and the name of an enterprise in the initial dataset of the th personal enterprise, which are respectively the name of a person and the name of an enterprise, text features, structured data, image data, and geographic information data;

[0099] The calculation method for data alignment of cross-modal data through the entity alignment algorithm is as follows:

[0100]

[0101] Among them, is the internal data type of the person-enterprise name data, , , is the cosine similarity;

[0102] If , it is determined that two data types are similar as an internal factor of the same person-enterprise name data;

[0103] Among them, is the similarity judgment threshold.

[0104] In this embodiment, the calculation formula for data fusion of the aligned data is as follows:

[0105]

[0106] Among them, is the data related to the name of a person and the name of an enterprise in the th personal enterprise initial dataset after fusion, are respectively the weights of text features, structured data, image data, and geographic information data, , , , , .

[0107] After data alignment, different modalities of data are combined through a data fusion process to generate a structured dataset. The structured dataset provides a standardized data format, enabling subsequent deep retrieval and name disambiguation algorithms to perform calculations on a more efficient basis. Structured data has stronger operability and consistency compared to the original unstructured data, which can improve the performance of subsequent models and reduce the impact of noise. In the existing technology, many systems rely on a single data source or fail to effectively fuse different modalities of data, so the accuracy and scope of application are relatively limited. In contrast, the structured dataset can more accurately reflect the multi-dimensional relationship between names and enterprises. Through the fusion and structuring of multi-modal data, the disambiguation model can perform deep retrieval and semantic analysis on a more efficient and accurate basis. This process not only improves the learning efficiency of the model but also makes the disambiguation results more accurate. Especially when dealing with complex name and enterprise relationships, the model can effectively identify potential matches and incorrect information.

[0108] Step 2: According to the person-enterprise structured dataset, a multi-relation deep retrieval model is established based on a pre-trained language model to generate semantic vectors, and semantic similarity is calculated based on the semantic vectors.

[0109] By using a pre-trained language model to generate semantic vectors and calculating semantic similarity, the semantic information of the text can be efficiently captured. RoBERTa can provide higher-quality representation vectors and can efficiently capture deeper semantic information. Especially for the complex multiple relationships between names and enterprises, the model can more accurately perform matching and disambiguation at the semantic level. By using the [CLS] token as the representative of the sentence and the [SEP] token to separate text fields, accurate context embeddings can be generated, enhancing the understanding of the relationships and meanings between texts.

[0110] Semantic vectors are a representation form that converts text into high-dimensional vectors through natural language processing models. These vectors are numerical expressions of the semantics of the text and can capture rich information about vocabulary, syntax, and context in the text. Semantic vectors not only represent the literal meaning of individual words but also reflect the semantic meaning in the context. Such vectors are usually high-dimensional and learn the deep connections between vocabulary and sentences through pre-trained language models, which can better express the diversity and complexity in language.

[0111] In the scenario of name disambiguation, semantic vectors can represent the semantic relationships between each candidate name and enterprise names, addresses, or other context information, thereby helping to identify the similarity between the two.

[0112] In this embodiment, the specific steps for generating the semantic vectors are as follows:

[0113] Use [CLS] and [SEP] tokens to define text boundaries in the person-enterprise structured dataset, separate fields with commas, and use RoBERTa to extract the semantic vector marked by [CLS]:

[0114]

[0115] Among them, , are the person name and enterprise name in the person-enterprise structured dataset respectively, is the language training model, is the sentence-level representation vector, , are the semantic vectors of the person name and enterprise name respectively.

[0116] In this embodiment, the calculation formula of the semantic similarity is:

[0117]

[0118] Among them, is the semantic similarity between the person name and the enterprise name.

[0119] By calculating the similarity between the generated semantic vectors, the semantic similarity between different texts can be compared more precisely. Traditional text similarity calculation methods (such as TF-IDF based on word frequency or traditional vector space models) are difficult to capture the deep semantic relationships in the text, while the semantic vectors generated using pre-trained models can identify more detailed similarities. For example, some enterprise names may have similar naming habits, and RoBERTa can identify these potential similarities and make more accurate judgments in the disambiguation process.

[0120] Semantic similarity measures the semantic proximity between two texts. In person name disambiguation, semantic similarity is used to calculate the matching degree between candidate person names and information such as enterprise names and addresses. For example, by calculating the semantic similarity between different candidate person names and a certain enterprise name, it can be judged whether the person name is associated with the enterprise or belongs to the employees of the enterprise.

[0121] Step 3: Calculate the co-occurrence frequency and spatio-temporal correlation features according to the person-enterprise structured dataset, generate a person embedding vector based on the semantic similarity, co-occurrence frequency, and spatio-temporal correlation features, establish a disambiguation recognition model based on the adversarial neural network, and perform person name disambiguation according to the person embedding vector.

[0122] The co-occurrence frequency refers to the frequency of two entities, such as a person's name and a company name, appearing together in the same text. In person name disambiguation, assuming a person's name co-occurs with certain companies or locations in multiple texts, this co-occurrence relationship can serve as evidence of the potential association between the person's name and the enterprise. Persons' names and enterprises with a higher co-occurrence frequency may have a stronger connection in reality.

[0123] The spatio-temporal correlation feature refers to analyzing the correlation between a person's name and other entities based on information in the time and space dimensions. For example, a person's name may have a closer cooperation with an enterprise at a specific time (such as a certain year), or the person's name may only appear in data of a specific location (such as a certain city or region). By combining time and geographical information, spatio-temporal features can help distinguish homonymous individuals who frequently appear in different time periods or locations.

[0124] In this embodiment, the calculation formula for the co-occurrence frequency is:

[0125]

[0126] Where, is the co-occurrence frequency, is the number of times the person's name and the enterprise name appear in the person-enterprise structured dataset, is the number of times the person's name appears in the person-enterprise structured dataset, is the number of times the enterprise name appears in the person-enterprise structured dataset;

[0127] The spatio-temporal correlation feature:

[0128]

[0129] Where, is the time correlation feature, , are respectively the start time and end time of the association between the person's name and the enterprise name, , are respectively the start time and end time of the association between the enterprise name and the person's name;

[0130] Where, is the spatial correlation feature, are respectively the latitude and longitude of the person's name, are respectively the latitude and longitude of the enterprise name, is the radius of the earth, is the spherical distance between the person's name and the enterprise name.

[0131] In this embodiment, the calculation formula for generating the person embedding vector is:

[0132]

[0133] Among them, is the combined feature vector, is the person embedding vector, is the projection matrix, is the bias term.

[0134] The person embedding vector is a vector representation comprehensively constructed based on semantic similarity, co-occurrence frequency, and spatio-temporal correlation features. This embedding vector fuses the multi-dimensional features of each candidate person name to form a multi-dimensional vector representation. It not only contains the basic semantic information of the person name (such as through the learning of semantic vectors), but also combines co-occurrence information and spatio-temporal relationships, and can comprehensively reflect the performance of this person name in different contexts.

[0135] In this embodiment, the disambiguation recognition model is established based on the adversarial neural network, and person name disambiguation is performed according to the person embedding vector:

[0136] The disambiguation recognition model is established based on the adversarial neural network. By inputting the person embedding vectors of different historical person names into the recognition model, whether different person names are the same person is used as the result of the recognition model to train the recognition model. The person embedding vector of the person name to be recognized is used as the input of the trained recognition model to perform recognition and disambiguation of the person name:

[0137] The disambiguation recognition model includes a generator and a discriminator:

[0138] The generator maps the input to the output of the person name disambiguation result through a neural network, and the calculation formula is:

[0139]

[0140] Among them, is the person name disambiguation result, are the person embedding vectors of different person names, are the parameters of the generator, is the mapping function of the generator;

[0141] The goal of the discriminator is to judge whether the person name disambiguation result is generated by the generator, and based on this judgment, it gives a probability value:

[0142]

[0143] Among them, represents the output person name disambiguation result, are the weights of the discriminator, is the bias of the discriminator.

[0144] An adversarial neural network is a neural network architecture that enhances the robustness of a model through adversarial training. Its basic idea is to construct two networks: a generative network and a discriminative network, and train them against each other to optimize the disambiguation effect. In this solution, the generative network may generate candidate disambiguation results, while the discriminative network evaluates the correctness of these results, ultimately forming an optimized disambiguation model.

[0145] Step 4: Using the obtained real-time updated person-enterprise structured dataset and person embedding vectors, establish a structure update model through a graph attention network. Taking person names as nodes, capture the high-order relationship changes between each node and node relationships to update the graph structure in real time.

[0146] A graph attention network is a deep learning model based on graph structure. It uses the attention mechanism to assign different weights to different adjacent nodes in the graph, thereby dynamically adjusting the flow of information according to the relationship strength between nodes during the information propagation process. Specifically for the problem of person name disambiguation, the graph attention network can model the relationships between nodes (person names) and adjacent nodes (such as companies, positions, etc.) in the graph, and weight these relationships through the attention mechanism to effectively capture high-order relationship changes.

[0147] The graph structure is dynamically updated in this step, which means that with the introduction of new structured data (such as the relationship data between companies and people), the graph structure will also be updated accordingly. This means that as the data changes, the node relationships in the graph will reflect new correlations in real time, being able to effectively handle changes and new information in the data, further enhancing the disambiguation effect.

[0148] The introduction of the real-time updated person-enterprise structured dataset and person embedding vectors means that we can timely capture the new relationships between person names and companies, as well as the interactions between people, especially in an environment where people or companies change frequently. Traditional static datasets often lead to information lag, affecting the model's effect. By updating the graph structure in real time, it can ensure that the graph structure and person embedding vectors always remain up-to-date, greatly enhancing the disambiguation effect. Graph neural networks can integrate global information from different nodes, which enables disambiguation not to rely solely on a single node or a single feature, but to comprehensively consider the relationship between a person name and other relevant nodes in the graph (such as companies, geographical locations, time, etc.). This graph-based modeling method is more comprehensive than traditional single context-based matching methods and can effectively solve the ambiguity phenomenon. Especially when faced with homonymous people or polysemous person names, graph neural networks can effectively distinguish through the relationships of adjacent nodes.

[0149] In this embodiment, the calculation formula for establishing a structure update model through a graph attention network is:

[0150]

[0151] Among them, is the feature vector after node update, is the set of neighbor nodes of the node, is the attention weight between the neighbor nodes of the node, is the self - feature of the node, is the activation function;

[0152] The calculation method of the attention weight between the neighbor nodes of the node is:

[0153]

[0154] Among them, is the attention score between node and node, is the learnable parameter vector, is the transpose operation, are the node feature vectors respectively;

[0155] The calculation formula of the attention score between node and node is:

[0156]

[0157] Among them, represents the th neighbor node of the node, represents th attention score of the neighbor node.

[0158] Step 5: Through the analysis of semantic similarity and co - occurrence frequency, calculate the confidence of semantic similarity and co - occurrence frequency according to Bayes' theorem, and update the weights of semantic similarity and co - occurrence frequency according to the confidence.

[0159] Bayes' theorem provides a method to calculate the probability of an event occurrence by combining prior knowledge and observed data. In this application, through Bayes' theorem, we can calculate the confidence of semantic similarity and co - occurrence frequency, and use these confidences to dynamically update the weights, so as to optimize the disambiguation process.

[0160] Combine semantic similarity and co - occurrence frequency, and use Bayes' theorem for dynamic update. This can not only evaluate the similarity between personal names from the semantic level, but also verify whether these personal names may refer to the same entity through the statistical information of co - occurrence frequency. Such multi - dimensional evaluation makes the disambiguation process more accurate and reliable.

[0161] This solution enables the model to adapt to data changes by dynamically updating weights. As new information enters the system, the system can recalculate the confidence levels of various factors through Bayes' theorem and flexibly adjust the weights according to the current context. This mechanism of dynamically adjusting weights helps to handle complex and variable text scenarios, especially in the case of homonymous or polysemous names, where it can more accurately distinguish different individuals.

[0162] In this embodiment, the specific steps for calculating the confidence levels of semantic similarity and co-occurrence frequency according to Bayes' theorem and updating the weights of semantic similarity and co-occurrence frequency are as follows:

[0163] The formula for the prior probability is:

[0164]

[0165] where in when it is the prior probability of semantic similarity, in when it is the prior probability of co-occurrence frequency, is the number of correct disambiguation times for semantic similarity and co-occurrence frequency, is the total number of disambiguation times for semantic similarity and co-occurrence frequency;

[0166] The formula for the likelihood probability is:

[0167]

[0168] where is the probability of correct disambiguation through semantic similarity and co-occurrence frequency, is the semantic similarity and co-occurrence probability of the person name and enterprise name during the current disambiguation, is the historical semantic similarity and co-occurrence probability of the person name and enterprise name during the disambiguation, is the variance of the semantic similarity and co-occurrence probability of the person name and enterprise name during the disambiguation;

[0169] The posterior probability is:

[0170]

[0171] where is the confidence level, is a normalization constant;

[0172] The formula for updating the weights is:

[0173]

[0174] where For updating weights in semantic matching and co-occurrence analysis. When then is which represents the probability of correct disambiguation of semantic similarity. When then is which represents the probability of correct disambiguation of co-occurrence frequency.

[0175] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0176] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software methods depends on the specific application and design constraints of the technical solution.

[0177] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, and may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0178] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered within the protection scope of this application.

Claims

1. A method for name disambiguation based on multi-relation deep retrieval text matching, characterized in that, The specific steps include: Step 1: Obtain cross-modal data related to enterprises and people, form an initial enterprise-person dataset based on person names and enterprise names, perform data alignment on the cross-modal data through an entity alignment algorithm, and perform data fusion on the aligned data to form an enterprise-person structured dataset; Step 2: According to the enterprise-person structured dataset, generate semantic vectors by establishing a multi-relation deep retrieval model based on a pre-trained language model, and calculate semantic similarity based on the semantic vectors; Step 3: Calculate the co-occurrence frequency and spatio-temporal correlation features according to the enterprise-person structured dataset, generate person embedding vectors based on semantic similarity, co-occurrence frequency, and spatio-temporal correlation features, establish a disambiguation recognition model based on an adversarial neural network, and perform person name disambiguation based on the person embedding vectors; Step 4: Through the obtained real-time updated enterprise-person structured dataset and person embedding vectors, establish a structure update model through a graph attention network, use person names as nodes, and capture high-order relationship changes between each node and node relationships to update the graph structure in real time; Step 5: Through the analysis of semantic similarity and co-occurrence frequency, calculate the confidence levels of semantic similarity and co-occurrence frequency according to Bayes' theorem, and update the weights of semantic similarity and co-occurrence frequency according to the confidence levels.

2. The method for name disambiguation based on multi-relationship depth retrieval text matching according to claim 1, wherein: The methods for obtaining the cross-modal data are respectively: through the enterprise publicity system, through social software, and news aggregation platforms; The cross-modal data includes: text data, structured data, image data, and geographic information data; The method for establishing the initial enterprise-person dataset is: The initial dataset of people and enterprises is represented as , where each element represents the data related to a person's name and an enterprise, including a person's name and multiple related modal information. The mathematical expression form is as follows: Among them, is the initial data set of people and enterprises, is the quantity of data related to personal names and enterprise names in the initial data set of people and enterprises; Among them, each element The represented expression form is: Among them, is the data related to the names of people and enterprise names in the initial data set of the personal enterprise, which are respectively the names of people and enterprise names, text features, structured data, image data, and geographic information data.

3. The method for disambiguating personal names based on multi-relationship depth retrieval text matching according to claim 2, wherein: The calculation method for performing data alignment on the cross-modal data through the entity alignment algorithm is: Among them, is the internal data type of the enterprise name data, , , is the cosine similarity; If When, determining that two data types are similar is an internal factor of the enterprise name data of the same person; wherein, is the similarity judgment threshold value.

4. The method for disambiguating personal names based on multi-relationship depth retrieval text matching according to claim 3, wherein: The calculation formula for performing data fusion on the aligned data is: Among them, is the relevant data of the name of the person and the enterprise name in the initial data set of the merged .

5. A method for disambiguating personal names based on multi-relationship depth retrieval text matching according to claim 1, characterized in that: The specific steps for generating the semantic vectors are: Use [CLS] and [SEP] to mark the text boundaries in the enterprise-person structured dataset, separate fields with commas, and use RoBERTa to extract the semantic vectors marked by [CLS]; Among them, , are the person name and enterprise name in the enterprise-person structured dataset respectively, is the language training model, is the sentence-level representation vector, , are the semantic vectors of the person name and enterprise name respectively.

6. The method for personal name disambiguation based on multi-relationship depth retrieval text matching according to claim 5, characterized in that: The calculation formula for the semantic similarity is: Among them, is the semantic similarity of personal names and company names.

7. A method for disambiguating personal names based on multi-relationship depth retrieval text matching according to claim 1, characterized in that: The calculation formula for the co-occurrence frequency is: Among them, is the co-occurrence frequency, is the number of times a person's name and a company name appear in the person-enterprise structured dataset, is the number of times a person's name appears in the person-enterprise structured dataset, is the number of times a company name appears in the person-enterprise structured dataset; The spatio-temporal correlation features: Among them, is the time correlation feature, , are respectively the start time and end time of the association between the person name and the enterprise name, , are respectively the start time and end time of the association between the enterprise name and the person name; Among them, is the spatial correlation characteristic, are the latitude and longitude of the person's name respectively, are the latitude and longitude of the enterprise name respectively, is the radius of the earth, is the spherical distance between the person's name and the enterprise name; The calculation formula for generating the person embedding vectors is: Among them, is the combined feature vector, is the person embedding vector, is the projection matrix, is the bias term.

8. A person name disambiguation method based on multi-relation deep retrieval text matching according to claim 1, characterized in that: Establish a disambiguation recognition model based on an adversarial neural network, and perform person name disambiguation based on the person embedding vectors: Establish a disambiguation recognition model based on an adversarial neural network, input the person embedding vectors of different historical person names into the recognition model, use whether different person names are the same person as the result of the recognition model to train the recognition model, and use the person embedding vectors of the person name to be recognized as the input of the trained recognition model to perform recognition and disambiguation on the person name: The disambiguation recognition model includes a generator and a discriminator: The generator maps the input to the output of the person name disambiguation result through a neural network, and the calculation formula is: Among them, is the result of person name disambiguation, is the person embedding vector of different person names, is the parameter of the generator, is the mapping function of the generator; The goal of the discriminator is to judge whether the person name disambiguation result is generated by the generator, and based on this judgment, it gives a probability value: Among them, represents the output result of name disambiguation, is the weight of the discriminator, is the bias of the discriminator.

9. A method for disambiguating personal names based on multi-relationship deep retrieval text matching according to claim 1, characterized in that: The calculation formula for establishing the structure update model through the graph attention network is: Among them, is the feature vector after node update, is the set of neighbor nodes of the node, is the attention weight between the neighbor nodes of the node, is the self - feature of the node, is the activation function; The calculation method for the attention weights between the neighbor nodes of the node is: Among them, is the attention score between nodes, is a learnable parameter vector, is the transpose operation, are node feature vectors respectively; The calculation formula for the attention score between nodes is as follows: Among them, represents the th neighbor node of the node, represents the th attention score of the neighbor node.

10. A method for name disambiguation based on multi-relationship depth retrieval text matching according to claim 1, characterized in that: The specific steps for calculating the semantic similarity, the confidence of the co-occurrence frequency according to Bayes' theorem, and updating the weights of the semantic similarity and the co-occurrence frequency are as follows: The calculation formula for the prior probability is: Among them, in is the prior probability of semantic similarity when in is the prior probability of co-occurrence frequency when is the number of correct disambiguations of semantic similarity and co-occurrence frequency, is the total number of disambiguations of semantic similarity and co-occurrence frequency; The calculation formula for the likelihood probability is: Among them, is the probability of correct disambiguation through semantic similarity and co-occurrence frequency, is the semantic similarity and co-occurrence probability of personal names and company names during current disambiguation, is the historical semantic similarity and co-occurrence probability of personal names and company names during disambiguation, is the variance of the semantic similarity and co-occurrence probability of personal names and company names during disambiguation; The posterior probability is: wherein, is the confidence level, is a normalization constant; The calculation formula for updating the weights is: Among them, is the updated weight for semantic matching and co-occurrence analysis. When is the case, is which represents the probability of correct disambiguation of semantic similarity. When is the case, is which represents the probability of correct disambiguation of co-occurrence frequency.

Citation Information

Patent Citations

  • Name disambiguation method and device, storage medium and electronic equipment

    CN114611516A

  • Homonymous author disambiguation method based on network representation and semantic representation

    CN111191466A

  • Intelligent multi-mode fusion electricity utilization inspection method, device, system, equipment and medium

    CN115114448A