Short text entity disambiguation-oriented candidate entity secondary screening method for multi-factor text characteristic fusion

By employing a two-stage candidate entity screening strategy that integrates multiple textual features, the problems of candidate entity redundancy and noise interference in short text entity disambiguation are solved, enabling the screening of high-quality candidate entity sets and improving the accuracy and robustness of entity disambiguation.

CN121009969APending Publication Date: 2025-11-25LIAONING UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511117668.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies for short text entity disambiguation suffer from redundancy in candidate entity sets and noise interference, making it difficult to effectively select high-quality candidate entity sets.

Method used

A two-stage candidate entity screening strategy based on multi-factor text feature fusion is adopted, including coarse screening and fine screening stages. Coarse screening utilizes the Wikipedia knowledge base and multi-dimensional feature filtering, while fine screening optimizes the candidate entity set by combining multi-dimensional feature measurement and comprehensive similarity scoring with information entropy, IDF, and word vector similarity.

Benefits of technology

It significantly reduces the size of the candidate entity set, improves the quality and relevance of the candidate entity set, and enhances the accuracy and robustness of short text entity disambiguation, especially showing strong adaptability when dealing with polysemy and new entities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009969A_ABST
    Figure CN121009969A_ABST
Patent Text Reader

Abstract

The invention relates to a short text entity disambiguation-oriented multi-factor text characteristic fusion candidate entity secondary screening method, and belongs to the field of entity disambiguation. According to actual application requirements of short text entity disambiguation, in order to further simplify the scale of candidate entities, candidate entity screening is divided into two stages of coarsening screening and refining screening. The method comprises the following steps: firstly, in a coarsening and screening stage, performing preliminary screening on candidate entities by utilizing a Wikipedia knowledge base and considering indexes such as a context local matching degree and an entity correlation degree; secondly, in the refining and screening stage, a keyword extraction method of multi-dimensional feature measurement is provided, prior information is introduced to calculate the similarity between candidate entities and entity references, and refining and screening of the candidate entities are completed through comprehensive similarity scores of the candidate entities.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of entity disambiguation, and relates to a candidate entity secondary screening strategy, in particular to a candidate entity secondary screening method for short text entity disambiguation based on multi-factor text feature fusion. BACKGROUND

[0002] With the iterative evolution of mobile Internet technology, the global data size is growing exponentially. According to the statistical prediction of Statista, the total amount of global data will reach 180 ZB level in 2025. In the face of massive non-standard text data, people need to accurately extract semantic entities and abstract concepts from the data to realize the structured representation of knowledge. Traditional information extraction technology is limited by the ability of shallow semantic understanding, and it is difficult to meet the demand of high-quality knowledge acquisition, which has given rise to the rise of knowledge graph and various natural language processing technologies. Structured semantic network promotes the improvement of cognitive ability in artificial intelligence by modeling the complex correlation between entities, and thus shows significant value in intelligent question answering, personalized recommendation and other scenarios.

[0003] Studies have shown that traditional manual annotation methods rely on expert experience and consume a lot of resources, and it is difficult to ensure the consistency of the annotation results. Moreover, simple extraction of entity surface information cannot completely solve the ambiguity problem of semantics. Therefore, the research on short text entity disambiguation technology has become one of the important research tasks in the field of natural language processing.

[0004] In the entity disambiguation task, candidate entity screening as a core preprocessing link, its research course reflects the paradigm shift of natural language processing technology from rule-driven to data-driven. Early research mainly relied on artificial rules and shallow feature matching, such as LinkCount algorithm which generates candidate entity set through string similarity, word frequency statistics and simple context co-occurrence pattern. This kind of method can realize the basic disambiguation function, but it is limited by the fixedness of rules and the surface of features, and it is difficult to deal with entity alias, abbreviation and cross-domain semantic drift problem.

[0005] With the development of statistical machine learning, the Concept Vector algorithm constructs a semantic vector by extracting keywords or Wikipedia concepts, gradually introduces a probability model and feature engineering, and uses models such as maximum entropy and support vector machine to fuse context bag-of-words features, entity popularity and type consistency indicators, and optimizes candidate sorting through supervised learning. However, traditional machine learning methods still face the challenges of subjective feature design and sparse data, especially in the short text scenario, the limited context leads to insufficient feature representation, which restricts the generalization ability of the model.

[0006] The popularity of knowledge base has driven the innovation of candidate entity selection technology. Based on structured knowledge graph, such as YAGO and Freebase, the semantic enhanced candidate generation framework is constructed by entity attribute association, ontology level constraint and relationship path reasoning. This kind of method uses the explicit semantic association between entities in the knowledge base, combines with the graph traversal algorithm to select the candidate entity that meets the logical constraint, and significantly improves the semantic rationality of the selection. However, the update lag of static knowledge base and the incomplete coverage of the field limit its applicability in dynamic language environment and emerging entity scene.

[0007] Researchers have solved some of the problems existing in current candidate entity selection using deep learning technology. Based on the representation learning method of neural network (such as word vector and sentence vector), the implicit association between entity reference and candidate entity is automatically captured through distributed semantic coding. The sequence model enables the model to dynamically focus on key context words, enhancing the sensitivity to local semantics. SUMMARY

[0008] In order to solve the problem of redundancy and noise interference of existing candidate entity set for short text entity disambiguation, the present application provides a multi-factor text feature fusion candidate entity secondary screening strategy for short text entity disambiguation, which can effectively reduce the size of the candidate entity set and optimize and improve the quality of the entity candidate set.

[0009] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows: a multi-factor text feature fusion candidate entity secondary screening method for short text entity disambiguation, which is divided into two stages of rough screening and refined screening. In the rough screening stage, the Wikipedia knowledge base is used to consider the local matching degree of context and the entity correlation degree and other indicators to preliminarily screen the candidate entities. In the refined screening stage, a multi-dimensional feature measurement keyword extraction method is proposed, and prior information is introduced to calculate the similarity between the candidate entity and the entity reference, and the refined screening of the candidate entity is completed through the comprehensive similarity score of the candidate entity.

[0010] Rough screening of candidate entity based on Wikipedia knowledge base: morphological normalization is realized through morphological merging technology, and variant expansion is realized relying on the alias index of knowledge base, so that the entities corresponding to the alias and variant of the entity reference are also used as initial candidate entities.

[0011] The morphological normalization is realized by the word form reduction technology, and the variant expansion is performed relying on the alias index of the knowledge base (Wikipedia knowledge base redirection) to expand the entity corresponding to the alias and variant of the entity designation as the initial candidate entity. The multi-level expansion mechanism significantly improves the completeness of the candidate entity coverage, and is especially suitable for processing professional terms or informal language scenarios. Meanwhile, in order to reduce the redundancy of the candidate entity set, a multi-dimensional feature screening strategy is proposed, which fully considers the popularity of the candidate entity, the context local matching degree and the entity correlation degree and other indicators. The popularity of the entity is obtained by counting the historical link frequency of the entity in the knowledge base, and the high-frequency entity in the general context is given a higher priority. The context local matching degree is obtained by extracting the core words of the sentence where the designation is located and the encyclopedia abstract of the candidate entity for simple bag-of-words model matching, so as to quickly exclude the entities with large semantic deviation. The entity correlation degree is calculated, and the multi-dimensional feature screening strategy proposed in the present application gradually removes the low correlation candidate entities in a hierarchical progressive manner, and then generates the coarse candidate entity set.

[0012] Comprehensive similarity score candidate entity refinement screening: further reduce noise interference from the aspects of semantic generalization and text characteristic similarity, and obtain a refined candidate entity set with high quality and close correlation.

[0013] Although the coarse screening can reduce the candidate entity set to a certain extent, due to the wide coverage of the existing knowledge base, in order to improve the recall rate of the candidate set, too many emerging entities and variants will be stored in the coarse candidate entity set, resulting in an excessively large candidate entity set. In view of the problem, the present application proposes a comprehensive similarity score candidate entity refinement screening strategy, which further reduces noise interference from the aspects of semantic generalization and text characteristic similarity, and obtains a refined candidate entity set with high quality and close correlation.

[0014] Step 2-1: key word extraction based on multi-dimensional feature measurement;

[0015] A key word extraction algorithm based on multi-dimensional feature measurement (MFMKE) is proposed, which reconstructs the edge weight calculation method by fusing multi-dimensional features such as part-of-speech constraint, information entropy measurement, semantic embedding and context distance, and significantly improves the accuracy and robustness of key word extraction. The information entropy theory is introduced into the graph model construction process for the first time, and the semantic association strength between words is dynamically adjusted by combining linguistic rules and statistical characteristics, so that the implicit theme of short text can be captured more accurately. The MFMKE algorithm introduces the information entropy theory as a weight adjustment factor. Specifically, the information entropy is defined as

[0016]

[0017] where P(c|w i ) is the conditional probability of the word w iThe frequency of its context c, the log base 2 logarithm operation for converting probability into information measure, and the negative sign for ensuring the information entropy is non-negative since the logarithm of probability is usually negative.

[0018] Based on the Shannon entropy theory, the vocabulary w i Distribution dispersion in different contexts. Statistics w i The position of occurrence in all sliding windows in the library, regarding each window as a context c, and estimating P(c|w i ) by frequency. A high entropy value indicates that w i has high contextual diversity, such as the function word "and" can appear anywhere; a low entropy value reflects the contextual specificity of the vocabulary, such as "Basketball" often appears in paragraphs related to basketball sports. The initial score of the node is set to

[0019] S(V i )=TF(W i )×(1-H(w i ))

[0020] where TF(w i ) represents the word frequency of w i , and H(w i ) is the information entropy of w i , which reduces the initial value of high-entropy words. Words with high word frequency and low information entropy will have higher initial scores.

[0021] MFMKE algorithm designs different edge weight calculation methods for different part-of-speech combinations. Traditional methods treat part-of-speech as static features, while MFMKE dynamically adjusts the strength of part-of-speech association to achieve deep perception of grammatical structure.

[0022] Step 2-2: Introduce similarity calculation of prior information;

[0023] Inverse Document Frequency (IDF) is an important statistical measure in the field of information retrieval, and plays a key role in short text analysis. Its core value lies in the quantitative evaluation of the discrimination ability of vocabulary in the corpus. By calculating the distribution characteristics of vocabulary in the document set, it effectively identifies key terms with high discriminability. In the entity disambiguation task, the entity description documents of the initial candidate entity set are constructed into a structured corpus. Based on the statistical characteristics of IDF, the fine selection of key features of candidate entities can be realized, and the specific formula is

[0024]

[0025] where N represents the cardinality of the document set, n wThe logarithmic function is introduced to effectively compress the numerical scale, so that the high-frequency words and low-frequency words are more distinguished.

[0026] It is worth noting that the introduction of the additive smoothing term effectively avoids the zero probability problem and enhances the numerical stability of the model.

[0027] The present application proposes a semantic strategy based on large-scale pre-training word vectors, which maps the short text keywords w and candidate entities E into the same high-dimensional vector space through the Word2Vec technology, respectively obtaining their vectors V W and V E . This strategy is based on the training mechanism of the neural probabilistic language model, which introduces text features through the co-occurrence relationship of words within the context window, and effectively captures the semantic association between words through the characteristics of the continuous vector space. The cosine similarity between word vectors is

[0028]

[0029] where V W ·V E represents the dot product of vectors V W and V E , |V W | and |V E | represent the modulus of vectors V W and V E ;

[0030] This operation represents the similarity between vectors, which eliminates the influence of dimension scaling through modulus length normalization. This measurement essentially reflects the projection relationship of words in the latent semantic space, and can more accurately represent the deep semantic association of words compared to the traditional bag-of-words model. It is particularly important to note that this similarity measure integrates multi-dimensional language features. At the local level, it captures syntactic structure information through n-gram syntax patterns; at the global level, it preserves semantic distribution characteristics based on large-scale word co-occurrence statistics. Combining the IDF weight with the semantic similarity can construct a composite feature selection mechanism. Specifically, for each candidate entity e∈E, the relevance of the target reference m is

[0031]

[0032] where IDF(w) is the inverse document frequency value of the word w, and sim(V m ,V e ) is the cosine similarity between vectors V W and V E ;

[0033] This scoring function has two advantages: first, IDF suppresses interference from common words and strengthens the weight of domain-specific names and entity features; second, the introduction of semantic similarity ensures that entities with different surface forms but similar semantics can be appropriately matched. This method, which combines statistics and semantics, effectively solves the problem of over-reliance on word form matching in traditional methods.

[0034] This method is adaptive to the size and quality of the corpus. When the document set N is small, the IDF mechanism can automatically adjust the feature weights to avoid interference from low-frequency noise; as the corpus size expands, the word vector model can fully learn subtle semantic differences. This dynamic balancing mechanism ensures the universality of the method in different domains and knowledge bases of different sizes.

[0035] Steps 2-3: Refined screening of candidate entities;

[0036] To achieve collaborative filtering of textual features and semantic information, this invention proposes a comprehensive similarity scoring function. Based on the association characteristics between referential and candidate entities, scoring is achieved through the interaction of semantic elements and textual features. The comprehensive scoring function is as follows:

[0037]

[0038] Among them, V E and V w Let E and w represent the vector representations of the candidate entity E and the keyword w, respectively. The normalized dot product operation represents their similarity in the semantic space, and IDF represents the inverse document frequency weight.

[0039] By dynamically adjusting keyword weights, interference from common words is suppressed while reinforcing the importance of domain-specific keywords. A linear weighted summation mechanism organically integrates local semantic matching degrees with global statistical features.

[0040] The P-value is set as the quantitative index of entity relevance, and the Top-K results are used as the final candidate entity set for entity referencing. This method achieves knowledge fusion at three levels: lexical level, quantifying the discriminative ability of terms through IDF; syntactic level, preserving contextual structural information with the help of word vectors; and semantic level, modeling concept associations based on distributed representation. While ensuring the recall rate of real entities, this method further compresses the size of the candidate entity set, providing a good data foundation for improving the disambiguation performance of short text entities, especially demonstrating strong robustness when dealing with polysemy and new entities.

[0041] The beneficial effects of this invention are as follows: This invention proposes a two-stage candidate entity screening strategy based on multi-factor text feature fusion. To further streamline the candidate entity scale according to practical application needs, the candidate entity screening is divided into two stages: coarse screening and fine screening. First, coarse screening based on the Wikipedia knowledge base expands candidate entities through surface form fuzzy matching and morphological normalization, and constructs an initial candidate set with high recall by combining alias indexing and word variant parsing. Second, in the fine screening stage, a keyword extraction method based on multi-dimensional feature measurement is proposed. By introducing part-of-speech constraints, information entropy adjustment, and context distance weighting, keywords in short texts are extracted; prior information is introduced to calculate the similarity between candidate entities and entity references, and a comprehensive similarity score of candidate entities is calculated from the perspectives of semantic generalization and text feature similarity, thereby completing the fine screening of candidate entities. This invention, through two-stage screening, further reduces the interference of data noise and obtains a high-quality, closely related candidate entity set, providing good data support for entity disambiguation. Attached Figure Description

[0042] Figure 1 Workflow diagram of this invention;

[0043] Figure 2 Schematic diagram of the coarsening screening method of this invention;

[0044] Figure 3 A detailed screening diagram of the present invention. Detailed Implementation

[0045] The present invention will be further described with reference to the accompanying drawings.

[0046] The steps of this invention are as follows:

[0047] Step 1: Coarse filtering of candidate entities based on Wikipedia knowledge base;

[0048] Morphological normalization is achieved through lemmatization, and variant expansion is performed using the alias index of the knowledge base (Wikipedia knowledge base redirection). Alternatives and variants of entity references are also included as initial candidate entities. For example, when searching for "guards," the entity with its singular form "guard," along with related aliases in the knowledge base such as "shooting guard," are added to the candidate entity set. This multi-level expansion mechanism significantly improves the completeness of candidate entity coverage, especially suitable for handling language scenarios involving specialized terminology or informal text. Simultaneously, to reduce the redundancy of the candidate entity set, a multi-dimensional feature selection strategy is proposed, fully considering indicators such as candidate entity popularity, local contextual matching, and entity relevance. Entity popularity is determined by statistically analyzing the historical link frequency of entities in the knowledge base, assigning higher priority to high-frequency entities in general contexts. For example, in short texts without a clear domain focus, "Apple" company and "Apple" apple have a higher prior probability than "Apple" song. Contextual local matching is achieved by extracting the core words of the sentence containing the reference and performing a simple bag-of-words model match with the encyclopedia summary of the candidate entities, thereby quickly eliminating entities with significant semantic deviations. Entity relevance is calculated based on the closeness between entities. For example, when "MichaelJordan" and "Chicago Bulls" appear simultaneously in the text, the entity "Point Guard," which is closely related to basketball, will be prioritized for retention. The multi-dimensional feature selection strategy proposed in this invention gradually eliminates low-relevance candidate entities in a hierarchical and progressive manner, thereby generating a coarsened candidate entity set.

[0049] Step 2: Refine and filter candidate entities based on comprehensive similarity scores;

[0050] While coarse filtering can reduce the candidate entity set to some extent, the broad coverage of existing knowledge bases means that coarse filtering, in order to improve the recall rate of the candidate set, will include too many emerging entities and variants, resulting in an excessively large candidate entity set. To address this issue, this invention proposes a comprehensive similarity scoring-based candidate entity refinement filtering strategy. This strategy further reduces noise interference from the perspectives of semantic generalization and textual feature similarity, obtaining a refined candidate entity set with higher quality and stronger relevance.

[0051] Step 2-1: Keyword extraction for multidimensional feature measurement;

[0052] This paper proposes a multi-dimensional feature-based keyword extraction algorithm (MFMKE). By integrating multi-dimensional features such as part-of-speech constraints, information entropy measurement, semantic embedding, and context distance, the algorithm reconstructs the edge weight calculation method, significantly improving the accuracy and robustness of keyword extraction. This algorithm is the first to introduce information entropy theory into the graph model construction process, combining linguistic rules and statistical properties to dynamically adjust the semantic association strength between words, thereby more accurately capturing the implicit themes of short texts.

[0053] The TextRank algorithm relies solely on word co-occurrence frequency and sliding window distance to construct a word graph, neglecting the impact of part-of-speech distribution, contextual uncertainty, and deep semantic relationships on keyword extraction. This invention proposes the MFMKE algorithm, which incorporates information entropy theory as a weight adjustment factor. Specifically, information entropy is defined as…

[0054]

[0055] Based on Shannon entropy theory, quantify vocabulary w i The dispersion of the distribution in different contexts. Statistical w i The occurrence position of each sliding window in the library is determined by treating each window as a context c and estimating P(c|w) by frequency. i High entropy indicates that w i High contextual diversity, such as the function word "and" appearing in any position; low entropy values ​​reflect the contextual specificity of words, such as "Basketball" appearing frequently in paragraphs related to basketball. The initial score of the node is set to...

[0056] S(V i ) = TF(W i )×(1-H(w i ))

[0057] TF(W i ) indicates w i Word frequency, H(w) i ) for w i Information entropy reduces the initial value of high-entropy words; words with higher frequency and lower information entropy will have higher initial scores.

[0058] The MFMKE algorithm employs differentiated edge weight calculation methods for different parts-of-speech combinations. Traditional methods treat parts-of-speech as static features, while MFMKE achieves deep perception of grammatical structure by dynamically adjusting the strength of part-of-speech associations.

[0059] Step 2-2: Similarity calculation using prior information;

[0060] Inverse Document Frequency (IDF), an important statistical measure in information retrieval, plays a crucial role in short text analysis. Its core value lies in the quantitative assessment of the word discrimination ability within a corpus. By calculating the distribution characteristics of words in a document set, it effectively identifies key terms with high discriminative power. In entity disambiguation tasks, the entity description documents of the initial candidate entity set are constructed into a structured corpus. Based on the statistical properties of IDF, refined screening of key features of candidate entities can be achieved. The specific formula is as follows:

[0061]

[0062] Where N represents the cardinality of the document set, n w This represents the frequency of documents containing the target word w. The introduction of the logarithmic function effectively compresses the numerical scale, making the distinction between high-frequency and low-frequency words more significant. Notably, the introduction of the additive smoothing term effectively avoids the zero-probability problem and enhances the numerical stability of the model.

[0063] This invention proposes a semantic strategy based on large-scale pre-trained word vectors. Using Word2Vec technology, it maps short text keywords w and candidate entities E to the same high-dimensional vector space, obtaining their respective vectors V. W and V E This strategy is based on the training mechanism of a neural probabilistic language model. It introduces text features through word co-occurrence relationships within a context window, and effectively captures semantic associations between words through the characteristics of continuous vector spaces. The cosine similarity formula between word vectors is...

[0064]

[0065] In this model, the dot product operation represents the similarity between vectors, and the effect of dimensionality scaling is eliminated through modulus normalization. This metric essentially reflects the projection relationship of words in the latent semantic space, and compared to the traditional bag-of-words model, it can more accurately represent the deep semantic connections of words. It is particularly noteworthy that this similarity metric integrates multi-dimensional language features. At the local level, it captures syntactic structure information through n-gram word order patterns; at the global level, it preserves semantic distribution characteristics based on vocabulary co-occurrence statistics from a large-scale corpus. Combining IDF weights with semantic similarity allows for the construction of a composite feature selection mechanism. Specifically, for each candidate entity e∈E, its relevance to the target referent m is...

[0066]

[0067] This scoring function has two advantages: first, IDF suppresses interference from common words and strengthens the weight of domain-specific names and entity features; second, the introduction of semantic similarity ensures that entities with different surface forms but similar semantics can be appropriately matched. This method, which combines statistics and semantics, effectively solves the problem of over-reliance on word form matching in traditional methods.

[0068] This method is adaptive to the size and quality of the corpus. When the document set N is small, the IDF mechanism can automatically adjust the feature weights to avoid interference from low-frequency noise; as the corpus size expands, the word vector model can fully learn subtle semantic differences. This dynamic balancing mechanism ensures the universality of the method in different domains and knowledge bases of different sizes.

[0069] Steps 2-3: Refined screening of candidate entities;

[0070] To achieve collaborative filtering of textual features and semantic information, this invention proposes a comprehensive similarity scoring function. Based on the association characteristics between referential and candidate entities, the comprehensive scoring function is as follows:

[0071]

[0072] This formula achieves scoring through the interaction of semantic elements and textual characteristics. V W and V E Let E and w be the vector representations of candidate entity E and keyword w, respectively. Their normalized dot product operation represents their similarity in the semantic space. IDF, as the inverse document frequency weight, dynamically adjusts the keyword weights to suppress interference from common words and strengthen the importance of domain keywords. A linear weighted summation mechanism organically integrates local semantic matching degree with global statistical features.

[0073] The P-value is set as the quantitative index of entity relevance, and the Top-K results are used as the final candidate entity set for entity referencing. This method achieves knowledge fusion at three levels: lexical level, quantifying the discriminative ability of terms through IDF; syntactic level, preserving contextual structural information with the help of word vectors; and semantic level, modeling concept associations based on distributed representation. While ensuring the recall rate of real entities, this method further compresses the size of the candidate entity set, providing a good data foundation for improving the disambiguation performance of short text entities, especially demonstrating strong robustness when dealing with polysemy and new entities.

Claims

1. A two-stage candidate entity selection method based on multi-factor text feature fusion for short text entity disambiguation, characterized in that: Candidate entity coarsening and filtering based on Wikipedia knowledge base: Morphological normalization is achieved through word form merging technology, and variant expansion is carried out by relying on the alias index of the knowledge base, and the entities corresponding to the aliases and variants of the entity referencing are also used as initial candidate entities. Candidate entity refinement screening based on comprehensive similarity score: A keyword extraction method based on multidimensional feature measurement is proposed, and prior information is introduced to calculate the similarity between candidate entities and entity references. The candidate entities are then refined and screened based on the comprehensive similarity score.

2. The candidate entity secondary screening method based on multi-factor text feature fusion for short text entity disambiguation according to claim 1, characterized in that: In the aforementioned candidate entity coarsening and filtering based on the Wikipedia knowledge base: A multi-level expansion mechanism and a multi-dimensional feature selection strategy are proposed, which fully consider three indicators: candidate entity popularity, contextual local matching degree, and entity relevance degree. Entity popularity is assigned higher priority to high-frequency entities in general contexts by statistically analyzing the historical link frequency of entities in the knowledge base. Contextual local matching degree is achieved by extracting the core words of the sentence in which the reference is located and performing a simple bag-of-words model matching with the encyclopedia summary of the candidate entities to exclude entities with significant semantic deviations. The entity relevance degree is used to calculate the closeness between entities. The multi-dimensional feature selection strategy gradually eliminates low-relevance candidate entities in a hierarchical and progressive manner to generate a coarse candidate entity set.

3. The method for secondary screening of candidate entities based on multi-factor text feature fusion for short text entity disambiguation according to claim 1, characterized in that: The specific steps in the candidate entity refinement and screening based on the comprehensive similarity score are as follows: Step 2-1: Extract keywords for multidimensional feature measurement; Step 2-2: Calculate the similarity of prior information; Steps 2-3: Refine and filter candidate entities.

4. The method for secondary screening of candidate entities based on multi-factor text feature fusion for short text entity disambiguation according to claim 3, characterized in that: In step 2-1, This paper proposes a keyword extraction algorithm, MFMKE, based on multidimensional feature measurement. By integrating multidimensional features such as part-of-speech constraints, information entropy measurement, semantic embedding, and context distance, it reconstructs the edge weight calculation method, introduces information entropy theory into the graph model construction process, and combines linguistic rules and statistical properties to dynamically adjust the semantic association strength between words, capturing the implicit themes of short texts. The MFMKE algorithm incorporates information entropy theory as a weight adjustment factor; information entropy is defined as… Where P(c|w) i ) is a given vocabulary w i In the case of , the frequency of occurrence of its context c, logarithm base 2, is used to convert probability into information content measure, and the negative sign is used to ensure that the information entropy is non-negative, because the logarithm of probability is usually negative; Based on Shannon entropy theory, quantify vocabulary w i Distribution dispersion in different contexts, statistical w i The occurrence position of each sliding window in the library is determined by treating each window as a context c and estimating P(c|w) by frequency. i High entropy values ​​indicate that w i The high contextual diversity and low entropy value reflect the contextual specificity of words. The initial score of the node is set to S(V). i ) = TF(w i )×(1-H(w i )) Among them, TF(w i ) indicates w i Word frequency, H(w) i ) for w i The information entropy is reduced, and the initial value of high-entropy words is lowered. Words with higher frequency and lower information entropy will have higher initial scores.

5. The candidate entity secondary screening method based on multi-factor text feature fusion for short text entity disambiguation according to claim 3, characterized in that: In step 2-2, Inverse Document Frequency (IDF) is an important statistical measure in information retrieval, used to quantitatively evaluate the word discrimination ability in a corpus. By calculating the distribution characteristics of words in a document set, it identifies key terms with high discriminative power. In entity disambiguation tasks, the entity description documents of the initial candidate entity set are constructed into a structured corpus. Based on the statistical properties of IDF, a refined selection of key features of candidate entities is achieved. The specific formula is as follows: Where N represents the cardinality of the document set, n w This indicates the frequency of documents containing the target word w; The Word2Vec technique is used to map the short text keyword w and the candidate entity E to the same high-dimensional vector space, obtaining their respective vectors V. W and V E Based on the training mechanism of a neural probabilistic language model, text features are introduced through word co-occurrence relationships within a context window. Semantic associations between words are captured through the properties of continuous vector spaces, and the cosine similarity between word vectors is... Among them, V W ·V E Represents vector V W and V E The dot product, |V W | and | V E | represents vector V W and V E The model; This operation characterizes the similarity between vectors, eliminating the impact of dimensional scaling through modulus normalization. This metric reflects the projection relationship of words in the latent semantic space. This similarity metric integrates multi-dimensional language features. At the local level, it captures syntactic structure information through n-gram word order patterns; at the global level, it preserves semantic distribution characteristics based on large-scale corpus-based word co-occurrence statistics, combining IDF weights with semantic similarity to construct a composite feature selection mechanism. For each candidate entity e∈E, its correlation formula with the target referent m is: Where IDF(w) is the inverse document frequency of word w, sim(V m V e ) is a vector V W and V E Cosine similarity; This method is adaptive to the size and quality of the corpus. When the document set N is small, the IDF mechanism can automatically adjust the feature weights to avoid interference from low-frequency noise. When the corpus size expands, the word vector model can fully learn the subtle differences in semantics.

6. The method for secondary screening of candidate entities based on multi-factor text feature fusion for short text entity disambiguation according to claim 3, characterized in that: In steps 2-3, To achieve collaborative filtering of textual features and semantic information, a comprehensive similarity scoring function is proposed. This function scores based on the association characteristics between referents and candidate entities, through the interaction of semantic elements and textual features. The comprehensive scoring function is as follows: Among them, V E and V w Let E and w represent the vector representations of the candidate entity E and the keyword w, respectively. The normalized dot product operation represents their similarity in the semantic space, and IDF represents the inverse document frequency weight. By dynamically adjusting the weight of keywords, interference from common words is suppressed and the importance of domain keywords is strengthened. Through a linear weighted summation mechanism, local semantic matching degree and global statistical features are organically integrated. The P-value is set as the quantitative index of entity relevance, and the Top-K results are used as the final candidate entity set for entity referencing.