Document vector similarity calculation method combined with semantic feature enhancement

By normalizing and filtering the document set, an initial TF-IDF matrix is ​​constructed. Then, terms are mapped to semantic structures composed of multiple semantic primitives. The semantic similarity matrix is ​​calculated and sparsified to generate a semantically enhanced term-document matrix. This solves the problem of insufficient semantic expression in existing technologies and achieves more accurate document similarity calculation.

CN121052237APending Publication Date: 2025-12-02CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511115574.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing document similarity calculation methods are insufficient in terms of semantic expression, domain adaptability, and accuracy, making it difficult to meet the needs of efficient retrieval and matching in large-scale, specialized document scenarios.

Method used

After normalizing and filtering the document set and constructing an initial TF-IDF matrix, the terms are mapped to a semantic structure composed of multiple semantic primitives. The semantic similarity between terms is calculated and a semantic similarity matrix is ​​constructed. After sparsification, a semantically enhanced term-document matrix is ​​generated and fused with the original matrix to calculate the final document similarity.

Benefits of technology

It enhances the semantic expressiveness of document similarity calculation, making the calculation results closer to the real semantic associations and having stronger accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121052237A_ABST
    Figure CN121052237A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of natural language processing and information retrieval, and provides a document vector similarity calculation method combined with semantic feature enhancement, which comprises the following steps of: firstly, carrying out standardization processing and vocabulary screening on an original document set, and constructing an initial TF-IDF matrix; then mapping the lexical items into a semantic structure formed by multiple classes of primitives, calculating semantic similarity among the lexical items and constructing a semantic similarity matrix; and after the matrix is subjected to sparse processing, generating a lexical item-document matrix with enhanced semantics, fusing the lexical item-document matrix with the original matrix, and calculating the final document similarity. In this way, the expression ability of document similarity calculation on the semantic level is improved, the calculation result is closer to the real semantic association, and higher accuracy and robustness are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of natural language processing and information retrieval technology, and in particular relates to a document vector similarity calculation method that combines semantic feature enhancement. Background Technology

[0002] With the rapid development of the internet and big data technologies, textual information has experienced explosive growth. How to efficiently and accurately understand and retrieve knowledge from massive amounts of documents has become an important research direction in the fields of natural language processing and information retrieval. Document similarity calculation, as one of the core technologies supporting artificial intelligence applications such as search engine optimization, intelligent question answering systems, and recommendation systems, aims to achieve accurate information matching by automatically analyzing the semantic or content similarity between different documents.

[0003] Traditional document similarity calculation methods are mainly based on the Vector Space Model (VSM), which represents documents as high-dimensional vectors weighted by TF-IDF and compares them using metrics such as cosine similarity. While this method offers advantages such as high computational efficiency and simplicity, it has significant limitations in practical applications: it cannot effectively capture the semantic relationships between words. To address these issues, researchers have proposed various semantic similarity calculation methods in recent years, primarily including knowledge-based methods (such as HowNet and WordNet) and corpus-based deep learning methods (such as Word2Vec and BERT). These methods have improved semantic understanding to some extent, but still suffer from problems such as limited semantic expression, poor adaptability to specialized terminology, and low accuracy.

[0004] However, existing document similarity calculation methods still have many shortcomings in terms of semantic expression ability, domain adaptability and accuracy, making it difficult to meet the needs of efficient retrieval and matching in current large-scale and specialized document scenarios. Summary of the Invention

[0005] This application provides a document vector similarity calculation method, terminal device, and storage medium that combines semantic feature enhancement. It can solve the problem that existing document similarity calculation methods still have many shortcomings in terms of semantic expression ability, domain adaptability, and accuracy, and are difficult to meet the needs of efficient retrieval and matching in current large-scale and specialized document scenarios.

[0006] In a first aspect, embodiments of this application provide a document vector similarity calculation method that combines semantic feature enhancement, comprising: S1, processing a first document set according to a preset normalization processing method to generate a second document set; S2, filtering a first vocabulary according to preset constraints to generate a second vocabulary V = {term1, term2, ..., term...} n}, the first vocabulary is generated by deduplicating all terms in the second document set; S3, based on the second vocabulary and the second document set, and combined with the TF-IDF weight calculation method, construct an initial term-inverse document matrix A = {W ij S4. Based on a preset knowledge base, map each term in the second vocabulary to a semantic structure composed of multiple semantic primitives. Determine the semantic similarity between terms according to the semantic primitive similarity calculation rules, and construct a semantic similarity matrix between terms. S5. Perform sparsification processing on the semantic similarity matrix to generate a sparse semantic similarity matrix Sem. sparse S6. Perform matrix operations on the sparse semantic similarity matrix and the initial term-inverse document matrix to obtain the semantically enhanced term-inverse document matrix; S7. Based on the initial term-inverse document matrix and the semantically enhanced term-inverse document matrix, and combined with the weighted average method, generate the final document similarity.

[0007] In one possible implementation of the first aspect, S1, processing the first document set according to a preset normalization processing method to generate a second document set, includes:

[0008] The first document set is segmented using a preset word segmentation tool to generate a corresponding word list.

[0009] Using a preset stop word list and combined with custom filtering rules based on the features of a preset corpus, meaningless or distracting words in each word list are removed to generate a second document set.

[0010] Optionally, in another possible implementation of the first aspect, S2, filtering the first vocabulary according to preset constraints to generate a second vocabulary, includes:

[0011] Retain terms that have appeared in at least 3 documents in the first vocabulary;

[0012] Filter out terms that appear in more than 95% of documents from the first vocabulary list;

[0013] The second vocabulary is generated by sorting each term according to its frequency of occurrence in the first document set and limiting the total number of retained terms to no more than 10,000.

[0014] Optionally, in another possible implementation of the first aspect, the above-mentioned S3, based on the second vocabulary and the second document set, and combined with the TF-IDF weight calculation method, constructs an initial term-inverse document matrix A = {W}. ij The specific formula is as follows:

[0015]

[0016] Among them, tf i,j idf represents the frequency of term i in document j within the second document set. i df represents the inverse document frequency of term i. i The second document set contains the number of documents with term i, and m represents the total number of documents in the second document set.

[0017] Optionally, in another possible implementation of the first aspect, S4 above, mapping each term in the second vocabulary to a semantic structure composed of multiple semantic primitives based on a preset knowledge base, determining the semantic similarity between terms according to the semantic primitive similarity calculation rules, and constructing a semantic similarity matrix between terms, includes:

[0018] S41. Map each term in the second vocabulary to one or more concepts in a preset knowledge base. Each concept is described by a semantic primitive structure consisting of independent semantic primitives, relational semantic primitives, and symbolic semantic primitives.

[0019] S42. Based on the depth information and path distance of each basic semantic primitive in the semantic primitive tree, the following method for calculating the similarity between two basic semantic primitives is used:

[0020]

[0021] Where s1 and s2 are two basic primitives, and the adjustment parameter α = 0.5. Let s1 and s2 represent the depths of the basic primitives s1 and s2 in their respective primitive trees. The path distance dis(s1, s2) is defined as the sum of the weights of the edges on the shortest path connecting s1 and s2.

[0022]

[0023] Where n represents the number of edges contained in the shortest path between semantic primitives s1 and s2; level(k) represents the edge connecting the k-th level node and the (k+1)-th level node; the level number of the root node of the semantic primitive tree is defined as 0, and the edge weight weight(i) is calculated as follows:

[0024]

[0025] Where m represents the maximum number of levels in the primitive tree, and the adjustment parameter θ = 4;

[0026] S43. Using the maximum value matching strategy, based on the basic semantic primitive similarity matrix, calculate the similarity between any two concepts in the independent semantic primitive dimension:

[0027] S431. Let any two independent primitives of meaning be X = x. i (i = 1, 2, ..., m), Y = y j (j=1,2,…,n), where m≤n, and create an m*n similarity matrix B:

[0028]

[0029] S432. Find the maximum value max(Sim(x) in matrix B. p ,y q Add it to the similarity list L;

[0030] S433. Set the element in row p and column q containing the maximum value to 0;

[0031] S434. Repeat steps S432-S433 above until all elements in matrix B are zero.

[0032] S435. Set nm unmatched elements y in the independent primitive Y. t (t = 1, 2, ..., nm) corresponds to an empty element And set the similarity between unmatched elements and empty elements. δ is a preset constant;

[0033] S436, will Add the data to the similarity list L, take the average of L, and obtain the similarity between the independent primitives X and Y:

[0034]

[0035] Among them, L k This represents the k-th element in list L;

[0036] S44. For any two concepts, based on the rules of attribute matching and value similarity calculation, obtain their relational semantic primitive similarity:

[0037] S441. Extract the relational primitive structure of each of the two concepts and represent it as a set composed of "attribute-value" pairs;

[0038] S442. Match the attributes in the two relational primitive structures; if any attribute exists in both concepts, extract the two corresponding attribute values ​​and calculate the similarity between the two attribute values ​​according to the basic primitive similarity calculation method; if any attribute does not exist in both concepts, set the attribute value corresponding to the attribute to an empty element, and set the similarity between the empty element and any other value to δ, where δ is a preset constant.

[0039] S443. Summarize the similarity of all attribute values ​​involved in the calculation, calculate their arithmetic mean, and use it as the similarity between the two concepts in the relational semantic primitive dimension.

[0040] S45. For any two concepts, based on the rules of attribute matching and value similarity calculation, obtain their symbolic semantic primitive similarity:

[0041] S451. Extract the relational primitive structure of each of the two concepts and represent it as a set composed of "attribute-value" pairs;

[0042] S452. Match the attributes in the two symbolic primitive structures; if any attribute exists in both concepts, extract the two corresponding attribute values ​​and calculate the similarity between the two attribute values ​​according to the basic primitive similarity calculation method; if any attribute does not exist in both concepts, set the attribute value corresponding to the attribute to an empty element, and set the similarity between the empty element and any other value to δ, where δ is a preset constant.

[0043] S453. Summarize the similarity of all attribute values ​​involved in the calculation, calculate their arithmetic mean, and use it as the similarity between the two concepts in the symbolic semantic primitive dimension.

[0044] S46. Combining the independent semantic primitives, relational semantic primitives, and symbolic semantic primitives, the similarity between any two concepts is calculated using a weighted average:

[0045]

[0046] Where i1, i2, and i3 correspond to independent primitives, relational primitives, and symbolic primitives, respectively, and β i =[0.7,0.17,0.13], Sim j (S1,S2) represents the similarity between the j-th semantic primitives of the two concepts;

[0047] S47. For each pair of terms, iterate through the combinations of their corresponding concepts, and select the concept pair with the highest similarity as the semantic similarity between the two terms, as shown below:

[0048]

[0049] Among them, S1i ,S 2j Each represents all the concepts corresponding to the two terms;

[0050] S48. Fill the semantic similarity values ​​between all term pairs into the corresponding matrix positions to construct a symmetric semantic similarity matrix indexed by the second vocabulary:

[0051]

[0052] Optionally, in another possible implementation of the first aspect, S5 above performs sparsification processing on the semantic similarity matrix to generate a sparse semantic similarity matrix Sem. sparse ,include:

[0053] Based on the constructed TF-IDF matrix A, the weight vector reflecting the importance of terms is obtained by summing the columns.

[0054] For each term i The upper limit K of the corresponding number of matches is calculated based on its weight vector. i ,in, w i This represents the value of the term in the weight vector, where β is an adjustable parameter used to control the overall proportion of the number of matches.

[0055] Filter word pairs based on the set similarity threshold μ;

[0056] For each term i Select the maximum number of K values ​​from the i-th row of Sem. i The maximum similarity value greater than the threshold μ is selected, and the similarity values ​​that are not selected or are lower than the threshold μ are set to 0 to generate a sparse semantic similarity matrix Sem. sparse .

[0057] Optionally, in another possible implementation of the first aspect, in step S6 above, the sparse semantic similarity matrix is ​​subjected to matrix operations with the initial term-inverse document matrix to obtain the semantically enhanced term-inverse document matrix, as follows:

[0058] A enhanced =Sem sparse ×A

[0059] Among them, A enhanced The semantically enhanced term-inverse document matrix has elements (A enhanced ) ij This represents the enhanced weight of term i in document j after considering semantic information.

[0060] Optionally, in another possible implementation of the first aspect, S7 above, generating the final document similarity based on the initial term-inverse document matrix and the semantically enhanced term-inverse document matrix, combined with a weighted average method, includes:

[0061] Let document D i and D j The vector representations in the traditional vector space model are A i and A j In the semantic enhancement model, they are represented as A. enhanced,i and A enhanced,j Then its corresponding cosine similarity can be expressed as:

[0062]

[0063] By introducing a fusion parameter α∈[0,1], the similarity results from the two models are weighted and fused to obtain the final document similarity:

[0064] Final Sim (D i D j )

[0065] =α*Cos Sim (A i A j )+(1-α)*Cos Sim (A enhanced,i A enhanced,j )

[0066] Where α represents the weight of the original statistical information in the final similarity; 1-α represents the weight of the semantic enhancement information; when α=1, it relies entirely on the traditional vector space model; when α=0, it is based entirely on the semantic enhancement model.

[0067] Secondly, embodiments of this application provide a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned document vector similarity calculation method that combines semantic feature enhancement.

[0068] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned document vector similarity calculation method that combines semantic feature enhancement.

[0069] Beneficial Effects: In this technical solution, the original document set is first normalized and vocabulary is filtered to construct an initial TF-IDF matrix. Then, terms are mapped to semantic structures composed of multiple semantic primitives, and the semantic similarity between terms is calculated to construct a semantic similarity matrix. After sparsification of this matrix, a semantically enhanced term-document matrix is ​​generated and fused with the original matrix to calculate the final document similarity. This improves the semantic expressiveness of document similarity calculation, making the calculation results closer to real semantic associations and exhibiting stronger accuracy and robustness. Attached Figure Description

[0070] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0071] Figure 1 This is a flowchart illustrating a document vector similarity calculation method that combines semantic feature enhancement, as provided in an embodiment of this application.

[0072] Figure 2 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation

[0073] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0074] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0075] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0076] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0077] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0078] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0079] The following description, with reference to the accompanying drawings, details a document vector similarity calculation method, terminal device, and storage medium that combine semantic features for enhancement, as provided in this application.

[0080] Figure 1 The diagram illustrates a flowchart of a document vector similarity calculation method that combines semantic feature enhancement, as provided in an embodiment of this application.

[0081] like Figure 1 As shown, this document vector similarity calculation method combining semantic feature enhancement includes the following steps:

[0082] S1. Process the first document set according to the preset normalization processing method to generate the second document set;

[0083] Furthermore, in this embodiment of the application, step S1 includes:

[0084] The first document set is segmented using a preset word segmentation tool to generate a corresponding word list.

[0085] Using a preset stop word list and combining with filtering rules customized according to the characteristics of the preset corpus, meaningless or interfering terms in each list of terms are removed to generate a second document set.

[0086] As a possible implementation, the word segmentation operation can use the Chinese word segmentation tool PKUSEG developed by the Language Computing and Machine Learning Research Group of Peking University to perform word segmentation on each document to generate a list of terms. For example, any document, after being segmented by PKUSEG, is converted into a list of terms, where ti represents a term, and duplicate terms are allowed to appear in the list of terms.

[0087] As a possible implementation, the preset stop word list can be the Sogou stop word list.

[0088] S2. Filter the first vocabulary according to preset constraint conditions to generate a second vocabulary. The first vocabulary is generated by de-duplicating all terms in the second document set;

[0089] Furthermore, in the embodiments of the present application, the above step S2 includes:

[0090] Step S21. Retain terms that appear in at least 3 documents in the first vocabulary;

[0091] Step S22. Filter out terms that appear in more than 95% of the documents in the first vocabulary;

[0092] Step S23. Sort each term according to the frequency of occurrence of each term in the first document set, and limit the total number of retained terms not to exceed 10,000 to generate a second vocabulary.

[0093] In the embodiments of the present application, the above step S21 can filter out extremely low-frequency terms (possibly spelling mistakes or special terms) to reduce noise. The above step S22 can exclude high-frequency general terms (such as stop words like "de", "shi", etc.) and retain distinctive features. The above step S23 can control the dimension of the feature space and reduce the computational complexity.

[0094] S3. According to the second vocabulary and the second document set, and combining with the TF-IDF weight calculation method, construct an initial term-inverse document matrix;

[0095] It should be noted that, in the embodiments of the present application, the above step S3 is specifically shown by the following formula:

[0096]

[0097] Where, tf i,j represents the term frequency of term i in document j in the second document set, and idf idf represents the inverse document frequency of term i. i The second document set contains the number of documents with term i, and m represents the total number of documents in the second document set.

[0098] S4. Based on the preset knowledge base, map each word in the second vocabulary to a semantic structure composed of multiple semantic primitives, determine the semantic similarity between words according to the semantic primitive similarity calculation rules, and construct a semantic similarity matrix between words.

[0099] For example, the aforementioned pre-defined knowledge base could be HowNet.

[0100] Furthermore, in this embodiment, step S4 is specifically as follows:

[0101] S41. Map each term in the second vocabulary to one or more concepts in a preset knowledge base. Each concept is described by a semantic primitive structure consisting of independent semantic primitives, relational semantic primitives, and symbolic semantic primitives.

[0102] S42. Based on the depth information and path distance of each basic semantic primitive in the semantic primitive tree, the following method for calculating the similarity between two basic semantic primitives is used:

[0103]

[0104] Where s1 and s2 are two basic primitives, and the adjustment parameter α = 0.5. Let s1 and s2 represent the depths of the basic primitives s1 and s2 in their respective primitive trees. The path distance dis(s1, s2) is defined as the sum of the weights of the edges on the shortest path connecting s1 and s2.

[0105]

[0106] Where n represents the number of edges contained in the shortest path between semantic primitives s1 and s2; level(k) represents the edge connecting the k-th level node and the (k+1)-th level node; the level number of the root node of the semantic primitive tree is defined as 0; and the weight of the edge weight(i) (i.e., the weight of the edge connecting the i-th level node and the (i+1)-th level node) is calculated as follows:

[0107]

[0108] Where m represents the maximum number of levels in the primitive tree, and the adjustment parameter θ = 4;

[0109] As one possible implementation, in HowNet, m=14, meaning the maximum height of the semantic primitive tree is 14.

[0110] In one embodiment of this application, the parent nodes and ancestor nodes of s1 and s2 can be found through the semantic primitive table, thereby obtaining their depths in their respective semantic primitive trees and the depths of all ancestor nodes. If the lowest common ancestor of two semantic primitives is found, their shortest path length can be calculated. If there is no lowest common ancestor, they belong to different semantic primitive trees, and the path length dis(s1,s2) is set to a large constant, such as 20. Furthermore, it is stipulated that the similarity between a specific word and a semantic primitive is set to 0.2, the similarity between two different specific words is 0, and the similarity between identical words is 1.

[0111] S43. Using the maximum value matching strategy, based on the basic semantic primitive similarity matrix, calculate the similarity between any two concepts in the independent semantic primitive dimension:

[0112] S431. Let any two independent primitives of meaning be X = x. i (i = 1, 2, ..., m), Y = y j (j=1,2,…,n), where m≤n, and create an m*n similarity matrix B:

[0113]

[0114] S432. Find the maximum value max(Sim(x) in matrix B. p ,y q Add it to the similarity list L;

[0115] S433. Set the element in row p and column q containing the maximum value to 0;

[0116] S434. Repeat steps S432-S433 above until all elements in matrix B are zero.

[0117] S435. Set nm unmatched elements y in the independent primitive Y. t (t = 1, 2, ..., nm) corresponds to an empty element And set the similarity between unmatched elements and empty elements. δ is a preset constant;

[0118] For example, δ can be set to a small constant, such as 0.2.

[0119] S436, will Add the data to the similarity list L, take the average of L, and obtain the similarity between the independent primitives X and Y:

[0120]

[0121] Among them, L k This represents the k-th element in list L;

[0122] S44. For any two concepts, based on the rules of attribute matching and value similarity calculation, obtain their relational semantic primitive similarity:

[0123] S441. Extract the relational primitive structure of each of the two concepts and represent it as a set composed of "attribute-value" pairs;

[0124] S442. Match the attributes in the two relational primitive structures; if any attribute exists in both concepts, extract the two corresponding attribute values ​​and calculate the similarity between the two attribute values ​​according to the basic primitive similarity calculation method; if any attribute does not exist in both concepts, set the attribute value corresponding to the attribute to an empty element, and set the similarity between the empty element and any other value to δ, where δ is a preset constant.

[0125] For example, δ can be set to a small constant, such as 0.2.

[0126] S443. Summarize the similarity of all attribute values ​​involved in the calculation, calculate their arithmetic mean, and use it as the similarity between the two concepts in the relational semantic primitive dimension.

[0127] S45. For any two concepts, based on the rules of attribute matching and value similarity calculation, obtain their symbolic semantic primitive similarity:

[0128] S451. Extract the relational primitive structure of each of the two concepts and represent it as a set composed of "attribute-value" pairs;

[0129] S452. Match the attributes in the two symbolic primitive structures; if any attribute exists in both concepts, extract the two corresponding attribute values ​​and calculate the similarity between the two attribute values ​​according to the basic primitive similarity calculation method; if any attribute does not exist in both concepts, set the attribute value corresponding to the attribute to an empty element, and set the similarity between the empty element and any other value to δ, where δ is a preset constant.

[0130] S453. Summarize the similarity of all attribute values ​​involved in the calculation, calculate their arithmetic mean, and use it as the similarity between the two concepts in the symbolic semantic primitive dimension.

[0131] S46. Combining the independent semantic primitives, relational semantic primitives, and symbolic semantic primitives, the similarity between any two concepts is calculated using a weighted average:

[0132]

[0133] Where i1, i2, and i3 correspond to independent primitives, relational primitives, and symbolic primitives, respectively, and β i=[0.7,0.17,0.13], Sim j (S1,S2) represents the similarity between the j-th semantic primitives of the two concepts;

[0134] S47. For each pair of terms, iterate through the combinations of their corresponding concepts, and select the concept pair with the highest similarity as the semantic similarity between the two terms, as shown below:

[0135]

[0136] Among them, S 1i ,S 2j Each represents all the concepts corresponding to the two terms;

[0137] S48. Fill the semantic similarity values ​​between all term pairs into the corresponding matrix positions to construct a symmetric semantic similarity matrix indexed by the second vocabulary:

[0138]

[0139] In this embodiment, the semantic similarity matrix is ​​an identity matrix because the similarity between any term and itself is always 1. Then, for all different term pairs in the vocabulary... i ,term j (where i = 1, ..., n, j = i + 1, ..., n), according to the aforementioned semantic similarity calculation method, calculate the corresponding semantic similarity Sim(term) sequentially. i ,term j Considering the symmetry of semantic similarity, that is:

[0140] Sem(i,j)=Sem(j,i)=Sim(term i ,term j ).

[0141] The semantic similarity matrix Sem constructed in this way is a symmetric matrix, and its form is as follows:

[0142]

[0143] S5. Sparsify the semantic similarity matrix to generate a sparse semantic similarity matrix;

[0144] Furthermore, in this embodiment of the application, step S5 includes:

[0145] Based on the constructed TF-IDF matrix A, the weight vector reflecting the importance of terms is obtained by summing the columns.

[0146] For each term iThe upper limit K of the corresponding number of matches is calculated based on its weight vector. i ,in, w i This represents the value of the term in the weight vector, where β is an adjustable parameter used to control the overall proportion of the number of matches.

[0147] Filter word pairs based on the set similarity threshold μ;

[0148] For each term i Select the maximum number of K values ​​from the i-th row of Sem. i The maximum similarity value greater than the threshold μ is selected, and the similarity values ​​that are not selected or are lower than the threshold μ are set to 0 to generate a sparse semantic similarity matrix Sem. sparse .

[0149] S6. Perform matrix operations on the sparse semantic similarity matrix and the initial term-inverse document matrix to obtain the semantically enhanced term-inverse document matrix;

[0150] Furthermore, in this embodiment of the application, step S6 includes:

[0151] A enhanced =Sem sparse ×A

[0152] Among them, A enhanced The semantically enhanced term-inverse document matrix has elements (A enhanced ) ij This represents the enhanced weight of term i in document j after considering semantic information.

[0153] S7. Based on the initial term-inverse document matrix and the semantically enhanced term-inverse document matrix, and combined with the weighted average method, generate the final document similarity.

[0154] Furthermore, in this embodiment of the application, step S7 includes:

[0155] Let document D i and D j The vector representations in the traditional vector space model are A i and A j In the semantic enhancement model, they are represented as A. enhanced,i and A enhanced,j Then its corresponding cosine similarity can be expressed as:

[0156]

[0157] By introducing a fusion parameter α∈[0,1], the similarity results from the two models are weighted and fused to obtain the final document similarity:

[0158] Final Sim (D i D j )

[0159] =α*Cos Sim (A i A j )+(1-α)*Cos Sim (A enhanced,i A enhanced,j )

[0160] Where α represents the weight of the original statistical information in the final similarity; 1-α represents the weight of the semantic enhancement information; when α=1, it relies entirely on the traditional vector space model; when α=0, it is based entirely on the semantic enhancement model.

[0161] This application provides a document vector similarity calculation method that combines semantic feature enhancement. First, the original document set is normalized and vocabulary is filtered to construct an initial TF-IDF matrix. Then, terms are mapped to semantic structures composed of multiple semantic primitives, and the semantic similarity between terms is calculated to construct a semantic similarity matrix. After sparsification of this matrix, a semantically enhanced term-document matrix is ​​generated and fused with the original matrix to calculate the final document similarity. This improves the semantic expressiveness of document similarity calculation, making the calculation results closer to real semantic associations and exhibiting stronger accuracy and robustness.

[0162] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0163] To implement the above embodiments, this application also proposes a terminal device.

[0164] Figure 2 This is a schematic diagram of the structure of a terminal device according to an embodiment of this application.

[0165] like Figure 2 As shown, the terminal device 200 includes:

[0166] The system includes a memory 210 and at least one processor 220, and a bus 230 connecting different components (including the memory 210 and the processor 220). The memory 210 stores a computer program, which, when executed by the processor 220, implements a document vector similarity calculation method combining semantic feature enhancement as described in the embodiments of this application.

[0167] Bus 230 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0168] Terminal device 200 typically includes various electronically readable media. These media can be any available media that can be accessed by terminal device 200, including volatile and non-volatile media, removable and non-removable media.

[0169] Memory 210 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 240 and / or cache memory 250. Terminal device 200 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 260 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 2 Not shown; usually referred to as a "hard drive"). Although Figure 2 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 230 via one or more data media interfaces. Memory 210 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this application.

[0170] A program / utility 280 having a set (at least one) of program modules 270 may be stored in, for example, memory 210. Such program modules 270 include—but are not limited to—an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 270 typically perform the functions and / or methods described in the embodiments of this application.

[0171] Terminal device 200 can also communicate with one or more external devices 290 (e.g., keyboard, pointing device, display 291, etc.), and with one or more devices that enable a user to interact with terminal device 200, and / or with any device that enables terminal device 200 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 292. Furthermore, terminal device 200 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 293. As shown, network adapter 293 communicates with other modules of terminal device 200 via bus 230. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with terminal device 200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0172] The processor 220 performs various functional applications and data processing by running programs stored in the memory 210.

[0173] It should be noted that the implementation process and technical principles of the terminal device in this embodiment are explained in the foregoing description of a document vector similarity calculation method combining semantic feature enhancement in the embodiments of this application, and will not be repeated here.

[0174] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0175] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0176] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0177] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0178] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0179] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0180] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0181] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A document vector similarity calculation method combining semantic feature enhancement, characterized in that, include: S1. Process the first document set according to the preset normalization processing method to generate the second document set; S2. Filter the first vocabulary according to preset constraints to generate a second vocabulary V = {term1, term2, ..., term...} n The first vocabulary is generated by deduplicating all terms in the second document set; S3. Based on the second vocabulary and the second document set, and using the TF-IDF weight calculation method, construct an initial term-inverse document matrix A = {W}. ij }; S4. Based on the preset knowledge base, map each word in the second vocabulary to a semantic structure composed of multiple semantic primitives, determine the semantic similarity between words according to the semantic primitive similarity calculation rules, and construct a semantic similarity matrix between words. S5. Sparsify the semantic similarity matrix to generate a sparse semantic similarity matrix Sem. sparse ; S6. Perform matrix operations on the sparse semantic similarity matrix and the initial term-inverse document matrix to obtain the semantically enhanced term-inverse document matrix; S7. Based on the initial term-inverse document matrix and the semantically enhanced term-inverse document matrix, and combined with the weighted average method, generate the final document similarity.

2. The method as described in claim 1, characterized in that, S1, processing the first document set according to a preset normalization processing method to generate a second document set, includes: The first document set is segmented using a preset word segmentation tool to generate a corresponding word list. Using a preset stop word list and combined with custom filtering rules based on the features of a preset corpus, meaningless or distracting words in each word list are removed to generate a second document set.

3. The method as described in claim 1, characterized in that, S2, filtering the first vocabulary according to preset constraints to generate a second vocabulary, includes: Retain terms that have appeared in at least 3 documents in the first vocabulary; Filter out terms that appear in more than 95% of documents from the first vocabulary list; The second vocabulary is generated by sorting each term according to its frequency of occurrence in the first document set and limiting the total number of retained terms to no more than 10,000.

4. The method as described in claim 1, characterized in that, Step S3 involves constructing an initial term-inverse document matrix A = {W} based on the second vocabulary and the second document set, and using the TF-IDF weight calculation method. ij The specific formula is as follows: Among them, tf i,j idf represents the frequency of term i in document j within the second document set. i df represents the inverse document frequency of term i. i The second document set contains the number of documents with term i, and m represents the total number of documents in the second document set.

5. The method as described in claim 1, characterized in that, S4 involves mapping each term in the second vocabulary to a semantic structure composed of multiple semantic primitives based on a preset knowledge base, determining the semantic similarity between terms according to the semantic primitive similarity calculation rules, and constructing a semantic similarity matrix between terms, including: S41. Map each term in the second vocabulary to one or more concepts in a preset knowledge base. Each concept is described by a semantic primitive structure consisting of independent semantic primitives, relational semantic primitives, and symbolic semantic primitives. S42. Based on the depth information and path distance of each basic semantic primitive in the semantic primitive tree, the following method for calculating the similarity between two basic semantic primitives is used: Where s1 and s2 are two basic primitives, and the adjustment parameter α = 0.

5. Let s1 and s2 represent the depths of the basic primitives s1 and s2 in their respective primitive trees. The path distance dis(s1, s2) is defined as the sum of the weights of the edges on the shortest path connecting s1 and s2. Where n represents the number of edges contained in the shortest path between semantic primitives s1 and s2; level(k) represents the edge connecting the k-th level node and the (k+1)-th level node; the level number of the root node of the semantic primitive tree is defined as 0, and the edge weight weight(i) is calculated as follows: Where m represents the maximum number of levels in the primitive tree, and the adjustment parameter θ = 4; S43. Using the maximum value matching strategy, based on the basic semantic primitive similarity matrix, calculate the similarity between any two concepts in the independent semantic primitive dimension: S431. Let any two independent primitives of meaning be X = x. i (i = 1, 2, ..., m), Y = y j (j=1,2,…,n), where m≤n, and create an m*n similarity matrix B: S432. Find the maximum value max(Sim(x) in matrix B. p ,y q Add it to the similarity list L; S433. Set the element in row p and column q containing the maximum value to 0; S434. Repeat steps S432-S433 above until all elements in matrix B are zero. S435. Set nm unmatched elements y in the independent primitive Y. t (t = 1, 2, ..., nm) corresponds to an empty element And set the similarity between unmatched elements and empty elements. δ is a preset constant; S436, will Add the data to the similarity list L, take the average of L, and obtain the similarity between the independent primitives X and Y: Among them, L k This represents the k-th element in list L; S44. For any two concepts, based on the rules of attribute matching and value similarity calculation, obtain their relational semantic primitive similarity: S441. Extract the relational primitive structure of each of the two concepts and represent it as a set consisting of "attribute-value" pairs; S442. Match the attributes in the two relational primitive structures; if any attribute exists in both concepts, extract the two corresponding attribute values ​​and calculate the similarity between the two attribute values ​​according to the basic primitive similarity calculation method; if any attribute does not exist in both concepts, set the attribute value corresponding to the attribute to an empty element, and set the similarity between the empty element and any other value to δ, where δ is a preset constant. S443. Summarize the similarity of all attribute values ​​involved in the calculation, calculate their arithmetic mean, and use it as the similarity between the two concepts in the relational semantic primitive dimension. S45. For any two concepts, based on the rules of attribute matching and value similarity calculation, obtain their symbolic semantic primitive similarity: S451. Extract the relational primitive structure of each of the two concepts and represent it as a set composed of "attribute-value" pairs; S452. Match the attributes in the two symbolic primitive structures; if any attribute exists in both concepts, extract the two corresponding attribute values ​​and calculate the similarity between the two attribute values ​​according to the basic primitive similarity calculation method; if any attribute does not exist in both concepts, set the attribute value corresponding to the attribute to an empty element, and set the similarity between the empty element and any other value to δ, where δ is a preset constant. S453. Summarize the similarity of all attribute values ​​involved in the calculation, calculate their arithmetic mean, and use it as the similarity between the two concepts in the symbolic semantic primitive dimension. S46. Combining the independent semantic primitives, relational semantic primitives, and symbolic semantic primitives, the similarity between any two concepts is calculated using a weighted average: Where i1, i2, and i3 correspond to independent primitives, relational primitives, and symbolic primitives, respectively, and β i =[0.7,0.17,0.13], Sim j (S1,S2) represents the similarity between the j-th semantic primitives of the two concepts; S47. For each pair of terms, iterate through the combinations of their corresponding concepts, and select the concept pair with the highest similarity as the semantic similarity between the two terms, as shown below: Among them, S 1i ,S 2j Each represents all the concepts corresponding to the two terms; S48. Fill the semantic similarity values ​​between all term pairs into the corresponding matrix positions to construct a symmetric semantic similarity matrix indexed by the second vocabulary:

6. The method as described in claim 1, characterized in that, Step S5 involves sparsifying the semantic similarity matrix to generate a sparse semantic similarity matrix Sem. sparse ,include: Based on the constructed TF-IDF matrix A, the weight vector reflecting the importance of terms is obtained by summing the columns. For each term i Calculate the upper limit K of the corresponding number of matches based on its weight vector. i ,in, w i This represents the value of the term in the weight vector, where β is an adjustable parameter used to control the overall proportion of the number of matches. Filter word pairs based on the set similarity threshold μ; For each term i Select the maximum number of K values ​​from the i-th row of Sem. i The maximum similarity value greater than the threshold μ is selected, and the similarity values ​​that are not selected or are lower than the threshold μ are set to 0 to generate a sparse semantic similarity matrix Sem. sparse .

7. The method as described in claim 1, characterized in that, In step S6, matrix operations are performed between the sparse semantic similarity matrix and the initial term-inverse document matrix to obtain the semantically enhanced term-inverse document matrix, as follows: THE enhanced =None sparse ×A Among them, A enhanced The semantically enhanced term-inverse document matrix has elements (A enhanced ) ij This represents the enhanced weight of term i in document j after considering semantic information.

8. The method as described in claim 1, characterized in that, Step S7 involves generating the final document similarity based on the initial term-inverse document matrix and the semantically enhanced term-inverse document matrix, combined with a weighted average method. include: Let document D i and D j The vector representations in the traditional vector space model are A i and A j The representations in the semantic enhancement model are A and B, respectively. enhanced,i and A enhanced,j Then its corresponding cosine similarity can be expressed as: By introducing a fusion parameter α∈[0,1], the similarity results from the two models are weighted and fused to obtain the final document similarity: Final Sim (D i ,D j ) =α*Cos Sim (A i ,A j )+(1-α)*Cos Sm (A enhanced,i ,A enhanced,j ) Where α represents the weight of the original statistical information in the final similarity; 1-α represents the weight of the semantic enhancement information; when α=1, it relies entirely on the traditional vector space model; when α=0, it is entirely based on the semantic enhancement model.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.