Word embedding model document screening query method and system

By using a pre-trained word embedding model to convert documents and vocabulary into high-dimensional numerical vectors, the problem of insufficient vocabulary semantic relationship capture in the prior art is solved, and more accurate document screening and information retrieval is achieved.

CN120256593APending Publication Date: 2025-07-04INST OF ECONOMIC & TECH STATE GRID HEBEI ELECTRIC POWER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510310839.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing document filtering query methods such as bag-of-word model and TF-IDF cannot effectively capture the semantic relationship between vocabulary, resulting in inaccurate filtering in the case of synonyms or polysynonyms.

Method used

Pre-trained word embedding models such as Word2Vec, GloVe or BERT are used to convert target vocabulary and documents into numerical vectors in high-dimensional space, filter out relevant documents by calculating similarity sorting, and display detailed information.

Benefits of technology

It improves the accuracy and relevance of document screening, adapts to the needs of different application scenarios, and enhances the efficiency and accuracy of information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256593A_ABST
    Figure CN120256593A_ABST
Patent Text Reader

Abstract

The invention discloses a word embedding model document screening query method and system. The method comprises the following steps: obtaining a target word to be queried and a document set to be screened; preprocessing the document set, wherein the preprocessing comprises stop word removal, word stem extraction or word form reduction; converting the target vocabulary and each document in the document set into a word embedding vector by using a pre-trained word embedding model; calculating the similarity between the word embedding vector of the target vocabulary and the word embedding vector of each document in the document set; sorting the document set according to the similarity, and selecting the first N documents with the highest similarity as screening results; and sorting the screened first N documents according to the similarity from high to low, and displaying the sorted documents to the user. According to the method, the target vocabularies and documents can be converted into numerical vectors in a high-dimensional space by using the pre-trained word embedding model, and the word embedding vectors not only capture the surface features of the vocabularies, but also reflect the complex semantic relationship between the vocabularies, so that the accuracy and correlation of document screening are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and specifically to a method and system for screening and querying documents based on a word embedding model. Background Art

[0002] With the rapid development of information technology, the number of documents on the Internet has increased explosively. How to quickly and accurately screen out the content related to the user's query from a large number of documents has become an urgent problem to be solved. The technology of screening and querying documents based on a word embedding model has emerged as the times require. It converts words and documents into numerical representations in a high-dimensional space by using a word embedding model, and then calculates the similarity between them, so as to realize the rapid screening and sorting of documents.

[0003] The word embedding model can capture and reflect the complex semantic relationships between words, and map semantically similar words to adjacent positions in a high-dimensional space. Therefore, the method of screening and querying documents based on a word embedding model can better understand the user's query intention, screen out the content most relevant to the user's query from a large number of documents, and greatly improve the efficiency and accuracy of information retrieval.

[0004] In the early screening and querying of documents, common methods such as the Bag-of-Words (BoW) or TF-IDF only considered the frequency of word occurrences, ignored the order and context relationships between words, and could not effectively capture the semantic connections between words. Traditional methods cannot handle the problems of synonyms or polysemous words well because they rely on exact word matching. When the target word is not exactly the same as the words in the document, relevant documents may be missed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the above technical defects and provide a method and system for screening and querying documents based on a word embedding model.

[0006] To solve the above problems, the technical solution of the present invention is a method for screening and querying documents based on a word embedding model, including the following steps:

[0007] S1. Obtain the target word to be queried and the document set to be screened;

[0008] S2. Preprocess the document set, including removing stop words, stemming or lemmatization;

[0009] S3. Use a pre-trained word embedding model to convert the target word and each document in the document set into word embedding vectors;

[0010] S4. Calculate the similarity between the word embedding vector of the target word and the word embedding vectors of each document in the document set;

[0011] S5. Sort the document set according to the similarity, and select the top N documents with the highest similarity as the screening results;

[0012] S6. Sort the top N screened documents from the highest similarity to the lowest and display them to the user.

[0013] Furthermore, the word embedding vector in step S3 is a numerical representation of a word or a document in a high-dimensional space, which can capture and reflect the complex semantic relationships between words. The pre-trained word embedding model in step S3 is obtained by an unsupervised learning method based on a large-scale corpus, such as Word2Vec, GloVe or BERT, etc.

[0014] Furthermore, the screening results displayed in step S5 may further include detailed information of each document, such as the author, the release time, the source link, etc., so that the user can understand the screened documents more comprehensively.

[0015] Furthermore, in step S1, the user can input or select the target word in various ways, such as manually inputting, selecting from a predefined word list, or extracting keywords from the sentence input by the user through natural language processing technology, etc.

[0016] Furthermore, in step S4, at least one of cosine similarity, Euclidean distance or Manhattan distance is used for similarity calculation. The cosine similarity measures the similarity degree of two vectors in direction, while the Euclidean distance and the Manhattan distance measure the difference degree of two vectors in value.

[0017] A word embedding model document screening and query system includes the following modules:

[0018] An input module, which is used to receive the target word input by the user and the document set to be screened;

[0019] A preprocessing module, which is used to preprocess the document set, including removing stop words, stemming or lemmatization;

[0020] A word embedding conversion module, which is used to convert the target word and each document in the document set into word embedding vectors by using a pre-trained word embedding model;

[0021] A similarity calculation module, which is used to calculate the similarity between the word embedding vector of the target word and the word embedding vectors of each document in the document set;

[0022] A screening result output module, which is used to sort the document set according to the similarity, select the top N documents with the highest similarity as the screening results and display them to the user.

[0023] Furthermore, the pre-trained word embedding model used by the word embedding conversion module is obtained through an unsupervised learning method based on a large-scale corpus, such as Word2Vec, GloVe, or BERT, etc. It can capture and reflect the complex semantic relationships between words, and represent words or documents as numerical vectors in a high-dimensional space.

[0024] Furthermore, the screening result output module can not only display the top N screened documents and their similarity scores, but also further include detailed information of each document, such as the author, publication time, source link, etc., to help users more comprehensively understand and evaluate the content of the screened documents.

[0025] Furthermore, the input module supports multiple ways to obtain target words, including but not limited to manual input, selection from a predefined word list, or automatic extraction of keywords from sentences input by users through natural language processing techniques.

[0026] Furthermore, the similarity calculation module adopts at least one similarity measurement method, including but not limited to cosine similarity, Euclidean distance, or Manhattan distance. Among them, cosine similarity measures the similarity degree of two vectors in direction, while Euclidean distance and Manhattan distance measure the difference degree of two vectors in terms of numerical values.

[0027] The advantages of the present invention compared with the existing technologies are as follows:

[0028] 1. The present invention provides a method and system for screening and querying documents using a word embedding model. By using a pre-trained word embedding model (such as Word2Vec, GloVe, or BERT), the present invention can convert target words and documents into numerical vectors in a high-dimensional space. These word embedding vectors not only capture the surface features of words, but also can reflect the complex semantic relationships between words, including synonyms, antonyms, and different meanings of polysemous words, thus greatly improving the accuracy and relevance of document screening.

[0029] 2. The present invention provides a method and system for screening and querying documents using a word embedding model. The pre-trained word embedding model is obtained through an unsupervised learning method based on a large-scale corpus, which ensures the strong generalization ability and wide application scope of the model. At the same time, the present invention allows adjusting or fine-tuning the pre-trained model according to the characteristics of a specific field to better meet the needs of different application scenarios and improve the relevance and professionalism of screening results. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a flowchart of a method and system for screening and querying documents using a word embedding model according to the present invention.

[0031] Figure 2It is a system diagram of a method and system for screening and querying documents of a word embedding model according to the present invention. Detailed implementation manners

[0032] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims.

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0034] As Figure 1 shown, this embodiment proposes a method for screening and querying documents of a word embedding model, including the following steps:

[0035] S1. Obtain the target vocabulary to be queried and the document set to be screened;

[0036] S2. Preprocess the document set, including removing stop words, stemming or lemmatization;

[0037] S3. Use a pre-trained word embedding model to convert the target vocabulary and each document in the document set into word embedding vectors;

[0038] S4. Calculate the similarity between the word embedding vector of the target vocabulary and the word embedding vectors of each document in the document set;

[0039] S5. Sort the document set according to the similarity, and select the top N documents with the highest similarity as the screening results;

[0040] S6. Sort the top N screened documents from high to low according to the similarity and display them to the user.

[0041] Furthermore, in step S1, the user can input or select the target vocabulary in various ways, such as manually inputting, selecting from a predefined vocabulary list, or extracting keywords from the sentence input by the user through natural language processing technology, etc.

[0042] Furthermore, the word embedding vectors in step S3 are numerical representations of words or documents in a high-dimensional space, which can capture and reflect the complex semantic relationships between words. The pre-trained word embedding model in step S3 is obtained through unsupervised learning methods based on large-scale corpora, such as Word2Vec, GloVe, or BERT, etc.

[0043] Furthermore, in step S4, the similarity calculation uses at least one of cosine similarity, Euclidean distance, or Manhattan distance. Cosine similarity measures the similarity degree of two vectors in terms of direction, while Euclidean distance and Manhattan distance measure the difference degree of two vectors in terms of numerical values.

[0044] Furthermore, the screening results presented in step S5 can further include detailed information of each document, such as author, publication time, source link, etc., so that users can understand the screened documents more comprehensively.

[0045] As Figure 2 shown, this embodiment proposes a word embedding model document screening and query system, including the following modules:

[0046] An input module, which is used to receive the target vocabulary input by the user and the document set to be screened;

[0047] A preprocessing module, which is used to preprocess the document set, including removing stop words, stemming, or lemmatization;

[0048] A word embedding conversion module, which is used to convert the target vocabulary and each document in the document set into word embedding vectors using a pre-trained word embedding model;

[0049] A similarity calculation module, which is used to calculate the similarity between the word embedding vector of the target vocabulary and the word embedding vectors of each document in the document set;

[0050] A screening result output module, which is used to sort the document set according to the similarity, select the top N documents with the highest similarity as the screening results and display them to the user.

[0051] Furthermore, the input module supports multiple ways to obtain the target vocabulary, including but not limited to manual input, selection from a predefined vocabulary list, or automatic extraction of keywords from the sentence input by the user through natural language processing technology.

[0052] Furthermore, the pre-trained word embedding model used by the word embedding conversion module is obtained through unsupervised learning methods based on large-scale corpora, such as Word2Vec, GloVe, or BERT, etc., which can capture and reflect the complex semantic relationships between words and represent words or documents as numerical vectors in a high-dimensional space.

[0053] Furthermore, the similarity calculation module adopts at least one similarity measurement method, including but not limited to cosine similarity, Euclidean distance or Manhattan distance. Among them, cosine similarity measures the similarity degree of two vectors in terms of direction, while Euclidean distance and Manhattan distance measure the difference degree of two vectors in terms of numerical values.

[0054] Furthermore, the screening result output module can not only display the top N screened documents and their similarity scores, but also further include the detailed information of each document, such as the author, publication time, source link, etc., to help users more comprehensively understand and evaluate the content of the screened documents.

[0055] When specifically used, refer to Figure 1 and Figure 2 As shown, the user inputs the target vocabulary, such as "machine learning", through the input module provided by the system. This module supports multiple input methods, including but not limited to manual input, selection from a predefined vocabulary list, or automatic extraction of keywords from the sentences input by the user using natural language processing technology. The user can upload a folder containing multiple documents or specify the system to access a document collection in a specific database. These documents may or may not be related to the target vocabulary and constitute the document pool to be screened.

[0056] The preprocessing module first removes the stop words (such as high-frequency meaningless words like "de" and "le") in the documents to reduce noise information and improve the effectiveness of subsequent analysis. According to statistics, about 5% of the vocabulary is identified as stop words and removed. Then, stemming or lemmatization is performed on the remaining vocabulary to normalize different morphological synonyms. For example, "running" will be reduced to "run", and "dogs" and "dog" will be regarded as the same vocabulary. This process reduces about 10% of the different morphological but semantically identical vocabulary forms and enhances the consistency and accuracy of document features.

[0057] The preprocessing module processes large-scale document collections through distributed storage and processing, inverted indexing, vector indexing, and hybrid processing modes. Using a distributed file system (such as HDFS) and a distributed computing framework (such as Apache Spark or Dask), it divides the document collection into multiple small chunks and distributes them to different computing nodes for preprocessing, transformation, and similarity calculation, which can significantly improve the processing speed. It constructs an inverted index (Inverted Index) to create a list for each term that contains all the document IDs where it appears and their location information. This enables quick location of relevant documents during queries, reducing unnecessary full-text scanning operations. It uses specially designed vector index structures (such as FAISS, Annoy, etc.), which can quickly retrieve the document vectors most similar to the target term in a high-dimensional space. For existing document collections, they can be processed in one go through batch processing; for continuously generated new documents, stream processing technologies (such as Apache Kafka + Flink) can be used to achieve real-time updates. When the document collection is so large that it cannot be fully loaded into memory, the system will adopt a paging loading method, only loading the data chunks that need to be processed currently, avoiding memory overflow problems.

[0058] The word embedding transformation module uses pre-trained word embedding models (such as Word2Vec, GloVe, or BERT, etc.) to transform the target term and each document in the document collection into a numerical vector representation in a high-dimensional space. Suppose we choose the Word2Vec model, which can map each term or document to a 300-dimensional vector space. For each document, instead of simply taking the average of all non-stop words, we calculate the weighted average according to the TF-IDF weights, which can better capture the core theme of the document. Finally, each document is represented as a 300-dimensional word embedding vector.

[0059] The similarity calculation module uses cosine similarity as the main measurement method because it can effectively measure the similarity degree of the directions of two vectors and is particularly suitable for text data comparison. Specifically, for the word embedding vector V of the target term "machine learning" target and the word embedding vector V of a certain document in the document collection doc , their cosine similarity can be calculated through the following formula:

[0060]

[0061] Among them, · represents the dot product operation, and ‖V‖ represents the norm of the vector. The system also supports other similarity measurement methods such as Euclidean distance or Manhattan distance, allowing users to select the most suitable calculation method according to actual needs. The calculated similarity score usually ranges between [-1, 1], where 1 indicates exactly the same direction, and -1 indicates the opposite direction.

[0062] The screening result output module sorts the document set according to the calculated similarity score and selects the top N documents with the highest similarity as the screening results. Here, N is a parameter specified by the user, which reflects the number of relevant documents the user hopes to obtain. To further optimize the result quality, the system also performs a secondary sorting by combining factors such as the novelty and authority of the documents.

[0063] Finally, the screening result output module not only displays the top N screened documents and their similarity scores, but also provides detailed information about each document, such as the author, publication time, source link, etc., to help users more comprehensively understand and evaluate the content of the screened documents. The display interface is user-friendly and supports multiple view modes such as list, chart, or summary to meet the preferences of different users.

[0064] With the development of technology and the change of application scenarios, we will continue to explore more efficient word embedding models and similarity calculation methods, such as pre-trained models based on the Transformer architecture like BERT, as well as more intelligent sorting algorithms. Future work also includes integrating more natural language processing technologies and machine learning models to enable the system to adapt to more diverse query needs and provide more personalized and accurate services.

[0065] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.

[0066] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

[0067] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.

Claims

1. A method for screening and querying word embedding model documents, characterized in that It includes the following steps: S1. Obtain the target vocabulary to be queried and the document set to be screened; S2. Preprocess the document set, including removing stop words, stemming or lemmatization; S3. Use a pre-trained word embedding model to convert the target vocabulary and each document in the document set into word embedding vectors; S4. Calculate the similarity between the word embedding vector of the target vocabulary and the word embedding vectors of each document in the document set; S5. Sort the document set according to the similarity, and select the top N documents with the highest similarity as the screening results; S6. Sort the top N screened documents from high to low similarity and display them to the user.

2. The method for screening and querying a word embedding model document according to claim 1, wherein: In step S1, the user can input or select the target vocabulary in various ways, such as manual input, selection from a predefined vocabulary list, or extraction of keywords from the sentence input by the user through natural language processing technology, etc.

3. The word embedding model document screening and query method according to claim 1, characterized in that: In step S3, the word embedding vector is a numerical representation of a vocabulary or document in a high-dimensional space, which can capture and reflect the complex semantic relationships between vocabularies. The pre-trained word embedding model in step S3 is obtained through an unsupervised learning method based on a large-scale corpus, such as Word2Vec, GloVe or BERT, etc.

4. A method for screening and querying a word embedding model document according to claim 1, characterized in that: In step S4, the similarity calculation adopts at least one of cosine similarity, Euclidean distance or Manhattan distance. Cosine similarity measures the similarity degree of two vectors in direction, while Euclidean distance and Manhattan distance measure the difference degree of two vectors in value.

5. A method for screening and querying a word embedding model document according to claim 1, characterized in that: In step S5, the displayed screening results can further include detailed information of each document, such as author, publication time, source link, etc., so that the user can understand the screened documents more comprehensively.

6. A word embedding model document screening and querying system, characterized in that, It includes the following modules: Input module, which is used to receive the target vocabulary input by the user and the document set to be screened; Preprocessing module, which is used to preprocess the document set, including removing stop words, stemming or lemmatization; Word embedding conversion module, which is used to use a pre-trained word embedding model to convert the target vocabulary and each document in the document set into word embedding vectors; Similarity calculation module, which is used to calculate the similarity between the word embedding vector of the target vocabulary and the word embedding vectors of each document in the document set; Screening result output module, which is used to sort the document set according to the similarity, select the top N documents with the highest similarity as the screening results and display them to the user.

7. The word embedding model document screening and querying system according to claim 6, characterized in that: The input module supports various ways to obtain the target vocabulary, including but not limited to manual input, selection from a predefined vocabulary list, or automatic extraction of keywords from the sentence input by the user through natural language processing technology.

8. The word embedding model document screening and querying system according to claim 6, characterized in that: The pre-trained word embedding model used by the word embedding conversion module is obtained through an unsupervised learning method based on a large-scale corpus, such as Word2Vec, GloVe or BERT, etc., which can capture and reflect the complex semantic relationships between vocabularies and represent the vocabulary or document as a numerical vector in a high-dimensional space.

9. The word embedding model document screening and querying system according to claim 6, wherein: The similarity calculation module uses at least one similarity measurement method, including but not limited to cosine similarity, Euclidean distance or Manhattan distance. Among them, cosine similarity measures the similarity degree of two vectors in direction, while Euclidean distance and Manhattan distance measure the difference degree of two vectors in numerical value.

10. A word embedding model document screening and querying system according to claim 6, characterized in that: The screening result output module can not only display the top N screened documents and their similarity scores, but also further include detailed information of each document, such as author, publication time, source link, etc., to help users more comprehensively understand and evaluate the content of the screened documents.