A text information retrieval method and device
By semantic generalization and screening of seed vocabulary, the scope and accuracy of text information retrieval is expanded, and the problem of low search efficiency in the existing technology is solved, and more efficient and accurate text information retrieval is achieved.
Patent Information
- Application Number
- CN202111438016.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-11-29
AI Technical Summary
The existing text information retrieval methods are inefficient and difficult to achieve user-satisfied search results, mainly because direct matching search cannot effectively expand the search range and accuracy.
By semantic generalization of the seed lexicon and filtering based on preset word selection conditions, the search scope and accuracy are expanded. At the same time, the words are filtered based on the quantity of text information, the vocabulary is optimized, and the search efficiency is improved.
The search scope is expanded, the retrieved text is made more comprehensive, the accuracy and efficiency of the retrieved text information is improved, and the problem of inaccurate text information is solved due to inaccurate words is solved.
Smart Images

Figure CN114003713B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data duplicate checking, and particularly to a text information retrieval method and device. Background Art
[0002] In the process of scientific research, text information retrieval is an important step. For example, before researching the current topic, it is necessary to retrieve text information such as literature, papers, and patents to check for duplicates of the current problem to be studied.
[0003] By checking for duplicates, enterprises can clarify the current scientific research trends, avoid problems of duplicate development and waste of funds, which is of great merit to enterprises. According to incomplete statistics, the losses caused by research topics losing value due to failure to consult papers, patents and other documents in various countries amount to billions, and the indirect losses are even more. Therefore, the retrieval of text information plays a crucial role in the growth of enterprises and the saving and improvement of global productivity.
[0004] However, the current conventional method of text information retrieval mainly directly matches and searches for the description information of the text information to be retrieved, with low retrieval efficiency and it is also difficult to achieve a retrieval effect satisfactory to users. Summary of the Invention
[0005] Embodiments of this application provide a text information retrieval method and device, which are used to solve the following technical problems: The current conventional method of text information retrieval mainly directly matches and searches for the description information of the text information to be retrieved, with low retrieval efficiency and it is also difficult to achieve a retrieval effect satisfactory to users.
[0006] Embodiments of this application adopt the following technical solutions:
[0007] Embodiments of this application provide a text information retrieval method. The method includes: determining a plurality of words according to the first technical field to which the text information to be retrieved belongs, and constructing a seed word library based on the plurality of words; performing semantic generalization on the seed word library, and screening the words after semantic generalization based on a first preset word selection condition to obtain a first-level generalized word library; performing text information retrieval based on the first-level generalized word library, determining the quantity of the text information corresponding to each word in the first-level generalized word library, and screening the words in the first-level generalized word library based on the quantity of the text information to obtain a second-level generalized word library; performing semantic generalization on the second-level generalized word library, and screening the generalized words based on a second preset word selection condition to obtain a third-level generalized word library, so as to perform text information retrieval based on the third-level generalized word library.
[0008] In the embodiments of the present application, semantic generalization is performed on the seed word library, and the words after semantic generalization are screened based on the first preset word selection condition to obtain a first-level generalized word library. Text information retrieval can be performed based on the generalized words, thereby expanding the retrieval scope and making the retrieved text more comprehensive. Secondly, in the embodiments of the present application, the words in the first-level generalized word library are screened based on the quantity of text information, so as to delete the text information with a small quantity, and solve the problem of inaccurate text information caused by inaccurate words. In addition, in the embodiments of the present application, semantic generalization is also performed on the second-level generalized word library, so as to further expand the words corresponding to the determined text information in the current field, thereby increasing the quantity of text information in the current technical field obtained, and making the content in the current technical field obtained more comprehensive. By continuously optimizing the word library, the embodiments of the present application can continuously improve the accuracy of the retrieved text information, and further improve the retrieval efficiency.
[0009] In one implementation manner of the present application, screening the words after semantic generalization based on the first preset word selection condition to obtain a first-level generalized word library specifically includes: determining first core words based on the words after semantic generalization and the words in the preset core corpus; determining the similarity value between the words after semantic generalization and the first core words, and determining the quantity of words with a similarity value greater than the first preset similarity value; when the quantity of words with a similarity value greater than the first preset similarity value meets the first preset word selection condition, forming a first-level generalized word library with the words corresponding to the similarity value.
[0010] In one implementation manner of the present application, screening the words in the first-level generalized word library based on the quantity of text information specifically includes: determining the quantity of retrieved text information corresponding to each word in the first-level generalized word library, and using the quantity of text information as the anti-weight coefficient corresponding to each word; when the quantity of text information corresponding to any word is greater than the first preset quantity value, adjusting the anti-weight coefficient to adjust the quantity of retrieved text information, and screening the words in the first-level generalized word library through the adjusted quantity of text information.
[0011] In the embodiments of the present application, by using the quantity of text information as the anti-weight coefficient corresponding to each word, the quantity of text information that can be retrieved by the current word can be adjusted by adjusting the anti-weight coefficient. Thus, the quantity of text information corresponding to the technical fields with more current research results is reduced to improve the retrieval efficiency of the user for the required text information.
[0012] In an implementation manner of the present application, the anti-weight coefficient is adjusted to adjust the quantity of the retrieved text information, and the words in the first-level generalization thesaurus are screened according to the quantity of the adjusted text information, specifically including: reducing the anti-weight coefficient corresponding to a word to reduce the quantity of the retrieved text information corresponding to the word; after adjusting the anti-weight coefficient, retrieving text information again, and when the quantity of the text information corresponding to any word in the first-level generalization thesaurus is less than the second preset quantity value, deleting the word.
[0013] In an implementation manner of the present application, semantic generalization is performed on the second-level generalization thesaurus, and the generalized words are selected based on the second preset word selection condition to obtain the third-level generalization thesaurus, specifically including: performing semantic generalization on the words in the second-level generalization thesaurus, and removing the repeated words that appear after semantic generalization; determining the second core words based on the remaining words after removal and the words in the preset core corpus; determining the similarity value between the remaining words and the second core words, and determining the quantity of the words whose similarity value is greater than the second preset similarity value; when the quantity of the words whose similarity value is greater than the second preset similarity value meets the second preset word selection condition, forming the third-level generalization thesaurus with the words corresponding to the similarity value.
[0014] In an implementation manner of the present application, after retrieving text information based on the third-level generalization thesaurus, the method further includes: determining the corresponding second technical field based on the words in the third-level generalization thesaurus, and determining the reference words corresponding to the second technical field based on the preset text corpus; wherein, the second technical field is any sub-field of the first technical field; the preset text corpus includes a plurality of reference words and the field categories corresponding to the plurality of words respectively; obtaining the retrieved text information corresponding to the third-level generalization thesaurus, and determining the reference words in the text information; comparing the reference words in the text information with the reference words corresponding to the second technical field to determine the field categories corresponding to the reference words in the text information respectively; classifying the retrieved text information according to the field categories corresponding to the reference words in the text information respectively.
[0015] The embodiments of the present application classify the retrieved multiple text information through the reference words corresponding to the retrieved text information, so as to classify the multiple text information corresponding to the current technical field according to sub-fields, which is convenient for users to analyze the text information of the required technical field and improves the efficiency of user retrieval.
[0016] In an implementation manner of the present application, the retrieved text information is classified according to the field categories respectively corresponding to the reference words in the text information, which specifically includes: counting the number of reference words in each text information, and sorting the reference words in descending order of the number; obtaining multiple reference words with serial numbers less than a preset serial number, and determining the field categories respectively corresponding to the multiple reference words, so as to use the field category corresponding to the reference word with the largest number as the field category of the current text information; clustering the text information with the same field category to realize the classification of the retrieved text information.
[0017] In an implementation manner of the present application, semantic generalization is performed on the seed word library, which specifically includes: performing vectorization calculation on multiple sample words through a pre-trained language model to obtain multiple feature vectors; calculating the correlation degree between every two of the multiple feature vectors to construct a concept tree according to the correlation degree; performing semantic generalization on the seed word library based on the concept tree.
[0018] In an implementation manner of the present application, semantic generalization is performed on the seed word library based on the concept tree, which specifically includes: performing vectorization calculation on the words in the seed word library to obtain token vectors; screening out word vectors in the concept tree whose correlation degree with the token vectors exceeds a preset correlation degree threshold, and obtaining the similar words corresponding to the word vectors, so as to perform semantic generalization on the seed word library through the similar words.
[0019] An embodiment of the present application provides a text information retrieval device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: determine multiple words according to the first technical field to which the text information to be retrieved belongs, and construct a seed word library based on the multiple words; perform semantic generalization on the seed word library, and screen the words after semantic generalization based on a first preset word selection condition to obtain a first-level generalized word library; perform text information retrieval based on the first-level generalized word library, determine the number of text information respectively corresponding to the words in the first-level generalized word library, and screen the words in the first-level generalized word library based on the number of text information to obtain a second-level generalized word library; perform semantic generalization on the second-level generalized word library, and screen the generalized words based on a second preset word selection condition to obtain a third-level generalized word library, so as to perform text information retrieval based on the third-level generalized word library.
[0020] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects: By semantic generalization of the seed word library and screening the words after semantic generalization based on the first preset word selection condition, the first-level generalization word library is obtained in the embodiments of the present application. Text information retrieval can be performed according to the generalized words, thereby expanding the retrieval scope and making the retrieved text more comprehensive. Secondly, in the embodiments of the present application, the words in the first-level generalization word library are screened based on the quantity of the text information, so as to delete the text information with a small quantity, and solve the problem of inaccurate text information caused by inaccurate words. In addition, the embodiments of the present application also perform semantic generalization on the second-level generalization word library, so as to further expand the words corresponding to the determined text information in the current field, thereby increasing the quantity of the text information in the current technical field obtained, and making the content in the current technical field obtained more comprehensive. By continuously optimizing the word library, the embodiments of the present application can continuously improve the accuracy of the retrieved text information, and further improve the retrieval efficiency. Brief Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the attached
[0022] In the figure:
[0023] Figure 1 is a flowchart of a text information retrieval method provided by an embodiment of the present application;
[0024] Figure 2 is a schematic structural diagram of a text information retrieval device provided by an embodiment of the present application. Detailed Embodiments
[0025] The embodiments of the present application provide a text information retrieval method and device.
[0026] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0027] In the process of scientific research, text information retrieval is an important step. For example, before researching the current topic, it is necessary to retrieve text information such as literature, papers, and patents to check for duplicate research on the problem to be studied currently.
[0028] By checking for duplicates, enterprises can clarify the current scientific research trends, avoid problems of duplicate development and waste of funds, which is of great benefit to enterprises. According to incomplete statistics, the losses caused by not consulting papers, patents and other literature in various countries, resulting in the loss of value of research topics, amount to billions, and the indirect losses are even more. Therefore, the retrieval of text information plays a crucial role in the growth of enterprises and the saving and improvement of global productivity.
[0029] However, the current conventional method of text information retrieval mainly directly matches and searches for the descriptive information of the text information to be retrieved, with low retrieval efficiency and it is also difficult to achieve a retrieval effect satisfactory to users.
[0030] To solve the above problems, the embodiments of the present application provide a text information retrieval method and device. By performing semantic generalization on the seed word library and screening the words after semantic generalization based on the first preset word selection condition, a first-level generalized word library is obtained. It is possible to perform text information retrieval based on the generalized words, thereby expanding the retrieval scope and making the retrieved text more comprehensive. Secondly, the embodiments of the present application screen the words in the first-level generalized word library based on the quantity of text information, thereby deleting the text information with a small quantity, and solving the problem of inaccurate text information caused by inaccurate words. In addition, the embodiments of the present application also perform semantic generalization on the second-level generalized word library, thereby further expanding the words corresponding to the determined text information in the current field, thereby increasing the quantity of text information in the current technical field obtained, and making the content in the current technical field obtained more comprehensive. The embodiments of the present application can continuously improve the accuracy of the retrieved text information and thus improve the retrieval efficiency by continuously optimizing the word library.
[0031] The technical solutions proposed by the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0032] Figure 1 It is a flowchart of a text information retrieval method provided by an embodiment of the present application. As Figure 1 shown, the text information retrieval method includes the following steps:
[0033] S101. The text information retrieval device determines a plurality of words according to the first technical field to which the text information to be retrieved belongs, and constructs a seed word library based on the plurality of words.
[0034] In one embodiment of the present application, the text information retrieval device determines a plurality of words according to the first technical field that the user needs to retrieve, and constructs a seed word library according to the determined words.
[0035] For example, if the user wants to retrieve text information in the field of real estate sales, after the user inputs the field of real estate sales, words such as sales amount, total sales amount, real estate, property, real estate market, real estate, price, etc. are generated according to the data information corresponding to the current field, and a seed word library is constructed according to the generated words.
[0036] S102. The text information retrieval device performs semantic generalization on the seed word library, and filters the words after semantic generalization based on the first preset word selection condition to obtain a first-level generalized word library.
[0037] In one embodiment of the present application, vectorization calculation is performed on a plurality of sample words through a pre-trained language model to obtain a plurality of feature vectors. The correlation degree between each two of the plurality of feature vectors is calculated to construct a concept tree according to the correlation degree. The seed word library is semantically generalized based on the concept tree.
[0038] Specifically, a plurality of sample words need to be obtained first. For example, text corpus can be obtained, and the text corpus can be segmented based on a grammar and rule-based word segmentation method to obtain a plurality of words. Or the text corpus can be segmented based on a mechanical word segmentation method (i.e., a dictionary) to obtain a plurality of words.
[0039] Furthermore, vectorization calculation is performed on each of the plurality of words to obtain a plurality of feature vectors. For example, a pre-trained language model can be obtained, and the pre-trained language model is used to perform vectorization calculation on each of the plurality of words. Here, the pre-trained language model can include an autoregressive language model or an autoencoder language model. And the pre-trained semantic models that can be used can be models such as GloVe, word2vec, and fastText. In addition, a trained dual TriNet model can be obtained, and the trained dual TriNet model is used to perform vectorization calculation on each of the plurality of words. The correlation degree or similarity between each two of the plurality of feature vectors is calculated, and a concept tree is constructed according to the correlation degree or similarity.
[0040] In one embodiment of the present application, vectorization calculation is performed on the words in the seed word library to obtain token vectors. Words vectors in the concept tree whose correlation degree with the token vectors exceeds a preset correlation degree threshold are filtered out, and the similar words corresponding to the words vectors are obtained, so as to semantically generalize the seed word library through the similar words.
[0041] Specifically, a pre-trained language model can also be used to perform vectorization calculations on the words in the seed word library to obtain the token vectors corresponding to each word. Calculate the correlation between the token vector corresponding to each word and the word vector in the concept tree, that is, calculate the similarity between the two. Determine the word vectors in the concept tree whose similarity is greater than the preset correlation threshold, and determine the corresponding words in the concept tree according to the word vectors, so as to use the words as the similar words corresponding to those in the seed word library, and perform semantic generalization on the seed word library through the obtained multiple similar words.
[0042] In an embodiment of the present application, based on the words after semantic generalization and the words in the preset core corpus, the first core words are determined. Determine the similarity value between the words after semantic generalization and the first core words, and determine the number of words whose similarity value is greater than the first preset similarity value. When the number of words greater than the first preset similarity value meets the first preset word selection condition, the words corresponding to the similarity value are combined into a first-level generalization word library.
[0043] Specifically, compare the words after semantic generalization with the words in the preset core corpus, and use the words that are the same as the preset core corpus as the first core words. And calculate the similarity value between the words after semantic generalization and the first core words. For example, the similarity value can be calculated through Ontology or Taxonomy. Based on the synonym dictionary, all words are organized in one or several tree structures, and the path length between two nodes can be used as the semantic distance, and the similarity value between different words can be calculated through the semantic distance.
[0044] Furthermore, calculate the similarity value between the words after semantic generalization and the first core words, and compare the calculated similarity value with the first preset similarity value to obtain the number of words whose similarity value is greater than the first preset similarity value. For example, the first preset similarity value can be set to 0.5, obtain the number of words whose similarity value is greater than 0.5. When the number is greater than or equal to 50, it indicates that the current word meets the first preset word selection condition. That is, the words whose similarity value is greater than 0.5 are combined into a first-level generalization word library.
[0045] It should be noted that in the embodiment of the present application, the first preset similarity value is preferably set to 0.5, but is not limited to 0.5 only. In applications, the first preset similarity value can be adjusted according to actual situations. And, in the embodiment of the present application, the quantity value is preferably set to 50, but is not limited to 50 only. In applications, the quantity value can be adjusted according to actual situations.
[0046] S103. The text information retrieval device performs text information retrieval based on the first-level generalization thesaurus, determines the quantity of text information corresponding to each word in the first-level generalization thesaurus, and filters the words in the first-level generalization thesaurus based on the quantity of text information to obtain the second-level generalization thesaurus.
[0047] In an embodiment of the present application, an automatic interaction is performed between the first-level generalization thesaurus and the retrieval system, and text information retrieval is performed in the retrieval system using the words in the first-level generalization thesaurus. And the quantity of text information corresponding to each word is obtained.
[0048] In an embodiment of the present application, it is determined the quantity of retrieved text information corresponding to each word in the first-level generalization thesaurus, and the quantity of text information is used as the anti-weight coefficient corresponding to each word. When the quantity of text information corresponding to any word is greater than the first preset quantity value, the anti-weight coefficient is adjusted to adjust the quantity of retrieved text information, and the words in the first-level generalization thesaurus are filtered using the adjusted quantity of text information.
[0049] Specifically, the quantity of text information corresponding to each word in the first-level generalization thesaurus is used as the anti-weight coefficient of the word. The larger the quantity, the higher the content heat in the technical field corresponding to the word and the more research results. At this time, in order to avoid the phenomenon of duplicate research and try to obtain text information corresponding to technical fields with not particularly high research heat, the anti-weight coefficient corresponding to the word can be reduced, thereby reducing the quantity of text information that can be retrieved corresponding to the word, so as to avoid the appearance of text information in this highly heated technical field as much as possible.
[0050] For example, the quantity of text information retrieved by any word in the first-level generalization thesaurus reaches 1000, which has reached the first preset quantity value at this time. Then the research heat in the technical field corresponding to the word is relatively high. Therefore, in order to avoid the problem of duplicate research, the anti-weight coefficient corresponding to the word can be reduced, so as to reduce the appearance of text information in this field as much as possible, enabling the user to obtain more text information corresponding to technical fields with not particularly high research heat.
[0051] In an embodiment of the present application, the anti-weight coefficient corresponding to the word is reduced to reduce the quantity of retrieved text information corresponding to the word. After adjusting the anti-weight coefficient, text information retrieval is performed again, and when the quantity of text information corresponding to any word in the first-level generalization thesaurus is less than the second preset quantity value, the word is deleted.
[0052] Specifically, after adjusting the anti-weight coefficient, retrieve again based on the words in the first-level generalization thesaurus. Since the anti-weight coefficients of one or more words have been changed, the number of retrieved text information will change. At this time, determine the words whose number of text information is less than the second preset quantity value. For example, the quantity value can be set to 10. Delete the words whose number of text information is less than 10 to screen the first-level generalization thesaurus and obtain the second-level generalization thesaurus.
[0053] In the embodiment of the present application, by reducing the anti-weight coefficient corresponding to a word, the number of retrieved text information corresponding to the word can be reduced, thereby reducing the number of retrieved text information corresponding to the technical field with a higher popularity. Furthermore, the risk of repeated research on the same topic is reduced, facilitating users to obtain the text information materials they need more conveniently.
[0054] It should be noted that in the embodiment of the present application, the quantity value is preferably set to 10, but is not limited to 10 only. In the application, the quantity value can be adjusted according to the actual situation.
[0055] S104. The text information retrieval device performs semantic generalization on the second-level generalization thesaurus, and screens the generalized words based on the second preset word selection condition to obtain a third-level generalization thesaurus for text information retrieval based on the third-level generalization thesaurus.
[0056] In an embodiment of the present application, perform semantic generalization on the words in the second-level generalization thesaurus, and eliminate the repeated words that appear after semantic generalization. Based on the remaining words after elimination and the words in the preset core corpus, determine the second core words. Determine the similarity value between the remaining words and the second core words, and determine the number of words whose similarity value is greater than the second preset similarity value. When the number of words whose similarity value is greater than the second preset similarity value meets the second preset word selection condition, form a third-level generalization thesaurus with the words corresponding to the similarity value.
[0057] Specifically, the words corresponding to the technical fields with a higher popularity in the second-level generalization thesaurus have been deleted. Therefore, the technical field directions corresponding to the words in the second-level generalization thesaurus have been further narrowed. Based on the pre-constructed concept tree, perform semantic generalization on the words in the second-level generalization thesaurus. And determine the repeated words after semantic generalization and delete them.
[0058] Further, compare the remaining words in the secondary generalization thesaurus with the words in the preset core corpus to determine the second core words. Similarly, the similarity value between the remaining words and the second core words can be calculated through Ontology or Taxonomy. And at the same time, obtain the number of words with a similarity value greater than the second preset similarity value. For example, the second preset similarity value can be set to 0.7, and the number of words with a similarity value greater than 0.7 is obtained. When the number is not greater than 10, it indicates that the current word meets the second preset word selection condition. That is, the words with a similarity value greater than 0.7 are grouped into a tertiary generalization thesaurus. If the number of words with a similarity value greater than 0.7 is more than 10, at this time, the preset second similarity value can be increased to reduce the number of words in the tertiary generalization thesaurus.
[0059] In an embodiment of the present application, based on the words in the tertiary generalization thesaurus, determine the corresponding second technical field, and based on the preset text corpus, determine the reference words corresponding to the second technical field. Wherein, the second technical field is any sub-field of the first technical field. The preset text corpus includes multiple reference words and the domain categories corresponding to the multiple words respectively. Obtain the retrieved text information corresponding to the tertiary generalization thesaurus, and determine the reference words in the text information. Compare the reference words in the text information with the reference words corresponding to the second technical field to determine the domain categories corresponding to the reference words in the text information respectively. Classify the retrieved text information according to the domain categories corresponding to the reference words in the text information respectively.
[0060] Specifically, according to the words in the tertiary generalization thesaurus, it can be determined which second technical fields, that is, sub-fields, belong to the first technical field, thereby narrowing the scope of text information retrieval. For example, the words in the tertiary generalization thesaurus can be compared with a preset database to determine the second technical fields to which the words in the tertiary generalization thesaurus belong. Wherein, the preset database contains the words corresponding to multiple technical fields respectively, and multiple words corresponding to the sub-fields to which each technical field belongs respectively. And, based on the preset text corpus, determine the reference words corresponding to the second technical field. Query the reference words for the text information corresponding to each word in the tertiary generalization thesaurus, and compare the queried reference words with the reference words corresponding to the second technical field to achieve the classification of the retrieved text information.
[0061] In one embodiment of the present application, the reference words in each text information are counted, and the reference words are sorted in descending order of the quantity. A plurality of reference words with serial numbers less than the preset serial number are obtained, and the corresponding field categories of the plurality of reference words are determined, so as to use the field category corresponding to the reference word with the largest quantity as the field category of the current text information. The text information with the same field category is clustered to classify the retrieved text information.
[0062] Specifically, the reference words in each retrieved text information are counted, and the quantity of the same reference words is counted, so as to sort the reference words in the same text information according to the quantity. The reference words with higher quantity values are obtained, that is, the reference words with more occurrences. And the field category corresponding to the reference word with more occurrences is determined, that is, the second technical field is further divided into fields to determine the divided technical field corresponding to the reference word. The quantity of the reference words corresponding to each divided technical field is counted, and the field with the largest quantity of reference words is used as the technical field to which the text information belongs. Through the above method, the fields of all the retrieved text information are determined, so as to classify the text information. It is convenient for users to analyze and process text information in different fields.
[0063] Figure 2 The figure is a schematic structural diagram of a text information retrieval device provided by an embodiment of the present application. As Figure 2 shown, the text information retrieval device includes:
[0064] At least one processor; and,
[0065] A memory communicatively connected to the at least one processor; wherein,
[0066] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can:
[0067] Determine a plurality of words according to the first technical field to which the text information to be retrieved belongs, and construct a seed word library based on the plurality of words;
[0068] Perform semantic generalization on the seed word library, and screen the words after semantic generalization based on the first preset word selection condition to obtain a first-level generalized word library;
[0069] Perform text information retrieval based on the first-level generalized word library, determine the quantity of the text information corresponding to each word in the first-level generalized word library, and screen the words in the first-level generalized word library based on the quantity of the text information to obtain a second-level generalized word library;
[0070] Semantically generalize the secondary generalization thesaurus, and screen the generalized words based on the second preset word selection condition to obtain a tertiary generalization thesaurus, so as to perform text information retrieval based on the tertiary generalization thesaurus.
[0071] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0072] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0073] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, the embodiments of the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included within the scope of the claims of the present application.
Claims
1. A text information retrieval method, characterized in that, The method includes: Determining a plurality of words according to the first technical field to which the text information to be retrieved belongs, and constructing a seed word library based on the plurality of words; Performing semantic generalization on the seed word library, and screening the semantically generalized words based on a first preset word selection condition to obtain a first-level generalized word library; Performing text information retrieval based on the first-level generalized word library, determining the number of text information corresponding to each word in the first-level generalized word library, and screening the words in the first-level generalized word library based on the number of the text information to obtain a second-level generalized word library; Performing semantic generalization on the second-level generalized word library, and screening the generalized words based on a second preset word selection condition to obtain a third-level generalized word library, so as to perform text information retrieval based on the third-level generalized word library.
2. The text information retrieval method according to claim 1, wherein The screening the semantically generalized words based on the first preset word selection condition to obtain a first-level generalized word library specifically includes: Determining first core words based on the semantically generalized words and the words in a preset core corpus; Determining the similarity value between the semantically generalized words and the first core words, and determining the number of words whose similarity value is greater than a first preset similarity value; When the number of words whose similarity value is greater than the first preset similarity value meets the first preset word selection condition, forming the words corresponding to the similarity value into the first-level generalized word library.
3. A text information retrieval method according to claim 1, characterized in that The screening the words in the first-level generalized word library based on the number of the text information specifically includes: Determining the number of retrieved text information corresponding to each word in the first-level generalized word library, and using the number of the text information as the anti-weight coefficient corresponding to each word; When the number of the retrieved text information corresponding to any word is greater than a first preset quantity value, adjusting the anti-weight coefficient to adjust the number of the retrieved text information, and screening the words in the first-level generalized word library through the adjusted number of the text information.
4. A text information retrieval method according to claim 3, characterized in that, The adjusting the anti-weight coefficient to adjust the number of the retrieved text information, and screening the words in the first-level generalized word library through the adjusted number of the text information specifically includes: Reducing the anti-weight coefficient corresponding to the word to reduce the number of the retrieved text information corresponding to the word; After adjusting the anti-weight coefficient, re-performing text information retrieval, and when the number of the retrieved text information corresponding to any word in the first-level generalized word library is less than a second preset quantity value, deleting the word.
5. A text information retrieval method according to claim 1, characterized in that, The performing semantic generalization on the second-level generalized word library, and selecting the generalized words based on a second preset word selection condition to obtain a third-level generalized word library specifically includes: Performing semantic generalization on the words in the second-level generalized word library, and removing the repeated words that appear after semantic generalization; Determining second core words based on the remaining words after removal and the words in a preset core corpus; Determining the similarity value between the remaining words and the second core words, and determining the number of words whose similarity value is greater than a second preset similarity value; When the number of words greater than the second preset similarity value meets the second preset word selection condition, the words corresponding to the similarity value are combined to form the third-level generalization thesaurus.
6. A text information retrieval method according to claim 1, characterized in that After performing text information retrieval based on the third-level generalization thesaurus, the method further includes: Based on the words in the third-level generalization thesaurus, determine the corresponding second technical field, and based on a preset text corpus, determine the reference words corresponding to the second technical field; wherein, the second technical field is any sub-field of the first technical field; the preset text corpus includes multiple reference words and the field categories corresponding to the multiple words respectively; Obtain the retrieved text information corresponding to the third-level generalization thesaurus, and determine the reference words in the text information; Compare the reference words in the text information with the reference words corresponding to the second technical field to determine the field categories corresponding to the reference words in the text information respectively; Classify the retrieved text information according to the field categories corresponding to the reference words in the text information respectively.
7. A text information retrieval method according to claim 6, characterized in that The classifying the retrieved text information according to the field categories corresponding to the reference words in the text information respectively specifically includes: Count the number of reference words in each text information, and sort the reference words in descending order of the number; Obtain multiple reference words with serial numbers less than a preset serial number, and determine the field categories corresponding to the multiple reference words respectively, so as to use the field category corresponding to the reference word with the largest number as the field category of the current text information; Cluster the text information with the same field category to realize the classification of the retrieved text information.
8. A text information retrieval method according to claim 1, characterized in that The semantic generalization of the seed thesaurus specifically includes: Perform vectorization calculation on multiple sample words through a pre-trained language model to obtain multiple feature vectors; Calculate the correlation degree between every two of the multiple feature vectors to construct a concept tree according to the correlation degree; Perform semantic generalization on the seed thesaurus based on the concept tree.
9. A text information retrieval method according to claim 8, wherein The performing semantic generalization on the seed thesaurus based on the concept tree specifically includes: Perform vectorization calculation on the words in the seed thesaurus to obtain token vectors; Screen out the word vectors in the concept tree whose correlation degree with the token vector exceeds a preset correlation degree threshold, and obtain the similar words corresponding to the word vectors, so as to perform semantic generalization on the seed thesaurus through the similar words.
10. A text information retrieval device, comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: Determine multiple words according to the first technical field to which the text information to be retrieved belongs, and construct a seed thesaurus based on the multiple words; Semantically generalize the seed word library, and screen the words after semantic generalization based on the first preset word selection condition to obtain a first-level generalized word library; Based on the first-level generalized word library, perform text information retrieval, determine the quantity of text information corresponding to each word in the first-level generalized word library, and screen the words in the first-level generalized word library based on the quantity of the text information to obtain a second-level generalized word library; Semantically generalize the second-level generalized word library, and screen the generalized words based on the second preset word selection condition to obtain a third-level generalized word library, so as to perform text information retrieval based on the third-level generalized word library.
Citation Information
Patent Citations
Full-text database accurate and efficient retrieval method for perfecting subject terms
CN111831786A
System and method for offering searching service basedon topics
KR1020070040162A