An intelligent recommendation method and device for search results and a unified search method

By calculating the weight values ​​of the original query keywords and expanded terms, and combining a unified retrieval syntax converter and a data resource deduplication model, the problems of retrieval efficiency and accuracy of multi-source heterogeneous tobacco science and technology literature resources were solved, and efficient knowledge recommendation was achieved.

CN116450772BActive Publication Date: 2025-12-16ZHENGZHOU TOBACCO RES INST OF CNTC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310151132.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-22
Publication Date
2025-12-16
Estimated Expiration
2043-02-22

AI Technical Summary

Technical Problem

Existing methods for retrieving tobacco science and technology literature cannot effectively integrate multi-source heterogeneous resources, resulting in low retrieval efficiency and poor accuracy, and failing to meet the knowledge needs of interdisciplinary and cross-domain fields.

Method used

We employ an intelligent recommendation method for search results based on word vectors. By calculating the combined positional weight and domain feature weight of the original query keywords and expanded terms, we achieve in-depth similarity calculation for query expansion tasks. Combined with a unified search syntax converter and a data resource deduplication model, we improve the relevance and accuracy of search results.

Benefits of technology

It significantly improves the accuracy of retrieval and recommendation of multi-source heterogeneous tobacco science and technology literature resources, helping users to quickly and accurately find the content they need, and solving the problems of resource redundancy and difficulty in discovering knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116450772B_ABST
    Figure CN116450772B_ABST
Patent Text Reader

Abstract

The application provides a retrieval result intelligent recommendation method, device and unified retrieval method, and the retrieval result intelligent recommendation method comprises the following steps: S1, acquiring a key word in a scientific literature field and calculating a field characteristic weight value of each key word in the scientific literature field; S2, acquiring a query sentence, determining a query key word of the query sentence, and acquiring a retrieval result from a scientific literature resource library according to the query key word; S3, calculating a comprehensive position weight value and a value of each query key word according to the retrieval result; S4, calculating a correlation degree of the retrieval result and the query sentence based on the comprehensive position weight value, the value and the field characteristic weight value of each query key word, and sorting the retrieval result according to the correlation degree.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of tobacco scientific literature resource retrieval, and in particular to a retrieval result intelligent recommendation method and device and a unified retrieval method. BACKGROUND

[0002] Scientific literature resources contain a large amount of information knowledge and are important knowledge bases. For the tobacco field, scientific literature such as papers and patent achievements contains 85-90% of scientific and technological information in the tobacco field. Effective use of the knowledge information contained in these literature materials can avoid repeated research, improve work efficiency, reduce work costs, and also provide theoretical and technical support for the development of new products and new technologies. The tobacco industry has always attached great importance to investment in scientific and technological innovation. In recent years, it has accumulated a large amount of scientific and technological literature information resources in scientific research, technical development and production and operation activities, such as tobacco scientific papers, tobacco patents, tobacco standards, tobacco scientific and technological achievements and other data, and the data volume has reached millions. The explosive growth of tobacco scientific literature resources has prompted the demand for scientific and technological information resources in the tobacco industry to evolve from simply obtaining resources to precise knowledge service needs, which has posed great challenges to the resource integration capability, information retrieval capability, knowledge precise positioning capability and knowledge analysis capability of the tobacco scientific literature platform. Deep integration of tobacco scientific literature resources from multiple sources, different structures and various data types, and provision of unified retrieval and precise recommendation services are important methods for tobacco researchers to obtain knowledge resources across disciplines, fields and retrieval libraries, and are effective support means for improving literature information resource sharing mechanism and improving literature information service level.

[0003] The tobacco field is a field spanning multiple disciplines, covering biology, chemistry, agriculture, process, etc. From the data source of tobacco technology literature, tobacco technology literature resources can be generally divided into two categories: self-owned literature resource library and purchased literature resource library. These tobacco technology literature resources are often huge in quantity, various in type and different in structure, requiring a large amount of manpower to configure and maintain data sources to provide literature retrieval services for users, which is very costly. In addition, since the purchased resources are provided by different literature data resource service providers, the data structure, storage means, publishing mechanism, retrieval method, display form, etc. of the data resources are very different, and various types of tobacco technology literature resources cannot achieve effective knowledge fusion and accurate knowledge retrieval service. In the face of huge amount of tobacco technology literature resources with extensive sources and different data structures and organization forms, how to deeply integrate, uniformly retrieve and accurately recommend the massive multi-source heterogeneous tobacco technology literature resources, realize the ordered organization, rapid positioning and effective disclosure of the tobacco technology literature resources, and uniformly load, uniformly display and uniformly analyze the retrieval results, help users efficiently and accurately find the retrieval content, so as to improve the retrieval efficiency and accuracy of tobacco technology literature retrieval, has become a problem that needs to be solved in the current tobacco technology literature retrieval field.

[0004] Due to the differences in service modes of various tobacco technology literature database providers, the commonly used uniform retrieval technology is generally for specific database types, and the quality and sorting effect of the retrieval results are not ideal when facing massive multi-source heterogeneous tobacco literature resources, and the interoperability of heterogeneous databases is poor. The existing tobacco literature retrieval method is usually based on the keyword-based way to match the retrieval results, and the limited keywords submitted to the search engine often cannot completely express the retrieval information requirements, and due to the differences between human language and machine language, the search engine usually loses semantic information when processing user queries. Moreover, the tobacco field spans multiple disciplines, involving a wide range of technology literature types and literature scope, and there are a large number of tobacco-specific terms and some abbreviations, synthetic words, etc. in the tobacco field. In the face of multi-source heterogeneous tobacco technology literature resources, the traditional retrieval method has low retrieval efficiency and performance. In addition, the retrieval method based on query keywords often sorts the retrieval results by counting the frequency of query keywords appearing in the retrieval literature, ignoring the user's retrieval intention and semantic environment, resulting in that the recall rate and precision rate of massive multi-source heterogeneous tobacco technology literature retrieval results often cannot achieve ideal results.

[0005] In order to solve the above problems, people have been seeking an ideal technical solution. SUMMARY

[0006] The purpose of this invention is to address the shortcomings of the existing technology. When searching for multi-source heterogeneous tobacco science and technology literature, an intelligent recommendation method for search results is introduced based on the search conditions. By calculating the weight of the original query keywords and query expansion terms, not only can the importance of the query keywords be reflected and the query expansion task be better completed, but also a deeper similarity calculation can be performed on the query expansion terms and search results, thereby improving the retrieval performance of science and technology literature.

[0007] This invention also provides a unified retrieval method for multi-source heterogeneous tobacco science and technology literature resources, enabling unified retrieval and intelligent recommendation of massive multi-source heterogeneous tobacco science and technology literature resources, thereby significantly improving the accuracy of multi-source heterogeneous tobacco science and technology literature resource retrieval and result recommendation, and helping users quickly and accurately find the search content.

[0008] To achieve the above objectives, the technical solution adopted by this invention is: an intelligent recommendation method for search results, comprising the following steps:

[0009] S1, Obtain keywords in the field of scientific and technological literature, and calculate the domain feature weight of each keyword in the field of scientific and technological literature;

[0010] S2, obtain the query statement, determine the query keywords of the query statement, and obtain the search results from the scientific and technological literature resource database based on the query keywords;

[0011] S3, Calculate the comprehensive positional weight value of each query keyword based on the search results and value;

[0012] S4, based on the comprehensive positional weight value of each query keyword, The search results are ranked according to the relevance of the search results to the query statement, using the domain feature weights and the domain feature weights.

[0013] Specifically, the steps for S1 are as follows:

[0014] A corpus of scientific and technological literature was constructed. After removing stop words and performing word segmentation on the corpus, the TF-IDF algorithm was used to extract keywords from the corpus.

[0015] Calculate the domain feature weights for each keyword: ;

[0016] in, As keywords, Keywords Inverse document frequency in scientific and technological literature corpora, Represents a logarithmic function.

[0017] Specifically, the steps for determining the query keywords in S2 are as follows:

[0018] The query statement is input in the search box, and the basic keywords of the query statement are obtained after word segmentation and stop word removal operations;

[0019] The trained scientific literature word vector model is used to calculate the similarity between each expansion keyword and the basic keyword in the pre-constructed scientific literature keyword expansion library using the cosine similarity calculation method.

[0020] The expansion keywords are sorted in descending order of similarity, and the top preset number of expansion keywords are selected for synonym merging to obtain the expansion keywords.

[0021] Each basic keyword and its corresponding expansion keyword are used as query keywords.

[0022] Further, after obtaining the expansion keywords, the word vectors of each basic keyword and its corresponding expansion keyword are normalized and regularized to obtain the word vector space model of each basic keyword and its corresponding expansion keyword.

[0023] Based on the word vector space model of each basic keyword and its corresponding expansion keyword, the feature weight of each basic keyword and its corresponding expansion keyword in the query statement is calculated.

[0024] i

[0025] wherein, represents the feature weight value, S represents the vector of the query statement, represents the vector of the basic keyword or the expansion keyword, represents the similarity order of the expansion keyword and the basic keyword, i=0, represents the basic keyword, represents the similarity of the basic keyword or the expansion keyword and the query statement;

[0026] In the calculation of the relevance between the retrieval results and the query statement based on the comprehensive position weight value of each query keyword, the value and the field feature weight value, the scientific literature retrieval result evaluation analysis model is used for calculation.

[0027] wherein, represents the frequency of each query keyword after word segmentation of the query statement in a retrieval result; represents the inverse document frequency of the query keyword in all retrieval results; representing the query keyword field feature weight value of the technical literature field, for the query keyword comprehensive position weight value.

[0028] The second aspect of the present application provides a retrieval result intelligent recommendation device, comprising:

[0029] The keyword feature weight value acquisition unit is used for acquiring the technical literature field keywords and calculating the field feature weight value of each keyword in the technical literature field.

[0030] The query unit is used for acquiring the query statement, determining the query keyword of the query statement, and acquiring the retrieval result from the technical literature resource library according to the query keyword.

[0031] The comprehensive position weight value acquisition unit is used for calculating the comprehensive position weight value of each query keyword according to the retrieval result.

[0032] The calculation unit is used for calculating the value of each query keyword in the retrieval result.

[0033] The sorting unit is used for calculating the relevance of the retrieval result and the query statement based on the comprehensive position weight value, value and the field feature weight value of each query keyword, and sorting the retrieval result according to the relevance size.

[0034] The third aspect of the present application provides a unified retrieval method for multi-source heterogeneous tobacco technical literature resources, comprising the following steps:

[0035] Step 1, according to the syntax characteristics and logical operation methods of the retrieval formula of each available tobacco technical literature resource library, a unified retrieval syntax converter is constructed; according to the resource type of the tobacco technical literature resource, a tobacco technical literature data resource deduplication model is constructed; according to the field richness of the tobacco technical literature and the content richness of the literature, a tobacco technical literature quality evaluation model is constructed;

[0036] Step 2, acquiring the technical literature field keywords and calculating the field feature weight value of each keyword in the technical literature field.

[0037] Step 3, inputting the query statement into the unified retrieval box, converting the query statement into a preset unified syntax through the unified retrieval syntax converter, determining the query keyword of the query statement through the word segmentation and stop word removal operation; acquiring the retrieval result from each tobacco technical literature resource library according to the query keyword.

[0038] Step 4, the structural formatting, content cleaning, data normalization processing are carried out on the search results, and based on the constructed tobacco science literature data resource deduplication model, the tobacco science literature information fingerprint is extracted for deduplication and integration operation; and based on the tobacco science literature quality evaluation model, the quality of the search results is evaluated, the low-quality search results are removed, and the high-quality search results are retained;

[0039] Step 5, the comprehensive position weight value of each query keyword is calculated according to the search results Value;

[0040] Step 6, based on the comprehensive position weight value of each query keyword, Value and the field feature weight value, the relevance of the search results and the query sentence is calculated, and the search results are sorted according to the relevance.

[0041] The application also provides a computer storage medium, and the computer readable storage medium stores a computer program, which is executed by a processor to realize the aforementioned search result intelligent recommendation method.

[0042] The application has outstanding substantial features and significant progress, specifically, the application introduces a search result matching method based on a word vector for a query sentence, through the comprehensive position weight value, Value and the field feature weight value calculation of the original basic keyword and the query expansion word, the importance of the query word can be embodied, the query expansion task can be better completed, and thus the search performance of the science literature is improved.

[0043] When the weight of the original basic keyword and the query expansion word is calculated, the feature weight of each basic keyword and the corresponding expansion keyword in the query sentence is further calculated, so that the similarity calculation of the query expansion word and the search result can be deeper, the importance of the query word can be further embodied, and thus the search performance of the science literature is improved.

[0044] The application also provides a unified search method for multi-source heterogeneous tobacco science literature resources, through the tobacco science literature data resource deduplication model, the unified search syntax converter and the intelligent recommendation algorithm, the unified search and intelligent recommendation of the massive multi-source heterogeneous tobacco science literature resources are realized, the accuracy of the multi-source heterogeneous tobacco science literature resource search and result recommendation is greatly improved, the user can quickly and accurately find the search content, and the problems of redundancy and difficult knowledge discovery of the multi-source heterogeneous tobacco science literature resources are solved. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 is a flowchart of the intelligent recommendation method of the embodiment 1 of the application.

[0046] Figure 2is a flowchart of the intelligent recommendation method according to Embodiment 2 of the present application.

[0047] Figure 3 is a flowchart of the intelligent recommendation method according to Embodiment 3 of the present application.

[0048] Figure 4 is a flowchart of the unified search method according to Embodiment 5 of the present application. DETAILED DESCRIPTION

[0049] The features and exemplary embodiments of the various aspects of the present application will be described in detail below with reference to the drawings. To make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are configured only to explain the present application and are not configured to limit the present application.

[0050] For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by showing examples of the present application.

[0051] It should be noted that, in this document, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element preceded by "comprises... " does not, without more limitations, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the stated elements.

[0052] In the present application, TF-IDF is involved, where TF-IDF (term frequency-inverse document frequency) is a statistical method used to evaluate the importance of a word to a document in a document set or corpus. TF is the number of occurrences of a word in an article, and IDF is the inverse of the number of occurrences of a word in all documents. The more a word appears in a document and the less it appears in all documents, the more it can represent the article, so the product of TF-IDF can be used to measure the importance of a word in a document. TF-IDF weighting can be used as a measure or rating of the relevance between the search results and the user query.

[0053] The formula for calculating TF is as follows:

[0054]

[0055] in, Represents a term in a text Number of times it appears This represents the total number of terms.

[0056] The formula for calculating IDF is as follows:

[0057]

[0058] Where Y is the total number of documents in the corpus, This is the number of documents containing the term "w". To avoid... For words that don't appear in any documents, resulting in a denominator of 0, the formula needs to be smoothed by adding one to the denominator so that even words not found in the corpus can obtain a suitable IDF value.

[0059] By defining TF and IDF, we can further calculate the TF-IDF value of a word w:

[0060] .

[0061] The technical solution of the present invention will be further described in detail below through specific embodiments.

[0062] Example 1

[0063] like Figure 1 As shown, this embodiment provides an intelligent recommendation method for search results, including the following steps:

[0064] S1, Obtain keywords in the field of scientific and technological literature, and calculate the domain feature weight of each keyword in the field of scientific and technological literature;

[0065] S2, obtain the query statement, determine the query keywords of the query statement, and obtain the search results from the scientific and technological literature resource database based on the query keywords;

[0066] S3, Calculate the comprehensive positional weight value of each query keyword based on the search results and value;

[0067] S4, based on the comprehensive positional weight value of each query keyword, The search results are ranked according to the relevance of the search results to the query statement, using the domain feature weights and domain feature weights.

[0068] Preferably, a scientific literature retrieval result evaluation and analysis model is adopted. , calculate the relevance of the search results and the query statement;

[0069] wherein, ;

[0070] In the formula, represents each query keyword after the query statement is tokenized The frequency of occurrence in a search result; represents the query keyword The inverse document frequency of all search results; represents the query keyword The field feature weight value of the keyword in the field of scientific literature, The comprehensive position weight value of the query keyword .

[0071] It can be seen that the embodiment introduces a search result matching method based on word vectors for the query statement. Through the calculation of the comprehensive position weight value of the original basic keyword and the query expansion word, The value, and the field feature weight value, the importance of the basic keyword and the query expansion word can be reflected, and the query expansion task can be better completed, thereby improving the search performance of scientific literature.

[0072] Embodiment 2

[0073] The embodiment gives a specific embodiment, as shown in Figure 2 The specific steps are as follows:

[0074] S1, obtain the keywords in the field of scientific literature, and calculate the field feature weight value of each keyword in the field of scientific literature, and the specific steps are as follows:

[0075] Construct a scientific literature corpus, after removing stop words and tokenizing the scientific literature corpus, extract the keywords of the scientific literature corpus by using the TF-IDF algorithm;

[0076] Calculate the field feature weight of each keyword: ;

[0077] wherein, is the keyword, represents the query keyword The inverse document frequency of the scientific literature corpus, represents the logarithmic function;

[0078] S2, obtain the query statement, determine the query keyword of the query statement, and obtain the search result from the scientific literature resource library according to the query keyword;

[0079] Wherein, the specific steps of determining the query keyword of the query statement are as follows:

[0080] The query statement is input in the search box, and the basic keywords of the query statement are obtained after word segmentation and stop word removal operations;

[0081] The trained sci-tech literature word vector model is used to calculate the similarity between each expansion word and the basic keyword in the pre-constructed sci-tech literature keyword expansion library by using the cosine similarity calculation method, to obtain the similarity between each expansion word and the basic keyword;

[0082] The expansion words are sorted in descending order of similarity, and the top preset number of expansion words are selected, and the same / synonymous word merging is performed to obtain the expansion keywords;

[0083] Each basic keyword and its corresponding expansion keyword are used as query keywords;

[0084] S3, according to the search result, the comprehensive position weight value of each query keyword and value are calculated;

[0085] S4, based on the sci-tech literature search result evaluation analysis model , the relevance of the search result and the query statement is calculated, and the search result is sorted according to the relevance;

[0086] Wherein, ;

[0087] In the formula, represents the frequency of each query keyword of the query statement after word segmentation in a search result; represents the inverse document frequency of the query keyword in all search results; represents the field feature weight of the query keyword in the field of sci-tech literature, is the comprehensive position weight value of the query keyword .

[0088] In specific implementation, since the keyword vocabulary belongs to human language, and the computer cannot understand human language, in order to facilitate computer operation, it is necessary to map the keyword vocabulary to the dimension that the computer can understand, that is, the word vector.

[0089] word2vec is a commonly used word vector classic model, its principle is that in a sentence, the surrounding words of a word have strong correlation with this word, and other words have poor correlation, according to this idea, a neural network is constructed to train the current word and its context word, and finally the word vector is obtained.

[0090] In this embodiment, Word2vec is also used to obtain word vectors. Specifically, the training steps of the scientific literature word vector model are as follows: constructing a scientific literature corpus, after removing stop words and performing word segmentation on the scientific literature corpus, training and learning the scientific literature corpus by using a Word2vec embedding model, generating a scientific literature word vector model, and obtaining the word vector form of each keyword.

[0091] In a specific implementation, the construction steps of the scientific literature keyword expansion library are as follows:

[0092] Constructing a scientific literature corpus, extracting literature keywords of the scientific literature corpus by using a TF-IDF algorithm, and obtaining a keyword expansion library through synonym and near-synonym expansion;

[0093] According to the literature keywords in the keyword expansion library, on the basis of the scientific and technological classification table and the field term table, combining the tobacco field thesaurus, a tobacco scientific keyword dictionary is constructed;

[0094] Using the trained scientific literature word vector model, the semantic similarity of tobacco vocabulary is calculated by using a cosine similarity calculation method in the scientific literature corpus, and a tobacco vocabulary semantic similarity matching model is constructed;

[0095] Based on the tobacco vocabulary semantic similarity matching model, synonym and near-synonym expansion is performed on each tobacco vocabulary in the tobacco scientific keyword dictionary, and a tobacco scientific literature keyword expansion library is obtained.

[0096] As can be seen, the training steps of the scientific literature word vector model, the construction steps of the scientific literature keyword expansion library, and the calculation of the field feature weight of each keyword in the scientific literature field all include the step of constructing a scientific literature corpus and extracting literature keywords of the scientific literature corpus by using a TF-IDF algorithm. Therefore, the scientific literature corpus can be constructed first, the literature keywords of the scientific literature corpus can be extracted by using a TF-IDF algorithm, and then the training steps of the scientific literature word vector model, the construction steps of the scientific literature keyword expansion library, and the calculation of the field feature weight of each keyword in the scientific literature field can be performed simultaneously.

[0097] Embodiment 3

[0098] The difference between this embodiment and embodiment 2 is that, as shown in Figure 3 after obtaining the expanded keywords, the word vectors of each basic keyword and its corresponding expanded keywords are normalized and regularized to obtain a word vector space model of each basic keyword and its corresponding expanded keywords;

[0099] Based on the word vector space model of each basic keyword and its corresponding extended keyword, the feature weight of each basic keyword and its corresponding extended keyword in the query sentence is calculated respectively:

[0100]

[0101] wherein, represents the feature weight value, S represents the vector of the query sentence, represents the vector of the basic keyword or the extended keyword, represents the similarity order of the extended keyword and the basic keyword, i=0, represents the basic keyword, represents the similarity of the basic keyword or the extended keyword and the query sentence;

[0102] Construction of scientific literature retrieval result evaluation analysis model is: .

[0103] In this embodiment, when calculating the weight of the original basic keyword and the query expansion word, the feature weight of each basic keyword and its corresponding extended keyword in the query sentence is further calculated, so that the similarity calculation of the query expansion word and the retrieval result can be deeper, the importance of the query word can be further reflected, and the retrieval performance of the scientific literature can be improved.

[0104] Embodiment 4

[0105] The embodiment provides a retrieval result intelligent recommendation device, which comprises:

[0106] A keyword feature weight value acquisition unit is configured to acquire keywords in the field of scientific literature and calculate the field feature weight value of each keyword in the field of scientific literature.

[0107] A query unit is configured to acquire a query sentence, determine query keywords of the query sentence, and acquire retrieval results from a scientific literature resource library according to the query keywords.

[0108] A comprehensive position weight value acquisition unit is configured to calculate the comprehensive position weight value of each query keyword according to the retrieval results.

[0109] A calculation unit is configured to calculate the value of each query keyword in the retrieval results.

[0110] A sorting unit is configured to calculate the relevance of the retrieval results and the query sentence based on the comprehensive position weight value, the value of each query keyword and the field feature weight value, and sort the retrieval results according to the relevance. ​​

[0111] In the implementation, the field feature weight value of each keyword in the field of scientific literature, the query keyword of the query sentence, the comprehensive position weight value of each query keyword, and The value and the specific calculation steps of calculating the relevance of the retrieval result and the query sentence are the same as those described in the foregoing embodiments 1-3.

[0112] Embodiment 5

[0113] The embodiment provides a unified retrieval method for multi-source heterogeneous tobacco scientific literature resources, which comprises the following steps as shown in the figure: Figure 4

[0114] Step 1, according to the syntax characteristics and logical operation methods of the retrieval formula of each available tobacco scientific literature resource library, a unified retrieval syntax converter is constructed; according to the resource type of the tobacco scientific literature resource, a tobacco scientific literature data resource deduplication model is constructed; according to the field richness of the tobacco scientific literature and the content richness of the literature, a tobacco scientific literature quality evaluation model is constructed;

[0115] Step 2, scientific literature field keywords are obtained, and the field feature weight value of each keyword in the field of scientific literature is calculated;

[0116] Step 3, a query sentence is input into the unified retrieval box, and after being converted into a preset unified syntax by the unified retrieval syntax converter, the query keywords of the query sentence are determined through word segmentation and stop word removal operation; retrieval results are obtained from each tobacco scientific literature resource library according to the query keywords;

[0117] Step 4, the retrieval results are subjected to structure formatting, content cleaning, data normalization processing, and based on the constructed tobacco scientific literature data resource deduplication model, the tobacco scientific literature information fingerprint is extracted for deduplication and integration operation; and based on the tobacco scientific literature quality evaluation model, the quality of the retrieval results is evaluated, the low-quality retrieval results are removed, and the high-quality retrieval results are retained;

[0118] Step 5, the comprehensive position weight value of each query keyword and the value are calculated according to the retrieval results;

[0119] Step 6, based on the comprehensive position weight value of each query keyword, the value and the field feature weight value, the relevance of the retrieval result and the query sentence is calculated, and the retrieval results are sorted according to the relevance.

[0120] ​The unified retrieval method for multi-source heterogeneous tobacco scientific literature resources provided in the embodiment, through a tobacco scientific literature data resource deduplication model, a unified retrieval syntax converter and an intelligent recommendation algorithm, realizes unified retrieval and intelligent recommendation of massive multi-source heterogeneous tobacco scientific literature resources, thereby greatly improving the accuracy of retrieval and result recommendation of multi-source heterogeneous tobacco scientific literature resources, helping users quickly and accurately find the retrieval content, and solving the problems of redundancy and difficulty in finding knowledge of multi-source heterogeneous tobacco scientific literature resources.

[0121] Embodiment 6

[0122] The application provides a specific embodiment of a unified retrieval method for multi-source heterogeneous tobacco scientific literature resources, specifically comprising the following steps:

[0123] Step 11, first, the interface service provided by the purchased tobacco scientific literature resource library is checked for health, whether the interface service is available is judged, further, the available purchased tobacco scientific literature resource library is determined, according to the syntax characteristics and logical operation methods of the retrieval formula of each available tobacco scientific literature resource library, a unified retrieval syntax converter is constructed.

[0124] Step 12, according to the resource type of the tobacco scientific literature resource, a tobacco scientific literature data resource deduplication model is constructed.

[0125] Step 13, based on the tobacco scientific literature field richness and the literature content richness, a tobacco scientific literature quality evaluation model is constructed;

[0126] Step 21, according to the characteristics of the tobacco scientific literature, a tobacco scientific literature corpus is constructed, the corpus is subjected to stop word removal and word segmentation operation, the TF-IDF algorithm is adopted to extract the keywords of the tobacco scientific literature text corpus, and the IDF values of all the keywords are calculated.

[0127] The IDF values and TF values of part of the keywords in the tobacco scientific literature corpus are shown in the following table.

[0128] Tobacco vocabulary IDF values TF values Tobacco 50.7745913 17679 Flue-cured tobacco 62.1350577 15697 Tobacco leaf 71.635847 4575 Float culture 155.243684 1227 Mainstream smoke 172.38371 1105 Oriental tobacco 172.696283 1103 Phyllocyanin 266.027972 786 Mosaic disease 324.273764 663 Whitefly 349.937008 554

[0129] Step 22, combining the characteristics of the tobacco scientific literature and the tobacco glossary word meaning features, the analytic hierarchy process is adopted to further calculate the field feature weight of the keywords obtained in step 21, and the calculation formula is as follows:

[0130]

[0131] Wherein, is the keyword, indicates the keyword The inverse document frequency of the scientific literature corpus, indicates the logarithmic function.

[0132] The feature weight values of the selected partial keywords in step 21 in the tobacco science and technology literature field are as follows:

[0133] Tobacco vocabulary IDF values TF values Feature weight Tobacco 50.7745913 17679 0.71 Flue-cured tobacco 62.1350577 15697 0.79 Tobacco leaf 71.635847 4575 0.86 Float culture 155.243684 1227 1.19 Mainstream smoke 172.38371 1105 1.24 Oriental tobacco 172.696283 1103 1.24 Phyllocyanin 266.027972 786 1.42 Nicotine 324.273764 663 1.51 Whitefly 349.937008 554 1.54

[0134] The higher the feature weight value is, the more important the keyword is to the tobacco science and technology literature field. For example, “flue-cured tobacco” and “tobacco leaf” are common words in the tobacco field, and although they appear frequently in documents, these words are relatively general and cover a wide range, and have little effect on retrieval, and cannot accurately match the deep knowledge that the user wants through the words.

[0135] In step 23, the Word2vec embedding model is used to train and learn the tobacco science and technology literature corpus to generate a tobacco literature word vector model and obtain the word vector form of each tobacco keyword.

[0136] In specific implementation, the training parameters of the Word2vec embedding model are as follows: the most similar word dimension topNSize=40, the context window size parameter Window=5, the random down-sampling configuration threshold for high-frequency words is 1e-3, the CBOW algorithm model is used to obtain the keyword vector, and the Softmax method is used for optimization, and the tobacco science and technology literature word vector model is generated.

[0137] In step 24, the tobacco science and technology literature keywords extracted in step 21 are expanded by synonym and near-synonym expansion to obtain a keyword expansion library.

[0138] In step 25, based on the existing authoritative science and technology thesaurus and tobacco terminology table, and combined with the tobacco field subject word library, a tobacco science and technology keyword dictionary is constructed according to the tobacco keywords in the keyword expansion library.

[0139] In step 26, the tobacco science and technology literature word vector model constructed in step 23 is used to calculate the semantic similarity of tobacco words on the tobacco science and technology literature corpus constructed in step 21 by using the cosine similarity calculation method to construct a tobacco word semantic similarity matching model.

[0140] In step 27, based on the tobacco word semantic similarity matching model obtained in step 26, the synonym and near-synonym expansion of each tobacco keyword in the tobacco keyword dictionary constructed in step 25 is performed to obtain a tobacco science and technology literature keyword expansion library.

[0141] In step 31, the user inputs a query statement, such as “how much nicotine is contained in tobacco”, in the unified retrieval input box, and the unified retrieval syntax converter constructed in step 11 is used for conversion to execute the unified retrieval syntax and realize logical operations such as and, or, and not, and retrieval priority logical operations.

[0142] Step 32, the query sentence converted by the unified search syntax converter in step 31 is segmented and stop words are removed, and words not representing concepts are filtered out, and finally the keywords obtained are "tobacco" and "nicotine", which are the basic keywords.

[0143] Step 33, for the basic keywords "tobacco" and "nicotine" segmented in step 32, combined with the tobacco technology literature keyword expansion library obtained in step 27, and using the tobacco literature word vector model obtained in step 23, the three expansion keywords "tobacco leaf", "cigarette", and "flue-cured tobacco" with the highest similarity to the basic keyword "tobacco" are calculated and obtained, and the three expansion keywords "nicotine", "smoke", and "nicotinic acid" with the highest similarity to the basic keyword "nicotine" are calculated and obtained.

[0144] Step 34, the word vectors of "tobacco" and "nicotine" and the expanded "tobacco leaf", "cigarette", "flue-cured tobacco", "nicotine", "smoke", and "nicotinic acid" obtained in step 33 are normalized and processed to obtain the word vector space model of each query keyword.

[0145] Step 35, the feature weight of each basic keyword and expansion keyword in the search formula "how much nicotine is contained in tobacco" is calculated , the calculation formula is as follows:

[0146] (5)

[0147] Where S represents the vector of the query sentence, represents the vector of the query keyword and the query expansion word, represents the order of the n most similar query expansion word items before the query keyword, i=0 represents the query keyword or the query synonym, represents the similarity of the query keyword or the query expansion word to the query sentence.

[0148] Step 35, the expansion keywords obtained in step 33 and the basic keywords obtained in step 32 are used as query keywords to initiate a query request to each tobacco technology literature resource library, and the self-owned literature resource search interface and the outsourced tobacco literature resource search interface identified as available in step 11 are called, and the search results are cached to the cache server.

[0149] Step 41, the search results obtained in step 35 are structured, formatted, content cleaned, and data normalized, and according to the structure characteristics of the tobacco technology literature data resources, combined with the tobacco technology literature data resource deduplication model constructed in step 12, the tobacco technology literature information fingerprint is extracted, and the deduplication and integration of the tobacco technology literature search results are realized.

[0150] Step 42, based on the tobacco science and technology literature quality evaluation model constructed in step 13, the quality of the search results processed in step 42 is evaluated, and low-quality search results are removed and high-quality search results are retained.

[0151] Step 51, on the basis of step 42, using the analytic hierarchy process, according to the position information of the query keywords in the search results, the correlation weight of different positions is determined.

[0152] The specific steps are: determine the respective weights of the query keywords matching the title, abstract and text of the tobacco science and technology literature, for example, the weight corresponding to the title is 0.8, the weight corresponding to the abstract is 0.5, and the weight corresponding to the text is 0.3, according to the number of query keywords appearing in the title, abstract and text, further can get the comprehensive weight value of the keywords matching in different positions , The calculation formula is as follows:

[0153] (6)

[0154] Where, i represents the similarity order of the extended keywords and the basic keywords, j is the position number of the keywords, wherein the number of the title is 1, the number of the abstract is 2, and the number of the text is 3, represents the basic keywords or extended keywords in the title, abstract and text, represents the weight of the basic keywords or extended keywords in each position.

[0155] Step 52, according to the search results, calculate the value of each query keyword;

[0156] Step 6, based on the comprehensive position weight value of each query keyword, value and field feature weight value, construct a tobacco science and technology literature search result evaluation and analysis model , by calculating the relevance of the search results and the query statement, according to the score calculated by the tobacco science and technology literature search result evaluation and analysis model, the search results are sorted.

[0157] Wherein, the calculation formula of the tobacco science and technology literature search result evaluation and analysis model is as follows:

[0158] (7)

[0159] Wherein, represents one of the query keywords after the query statement is segmented frequency of occurrence in a search result; representing query keywords inverse document frequency of all search results; representing query keywords feature weight in the field of tobacco science literature; representing query keywords feature weight in the search formula; representing query keywords comprehensive weight value of matching in different positions; comprehensive weight of all query keywords in the query statement.

[0160] Further, the list of the query statements can be outputted according to the matching degree of the query statements and the tobacco science literature, so that the user can check, and unified search and intelligent recommendation of massive multi-source heterogeneous tobacco science literature resources can be realized, and the user can quickly and accurately find the search content.

[0161] Embodiment 7

[0162] The embodiment also provides a computer storage medium, and the computer readable storage medium stores a computer program, and the program is executed by a processor to realize the search result intelligent recommendation method in any one of the embodiments 1-3.

[0163] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application rather than limit them; although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that: the specific embodiments of the present application can be modified or some technical features can be replaced by equivalent ones; without departing from the spirit of the technical solutions of the present application, they should be covered in the technical solution range of the present application claimed.

Claims

1. An intelligent recommendation method of search results, characterized by, The method comprises the following steps: S1, acquiring the keywords in the field of scientific and technological literature and calculating the field feature weight value of each keyword in the field of scientific and technological literature; S2, acquiring a query statement, determining the query keywords of the query statement, and acquiring a retrieval result from the scientific and technological literature resource library according to the query keywords; The specific steps of determining the query keywords of the query statement in S2 are as follows: The query statement is input into a retrieval box, and after word segmentation and stop word removal operations, the basic keywords of the query statement are obtained; The trained scientific and technological literature word vector model is used to calculate the similarity between each extended keyword and the basic keyword by using a cosine similarity calculation method in the pre-constructed scientific and technological literature keyword expansion library; The extended keywords are sorted in descending order of similarity, and a preset number of extended keywords with high similarity are selected, and the same / synonymous keyword merging is performed to obtain the expanded keywords; Each basic keyword and its corresponding expanded keyword are taken as the query keywords; After the expanded keywords are acquired, the word vectors of each basic keyword and its corresponding expanded keyword are normalized and regularized to obtain the word vector space model of each basic keyword and its corresponding expanded keyword; Based on the word vector space model of each basic keyword and its corresponding expanded keyword, the feature weight of each basic keyword and its corresponding expanded keyword in the query statement is calculated respectively: The specific steps of S1 are as follows: i After the stop word removal and word segmentation operations are performed on the scientific and technological literature corpus, the keywords of the scientific and technological literature corpus are extracted by using a TF-IDF algorithm. ( S , w i ) (0 ≤ i ≤ n) wherein, The specific steps of S1 are as follows: i represents the characteristic weight value, S represents the vector of the query statement, w i represents the vector of the basic keyword or the extended keyword, i represents the similarity order of the extended keyword and the basic keyword, when i=0, w 0 represents the basic keyword, 3. The retrieval result intelligent recommendation method according to claim 1 or 2, characterized in that, S , w i represents the similarity of the basic keyword or the extended keyword and the query statement;​ S3, calculating the comprehensive position weight value of each query keyword according to the search result, and The training steps of the scientific and technological literature word vector model are as follows: after the stop word removal and word segmentation operations are performed on the scientific and technological literature corpus, the scientific and technological literature corpus is trained and learned by using a Word2vec embedding model to generate a scientific and technological literature word vector model, and the word vector form of each keyword is obtained; value; S4, based on the comprehensive positional weight value of each query keyword, The construction steps of the scientific and technological literature keyword expansion library are as follows: The search results are ranked according to the relevance of the search results to the query statement, using the domain feature weights and domain feature weights. In the calculation of the relevance of the retrieval results and the query statement based on the comprehensive location weight value of each query keyword, After the stop word removal and word segmentation operations are performed on the scientific and technological literature corpus, the literature keywords of the scientific and technological literature corpus are extracted by using a TF-IDF algorithm, and the synonym and synonymous expansion is performed; and the field feature weight value, the retrieval result evaluation analysis model based on scientific literature is used to perform the calculation. wherein, Based on the literature keywords in the keyword expansion library, a scientific and technological keyword dictionary is constructed based on the scientific and technological thesaurus and the term table in combination with the field subject heading library; wi represents each query keyword after tokenization of the query statement w i frequency of occurrence in a search result; By using the trained scientific and technological literature word vector model, a keyword semantic similarity matching model is constructed by using a cosine similarity calculation method in the scientific and technological literature corpus; wi represents a query keyword w i inverse document frequency of all search results; f wi represents a query keyword w i field characteristic weight value in the field of scientific literature, Based on the keyword semantic similarity matching model, the synonym and synonymous expansion of each keyword in the scientific and technological keyword dictionary is performed to obtain the scientific and technological literature keyword expansion library. wi is a query keyword w i comprehensive position weight value.

2. The intelligent recommendation method for search results according to claim 1, characterized in that, The method comprises the following steps: The keyword feature weight acquisition unit is configured to acquire the keywords in the field of scientific and technological literature and calculate the field feature weight value of each keyword in the field of scientific and technological literature. Calculate the field feature weight of each keyword: f wl =| lg( ​ wl )-1 | ; wherein, f wl is a keyword, ​ wl denotes a keyword w l inverse document frequency in a corpus of scientific literature, lg( ) denotes the logarithm function. ​ ​ ​ ​ ​ ​ ​ 4. An intelligent recommendation of search results apparatus, characterized by comprising: ​ ​ The query unit is configured to obtain a query statement, determine a query keyword of the query statement, and obtain a search result from the scientific literature resource library according to the query keyword. The specific steps of determining the query keyword of the query statement are as follows: The query statement is input into the search box, and after word segmentation and stop word removal operations, the basic keywords of the query statement are obtained. The trained scientific literature word vector model is used to calculate the similarity between each extended keyword and the basic keyword in the pre-constructed scientific literature keyword expansion library by using the cosine similarity calculation method. The extended keywords are sorted in descending order of similarity, and a preset number of extended keywords with high similarity are selected. Each basic keyword and its corresponding extended keyword are taken as a query keyword. After obtaining the extended keywords, the word vectors of each basic keyword and its corresponding extended keyword are normalized to obtain the word vector space model of each basic keyword and its corresponding extended keyword. Based on the word vector space model of each basic keyword and its corresponding extended keyword, the feature weight of each basic keyword and its corresponding extended keyword in the query statement is calculated. The comprehensive position weight value acquisition unit is configured to calculate the comprehensive position weight value of each query keyword according to the search result. i TF-IDF ( S , w i ) (0 ≤ i ≤ n) wherein, TF-IDF i represents the characteristic weight value, S represents the vector of the query statement, w i represents the vector of the basic keyword or the extended keyword, i represents the similarity order of the extended keyword and the basic keyword, when i=0, w 0 represents the basic keyword, TF-IDF S , w i represents the similarity of the basic keyword or the extended keyword and the query statement;​ TF-IDF TF-IDF a computing unit for computing a value of each query keyword in the search result The method comprises the following steps: value; The sorting unit is used to sort based on the comprehensive positional weight value of each query keyword. Step 1: According to the syntax characteristics and logical operation methods of various available tobacco scientific literature resource libraries, a unified search syntax converter is constructed. The search results are ranked according to the relevance of the search results to the query statement, using the domain feature weights and domain feature weights. In the calculation of the relevance of the retrieval results and the query statement based on the comprehensive location weight value of each query keyword, According to the resource type of the tobacco scientific literature resources, a tobacco scientific literature data resource deduplication model is constructed; according to the field richness of the tobacco scientific literature and the content richness of the literature, a tobacco scientific literature quality evaluation model is constructed. and the field feature weight value, the retrieval results evaluation analysis model based on scientific literature is used to perform the calculation. wherein, Step 2: Obtain the keywords in the scientific literature field, and calculate the field feature weight of each keyword in the scientific literature field. wi denotes each query keyword after tokenization of the query statement w i frequency of occurrence in a search result; Step 3: Input the query statement into the unified search box, and convert it into a preset unified syntax by using the unified search syntax converter to determine the query keyword of the query statement; according to the query keyword, search results are obtained from each tobacco scientific literature resource library; the specific steps of determining the query keyword of the query statement are as follows: wi denotes query keyword w i inverse document frequency over all search results; f wi denotes query keyword w i field characteristic weight value in the field of scientific literature, The query statement is input into the search box, and after word segmentation and stop word removal operations, the basic keywords of the query statement are obtained. wi is query keyword w i comprehensive position weight value.

5. A unified search method for multi-source heterogeneous tobacco science literature resources, characterized in that, The trained scientific literature word vector model is used to calculate the similarity between each extended keyword and the basic keyword in the pre-constructed scientific literature keyword expansion library by using the cosine similarity calculation method. The extended keywords are sorted in descending order of similarity, and a preset number of extended keywords with high similarity are selected. Each basic keyword and its corresponding extended keyword are taken as a query keyword. After obtaining the extended keywords, the word vectors of each basic keyword and its corresponding extended keyword are normalized to obtain the word vector space model of each basic keyword and its corresponding extended keyword. ​ ​ ​ ​ ​ ​ Based on the word vector space model of each basic keyword and its corresponding extended keyword, the feature weight of each basic keyword and its corresponding extended keyword in the query statement is calculated respectively: μ i =Sim ( S , w i ) (0 ≤ i ≤ n) wherein, μ i represents a feature weight value, S represents a vector of a query statement, w i represents a vector of a basic keyword or an extended keyword, i represents a similarity order of the extended keyword to the basic keyword, when i=0, w 0 represents a basic keyword, Sim S , w i represents a similarity of the basic keyword or the extended keyword to the query statement;​ Step 4, the structure of the retrieval result is formatted, the content is cleaned, the data is normalized, and based on the constructed tobacco science and technology literature data resource deduplication model, the tobacco science and technology literature information fingerprint is extracted for deduplication and integration operation; and based on the tobacco science and technology literature quality evaluation model, the quality of the retrieval result is evaluated, the low-quality retrieval result is removed, and the high-quality retrieval result is retained; Step 5, calculating the integrated position weight value of each query keyword according to the search result and TF-IDF value; Step 6, calculating the relevancy of the search result to the query sentence based on the integrated position weight value of each query keyword, the value and the field characteristic weight value, and sorting the search result according to the relevancy size; TF-IDF value and the field characteristic weight value, and sorting the search result according to the relevancy size; In the calculation of the relevance of the retrieval results and the query statement based on the comprehensive location weight value of each query keyword, TF-IDF the value and the field feature weight value, the retrieval result evaluation analysis model based on scientific literature is used to perform the calculation; wherein, TF wi represents each query keyword after tokenization of the query statement w i frequency of occurrence in a search result; IDF wi represents a query keyword w i inverse document frequency in all search results; f wi represents a query keyword w i field characteristic weight value in the field of scientific literature, φ wi is a query keyword w i comprehensive position weight value.

6. The method according to claim 5, wherein, The training steps of the tobacco science and technology literature word vector model are as follows: constructing a tobacco science and technology literature corpus, after removing stop words and performing word segmentation on the tobacco science and technology literature corpus, using the Word2vec embedding model to train and learn the tobacco science and technology literature corpus, generating a tobacco science and technology literature word vector model, and obtaining the word vector form of each keyword; The construction steps of the tobacco science and technology literature keyword expansion library are as follows: After constructing the tobacco scientific literature corpus, removing stop words and segmenting the tobacco scientific literature corpus, the tobacco keywords of the tobacco scientific literature corpus are extracted by using the algorithm, and the keyword expansion library is obtained through synonym and near-synonym expansion. TF-IDF After constructing the tobacco scientific literature corpus, removing stop words and segmenting the tobacco scientific literature corpus, the tobacco keywords of the tobacco scientific literature corpus are extracted by using the algorithm, and the keyword expansion library is obtained through synonym and near-synonym expansion. According to the tobacco keywords in the keyword expansion library, on the basis of the scientific and technological thesaurus and the tobacco terminology table, combined with the tobacco field subject heading library, a tobacco science and technology keyword dictionary is constructed; Using the trained tobacco science and technology literature word vector model, the cosine similarity calculation method is used to calculate in the tobacco science and technology literature corpus, and a tobacco vocabulary semantic similarity matching model is constructed; Based on the tobacco vocabulary semantic similarity matching model, the synonyms and near-synonyms of each keyword in the science and technology keyword dictionary are expanded, and a tobacco science and technology literature keyword expansion library is obtained.

7. A computer storage medium, characterized in that: The computer readable storage medium stores a computer program, which is executed by the processor to implement the retrieval result intelligent recommendation method of any one of claims 1-3.

Citation Information

Patent Citations

  • Method for indexing and acquiring semantic net information

    CN101030217A

  • Biomedicine literature retrieval method based on ranking learning algorithm

    CN108520038A