A Patent Recommendation Method Based on User Requirements and Combined with an Inverted Index

The method constructs an inverted index table based on user needs to provide both precise and broad patent recommendations, addressing speed and precision issues in existing systems by using a refined similarity word mechanism and text compression, enhancing user satisfaction and platform engagement.

CN116932736BActive Publication Date: 2025-07-15QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310882424.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-18
Publication Date
2025-07-15
Estimated Expiration
2043-07-18

AI Technical Summary

Technical Problem

The existing patent recommendation technology has problems such as slow recommendation speed, insufficient accurate and extensive fields, and it is difficult to meet the needs of users of varying degrees.

Method used

By constructing an inverted table and combining user needs into accurate requirements and extensive requirements, patent recommendations are used using improved similar word mechanisms and GPT2 requirements compression model to generate accurate and extensive recommendation lists respectively.

Benefits of technology

It realizes fast and accurate patent recommendations, which can meet users' in-depth understanding in specific fields and cross-domain new technical references, and improves the accuracy of user experience and recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116932736B_ABST
    Figure CN116932736B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer information for data recommendation, and provides a patent recommendation method based on user requirements combined with an inverted index table, including constructing an initial inverted index table for a patent dataset according to user requirements and adding a similar word mechanism to form a final inverted index table; the inverted index table includes: word number, word, and patent number list; numbering the patent information in the patent dataset to form a document list, and obtaining sentence vector representations for the patent information of each patent in the document list using the bert model; the document list includes: patent number, patent information, patent information sentence vector representation; dividing the user requirement information into precise requirements and broad requirements for dual-track recommendation. The present invention solves the problems in the prior art that due to the fact that patent recommendation involves patents in various fields and the quantity is huge, using patent information in a single field for recommendation results in poor recommendation effects and inaccurate patent recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer information of data recommendation, and more specifically, relates to a patent recommendation method based on user requirements combined with an inverted index table. Background Art

[0002] With the development of society and technology, intellectual property is increasingly valued in China. Intellectual property is a key part of the core competitiveness of enterprises and countries. It represents the core competitiveness of enterprises and the comprehensive national strength of countries. Patents are crucial for the protection of the core technologies of enterprises and countries, and are also crucial for the survival and competition strategies of enterprises. Recommending patents related to user information and user requirements to users on the platform can, on the one hand, improve users' interest in the website, increase users' reading volume of patents and the duration of their stay on the platform, which is conducive to users understanding the functions of the platform and discovering its advantages, attracting more users to register, and playing a positive role in the development of the platform; on the other hand, patent recommendation can improve users' work efficiency, and the automatic recommendation technology can provide accurate and extensive patent recommendations for users based on their basic information and demand information. According to investigations, without patent recommendation technology, the customer traffic and customer loyalty on the platform will both decline.

[0003] Chinese Patent Document CN107943910A discloses a personalized book recommendation method based on a combination algorithm, including the following steps: extracting keywords from the content information of books to obtain the feature vectors of books; obtaining the rating value of a user for a certain new book; generating an "inverted index table and query index of user behavior" and an "inverted index table and query index of book behavior"; generating a similarity file and query indexes: similar user indexes and query indexes, similar book indexes and query indexes; calculating the book recommendation index for the user according to the similar user indexes and query indexes, and similar book indexes and query indexes.

[0004] Most of the existing data recommendation methods generate recommendation sets based on data feature vectors, and all have certain limitations, making it difficult to quickly and accurately recommend patents related to users to users. For the existing patent recommendation technology, in the original recommendation process, the original patent information, including the name and abstract of the patent, etc., is input. The patent portrait is constructed for each patent in the system using word segmentation technology and keyword technology, the user portrait is constructed using the list of patents collected by the user and the list of search keywords, the neural network model is used to convert all the patent portraits and the user portrait of the user into sentence vector representations respectively, the similarity between the two is calculated, and the recommendation list is output to the user in the order of similarity. Since patent recommendation involves patents in various fields and the quantity is huge, in order to improve the feasibility of recommendation, using patent information in a single field for recommendation will result in slow recommendation speed, inaccurate and insufficiently extensive recommendation fields. Summary of the Invention

[0005] The present invention aims to overcome at least one defect of the above-mentioned prior art, and provides a patent recommendation method based on user needs combined with an inverted index table.

[0006] The detailed technical solution of the present invention is as follows:

[0007] In order to solve the above technical problems, the present invention provides a patent recommendation method based on user needs combined with an inverted index table to solve problems such as slow recommendation speed, inaccurate and narrow recommendation fields in the prior art.

[0008] First, construct an inverted index table and a document list for the patent dataset, and divide the user demand information into precise demand and broad demand; secondly, compress and preprocess the precise demand and then combine it with the inverted index table and the document list to give precise recommendation results; furthermore, segment the broad demand, and then combine each segment with the inverted index table and the document list to give recommendation results, and summarize the recommendation results of each segment to form the final broad recommendation result, specifically as follows:

[0009] A patent recommendation method based on user needs combined with an inverted index table, characterized by including the following steps:

[0010] S1. Construct an initial inverted index table for the patent dataset according to user needs and add a synonym mechanism to form a final inverted index table;

[0011] The inverted index table includes: word number, word, and patent number list;

[0012] S2. Number the patent information in the patent dataset to form a document list, and use the bert model to obtain sentence vector representations for the patent information of each patent in the document list;

[0013] The document list includes: patent number, patent information, and sentence vector representation of patent information;

[0014] S3. Divide the user demand information into precise demand and broad demand according to the user needs, and perform dual-track recommendation, where the dual-track recommendation includes precise recommendation and broad recommendation;

[0015] The precise recommendation is to compress and preprocess the user needs and then generate a precise candidate set in combination with the inverted index table, then look up the document list from the patent numbers in the precise candidate set to obtain the patent information corresponding to each patent number, and finally match the user needs with each patent information to obtain a precise recommendation list;

[0016] The above-mentioned extensive recommendation segments the user requirements. After preprocessing each segment, it combines with an inverted index to generate a corresponding candidate set. Then, it looks up the document list according to the patent numbers in the corresponding candidate set to obtain the patent information corresponding to each patent number. Finally, it matches the user requirements of this segment with each patent information to obtain the recommendation list for this segment. Finally, the recommendation lists of each paragraph are merged to form an extensive recommendation list.

[0017] By dividing the user requirement information into precise requirements and extensive requirements, it can meet the user's needs at different levels. Precise recommendation can focus on patent recommendation in the user's technical field and several similar fields to meet the user's requirements for specific technologies. While extensive requirements can achieve cross-field recommendation, which can provide patents in different fields that may be helpful to the user, and can provide reference and ideas for the user's new technologies.

[0018] The specific steps of S1 include:

[0019] S11: Segment the data of the user requirement part in the patent dataset to obtain words. Number the words, create an index with the words, and then record the numbers corresponding to all the patents containing the words to form an initial inverted index.

[0020] The user requirement part is composed of the abstract part of the patent specification, the claim part, and the beneficial effect part of the specification, which contains multiple channels of adaptable words through technical points and effect points, and can recommend adaptable patents for the user in many aspects.

[0021] S12: Add an improved similar word mechanism to the words, combine with a pre-trained Chinese word vector file to construct the similarity relationship of the words, and form a final inverted index.

[0022] The specific improved similar word mechanism is as follows: Traverse each word in the initial inverted index, and combine with the pre-trained Chinese word vector file (the pre-trained Chinese word vector file sgns.zhihu.word downloaded from an external known channel, preferably https: / / github.com / Embedding / Chinese-Word-Vectors) to obtain the top d words with high similarity. Then use the method of comprehensive similarity sorting to select the top c' similar words from them. Traverse these c' similar words. If the inverted index contains the similar word of this word, add the patent numbers of this similar word to the patent number list of the inverted index of this word. After the above traversal process, a final inverted index is formed.

[0023] The method of comprehensive similarity ranking refers to selecting other words that have duplicate characters with the word and calculating the duplication degree p, combining the similarity h between the other word and the single word, calculating the comprehensive similarity f of the word, where γ is an adjustable parameter, that is, a suitable value is adjusted according to the similarity and duplication degree during debugging. The purpose of adjustment is to make the result output by the comprehensive similarity calculation more literally conform to the required text input by the user:

[0024] f = (1 - γ)p + γh, where γ ∈ (0, 1) (1).

[0025] The precise recommendation specifically includes:

[0026] S311. Use an improved GPT2-based demand compression model to compress the user demand information;

[0027] S312. Preprocess the compressed user demand information, including word segmentation, stop word removal, and special stop word removal operations; for example, after the operation of "controllable benchmarked text generation machine learning model", it becomes "controllable benchmark text generation machine learning model", and use the preprocessed user demand information to search the inverted list and generate a precise candidate set;

[0028] S313. Use the bert model to obtain the sentence vector representation of the compressed user demand information;

[0029] S314. Search the document list by the patent number in the precise candidate set to obtain the patent information sentence vector corresponding to each patent number;

[0030] S315. Calculate the cosine similarity between the sentence vector of the compressed user demand information and the patent information sentence vector of each patent in the precise candidate set, and select the top n patent information with the highest similarity as the precise recommendation result according to the cosine similarity calculation result.

[0031] The improved GPT2-based demand compression model includes:

[0032] On the basis of the original GPT2 model, an encoder is connected in parallel. The encoder includes a word probability distribution and a multi-head attention mechanism. The original GPT2 model includes m layers of decoders;

[0033] After the input data, the data flows to both the GPT2 model and the encoder at the same time. Use the word probability distribution on the encoder and the decoder state inside the GPT2 model to calculate the weight G, and then use the weight G to calculate the word probability distribution at the moment, and finally output the predicted word with the highest probability value.

[0034] Before the improved GPT2-based requirement compression model generates the final predicted word probabilities, the multi-head attention mechanism in the encoder is used to extract the word probabilities of the original input, optimize the predicted words of the final output, solve the out-of-vocabulary problem, and make the predicted words more in line with the original text semantics; an encoder is added, and the multi-head attention mechanism of the encoder can obtain the word probabilities in the source text, and combining the word probability distribution on the encoder can better predict the words without deviating from the original text semantics, making the recommendation results more in line with user needs;

[0035] Use the pre-processed user requirement information as the input data of the improved GPT2-based requirement compression model, and generate the decoder state s of each layer through the 10-layer decoder of the improved GPT2-based requirement compression model i , the attention distribution generated by the data passing through the encoder is used as the probability distribution of the words on the source text and denoted as a; use the decoder state s of each layer i and the word probability distribution a on the source text generated by the encoder to calculate the weight G:

[0036]

[0037] where G ∈ [0, 1], Sigmoid is the activation function, W and b are adjustable parameters, and S1 - S 10 is the decoder state value of each layer in the improved GPT2-based requirement compression model at time t i ;

[0038]

[0039] where P(w) is the final distribution of the word w predicted by the improved GPT2-based requirement compression model at time t in the vocabulary, and the vocabulary is generated during pre-training; if w is a word outside the vocabulary, then P(w) = 0, represents the attention distribution of the word w at time t on the source text; if the word w does not appear, then T(w) refers to the final distribution of the word w predicted by the improved GPT2-based requirement compression model at time t in the vocabulary and the source text;

[0040] The compressed text of the user requirement information is updated over time t until all user requirement information is compressed.

[0041] The so-called extensive recommendation specifically is:

[0042] S321. Perform requirement segmentation processing on the user requirement information, search the inverted index for each processed segment of requirement information, and obtain the candidate set corresponding to each segment of requirement information;

[0043] S322. Obtain the sentence vector representation of each segmented requirement information using the BERT model;

[0044] S323. Look up the document list by the patent numbers in the candidate set to obtain the patent information sentence vectors corresponding to each patent number in the candidate set. Then, calculate the similarity between the sentence vector of each segmented requirement information and each patent information sentence vector in its candidate set, and obtain the top k recommended results for each segmented requirement information according to the similarity calculation results; Merge the top k recommended results of different paragraphs into a wide recommendation list.

[0045] The requirement segmentation specifically includes:

[0046] 1) Obtain the user requirement information and divide it into z segments according to two symbols: semicolon and period;

[0047] 2) Preprocess the segmented requirement information, and then extract keywords from each segmented requirement information through the TF-IDF keyword extraction mechanism; The preprocessing includes word segmentation, stop word removal, punctuation removal, etc.;

[0048] 3) Look up the inverted index according to the keywords in the first paragraph to generate candidate set 1;

[0049] 4) Obtain the sentence vector representation of the first segmented requirement information using the BERT model;

[0050] 5) Look up the document list by the patent numbers in candidate set 1 to obtain the patent information sentence vectors corresponding to each patent number in the candidate set. Calculate the cosine similarity between the sentence vector of the first segmented requirement information and each patent information sentence vector, and select the top j patent information with the highest similarity to generate recommendation list 1;

[0051] 6) Repeat steps 2)-5) for the remaining z - 1 segments of requirement information respectively, and merge all the generated recommendation lists to finally form a wide recommendation list.

[0052] Compared with the prior art, the beneficial effects of the present invention are:

[0053] (1) A recommendation method based on user requirements using text compression combined with an inverted index provided by the present invention can quickly achieve wide recommendation and accurate recommendation for users. Among them, wide recommendation can achieve cross-domain recommendation and give users more ideas for writing patents; Accurate recommendation can recommend patents with high similarity in the same field to users according to the overall user requirements, so that users can deeply understand the relevant information in this field.

[0054] (2) The patent recommendation method based on user requirements combined with an inverted index provided by the present invention uses an improved similar word mechanism for the words in the inverted index to construct a similarity relationship. This method not only makes the recommended results more flexible, but also fully considers the user's intuitive acceptance of the recommended results, and preferentially recommends patents containing the user's required words in the results, making the user more satisfied with the recommended results and having a better experience on this platform.

[0055] (3) The patent recommendation method based on user requirements combined with an inverted index provided by the present invention uses an improved GPT2-based requirement compression model to compress the requirement text. This method has the ability to generate out-of-vocabulary words, making each predicted word more accurate, optimizing the predicted words in the final output to solve the problem of out-of-vocabulary overflow, making the predicted words more in line with the original text semantics, improving the accuracy of the model, and also making the results of the precise recommendation process more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a schematic diagram of the method flow described in the present invention.

[0057] Figure 2 is a detailed processing scheme flowchart of the method described in the present invention.

[0058] Figure 3 is a schematic diagram of the inverted index file generated in Embodiment 1 of the present invention.

[0059] Figure 4 is the improved GPT2-based requirement compression model in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The following further describes the present disclosure in conjunction with the accompanying drawings and embodiments.

[0061] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0062] In the case of no conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0063] Embodiment 1

[0064] This embodiment provides a patent recommendation method based on user requirements combined with an inverted index, as Figure 1 、 Figure 2 shown;

[0065] S1. Construct an initial inverted index for the patent dataset according to user requirements and add a similar word mechanism to form a final inverted index.

[0066] Specifically, the S1 step specifically includes:

[0067] S11. Segment the data of the user requirement part in the patent dataset to obtain words; number the words, create an index with the words, and then record the numbers corresponding to all the patents containing the words to form an initial inverted list;

[0068] The user requirement part is composed of the abstract part of the patent specification, the claim part, and the beneficial effect part of the specification, and contains multiple paths of adaptable words through technical key points and effect key points, which can recommend adaptable patents to users in multiple aspects;

[0069] S12. Add an improved similar word mechanism to the words, combine the pre-trained Chinese word vector file to construct the similarity relationship of the words, and form the final inverted list. As Figure 3 shown, the inverted list specifically includes: word number, word, and patent number list;

[0070] The improved similar word mechanism is specifically: traverse each word in the initial inverted list, and combine the use of the pre-trained Chinese word vector file (the pre-trained Chinese word vector file sgns.zhihu.word downloaded from an external known channel, preferably https: / / github.com / Embedding / Chinese-Word-Vectors) to obtain the top d words with high similarity, and then use the method of comprehensive similarity sorting to select the top c' similar words from them. Traverse these c' similar words. If the inverted list contains the similar word of this word, add the patent number of this similar word to the patent number list of the inverted list of this word. After the above traversal process, the final inverted list is formed; preferably, the words can be de-duplicated finally, that is, delete multiple words that are similar words to each other;

[0071] The method of comprehensive similarity sorting refers to selecting other words with duplicate characters with the word and calculating the duplication degree p, combining the similarity h between the other word and the word, and calculating the comprehensive similarity f of the word. γ is an adjustable parameter, that is, a suitable value is adjusted according to the similarity and duplication degree during debugging. The purpose of adjustment is to make the result output by the comprehensive similarity calculation more conform to the required text input by the user literally:

[0072] f=(1 - γ)p+γh, where γ∈(0,1) (1).

[0073] S2. Number the patent information in the patent dataset to form a document list, and use the bert model to obtain the sentence vector representation for the patent information of each patent in the document list;

[0074] The document list includes: id (patent number), patent information, and sentence vector representation of patent information; the patent numbers in the inverted index are the ids of the document list; after obtaining the candidate set, patent information is retrieved from the patent list according to the patent numbers in the candidate set; the patent information refers to the title and abstract of each patent.

[0075] S3. Classify into precise recommendation and broad recommendation according to user demand information, and perform dual-track recommendation.

[0076] The precise recommendation is to generate a precise candidate set by compressing and preprocessing the user demand and then combining it with the inverted index, retrieve the patent information corresponding to each patent number from the document list according to the patent numbers in the precise candidate set, and finally match the user demand with each patent information to obtain a precise recommendation list.

[0077] The broad recommendation is to segment the user demand, generate a candidate set by preprocessing each segment and then combining it with the inverted index, retrieve the patent information corresponding to each patent number from the document list according to the patent numbers in the candidate set, and finally match the user demand of this segment with each patent information to obtain the recommendation list of this segment, and finally merge the recommendation lists of each paragraph to form a broad recommendation list.

[0078] Generating a candidate set by combining the user demand with the inverted index means traversing the inverted index to find all the patent numbers corresponding to each word in the preprocessed user demand in the inverted index, and forming a candidate set with the patent numbers corresponding to each word in the user demand. Each word corresponds to one or more patents.

[0079] The specific steps of the precise recommendation are as follows:

[0080] S311. Compress the user demand information using an improved GPT2-based demand compression model.

[0081] Perform preprocessing on the compressed user demand information, including word segmentation, stop word removal, and special stop word removal operations; for example, after processing "controllable text generation machine learning model with a benchmark", it becomes "controllable text generation machine learning model". Use the preprocessed user demand information to search the inverted index and generate a precise candidate set.

[0082] Use the bert model to obtain the sentence vector representation of the compressed user demand information.

[0083] Retrieve the sentence vectors of the patent information corresponding to each patent number from the document list according to the patent numbers in the precise candidate set.

[0084] S315. Calculate the cosine similarity between the sentence vectors of the compressed user requirement information and the sentence vectors of the patent information of each patent in the precise candidate set, and select the top 5 patent information with the highest similarity as the precise recommendation result according to the cosine similarity calculation result.

[0085] The improved GPT2-based requirement compression model includes:

[0086] As Figure 4 shown, a parallel encoder is added to the original GPT2 model. The encoder includes a word probability distribution and a multi-head attention mechanism. The original GPT2 model includes m layers of decoders; preferably, the original GPT2 model includes m10 layers of decoders;

[0087] After the input data, the data flows to both the GPT2 model and the encoder simultaneously. Use the word probability distribution on the encoder and the decoder state inside the GPT2 model to calculate the weight G, and then use the weight G to calculate the word probability distribution at the moment, and finally output the predicted word with the highest probability value;

[0088] Before the improved GPT2-based requirement compression model generates the final predicted word probability, use the multi-head attention mechanism in the encoder to extract the word probability of the original input, optimize the finally output predicted word to solve the out-of-vocabulary problem, and make the predicted word more in line with the original text semantics;

[0089] Use the pre-processed user requirement information as the input data of the improved GPT2-based requirement compression model, and generate the decoder state s of each layer through the 10 layers of decoders of the improved GPT2-based requirement compression model i , and the attention distribution generated by the data passing through the encoder can be used as the probability distribution of the words on the source text and is denoted as a; use the decoder state s of each layer i and the probability distribution a of the words on the source text generated by the encoder to calculate the weight G:

[0090]

[0091] where G ∈ [0, 1], Sigmoid is the activation function, W and b are adjustable parameters, S1 - S 10 is the decoder state value of each layer in the improved GPT2-based requirement compression model at time t i , T is the transpose calculation symbol, which is a commonly used mathematical notation;

[0092]

[0093] Where P(w) is the final distribution of word w predicted by the improved GPT2-based requirement compression model at time t in the vocabulary, and the vocabulary is generated during pre-training; if w is a word outside the vocabulary, then P(w)=0. represents the attention distribution of word w at time t in the source text; if word w does not appear in the source text, then T(w) refers to the final distribution of word w predicted by the improved GPT2-based requirement compression model at time t in the vocabulary and the source text.

[0094] The compressed text of the user requirement information is updated over time t until all the text is compressed.

[0095] The steps of requirement compression of the improved GPT2-based requirement compression model are as follows:

[0096] (1) Input user requirement information: "A controllable benchmark response generation framework includes a machine learning model, a benchmark interface, and a control interface. The machine learning model is trained to output computer-generated text based on the input text. The benchmark interface can be used by the machine learning model to access a benchmark source including information related to the input text. The control interface can be used by the machine learning model to identify control signals. The machine learning model is configured to include information from the benchmark source in the computer-generated text and focus the computer-generated text based on the control signal."

[0097] (2) Invoke the model;

[0098] (3) Generate compressed text: "Controllable text generation with benchmark machine learning model".

[0099] The extensive recommendation is specifically as follows:

[0100] S321. Perform requirement segmentation on the user requirement information, search the inverted index for each processed segment of requirement information, and obtain the candidate set corresponding to each segment of requirement information;

[0101] S322. Use the bert model to obtain the sentence vector representation of each segment of requirement information after segmentation

[0102] S323. Search the document list by the patent numbers in the candidate set to obtain the patent information sentence vectors corresponding to each patent number in the candidate set. Then calculate the cosine similarity between the sentence vector of each segment of requirement information and each patent information sentence vector, and obtain the top 5 recommendation results for each segment according to the similarity calculation results; merge the top 5 recommendation results of different segments into an extensive recommendation list.

[0103] The requirement segmentation specifically includes:

[0104] 1) Obtain user requirement information and divide it into z segments according to two symbols, semicolons and full stops. Preferably, divide it into 5 segments according to semicolons and full stops;

[0105] 2) Preprocess the segmented requirement information, and then extract keywords from each segment of requirement information through the TF-IDF keyword extraction mechanism; For example:

[0106] "A controllable reference response generation framework includes a machine learning model, a reference interface, and a control interface. The machine learning model is trained to output computer-generated text based on the input text. The reference interface can be used by the machine learning model to access a reference source including information related to the input text. The control interface is used by the machine learning model to identify control signals. The machine learning model is configured to include information from the reference source in the computer-generated text and focus the computer-generated text based on the control signal." There are five segments of user requirement information in total; First, the preprocessed user requirements after segmentation:

[0107] ['Controllable reference response generation framework, machine learning model, reference interface, control interface',

[0108] 'Machine learning model, trained on input text, outputs computer-generated text',

[0109] 'Reference interface, used by machine learning model to access input text information, reference source',

[0110] 'Control interface, used by machine learning model to identify control signals',

[0111] 'Machine learning model configured with information from reference source, computer-generated text, control signal focuses on computer-generated text'], There are five segments of segmented data in total.

[0112] Then, the data after using TF-IDF keyword extraction is:

[0113] ['Controllable reference response framework, interface',

[0114] 'Trained on input text, outputs computer-generated',

[0115] 'Reference interface, accesses input information, reference source',

[0116] 'Control interface, identifies signal',

[0117] 'Configured with information from reference source, computer focuses'], There are five segments of keyword-extracted data in total.

[0118] Using TF-IDF keyword extraction solves the problem of repeated words in a passage on the one hand, and weakens the impact of repeatedly appearing words in each sentence on the retrieval results on the other hand. For example, the two words "machine learning" and "model" appear repeatedly in the above 5 passages. If both of these repeated words are placed in the keywords of each paragraph for retrieval, the retrieval results will all tend to patents of the machine learning model type, which goes against the original intention of making extensive recommendations after segmentation.

[0119] 3) Search the inverted list according to the keywords in the first paragraph to generate candidate set 1;

[0120] 4) Use the bert model to obtain the sentence vector representation of the demand information in the first paragraph after segmentation;

[0121] 5) Search the document list by the patent numbers in candidate set 1 to obtain the patent information sentence vectors corresponding to each patent number in the candidate set. Calculate the cosine similarity between the sentence vector of the demand information in the first paragraph and each patent information sentence vector, and select the top 5 patent information with the highest similarity to generate recommendation list 1;

[0122] 6) Repeat steps 2)-5) for the remaining 4 paragraphs of demand information respectively, and merge all the generated recommendation lists to finally form an extensive recommendation list.

[0123] Taking the user demand information in this embodiment as an example, the recommendation results are as follows in the table:

[0124]

[0125]

[0126]

[0127] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A patent recommendation method based on user requirements combined with an inverted index, characterized in that, including S1. Construct an initial inverted index table for the patent dataset according to user requirements and add a similar word mechanism to form the final inverted index table; The inverted index table includes: word number, word, and patent number list; S2. Number the patent information in the patent dataset to form a document list, and use the bert model to obtain the sentence vector representation for the patent information of each patent in the document list; The document list includes: patent number, patent information, and sentence vector representation of patent information; S3. Divide the user requirement information into precise requirements and broad requirements, and perform dual-track recommendation. The dual-track recommendation includes precise recommendation and broad recommendation; The precise recommendation is to compress and preprocess the user requirements, combine with the inverted index table to generate a precise candidate set, then find the patent information corresponding to each patent number in the precise candidate set from the document list, and finally match the user requirements with each patent information to obtain a precise recommendation list; The broad recommendation is to segment the user requirements, preprocess each segment, combine with the inverted index table to generate a corresponding candidate set, then find the patent information corresponding to each patent number in the corresponding candidate set from the document list, and finally match the user requirements of this segment with each patent information to obtain the recommendation list of this segment. Finally, merge the recommendation lists of each paragraph to form a broad recommendation list; The specific steps of S1 include: S11. Segment the data of the user requirement part in the patent dataset to obtain words; number the words, create an index with the words, and then record the numbers corresponding to all the patents containing the words to form an initial inverted index table; S12. Add an improved similar word mechanism to the words, combine with the pre-trained Chinese word vector file to construct the similarity relationship of the words, and form the final inverted index table; The improved similar word mechanism is: Traverse each word in the initial inverted index table, combine it with the pre-trained Chinese word vector file to obtain the top d words with high similarity, and then use the method of comprehensive similarity ranking to select the top c , similar words. Traverse these c , similar words. If the inverted index table contains similar words of this word, add the patent numbers of the similar words to the patent number list of the inverted index table of this word. After the above traversal process, the final inverted index table is formed; The method of comprehensive similarity ranking refers to selecting other words that have repeated characters with the word and calculating the repetition degree p, combining the similarity h between the other word and the single word, and calculating the comprehensive similarity of the single word , is an adjustable parameter: , among which (1).

2. The patent recommendation method based on user requirements combined with an inverted index according to claim 1, characterized in that The specific steps of the precise recommendation include: S311. Use an improved GPT2-based requirement compression model to compress the user requirement information; S312. Preprocess the compressed user requirement information, including word segmentation, stop word removal, and special stop word removal operations, and use the preprocessed user requirement information to search the inverted index table and generate a precise candidate set; S313. Use the bert model to obtain the sentence vector representation of the compressed user requirement information; S314. Find the sentence vector of the patent information corresponding to each patent number from the document list according to the patent numbers in the precise candidate set; S315. Calculate the cosine similarity between the sentence vector of the compressed user requirement information and the sentence vector of the patent information of each patent in the precise candidate set, and select the top n patent information with the highest similarity as the precise recommendation result according to the cosine similarity calculation result.

3. According to the patent recommendation method based on user requirements combined with an inverted index table described in claim 2, characterized in that The improved GPT2-based requirement compression model includes: On the basis of the original GPT2 model, an encoder is connected in parallel. The encoder includes a word probability distribution and a multi-head attention mechanism. The original GPT2 model includes m layers of decoders; After the input data, the data flows to the GPT2 model and the encoder simultaneously. The weight G is calculated using the word probability distribution on the encoder and the decoder state inside the GPT2 model. Then, the word probability distribution at the moment is calculated using the weight G, and finally, the predicted word with the highest probability value is output.

4. A patent recommendation method based on user needs combined with an inverted index according to claim 3, characterized in that Before the improved GPT2-based requirement compression model generates the final predicted word probability, the multi-head attention mechanism in the encoder is used to extract the word probability of the original input and optimize the predicted word of the final output. The pre - processed user requirement information is used as the input data of the improved GPT2 - based requirement compression model. The m - layer decoder of the improved GPT2 - based requirement compression model generates the decoder state s for each layer i , the attention distribution generated by the data through the encoder can be used as the probability distribution of words on the source text, denoted as a; Using the decoder state s for each layer i and the probability distribution a of words on the source text generated by the encoder to calculate the weight G: (2); where G ∈ [0, 1], is the activation function, and W and b are adjustable parameters, - is the decoder state value of each layer in the GPT2 model at time i t; (3); where P(w) is the final distribution of word w predicted by the improved GPT2-based requirement compression model at time t in the vocabulary, and the vocabulary is generated during pre-training; if w is a word outside the vocabulary, then P(w)=0. represents the attention distribution of word w at time t in the source text. If the word w does not appear, then ; T(w) is the final distribution of the word w predicted by the improved GPT2-based requirement compression model at time t in the vocabulary and the source text.

5. A patent recommendation method based on user needs combined with an inverted index according to claim 1, characterized in that, The extensive recommendation specifically is as follows: S321. Perform requirement segmentation processing on the user requirement information, search the inverted index for each processed segment of requirement information, and obtain the candidate set corresponding to each segment of requirement information. S322. Use the bert model for each segmented segment of requirement information to obtain its sentence vector representation. S323. Search the document list by the patent numbers in the candidate set to obtain the patent information sentence vectors corresponding to each patent number in the candidate set; then calculate the cosine similarity between the sentence vector of each segment of requirement information and each patent information sentence vector, and obtain the top k recommendation results for each segment according to the similarity calculation results; merge the top k recommendation results of different segments into an extensive recommendation list.

6. The patent recommendation method based on user requirements combined with an inverted index according to claim 5, wherein, The requirement segmentation specifically includes: 1) Obtain the user requirement information and divide it into z segments according to two symbols, semicolon and period. 2) Preprocess the segmented requirement information, and then extract keywords from each segment of requirement information through the TF-IDF keyword extraction mechanism. 3) Search the inverted index according to the keywords in the first segment to generate candidate set 1. 4) Use the bert model for the first segmented segment of requirement information to obtain its sentence vector representation. 5) Search the document list by the patent numbers in candidate set 1 to obtain the patent information sentence vectors corresponding to each patent number in the candidate set; calculate the cosine similarity between the sentence vector of the first segment of requirement information and the patent information sentence vectors in the candidate set, and select the top j patent information with the highest similarity to generate recommendation list 1. 6) Repeat steps 2)-5) for the remaining z - 1 segments of requirement information respectively, and merge all the generated recommendation lists to finally form an extensive recommendation list.

7. A patent recommendation method based on user needs combined with an inverted index according to claim 2, characterized in that, The preprocessing includes word segmentation, stop word removal, and special stop word removal operations.

Citation Information

Patent Citations

  • Personalized book recommendation method based on combinational algorithm

    CN107943910A