Patent recommendation method and system

By obtaining the number of recommended needs from potential buyers, determining the target extraction model, and constructing semantic profiles, the problem of low accuracy in existing patent recommendation methods is solved, achieving more efficient patent recommendation.

CN120975879APending Publication Date: 2025-11-18WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511067792.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing patent recommendation methods rely on text information input by the demander, resulting in low accuracy of matching results, difficulty in portraying the demander's long-term technical interests, and impact on matching speed and accuracy when input information is insufficient or excessive.

Method used

By obtaining the number of recommended needs of potential buyers, a target extraction model is determined. The model is then used to construct a semantic profile of potential buyers by calling the keywords pre-selected from multiple semantic extraction models. The semantic similarity between the profile and the patent text is calculated, and patent recommendation results are generated.

Benefits of technology

It improves the accuracy and efficiency of patent recommendations, fully utilizes the performance differences of different semantic extraction models in different scenarios, and generates more accurate patent recommendation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975879A_ABST
    Figure CN120975879A_ABST
Patent Text Reader

Abstract

The invention provides a patent recommendation method and system, and the method comprises the steps: obtaining a recommendation demand number selected by a potential purchaser at a first moment, and enabling a plurality of different recommendation demand numbers to serve as options for pre-configuration; determining associated patents of potential purchasers, calling keywords of the associated patents under the target extraction model, performing weighted fusion on vector representations of the called keywords based on the frequency of each keyword, and constructing semantic portraits of the potential purchasers; calling the semantic portrait of each patent in the patent library extracted by the target extraction model at the second moment; and calculating the semantic similarity between the semantic portraits of the potential purchasers and the semantic portraits of the patent text, and outputting a patent recommendation result according to the similarity ranking and the recommendation demand number selected by the potential purchasers at the first moment. According to the method, the expression difference of different semantic extraction models in different scenes is fully utilized, so that a more accurate patent recommendation result is more efficiently generated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of patent transaction, and particularly relates to a patent recommendation method and system. BACKGROUND

[0002] Under the background of increasing patent information and growing demand for technology transaction, how to efficiently mine the deep semantic content in the patent text has become an important basis for supporting key tasks such as technology retrieval, innovation evaluation and achievement transformation.

[0003] However, corresponding to this, the current patent transaction recommendation scene also pays more attention to the recommendation of patent technical solutions under semantic matching, that is, the demander inputs the theme or general technical direction of the patent to be purchased, and then performs semantic matching between the content input by the demander and the patent files in the patent library, so as to obtain patent files similar to the technical solution required by the demander.

[0004] On the one hand, such a matching method depends on the input of the demander as the semantic benchmark, and the matching result also depends on the input of the demander. If the effective information in the input text is less, it will have a greater impact on the accuracy of the matching result. If the content of the input text is more, it will have a greater impact on the matching speed. On the other hand, such a method pays more attention to the current technical demand of the demander, and is difficult to depict the long-term technical interest of the demander. In the case of single information, it is difficult to accurately recommend patents to the demander. SUMMARY

[0005] The present application provides a patent recommendation method and system to solve the problem of low accuracy of existing patent recommendation, and realizes a patent recommendation method with high accuracy and explainability.

[0006] The present application provides a patent recommendation method, comprising: obtaining a number of recommendation requirements selected by a potential purchaser at a first time, wherein a plurality of different recommendation requirements are pre-configured as options; determining the associated patents of the potential purchaser, calling the keywords of the associated patents under a target extraction model, weighting and fusing the vector representation of the called keywords based on the frequency of each keyword, and constructing the semantic portrait of the potential purchaser, wherein the target extraction model is pre-selected from a plurality of semantic extraction models according to the number of recommendation requirements selected by the potential purchaser at the first time and the matching hit result determined at the second time, and the second time is earlier than the first time; calling the semantic portrait of each patent in the patent library extracted by the target extraction model at the second time; Calculate semantic similarity between the semantic portrait of the potential buyer and the semantic portrait of the patent text, and output a patent recommendation result according to the similarity ranking and the number of recommendation requirements selected by the potential buyer at the first time.

[0007] According to the patent recommendation method provided by the application, the target extraction model selects a step in advance from a plurality of semantic extraction models according to the number of recommendation requirements selected by the potential buyer at the first time and the matching hit result at the second time, and specifically comprises: Collecting patent samples on the authorization day in a third period, and dividing the collected patent samples into a demand side set and a supply side set, wherein the third period is earlier than the second time; Using each semantic extraction model to extract the portrait of each demand side on the demand side set respectively, and obtaining a demand side portrait set corresponding to each semantic extraction model; Using each semantic extraction model to extract the portrait of each patent sample on the supply side set respectively, and obtaining a supply side portrait set corresponding to each semantic extraction model; For each semantic extraction model, match each data in the demand side portrait set in the supply side portrait set, and calculate the matching success rate under each number of recommendation requirements; The semantic extraction model with the highest matching success rate under the number of recommendation requirements selected by the potential buyer at the first time is selected as the target extraction model.

[0008] According to the patent recommendation method provided by the application, the step of collecting patent samples on the authorization day in a third period, and dividing the collected patent samples into a demand side set and a supply side set, specifically comprises: Removing patents that have not yet occurred transfer behavior and are valid from all patents in the third period on the authorization day to obtain first data; Reserving patents that have occurred at least once in the first data as second data; Identifying patents that do not have market transaction characteristics in the transfer behavior in the second data as third data, and reserving patents that are still in the effective period in the third data as fourth data; Reserving patents that are still reserved after the first data is removed from the fourth data as the collected patent samples, wherein the patents that are still reserved after the second data is removed from the third data are reserved as positive samples, and the remaining patents are reserved as negative samples; Dividing the collected patent samples into a demand side set and a supply side set, and the supply side set includes the negative samples.

[0009] According to the patent recommendation method provided by the application, the step of identifying patents that do not have market transaction characteristics in the transfer behavior in the second data, specifically comprises: The patents with a transfer time after the effective date of the patent right, the patents transferred within an enterprise, the patents with the assignee being one of the inventors, and the patents with the assignee being a potential associated party of the applicant are identified as the patents without market transaction characteristics.

[0010] According to the patent recommendation method provided by the application, the step of matching each piece of data in the demand side portrait set in the supply side portrait set and calculating the matching success rate under each recommended demand is specifically as follows: Each piece of data in the demand side portrait set is matched in the supply side portrait set to obtain a semantic similarity ranking; In the supply side portrait ranked in the front recommended demand number, if there is a patent portrait that the applicant of the demand side portrait has actually purchased, it is recorded as a recommendation success, otherwise, it is recorded as a recommendation failure as a matching result; According to the matching results of all demand side portraits, the matching success rate of each semantic extraction model under each recommended demand is calculated.

[0011] According to the patent recommendation method provided by the application, the semantic portrait of the potential buyer and the semantic portrait of the patent are both preset number of keywords.

[0012] The application also provides a patent recommendation system, comprising: The acquisition module is used to acquire the recommended demand number selected by the potential buyer at the first time, wherein a plurality of different recommended demand numbers are pre-configured as options; The extraction module is used to determine the associated patent of the potential buyer, call the keywords of the associated patent under the target extraction model, weight and fuse the vector representation of the called keywords based on the frequency of each keyword, and construct the semantic portrait of the potential buyer, wherein the target extraction model is pre-selected from a plurality of semantic extraction models according to the recommended demand number selected by the potential buyer at the first time and the matching hit result determined at the second time, and the second time is earlier than the first time; The calling module is used to call the semantic portrait of each patent in the patent library extracted by the target extraction model at the second time; The recommendation module is used to calculate the semantic similarity between the semantic portrait of the potential buyer and the semantic portrait of the patent text, and output the patent recommendation result according to the similarity ranking and the recommended demand number selected by the potential buyer at the first time.

[0013] The application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the patent recommendation method of any of the above when executing the program.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method recommended in any of the above-described patents.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the method recommended in any of the above-described patents.

[0016] The patent recommendation method and system provided by this invention determine the target extraction model by the number of recommendation needs selected by the potential buyer at the first moment, and generate a semantic profile of the potential buyer by calling the keywords pre-extracted by the target extraction model. Based on the matching result of the semantic profile of the potential buyer and the semantic profile of the called patent, a patent recommendation result is generated. This fully utilizes the performance differences of different semantic extraction models in different scenarios to generate more efficient and accurate patent recommendation results. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is one of the flowcharts illustrating the patent recommendation method provided by this invention; Figure 2 This is a schematic diagram of the patent sample collection process in the patent recommendation method provided by the present invention; Figure 3 This is the second flowchart of the patent recommendation method provided by the present invention; Figure 4 This is a schematic diagram of the patent recommendation system provided by the present invention; Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0020] The following is combined Figures 1 to 3 This invention patent recommends a method, such as... Figure 1As shown, comprising: Step 101, obtaining the number of recommended patents selected by the potential purchaser at the first time, wherein a plurality of different numbers of recommended patents are pre-configured as options; The different numbers of recommended patents are pre-configured as selectable options for the potential purchaser, and the number of recommended patents refers to the number of patents that the potential purchaser needs to recommend.

[0021] The first time represents the time when the potential purchaser actually uses the patent recommendation system to obtain recommended patents, and the number of recommended patents selected by the potential purchaser is obtained, for example, the options of recommending 30 patents, recommending 50 patents, and recommending 100 patents are pre-defined, the selection result of the user is received, and it is determined that the user selects to recommend 30 patents, so the number of recommended patents selected by the potential purchaser at the first time is 30.

[0022] Step 102, determining the associated patents of the potential purchaser, calling the keywords of the associated patents under the target extraction model, weighting and fusing the vector representation of the called keywords based on the frequency of each keyword, and constructing the semantic portrait of the potential purchaser, wherein the target extraction model is pre-selected from a plurality of semantic extraction models according to the number of recommended patents selected by the potential purchaser at the first time and the matching hit result determined at the second time, and the second time is earlier than the first time; The associated patents of the potential purchaser are the patents representing the preferences of the potential purchaser within a certain period of time.

[0023] Optionally, in the case that the potential purchaser has patent application records, at least the patents applied by the potential purchaser within a certain period of time are taken as the associated patents of the potential purchaser.

[0024] Optionally, in the case that the potential purchaser has no patent application records, the patents browsed, collected and / or downloaded by the potential purchaser within a certain period of time can be taken as the associated patents of the potential purchaser.

[0025] On this basis, the keywords pre-extracted by the target extraction model are called for the associated patents of the potential purchaser. Optionally, the keywords can be de-duplicated and sorted into a keyword set of the potential purchaser: ; Wherein, is the jth keyword, represents the total frequency of occurrence in the patent text of all associated patents.

[0026] Then, each keyword that does not coincide is vectorized to obtain its vector representation , and then the vector representation of the keyword is weighted and fused based on the frequency of occurrence of each keyword, and the result is taken as the semantic portrait of the potential purchaser : .

[0027] Further, the target extraction model is selected from a plurality of pre-determined semantic extraction models according to the recommended number of the potential purchaser, and the selected model is the one with the best performance under the current recommended number.

[0028] Specifically, the second time represents the historical time, that is, before the patent recommendation system goes online, the accuracy of each semantic extraction model under the recommended number corresponding to each option is determined, and the semantic extraction model with the highest accuracy under each recommended number is determined.

[0029] Because different semantic extraction models have different representation granularity and generalization ability for text, for example, models that are good at capturing local precise matching and models that are good at capturing deep semantic association, the performance under different recommended numbers is different. Therefore, before the patent recommendation system goes online, the semantic extraction model with the best matching hit result after keyword extraction under different recommended numbers is determined, and the recommended number option is pre-configured according to the determined recommended number. For example, when the recommended number is determined to be 20, 50 and 100, the optimal semantic extraction model corresponding to each recommended number is determined, and three recommended number options of recommending 20 patents, 50 patents and 100 patents are configured.

[0030] On this basis, according to the test results at the historical time and the recommended number selected by the potential purchaser at the current time, the semantic extraction model with the best performance under the recommended number can be used as the target extraction model, and the keywords and keyword frequencies of the associated patents pre-extracted and stored in the target extraction model are called to generate the semantic portrait of the potential purchaser.

[0031] Step 103, calling the semantic portrait of each patent in the patent library extracted by the target extraction model at the second time; It can be understood that at the second time, that is, when the semantic extraction model with the best performance under each recommended number is determined, the determined semantic extraction model can be used to input the patent text of the patent in the patent library into the semantic extraction model to obtain the keywords of each patent under the semantic extraction model.

[0032] Alternatively, the patent text can be the full text of the original application file of the patent, or the specification of the patent, and in the present embodiment, the text is the combination of the title, abstract and claims of the patent.

[0033] The patent text of the patent is input into the determined several semantic extraction models, and the pre-set number of keywords output by the model are obtained to represent the semantic portrait of the patent.

[0034] Optionally, the number of keywords output by each semantic extraction model can be the same or different. In the embodiment, each semantic extraction model outputs 15 keywords to construct the patent image of a patent, so as to achieve an effective balance between semantic expressiveness and representation sparsity.

[0035] Then, the arithmetic mean of the keyword vector corresponding to each patent is calculated as the semantic image of the patent: ; In the formula, denotes the semantic image of the patent, n denotes the number of keywords, is the vector representation of the i-th keyword. l

[0036] In the embodiment, after determining the semantic extraction model with the best performance under each recommendation number at the second time, a database is established for each semantic extraction model, and each piece of data in the database includes the keyword vector and the word frequency of the patent, and the patent image of the patent.

[0037] Optionally, the patent image of the granted patent still in the effective period can also be stored separately for matching to obtain the recommendation result.

[0038] On this basis, after determining the recommendation requirement number selected by the potential purchaser at the first time, the target recommendation model can be determined, and then the database of the template recommendation model is called to read the keyword vector and the word frequency of the associated patent for constructing the potential purchaser image, and the patent image is read for matching with the potential purchaser image, so as to effectively improve the matching speed.

[0039] Step 104, calculating the semantic similarity between the semantic image of the potential purchaser and the semantic image of the patent text, and outputting the patent recommendation result according to the similarity ranking and the recommendation requirement number selected by the potential purchaser at the first time.

[0040] In the embodiment, the cosine similarity is used to calculate the semantic image of the potential purchaser and the semantic image of each patent text in the patent library to measure the semantic correlation: ; In the formula, is the similarity value.

[0041] According to the ranking from high to low of the semantic similarity calculation result and the recommendation requirement number selected by the potential purchaser, the top recommendation requirement number patents with the highest similarity ranking are output as the patent recommendation result for the potential purchaser.

[0042] ​The application determines a target extraction model according to the number of recommendation requirements selected by the potential purchaser at the first moment, and generates a semantic portrait of the potential purchaser by calling the keywords extracted in advance by the target extraction model, and generates a patent recommendation result based on the matching result of the semantic portrait of the potential purchaser and the semantic portrait of the called patent, so as to fully utilize the performance difference of different semantic extraction models in different scenarios, and more efficiently generate more accurate patent recommendation results.

[0043] In the patent recommendation method, the step of preselecting the target extraction model from a plurality of semantic extraction models according to the number of recommendation requirements selected by the potential purchaser at the first moment and the matching hit result at the second moment specifically comprises: Collecting patent samples in a third period on the authorization date, and dividing the collected patent samples into a demand side set and a supply side set, wherein the third period is earlier than the second moment; In order to determine the performance difference of different semantic extraction models in different scenarios, in this embodiment, patent data in a historical period is collected to construct a data set, so as to determine the performance of different semantic extraction models under different recommendation requirements based on real historical data.

[0044] Specifically, the third period and the historical period earlier than the second moment are set to January 1, 2020 to December 31, 2011 in this embodiment, and 200,000 authorized patents with IPC classification number G are further collected as patent samples.

[0045] It can be understood that since a patent is a kind of unlabeled data, it is difficult to directly evaluate the performance difference of different semantic extraction models in keyword representation, therefore, in this embodiment, a simulation matching evaluation framework based on historical patent data is designed. The framework divides the collected patent samples in the historical period into a demand side set and a supply side set, which are used to construct demand side portraits and patent portraits respectively, and generates a patent recommendation sequence for each demand side by matching the demand side portrait and the supply side patent portrait. According to whether there is a patent that the demand side has actually purchased in the patent recommendation sequence corresponding to each recommendation requirement, it is determined whether the recommendation is successful, and the performance of each semantic extraction model under each recommendation requirement is defined according to the recommendation success rate.

[0046] Therefore, in this embodiment, the collected samples are divided into a demand side set and a supply side set, wherein the demand side set includes demand sides that have generated real patent transaction behaviors.

[0047] The portrait of each demand side is extracted by using each semantic extraction model on the demand side set, and the demand side portrait set corresponding to each semantic extraction model is obtained; An image of each patent sample is extracted from the supply side set using each semantic extraction model, to obtain a supply side image set corresponding to each semantic extraction model; In this embodiment, four semantic extraction models, TF-IDF, TextRank, EmbedRank and DeepSeek V3, are determined in advance to construct multi-level and multi-angle semantic representation of patent texts.

[0048] Specifically, TF-IDF constructs a sparse vector based on word frequency and inverse document frequency, and is suitable for extracting significant text keywords; TextRank constructs a graph structure based on word co-occurrence, and uses a graph ranking algorithm to identify local semantic center words; EmbedRank uses sentence vectors generated by a pre-trained language model to project the entire text and candidate phrases into a unified embedding space, and sorts and filters keywords by calculating semantic similarity, which can efficiently obtain highly relevant keyword expressions under unsupervised conditions; DeepSeek V3 is a pre-trained language model based on a hybrid expert mechanism (MoE), which has the ability to abstractly model semantic associations, technical term structures and semantic evolution paths in long texts, and is suitable for extracting deep semantic embedding representations to reveal the concept kernel and structural logic of technical content.

[0049] In the demand side set, the applicants as demanders are counted, and the images of the demanders are constructed according to the patent texts of the demanders using the above four semantic extraction models, to obtain the demander image sets of the four semantic extraction models.

[0050] In the supply side set, the images of each patent sample are constructed using the four semantic extraction models, to obtain the supply side image sets corresponding to the four semantic extraction models.

[0051] In a specific embodiment, TF-IDF uses TfidfVectorizer in scikit-learn, the vocabulary limit is 3000, the minimum word frequency is 5, and stop words are uniformly removed. TextRank is sorted based on the co-occurrence graph, and the window size is set to 5. EmbedRank is an unsupervised keyword extraction method based on embedding space semantic alignment, which uses a Chinese pre-trained BERT model to generate context embedding representations of text and candidate phrases, and sorts and filters candidate phrases by calculating the cosine similarity between embedding vectors. In the keyword extraction process, the maximum n-gram length of the candidate phrase is set to 5, and the top 15 keywords with the highest similarity scores are finally output. DeepSeek V3 uses corresponding instructions to extract semantic keywords of patents, limits the text length to 512 tokens, and completes the inference in a GPU environment. To support the above semantic keyword extraction process and the processing of large-scale patent data, this study uses Python 3.11 and CUDA 11.8 as the development and running platform, and all model training and matching recommendation tasks are completed on a high-performance server with 16GB of memory, 64-thread CPU, NVIDIA GeForce RTX 4090 GPU, and 6TB of storage space.

[0052] For each semantic extraction model, each data in the demand side portrait set is matched in the supply side portrait set, and the matching success rate under each recommended demand is calculated; The semantic extraction model with the highest matching success rate under the recommended demand selected by the potential purchaser at the first time is selected as the target extraction model.

[0053] For each semantic extraction model, each data in its demand side portrait set, that is, each demand side portrait, is sequentially matched in its supply side portrait set, and the matching success rate under each recommended demand is calculated according to the matching result, and the semantic extraction model with the highest matching success rate under each recommended demand is determined.

[0054] After obtaining the recommended demand selected by the potential purchaser at the current time, the semantic extraction model with the highest matching success rate under the recommended demand is determined as the target extraction model.

[0055] In the patent recommendation method of the application, the step of dividing the collected patent samples into demand side set and supply side set in the third period of the authorized day, specifically includes: Remove the patents that have not occurred transfer behavior and are effective from all patents in the third period of the authorized day to obtain the first data; Keep the patents that have occurred at least once transfer behavior in the first data as the second data; identify the patents without market transaction characteristics in the second data as third data, and the patents still in the effective period in the third data as fourth data; retain the patents still remaining after the first data is removed from the fourth data as the collected patent samples, wherein the patents still remaining after the second data is removed from the third data are retained as positive samples, and the rest are negative samples; divide the collected patent samples into a demand side set and a supply side set, and the negative samples are included in the supply side set.

[0056] Further, in order to reasonably divide the demand side set and the supply side set, in the embodiment, data cleaning is performed on the patent samples when the patent samples are collected.

[0057] Specifically, still taking the 200,000 granted patents collected in the G part as an example, first, 23,922 patents still in the effective period and without transfer behavior are removed, and 176,078 patents are obtained as the first data.

[0058] It can be understood that the removed patents do not have transfer behavior, but since they are still in the effective period, it cannot be determined whether the subsequent transfer behavior will occur, and therefore it cannot be determined whether such patents are patent samples with market value, and therefore they are removed.

[0059] On this basis, 38,759 patents that have at least one transfer behavior are further screened out from the first data as the second data, representing patents that may have market value.

[0060] This is because although some patents have transfer behavior, this behavior may be a transfer within a group or a non-market transaction behavior such as a company name change or address change, and such patents do not represent real market transaction characteristics, and therefore not all second data represents market transaction characteristics.

[0061] On this basis, further identify the patents without market transaction characteristics in the second data as third data, and the patents still remaining after the second data is removed from the third data, i.e., 18,488 patents with real market transaction characteristics, are retained as positive samples in the patent samples used to construct the data set.

[0062] It should be noted that although the third data is patents without market transaction characteristics, for some of them still in the effective period of the patent, it cannot be determined whether other real transactions will occur in the future, and therefore 6,841 patents still in the effective period in the third data are identified as fourth data, and the remaining part is the collected patent samples.

[0063] The collected patent samples are divided into a demand side set and a supply side set, so that the samples in the demand side set are all positive samples, and the samples in the supply side set contain both positive samples and negative samples, that is, the demand side in the demand side set all has the real patent transaction behavior, and the samples in the supply side set include patents that are actually purchased by the demand side and other real transaction patents, and also include negative samples that do not have market value.

[0064] In the patent recommendation method, the step of identifying the patent without market transaction characteristics in the second data specifically includes: The patents with a transfer time after the validity date of the patent right, the patents transferred within the enterprise, the patents with the assignee being one of the inventors, and the patents with the assignee being a potential associated party of the applicant are identified as the patents without market transaction characteristics.

[0065] In order to accurately identify the patents without real market transaction characteristics, the following screening framework is proposed in the embodiment: In the second data, first, 250 patents with a transfer occurring after the termination of the legal state, that is, the patents with a transfer occurring at the time of invalidation or expiration, are removed.

[0066] Secondly, 4413 patents transferred within the enterprise and 294 patents with the assignee being one of the inventors are removed.

[0067] Further, in order to more accurately remove the patents without real market transaction, in the remaining 33802 patents, 15314 potential associated party transfer records, that is, the patents with the assignee being a potential associated party of the applicant, are cleaned in the following way: On the one hand, whether there is a stock ownership, employment or management association between the appropriate transfer parties through a third-party platform is determined, and if there is, it is considered that the patent with the assignee being a potential associated party of the applicant; on the other hand, the text similarity of the names and addresses of the transfer parties is evaluated by using the Levenshtein distance algorithm, and the patents with a similarity greater than a preset threshold are identified as the patents with the assignee being a potential associated party by combining manual verification auxiliary judgment.

[0068] Through the above-mentioned way, 18488 patents are finally identified as positive samples. The complete patent sample collection and screening process is as shown in Figure 2 .

[0069] In the patent recommendation method, the step of using each data in the demand side portrait set to match in the supply side portrait set and calculating the matching success rate under each recommended demand specifically includes: Each data in the demand side portrait set is used to match in the supply side portrait set to obtain a semantic similarity ranking; In the ranked top recommended number of supply side images, if there is a patent image that the applicant of the demand side image has actually purchased, it is recorded as a recommendation success, otherwise, it is recorded as a recommendation failure as a matching result; The matching success rate of each semantic extraction model under each recommended number of requirements is calculated according to the matching results of all demand side images.

[0070] In this embodiment, the cosine similarity is also used to calculate the semantic similarity between the demand side image and the patent image of the supply side, and a corresponding semantic similarity ranking is generated for each demand side.

[0071] Since the recommended number of requirements is predetermined as 20, 50 and 100, a list of top 100 patents ranked by semantic similarity is generated for each demand side image.

[0072] In the case of a recommended number of requirements of 20, the top 20 patents with the highest semantic similarity for each demand side image are obtained, and if there is a patent that the demand side has actually purchased, it is considered to be a recommendation success, otherwise, it is considered to be a recommendation failure.

[0073] The results of each semantic extraction model for each demand side recommendation success and recommendation failure are counted, and the ratio of the number of recommendation successes to the total number of demand sides is used as the matching success rate under the recommended number of requirements.

[0074] In one specific embodiment, using the above data set and the above four semantic extraction models, the matching success rates under the recommended number of requirements of 20, 50 and 100 are calculated as shown in Table 1:

[0075] It can be seen that under the recommended number of requirements of 20, the best performing model is TF-IDF, which is more sensitive locally, and has higher precision. When the recommended number of requirements is 50 and 100, the best performing model is EmbedRanK, which pays more attention to deep semantic connections, and has better performance in recall rate.

[0076] Therefore, in this embodiment, when the recommended number of requirements selected by the potential buyer is 20, TF-IDF is used to generate the patent recommendation results of the potential buyer, and when the recommended number of requirements selected by the potential buyer is 50 or 100, EmbedRanK is used to generate the patent recommendation results of the potential buyer, so as to improve the accuracy of the recommendation results.

[0077] It should be noted that if a new semantic extraction model appears in iteration, the new semantic extraction model can also be used to extract and match on the demand side dataset and the supply side dataset constructed by the above-mentioned manner, and the matching accuracy under each recommended demand number is calculated, and the semantic extraction model with the best performance under each recommended demand number is updated based on the matching accuracy. Such a way also decouples the updating process of the patent recommendation system from the user's use process, so as to ensure that the user only needs to call the data stored in the maintained database when obtaining recommended patents, effectively improves the matching rate, and thus improves the user's use experience.

[0078] On the basis of the above, a complete semantic portrait construction and matching process is as shown in Figure 3 Based on the significant differences between different model adaptation scenarios, the present application fully utilizes the advantages of shallow models in processing high-frequency features and concentrated technical needs, and the better adaptability and robustness of deep semantic models in processing semantically heterogeneous and multi-topic cross patent texts, and uses different semantic extraction models for recommendation for users in small-range accurate scenarios and large-range coverage scenarios, thereby improving the recommendation accuracy.

[0079] The patent recommendation system provided by the present application will be described below. The patent recommendation system described below can be referred to in conjunction with the patent recommendation method described above.

[0080] As Figure 4 The patent recommendation system of the present application comprises an acquisition module 401, an extraction module 402, a calling module 403 and a recommendation module 404. The acquisition module 401 is used to acquire the number of recommendations selected by the potential purchaser at the first time, wherein a plurality of different numbers of recommendations are pre-configured as options. The different numbers of recommendations are pre-configured as optional items for the potential purchaser. The number of recommendations is the number of patents that the potential purchaser needs to recommend.

[0081] The first time represents the time when the potential purchaser actually uses the patent recommendation system to obtain recommended patents. The number of recommendations selected by the potential purchaser is acquired. For example, the options of recommending 30 patents, recommending 50 patents and recommending 100 patents are pre-defined, the selection result of the user is received, it is determined that the user selects to recommend 30 patents, and the number of recommendations selected by the potential purchaser at the first time is 30.

[0082] The extraction module 402 is configured to determine the associated patents of the potential purchaser, call the keywords of the associated patents under a target extraction model, weight and fuse the vector representation of the called keywords based on the frequency of each keyword, and construct a semantic portrait of the potential purchaser, wherein the target extraction model is pre-selected from a plurality of semantic extraction models according to the number of recommended requirements selected by the potential purchaser at a first time and the matching hit result determined at a second time, and the second time is earlier than the first time. The associated patents of the potential purchaser are the patents representing the preferences of the potential purchaser within a certain period of time.

[0083] Optionally, in the case that the potential purchaser has patent application records, at least the patents applied by the potential purchaser within a certain period of time are taken as the associated patents of the potential purchaser.

[0084] Optionally, in the case that the potential purchaser has no patent application records, the patents browsed, collected and / or downloaded by the potential purchaser within a certain period of time can be taken as the associated patents of the potential purchaser.

[0085] On this basis, the keywords of the associated patents of the potential purchaser are called using the keywords pre-extracted by the target extraction model. Optionally, the keywords can be de-duplicated and arranged into a keyword set of the potential purchaser: ; wherein, is the jth keyword, and represents the total frequency of occurrence in the patent texts of all associated patents.

[0086] Then, each keyword that does not coincide is vectorized to obtain its vector representation , and then the vector representation of the keyword is weighted and fused based on the frequency of occurrence of each keyword, and the result is taken as the semantic portrait of the potential purchaser : .

[0087] Further, the target extraction model selects the model with the best performance under the current recommended requirement number from a plurality of pre-rated semantic extraction models according to the number of recommended requirements selected by the potential purchaser, and uses the model as the target extraction model for this recommendation.

[0088] Specifically, the second time represents a historical time, that is, before the patent recommendation system is put into operation, the accuracy rate of each semantic extraction model under the recommended requirement number corresponding to each option is determined, and the semantic extraction model with the highest accuracy rate under each recommended requirement number is determined.

[0089] Because different semantic extraction models have fundamentally different representational granularities and generalization abilities for text—for example, models that excel at capturing precise local matches and models that excel at capturing deep semantic relationships—their performance varies depending on the number of recommendation requests. Therefore, before the patent recommendation system goes live, the semantic extraction model with the optimal keyword extraction matching results for different recommendation numbers is pre-determined, and recommendation request number options are pre-configured based on the determined recommendation numbers. For example, if the optimal semantic extraction models for recommendation numbers of 20, 50, and 100 are determined, then three recommendation request number options—20 patents, 50 patents, and 100 patents—are configured.

[0090] Based on this, according to the test results at historical moments and the number of recommended needs selected by potential buyers at the current moment, the semantic extraction model that performs best under the number of recommended needs can be used as the target extraction model, and the keywords and keyword frequencies pre-extracted and stored by the associated patent under the target extraction model can be called to generate a semantic profile of potential buyers.

[0091] The calling module 403 is used to call the semantic profile of each patent in the patent library extracted by the target extraction model at the second time step; Understandably, at the second moment, when the optimal semantic extraction model for each number of recommendations is determined, the determined semantic extraction model can be used to input the patent texts of the patents in the patent database into the semantic extraction model to obtain the keywords of each patent under that semantic extraction model.

[0092] Optionally, the patent text can be the full text of the original patent application documents or the patent specification. In this embodiment, it is a combination of the patent title, abstract, and claims.

[0093] The patent text is input into several defined semantic extraction models to obtain a preset number of keywords output by the models, which are used to characterize the semantic profile of the patent.

[0094] Optionally, the number of keywords output by each semantic extraction model can be the same or different. In this embodiment, each semantic extraction model outputs 15 keywords to construct a patent profile for a patent, thereby achieving an effective balance between semantic expressiveness and representation sparsity.

[0095] Then, the arithmetic mean of the keyword vectors corresponding to each patent is calculated to form the semantic profile of that patent: ; In the formula, A semantic profile representing a patent. n Indicates the number of keywords. For the first l Vector representation of each keyword.

[0096] In the embodiment, after determining the semantic extraction model with the best performance under each recommended number at the second time, a database is established for each semantic extraction model, and one piece of data in the database includes the keyword vector and the word frequency of the patent, and the patent image of the patent.

[0097] Optionally, the patent images of the granted patents still in the effective period can also be stored separately for matching to obtain the recommended results.

[0098] On this basis, after determining the recommended number selected by the potential purchaser at the first time, the target recommendation model can be determined, and then the database of the template recommendation model is called to read the keyword vector and the word frequency of the associated patent for constructing the potential purchaser image, and the patent image is read for matching with the potential purchaser image, thereby effectively improving the matching speed.

[0099] The recommendation module 404 is configured to calculate the semantic similarity between the semantic image of the potential purchaser and the semantic image of the patent text, and output the patent recommendation results according to the similarity ranking and the recommended number selected by the potential purchaser at the first time.

[0100] In the embodiment, the cosine similarity is used to calculate the semantic image of the potential purchaser and the semantic image of each patent text in the patent library to measure the semantic correlation: In the formula, is the similarity value.

[0101] According to the ranking from high to low of the semantic similarity calculation result and the recommended number selected by the potential purchaser, the recommended number of patents with the highest similarity ranking are output as the patent recommendation results for the potential purchaser.

[0102] The present application determines the target extraction model according to the recommended number selected by the potential purchaser at the first time, generates the semantic image of the potential purchaser by using the keywords extracted by the target extraction model in advance, and generates the patent recommendation results based on the matching results of the semantic image of the potential purchaser and the semantic image of the called patent, so as to fully utilize the performance differences of different semantic extraction models in different scenarios, and more efficiently generate more accurate patent recommendation results.

[0103] Figure 5 An example of an entity structure diagram of an electronic device is shown in FIG. 1. Figure 5 ​As shown, the electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 complete mutual communication through the communications bus 540. The processor 510 can call the logic instructions in the memory 530 to execute the patent recommendation method, which includes: acquiring a recommendation requirement number selected by a potential purchaser at a first time, wherein a plurality of different recommendation requirement numbers are pre-configured as options; determining an associated patent of the potential purchaser, calling keywords of the associated patent under a target extraction model, weighting and fusing vector representations of the called keywords based on frequencies of each keyword, and constructing a semantic portrait of the potential purchaser, wherein the target extraction model is pre-selected from a plurality of semantic extraction models according to the recommendation requirement number selected by the potential purchaser at the first time and a matching hit result determined at a second time, and the second time is earlier than the first time; calling semantic portraits of each patent in a patent library extracted by the target extraction model at the second time; calculating semantic similarity between the semantic portrait of the potential purchaser and semantic portraits of the patent texts, and outputting a patent recommendation result according to the similarity ranking and the recommendation requirement number selected by the potential purchaser at the first time.

[0104] In addition, the logic instructions in the memory 530 described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part of the technical solutions that essentially contribute to the prior art or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0105] In another aspect, the present application also provides a computer program product comprising a computer program, the computer program being stored in a non-transitory computer-readable storage medium, and the computer program being capable of executing the patent recommendation method provided by the above method when executed by a processor, the method comprising: obtaining a number of recommendation requirements selected by a potential purchaser at a first time, wherein a plurality of different numbers of recommendation requirements are pre-configured as options; determining associated patents of the potential purchaser, calling keywords of the associated patents under a target extraction model, performing weighted fusion on vector representations of the called keywords based on frequencies of each keyword, and constructing a semantic portrait of the potential purchaser, wherein the target extraction model is pre-selected from a plurality of semantic extraction models according to the number of recommendation requirements selected by the potential purchaser at the first time and a matching hit result determined at a second time, the second time being earlier than the first time; calling semantic portraits of each patent in a patent library extracted by the target extraction model at the second time; calculating semantic similarity between the semantic portrait of the potential purchaser and semantic portraits of the patent texts, and outputting a patent recommendation result according to a similarity ranking and the number of recommendation requirements selected by the potential purchaser at the first time.

[0106] In another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, the computer program being capable of implementing the patent recommendation method provided by the above method when executed by a processor, the method comprising: obtaining a number of recommendation requirements selected by a potential purchaser at a first time, wherein a plurality of different numbers of recommendation requirements are pre-configured as options; determining associated patents of the potential purchaser, calling keywords of the associated patents under a target extraction model, performing weighted fusion on vector representations of the called keywords based on frequencies of each keyword, and constructing a semantic portrait of the potential purchaser, wherein the target extraction model is pre-selected from a plurality of semantic extraction models according to the number of recommendation requirements selected by the potential purchaser at the first time and a matching hit result determined at a second time, the second time being earlier than the first time; calling semantic portraits of each patent in a patent library extracted by the target extraction model at the second time; calculating semantic similarity between the semantic portrait of the potential purchaser and semantic portraits of the patent texts, and outputting a patent recommendation result according to a similarity ranking and the number of recommendation requirements selected by the potential purchaser at the first time.

[0107] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0108] Those skilled in the art can clearly understand the implementation of the embodiments by the description of the above embodiments, and the embodiments can be implemented by means of software and necessary universal hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the method described in each embodiment or some parts of the embodiment.

[0109] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for some technical features thereof; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A patent recommendation method, characterized in that, include: Obtain the number of recommended needs selected by potential buyers at the first moment, where multiple different recommended needs are pre-configured as options; Identify the associated patents of potential buyers, call the keywords of the associated patents under the target extraction model, and perform weighted fusion of the vector representations of the called keywords based on the frequency of each keyword to construct a semantic profile of potential buyers. The target extraction model is pre-selected from multiple semantic extraction models based on the number of recommended needs selected by the potential buyer at the first time and the matching hit result determined at the second time. The second time is earlier than the first time. The semantic profile of each patent in the patent database extracted by the target extraction model at the second time step is invoked; Calculate the semantic similarity between the semantic profile of the potential buyer and the semantic profile of the patent text, and output the patent recommendation result based on the similarity ranking and the number of recommended needs selected by the potential buyer at the first moment.

2. The patent recommendation method according to claim 1, characterized in that, The target extraction model, based on the number of recommended needs selected by the potential buyer at the first moment and the matching results at the second moment, pre-selects from multiple semantic extraction models, specifically including: Patent samples from the third time period on the grant date are collected, and the collected patent samples are divided into demand-side set and supply-side set, wherein the third time period is earlier than the second time period; Each semantic extraction model is used to extract a profile of each demander on the demand side set, resulting in a demander profile set corresponding to each semantic extraction model. Each semantic extraction model is used to extract a profile of each patent sample on the supply-side set, resulting in a supply-side profile set corresponding to each semantic extraction model. For each semantic extraction model, each piece of data in the demand-side profile set is matched in the supply-side profile set, and the matching success rate is calculated for each number of recommended demands. The semantic extraction model with the highest matching success rate among the recommended needs selected by potential buyers at the first moment is selected as the target extraction model.

3. The patent recommendation method according to claim 2, characterized in that, The step of collecting patent samples from the third time period on the authorization date and dividing the collected patent samples into demand-side and supply-side sets specifically includes: The first data is obtained by removing valid patents that have not yet been transferred from all patents in the third period of the authorization date. Patents that have been transferred at least once are retained in the first data as the second data; Patents whose transfer transactions do not exhibit market transaction characteristics are identified in the second data and used as the third data; patents that are still valid in the third data are used as the fourth data. The patents that remain after removing the fourth data from the first data are used as the collected patent samples, wherein the patents that remain after removing the third data from the second data are used as positive samples, and the remaining patents are used as negative samples. The collected patent samples are divided into a demand-side set and a supply-side set, and the supply-side set includes the negative samples.

4. The patent recommendation method according to claim 3, characterized in that, The step of identifying patents in the second data whose transfer behavior does not possess market transaction characteristics specifically includes: Patents that are transferred after the patent's expiry date, patents transferred within an enterprise, patents where the transferee is one of the inventors, and patents where the transferee is a potential affiliate of the applicant are identified as patents that do not have market transaction characteristics.

5. The patent recommendation method according to claim 2, characterized in that, The step of matching each piece of data in the demand-side profile set with the supply-side profile set and calculating the matching success rate for each recommended demand number specifically includes: Each piece of data in the demand-side profile set is matched against the supply-side profile set to obtain a semantic similarity ranking. If, among the top-ranked supplier profiles with the highest number of recommended demands, there is a patent profile that the applicant has actually purchased, the recommendation is considered successful; otherwise, it is considered a failed recommendation and is recorded as a matching result. The matching success rate of each semantic extraction model is calculated based on the matching results of all demand profiles for each recommended demand.

6. The patent recommendation method according to claim 1, characterized in that, Both the semantic profile of the potential buyer and the semantic profile of the patent consist of a preset number of keywords.

7. A patent recommendation system, characterized in that, include: The acquisition module is used to acquire the number of recommended needs selected by potential buyers at the first moment, where multiple different numbers of recommended needs are pre-configured as options; The extraction module is used to determine the associated patents of potential buyers, call the keywords of the associated patents under the target extraction model, and perform weighted fusion of the vector representation of the called keywords based on the frequency of each keyword to construct a semantic profile of potential buyers. The target extraction model is pre-selected from multiple semantic extraction models based on the number of recommended needs selected by the potential buyer at the first time and the matching hit result determined at the second time. The second time is earlier than the first time. The calling module is used to call the semantic profile of each patent in the patent database extracted by the target extraction model at the second time step; The recommendation module is used to calculate the semantic similarity between the semantic profile of the potential buyer and the semantic profile of the patent text, and output the patent recommendation result based on the similarity ranking and the number of recommendation requests selected by the potential buyer at the first moment.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the patent recommendation method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the patent recommendation method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the patent recommendation method as described in any one of claims 1 to 6.