Patent information push management system and method

By designing a patent information push management system, using real-time data collection, user image analysis and semantic processing technologies, the problems of information overload and unrelated information interference in the existing system are solved, and efficient and accurate patent information push is achieved.

CN120067455APending Publication Date: 2025-05-30GUANGDONG JUZHICHENG TECH

Patent Information

Application Number
CN202510297001.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing patent information acquisition system has problems of information overload and unrelated information interference, making it difficult to provide efficient and accurate patent information push.

Method used

A patent information push management system is designed to capture policy and regulation data and user portrait units in real time through the data acquisition module to analyze user behavior data, and combine the two-way LSTM word segmentation model of the semantic processing module and the multi-level matching strategy of the rule knowledge module to achieve accurate patent information push.

Benefits of technology

It improves the timeliness and integrity of patent information, and the push content is more in line with user needs, reduces information overload and unrelated information interference, and improves the accuracy and user satisfaction of push.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067455A_ABST
    Figure CN120067455A_ABST
Patent Text Reader

Abstract

The invention relates to a patent information push management system and method. The system comprises a data acquisition module, a semantic processing module, a feature extraction module, a rule knowledge module, a matcher module, a feedback optimization module and a data storage module. The data acquisition module is used for acquiring policy and regulation, industrial dynamics and user related text data sets; the semantic processing module performs word segmentation on the data set and generates word vectors; the feature extraction module extracts a reference word set and a description word set; the rule knowledge module stores a push rule tree with an IPC classification number as a root node; the matcher executes a multi-stage matching strategy; the feedback optimization module optimizes the system according to user feedback; and the data storage module adopts a time sequence database to store related records. The method comprises the steps of establishing a dynamic strategy library, constructing a user portrait matrix, implementing double-layer matching, generating a push decision and establishing a feedback closed loop. According to the invention, accurate and efficient patent information pushing can be realized, the integrating degree of the pushed content and the user demand is improved, and the system performance can be continuously optimized according to the user feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data management, and particularly relates to a patent information push management system and method. Background Art

[0002] With the rapid development of global scientific and technological innovation, the number of patents has increased explosively. Researchers, enterprise R & D teams, and patent investors in various fields need to constantly pay attention to the latest patent information closely related to their own businesses in order to obtain the latest technological trends, avoid infringement risks, and explore potential cooperation and investment opportunities. However, there are many problems in obtaining patent information at present: on the one hand, patent data sources are extremely scattered, including both the databases of national official patent offices and numerous commercial patent information platforms. The formats and classification standards of different data sources vary greatly, and it is difficult for users to integrate the required information in one stop; on the other hand, the traditional information push mode is too extensive, mostly using a universal push with a unified template and fixed frequency, without considering the individual differences of users. As a result, the pushed content often has a low degree of fit with the actual needs of users, and a large amount of irrelevant information interferes with users, causing information overload. For example, a researcher focusing on new energy vehicle battery technology may frequently receive patent pushes in unrelated fields such as software algorithms and mechanical manufacturing, which not only wastes their precious time but also easily causes them to miss key battery technology updates. In addition, there is a lack of a precise push evaluation mechanism, and it is impossible to optimize the push strategy in a timely manner based on user feedback, so the push effect is difficult to improve, and it is difficult to meet the urgent need for efficient and precise patent information services at present.

[0003] In view of this, there is an urgent need to propose a patent information push management system and method. Summary of the Invention

[0004] To this end, the present invention provides a patent information push management system to achieve precise patent information push and management according to policy guidance.

[0005] In the first aspect of the present invention, a patent information push management system is provided, including:

[0006] A data collection module, configured to include a policy collection unit and a user portrait unit;

[0007] The policy collection unit captures policy regulations and industrial dynamic data of official patent agency websites and industry databases in real time through an API interface to form a first text data set;

[0008] The user portrait unit extracts a second text data set by analyzing the historical behavior logs, project application materials, and collection records of users on the patent retrieval platform;

[0009] The semantic processing module is configured to be a word segmentation model with a built-in bidirectional LSTM neural network, which performs dynamic word segmentation processing on the first and second text datasets and generates word vectors; the word segmentation model is pre-trained based on a patent domain corpus and includes a part-of-speech tagging layer and a stop word filtering layer;

[0010] The feature extraction module extracts a first set of benchmark words composed of core nouns and a first set of descriptive words composed of restrictive phrases from the segmented first text data, and extracts a second set of benchmark words composed of user feature nouns and a second set of descriptive words composed of associated adjectives from the second text data; where the set of benchmark words satisfies a preset word length rule, and the set of descriptive words uses a sliding window algorithm to extract prefix and suffix phrases;

[0011] The rule knowledge module stores a push rule tree with the IPC classification number as the root node, and the push rule tree includes multiple rule nodes, where each of the rule nodes includes:

[0012] The first matching condition describes the Jaccard similarity between the user's benchmark words and the policy benchmark words;

[0013] The second matching condition calculates the cosine similarity between words through an attention mechanism;

[0014] The third matching condition establishes a dependency relationship graph between parent and child rule nodes through the paragraph position encoding of the policy text;

[0015] The matcher is configured to execute a multi-level matching strategy, including:

[0016] The first matching is to match the user's second set of benchmark words with the rule nodes to generate a candidate rule set;

[0017] The second matching calculates the weighted matching degree in the candidate rule set, and the weighted matching degree includes a benchmark matching degree, a description matching degree, and a logical association degree with weights set respectively;

[0018] Among them,

[0019] The benchmark degree = the number of matching benchmark words / the total number of benchmark words;

[0020] The description matching degree = Σ(description word similarity × TF-IDF weight);

[0021] The logical integrity degree = the connectivity of the matching rule node in the dependency relationship graph;

[0022] The feedback optimization module includes a feedback collection unit and a model tuning unit, and the feedback collection unit collects the user's click-through rate and reading duration after each push;

[0023] The model tuning module is configured to perform the following optimizations when the target click-through rate and target reading time of the user do not reach the preset threshold after continuously performing the preset number of pushes:

[0024] Adjust the word length rule threshold of the benchmark word set;

[0025] Calculate the attention weight distribution for describing word similarity;

[0026] Perform incremental training on the LSTM word segmentation model based on the user's click-through rate and reading duration;

[0027] The data storage module is configured to store the policy version change records and user behavior records using a time-series database.

[0028] As a preferred method, when constructing the rule knowledge base, at least the following steps are included:

[0029] Perform IPC classification mapping on the first benchmark word set to generate a networked or tree-like storage structure with the main classification number as the root node;

[0030] Perform syntactic analysis on the policy texts corresponding to all classification numbers included under each node, and extract the normative verbs therein as strong association rules;

[0031] Analyze the citation relationship of the policy texts through the GNN graph neural network, and construct cross-classification association edges when there is a two-way citation between two nodes.

[0032] As a preferred method, when the feedback collection unit detects that the user ignores the push content of the same classification after a preset number of consecutive times, the following steps are performed:

[0033] Freeze the push rule node of this classification;

[0034] Send a confirmation request containing a list of alternative classifications to the user;

[0035] Reconstruct the association weights of the rule knowledge base according to the user confirmation result.

[0036] As a preferred method, when the matcher performs matching, the following steps are included:

[0037] Perform IPC classification mapping on the user's second benchmark word;

[0038] Generate an initial candidate set S 0 ={r 1 ,r 2 ,...,r n}, where r n represents the rule node in the rule knowledge base that has an intersection with the user's benchmark word;

[0039] Use a Bloom filter for S0 Perform deduplication and set the misjudgment rate p < 0.01;

[0040] For each candidate rule node r k ∈S 0 , calculate the precise matching degree:

[0041]

[0042] where,

[0043] V user is the user's reference word set, and V rule is the reference word set of the rule node;

[0044] α and β are empirical coefficients;

[0045] N is the total number of words in the rule knowledge base, and IDF(w) is the inverse document frequency;

[0046] Retain the rule nodes with a precise matching degree above the preset value to generate the second candidate set S 1 ;

[0047] For each r in the second candidate set k Construct a word vector matrix;

[0048] Convert the user's second description word set into a matrix

[0049] Convert the first description word set of the rule node into a matrix where both A and B are the dimensions of the word vector;

[0050] Calculate the attention weight matrix

[0051]

[0052] where, B T represents the transpose matrix of B;

[0053] Calculate the description matching degree:

[0054]

[0055] where, K = min(m, n);

[0056] Extract the depth feature of the rule node r k in the dependency graph, and generate a node vector through the graph embedding algorithm Node2Vec Calculate the LSTM hidden state of the user's historical behavior sequence

[0057] Calculate the logical correlation degree:

[0058]

[0059] Among them, PathCount is the number of paths from r k to the root node, and MaxPath is the maximum number of paths in the knowledge base;

[0060] Adaptively adjust each weight in the weighted matching degree according to the user terminal type;

[0061] Perform negative feedback suppression on the candidate results; if there are rule nodes in the historical rejection records, set the final matching degree

[0062]

[0063] Among them, M orig is the weighted matching degree.

[0064] In the second aspect of the present invention, a patent information push management method is provided, including the following steps:

[0065] Step 1: Establish a dynamic policy library, crawl the announcement data of patent offices in various countries through a distributed crawler cluster, use the Diff algorithm to identify policy version differences, and generate a first text data set with version tags;

[0066] Step 2: Construct a user portrait matrix, integrate the patent keywords input by the user, the IPC classifications concerned, the domain focus curve calculated based on the historical browsing path, and the second text data set obtained from the depth index of the comparative document access;

[0067] Step 3: Implement two-layer matching:

[0068] Coarse matching layer: Perform inverted index matching between the user portrait matrix and the policy library, and screen out the top 50 policy entries with the highest relevance;

[0069] Fine matching layer: Calculate for the screened policy entries:

[0070] The first matching condition, which describes the Jaccard similarity between the user's base word and the policy's base word;

[0071] The second matching condition, which calculates the cosine similarity between words through the attention mechanism;

[0072] The third matching condition, which establishes a dependency graph of parent and child rule nodes through the paragraph position encoding of the policy text;

[0073] Step 4: Generate a push decision:

[0074] Push is triggered if and only if the following conditions are met:

[0075] (The first matching condition ≥ threshold T1) ∧ (the second matching condition ≥ T2) ∨ (the third matching condition ≥ T3)

[0076] where T1, T2, and T3 are dynamically adjusted according to the user's historical feedback data;

[0077] Step Five: Establish a feedback loop:

[0078] Correlate each push result with the user's patent application behavior within the next 6 months to verify the effectiveness of the push, and mark and eliminate ineffective push strategies.

[0079] As a preferred method, when generating a push decision, it specifically includes the following steps:

[0080] Extract the first set of reference words composed of core nouns and the first set of descriptive words composed of restrictive phrases from the segmented first text data, and extract the second set of reference words composed of user feature nouns and the second set of descriptive words composed of associated adjectives from the second text data; where the set of reference words meets the preset word length rule, and the set of descriptive words is extracted using the sliding window algorithm for prefix and suffix phrases;

[0081] Perform IPC classification mapping on the user's second set of reference words;

[0082] Generate the initial candidate set S 0 ={r 1 ,r 2 ,...,r n}, where r n represents the rule node in the rule knowledge base that has an intersection with the user's reference words;

[0083] Use a Bloom filter to deduplicate S 0 , and set the false positive rate p < 0.01;

[0084] For each candidate rule node r k ∈S 0 , calculate the precise matching degree:

[0085]

[0086] where,

[0087] V user is the user's reference word set, and V rule is the rule node reference word set;

[0088] α and β are empirical coefficients;

[0089] N is the total number of words in the rule knowledge base, and IDF(w) is the inverse document frequency;

[0090] Retain the rule nodes with the precision match degree above the preset value, and generate the second candidate set S 1 ;

[0091] For each r in the second candidate set k Construct a word vector matrix;

[0092] Convert the user's second descriptor set into a matrix

[0093] Convert the first descriptor set of the rule node into a matrix where both A and B are the dimensions of the word vector;

[0094] Calculate the attention weight matrix

[0095]

[0096] where, B T represents the transpose matrix of B;

[0097] Calculate the description match degree:

[0098]

[0099] where, K = min(m, n);

[0100] Extract the depth feature of the rule node r k in the dependency graph, and generate a node vector through the graph embedding algorithm Node2Vec Calculate the LSTM hidden state of the user's historical behavior sequence

[0101] Calculate the logical correlation degree:

[0102]

[0103] where, PathCount is the number of paths from r k to the root node, and MaxPath is the maximum number of paths in the knowledge base;

[0104] Adaptively adjust each weight in the weighted match degree according to the user's terminal type;

[0105] Perform negative feedback suppression on the candidate results; if there is a rule node in the historical rejection record, set the final match degree

[0106]

[0107] where, M orig is the weighted match degree.

[0108] The above technical solutions of the present invention have the following advantages compared with the prior art:

[0109] The policy collection unit in the data collection module retrieves policy regulations and industrial dynamics data from official patent agency websites and industry databases in real time through API interfaces. This ensures that the system can obtain the latest patent-related policy information in a timely manner. For example, when the patent examination standards are updated or the patent subsidy policies are adjusted, this information can be incorporated into the system for subsequent processing immediately. Compared with traditional manual collection or regular batch collection methods, it greatly improves the timeliness and integrity of policy information, provides users with the latest policy guidelines, and helps enterprises and individuals adjust their patent strategies and business plans in a timely manner.

[0110] The user profiling unit extracts the second text dataset by parsing the historical behavior logs, project application materials, and collection records of users on the patent retrieval platform. This multi-dimensional data extraction method can comprehensively and accurately depict users' patent needs, interest preferences, and behavioral characteristics. For example, by analyzing users' historical search keywords and browsing records, it can accurately determine the technical fields and patent types that users are interested in; from project application materials, it can understand users' business directions and R & D priorities. This lays a solid foundation for subsequent precise push, making the pushed content more in line with users' actual needs and avoiding interference from invalid information.

[0111] The semantic processing module incorporates a word segmentation model of a bidirectional LSTM neural network, which is pre-trained based on a patent domain corpus and includes a part-of-speech tagging layer and a stop word filtering layer. For the complex professional terms and sentence structures in patent texts, it can perform accurate dynamic word segmentation and generate effective word vectors. For example, for a text like "Patent for a new display device based on quantum dot technology", the model can accurately identify key terms such as "quantum dot technology" and "new display device", and generate word vectors reflecting their semantic features. This helps to better understand the text meaning in subsequent matching and analysis, improving the matching accuracy and efficiency.

[0112] The feature extraction module extracts a set of benchmark words composed of core nouns and a set of descriptive words composed of restrictive phrases from the segmented text data. It conducts targeted feature extraction for policy texts and user texts respectively. The set of benchmark words meets the preset word length rule, and the set of descriptive words uses a sliding window algorithm to extract prefix and suffix phrases. This method can accurately extract the key feature information in the text. For example, from the policy text "Policy on patent application fee reduction for high-tech enterprises", it accurately extracts the core noun "high-tech enterprises" as the benchmark word and "patent application fee reduction" as the descriptive word. In the process of matching users with policies, these accurately extracted features can more effectively measure the correlation between the two, improving the accuracy of matching.

[0113] The rule knowledge module stores a push rule tree with the IPC classification number as the root node. By performing IPC classification mapping on the first set of benchmark words, a networked or tree-like storage structure is generated. Syntactic analysis is carried out on the policy text to extract normative verbs as strong association rules, and the citation relationship of the policy text is analyzed using the GNN graph neural network to construct cross-classification association edges. This enables the rule knowledge base to comprehensively and systematically cover various policy rules and logical relationships in the patent field. For example, in patent policies involving multiple cross-cutting technical fields, relevant policy rules under different classifications can be accurately associated through cross-classification association edges, providing strong support for precise push in complex situations.

[0114] The multi-level matching strategy improves the matching effect: The matcher executes a multi-level matching strategy. First, it performs the first match, matching the user's second set of benchmark words with the rule nodes to generate a candidate rule set. Then, it calculates the weighted matching degree in the candidate rule set, including the benchmark matching degree, description matching degree, and logical association degree. This multi-level matching strategy comprehensively considers the relevance between the user and the policy text at multiple levels. For example, when judging the matching degree of a corporate user with a "patent subsidy policy for intelligent transportation systems", it not only considers the matching situation of the benchmark words related to "intelligent transportation" that the user is concerned about with the policy benchmark words (benchmark matching degree), but also takes into account the matching degree of descriptive words such as subsidy conditions and scope of application in the policy description with the user characteristics (description matching degree), as well as the logical association position of this policy in the entire rule knowledge base (logical association degree). Through this comprehensive matching method, the fit between the push result and the user's needs is greatly improved, and patent information that meets the user's needs is accurately pushed.

[0115] The feedback collection unit collects the click-through rate and reading duration of the user after each push. And when it detects that the user ignores the push content of the same classification after a preset number of consecutive times, it will perform a series of operations, such as freezing the push rule nodes of this classification, sending a confirmation request containing a list of alternative classifications to the user, and reconstructing the association weights of the rule knowledge base according to the user's confirmation result. This comprehensive feedback collection method can timely understand the actual reaction of the user to the push content. It not only pays attention to whether the user clicks on the push information, but also deeply analyzes the user's reading behavior. For classifications that the user is not interested in for a long time, the push strategy can be adjusted in a timely manner to avoid wasting resources.

[0116] When the model optimization module has continuously executed a preset number of pushes and the target click-through rate and target reading duration of the user do not reach the preset threshold, it will perform a series of optimization operations, such as adjusting the word length rule threshold of the benchmark word set, calculating the attention weight distribution for describing word similarity, and performing incremental training on the LSTM word segmentation model based on the user's click-through rate and reading duration. Through these optimization operations, the performance of the system and the accuracy of the push can be continuously improved. For example, over time and with changes in user behavior, by adjusting the word length rule threshold of the benchmark word set, it can better adapt to changes in the focus of information attention of users at different stages; performing incremental training on the LSTM word segmentation model enables the model to continuously learn new user behavior patterns and text features, thereby continuously optimizing semantic processing and matching effects and improving the overall service quality of the system. Brief Description of the Drawings

[0117] Figure 1 It is a structural block diagram of the patent information push management system provided by the present invention. Detailed Embodiments

[0118] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0119] In the first aspect of the embodiments of the present disclosure, a patent information push management system is provided, as Figure 1 shown, including:

[0120] A data collection module, configured to include a policy collection unit and a user portrait unit;

[0121] The policy collection unit captures policy regulations and industrial dynamic data of official patent agency websites and industry databases in real time through an API interface to form a first text data set;

[0122] The user portrait unit extracts a second text data set by parsing the historical behavior logs, project application materials, and collection records of users on the patent retrieval platform;

[0123] A semantic processing module, configured to have a word segmentation model of a bidirectional LSTM neural network, perform dynamic word segmentation processing on the first and second text data sets, and generate word vectors; the word segmentation model is pre-trained based on a patent domain corpus and includes a part-of-speech tagging layer and a stop word filtering layer;

[0124] The feature extraction module extracts a first set of benchmark words composed of core nouns and a first set of descriptive words composed of restrictive phrases from the first text data after word segmentation, and extracts a second set of benchmark words composed of user feature nouns and a second set of descriptive words composed of associated adjectives from the second text data; wherein the set of benchmark words satisfies a preset word length rule, and the set of descriptive words extracts prefix and suffix phrases using a sliding window algorithm.

[0125] The rule knowledge module stores a push rule tree with the IPC classification number as the root node. The push rule tree includes multiple rule nodes, and each of the rule nodes includes:

[0126] The first matching condition describes the Jaccard similarity between the user's benchmark words and the policy benchmark words;

[0127] The second matching condition calculates the cosine similarity between words through an attention mechanism;

[0128] The third matching condition establishes a dependency relationship graph between parent and child rule nodes through the paragraph position encoding of the policy text;

[0129] The matcher is configured to execute a multi-level matching strategy, including:

[0130] The first matching matches the user's second set of benchmark words with the rule nodes to generate a candidate rule set;

[0131] The second matching calculates the weighted matching degree in the candidate rule set. The weighted matching degree includes a benchmark matching degree, a description matching degree, and a logical association degree with weights set respectively;

[0132] Among them,

[0133] The benchmark degree = the number of matching benchmark words / the total number of benchmark words;

[0134] The description matching degree = Σ(description word similarity × TF-IDF weight);

[0135] The logical integrity degree = the connectivity degree of the matching rule nodes in the dependency relationship graph;

[0136] The feedback optimization module includes a feedback collection unit and a model tuning unit. The feedback collection unit collects the click-through rate and reading duration of the user after each push;

[0137] The model tuning module is configured to perform the following optimizations when the target click-through rate and target reading duration of the user do not reach the preset threshold after continuously executing the push for a preset number of times:

[0138] Adjust the word length rule threshold of the set of benchmark words;

[0139] Calculate the attention weight distribution of the description word similarity;

[0140] Perform incremental training on the LSTM word segmentation model based on the user's click-through rate and reading duration;

[0141] A data storage module, configured to store policy version change records and user behavior records using a time-series database.

[0142] As a preferred method, when constructing the rule knowledge base, at least the following steps are included:

[0143] Perform IPC classification mapping on the first set of reference words to generate a networked or tree-like storage structure with the main classification number as the root node;

[0144] Perform syntactic analysis on the policy texts corresponding to all classification numbers included under each node, and extract the prescriptive verbs therein as strong association rules;

[0145] Analyze the citation relationships of policy texts through a GNN graph neural network, and construct cross-classification association edges when there is a two-way citation between two nodes.

[0146] As a preferred method, when the feedback collection unit detects that the user ignores the push content of the same classification after a preset number of consecutive times, it performs the following steps:

[0147] Freeze the push rule node of this classification;

[0148] Send a confirmation request containing a list of alternative classifications to the user;

[0149] Reconstruct the association weights of the rule knowledge base according to the user's confirmation result.

[0150] As a preferred method, when the matcher performs matching, it includes the following steps:

[0151] Perform IPC classification mapping on the user's second set of reference words;

[0152] Generate an initial candidate set S 0 ={r 1 ,r 2 ,...,r n}, where r n represents the rule node in the rule knowledge base that has an intersection with the user's reference words;

[0153] Use a Bloom filter to perform deduplication processing on S 0 , and set the false positive rate p < 0.01;

[0154] For each candidate rule node r k ∈S 0 , calculate the precise matching degree:

[0155]

[0156] Among them,

[0157] V user is the user's benchmark word set, and V rule is the rule node benchmark word set;

[0158] α and β are empirical coefficients;

[0159] N is the total number of words in the rule knowledge base, and IDF(w) is the inverse document frequency;

[0160] Retain the rule nodes with a precise matching degree above the preset value to generate the second candidate set S 1 ;

[0161] For each r in the second candidate set k Construct a word vector matrix;

[0162] Convert the user's second description word set into a matrix

[0163] Convert the first description word set of the rule node into a matrix where both A and B are the dimensions of the word vector;

[0164] Calculate the attention weight matrix

[0165]

[0166] Among them, B T represents the transpose matrix of B;

[0167] Calculate the description matching degree:

[0168]

[0169] where K = min(m, n);

[0170] Extract the depth feature of the rule node r k in the dependency graph, and generate a node vector through the graph embedding algorithm Node2Vec Calculate the LSTM hidden state of the user's historical behavior sequence

[0171] Calculate the logical correlation degree:

[0172]

[0173] where PathCount is the number of paths from rk to the root node, and MaxPath is the maximum number of paths in the knowledge base;

[0174] Adaptively adjust each weight in the weighted matching degree according to the user's terminal type;

[0175] Perform negative feedback suppression on the candidate results; if there are rule nodes in the historical rejection records, set the final matching degree

[0176]

[0177] where M orig is the weighted matching degree.

[0178] In the second aspect of the embodiments of the present disclosure, a method for patent information push management is provided, including the following steps:

[0179] Step 1: Establish a dynamic policy library, crawl the announcement data of patent offices in various countries through a distributed crawler cluster, use the Diff algorithm to identify policy version differences, and generate a first text dataset with version tags;

[0180] Step 2: Construct a user portrait matrix, integrate the patent keywords input by the user, the IPC classifications concerned, the domain focus curve calculated based on the historical browsing path, and the second text dataset obtained from the depth index of the comparison document access;

[0181] Step 3: Implement two-layer matching:

[0182] Coarse matching layer: Perform inverted index matching between the user portrait matrix and the policy library, and screen out the top 50 policy entries with the highest relevance;

[0183] Fine matching layer: Calculate for the screened policy entries:

[0184] The first matching condition, which describes the Jaccard similarity between the user's reference word and the policy reference word;

[0185] The second matching condition, which calculates the cosine similarity between words through the attention mechanism;

[0186] The third matching condition, which establishes a dependency graph of parent-child rule nodes through the paragraph position encoding of the policy text;

[0187] Step 4: Generate a push decision:

[0188] Push is triggered if and only if the following conditions are met:

[0189] (The first matching condition ≥ threshold T1) ∧ (the second matching condition ≥ T2) ∨ (the third matching condition ≥ T3)

[0190] where T1, T2, and T3 are dynamically adjusted according to the user's historical feedback data;

[0191] Step 5: Establish a feedback loop:

[0192] Associate the results of each push with the user's patent application behavior within the next 6 months, verify the effectiveness of the push, and mark and eliminate ineffective push strategies.

[0193] As a preferred method, when generating a push decision, it specifically includes the following steps:

[0194] Extract the first set of reference words composed of core nouns and the first set of descriptive words composed of restrictive phrases from the first text data after word segmentation, and extract the second set of reference words composed of user feature nouns and the second set of descriptive words composed of associated adjectives from the second text data; among them, the set of reference words meets the preset word length rule, and the set of descriptive words uses the sliding window algorithm to extract prefix and suffix phrases.

[0195] Perform IPC classification mapping on the user's second set of reference words.

[0196] Generate an initial candidate set S 0 ={r 1 ,r 2 ,...,r n}, where r n represents the rule node in the rule knowledge base that has an intersection with the user's reference words.

[0197] Use a Bloom filter to perform deduplication processing on S 0 , and set the false positive rate p < 0.01.

[0198] For each candidate rule node r k ∈ S 0 , calculate the precise matching degree:

[0199]

[0200] Among them,

[0201] V user is the user's reference word set, and V rule is the reference word set of the rule node;

[0202] α and β are empirical coefficients;

[0203] N is the total number of words in the rule knowledge base, and IDF(w) is the inverse document frequency;

[0204] Retain the rule nodes with a precise matching degree above the preset value to generate a second candidate set S 1 ;

[0205] Construct a word vector matrix for each r k in the second candidate set;

[0206] Convert the user's second set of descriptive words into a matrix

[0207] Convert the first descriptor set of the rule node into a matrix where both A and B are the dimensions of the word vectors;

[0208] Calculate the attention weight matrix

[0209]

[0210] where, B T represents the transpose matrix of B;

[0211] Calculate the description matching degree:

[0212]

[0213] where, K = min(m, n);

[0214] Extract the depth feature of the rule node r k in the dependency graph, and generate the node vector through the graph embedding algorithm Node2Vec Calculate the LSTM hidden state of the user's historical behavior sequence

[0215] Calculate the logical correlation degree:

[0216]

[0217] where, PathCount is the number of paths from r k to the root node, and MaxPath is the maximum number of paths in the knowledge base;

[0218] Adaptively adjust each weight in the weighted matching degree according to the user terminal type;

[0219] Perform negative feedback suppression on the candidate results; if there is a rule node in the historical rejection record, set the final matching degree

[0220]

[0221] where, M orig is the weighted matching degree.

[0222] The policy collection unit in the data collection module captures the policy regulations and industrial dynamic data of the official patent agency website and industry databases in real time through the API interface. This ensures that the system can obtain the latest patent-related policy information in a timely manner. For example, when the patent examination standards are updated or the patent subsidy policies are adjusted, these information can be incorporated into the system for subsequent processing in the first time. Compared with the traditional manual collection or regular batch collection methods, it greatly improves the timeliness and integrity of the policy information, provides the latest policy guidance for users, and helps enterprises and individuals adjust their patent strategies and business plans in a timely manner.

[0223] The user profile unit extracts the second text dataset by parsing the user's historical behavior logs, project application materials, and collection records on the patent retrieval platform. This multi-dimensional data extraction method can comprehensively and accurately depict the user's patent needs, interest preferences, and behavior characteristics. For example, by analyzing the user's historical search keywords and browsing records, it is possible to accurately judge the technical fields, patent types, etc. that the user is interested in; the user's business direction and R & D focus can be understood from the project application materials. This lays a solid foundation for subsequent precise push, making the pushed content more in line with the actual needs of users and avoiding interference from invalid information.

[0224] The semantic processing module incorporates a word segmentation model of a bidirectional LSTM neural network, and this model is pre-trained based on a patent-domain corpus, including a part-of-speech tagging layer and a stop-word filtering layer. For the complex professional terms and sentence structures in patent texts, it can perform accurate dynamic word segmentation processing and generate effective word vectors. For example, for a text like "Patent for a new display device based on quantum dot technology", the model can accurately identify key terms such as "quantum dot technology" and "new display device", and generate word vectors reflecting their semantic features. This helps to better understand the text meaning in subsequent matching and analysis, improving the matching accuracy and efficiency.

[0225] The feature extraction module extracts a set of benchmark words composed of core nouns and a set of descriptive words composed of restrictive phrases from the segmented text data. It conducts targeted feature extraction for policy texts and user texts respectively, and the set of benchmark words meets the preset word length rule, while the set of descriptive words uses a sliding window algorithm to extract prefix and suffix phrases. This method can accurately extract the key feature information in the text. For example, from the policy text "Policy on the reduction of patent application fees for high-tech enterprises", it can accurately extract the core noun "high-tech enterprises" as a benchmark word and "reduction of patent application fees" as a descriptive word. In the process of matching users with policies, these precisely extracted features can more effectively measure the correlation between the two, improving the accuracy of matching.

[0226] The rule knowledge module stores a push rule tree with the IPC classification number as the root node. By performing IPC classification mapping on the first set of benchmark words, it generates a networked or tree-like storage structure, and conducts syntactic analysis on the policy text to extract normative verbs as strong association rules, as well as uses a GNN graph neural network to analyze the citation relationships in the policy text to construct cross-classification association edges. This enables the rule knowledge base to comprehensively and systematically cover various policy rules and logical relationships in the patent field. For example, in patent policies involving intersections of multiple technical fields, relevant policy rules under different classifications can be accurately associated through cross-classification association edges, providing strong support for precise push in complex situations.

[0227] Multi-level matching strategy improves matching effect: The matcher executes a multi-level matching strategy. First, it performs the first-level matching, matching the user's second reference word set with the rule nodes to generate a candidate rule set. Then, it calculates the weighted matching degrees in the candidate rule set, including the reference matching degree, the description matching degree, and the logical association degree. This multi-level matching strategy comprehensively considers the relevance between the user and the policy text at multiple levels. For example, when judging the matching degree of an enterprise user with a "patent subsidy policy for intelligent transportation systems", it not only considers the matching situation of the reference words related to "intelligent transportation" that the user is concerned about with the policy reference words (reference matching degree), but also examines the matching degree between the descriptive words such as subsidy conditions and applicable scopes in the policy description and the user characteristics (description matching degree), as well as the logical association position of this policy in the entire rule knowledge base (logical association degree). Through this comprehensive matching method, the fit between the push result and the user's needs is greatly improved, and patent information that meets the user's needs is accurately pushed.

[0228] The feedback collection unit collects the user's click-through rate and reading duration after each push. And when it detects that the user ignores the push content of the same classification after a preset number of consecutive times, it will perform a series of operations, such as freezing the push rule nodes of this classification, sending a confirmation request containing a list of alternative classifications to the user, and reconstructing the association weights of the rule knowledge base according to the user's confirmation result. This comprehensive feedback collection method can timely understand the actual reaction of the user to the push content. It not only focuses on whether the user clicks on the push information, but also deeply analyzes the user's reading behavior. For classifications that the user is not interested in for a long time, it can timely adjust the push strategy to avoid wasting resources.

[0229] After the model tuning module continuously executes the push for a preset number of times, if the user's target click-through rate and target reading duration do not reach the preset thresholds, it will perform a series of optimization operations, such as adjusting the word length rule threshold of the reference word set, calculating the attention weight distribution for the description word similarity, and performing incremental training on the LSTM word segmentation model based on the user's click-through rate and reading duration. Through these optimization operations, the performance of the system and the accuracy of the push can be continuously improved. For example, as time goes by and the user's behavior changes, by adjusting the word length rule threshold of the reference word set, it can better adapt to the changes in the user's focus on information at different stages; performing incremental training on the LSTM word segmentation model enables the model to continuously learn new user behavior patterns and text features, thereby continuously optimizing semantic processing and matching effects and improving the overall service quality of the system.

[0230] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion thereof, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based device that performs the specified function or action, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A patent information push management system, characterized in that: include: A data collection module is configured to include a policy collection unit and a user profiling unit; The policy collection unit captures policy and regulations and industry dynamic data from official patent agency websites and industry databases in real time through an API interface to form a first text data set; The user portrait unit extracts a second text data set by analyzing the user's historical behavior log, project application materials and collection records on the patent search platform; The semantic processing module is configured as a word segmentation model with a built-in bidirectional LSTM neural network, which performs dynamic word segmentation processing on the first and second text data sets and generates word vectors; the word segmentation model is pre-trained based on a patent field corpus and includes a part-of-speech tagging layer and a stop word filtering layer; A feature extraction module extracts a first reference word set consisting of core nouns and a first description word set consisting of restrictive phrases from the first text data after word segmentation, and extracts a second reference word set consisting of user characteristic nouns and a second description word set consisting of related adjectives from the second text data; wherein the reference word set meets a preset word length rule, and the description word set uses a sliding window algorithm to extract prefix and suffix phrases; The rule knowledge module stores a push rule tree with the IPC classification number as the root node, wherein the push rule tree includes a plurality of rule nodes, wherein each of the rule nodes includes: The first matching condition describes the Jaccard similarity between the user benchmark word and the policy benchmark word; The second matching condition is to calculate the cosine similarity between words through the attention mechanism; The third matching condition is to build a dependency graph of parent-child rule nodes through the paragraph position encoding of the policy text; Matchers are configured to perform multi-level matching strategies, including: First matching, matching the user's second benchmark word set with the rule node to generate a candidate rule set; The second matching is to calculate the weighted matching degree in the candidate rule set, wherein the weighted matching degree includes a reference matching degree, a description matching degree, and a logical association degree, each of which has a weight set; in, The benchmark degree = matching benchmark word / total number of benchmark words; The description matching degree = Σ(description word similarity × TF-IDF weight); The logic completeness = the connectivity of the matching rule nodes in the dependency graph; A feedback optimization module, including a feedback collection unit and a model tuning unit, wherein the feedback collection unit collects the click rate and reading time of the user after each push; The model tuning module is configured to perform the following optimization when the user's target click rate and target reading market do not reach the preset threshold after the preset number of pushes are executed continuously: Adjust the word length rule threshold of the benchmark word set; Calculate the attention weight distribution of description word similarity; Perform incremental training on the LSTM word segmentation model based on the user's click rate and reading time; The data storage module is configured to use a time series database to store policy version change records and user behavior records.

2. A patent information push management system according to claim 1, characterized in that: When constructing the rule knowledge base, at least the following steps are included: Performing IPC classification mapping on the first reference word set to generate a mesh or tree storage structure with the main classification number as the root node; Perform syntactic analysis on the policy texts corresponding to all the classification numbers under each node, and extract the normative verbs as strong association rules; The GNN graph neural network is used to analyze the citation relationship of the policy text, and a cross-classification association edge is constructed when there is a bidirectional reference between two nodes.

3. A patent information push management system according to claim 2, characterized in that: The feedback collection unit, when detecting that the user ignores the push content of the same category after a preset number of consecutive times, executes the following steps: Freeze the push rule node of this category; Sending a confirmation request to the user containing a list of alternative categories; The association weight of the rule knowledge base is reconstructed according to the user confirmation result.

4. A patent information push management system according to claim 3, characterized in that: The matching process includes the following steps: Perform IPC classification mapping on the user's second benchmark word; Generate the initial candidate set S0 = {r1, r2, ..., r n }, where r n Indicates the rule nodes in the rule knowledge base that have intersections with the user's benchmark words; Bloom filter is used to remove duplicates from S0, and the false positive rate is set to p<0.01; For each candidate rule node r k ∈S0, calculate the exact match: in, V user is the user benchmark word set, V rule is the benchmark word set for rule nodes; α and β are empirical coefficients; N is the total number of words in the rule knowledge base, and IDF(w) is the inverse document frequency; The rule nodes with an exact matching degree above a preset value are retained to generate a second candidate set S1; Each r in the second candidate set k Construct word vector matrix; Convert the user's second description word set into a matrix Convert the first description word set of the rule node into a matrix Among them, A and B are word vector dimensions; Calculate the attention weight matrix Among them, B T represents the transposed matrix of B; Calculate description matching: Where, K = min(m,n); Extract rule node r k Deep features in the dependency graph, generating node vectors through the graph embedding algorithm Node2Vec Calculate the LSTM hidden state of the user's historical behavior sequence Calculate logical association: Among them, PathCount is the number of paths from rk to the root node, and MaxPath is the maximum number of paths in the knowledge base; Adaptively adjusting each weight in the weighted matching degree according to the user terminal type; Negative feedback suppression is performed on candidate results; if there is a rule node in the historical rejection record, the final matching degree is set Among them, M orig is the weighted matching degree.

5. A patent information push management method, characterized in that: The steps include: Step 1: Establish a dynamic policy library, crawl the announcement data of patent offices in various countries through a distributed crawler cluster, use the Diff algorithm to identify policy version differences, and generate the first text data set with version labels; Step 2: Build a user portrait matrix, integrate the patent keywords entered by the user, the IPC classification of interest, the field focus curve calculated based on the historical browsing path, and the comparison document review depth index to obtain the second text data set; Step 3: Implement double-layer matching: Coarse matching layer: Perform inverted index matching between the user portrait matrix and the policy library to filter out the top 50 policy items with the highest relevance. Fine matching layer: Calculate the filtered policy items: The first matching condition describes the Jaccard similarity between the user benchmark word and the policy benchmark word; The second matching condition is to calculate the cosine similarity between words through the attention mechanism; The third matching condition is to build a dependency graph of parent-child rule nodes through the paragraph position encoding of the policy text; Step 4: Generate push decision: The push is triggered if and only if the following conditions are met: (first matching condition ≥ threshold T1) ∧ (second matching condition ≥ T2) ∨ (third matching condition ≥ T3) where T1, T2, and T3 are dynamically adjusted based on historical user feedback data; Step 5: Establish a feedback loop: The results of each push are correlated with the user’s patent application behavior in the following 6 months to verify the effectiveness of the push and mark and eliminate invalid push strategies.

6. A patent information push management method according to claim 5, characterized in that: When generating a push decision, the following steps are included: Extracting a first reference word set consisting of core nouns and a first description word set consisting of restrictive phrases from the first text data after word segmentation, and extracting a second reference word set consisting of user characteristic nouns and a second description word set consisting of related adjectives from the second text data; wherein the reference word set meets the preset word length rule, and the description word set uses a sliding window algorithm to extract prefix and suffix phrases; Perform IPC classification mapping on the user's second benchmark word; Generate the initial candidate set S0 = {r1, r2, ..., r n }, where r n Indicates the rule nodes in the rule knowledge base that have intersections with the user's benchmark words; Bloom filter is used to remove duplicates from S0, and the false positive rate is set to p<0.01; For each candidate rule node r k ∈S0, calculate the exact match: in, V user is the user benchmark word set, V rule is the benchmark word set for rule nodes; α and β are empirical coefficients; N is the total number of words in the rule knowledge base, and IDF(w) is the inverse document frequency; The rule nodes with an exact matching degree above a preset value are retained to generate a second candidate set S1; Each r in the second candidate set k Construct word vector matrix; Convert the user's second description word set into a matrix Convert the first description word set of the rule node into a matrix Among them, A and B are word vector dimensions; Calculate the attention weight matrix Among them, B T represents the transposed matrix of B; Calculate description matching: Where, K = min(m,n); Extract the deep features of the regular node rk in the dependency graph and generate the node vector through the graph embedding algorithm Node2Vec Calculate the LSTM hidden state of the user's historical behavior sequence Calculate logical association: Among them, PathCount is the number of paths from rk to the root node, and MaxPath is the maximum number of paths in the knowledge base; Adaptively adjusting each weight in the weighted matching degree according to the user terminal type; Negative feedback suppression is performed on candidate results; if there is a rule node in the historical rejection record, the final matching degree is set Among them, M orig is the weighted matching degree.

Citation Information

Patent Citations

  • Patent data analysis system

    CN109492117A

  • Patent recommendation method based on Transform encoder and regularization strategy

    CN117370648A

  • Deep learning-based science and technology project innovation potential estimation method and device

    CN118333033A

  • Policy information integration pushing system based on information accuracy

    CN118467817A

  • KR20230048198A

Cited By

  • Information message matching and pushing method and system based on big data driving

    CN120277271A