Intelligent model-based domain high-revolution patent prediction method and system

By combining SVM and LDA topic models with analytic hierarchy process and entropy weight method, a disruptive technology measurement model is constructed, which solves the problem of insufficient accuracy and depth in the identification of disruptive technologies in the existing technology, and realizes efficient identification and patent transformation of potential disruptive technologies.

CN117851901BActive Publication Date: 2026-07-21WUHAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2023-12-28
Publication Date
2026-07-21

Smart Images

  • Figure CN117851901B_ABST
    Figure CN117851901B_ABST
Patent Text Reader

Abstract

The application provides a kind of field high subversion patent prediction method and system based on intelligent model, comprising: obtaining the subversive technology prediction task released by user;Patent dataset is constructed, and the sub-datasets of each category are obtained by classifying patent dataset using SVM classification algorithm, LDA topic model is constructed, and the technology topics of each category of sub-dataset are extracted using LDA topic model;Subversive technology measure scoring model is constructed, and the extracted technology topics are comprehensively scored, the technology topics with top comprehensive score are screened out, and the corresponding patent text is sorted;Key words in patent text are extracted, and the topic words of each technology topic corresponding category are counted, the similarity between the key words of each patent text and the topic words of each corresponding category is calculated, and the patent text with the maximum similarity to category topic word and the technical field thereof are selected as the prediction result and output.The application improves the accuracy of potential subversive technology measure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of technological advancement assessment, specifically relating to a method and system for predicting highly disruptive patents in a field based on intelligent models. Background Technology

[0002] Disruptive technologies, as a crucial driver of technological innovation, require careful consideration of their development direction. Identifying potential disruptive technologies, disciplines, industries, and sectors is of significant theoretical and practical importance for enhancing national comprehensive scientific and technological strength, improving enterprise core competitiveness, and boosting the effectiveness of technology transfer and commercialization in universities. Currently, the identification of cutting-edge technologies both domestically and internationally primarily relies on qualitative methods such as the Delphi method based on expert opinions, technology roadmaps, and scenario analysis. These approaches mainly focus on identifying emerging technologies, with limited research on disruptive technologies. For example, patent literature CN... Patent citation analysis method 107220320B discloses a method for identifying emerging technologies based on patent citations. The method comprises: S1 characterizing a patent citation database; S2 grouping each patent published in year T+1 according to its main classification number, denoting the group as Gy; S3 labeling Gy as a new technology group if the main classification number was newly established in year T+1, otherwise labeling it as a non-new technology group; S4 clustering all patents in year T based on their patent citation feature vectors, denoting the cluster as Cx; and S5 calculating the relationship between any C′x in year T and all patents in group Gy in year T+1. S6 Find the group G′y with the highest coupling degree with patent C′x; S7 If G′y is an emerging technology group, mark cluster C′x as an emerging technology, otherwise mark it as a non-emerging technology; S8 Jump to step S4 until all clusters Cx in year T have been marked; S9 Jump to step S1 until all patents except the one with the largest year have been clustered and labeled; S10 Train a classifier using labeled data; S11 Use the classifier to determine whether the clusters based on patent citation feature vectors are emerging technologies.

[0003] Furthermore, patent document CN 114969251 A discloses a method and apparatus for identifying emerging technologies based on a large-scale corpus. The method includes: determining a research field and constructing a candidate document set, and extracting keywords from the candidate document set to obtain a candidate keyword dataset; filtering the candidate keyword dataset based on the number of candidate documents and the relevant information of the keywords to obtain a candidate keyword filter set; calculating the emerging score of each keyword in the candidate keyword filter set; screening the candidate keyword filter set based on the emerging score of each keyword and a set emerging score threshold to obtain a candidate emerging technology keyword dataset; and processing the candidate emerging technology keyword dataset using a dynamic backtracking method to obtain a target emerging technology keyword dataset, thereby improving the accuracy of emerging technology identification.

[0004] While the aforementioned methods for identifying cutting-edge technologies are simple and highly operable, their identification methods are relatively singular and prone to misidentification, resulting in low accuracy of predictions. Some inventors have attempted quantitative analysis using econometric models, data mining, and statistics, but these methods suffer from shortcomings such as insufficient analytical perspectives and the need for algorithmic model optimization. Furthermore, most inventors remain at the technology identification stage with shallow analytical depth and limited research on technology measurement. Overall, a systematic, universal, and highly operable methodological framework for measuring disruptive technologies has not yet been established. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing a method for predicting highly disruptive patents in a domain based on intelligent models. This method is effective in identifying disruptive technologies in complex and uncertain environments and improves the accuracy of measuring potential disruptive technologies.

[0006] To address the aforementioned technical problems, this invention provides a method for predicting highly disruptive patents in a domain based on intelligent models, comprising the following steps:

[0007] Step 1: Obtain disruptive technology prediction tasks posted by users;

[0008] Step 2: Based on the technical field to which the disruptive technology prediction task belongs, obtain relevant patents from the patent database and construct a patent dataset. Use the SVM classification algorithm to classify the patent dataset to obtain sub-datasets of each category. Construct an LDA topic model, use the LDA topic model to extract technical topics from each category of sub-datasets, and output the technical topic extraction results.

[0009] Step 3: Construct a disruptive technology measurement and scoring model. Use the disruptive technology measurement and scoring model to comprehensively score the extracted technology topics, select the top-ranked technology topics, and organize their corresponding patent texts.

[0010] Step 4: Extract keywords from the patent texts selected in Step 3, and count the subject words corresponding to each technical topic selected in Step 3. Calculate the similarity between the keywords of each patent text and the subject words of each corresponding category. Select the patent text with the highest similarity to the subject words of the category and its corresponding technical field as the prediction result for output.

[0011] Furthermore, step 2 involves classifying the patent dataset using the SVM classification algorithm as follows: patents related to the technical subject area of ​​the disruptive technology prediction task are acquired and constructed into a patent dataset. The technical basic category classification standard file corresponding to the technical subject area is called, and the SVM classification algorithm is used to classify the patent dataset according to the technical basic category classification standard file to obtain sub-datasets of each category.

[0012] Furthermore, the method for constructing the disruptive technology measurement and scoring model is as follows:

[0013] First, measurement indicators are constructed from multiple dimensions based on the results of technology topic extraction;

[0014] Secondly, the importance of each metric for each technical topic is evaluated based on the analytic hierarchy process (AHP), resulting in the priority weight W for each metric. H ;

[0015] Next, the contribution of each measurement index is evaluated based on the entropy weight method, and the weight W of each measurement index is obtained. j ;

[0016] Then, the weight W is calculated based on the principle of minimizing the overall deviation. H Weight W j The combined weight vector w;

[0017] Finally, based on the combined weight vector w, each measurement index is assigned a corresponding weight, and then the weighted sum of each measurement index is performed to obtain the comprehensive score of each technical topic.

[0018] Furthermore, measurement indicators are constructed from four dimensions: technological integration, technological innovation, technological importance, and technological breakthrough.

[0019] Furthermore, technological convergence is measured from proximity centrality and the average number of IPC categories, where proximity centrality is calculated using the following formula:

[0020]

[0021] AA i d represents the proximity centrality of technical topic i. ij Let N represent the shortest distance between technology topics i and j, and N represent the total number of technology topics.

[0022] Technological innovation is measured using the structural void index. Constraint degree, grade degree, and effective size are typical indicators of the structural void index, and their formulas are shown below:

[0023]

[0024] In the formula, C ijRepresents the degree of constraint; node q is a common adjacency of nodes i and j; P ij P represents the weight proportion of node j among all the adjacent nodes of node i. iq P represents the weight ratio of node q within node i. ij This represents the weight ratio of node q among the adjacent nodes of node j;

[0025]

[0026] In the formula, YI i The degree of constraint is represented by N; N represents the individual network size of technical subject i; C / N represents the mean of the degree of constraint of each technical subject; and C represents the sum of the degrees of constraint.

[0027]

[0028] In the formula, YX i represents the effective size; j represents the adjacent nodes of node i, and q represents the common adjacent nodes of nodes i and j; P iq P represents the weight ratio of node q within node i. ij This represents the weight ratio of node q among the neighboring nodes of node j.

[0029] Furthermore, technological importance is quantified using degree centrality and proximity centrality to determine technological subject rights. The formula for calculating degree centrality is:

[0030] DC i =K i / (g-1)

[0031] In the formula, DC i K represents the degree centrality of technical topic i in the network. i The degree of technical topic i is represented by g, and the network size is represented by g.

[0032] The formula for calculating proximity centrality is:

[0033]

[0034] In the formula, CC i d represents the proximity centrality of technical topic i in the network. ij Let represent the shortest distance between technology topic i and technology topic j, and n represent the degree of node i;

[0035] Technological breakthroughs are determined using the K-means method for detecting technological anomalies. The calculation formula is as follows:

[0036]

[0037] In the formula, dist(x,y)i ) represents technology topic x and technology topic y i The distance between them, the technical topic y i It is one of the K nearest neighbor technical topics x, where K represents the number of nearest neighbors of technical topic x.

[0038] Furthermore, the priority weight W is obtained based on the analytic hierarchy process. H The method is as follows:

[0039] The first step is to construct pairwise comparison judgment matrices; establish judgment matrix X based on the relative importance of each indicator; then use the eigenvalue method to calculate the largest eigenvalue λ of judgment matrix X. max And the corresponding feature vector N, with m indicators, the judgment matrix X = (b ij ) m×m b in the matrix ij This indicates the degree of importance of indicator i compared to indicator j;

[0040] The second step is to perform a consistency check; the judgment matrix must have consistency and transitivity, and the consistency index CI = (λ) / (λ) max -m) / m-1;

[0041] The third step is to calculate the consistency ratio. If the consistency ratio CR < 0.1, the consistency of the judgment matrix is ​​considered acceptable; otherwise, the judgment matrix needs to be corrected. CR = CI / RI, where RI is the random consistency index.

[0042] Step 4: Find the largest eigenvalue and the corresponding eigenvector of the judgment matrix; when the eigenvalue is n, the corresponding eigenvector is... Where k is a non-zero constant;

[0043] The fifth step is normalization; the obtained eigenvectors are normalized to obtain the weights W. H =(w1,w2,...w n ) T .

[0044] Furthermore, the weight W is obtained based on the entropy weight method. j The method is as follows:

[0045] The first step is to standardize the indicator data, and the calculation formula is as follows:

[0046]

[0047] In the formula x i As the initial value, y ij These are the standardized values ​​of each indicator.

[0048] The second step is to normalize the indicator data; after normalizing the data, the probability value is calculated using the following formula:

[0049]

[0050] The third step is to calculate the information entropy value of each indicator; according to the definition of information entropy in information theory, information entropy E j The formula is shown below:

[0051]

[0052] Step 4: Determine the weight of each indicator, using the following formula:

[0053]

[0054] Furthermore, step 4 includes the following steps:

[0055] The selected patent text abstracts are processed to extract the keyword set S;

[0056] The LDA model is used to extract the set T of topic word vectors corresponding to the categories of the technical topics selected in step 3, and the TF-IDF value of the topic words is used as the contribution of the topic words to the categories.

[0057] The extracted keywords are represented by word vectors. Assuming the keyword vector of the patent text is Si, and its corresponding technical topic category word vector is T. j The formula for calculating the similarity between patent text keywords and subject terms is:

[0058]

[0059] Among them, ST ij This represents the similarity between the i-th keyword and the j-th subject term in the patent text.

[0060] The similarity between each keyword and category is calculated using the following formula:

[0061] h(ST i )=w1ST i1+ w2ST i2 ...+w n ST in ;

[0062] Among them, h(ST) i ) represents the similarity between each keyword and the category, w n This represents the contribution of a keyword to a category, where n is the number of keywords in each category.

[0063] Therefore, the similarity between the patent text and the category is:

[0064]

[0065] Where Z(ST) represents the similarity between the patent text and the category, and m is the number of keywords;

[0066] Finally, the patent text with the highest similarity to the category keywords and its corresponding technical field are selected as the prediction results for output.

[0067] Another object of the present invention is to provide a system based on the above-described intelligent model-based method for predicting highly disruptive patents in a domain, comprising:

[0068] User terminals are used by task publishers to publish disruptive technology prediction tasks via communication networks;

[0069] The technology topic extraction module is used to obtain relevant patents from the patent database based on the technology field to which the disruptive technology prediction task belongs and construct a patent dataset. The SVM classification algorithm is used to classify the patent dataset to obtain sub-datasets of each category. An LDA topic model is constructed, and the technology topic is extracted from each category of sub-datasets using the LDA topic model and the technology topic extraction results are output.

[0070] The technology topic screening module is used to construct a disruptive technology measurement and scoring model, and to use the disruptive technology measurement and scoring model to comprehensively score the extracted technology topics, screen out the technology topics with the highest comprehensive scores, and organize their corresponding patent texts.

[0071] The prediction result acquisition module is used to extract keywords from the selected patent texts, count the subject words of each selected technical topic corresponding to the category, calculate the similarity between the keywords of each patent text and the subject words of each corresponding category, and select the patent text with the highest similarity to the subject words of the category and its technical field as the prediction result for output.

[0072] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0073] 1. This invention first classifies the patent dataset and extracts technical themes using machine learning algorithms, then constructs a disruptive technology measurement and scoring model, and uses the disruptive technology measurement and scoring model to comprehensively score the extracted technical themes. From the comprehensive scores, the top 10% of technical themes and their corresponding patent texts are selected. Finally, the final prediction result is obtained by calculating the similarity between the theme words of the category and the keywords of the patent text. It follows the idea of ​​"identification → measurement" and combines three research methods: SVM-LDA, indicator system construction model, and cosine similarity algorithm. It innovatively proposes a method system for identifying disruptive technologies, which has a good effect on identifying potential disruptive technologies in complex and uncertain environments and improves the accuracy of disruptive technology identification.

[0074] 2. In addition, the present invention uses intelligent models to mine patents related to potential disruptive technologies, which has important practical value in helping to solve the problem of patent transfer and transformation difficulties for relevant institutions. Attached Figure Description

[0075] Figure 1 This is a flowchart of a domain-highly disruptive patent prediction method based on an intelligent model, as described in an embodiment of the present invention.

[0076] Figure 2 This is a flowchart illustrating the prediction results based on similarity filtering in an embodiment of the present invention. Detailed Implementation

[0077] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0078] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0079] The present invention will be further described below with reference to specific embodiments, but these are not intended to limit the scope of the invention.

[0080] like Figure 1 As shown, this embodiment of the invention provides a method for predicting highly disruptive patents in a domain based on an intelligent model, comprising the following steps:

[0081] Step 1: Obtain disruptive technology prediction tasks posted by users;

[0082] Users can publish one or more disruptive technology prediction tasks (such as technology prediction in the field of artificial intelligence) according to their own needs.

[0083] Step 2: Based on the technical field to which the disruptive technology prediction task belongs, obtain relevant patents from the patent database and construct a patent dataset. Use the SVM classification algorithm to classify the patent dataset to obtain sub-datasets of each category. Construct an LDA topic model, use the LDA topic model to extract technical topics from each category of sub-datasets, and output the technical topic extraction results.

[0084] In this step, patents related to the technical subject areas of the disruptive technology prediction task are acquired to construct a patent dataset. A technical foundation category classification standard file corresponding to the technical subject areas is then invoked. This standard file is formulated based on the patent IPC classification number. The patent dataset is classified using an SVM classification algorithm according to the technical foundation category classification standard file to obtain sub-datasets for each category. Then, the patents in each category's sub-dataset are processed into text vectors. In this embodiment, a text feature vectorization method is used to vectorize the patent text. The processing logic is to evaluate the importance of a word to a document in a document set or corpus. The calculation formula is:

[0085]

[0086] An LDA topic extraction model is constructed. A subset of the text vectorized data is input into the LDA topic extraction model for technical topic extraction, and the extracted technical topics are output. In this embodiment, the LDA topic extraction model can provide the topic of each document in the document set in the form of a probability distribution. This algorithm assumes that each word in an article is obtained through a process where "a document selects a certain topic with a certain probability, and this topic selects a certain word with a certain probability." The formula for LDA topic extraction is shown below:

[0087]

[0088] Step 3: Construct a disruptive technology measurement and scoring model. Use the disruptive technology measurement and scoring model to comprehensively score the extracted technology topics, select the top-ranked technology topics, and organize their corresponding patent texts.

[0089] In this embodiment, constructing a disruptive technology measurement and scoring model includes the following steps:

[0090] First, based on the results of technology theme extraction, measurement indicators are constructed from multiple dimensions. In this embodiment, measurement indicators are constructed from four dimensions: technological integration, technological innovation, technological importance, and technological breakthrough. The calculation formulas for each indicator are as follows:

[0091] 1) Technological convergence is measured from two dimensions: proximity centrality and the average number of IPC categories. The formula for calculating proximity centrality is:

[0092]

[0093] In the formula, AA i Denotes the proximity centrality of technical topic i in the middle, d ij Let N represent the shortest distance between technology topics i and j, and N represent the total number of technology topics.

[0094] The average number of IPC categories is calculated by dividing the number of IPC categories under each technology topic by the total number of IPC categories.

[0095] 2) Technological innovation is measured using the structural hole index. Constraint degree, hierarchy, and effective size are typical indicators of the structural hole index. The lower the hierarchy of a technological topic, the stronger its network capability, the less dependent it is on other technological topics, and the stronger its innovativeness. Conversely, the effective size measure shows the opposite: the larger the effective size, the stronger the technological topic's ability to acquire non-redundant information, and the higher the likelihood of achieving technological innovation. The specific formulas for constraint degree, hierarchy, and effective size are shown below:

[0096]

[0097] In the formula, C ij Represents the degree of constraint; node q is a common adjacency of nodes i and j; P ij P represents the weight proportion of node j among all the adjacent nodes of node i. iq P represents the weight ratio of node q within node i. ij This represents the weight ratio of node q among the adjacent nodes of node j;

[0098]

[0099] In the formula, YI i The degree of constraint is represented by N; N represents the individual network size of technical subject i; C / N represents the mean of the degree of constraint for each technical subject; and C represents the sum of the degrees of constraint.

[0100]

[0101] In the formula, YX i The effective size is represented by n; n represents the degree of node i, j represents the adjacent nodes of node i, and q represents the common adjacent nodes of nodes i and j; P iq and P jq These represent the weight proportions of node q among the adjacent nodes of node i and node j, respectively.

[0102] 3) The approach to determining technological importance is to use degree centrality and proximity centrality to quantify the power of technological topics and identify those with significant positions. Degree centrality describes the core position of a single technological topic within the network; proximity centrality describes the ability of nodes to transmit information. Higher proximity centrality indicates stronger information transmission capabilities, a greater likelihood of being at the network center, and thus, greater importance. The formula for calculating degree centrality is:

[0103] DC i =Ki / (g-1)

[0104] DC i K represents the degree centrality of technology topic i in the list of disruptive technologies. i The degree of technical topic i is represented by g, and the network size is represented by g.

[0105] The formula for calculating proximity centrality is:

[0106]

[0107] CCi represents the proximity centrality of technology topic i in the list of disruptive technologies, d ij This represents the shortest distance between technology topic i and technology topic j.

[0108] 4) Technological breakthroughs are determined using the K-means method for detecting technological anomalies:

[0109]

[0110] In the formula, dist(x,y) i () represents technology theme x and technology theme y in the list of disruptive technologies. i The distance between them, the technical topic y i It is one of the K nearest neighbor technical topics x, where K represents the number of nearest neighbors of technical topic x. For a specific value of K, the technical topics are clustered, and based on the clustering results, the distance from each point in the cluster to the cluster center is calculated. Finally, the distance is compared with a threshold. If it is greater than the threshold, it is considered abnormal; otherwise, it is normal.

[0111] Secondly, the importance of each metric for each technical topic is evaluated based on the analytic hierarchy process (AHP), resulting in the priority weight W for each metric. H The method for assigning weights to each indicator based on the Analytic Hierarchy Process (AHP) is as follows:

[0112] The first step is to construct pairwise comparison judgment matrices; based on the relative importance of each indicator and using a 1-9 scale, establish judgment matrix X; then, use the eigenvalue method to calculate the largest eigenvalue λ of judgment matrix X. max And the corresponding feature vector N, with m indicators, the judgment matrix X = (b ij ) m×m b in the matrix ij This indicates the degree of importance of indicator i compared to indicator j;

[0113] The second step is to perform a consistency check. The judgment matrix is ​​required to have consistency and transitivity. In this embodiment, the purpose of performing a consistency check on the judgment matrix is ​​to ensure its logical rationality. The consistency index CI = (λ) max -m) / m-1;

[0114] The third step is to calculate the consistency ratio. If the consistency ratio CR < 0.1, the consistency of the judgment matrix is ​​considered acceptable; otherwise, the value of the judgment matrix needs to be corrected. CR = CI / RI, where RI is the random consistency index.

[0115] The fourth step is to find the largest eigenvalue of the judgment matrix and its corresponding eigenvector; when the eigenvalue is n, the corresponding eigenvector is... Where k is a non-zero constant;

[0116] The fifth step is normalization; the obtained eigenvectors are normalized to obtain the weights W. H =(w1,w2,...w n ) T ;

[0117] Next, the contribution of each measurement index is evaluated based on the entropy weight method, and the weight W of each measurement index is obtained. j Specifically, it includes the following steps:

[0118] The first step is to standardize the indicator data, and the calculation formula is as follows:

[0119]

[0120] In the formula, x i As the initial value, y ij These are the standardized values ​​of each indicator.

[0121] The second step is to normalize the indicator data; after normalizing the data, the probability value is calculated using the following formula:

[0122]

[0123] The third step is to calculate the information entropy value of each indicator; according to the definition of information entropy in information theory, information entropy E j The formula is shown below:

[0124]

[0125] Step 4: Determine the weight of each indicator, using the following formula:

[0126]

[0127] Then, the weight W is calculated based on the principle of minimizing the overall deviation.H Weight W j The combined weight vector w; This embodiment obtains the subjective and objective weight vectors based on the analytic hierarchy process and the entropy weight method, and then calculates the combined weight vector based on the principle of minimizing overall deviation. This effectively solves the problems of a single weighting method and strong subjectivity, making the indicator weights more scientific and reliable. The calculation method of the combined weight is as follows:

[0128] w = Wα;

[0129] Where W = (W H W j ) is a matrix composed of the two weight vectors mentioned above, α=(α1,α2,…α n Let be the linear combination coefficient vector, satisfying α T α = 1. Therefore, the combined weighting model based on the principle of minimizing total deviation can be expressed as:

[0130]

[0131] Solving the combined weighting model yields the linear combination coefficient vector α, and then the combined weight vector w is obtained according to the combined weight calculation method.

[0132] Finally, based on the combined weight vector w, each measurement index is assigned a corresponding weight, and then the weighted sum of each measurement index is performed to obtain the comprehensive score for each technical topic. The technical topics are then sorted in descending order of comprehensive score. The top 10% of technical topics in terms of comprehensive score are selected as potential disruptive technical topics, and the corresponding patent texts for these potential disruptive technical topics are compiled.

[0133] Step 4: Extract keywords from the patent texts selected in Step 3, and statistically analyze the subject terms corresponding to each technical topic selected in Step 3. Calculate the similarity between the keywords of each patent text and the subject terms of each corresponding category. Select the patent text with the highest similarity to the category subject terms and its corresponding technical field as the prediction result for output; for example... Figure 2 As shown, this step specifically includes:

[0134] 1) Keyword extraction

[0135] The patent text abstracts selected in step 3 are processed using TF-IDF text mining technology to extract the keyword set S. The specific implementation method is as follows: First, calculate the term frequency (TF); then, calculate the inverse document frequency (IDF); finally, calculate the TF-IDF. Based on this, the top-N keywords are selected in descending order.

[0136] 2) Category keyword statistics and contribution

[0137] The LDA model is used to extract the set of topic word vectors T corresponding to the categories of the technical topics selected in step 3. The TF-IDF value of the topic words is used as the contribution of the topic words to the categories, laying the groundwork for subsequent weighted similarity calculation. In this step, the method for obtaining the topic word set T is as follows: the datasets of each subcategory selected in step 3 are input into the LDA topic model. The LDA topic model will create a TF-IDF model, convert words into word vector matrices, and finally output each technical topic and its corresponding set of topic word vectors T.

[0138] 3) Word vector representation

[0139] After keyword extraction, word vector representation of the keywords is needed to lay the groundwork for weighted similarity calculation from keywords to category-specific terms. Here, the word2vector model is chosen for word vector representation. This language model is trained on words using either the CBOW model or the Skip-gram model, focusing on predicting intermediate words based on contextual words or inferring related words in the context using the current word, respectively.

[0140] 4) Weighted similarity calculation from keywords to categories

[0141] Assume the keyword vector of the patent text is S. i The topic word vector corresponding to its technical topic category is T. j The formula for calculating the similarity between patent text keywords and subject terms is:

[0142]

[0143] Among them, ST ij This represents the similarity between the i-th keyword and the j-th subject term in the patent text.

[0144] The similarity between each keyword and the category to which its patent text belongs is calculated using the following formula:

[0145] h(ST i )=w1ST i1+ w2ST i2...+ w n ST in ;

[0146] Among them, h(ST) i ) represents the similarity between each keyword and the category, w n This represents the contribution of a keyword to a category, where n is the number of keywords in each category.

[0147] The similarity between the patent text and its category is then obtained as follows:

[0148]

[0149] Where Z(ST) represents the similarity between the patent text and the category, and m is the number of keywords;

[0150] Finally, the patent text with the highest similarity to the category keywords and its corresponding technical field are selected as the prediction results for output.

[0151] This invention also provides a system based on the above-described intelligent model-based method for predicting highly disruptive patents in a domain, comprising:

[0152] User terminals are used by task publishers to publish disruptive technology prediction tasks via communication networks;

[0153] The technology topic extraction module is used to obtain relevant patents from the patent database based on the technology field to which the disruptive technology prediction task belongs and construct a patent dataset. The SVM classification algorithm is used to classify the patent dataset to obtain sub-datasets of each category. An LDA topic model is constructed, and the technology topic is extracted from each category of sub-datasets using the LDA topic model and the technology topic extraction results are output.

[0154] The technology topic screening module is used to construct a disruptive technology measurement and scoring model, and to use the disruptive technology measurement and scoring model to comprehensively score the extracted technology topics, select the top 10% of technology topics in terms of comprehensive score, and organize their corresponding patent texts.

[0155] The prediction result acquisition module is used to extract keywords from the selected patent texts, count the subject words of each selected technical topic corresponding to the category, calculate the similarity between the keywords of each patent text and the subject words of each corresponding category, and select the patent text with the highest similarity to the subject words of the category and its technical field as the prediction result for output.

[0156] The above are merely preferred embodiments of the present invention and are not intended to limit the implementation methods and protection scope of the present invention. Those skilled in the art should recognize that any equivalent substitutions and obvious changes made based on the content of this specification should be included within the protection scope of the present invention.

Claims

1. A method for predicting highly disruptive patents in a domain based on an intelligent model, characterized in that, Includes the following steps: Step 1: Obtain disruptive technology prediction tasks posted by users; Step 2: Based on the technical field to which the disruptive technology prediction task belongs, obtain relevant patents from the patent database and construct a patent dataset. Use the SVM classification algorithm to classify the patent dataset to obtain sub-datasets of each category. Construct an LDA topic model, use the LDA topic model to extract technical topics from each category of sub-datasets, and output the technical topic extraction results. Step 3: Construct a disruptive technology measurement and scoring model. Use the disruptive technology measurement and scoring model to comprehensively score the extracted technology topics, select the top-ranked technology topics, and organize their corresponding patent texts. Step 4: Extract keywords from the patent texts selected in Step 3, and count the subject words corresponding to each technical topic selected in Step 3. Calculate the similarity between the keywords of each patent text and the subject words of each corresponding category. Select the patent text with the highest similarity to the subject words of the category and its corresponding technical field as the prediction result for output.

2. The method for predicting highly disruptive patents in a domain based on an intelligent model according to claim 1, characterized in that, Step 2 involves classifying the patent dataset using the SVM classification algorithm as follows: Patents related to the technology subject areas of the disruptive technology prediction task are acquired and constructed into a patent dataset. The technology basic category classification standard file corresponding to the technology subject area is called, and the SVM classification algorithm is used to classify the patent dataset according to the technology basic category classification standard file to obtain sub-datasets of each category.

3. The method for predicting highly disruptive patents in a domain based on an intelligent model according to claim 1, characterized in that, The method for constructing a disruptive technology measurement and scoring model is as follows: First, measurement indicators are constructed from multiple dimensions based on the results of technology topic extraction; Secondly, the importance of each metric for each technical topic is evaluated based on the analytic hierarchy process (AHP), resulting in the priority weight W for each metric. H ; Next, the contribution of each measurement index is evaluated based on the entropy weight method, and the weight W of each measurement index is obtained. j ; Then, the weights W are calculated based on the principle of minimizing the overall deviation. H Weight W j The combined weight vector w; Finally, based on the combined weight vector w, each measurement index is assigned a corresponding weight, and then the weighted sum of each measurement index is performed to obtain the comprehensive score of each technical topic.

4. The method for predicting highly disruptive patents in a domain based on an intelligent model according to claim 3, characterized in that, The measurement indicators are constructed from four dimensions: technological integration, technological innovation, technological importance, and technological breakthrough.

5. The method for predicting highly disruptive patents in a domain based on an intelligent model according to claim 4, characterized in that, Technological importance is quantified using degree centrality and proximity centrality to determine technological subject rights. The formula for calculating degree centrality is: DC i =K i / (g-1) In the formula, DC i K represents the degree centrality of technical topic i in the network. i The degree of technical topic i is represented by g, and the network size is represented by g. The formula for calculating proximity centrality is: ; In the formula, CC i d represents the proximity centrality of technical topic i in the network. ij Let represent the shortest distance between technology topic i and technology topic j, and n represent the degree of node i; Technological breakthroughs are determined using the K-means method for detecting technological anomalies. The calculation formula is as follows: In the formula, dist(x,y) i ) represents technology topic x and technology topic y i The distance between them, the technical topic y i It is one of the K nearest neighbor technical topics x, where K represents the number of nearest neighbors of technical topic x.

6. The method for predicting highly disruptive patents in a domain based on an intelligent model according to claim 3, characterized in that, Priority weights are obtained based on the analytic hierarchy process. W H The method is as follows: The first step is to construct pairwise comparison judgment matrices; establish judgment matrix X based on the relative importance of each indicator; then use the eigenvalue method to calculate the largest eigenvalue λ of judgment matrix X. max And the corresponding feature vector N, with m indicators, the judgment matrix X = (b ij ) m×m b in the matrix ij This indicates the degree of importance of indicator i compared to indicator j; The second step is to perform a consistency check; the judgment matrix must have consistency and transitivity, and the consistency index CI = (λ) / (λ) max -m) / m-1; The third step is to calculate the consistency ratio. If the consistency ratio CR < 0.1, the consistency of the judgment matrix is ​​considered acceptable; otherwise, the judgment matrix needs to be corrected. CR = CI / RI, where RI is the random consistency index. Step 4: Find the largest eigenvalue and the corresponding eigenvector of the judgment matrix; when the eigenvalue is n, the corresponding eigenvector is... Where k is a non-zero constant; The fifth step is normalization; the obtained eigenvectors are normalized to obtain the weights. .

7. The method for predicting highly disruptive patents in a domain based on an intelligent model according to claim 3, characterized in that, The weight W is obtained based on the entropy weight method. j The method is as follows: The first step is to standardize the indicator data, and the calculation formula is as follows: y ij = ; In the formula x i As the initial value, y ij These are the standardized values ​​of each indicator. The second step is to normalize the indicator data; after normalizing the data, the probability value is calculated using the following formula: P ij = ; The third step is to calculate the information entropy value of each indicator; according to the definition of information entropy in information theory, information entropy E j The formula is shown below: E j = ; Step 4: Determine the weight of each indicator, using the following formula: 。 8. The method for predicting highly disruptive patents in a domain based on an intelligent model according to claim 1, characterized in that, Step 4 includes the following steps: The selected patent text abstracts are processed to extract the keyword set S; The LDA model is used to extract the set of topic word vectors T corresponding to the categories of the technical topics selected in step 3, and the TF-IDF value of the topic words is used as the contribution of the topic words to the categories. The extracted keywords are represented by word vectors. Assuming the keyword word vectors in the patent text are... Si The word vectors of the corresponding technical topic categories are: T j The formula for calculating the similarity between patent text keywords and subject terms is: ; in, ST ij The first part of the patent text i The first keyword and the first j Similarity between keywords; The similarity between each keyword and category is calculated using the following formula: ; in, h(ST i ) This indicates the similarity between each keyword and the category. w n This represents the contribution of a keyword to a category, where n is the number of keywords in each category. Therefore, the similarity between the patent text and the category is: ; in, Z(ST) This represents the similarity between the patent text and the category, where m is the number of keywords; Finally, the patents with the highest similarity to the category keywords and their respective technical fields are selected as the prediction results for output.

9. A system for predicting highly disruptive patents in a domain based on an intelligent model according to any one of claims 1-8, characterized in that, include: User terminals are used by task publishers to publish disruptive technology prediction tasks via communication networks; The technology topic extraction module is used to obtain relevant patents from the patent database based on the technology field to which the disruptive technology prediction task belongs and construct a patent dataset. The SVM classification algorithm is used to classify the patent dataset to obtain sub-datasets of each category. An LDA topic model is constructed, and the technology topic is extracted from each category of sub-datasets using the LDA topic model and the technology topic extraction results are output. The technology topic screening module is used to construct a disruptive technology measurement and scoring model, and to use the disruptive technology measurement and scoring model to comprehensively score the extracted technology topics, screen out the technology topics with the highest comprehensive scores, and organize their corresponding patent texts. The prediction result acquisition module is used to extract keywords from the selected patent texts, count the subject words of each selected technical topic corresponding to the category, calculate the similarity between the keywords of each patent text and the subject words of each corresponding category, and select the patent text with the highest similarity to the subject words of the category and its technical field as the prediction result for output.