A patent value evaluation method based on depth map and semantic learning
By constructing a patent citation network and combining it with deep graphs and semantic learning, effective indicators are screened, the semantic novelty of patents is calculated, and the XGBoost algorithm is used to predict patent value. This solves the problem that traditional methods are difficult to identify high-value patents in the context of big data, and achieves higher evaluation accuracy and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DALIAN UNIV OF TECH
- Filing Date
- 2023-01-09
- Publication Date
- 2026-06-23
AI Technical Summary
Existing patent valuation methods struggle to quickly and effectively identify high-value patents in the context of big data, and traditional methods fail to fully consider the semantic novelty of patents.
By combining depth maps and semantic learning, a patent citation network is constructed to screen effective indicators, calculate the semantic novelty of patents, and use the XGBoost algorithm to predict patent value.
This paper presents an objective and fair method for evaluating patent value, which can effectively integrate multiple indicators and semantic information, improve the accuracy and reliability of the evaluation, and overcome the shortcomings of traditional methods.
Smart Images

Figure CN115983877B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of patent evaluation technology, and specifically relates to a patent value evaluation method based on depth maps and semantic learning. Background Technology
[0002] "High-value patents" are a hot topic in the industry, and cultivating high-value patents has become a consensus of the times for innovation-driven high-quality development. The State Intellectual Property Office has also made cultivating high-value patents and improving patent quality one of its key tasks. Therefore, how to assess patent value and identify high-value patents has become a critical issue that urgently needs to be addressed. However, with the deepening and implementation of the intellectual property strategy, the number of patents in my country has increased significantly, and traditional patent valuation methods are gradually becoming unable to meet the needs of assessing a large number of patents. Therefore, constructing a patent valuation model suitable for big data environments, and quickly and effectively identifying high-value patents from a large number of patents, has become a key issue in improving the quality of innovation and development.
[0003] Current research on patent value mainly explores the influencing factors of patent value from the perspective of single indicators, such as "Hall B, Trajtenberg M. Market value and patent citations[J].The Rand Journal of Economics,2005,36(1):16–38", "Lerner J, The importance of patent scope: an empirical analysis[J].The Rand Journal of Economics,1994,25,319-333.", "Harhoff D, Scherer FM, Vopel K. Citations, family size, opposition and the value of patent rights[J].Research Policy,2003,32(8).", and "Lanjouw J O, Schinkerman M. Patent quality and research productivity: measuring innovation with multiple indicators[J].Economic Journal, 2004, 114(495): 441–465., or from the perspective of evaluating patent value from multiple indicators, such as “Wan Xiaoli, Zhu Xuezhong. Evaluation index system and fuzzy comprehensive evaluation of patent value [J]. Scientific Research Management, 2008(02): 185-191., “Song Hefeng, Mu Rongping, Chen Fang. Research on patent quality and its measurement methods and measurement index system [J]. Science of Science and Management of Science and Technology, 2010, 31(04): 21-27., and “Guo Lei, Cai Hong, Zhang Yue. Analysis of the status of core patents in the context of patent strategy [J]. Science of Science Research, 2016, 34(11): 1663-1671+1757.,”).For example, Hall et al. first proposed using patent citation frequency to reflect patent value, and Lerner's research found that the scope of technology involved in a patent has a significant impact on patent value. However, these methods are difficult to objectively reflect the economic value of a patent. Secondly, many existing studies focus on relying on patent indicators to assess patent value, such as the number of patent citations and patent litigation. Wan Xiaoli et al. established an indicator system including 17 indicators such as innovation and technological content through hierarchical analysis and fuzzy comprehensive evaluation, proposing a new approach to patent value assessment by combining qualitative and quantitative methods. Guo Lei et al. found a significant positive relationship between the breadth of rights, the scope of technology, and self-citation behavior and patent value. However, it can be seen that the indicators in the study are characteristic information of the patent, and the indicators and their weights vary in the model. There is no consensus in academia on the selection of indicators. At the same time, the textual information of the patent is an important factor reflecting the novelty of the patent, and this semantic novelty has not been considered in existing research. Therefore, researchers need to propose a patent value assessment method that can effectively integrate multiple indicators and measure patent value from a semantic perspective. Summary of the Invention
[0004] This invention addresses the shortcomings of existing research by proposing a patent valuation method tailored to the characteristics of patents. First, patent features are screened. Then, deep semantic learning is used to measure the semantic novelty of the patent. Simultaneously, to effectively integrate external indicators and semantic information, node representations are learned based on mutual information maximization, preserving both local node information and global network information. Finally, the XGBoost algorithm is used to estimate the patent's value. This invention is the first to utilize semantic learning and deep graph learning to provide a method for evaluating the economic value of patents for big data applications.
[0005] The technical solution of this invention is a patent valuation method based on depth maps and semantic learning. It establishes a comprehensive evaluation model that effectively integrates multiple indicators and semantic novelty using existing patent datasets. This comprehensive evaluation model is then applied to the patent dataset to be evaluated to predict the value of the patents. The method includes the following steps:
[0006] Step 1. Obtain the citation relationships between the attribute features of the patent and the patent, and construct a patent citation network;
[0007] Step 2. Using patent transfer as the standard for high economic value patents, determine the preliminary selection indicators for patent valuation and the criteria to which the indicators belong;
[0008] The criteria for constructing the preliminary selection indicators for patent valuation include: technical indicators, citation indicators, IPC indicators, internationalization indicators, time indicators, rights indicators, and patentee indicators; the construction of the preliminary selection indicators is shown in Table 1.
[0009] Table 1. Criteria Layer and Selection Indicator System
[0010]
[0011] Step 3. Based on the KS method, screen the preliminary indicators for patent valuation and construct an indicator system for patent valuation;
[0012] Step 3.1: Standardization of preliminary index data for patent value assessment;
[0013] Data standardization processing employs the maximum-minimum standardization method to process the sample data of the preliminary indicators for patent value assessment, eliminating the influence of dimensions.
[0014] Step 3.2: Calculate the D value of a single indicator;
[0015] By calculating the maximum value of the cumulative frequency difference between transferred and non-transferred patents corresponding to each patent valuation screening index in the existing patent dataset, the KS test statistic D value of the patent valuation screening index is obtained.
[0016] Step 3.3: Calculate the correlation coefficients of indicators within the same criterion layer;
[0017] Calculate the correlation coefficient between any two indicators within the same criterion layer, identify the indicator pairs that reflect information duplication in the preliminary selection indicators for patent value assessment, delete the indicators with small D values from the indicator pairs with correlation coefficients greater than 0.7, and complete the first screening of the preliminary selection indicators for patent value assessment; the remaining K preliminary selection indicators for patent value assessment constitute the indicator system.
[0018] Step 3.4: Calculate the patent economic value score;
[0019] The remaining patent value assessment indicators are weighted according to the KS test statistic D value to ensure that the larger the D value, the greater the weight of the indicator; the economic value score of the patent is calculated by linear weighting; the weights of the patent value assessment indicators are calculated using formula (1):
[0020]
[0021] Calculate the patent economic value score using formula (2):
[0022]
[0023] Among them, w j D represents the weight of the preliminary selection indicators for assessing the value of the j-th patent. jLet D be the KS test statistic for the j-th indicator; k is the number of preliminary indicators for patent value assessment that need to be weighted: k = 1, 2, ..., K; K is the number of preliminary indicators for patent value assessment remaining after the first screening; Z is the patent economic value score; x j Let be the standardized value of the preliminary selection index for the value assessment of the j-th patent to be evaluated;
[0024] Step 3.5: Calculate the KS test statistic D value for the indicator system;
[0025] By analogy with the calculation of the D value of the preliminary selection index for a single value assessment, the KS test statistic D value of the patent economic value score derived from the index system is calculated.
[0026] Step 3.6: After calculating the index system D value composed of the K remaining patent value assessment indicators after the first screening, delete one patent value assessment indicator in turn, calculate the maximum value of D value in the combination of the remaining K-1 patent value assessment indicators, and compare the change of D value before and after deleting the patent value assessment indicator. If the D value of the remaining indicator combination is larger than before the deletion after deleting the patent value assessment indicator, then delete the patent value assessment indicator.
[0027] Step 3.7, repeat step 3.6 until any patent valuation index is deleted. If the D value of the remaining index combination is less than the D value before deleting the patent valuation index, then stop deleting the patent valuation index and complete the second screening of the patent valuation index. The remaining patent valuation index is the optimal patent valuation index combination.
[0028] Step 4. Calculate the semantic novelty of the patent, including the following steps;
[0029] Step 4.1: Establish a corpus set T = {t1, t2, ..., t} based on the invention title and abstract of the patent. i}, where t i It is the set of text information for patent i, namely the text consisting of the invention title and the patent specification abstract; the unique column vectors of the paragraph vector matrix V represent the text paragraphs of each patent, and the unique column vectors of the word vector matrix W represent each word in the patent text paragraphs;
[0030] Step 4.2: Predict text segment t based on the unique column vectors in the paragraph vector matrix and word vector matrix, which are the average values of the text segment and the words. i The probability of the next word appearing is used to derive the text paragraph representation and word representation; based on the training word sequence w1, w2, ..., w... |T| and paragraph v iMaximize the following objectives in a fixed-length window (Windows):
[0031]
[0032] Where M is the total number of training words, v i It is a text paragraph representation vector containing the words of the current window context; the prediction task is performed by hierarchical softmax:
[0033]
[0034] Where, N w is the total number of words in the training word sequence, and Pr is the log probability of the output, calculated using the following formula:
[0035] Pr = Ua(w t-|win| ,...,w t-1 ,w t+1 ,…,w t+|win| ,v i ;W,V)+b (5)
[0036] Where U and b are softmax parameters, and a is determined by w t and v i The average is obtained by using the PV-DM model in the latent space R. k The text segment representation of each patent is vectorized to obtain the final patent text representation matrix V;
[0037] Step 4.3: Calculate the Euclidean distance between the text paragraph representation vector of the patent and the text paragraph representation vector of the patent it cites:
[0038]
[0039] Step 4.4: Summarize and rank the Euclidean distances between all patent citation pairs |R| in the patent citation network, and calculate the semantic novelty S of the patent. i :
[0040]
[0041] Step 5: Based on the optimal combination of patent value assessment indicators obtained in Step 3 and the semantic novelty calculated in Step 4, generate a node feature matrix. Where n1 = |V|, an adjacency matrix for patent citations is established. To save the reference information between nodes, use an encoder. Obtaining the final node feature representation includes the following steps:
[0042] Step 5.1: Input the node feature matrix X, and obtain the local representation of the node in the positive sample by integrating the neighborhood information of the target node through a graph convolutional network ε; the information integration process is as follows:
[0043]
[0044] in, yes The degree matrix, H l It is the feature representation learned at each layer; W l These are the learning parameters of the l-th layer in the convolutional neural network; for the input layer l=0, H0=X, and σ is a non-linear activation function.
[0045] Step 5.2, Use the function To obtain negative samples, the nodes in the convolutional neural network are modified using the same information integration method as in step 5.1 to generate local node representations for the negative samples.
[0046] Step 5.3, through the transfer function Transmitting the local representation h of nodes in positive samples i Calculate the global representation of the network:
[0047]
[0048] Where N represents the number of positive samples;
[0049] Step 5.4: Use a discriminator Distinguish between local positive sample representations and negative sample representations:
[0050]
[0051] Step 5.5: Minimize the final loss function L n Update the final representation h of each patent node in the generated positive samples. i :
[0052]
[0053] Where, N n It is the number of negative samples; s is the negative sample representation; s is the network global representation; E (.) [.] represents the expected value of the function [.]. This represents the logarithm of formula (10);
[0054] Step 6: Patent Value Prediction; Input the final representation of the patent node into the XGBoost machine learning model to predict the value of the patent and obtain the score prediction result. For a given patent sample i, input its final patent node representation h.i The prediction result is obtained using the following formula:
[0055]
[0056] Among them, f k Let f be the k-th decision tree in the XGBoost model, where K is the number of trees in the model. k (h i ) represents the predicted value of patent sample i on the k-th tree.
[0057] The beneficial effects of this invention are as follows: This invention provides a patent valuation method based on depth graphs and semantic learning. In the indicator selection process, it combines patent transfer with the construction of a patent valuation indicator system, providing an objective, fair, and highly operable evaluation method for feature selection. Secondly, it calculates the novelty of the patent through textual semantic learning, measuring patent value from a semantic perspective. Furthermore, it utilizes depth graph learning to maximize the information integration node feature representation between local and global representations to evaluate patent value. This method overcomes the shortcomings of traditional methods in patent valuation, while introducing patent text novelty to measure patent value. Experimental results show that the proposed method has high accuracy and reliability. This invention provides a novel method for patent valuation and offers a new solution for patent value research. Attached Figure Description
[0058] Figure 1 This is a flowchart of the patent value assessment method based on depth graphs and semantic learning of the present invention.
[0059] Figure 2 This is a flowchart for the indicator selection process. Detailed Implementation
[0060] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.
[0061] This embodiment uses 2209 biopharmaceutical patents with a publication date greater than 5 years as examples. It employs indicators and criteria based on publication dates greater than 5 years to construct a patent valuation model and verify its effectiveness. 1473 patent samples were selected for constructing the valuation model, and 736 patent samples were used for patent valuation and model effectiveness verification. The implementation steps of this invention are as follows:
[0062] 1. Construct a patent citation network based on actual patent publications and citation information.
[0063] 2. Based on the characteristics of different patent indicators at different publication times, select preliminary selection indicators and construct a criterion layer.
[0064] 3. The index data of the patent sample are standardized by using the maximum-minimum standardization method to eliminate the influence of dimensions.
[0065] 4. Calculation of the KS test statistic D value for a single indicator.
[0066] The D-value of the preliminary selection indicator is used to measure the ability of the indicator to distinguish the patent transfer status. The larger the D-value, the greater the difference between transferred and non-transferred patents on that indicator, meaning that the indicator is better able to identify whether a patent has been transferred. The following uses the indicator "number of pages in the specification" as an example to illustrate the calculation steps for the D-value of a single indicator. For ease of understanding, we assume that the standardized values for "number of pages in the specification" are 1, 0.5, and 0.
[0067] (4.1) Each "Pages of Specification" indicator value corresponds to one or more patents. Patents with the same indicator value constitute a patent group. These patent groups are arranged in descending order according to the value of the "Pages of Specification" indicator. List them in row 2 of Table 2. Row 1 of Table 2 contains the patent group numbers.
[0068] (4.2) Calculate the number of transferred patents and the number of untransferred patents in each patent group and list them in rows 3 and 4 of Table 2, respectively.
[0069] (4.3) Calculate the number of transferred patents and the number of untransferred patents in each accumulated patent group.
[0070] The patent group with the highest index value is designated as the first accumulated patent group. Then, the next patent group with a lower index value is accumulated, meaning the first two patent groups form the second accumulated patent group, and the first three patent groups form the third accumulated patent group. The number of transferred patents and the number of untransferred patents in each accumulated patent group are calculated and listed in rows 5 and 6 of Table 2, respectively.
[0071] (4.4) Calculate the cumulative frequency of transferred patents and the cumulative frequency of untransferred patents in each accumulated patent group.
[0072] The cumulative number of transferred patents in row 5 of Table 2 is divided by the cumulative total number of transferred patents in the last column of row 5 of Table 2 to obtain the cumulative frequency of transferred patents, which is listed in row 7 of Table 2. Similarly, the cumulative number of untransferred patents is divided by the cumulative total number of untransferred patents to obtain the cumulative frequency of untransferred patents, which is listed in row 8 of Table 2.
[0073] (4.5) Calculate the difference d between the cumulative frequency of transferred patents and the cumulative frequency of untransferred patents in each accumulated patent group, d = |cumulative frequency of transferred patents - cumulative frequency of untransferred patents|, and list them in row 9 of Table 2.
[0074] (4.6) Determine the KS test statistic D value for a single indicator.
[0075] The KS test statistic D value is the maximum value of the difference d between the cumulative frequency of transferred patents and the cumulative frequency of untransferred patents, i.e., D = max(d). The obtained D value is listed in row 10 of Table 2.
[0076] Table 2. Calculation of the D-value of the KS test statistic.
[0077]
[0078] 5. Delete indicators that reflect duplicate information, and perform the first screening of indicators.
[0079] Calculate the correlation coefficient between any two indicators within the same criterion layer. For indicator pairs with a correlation coefficient greater than 0.7, delete the indicator with the smaller D value. This avoids information redundancy in the indicator system and also prevents the accidental deletion of indicators with strong transferability. The formula for calculating the correlation coefficient between indicator q and indicator j is:
[0080]
[0081] Where, r qj x represents the correlation coefficient between the q-th and j-th indicators; iq It is the q-th index value of the i-th patent; x represents the average value of the q-th indicator; ij It is the j-th index value of the i-th patent; It is the average value of the j-th indicator.
[0082] Through relevant analysis, in the indicator system where the patent publication time is greater than 5 years, a total of 9 indicators, including "number of domestic patents cited" and "number of foreign patents cited", were deleted, leaving 20 indicators.
[0083] 6. Assigning weights to indicators based on D-values
[0084] The indicators are weighted according to the principle that "the larger the D value of the KS test statistic of the indicator's transferability and discrimination ability, the greater the indicator's weight." The weighting formula is as follows:
[0085]
[0086] Among them, w j D represents the weight of the j-th indicator. j is the KS test statistic D value of the j-th indicator, representing the indicator's transferability; k is the number of indicators to be weighted, k = 1, 2, ..., 20.
[0087] 7. Calculate the patent value score
[0088] The economic value score of a patent is calculated using a linear weighting method, with the following weighting formula:
[0089]
[0090] Where Z represents the patent value score; w j Let x be the weight of the j-th indicator; k is the number of indicators to be weighted, k = 1, 2, ..., 20; j Let be the standardized value of the j-th indicator of the patent to be evaluated.
[0091] 8. Calculate the D-value of the patent value score and conduct a second screening of the indicator system.
[0092] (8.1) Calculate the D of the rating index system consisting of the remaining 20 indicators after the first screening. 20 .
[0093] Based on the calculation method for the individual indicator D value, the D value score of the patent value system composed of 20 indicators is calculated. 20 Value. Where D 20 The calculation is similar to the calculation of the D value of a single indicator. When inputting data, the standardized value of a single indicator needs to be replaced with the "patent value score".
[0094] (8.2) Determine the maximum value
[0095] After deriving 20 indicators D 20 After calculating the values, one indicator is removed at a time, and the remaining 19 indicators are used to calculate the system. Values are selected from 20 indicator combinations, after removing one indicator. The maximum value in
[0096] (8.3) Screen out an index system with a strong patent transfer differentiation ability D value.
[0097] When D 20 This indicates that the indicator system consisting of 19 indicators remaining after removing one indicator from the initial 20 indicators becomes more effective at distinguishing between transferred and non-transferred patents. Therefore, the rating system with 19 indicators is retained.
[0098] (8.4) Repeat steps (2) and (3) to continue deleting indicators until the time is reached. Stop screening indicators when that time comes.
[0099] This means that if any one indicator is removed from the k indicators, the remaining k-1 indicators will have a weaker ability to distinguish patent transfers. In this case, the k-indicator system should be retained and the indicator screening should be terminated.
[0100] After the second round of indicator screening, nine indicators, including "number of IPC subclasses" and "number of attached drawings," were removed from the indicator system with a patent publication period of more than 5 years, leaving 11 indicators. The indicator system composed of the remaining indicators is the one with strong patent transfer differentiation capabilities.
[0101] 9. Calculate the semantic novelty of the patent.
[0102] (9.1) Establish a corpus set T = {t1, t2, ..., t3} based on the invention title and abstract of the patent. i}, where t i It is a collection of text information for patent i. The unique column vectors of matrix V represent each text paragraph, and the unique column vectors of matrix W represent each word in a sentence. Maximize the following objective within a fixed-length window win:
[0103]
[0104] Where M is the total number of training words, v i This is a document representation vector containing the context words of the current window. Hierarchical softmax is used to predict the probability of the next word appearing in the document:
[0105]
[0106] Calculate the log probability of each paper's output:
[0107] Pr = Ua(w t-|win| ,...,w t-1 ,w t+1 ,…,w t+|win| ,v i ;W,V)+b
[0108] Where U and b are softmax parameters, and a is determined by w i and d j The average is obtained by using the PV-DM model in the latent space R. k The text representation matrix V of the patent is obtained by vectorization.
[0109] (9.2) Calculate the distance between the vector of a patent and the vector of the patents it references:
[0110]
[0111] (9.3) Summarize and rank the distances between all citation pairs, and calculate the semantic novelty score S of the patent. i :
[0112]
[0113] 10. Generate a node feature matrix based on screening criteria and calculated semantic novelty. Where n1 = |V|, establish the matrix To save the reference information between nodes, use an encoder. The final node feature representation acquisition process includes the following steps:
[0114] (10.1) Input the feature matrix X, and obtain the node representation in the positive sample by integrating the neighborhood information of the target node through the graph convolutional network ε:
[0115]
[0116] in yes The degree matrix, H l It is the feature representation learned at each layer.
[0117] (10.2) Using functions Modify the nodes in the network to obtain negative samples, and generate representations for the negative samples using the same method as in step (10.1).
[0118] (10.3) Through the transfer function Transmit local node representations and compute the global network representation:
[0119]
[0120] Where N represents the number of positive samples.
[0121] (10.4) Using a discriminator Using the distinction between local positive and negative samples:
[0122]
[0123] (10.5) Calculate the final loss function:
[0124]
[0125] Where, N n It represents the number of negative samples.
[0126] (10.6) Minimize the loss function to generate the representation h of each patent node. i .
[0127] 11. Patent Value Prediction. The patent node representation is input into the value prediction model XGBoost to obtain the score prediction result. For a sample i, its feature representation h is input. i The prediction result is obtained using the following formula:
[0128]
[0129] Among them, f k Let k be the decision tree.
[0130] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention shall fall within the scope of protection claimed by the present invention, and the scope of protection of the present invention shall be determined by the scope of protection of the claims.
Claims
1. A patent valuation method based on depth maps and semantic learning, characterized in that, Includes the following steps: Step 1. Obtain the citation relationships between the attribute features of the patent and the patent, and construct a patent citation network; Step 2. Using patent transfer as the standard for high economic value patents, determine the preliminary selection indicators for patent valuation and the criteria to which the indicators belong; The criteria for constructing preliminary indicators for patent valuation include: technical indicators, cited indicators, IPC indicators, internationalization indicators, time indicators, rights indicators, and patentee indicators. Step 3. Based on the KS method, screen the preliminary indicators for patent valuation and construct an indicator system for patent valuation; Step 3.1: Standardization of preliminary index data for patent value assessment; Data standardization processing employs the maximum-minimum standardization method to process the sample data of the preliminary selection indicators for patent value assessment, eliminating the influence of dimensions. Step 3.2: Calculate individual indicators value; By calculating the maximum cumulative frequency difference between transferred and non-transferred patents corresponding to each patent valuation screening index in the existing patent dataset, the KS test statistic for the patent valuation screening index is obtained. value; Step 3.3: Calculate the correlation coefficients of indicators within the same criterion layer; Calculate the correlation coefficient between any two indicators within the same criterion layer to identify indicator pairs reflecting information overlap in the initial screening indicators for patent valuation. Remove indicators with smaller D-values from indicator pairs with correlation coefficients greater than 0.7, completing the first screening of the initial screening indicators for patent valuation; the remaining... The indicator system is composed of preliminary selection indicators for evaluating the value of individual patents; Step 3.4: Calculate the patent economic value score; According to the KS test statistic The value is used to assign weights to the preliminary selection indicators for assessing the value of remaining patents, ensuring... The larger the value of the indicator, the greater its weight; the economic value score of the patent is calculated by linear weighting; the weight of the preliminary selection indicators for patent valuation is calculated using formula (1): (1) Calculate the patent economic value score using formula (2): (2) in, For the first Weighting of selection indicators for patent valuation; For the first KS test statistic for each indicator value; These are the initial selection indicators for patent value assessment that need to be assigned weights: ; The number of preliminary selection indicators for assessing the value of the remaining patents after the first screening. Score the economic value of the patent; For the patent to be evaluated Standardized values of the preliminary selection indicators for patent valuation; Step 3.5: Calculate the KS test statistic for the indicator system. value; Analogous to the initial screening criteria for a single value assessment The KS test statistic for calculating the patent economic value score derived from the indicator system is then calculated. value; Step 3.6: Calculate the remaining values after the first screening. An indicator system composed of preliminary selection indicators for patent value assessment. After determining the value, delete one patent valuation index at a time, and calculate the remaining values. In the preliminary selection of indicators for patent valuation The maximum value is compared with the value before and after removing the preliminary selection indicators for the patent's valuation. The change in value, after removing the preliminary selection indicators for the patent valuation, is reflected in the remaining indicator combinations. If the value increases compared to before deletion, then the preliminary selection index for the patent value assessment will be deleted. Step 3.7, repeat step 3.6 until any one of the initial selection indicators for patent valuation is removed, leaving the remaining indicator combinations. The values are all lower than before the removal of the preliminary selection criteria for the patent's valuation. At this point, the deletion of preliminary indicators for patent valuation is stopped, and the second screening of preliminary indicators for patent valuation is completed; the remaining preliminary indicators for patent valuation are the optimal combination of preliminary indicators for patent valuation. Step 4. Calculate the semantic novelty of the patent, including the following steps; Step 4.1: Establish a corpus set based on the invention title and abstract of the patent. ,in, It is a patent The text information set, namely the text consisting of the invention title and the patent specification abstract; paragraph vector matrix. The unique column vectors represent the text paragraphs of each patent, and the word vector matrix. The unique column vector represents each word in the patent text paragraph; Step 4.2: Predict the text paragraph based on the unique column vectors in the paragraph vector matrix and word vector matrix, which are the average values of the text paragraphs and words. The probability of the next word appearing is used to derive text paragraph and word representations; based on the training word sequence... and paragraphs In a fixed-length window Maximize the following objectives: (3) in, It is the total number of training words. It is a text paragraph representation vector containing the words of the current window context; the prediction task is performed by hierarchical softmax: (4) in, It is the total number of words in the training word sequence. It is the logarithmic probability of the output, calculated using the following formula: (5) in, and It is the softmax parameter. It is by and Averaging is obtained using the PV-DM model in the latent space. Vectorizing the text paragraph representation of each patent yields the final text representation matrix of the patent. ; Step 4.3: Calculate the Euclidean distance between the text paragraph representation vector of the patent and the text paragraph representation vector of the patent it cites: (6) Step 4.4: Summarize all patent citation pairs in the patent citation network. Calculate the semantic novelty of a patent by ranking the Euclidean distances between them. : (7) Step 5: Based on the optimal combination of patent value assessment indicators obtained in Step 3 and the semantic novelty calculated in Step 4, generate a node feature matrix. ,in Establish an adjacency matrix for patent citations. To save the reference information between nodes, use an encoder. Obtaining the final node feature representation includes the following steps: Step 5.1: Input the node feature matrix Through graph convolutional networks Integrating neighborhood information of the target node yields local representations of nodes in positive samples; the information integration process is as follows: (8) in, , yes The degree matrix, It is the feature representation learned at each layer; It is the first in convolutional neural networks The learning parameters of the layer; for the input layer In other words, , It is a non-linear activation function; Step 5.2, Use the function To obtain negative samples, the nodes in the convolutional neural network are modified using the same information integration method as in step 5.1 to generate local node representations for the negative samples. ; Step 5.3, through the transfer function Transmitting local representations of nodes in positive samples Calculate the global representation of the network: (9) in, Represents the number of positive samples; Step 5.4: Use a discriminator Distinguish between local positive sample representations and negative sample representations: (10) Step 5.5: Minimize the final loss function Update the final representation of each patent node in the generated positive samples. : (11) in, It is the number of negative samples; It represents negative samples; It is a global representation of the network; Represents the expected value of the function [.]. Represents the logarithm of formula (10); Step 6: Patent Value Prediction; Input the final representation of the patent node into the XGBoost machine learning model to predict the value of the patent and obtain the score prediction result. For a certain patent sample The final representation of its patent node is input. The prediction result is obtained using the following formula: (12) in, In the XGBoost model, the first Decision tree, The number of trees in the model. Patent sample In the Predicted values on each tree.