Multi-index fused chinese patent value evaluation method

By using a multi-indicator fusion method, combining a large language model and an XGBoost model, features are extracted from patent texts and basic information, solving the accuracy problem of Chinese patent valuation on small sample datasets and achieving more accurate patent valuation.

CN119991157BActive Publication Date: 2025-12-30MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY +1

Patent Information

Application Number
CN202411597818.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-12-30
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately assess the value of Chinese patents on small sample datasets, limited by data scarcity and the diversity of assessment standards.

Method used

This study employs a multi-indicator fusion approach, combining large language models to extract patent text features, utilizing GLM and HBERT models to extract short and long text features from the specification abstract and claims, and combining XGBoost models to assess patent value levels, integrating technical, economic, legal, and corporate dimensions.

Benefits of technology

A relatively objective and accurate patent value assessment was achieved on a small sample dataset, improving the assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991157B_ABST
    Figure CN119991157B_ABST
Patent Text Reader

Abstract

The application discloses a multi-index fused Chinese patent value evaluation method and belongs to the technical field of patent value evaluation. The method comprises the following steps: step 1, extracting patent text dimension features; including: extracting short text features from the abstract based on a large language model (GLM), and extracting long text features from the claim based on an HBert model; step 2, extracting patent basic information index features; including technical dimension indexes, economic dimension indexes, legal dimension indexes and enterprise dimension indexes; and step 3, performing patent value grade evaluation by using an XGBoost model based on the patent text dimension features and the patent basic information index features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of patent valuation technology, and in particular relates to a method for valuing Chinese patents by integrating multiple indicators. Background Technology

[0002] With the deepening of the new round of technological revolution and industrial transformation, the importance of technological innovation is becoming increasingly prominent. Accurately and objectively identifying high-value patents is of significant importance. Patent valuation methods are constantly being updated and improved. Initially, valuation primarily relied on the basic attributes of the patent combined with expert opinions in the field. Subsequently, with the rise of artificial intelligence technology, many scholars began to use machine learning and deep learning methods to research Chinese patent valuation. However, due to the scarcity of Chinese patent valuation data and the diversity of valuation standards, it remains impossible to achieve a relatively accurate valuation based on small sample datasets. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention proposes a multi-indicator fusion method for Chinese patent valuation, which combines a large language model with small sample datasets; namely, a multi-indicator fusion Chinese patent valuation scheme. This scheme can achieve relatively objective and accurate results even with scarce data.

[0004] The first aspect of this invention discloses a method for evaluating the value of Chinese patents by integrating multiple indicators, the method comprising:

[0005] Step 1: Extract dimensional features from the patent text; including: extracting short text features from the specification abstract based on the Large Language Model (GLM), and extracting long text features from the claims based on the HBERT model;

[0006] Step 2: Extract the basic information indicators and features of the patent; including technical, economic, legal, and corporate indicators.

[0007] Step 3: Based on the dimensional features of the patent text and the basic information indicators of the patent, use the XGBoost model to evaluate the patent value level.

[0008] According to the method of the first aspect of the present invention, in step 1, short text features are extracted from the specification summary based on the Large Language Model (GLM); wherein:

[0009] For the instruction manual summary, the Large Language Model (GLM) is used for feature embedding. The GLM model is based on the Transformer model, which is based on a 28-layer Transformer Encoder module. The order of normalization and residual connections is rearranged, and a single linear layer is used to predict the output words. GeLUs are used instead of ReLU activation functions.

[0010] The LoRa algorithm is used to fine-tune the GLM model. In LoRa, the parameters of the original large language model are frozen, the intrinsic rank of the parameters of the bypass simulation model is increased, only the parameters in the bypass matrix are trained, and the parameters in the bypass matrix are superimposed with the original large language model. The text feature encoding of the instruction manual summary is performed through the fine-tuned GLM model and a fully connected layer. The calculation formula is as follows:

[0011] GLM(W0+ΔW)=Lora(GLM(W0))

[0012] V A =GLM(A,W0+ΔW)

[0013] V AL =Linear(V A )

[0014] Where W0 represents the original parameters of the GLM model, ΔW represents the parameters of the bypass matrix fine-tuned by the Lora algorithm, A represents the input instruction summary text, and V AL This represents the final obtained feature word vectors, which serve as short text features.

[0015] According to the method of the first aspect of the present invention, in step 1, long text features are extracted from the claims based on the HBERT model; wherein:

[0016] The HBERT model acquires long text features through a hierarchical approach; by dividing the long text into sentences, the long text is represented as D = (sen1, sen2, ..., sen...). n ), D represents long text, sen i Let represent the i-th sentence of a long text containing n sentences. By splitting each sentence character by character, we can obtain the sentence sen. i Represented as sen i =(w i1 ,w i2 ,…,w im ), w ij The sentence sen i The j-th character, sen i There are m characters in total;

[0017] After dividing the long text, the sentence sen i =(w i1 ,w i2 ,...,w im As input to the BERT model, word vectors H are obtained at the sentence level. CLS Word vectors at the word granularity H CLSAttention weights are calculated for each word vector, and the representative sentence sen is obtained by weighted summation. i word vector S i The calculation formula is:

[0018]

[0019] The long text D = (sen1, sen2, ..., sen) n Each sentence in the text is input into the BERT model, which generates a set of word vectors {S1, S2, ..., S} for each sentence in the long text. n}, prepend a randomly initialized word vector [SCLS] before the set of word vectors, and set {[SCLS], S1, S2, ..., S...} to the set of word vectors. n As input, the words are encoded by the Transformer Encoder model to obtain word vectors H containing long text granularity. SCLS and sentence-level word vectors H SCLS Attention weights are calculated for the word vectors of each sentence, and the word vector V representing the long text D is obtained by weighted summation. D The calculation formula is:

[0020]

[0021] According to the method of the first aspect of the present invention, in step 2:

[0022] Technical dimension indicators include: technical efficacy matrix, number of independent claims, number of dependent claims, number of pages in the specification, number of embodiments, number of citations, number of cited documents, number of priorities, number of technical classifications, number of divisional and sub-cases, number of patent families, patent awards, and declassification of classified patents;

[0023] Economic indicators include: number of applicants, applicant types, authorization period, number of license filings, number of rights pledges, number of rights transfers, and customs filing status.

[0024] Legal dimension indicators include: legal status, lifespan / maintenance time, remaining lifespan, litigation status, and number of invalid claims;

[0025] Enterprise-level indicators include: company registered capital, number of insured employees, number of self-risks, number of related risks, number of historical risks, and number of sensitive public opinion incidents; university level, operating capital, number of self-risks, number of related risks, number of historical risks, and number of sensitive public opinion incidents.

[0026] The above indicators are divided into numerical indicators, categorical indicators, and matrix indicators. The representation forms of numerical indicators, categorical indicators, and matrix indicators are obtained respectively. The vector representation forms of all indicators are concatenated to obtain the basic information indicator features of the patent.

[0027] According to the method of the first aspect of the present invention, in step 3:

[0028] The XGBoost model is adopted as the patent value assessment model. XGBoost employs a massively parallel gradient boosting tree-based ensemble learning algorithm, performs a second-order Taylor expansion on the loss function, and introduces L1 and L2 regularization terms. The obtained patent text dimensional features and basic patent information index features are concatenated to obtain a patent feature vector, which is used as the input to the XGBoost model to perform a classification task and output the patent value level. The calculation formula is as follows:

[0029] V = concat(V) AL V D V info )

[0030] H = XGBoost(V)

[0031] Among them, V AL V represents the short text feature in the patent text dimension features. D V represents the long text feature in the patent text dimension features, V represents the patent feature vector, and H represents the patent value level result calculated by the model.

[0032] A second aspect of this invention discloses a multi-indicator fusion system for evaluating the value of Chinese patents. The system includes a processing unit configured to perform the following steps:

[0033] Extracting dimensional features from patent texts; including: extracting short text features from specification abstracts based on the Large Language Model (GLM), and extracting long text features from claims based on the HBERT model;

[0034] Extract the basic information indicators and characteristics of patents, including technical, economic, legal, and corporate indicators.

[0035] Based on the dimensional features of patent text and the basic information indicators of patents, the XGBoost model is used to evaluate the value level of patents.

[0036] According to a system of a second aspect of the present invention, the processing unit is specifically configured to perform the extraction of short text features from a specification summary based on a Large Language Model (GLM); wherein:

[0037] For the instruction manual summary, the Large Language Model (GLM) is used for feature embedding. The GLM model is based on the Transformer model, which is based on a 28-layer Transformer Encoder module. The order of normalization and residual connections is rearranged, and a single linear layer is used to predict the output words. GeLUs are used instead of ReLU activation functions.

[0038] The LoRa algorithm is used to fine-tune the GLM model. In LoRa, the parameters of the original large language model are frozen, the intrinsic rank of the parameters of the bypass simulation model is increased, only the parameters in the bypass matrix are trained, and the parameters in the bypass matrix are superimposed with the original large language model. The text feature encoding of the instruction manual summary is performed through the fine-tuned GLM model and a fully connected layer. The calculation formula is as follows:

[0039] GLM(W0+ΔW)=Lora(GLM(W0))

[0040] V A =GLM(A,W0+ΔW)

[0041] V AL =Linear(V A )

[0042] Where W0 represents the original parameters of the GLM model, ΔW represents the parameters of the bypass matrix fine-tuned by the Lora algorithm, A represents the input instruction summary text, and V AL This represents the final obtained feature word vectors, which serve as short text features.

[0043] According to the system of the second aspect of the present invention, the processing unit is specifically configured to perform the extraction of long text features from the claims based on the HBERT model; wherein:

[0044] The HBERT model acquires long text features through a hierarchical approach; by dividing the long text into sentences, the long text is represented as D = (sen1, sen2, ..., sen...). n ), D represents long text, sen i Let represent the i-th sentence of a long text containing n sentences. By splitting each sentence character by character, we can obtain the sentence sen. i Represented as sen i =(w i1 ,w i2 ,...,w im ), w ij The sentence sen i The j-th character, sen i There are m characters in total;

[0045] After dividing the long text, the sentence seni =(w i1 ,w i2 ,...,w im As input to the BERT model, word vectors H are obtained at the sentence level. CLS Word vectors at the word granularity H CLS Attention weights are calculated for each word vector, and the representative sentence sen is obtained by weighted summation. i word vector S i The calculation formula is:

[0046]

[0047] The long text D = (sen1, sen2, ..., sen) n Each sentence in the text is input into the BERT model, which generates a set of word vectors {S1, S2, ..., S} for each sentence in the long text. n}, prepend a randomly initialized word vector [SCLS] before the set of word vectors, and set {[SCLS], S1, S2, ..., S n As input, the words are encoded by the Transformer Encoder model to obtain word vectors H containing long text granularity. SCLS and sentence-level word vectors H SCLS Attention weights are calculated for the word vectors of each sentence, and the word vector V representing the long text D is obtained by weighted summation. D The calculation formula is:

[0048]

[0049] According to the system of the second aspect of the present invention, the processing unit is specifically configured to perform:

[0050] Technical dimension indicators include: technical efficacy matrix, number of independent claims, number of dependent claims, number of pages in the specification, number of embodiments, number of citations, number of cited documents, number of priorities, number of technical classifications, number of divisional and sub-cases, number of patent families, patent awards, and declassification of classified patents;

[0051] Economic indicators include: number of applicants, applicant types, authorization period, number of license filings, number of rights pledges, number of rights transfers, and customs filing status.

[0052] Legal dimension indicators include: legal status, lifespan / maintenance time, remaining lifespan, litigation status, and number of invalid claims;

[0053] Enterprise-level indicators include: company registered capital, number of insured employees, number of self-risks, number of related risks, number of historical risks, and number of sensitive public opinion incidents; university level, operating capital, number of self-risks, number of related risks, number of historical risks, and number of sensitive public opinion incidents.

[0054] The above indicators are divided into numerical indicators, categorical indicators, and matrix indicators. The representation forms of numerical indicators, categorical indicators, and matrix indicators are obtained respectively. The vector representation forms of all indicators are concatenated to obtain the basic information indicator features of the patent.

[0055] According to the system of the second aspect of the present invention, the processing unit is specifically configured to perform:

[0056] The XGBoost model is adopted as the patent value assessment model. XGBoost employs a massively parallel gradient boosting tree-based ensemble learning algorithm, performs a second-order Taylor expansion on the loss function, and introduces L1 and L2 regularization terms. The obtained patent text dimensional features and basic patent information index features are concatenated to obtain a patent feature vector, which is used as the input to the XGBoost model to perform a classification task and output the patent value level. The calculation formula is as follows:

[0057] V = concat(V) AL V D V info )

[0058] H = XGBoost(V)

[0059] Among them, V AL V represents the short text feature in the patent text dimension features. D V represents the long text feature in the patent text dimension features, V represents the patent feature vector, and H represents the patent value level result calculated by the model.

[0060] A third aspect of this invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the multi-index fusion method for evaluating the value of Chinese patents described in the first aspect of this disclosure.

[0061] A fourth aspect of this invention discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multi-index fusion method for evaluating the value of Chinese patents described in the first aspect of this disclosure.

[0062] In summary, this invention proposes a patent valuation index standard, encompassing five dimensions: patent text, technology, economics, law, and corporate information. Each type of index contains various subtle features, and for the first time, it combines short and long texts, along with a patent technology efficacy matrix, into the valuation index. This invention also proposes a small-sample Chinese patent valuation method incorporating a large language model. This method integrates all the aforementioned patent indicators and, for the first time, applies a large language model to the field of patent valuation, effectively improving the valuation results on small-sample datasets. Based on artificial intelligence patent data, this invention has conducted extensive experiments to verify the accuracy of the small-sample Chinese patent valuation method based on the large language model and the effectiveness of each index component. Attached Figure Description

[0063] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0064] Figure 1 A framework diagram for a multi-indicator fusion model for assessing the value of Chinese patents;

[0065] Figure 2 This is a diagram of the Hbert model framework. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0067] Patent value varies under different assessment contexts because patent holders have different orientations and research perspectives on patent value, and different scholars also have different assessment standards for patent value. Patent valuation is conducted by qualified asset appraisers from third-party valuation agencies, following relevant regulations, procedures, and methods, to quantify the value of patents owned by a specific entity to meet the requirements of all parties involved in a patent transaction. Patent valuation has always been a key focus of research in the patent field, and current patent valuation methods can be divided into three main categories: economic methods, comprehensive evaluation methods, and other emerging valuation methods.

[0068] 1) Economic Methods

[0069] This paper analyzes and demonstrates the applicability of the replacement cost method to patent value analysis in equipment procurement. Based on the patent circumvention design concept, the concept of patent replacement cost is appropriately extended, and the implementation steps of the replacement cost assessment method are given. A case study is used to demonstrate the patent value analysis process based on patent replacement cost. It is proposed that fully integrating digital technology to establish a database of intellectual property valuation parameters, improving the comparable indicator system of the market approach to intellectual property valuation, and reasonably determining the comparable factor correction coefficient are key points for future application of the market approach to intellectual property valuation.

[0070] 2) Comprehensive evaluation method

[0071] The comprehensive evaluation method mainly employs a combination of fuzzy comprehensive evaluation, analytic hierarchy process (AHP), binary classification evaluation, principal component analysis, and entropy weight method to assess patent value. The key to this method lies in constructing a patent value assessment index system, making it a commonly used approach. Based on the value distribution characteristics of patent assets, this study constructs a binary classification evaluation method for patent value by analyzing and comparing key indicators and influencing factors of domestic and international patent value assessments. A study of patents granted to the Chinese Academy of Sciences in the United States over 28 years was conducted, and the weights of the evaluation indicators were calculated using the principal component method. This method can overcome the influence of subjective factors on the evaluation results.

[0072] 3) New evaluation methods

[0073] In recent years, many scholars have conducted research on improvements to patent valuation methods. In the context of big data and the urgent need for evolving valuation methods, other emerging methods have also emerged, mainly including machine learning and system dynamics methods. This paper uses rough set theory to simplify the patent valuation index system, and then establishes a patent valuation model based on a backpropagation (BP) neural network. From the perspective of machine learning technology, the patent valuation indicators are first analyzed and selected. Then, three algorithms from machine learning methods—decision trees, support vector machines, and neural networks—are used to train and test the samples. Finally, the test results are analyzed.

[0074] Currently, traditional methods for automatic patent valuation suffer from low accuracy, while newer methods are still in their early stages, mostly relying on formula calculations. Even newer methods using artificial intelligence haven't considered the impact of the original text on patent value or the limitations of sample data. A new automatic Chinese patent valuation method that combines a patent technology efficacy matrix with other numerical indicators, incorporating prompt learning and a large language model, can overcome both the challenges of multi-dimensional fusion of evaluation indicators and insufficient data samples, thus possessing significant research value.

[0075] Language modeling is a fundamental task in the field of natural language processing (NLP). It aims to predict and generate continuous text sequences that conform to grammatical and semantic rules through learning from text data, laying the foundation for other NLP tasks. The introduction of pre-trained language models based on the Transformer framework has led to significant success in language modeling.

[0076] With the release of ChatGPT in November 2022, large language models began to gain public attention and demonstrated remarkable capabilities in a range of natural language processing tasks. Large language models refer to Transformer language models trained on massive general-purpose corpora containing hundreds of billions or more parameters and spanning hundreds of megabytes. The mainstream architectures of large language models can be categorized into three types: encoder-decoder architecture, causal decoder architecture, and prefix decoder architecture. Encoder-decoder-based large language models exhibit stronger understanding and encoding capabilities for input text but weaker long text generation capabilities. Causal decoder-based models demonstrate stronger text generation capabilities and higher training efficiency. Prefix decoder-based models represent a compromise between the capabilities of the two aforementioned architectures. Due to the sheer power of their parameters and training data, all three types of large language models can demonstrate powerful semantic understanding and generation capabilities in general knowledge domains, showcasing reasoning abilities previously unavailable in neural network models.

[0077] Currently, numerous scholars both domestically and internationally have applied large language models to various downstream tasks. Inspired by the potential applications of ChatGPT in smart agriculture, this paper proposes an agricultural text classification method based on ChatGPT, demonstrating that large language models can effectively improve the accuracy of agricultural text classification and alleviate the problem of insufficient data samples in this field, showing high feasibility in advancing agricultural practices. A few-sample classification method for financial texts is proposed by combining the SetFit method with GPT3.5 and GPT4 models. This method achieves state-of-the-art results and provides a practical solution for few-sample tasks. An entity language model is proposed, integrating continuous sensor modalities from the real world into a large language model to achieve robot operation planning, visual question answering, and explanation functions.

[0078] Practice has proven that large language models can handle various complex artificial intelligence tasks, and due to the complexity of their model framework and the richness of their general training corpora, they can achieve superior results even on small-sample or zero-sample data. However, no scholars have yet applied them to the field of patent valuation. This paper will, for the first time, use large language models and integrate various patent indicators to conduct small-sample Chinese patent valuation.

[0079] A framework for a multi-indicator integrated Chinese patent valuation model, such as... Figure 1 As shown, the model mainly consists of three parts.

[0080] The first part is the patent text dimension feature extraction section, including short text feature extraction based on the Large Language Model (GLM) and long text feature extraction based on the HBERT model. The second part is the patent basic information indicator extraction section, including four dimensions: technical, economic, legal, and corporate information, with each type of indicator containing multiple subtle features. The third part is the patent value rating model based on the XGBoost model, which vectorizes and concatenates all the above patent evaluation indicators to predict the final value rating.

[0081] The abstract and claims are crucial components of an invention patent application. The abstract provides a concise and accurate summary of the entire patent application, outlining the core features, technical problems, solutions, and advantages of the invention or innovation in simple language. The claims define and establish the scope and limits of the patent rights, specifically listing the various elements, features, and combinations of the invention or innovation. To a certain extent, the abstract can provide a preliminary assessment of the patent's feasibility and innovative value, while the claims provide the patent's specific technical features and application areas, further determining its commercial value. However, previous scholars, limited by the difficulty of processing long texts, found that using only the abstract as a textual indicator for patent valuation insufficient. Therefore, this paper, based on previous research, incorporates claims information for the first time, simultaneously considering both the abstract and claims as textual features for patent valuation.

[0082] Based on the specification summary, this paper employs the Large Language Model (GLM) for feature embedding. The GLM model is based on the Transformer model, and its model structure and training data represent significant improvements over pre-trained language models. The GLM model is based on a 28-layer Transformer Encoder module. In addition to the increased number of layers, three improvements were made: 1) the order of layer normalization and residual connections was rearranged; 2) a single linear layer was used for output word prediction; and 3) GeLUs were used instead of ReLU activation functions. Furthermore, the GLM model utilizes more diverse and comprehensive training data, resulting in a more powerful language modeling capability.

[0083] To make the feature vector generation of the GLM model more suitable for subsequent value assessment tasks, instruction fine-tuning is required. However, the training cost of large language models is extremely high, making it impossible to retrain all model parameters. Therefore, the LoRa method is used to fine-tune the GLM model. The LoRa method freezes the parameters of the original large language model, adds a bypass to simulate the intrinsic rank of the model parameters, trains only the parameters in the bypass matrix, and finally superimposes the parameters in the bypass matrix with the original large language model. The fine-tuned GLM model and fully connected layers are then used for text feature encoding of the instruction manual summary, calculated using the following formula:

[0084] GLM(W0+ΔW)=Lora(GLM(W0)) (1)

[0085] V A =GLM(A,W0+ΔW) (2)

[0086] V AL =Linear(V A (3)

[0087] Where W0 represents the original parameters of the GLM model, ΔW represents the parameters of the bypass matrix fine-tuned by the Lora method, A represents the input instruction summary text, and V AL This represents the final feature word vector obtained.

[0088] For the patent application specification, since its text length is not limited, this paper adopts the HBERT model for handling long texts, as proposed in previous studies, for feature embedding. The HBERT model framework diagram is shown below. Figure 2 As shown.

[0089] The model obtains text vectors in a hierarchical manner. According to language conventions, a sentence typically expresses a complete semantic and grammatical structure. If long texts are segmented according to a fixed character length, the segmentation position cannot be guaranteed, leading to disruption of the semantic structure. Therefore, long texts are first divided into sentences, which can then be represented as D = (sen1, sen2, ..., sen...). n ), D represents long text, sen i Let represent the i-th sentence of a long text containing n sentences. Then, by splitting each sentence character by character, we can divide the sentence into segments. i Represented as sen i =(w i1 ,w i2 ,…,w im ), w ij The sentence sen i The j-th character, sen i There are m characters in total.

[0090] After dividing the long text according to the above method, first divide the sentence sen i =(w i1 ,w i2 ,…,w im As input to the BERT model, word vectors H are obtained at the sentence level. CLS Word vectors at the word granularity H CLS Attention weights are calculated for each word vector, and then a weighted sum is performed to obtain the representative sentence sen. i word vector S i The calculation formula is as follows:

[0091]

[0092] The long text D = (sen1, sen2, ..., sen) n Each sentence in the text is input into the BERT model in the manner described above, which yields the sentence vector set {S1, S2, ..., S...} for each sentence in the long text. n Then, a randomly initialized word vector [SCLS] is concatenated before this set of word vectors, and {[SCLS], S1, S2, ..., S...} is defined. n As input, the words are encoded using the Transformer Encoder model to obtain word vectors H that contain long text granularity. SCLS and sentence-level word vectors Then H SCLS Attention weights are calculated for the word vectors of each sentence, and then a weighted sum is performed to obtain the word vector V representing the long text D. D The calculation formula is as follows:

[0093]

[0094] Many scholars have proposed different indicators or combinations of indicators for basic information related to patent valuation. Based on previous research, this paper proposes four dimensions: technology dimension, economic dimension, legal dimension, and enterprise information dimension. The specific patent characteristic indicators included in each category are shown in Table 1.

[0095] Table 1. Basic Patent Information Indicators

[0096]

[0097]

[0098] The technological dimension primarily reflects the technological advancement and innovation of the patent solution. It is domain-specific and reflects the technological content and advantages of an invention patent, thus influencing its commercial application and market potential. This paper introduces a technology efficacy matrix into the patent valuation index system for the first time. The technology efficacy matrix can largely reflect the saturation and innovation level of the patent market. Its values ​​can reflect the concentration of patent applications in key technological points and efficacy, playing a positive role in the assessment of patent value.

[0099] The economic dimension primarily reflects the patentee's investment in the invention patent process and the revenue and profit obtained from utilizing the patent. Characteristics include the number of applicants, applicant types (enterprises, universities, research institutions, individuals, government agencies, others), grant period, number of license filings, number of rights pledges, number of rights transfers, and customs filing status. These characteristics can help determine the commercial feasibility and value of a patent, assisting companies in making decisions, including patent commercialization, licensing, and market positioning.

[0100] The legal dimension primarily reflects the scope, validity, and compliance of patent legal protection. Its characteristics include legal status (granted - approved, under examination - substantive examination, under examination - publication of invention, rejected - dismissed, deemed withdrawn, voluntarily withdrawn, unpaid annual fees, expired, deemed abandoned, voluntarily abandoned, abandoned due to overlap), lifespan / maintenance time, remaining lifetime, litigation status, and number of invalidation requests. These characteristics help determine the strength of patent legal protection and the stability of patent rights, further influencing the patent's commercial value and legal risks.

[0101] The enterprise information dimension primarily reflects the background, strength, and strategy of the company owning the patent. Enterprises are mainly divided into two categories: corporations and universities. Corporation characteristics include registered capital, number of employees covered by social insurance, number of inherent risks, number of associated risks, number of historical risks, and number of sensitive public opinion events. University characteristics include university level, initial capital, number of inherent risks, number of associated risks, number of historical risks, and number of sensitive public opinion events. These characteristics can help determine the importance of the patent in the company's strategy and development, as well as the company's capabilities and resource support in the patent commercialization process.

[0102] Taking into account the evaluation results of these dimensions can help determine the value and potential of a patent, and guide patent management and commercialization decisions.

[0103] The first part, the feature extraction of patent text dimensional information, obtains the specification summary feature vector V using methods based on the GLM model and the HBERT model, respectively. AL and the feature vector V of the claims specification D .

[0104] The second part, the extraction of basic patent information indicators, includes indicators across three dimensions: technical, economic, legal, and corporate information. These indicators can be categorized into three types of data:

[0105] First, there are numerical indicators, such as the number of applicants and registered capital, which are fixed numerical values. For these indicators, this paper uses the Z-score for quantification, as shown in the formula below:

[0106]

[0107] Where μ is the mean and σ is the variance. The Z-score transforms the original data into a standard normal distribution, giving the data a mean of 0 and a standard deviation of 1. This eliminates dimensional differences between different data points, converting them into comparable standard scores.

[0108] Second, categorical indicators, such as customs filing status and applicant type, are elements in a fixed-type array. For these indicators, this paper uses one-hot encoding to represent categorical variables as binary vectors, using the number 1 to indicate the presence of the category and the number 0 to indicate the absence of the category. This encoding provides a clear and interpretable representation of categorical variables.

[0109] Thirdly, there are matrix-type indicators, such as the technical effectiveness matrix. For this type of indicator, this paper uses a matrix normalization method to map the value range of the matrix to [0,1], eliminating the difference in dimensions and avoiding the excessive influence of certain features on the model.

[0110] Using the methods described above, we can obtain the representations of numerical, categorical, and matrix indicators. Then, by concatenating the vectors of all indicators, we can obtain the feature vector V of the patent basic information indicators. info .

[0111] The third part is the patent value assessment. Taking into account issues such as overfitting, multicollinearity, and training speed, the XGBoost model is adopted as the patent value assessment model. XGBoost is a massively parallel ensemble learning algorithm based on gradient boosting trees. XGBoost performs a second-order Taylor expansion of the loss function and introduces L1 and L2 regularization terms, which can control the model's complexity, accelerate the model's convergence speed, reduce the risk of overfitting, and ensure the model's robustness. Furthermore, XGBoost is unaffected by highly correlated features, reducing the problem of feature multicollinearity. All the vectors obtained above are concatenated to obtain a patent vector containing rich patent feature information, which is used as input to the XGBoost model for classification tasks, ultimately outputting the patent's value level. The formula is shown below:

[0112] V = concat(V) AL V D Vinfo (9)

[0113] H = XGBoost(V) (10)

[0114] Where V represents the concatenated vector of all feature vectors obtained through various feature extraction and encoding methods, and H represents the patent value level result calculated by the model.

[0115] Specific experiments and analysis

[0116] The basic data used in the experiments came from 2,000 patents in the field of artificial intelligence in the database of the China National Intellectual Property Administration, covering applications of AI technology in various fields. In addition, some experiments included 8,000 data entries augmented using the previously proposed data augmentation methods based on these patents. The value levels and corresponding tag values ​​of all patents are shown in Table 2.

[0117] Table 2. Statistical Information on Patent Value Levels

[0118] Patent Value Rating quantity High-value patents - 4 469 Second highest value patent - 3 2087 Medium-value patents-2 5692 Lowest Value Patent - 1 1417 Low-value patents - 0 335

[0119] Based on the data analysis in the table, the distribution of the number of patents at different value levels is more in line with a normal distribution, indicating that this dataset enables the model to fit the data better and has a wider range of applicability in statistical inference and estimation.

[0120] The experimental development environment for the research on the multi-indicator fusion method for evaluating the value of Chinese patents is based on Python 3.10.6. The specific experimental development environment parameters are shown in Table 3.

[0121] Table 3 Experimental Development Environment Parameters

[0122]

[0123] The relevant parameter settings for the Chinese patent value assessment experiment are shown in Table 4. The batch size was 4, the number of epochs was 20, the number of early stops was 10, the initial learning rate was 5e-5, the weight decay rate was 0.01, the number of warmup epochs was 4, the feature vector dimension was 768, and the dropout probability was 0.2. Due to the extremely large amount of training required for the experimental model parameters, the entire experiment was conducted on a GPU.

[0124] Table 4 Experimental Parameter Settings

[0125] Parameter name Parameter value batchsize 4 epoch 50 early stop 10 learningrate 5e-5 weightdecay 0.01 warmup 4 dim 768 dropout 0.2

[0126] The experiment uses commonly used evaluation metrics for text classification tasks, including precision (P), recall (R), and F1 score (F1-score), to evaluate and compare the experimental results. The calculation formulas for each metric are as follows:

[0127]

[0128] Where TP represents the number of samples that are actually positive and are predicted as positive, FP represents the number of samples that are actually negative but are predicted as positive, and FN represents the number of samples that are actually positive but are predicted as negative.

[0129] To verify the effectiveness of the proposed multi-indicator fusion method for patent valuation, comparative and ablation experiments were conducted. These experiments primarily included other commonly used machine learning text classification models and methods for combining different indicators.

[0130] To verify that our proposed method is significantly more effective than other patent valuation methods, we compared it with Logistic Regression (LR), K-Nearest Neighbor (KNN), Bayesian Model (BM), Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), and a fully connected layer plus BERT model. Furthermore, to demonstrate the effectiveness of the proposed data augmentation method for small-sample Chinese patent valuation, both the original dataset of 2000 patents and the augmented dataset of 10000 patents were evaluated using the aforementioned classification methods. The experimental results are shown in Table 5.

[0131] Table 5 Comparison of experimental results

[0132]

[0133] The experimental results in Table 5.5 using the original dataset of 2000 records demonstrate that the multi-indicator fusion method for assessing the value of Chinese patents proposed in this chapter significantly outperforms other traditional models. In the Chinese patent value assessment prediction experiments, both the method in this chapter and other traditional classification models used the extracted indicator feature vectors, quantized and concatenated, as input. Analysis shows that the method in this chapter improves the F1 score by 39.11% compared to the logistic regression model, 58.87% compared to the K-nearest neighbor classification algorithm, 43.77% compared to the Bayesian model, 42.71% compared to the support vector machine model, 20.10% compared to the decision tree model, and 13.84% compared to the random forest algorithm. Logistic regression is a statistical model used to solve classification problems by mapping experimental data to a probability space and then classifying data according to a threshold. However, logistic regression has limited ability to model nonlinear relationships and is sensitive to outliers, resulting in poor prediction results. The K-Nearest Neighbors (KNN) classification algorithm calculates the distance between the test set data and all samples in the training set, and determines the category of the test sample through a voting method. However, the KNN classification algorithm is prone to underfitting, can introduce bias when dealing with imbalanced datasets, and is sensitive to noisy data, resulting in poor prediction results. Bayesian models, based on Bayes' theorem, update model parameter estimates using prior probabilities and training sample data, and then infer the classification result of the predicted sample using posterior probabilities. Bayesian models assume that data features are independent, but the features in patent value assessment have strong correlations; this independence assumption leads to a decline in classification performance. Support Vector Machines (SVMs) models obtain a hyperplane in the training dataset to separate samples of different categories, and then predict the category of the data. SVM models are binary classification models, requiring extensions for multi-class tasks, and are sensitive to missing data, resulting in poor prediction results. Decision tree models construct a tree structure to segment input features, selecting different paths based on feature values ​​to ultimately obtain the prediction result. Random forest is an ensemble learning algorithm based on decision tree models. It classifies data by constructing multiple decision trees and combining their predictions. Decision tree models and random forests share a similar basic structure. When dealing with relationships between features, they recursively select the optimal features and thresholds to build tree branches, capturing simple interrelationships between data features. Random forests can comprehensively consider relationships between multiple feature subsets, thus better capturing complex interrelationships and achieving higher classification results than a single decision tree model. However, in random forests, each tree's feature selection and node splitting are performed independently, leading to discontinuous decision boundaries—gaps between sample points—resulting in poor final classification results.The XGBoost model used in this chapter differs from the Random Forest algorithm; it's an ensemble learning algorithm based on gradient boosting trees. Multiple decision trees are trained iteratively, with each tree incorporating the prediction bias of the previous tree and correcting for the errors. Finally, the predictions from all decision trees are summed to obtain the classification result. Therefore, the XGBoost model can handle high-dimensional and non-linear data, and its performance is progressively optimized through gradient boosting. Ultimately, the experimental results outperform other traditional machine learning classification models compared to this model.

[0134] Unlike traditional classification models that use word vectors as input, classification methods based on large language models can only achieve classification through end-to-end natural language dialogue via API calls. Inspired by agricultural news text classification methods based on large language models, this chapter uses a template prompting method based on Zhipu Qingyan to predict the value level of Chinese patents. Analysis of the experimental results shows that the large language model performs the worst, failing to achieve its superior performance in general domain tasks. This is because while large language models do possess powerful language understanding and reasoning abilities, their training datasets are all in general domains, making them unable to accurately adapt to the patent domain. Their professional knowledge and data correspondence capabilities are relatively lacking, leading to biases in professional citation. Furthermore, large language models adapt their input interfaces to data, requiring all patent value assessment indicators to be concatenated as natural language text before inputting into the large model. It can only process all indicators using the same embedding and decoding methods, failing to effectively integrate multiple types of indicators. Simultaneously, due to the limitation of input text length, only one sample data point can be input as a prompt command for each patent level, failing to effectively activate the potential of the large language model. Therefore, large language models cannot effectively predict the value of Chinese patents.

[0135] The analysis of the experimental results of the original dataset and the data-augmented dataset in Table 5 shows that using the data-augmented dataset has a positive effect on the classification result prediction of all models.

[0136] The multi-indicator fusion method for assessing the value of Chinese patents includes various types of features. To verify the impact of these features on the assessment results, this paper explores different feature combinations and tests them based on the proposed classification method. The traditional patent valuation indicators are combinations of indicators proposed by other scholars for assessing the value level of Chinese patents, including patent title, abstract, basic technical information (number of independent claims, number of dependent claims, etc.), basic economic information, and basic legal information. The experimental results are shown in Table 6.

[0137] Table 6 Ablation Experiment Results

[0138]

[0139] Analysis of the experimental results in Table 6 shows that each component of the evaluation indicators proposed in this chapter plays a positive role in the patent value assessment task, and the combined indicators achieve the best results. Traditional patent value assessment indicators (a) and (b) contain the same indicator characteristics, differing only in the method of obtaining the text feature vectors of the patent title and specification abstract. Traditional patent value assessment indicator (a) uses the BERT model for text encoding, while traditional patent value assessment indicator (b) uses the Large Language Model (GLM). Although it is impossible to achieve good experimental results for the value assessment task through question-and-answer methods using the large language model, its encoding capability is far superior to the pre-trained language model BERT. Its deeper model structure has more parameters and levels, allowing the GLM model to perform language modeling of the input text through multiple levels of abstract representation, thereby improving its representation ability. More training data helps the model better learn the statistical regularities and potential structures of the data, making the model more adaptable to a wider input space and less reliant on specific samples. Therefore, the GLM model's language modeling ability for small samples or even zero samples is superior to current pre-trained language models. Therefore, the GLM model was used for embedding short text features, which improved the F1 score by 2.44%. Basic enterprise information reflects the market size and market share of the patent applicant, helping to determine the commercial value of the patented technology in the market; this indicator improved the F1 score by 3.84%. The technology efficacy matrix, compared to traditional technology-related numerical indicators, extracts technology-related keywords and efficacy phrases from a textual perspective to determine which key technical indicators the patent has a high advantage in. Simultaneously, the technology efficacy matrix provides a structured framework within the field, allowing for an objective assessment of the value of the technology covered by the patent; this indicator improved the F1 score by 3.93%. The claims, as an important component of the patent document, describe in detail the technical fields and protection requirements covered by the patent, providing the specific boundaries and scope of patent rights. Through its content, one can understand the core technology and commercial application areas of the patent, thereby determining the patent's transaction and licensing potential; this indicator improved the F1 score by 4.55%.

[0140] The combined indicators presented in this chapter form the value assessment index system, which resulted in the best experimental results, with the F1 score increasing by 6.37%.

[0141] In summary, this invention proposes a Chinese patent value assessment index system. Building upon traditional Chinese patent value assessment methods, it adds claims and company information, as well as the technical efficacy matrix constructed in Chapter 3. Furthermore, for small sample texts, the semantic feature extraction modeling method for short texts is changed to a large language model (GLM). Then, for the acquired multi-type indicators, suitable quantification methods are used to concatenate the feature vectors. Finally, a classification model based on the XGBoost model is used to assess the patent value.

[0142] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. The above embodiments only illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be pointed out that for those skilled in the art, several modifications and improvements can be made without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A multi-index fused Chinese patent value evaluation method, characterized in that, The method comprises: Step 1, extracting patent text dimension features; including: extracting short text features from the abstract based on a large language model GLM, and extracting long text features from the claims based on an HBert model; Step 2, extracting patent basic information index features; including technical dimension indexes, economic dimension indexes, legal dimension indexes and enterprise dimension indexes; Step 3, based on the patent text dimension features and the patent basic information index features, using the XGBoost model to evaluate the patent value grade; In step 1, short text features are extracted from the abstract based on a large language model GLM; wherein: For the abstract, a large language model GLM model is used for feature embedding; the GLM model is based on the Transformer model, the GLM model is based on the 28-layer Transformer Encoder module, the order of normalization and residual connection is rearranged, a single linear layer is used for output word prediction, and GeLUs is used instead of ReLU activation function; The Lora algorithm is used to fine-tune the GLM model, in Lora, the parameters of the original large language model are frozen, the internal rank of the bypass simulation model parameters is increased, only the parameters in the bypass matrix are trained, and the parameters in the bypass matrix are superimposed with the original large language model; the text feature coding of the abstract is performed through the fine-tuned GLM model and the full connection layer, and the calculation formula is: GLM(W0+ΔW)=Lora(GLM(W0)) V A = GLM(A, W0+ ΔW) V AL = Linear(V A ) Wherein, W0 represents the original parameter of GLM model, AW represents the parameter of the instruction fine-tuning bypass matrix through the Lora algorithm, A represents the input abstract text, V AL represents the final obtained feature word vector, as a short text feature; In step S1, long text features are extracted from the claims based on an HBert model; wherein: Each sentence in long text D = (sen1, sen2, …, sen n ) is input into the Bert model to obtain a set of word vectors {S1, S2, …, S n} for each sentence in the long text. A randomly initialized word vector [SCLS] is concatenated in front of the set of word vectors, and {[SCLS], S1, S2, …, S n} is input into the Transformer Encoder model for encoding to obtain a word vector H SCLS containing the granularity of the long text and a word vector H SCLS The attention weight of the word vector of each sentence is calculated, and the weighted sum is obtained to represent the word vector V D of the long text D, and the calculation formula is:

2. The multi-index fused Chinese patent value evaluation method according to claim 1, characterized in that, In step 1, long text features are extracted from the claims based on an HBert model; wherein: The HBert model obtains long text features in a hierarchical manner; when the long text is divided according to sentences, the long text is represented as D=(sen1, sen2,..., sen n ), D represents the long text, sen i represents the i-th sentence of the long text, and the long text contains n sentences in total; when each sentence is divided according to characters, the sentence sen i is represented as sen i =(w i1 ,w i2 ,...,w im ), w ij represents the j-th character of the sentence sen i , and sen i has m characters in total; After dividing the long text, the sentence sen i =(w i1 ,w i2 ,...,w im ) is taken as the input of the Bert model, and the word vector H CLS of the sentence granularity and the word vector of the word granularity are obtained. CLS Attention weight calculation is performed on each word vector, and the word vector S i representing the sentence sen i is obtained by weighted summation, and the calculation formula is:

3. The multi-index fused Chinese patent value evaluation method according to claim 2, characterized in that, In step 2: The technical dimension indexes include: technical efficacy matrix, number of independent claims, number of dependent claims, number of pages of specification, number of examples, number of citations, number of cited documents, number of priority, number of technical classifications, number of division subcases, number of same family patents, patent award situation, decryption of confidential patent; The economic dimension indexes include: number of applicants, applicant type, authorization period, number of license records, number of right pledges, number of right transfers, customs record situation; The legal dimension indexes include: legal status, survival period / maintenance time, remaining life, litigation situation, number of invalidation requests; For the enterprise dimension indexes: If the enterprise is a company, the enterprise dimension indexes include: registered capital, number of insured persons, own risk number, associated risk number, historical risk number, and sensitive public opinion number; If the enterprise is a university, the enterprise dimension indexes include: university level, starting capital, own risk number, associated risk number, historical risk number, and sensitive public opinion number; The above indexes are divided into digital indexes, type indexes and matrix indexes, the representation forms of the digital indexes, type indexes and matrix indexes are obtained respectively, and the performance form vectors of all indexes are spliced to obtain the patent basic information index features.

4. The multi-index fused Chinese patent value evaluation method according to claim 3, characterized in that, In step 3: The XGBoost model is used as the patent value evaluation model. The XGBoost model adopts large-scale parallel gradient boosting tree-based ensemble learning algorithm, performs second-order Taylor expansion on the loss function, and introduces L1 and L2 regularization terms. The obtained patent text dimension features and patent basic information index features are spliced to obtain a patent feature vector, which is used as the input of the XGBoost model to perform a classification task and output the patent value grade. The calculation formula is: V = concat(V AL ,V D ,V info ) H=XGBoost(V) wherein, V AL represents the short text feature in the patent text dimension feature, V D represents the long text feature in the patent text dimension feature, V represents the patent feature vector, and H represents the patent value grade result calculated by the model.

5. A multi-index fused Chinese patent value evaluation system, characterized in that, The system comprises a processing unit configured to perform the following steps: Extracting patent text dimension features, including extracting short text features from the abstract based on a large language model GLM and extracting long text features from the claims based on an HBert model; Extracting patent basic information index features, including technical dimension indexes, economic dimension indexes, legal dimension indexes, and enterprise dimension indexes; Using the XGBoost model to evaluate the patent value grade based on the patent text dimension features and the patent basic information index features; The processing unit is specifically configured to perform the following steps: For the abstract, a large language model GLM model is used for feature embedding. The GLM model is based on the Transformer model. The GLM model rearranges the order of normalization and residual connection based on a 28-layer Transformer Encoder module, uses a single linear layer for output word prediction, and replaces the ReLU activation function with GeLUs. The Lora algorithm is used to fine-tune the GLM model. In Lora, the parameters of the original large language model are frozen, the intrinsic rank of the bypass model parameters is increased, only the parameters in the bypass matrix are trained, and the parameters in the bypass matrix are superimposed with the original large language model. The text feature encoding of the abstract is performed through the fine-tuned GLM model and the full connection layer, and the calculation formula is: GLM(W0+ΔW)=Lora(GLM(W0)) V A = GLM(A, W0+ ΔW) V AL = Linear(V A ) Wherein, W0 represents the original parameter of GLM model, AW represents the parameter of the instruction fine-tuning bypass matrix through the Lora algorithm, A represents the input abstract text, V AL represents the final obtained feature word vector, as a short text feature; The processing unit is specifically configured to perform the following steps: The long text D = (sen1, sen2, ..., sen) n Each sentence in the text is input into the BERT model to obtain a set of word vectors {S1, S2, ..., S} for each sentence in the long text. n }, prepend a randomly initialized word vector [SCLS] before the set of word vectors, and set {[SCLS], S1, S2, ..., S...} to the set of word vectors. n As input, the words are encoded by the Transformer Encoder model to obtain word vectors H containing long text granularity. SCLS and sentence-level word vectors H SCLS Attention weights are calculated for the word vectors of each sentence, and the word vector V representing the long text D is obtained by weighted summation. D The calculation formula is:

6. The multi-index fused Chinese patent value evaluation system according to claim 5, characterized in that, The processing unit is specifically configured to perform the following steps: The HBert model obtains long text features in a hierarchical manner; the long text is divided according to sentences, and the long text is represented as D=(sen1, sen2,..., sen n ), D represents the long text, sen i represents the i-th sentence of the long text, and the long text contains n sentences in total; each sentence is divided according to characters, and the sentence sen i is represented as sen i =(w i1 ,w i2 ,...,w im ), w ij represents the j-th character of the sentence sen i , and sen i has m characters in total; After dividing the long text, the sentence sen i =(w i1 ,w i2 ,...,w im ) is taken as the input of the Bert model to obtain the word vector H CLS of the sentence granularity and the word vector of the word granularity. CLS Attention weight calculation is performed on each word vector, and the word vector S i representing the sentence sen i is obtained by weighted summation, and the calculation formula is:

7. The multi-index fused Chinese patent value evaluation system according to claim 6, characterized in that, The processing unit is specifically configured to perform the following steps: The technical dimension indexes include: technical efficacy matrix, number of independent claims, number of dependent claims, number of pages of specification, number of examples, number of citations, number of cited documents, number of priorities, number of technical classifications, number of division subcases, number of same family patents, patent award situation, and decryption of confidential patents. The economic dimension indexes include: number of applicants, type of applicant, authorization period, number of license filings, number of right pledges, number of right transfers, and customs filing situation. The legal dimension indexes include: legal status, survival period / maintenance time, remaining life, litigation situation, and number of invalidation requests. For the enterprise dimension indexes: If the enterprise is a company, the enterprise dimension indicators include: registered capital, number of insured persons, self-risk number, associated risk number, historical risk number, and sensitive public opinion number; If the enterprise is a college, the enterprise dimension indicators include: college level, starting capital, self-risk number, associated risk number, historical risk number, and sensitive public opinion number; The above indicators are divided into digital indicators, type indicators, and matrix indicators, and the representation forms of the digital indicators, the type indicators, and the matrix indicators are obtained respectively, and the performance form vectors of all the indicators are spliced to obtain patent basic information indicator characteristics.

8. The multi-index fused Chinese patent value evaluation system according to claim 7, characterized in that, The processing unit is specifically configured to perform: An XGBoost model is adopted as the patent value evaluation model, XGBoost adopts a large-scale parallel gradient boosting tree-based ensemble learning algorithm, performs a second-order Taylor expansion on a loss function, and simultaneously introduces L1 and L2 regularization terms; the obtained patent text dimension characteristics and patent basic information indicator characteristics are spliced to obtain a patent feature vector as an input of the XGBoost model, a classification task is performed, and a patent value level is output; the calculation formula is: V = concat(V AL ,V D ,V info ) H=XGBoost(V) wherein, V AL represents the short text feature in the patent text dimension feature, V D represents the long text feature in the patent text dimension feature, V represents the patent feature vector, and H represents the patent value grade result calculated by the model.

Citation Information

Patent Citations

  • Early evaluation method of patent value

    CN115186982A

Cited By

  • Method and device for evaluating patent value in biomedicine field based on transformer

    CN122471333A