Multi-index fusion Chinese patent value evaluation method

By combining the large language model and the XGBoost model, patent text and basic information indicators are extracted, and the accuracy of Chinese patent value evaluation on the small sample data set is solved, achieving relatively objective and accurate evaluation results.

CN119991157AActive Publication Date: 2025-05-13MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY +1

Patent Information

Application Number
CN202411597818.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-05-13
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate Chinese patent value assessment on small sample data sets, mainly due to the scarcity of Chinese patent value data and the diverse evaluation criteria.

Method used

A multi-index fusion of Chinese patent value evaluation method is proposed, combining a large language model, and extracting patent text dimension features and patent basic information index features, and using the XGBoost model to evaluate patent value levels.

Benefits of technology

Under the conditions of scarce data, relatively objective and accurate patent value evaluation results are achieved, effectively improving the evaluation effect on small sample data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991157A_ABST
    Figure CN119991157A_ABST
Patent Text Reader

Abstract

The invention discloses a Chinese patent value evaluation method based on multi-index fusion, and belongs to the technical field of patent value evaluation. The method comprises the steps of 1, extracting patent text dimension features; comprising the steps of extracting short text features from a specification abstract based on a large language model GLM, and extracting long text features from a claim based on an HBert model; 2, extracting patent basic information index characteristics; comprising a technical dimension index, an economic dimension index, a legal dimension index and an enterprise dimension index; and step 3, based on the patent text dimension features and the patent basic information index features, performing patent value grade evaluation by using an XGBoost model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of patent value assessment, and in particular to a multi-index fusion Chinese patent value assessment method. Background Art

[0002] With the continuous deepening of a new round of technological revolution and industrial transformation, the importance of scientific and technological innovation has become increasingly prominent. It is of great significance to accurately and objectively identify high-value patents. The means of patent value assessment are constantly being updated and improved. Initially, the value assessment method mainly relied on the basic attributes of the patent combined with the opinions of experts in the field. Later, with the rise of artificial intelligence technology, many scholars began to use machine learning and deep learning to study the value assessment methods of Chinese patents. However, due to the scarcity of value data of Chinese patents and the diversity of evaluation standards, it has not been possible to achieve a more accurate value assessment based on a small sample data set. Summary of the invention

[0003] In view of the above technical problems, the present invention proposes a Chinese patent value assessment method based on a small sample data set and a multi-index fusion combined with a large language model, namely a multi-index fusion Chinese patent value assessment scheme. This scheme can achieve more objective and accurate results under the condition of scarce data.

[0004] The first aspect of the present invention discloses a multi-index fusion Chinese patent value assessment method, the method comprising:

[0005] Step 1: Extracting patent text dimension features; including: extracting short text features from the abstract of the specification based on the large language model GLM, and extracting long text features from the claims based on the HBert model;

[0006] Step 2: Extract the basic information indicator characteristics of the patent, including technical dimension indicators, economic dimension indicators, legal dimension indicators and enterprise dimension indicators;

[0007] Step 3: Based on the patent text dimension characteristics and patent basic information indicator characteristics, use the XGBoost model to evaluate the patent value level.

[0008] According to the method of the first aspect of the present invention, in step 1, short text features are extracted from the specification abstract based on the large language model GLM; wherein:

[0009] For the summary of the instruction manual, the large language model GLM model is used for feature embedding; the GLM model is based on the Transformer model. The GLM model is based on a 28-layer Transformer Encoder module, rearranges the order of normalization and residual connections, uses a single linear layer to predict the output word, and replaces the ReLU activation function with GeLUs;

[0010] The Lora algorithm is used to fine-tune the GLM model. In Lora, the parameters of the original large language model are frozen, the intrinsic rank of the bypass simulation model parameters is increased, only the parameters in the bypass matrix are trained, and the parameters in the bypass matrix are superimposed on the original large language model; the text feature encoding of the instruction manual summary is performed through the fine-tuned GLM model and the fully connected layer, and the calculation formula is:

[0011] GLM(W0+ΔW)=Lora(GLM(W0))

[0012] V A =GLM(A,W0+ΔW)

[0013] V AL =Linear(V A )

[0014] Where W0 represents the original parameters of the GLM model, ΔW represents the parameters of the instruction fine-tuning bypass matrix through the Lora algorithm, A represents the input instruction summary text, V AL Represents the final feature word vector obtained as a short text feature.

[0015] According to the method of the first aspect of the present invention, in step 1, long text features are extracted from the claims based on the HBert model; wherein:

[0016] The HBert model obtains long text features in a hierarchical manner; the long text is divided into sentences, and the long text is represented as D = (sen1, sen2, ..., sen n ), D indicates long text, sen i represents the i-th sentence of a long text, which contains n sentences in total; if each sentence is split by characters, the sentence sen i Expressed as sen i =(w i1 ,w i2 ,…,w im ), w ij Indicates the sentence sen i The jth character of sen i There are m characters in total;

[0017] After dividing the long text into sentences, i =(w i1 ,w i2 ,...,w im ) as the input of the Bert model to obtain the word vector H at the sentence granularity CLS And word vectors at word granularity H CLSCalculate the attention weight for each word vector and get the representative sentence sen by weighted summation i The word vector S i , the calculation formula is:

[0018]

[0019] The long text D = (sen1, sen2, ..., sen n ) is input into the Bert model, and the word vector set {S1,S2,...,S n}, concatenate a randomly initialized word vector [SCLS] before the word vector set, and replace {[SCLS], S1, S2, ..., S n} is used as input and encoded through the Transformer Encoder model to obtain a word vector H containing the granularity of long text SCLS and sentence-level word vectors H SCLS Calculate the attention weight for the word vector of each sentence, and obtain the word vector V representing the long text D by weighted summation D , the calculation formula is:

[0020]

[0021] According to the method of the first aspect of the present invention, in step 2:

[0022] Technical dimension indicators include: technical efficacy matrix, number of independent claims, number of dependent claims, number of pages in the specification, number of embodiments, number of citations, number of cited documents, number of priorities, number of technical classifications, number of sub-cases, number of patent families, patent awards, and decryption of confidential patents;

[0023] Economic dimension indicators include: number of applicants, type of applicants, authorization cycle, number of license filings, number of rights pledges, number of rights transfers, and customs filing status;

[0024] Legal dimension indicators include: legal status, survival period / maintenance time, remaining life, litigation status, and number of invalidation requests;

[0025] Enterprise dimension indicators include: company registered capital, number of insured persons, number of own risks, number of related risks, number of historical risks, number of sensitive public opinions; university level, start-up capital, number of own risks, number of related risks, number of historical risks, number of sensitive public opinions;

[0026] The above indicators are divided into digital indicators, categorical indicators and matrix indicators, and the representation forms of digital indicators, categorical indicators and matrix indicators are obtained respectively. The representation vectors of all indicators are spliced ​​to obtain the basic information indicator characteristics of the patent.

[0027] According to the method of the first aspect of the present invention, in step 3:

[0028] The XGBoost model is used as the patent value evaluation model. XGBoost uses a large-scale parallel ensemble learning algorithm based on gradient boosting trees, performs a second-order Taylor expansion on the loss function, and introduces L1 and L2 regularization terms. The obtained patent text dimension features and patent basic information indicator features are concatenated to obtain the patent feature vector as the input of the XGBoost model, perform classification tasks, and output the patent value level. The calculation formula is:

[0029] V = concat(V AL ,V D ,V info )

[0030] H=XGBoost(V)

[0031] Among them, V AL Represents the short text feature in the patent text dimension feature, V D It represents the long text feature in the patent text dimension feature, V represents the patent feature vector, and H represents the patent value level result calculated by the model.

[0032] The second aspect of the present invention discloses a multi-index fusion Chinese patent value assessment system, the system comprising a processing unit, and the processing unit is configured to perform the following steps:

[0033] Extracting patent text dimension features; including: extracting short text features from the abstract of the specification based on the large language model GLM, and extracting long text features from the claims based on the HBert model;

[0034] Extract the basic information indicator characteristics of patents, including technical dimension indicators, economic dimension indicators, legal dimension indicators and enterprise dimension indicators;

[0035] Based on the patent text dimension characteristics and patent basic information indicator characteristics, the XGBoost model is used to evaluate the patent value level.

[0036] According to the system of the second aspect of the present invention, the processing unit is specifically configured to perform, based on the large language model GLM, extracting short text features from the specification abstract; wherein:

[0037] For the summary of the instruction manual, the large language model GLM model is used for feature embedding; the GLM model is based on the Transformer model. The GLM model is based on a 28-layer Transformer Encoder module, rearranges the order of normalization and residual connections, uses a single linear layer to predict the output word, and replaces the ReLU activation function with GeLUs;

[0038] The Lora algorithm is used to fine-tune the GLM model. In Lora, the parameters of the original large language model are frozen, the intrinsic rank of the bypass simulation model parameters is increased, only the parameters in the bypass matrix are trained, and the parameters in the bypass matrix are superimposed on the original large language model; the text feature encoding of the instruction manual summary is performed through the fine-tuned GLM model and the fully connected layer, and the calculation formula is:

[0039] GLM(W0+ΔW)=Lora(GLM(W0))

[0040] V A =GLM(A,W0+ΔW)

[0041] V AL =Linear(V A )

[0042] Where W0 represents the original parameters of the GLM model, ΔW represents the parameters of the instruction fine-tuning bypass matrix through the Lora algorithm, A represents the input instruction summary text, V AL Represents the final feature word vector obtained as a short text feature.

[0043] According to the system of the second aspect of the present invention, the processing unit is specifically configured to perform, based on the HBert model, extracting long text features from the claims; wherein:

[0044] The HBert model obtains long text features in a hierarchical manner; the long text is divided into sentences, and the long text is represented by D = (sen1, sen2, ..., sen n ), D indicates long text, sen i represents the i-th sentence of a long text, which contains n sentences in total; if each sentence is split by characters, the sentence sen i Expressed as sen i =(w i1 ,w i2 ,...,w im ), w ij Indicates the sentence sen i The jth character of sen i There are m characters in total;

[0045] After dividing the long text into sentences,i =(w i1 ,w i2 ,...,w im ) as the input of the Bert model to obtain the word vector H at the sentence granularity CLS And word vectors at word granularity H CLS Calculate the attention weight for each word vector and get the representative sentence sen by weighted summation i The word vector S i , the calculation formula is:

[0046]

[0047] The long text D = (sen1, sen2, ..., sen n ) is input into the Bert model, and the word vector set {S1,S2,…,S n}, concatenate a randomly initialized word vector [SCLS] before the word vector set, and replace {[SCLS], S1, S2, …, S n} is used as input and encoded through the Transformer Encoder model to obtain a word vector H containing the granularity of long text SCLS and sentence-level word vectors H SCLS Calculate the attention weight for the word vector of each sentence, and obtain the word vector V representing the long text D by weighted summation D , the calculation formula is:

[0048]

[0049] According to the system of the second aspect of the present invention, the processing unit is specifically configured to execute:

[0050] Technical dimension indicators include: technical efficacy matrix, number of independent claims, number of dependent claims, number of pages in the specification, number of embodiments, number of citations, number of cited documents, number of priorities, number of technical classifications, number of sub-cases, number of patent families, patent awards, and decryption of confidential patents;

[0051] Economic dimension indicators include: number of applicants, type of applicants, authorization cycle, number of license filings, number of rights pledges, number of rights transfers, and customs filing status;

[0052] Legal dimension indicators include: legal status, survival period / maintenance time, remaining life, litigation status, and number of invalidation requests;

[0053] Enterprise dimension indicators include: company registered capital, number of insured persons, number of own risks, number of related risks, number of historical risks, number of sensitive public opinions; university level, start-up capital, number of own risks, number of related risks, number of historical risks, number of sensitive public opinions;

[0054] The above indicators are divided into digital indicators, categorical indicators and matrix indicators, and the representation forms of digital indicators, categorical indicators and matrix indicators are obtained respectively. The representation vectors of all indicators are spliced ​​to obtain the basic information indicator characteristics of the patent.

[0055] According to the system of the second aspect of the present invention, the processing unit is specifically configured to execute:

[0056] The XGBoost model is used as the patent value evaluation model. XGBoost uses a large-scale parallel ensemble learning algorithm based on gradient boosting trees, performs a second-order Taylor expansion on the loss function, and introduces L1 and L2 regularization terms. The obtained patent text dimension features and patent basic information indicator features are concatenated to obtain the patent feature vector as the input of the XGBoost model, perform classification tasks, and output the patent value level. The calculation formula is:

[0057] V = concat(V AL ,V D ,V info )

[0058] H=XGBoost(V)

[0059] Among them, V AL Represents the short text feature in the patent text dimension feature, V D It represents the long text feature in the patent text dimension feature, V represents the patent feature vector, and H represents the patent value level result calculated by the model.

[0060] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the multi-index fusion Chinese patent value assessment method described in the first aspect of the present disclosure is implemented.

[0061] The fourth aspect of the present invention discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multi-index fusion Chinese patent value assessment method described in the first aspect of the present disclosure is implemented.

[0062] In summary, the present invention proposes an indicator standard for patent value assessment, including five dimensions: patent text dimension, technical dimension, economic dimension, legal dimension and enterprise information dimension, and each type of indicator contains a variety of subtle features, and for the first time combines short text and long text, and patent technology efficacy matrix in the value assessment indicator. The present invention proposes a small sample Chinese patent value assessment method combined with a large language model. This method integrates and utilizes all the above-mentioned patent indicators, applies the large language model to the field of patent value assessment for the first time, and effectively improves the evaluation effect of patent value on small sample data sets. Based on artificial intelligence patent data, the present invention has conducted a large number of experiments to verify the accuracy of the small sample Chinese patent value assessment method based on the large language model and the effectiveness of each part of the indicators. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0064] Figure 1 This is a framework diagram of the Chinese patent value assessment model integrating multiple indicators;

[0065] Figure 2 This is the framework diagram of the Hbert model. DETAILED DESCRIPTION

[0066] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0067] Patent value varies in different evaluation contexts, because the right holders have different orientations and research angles on patent value, and different scholars have different evaluation standards for patent value. Patent value evaluation is a process in which a third-party evaluation agency, with relevant qualifications, uses asset appraisers to quantify the value of patents owned by a specific entity in accordance with relevant procedures and methods to meet the requirements of all parties in patent transactions. Patent value evaluation has always been the focus of research in the patent field. Current patent value evaluation methods can be divided into three categories: economic methods, comprehensive evaluation methods, and other emerging evaluation methods.

[0068] 1) Economic Method

[0069] The analysis demonstrates the applicability of the replacement cost method to patent value analysis in equipment procurement. According to the patent avoidance design concept, the concept of patent replacement cost is appropriately extended, the implementation steps of the replacement cost evaluation method are given, and the patent value analysis process based on patent replacement cost is demonstrated with examples. It is proposed that the establishment of a database of intellectual property value evaluation parameters by fully combining digital technology, the improvement of the comparable indicator system of the intellectual property value evaluation market method, and the reasonable determination of the comparable factor correction coefficient are the key points and key points of using the market method to evaluate intellectual property in the future.

[0070] 2) Comprehensive evaluation method

[0071] The comprehensive evaluation method mainly uses a combination of fuzzy comprehensive evaluation method, hierarchical analysis method, binary classification evaluation method, principal component analysis method, entropy weight method, etc. to evaluate the value of patents. The key to the application of this method is to construct a patent value evaluation index system, which is a relatively common method. Based on the value distribution characteristics of patent assets, by analyzing and comparing the key indicators of domestic and foreign patent value evaluation and the influencing factors of patent value, a binary classification evaluation method for patent value is constructed. The patents authorized by the Chinese Academy of Sciences in the United States in the past 28 years are studied, and the evaluation index weights are calculated in combination with the principal component method. This method can overcome the influence of subjective factors on the evaluation results.

[0072] 3) New evaluation methods

[0073] In recent years, many scholars have conducted relevant research on the improvement of patent value assessment methods. Under the background of big data and the urgent need for assessment methods to keep pace with the times, some other emerging methods have also emerged, mainly machine learning methods, system dynamics methods, etc. The rough set theory is used to simplify the patent value assessment index system, and then a patent value assessment model based on BP neural network is established. From the perspective of machine learning technology, the patent value assessment indicators are first analyzed and selected, and then the three algorithms of decision tree, support vector machine and neural network in machine learning methods are used to train and test the samples, and finally the test results are analyzed.

[0074] At present, the traditional methods of automatic patent value evaluation have low accuracy, and the new methods are still in the early stages. Most of them use formula calculations. The new methods using artificial intelligence technology do not take into account the impact of the original content on the patent value and the problem of sample data volume. The automatic evaluation method of Chinese patent value using the combination of prompt learning and large language model that integrates the patent technology efficacy matrix with other digital indicators can not only overcome the problem of multi-dimensional fusion of evaluation indicators, but also overcome the problem of insufficient data samples, so it has important research value.

[0075] Language modeling is a basic task in the field of natural language processing. It aims to predict and generate continuous text sequences that conform to grammatical and semantic rules through learning text data, and lay the foundation for other natural language processing tasks. The introduction of pre-trained language models based on the Transformer framework has made language modeling a great success.

[0076] With the release of ChatGPT in November 2022, large language models began to enter the public eye and demonstrated amazing capabilities in a range of natural language processing tasks. A large language model refers to a Transformer language model that contains hundreds of billions or more parameters and is trained on a large-scale general corpus of hundreds of megabytes. The existing mainstream architectures of large language models can be divided into three types, including encoder-decoder architecture, causal decoder architecture, and prefix decoder architecture. The large language model based on the encoder-decoder architecture has stronger understanding and encoding capabilities for the input text but poor long text generation capabilities. The large language model based on the causal decoder architecture has stronger text generation capabilities and higher training efficiency. The capabilities of the model based on the prefix decoder architecture are a compromise between the capabilities of the above two architecture models. Due to the powerful parameters and training data of the large language model itself, these three types of large language models can all exert strong semantic understanding and generation capabilities in the field of general knowledge, demonstrating reasoning capabilities that previous neural network models did not have.

[0077] At present, many scholars at home and abroad have applied large language models to various downstream tasks. Inspired by the potential application of ChatGPT in the field of smart agriculture, an agricultural text classification method based on ChatGPT is proposed, which proves that large language models can effectively improve the accuracy of agricultural text classification and alleviate the problem of insufficient data samples in this field, and is highly feasible in promoting agricultural practice. It is proposed to combine the SetFit method with the GPT3.5 and GPT4 models to perform small sample classification of text in the financial field. This method obtains state-of-the-art results and provides a practical solution for small sample tasks. An entity language model is proposed to integrate the continuous sensor modality of the real world into the large language model, thereby realizing robot operation planning, visual question answering and description functions.

[0078] Practice has proven that large language models can handle various complex AI tasks, and due to the complexity of their model framework and the richness of their general training corpus, large language models can achieve superior results on small or even zero sample data. However, no scholar has yet applied it to the field of patent value assessment. This paper will be the first to use a large language model to integrate and utilize various patent indicators to conduct small sample Chinese patent value assessment.

[0079] The framework of the Chinese patent value assessment model based on multi-index integration is as follows: Figure 1 As shown, the model mainly consists of three parts.

[0080] The first part is the feature extraction of patent text dimensions, including short text feature extraction based on the large language model GLM model and long text feature extraction based on the HBert model. The second part is the extraction of basic patent information indicators, including four dimensions: technical dimension, economic dimension, legal dimension and corporate information dimension, and each type of indicator contains a variety of subtle features. The third part is a patent value rating evaluation model based on the XGBoost model as the basic classification model, which vectorizes and splices all the above patent evaluation indicators, and predicts the final value level through this model.

[0081] The abstract of the specification and the claims are important components of the application documents of invention patents. The abstract of the specification provides a concise and accurate summary of the entire patent application, summarizing the core features, technical problems, solutions and advantages of the invention or innovation in simple language. The claims define and determine the scope and limits of the patent rights, and specifically list the various elements, features and combinations of the invention or innovation. To a certain extent, the abstract of the specification can preliminarily evaluate the feasibility and innovative value of the patent, and the claims can provide the specific technical features and application fields of the patent, which can further determine the commercial value of the patent. However, due to the difficulty of processing long texts, scholars in the past have obviously not been sufficient to use only the abstract of the specification as a text indicator for patent value assessment. Therefore, based on previous research, this article adds the claims information for the first time, and uses the abstract of the specification and the claims as text features for patent value assessment.

[0082] For the instruction manual summary, this article uses the large language model GLM model for feature embedding. The GLM model is based on the Transformer model, and the model structure and training data have been greatly improved compared to the pre-trained language model in terms of model structure and training data set. The GLM model is based on a 28-layer Transformer Encoder module. In addition to the increase in model layers, three improvements have been made: 1) The order of layer normalization and residual connection has been rearranged; 2) A single linear layer is used to predict the output word; 3) The ReLU activation function is replaced by GeLUs. In addition, the GLM model uses more diverse and sufficient training data, making the model language modeling capabilities more powerful.

[0083] In order to make the feature vector generation of the GLM model more suitable for subsequent value assessment tasks, instruction fine-tuning operations are required. However, the training cost of the large language model is extremely high, and it is impossible to retrain all model parameters. Therefore, the Lora method is used to fine-tune the instructions of the GLM model. The Lora method freezes the parameters of the original large language model, simulates the intrinsic rank of the model parameters by adding a bypass, only trains the parameters in the bypass matrix, and finally superimposes the parameters in the bypass matrix with the original large language model. The text feature encoding of the instruction manual summary is then performed through the fine-tuned GLM model and the fully connected layer. The calculation formula is as follows:

[0084] GLM(W0+ΔW)=Lora(GLM(W0)) (1)

[0085] V A =GLM(A,W0+ΔW) (2)

[0086] V AL =Linear(V A ) (3)

[0087] Where W0 represents the original parameters of the GLM model, ΔW represents the parameters of the instruction fine-tuning bypass matrix through the Lora method, A represents the input instruction summary text, V AL Represents the final feature word vector obtained.

[0088] For the rights specification, since its text length is not limited, this paper adopts the HBert model proposed in previous studies to process long texts for feature embedding. The HBert model framework is shown in the figure below: Figure 2 shown.

[0089] The model obtains text vectors in a hierarchical manner. According to language habits, a sentence usually expresses a complete semantic and sentence structure. If long text is segmented according to a fixed character length, the segmentation position cannot be guaranteed, which will cause the semantic structure to be destroyed. Therefore, the long text is first divided into sentences, and the long text can be represented as D = (sen1, sen2, ..., sen n ), D indicates long text, sen i represents the i-th sentence of a long text, which contains n sentences in total. Then, each sentence can be split into i Expressed as sen i =(w i1 ,w i2 ,…,w im ), w ij Indicates the sentence sen i The jth character of sen i There are m characters in total.

[0090] After dividing the long text into the above method, first divide the sentences into i =(w i1 ,w i2 ,…,w im ) as the input of the Bert model to obtain the word vector H at the sentence granularity CLS And word vectors at word granularity H CLS Calculate the attention weight for each word vector, and then perform weighted summation to get the representative sentence sen i The word vector S i , the calculation formula is as follows:

[0091]

[0092] The long text D = (sen1, sen2, ..., sen n ) is input into the Bert model in the above way, and the sentence vector set {S1,S2,...,S n}. Then, a randomly initialized word vector [SCLS] is concatenated before the word vector set, and {[SCLS], S1, S2, …, S n} as input, and then encoded by the Transformer Encoder model to obtain the word vector H containing the long text granularity. SCLS and sentence-level word vectors Then H SCLS Calculate the attention weights for the word vectors of each sentence, and then perform weighted summation to obtain the word vector V representing the long text D D , the calculation formula is as follows:

[0093]

[0094] Many scholars have proposed different indicators or combinations of indicators for basic information indicators related to patent value assessment. Based on previous research, this article proposes four dimensions: technical dimension, economic dimension, legal dimension and enterprise information dimension. The specific patent characteristic indicators included in each category are shown in Table 1.

[0095] Table 1 Basic information indicators of patents

[0096]

[0097]

[0098] The technical dimension mainly reflects the technological advancement and innovation of the patent solution. It is field-oriented and can reflect the technical content and technical advantages of the invention patent, thereby affecting the commercial application and market potential of the patent. This paper introduces the technical efficacy matrix into the patent value evaluation index system for the first time. The technical efficacy matrix can largely reflect the saturation and innovation of the patent market. The values ​​in it can reflect the concentration of patent applications in key technical points and efficacy, which can play a positive role in the evaluation of patent value to a certain extent.

[0099] The economic dimension mainly reflects the investment of the patentee in the invention patent process and the income and profit obtained from the use of the patent. The characteristics include the number of applicants, the type of applicants (enterprises, colleges and universities, scientific research institutions, individuals, government agencies, and other), the authorization cycle, the number of license filings, the number of rights pledges, the number of power transfers, and the customs filing status. These characteristics can help determine the commercial feasibility and commercial value of patents and help enterprises make decisions, including the commercial use of patents, licensing, market positioning, etc.

[0100] The legal dimension mainly reflects the scope of legal protection, validity and compliance of the patent, and its characteristics include legal status (rights-approval, substantive examination, invention disclosure, no rights-rejection, no rights-deemed withdrawal, no rights-voluntary withdrawal, no rights-non-payment of annual fees, no rights-expiration termination, no rights-deemed abandonment, no rights-voluntary abandonment, no rights-avoidance abandonment), survival period / maintenance time, remaining life, litigation status and number of invalidation requests. These characteristics can help determine the strength of legal protection of patents and the stability of patent rights, and further affect the commercial value and legal risks of patents.

[0101] The enterprise information dimension mainly reflects the background, strength and strategy of the enterprise to which the patent belongs. Enterprises are mainly divided into two categories: companies and universities. Company characteristics include registered capital, number of insured persons, number of own risks, number of associated risks, number of historical risks, and number of sensitive public opinions; university characteristics include university level, start-up capital, number of own risks, number of associated risks, number of historical risks, and number of sensitive public opinions. These characteristics can help determine the importance of patents in corporate strategy and development, as well as the ability and resource support of enterprises in the process of patent commercialization.

[0102] Comprehensive consideration of the evaluation results of these dimensions can help determine the value and potential of patents and guide patent management and commercialization decisions.

[0103] The first part is the feature extraction of patent text dimension information. The feature vector V of the specification abstract is obtained by using the GLM model and HBert model. AL and the claim description feature vector V D .

[0104] The second part is the feature extraction of basic patent information indicators. The indicators in the technical, economic, legal and corporate information dimensions can be divided into three types of data:

[0105] The first is the numerical indicators, such as the number of applicants, registered capital, etc., which are fixed numerical values. For this type of indicators, this article uses Z-score to quantify the numerical indicators, and the formula is as follows:

[0106]

[0107] Where μ is the mean and σ is the variance. Z-score transforms the original data into a standard normal distribution, so that the data has a mean of 0 and a standard deviation of 1. It can eliminate the dimensional differences between different data and transform them into comparable standard scores.

[0108] The second is the categorical indicators, such as customs filing status, applicant type, etc., which are the values ​​of an element in a fixed type array. For this type of indicators, this paper uses One-hot encoding to represent categorical variables as binary vectors, using the number 1 to indicate the existence of the category and the number 0 to indicate the absence of the category. This encoding provides a clear and interpretable representation of categorical variables.

[0109] The third is matrix indicators, such as the technical effectiveness matrix. For this type of indicators, this paper uses the matrix normalization method to map the value range of the matrix to [0,1], eliminate dimensional differences, and avoid certain features from having too much influence on the model.

[0110] Through the above method, we can obtain the representation of digital, categorical and matrix indicators, and then concatenate the vectors of all indicators to obtain the patent basic information indicator characteristic vector V info .

[0111] The third part is the patent value assessment part. Taking into account the problems of overfitting, multicollinearity and training speed, the XGBoost model is used as the patent value assessment model. XGBoost is a large-scale parallel ensemble learning algorithm based on gradient boosting tree. XGBoost performs a second-order Taylor expansion on the loss function and introduces L1 and L2 regularization terms. It can control the complexity of the model, speed up the convergence of the model, reduce the risk of overfitting, and ensure the robustness of the model. In addition, XGBoost is not affected by highly correlated features, which reduces the problem of multicollinearity of features. All the vectors obtained above are concatenated to obtain a patent vector containing rich patent feature information, which is used as the input of the XGBoost model for classification tasks, and finally the value level of the patent is output. The formula is as follows:

[0112] V = concat(V AL ,V D ,Vinfo ) (9)

[0113] H=XGBoost(V) (10)

[0114] Among them, V represents the concatenated vector of all feature vectors obtained by various feature extraction and encoding methods, and H represents the patent value level result obtained by model calculation.

[0115] Specific experiments and analysis

[0116] The basic data used in the experiment comes from 2,000 patents in the field of artificial intelligence in the database of the State Intellectual Property Office of China, including the application of artificial intelligence technology in various directions. In addition, some experiments also include 8,000 pieces of data based on these patents using the data enhancement method proposed in the previous section. The value level of all patents and the amount of data corresponding to the label value are shown in Table 2.

[0117] Table 2 Statistical information on patent value levels

[0118] Patent value level quantity High value patents-4 469 Second Highest Value Patent-3 2087 Medium value patent-2 5692 Second lowest value patent -1 1417 Low value patents - 0 335

[0119] According to the data analysis in the table, the distribution of the number of patent value levels is more consistent with the normal distribution, which means that this data set can make the model fit the data better and has a wider applicability when performing statistical inference and estimation.

[0120] The experimental development environment for the study of the multi-index fusion Chinese patent value assessment method is based on python3.10.6. The specific experimental development environment parameters are shown in Table 3:

[0121] Table 3 Experimental development environment parameters

[0122]

[0123] The relevant parameter settings for the evaluation of the value level of Chinese patents in the experiment are shown in Table 4. The batch size (batchsize) is 4, the number of learning rounds (epoch) is 20, the number of early stop rounds (early stop) is 10, the initial learning rate (learning rate) is 5e-5, the learning rate decay rate (weight decay) is 0.01, the number of learning rate scheduling warmup rounds (warmup) is 4, the feature vector dimension (dim) is 768, and the random inactivation probability (dropout) is 0.2. Due to the large amount of experimental model parameter training, the entire experiment process is carried out on the GPU.

[0124] Table 4 Experimental parameter settings

[0125] Parameter name Parameter Value batchsize 4 epoch 50 earlystop 10 learningrate 5e-5 weightdecay 0.01 warmup 4 dim 768 dropout 0.2

[0126] The experiment uses commonly used evaluation indicators for text classification tasks, including precision (P, precision), recall (R, recall), and F1 score (F1, f1-score), to evaluate and compare the experimental results. The calculation formulas for each indicator are as follows:

[0127]

[0128] Among them, TP represents the number of samples that are actually positive examples and predicted as positive examples, FP represents the number of samples that are actually negative examples but predicted as positive examples, and FN represents the number of samples that are actually positive examples but predicted as negative examples.

[0129] In order to verify the effectiveness of the multi-index fusion proposed in this paper for patent value assessment tasks, this paper conducted comparative experiments and ablation experiments, mainly including other commonly used machine learning text classification models and methods of combining different indicators.

[0130] In order to verify that the method proposed in this paper is more effective than other patent value assessment methods in the experiment, the method proposed in this paper is compared with the logistic regression model (LR), K-nearest neighbor classification algorithm (KNN), Bayesian model (BM), support vector machine model (SVM), decision tree model (DT), random forest algorithm (RF) and full connection layer plus Bert model. At the same time, in order to prove the effectiveness of the data enhancement method proposed in the previous article on the value assessment of small sample Chinese patents, the original 2000 dataset (original dataset) and the 10000 dataset (augmented dataset) after data enhancement are predicted and evaluated by the above classification methods. The experimental results are shown in Table 5.

[0131] Table 5 Comparative experimental results

[0132]

[0133] The experimental results of the original 2000 data sets in Table 5.5 show that the multi-index fusion Chinese patent value level evaluation method proposed in this chapter is far superior to other traditional models. The method in this chapter and other traditional classification models quantized and spliced ​​the above-extracted indicator feature vectors as input in the Chinese patent value level evaluation prediction experiment. The analysis results show that the F1 value of this chapter method is 39.11% higher than that of the logistic regression model, 58.87% higher than that of the K nearest neighbor classification algorithm, 43.77% higher than that of the Bayesian model, 42.71% higher than that of the support vector machine model, 20.10% higher than that of the decision tree model, and 13.84% higher than that of the random forest algorithm. The logistic regression model is a statistical model used to solve classification problems. It maps experimental data to probability space and then divides the classification according to the threshold. However, the logistic regression model has very limited modeling capabilities for nonlinear relationships and is sensitive to outliers, so the prediction results are not good. The K-nearest neighbor classification algorithm calculates the distance between the data in the test set and all samples in the training set, and determines the category of the test sample by voting. However, the K-nearest neighbor classification algorithm is prone to underfitting, bias when dealing with unbalanced data sets, and is sensitive to noise data, so the prediction results are poor. The Bayesian model is based on the Bayesian theorem, updates the estimation of the model parameters through the prior probability and the training sample data, and then infers the classification results of the predicted samples through the posterior probability. The Bayesian model assumes that the data features are independent of each other, but the various features of the patent value assessment have a strong correlation. This independence assumption leads to a decrease in classification performance. The support vector machine model separates samples of different categories by obtaining a hyperplane in the training data set, and then classifies the predicted data. The support vector machine model itself is a binary classification model, which needs to be expanded for multi-classification tasks and is sensitive to missing data, resulting in poor prediction results. The decision tree model divides the input features by constructing a tree structure, selects different paths according to the value of the feature, and finally obtains the prediction result. The random forest algorithm is an ensemble learning algorithm based on the decision tree model. It performs classification by constructing multiple decision trees and combining their prediction results. The decision tree model and the random forest algorithm have the same basic structure. When dealing with the relationship between features, the best features and thresholds are selected recursively to construct the branches of the tree, which can capture the simple relationships between data features. The random forest can comprehensively consider the relationship between multiple feature subsets, so as to better capture the complex relationships between features, so the classification results are higher than those of a single decision tree model. However, each tree in the random forest performs feature selection and node partitioning independently, resulting in a discontinuous decision boundary, that is, there are some discontinuous gaps between sample points, resulting in poor final classification results.The XGBoost model used in this chapter is different from the random forest algorithm. It is an ensemble learning algorithm based on gradient boosting trees. Multiple decision trees are trained iteratively. Each tree takes into account the prediction bias of the previous tree and corrects the prediction error. Finally, the prediction results of all decision trees are added together to obtain the classification result. Therefore, the XGBoost model can handle high-dimensional data and data with nonlinear relationships, and gradually optimize the model performance through gradient boosting. The final experimental results are better than other traditional machine learning classification models.

[0134] Unlike traditional classification models that use word vectors as input for classification, classification methods based on large language models can only achieve classification through end-to-end natural language dialogue by calling API interfaces. Inspired by the agricultural news text classification method based on large language models, this chapter uses the template prompt method to predict the value level of Chinese patents based on Zhipu Qingyan. From the analysis of the experimental results, it can be seen that the prediction results of the large language model are the worst, and it has not achieved its superior effect in general field tasks. The reason is that the large language model does have strong language understanding and reasoning capabilities, but because its training data sets all belong to the general field, it cannot be accurately adapted to the patent field, and its professional knowledge and data corresponding capabilities are relatively lacking, and there is a deviation in professional citation. In addition, the large language model uses data to adapt the model input interface. All the indicators for patent value evaluation mentioned above need to be spliced ​​as natural language text and then input into the large model. It can only process all indicators with the same embedding method and decoding method, and cannot effectively integrate multiple types of indicators. At the same time, due to the limitation of the length of the input text, only one sample data can be input as a prompt command for each level of patent, and the potential of the large language model cannot be effectively activated. Therefore, large language models cannot make effective predictions in the evaluation of the value of Chinese patents.

[0135] Through the analysis of the experimental results of the original data set and the data augmented data set in Table 5, it can be seen that the use of the data augmented data set has a certain positive effect on the classification result prediction of all models.

[0136] The multi-index fusion Chinese patent value assessment method includes many different types of features. In order to verify the impact of various features on the results of Chinese patent value assessment, this paper tries different feature combinations and tests them based on the classification method proposed in this paper. Among them, the traditional patent value assessment index is an index combination proposed by other scholars for evaluating the value level of Chinese patents, including patent title, specification abstract, basic patent technical information (number of independent claims, number of dependent claims, etc.), basic patent economic information and basic patent legal information. The experimental results are shown in Table 6.

[0137] Table 6 Ablation experiment results

[0138]

[0139] Through the analysis of the experimental results in Table 6, it can be seen that each part of the evaluation index proposed in this chapter can play a positive role in the patent value level evaluation task, and its combined index can achieve the best results. The traditional patent value evaluation index (a) and (b) contain the same indicator features, the difference is that the text feature vectors of the patent title and the abstract of the specification are obtained in different ways. The traditional patent value evaluation index (a) uses the Bert model to encode the text, and the traditional patent value evaluation index (b) uses the large language model GLM to encode. Although it is impossible to obtain good experimental results through the large language model in the value evaluation task through the question-answering method, the encoding ability of the large language model is much higher than that of the pre-trained language model Bert. Its deeper model structure has more parameters and levels, so that the GLM model can perform language modeling of the input text through multiple levels of abstract representation, thereby improving the representation ability; more training data can help the model better learn the statistical laws and potential structures of the data, making the model more adaptable to a wide range of input spaces and not overly dependent on special samples. Therefore, the GLM model's language modeling ability for small samples or even zero samples is better than the current pre-trained language model. Therefore, the GLM model is used to embed short text features, and its F1 value is improved by 2.44%. Basic enterprise information reflects the market size and market share of the patent application enterprise organization, which helps to determine the commercial value of the patent technology in the market. This indicator increases the F1 value by 3.84%. Compared with traditional technology-related digital indicators, the technical efficacy matrix extracts technology-related keywords and efficacy phrases from the perspective of text to determine which key technical indicators the patent has a higher advantage in. At the same time, the technical efficacy matrix provides a structured framework within the field, which can objectively evaluate the value of the technology covered by the patent. This indicator increases the F1 value by 3.93%. As an important part of the patent document, the claims describe in detail the technical fields and protection requirements covered by the patent, and provide the specific boundaries and scope of the patent rights. The technical core and commercial application fields of the patent can be understood through its content, thereby determining the transaction and licensing potential of the patent. This indicator increases the F1 value by 4.55%.

[0140] The above indicators are combined to form the value evaluation indicator system proposed in this chapter, which makes the final experimental results achieve the best effect, and the F1 value is improved by 6.37%.

[0141] In summary, this paper proposes a Chinese patent value evaluation index system, which adds claims and corporate information to the traditional Chinese patent value evaluation, as well as the technical efficacy matrix constructed in Chapter 3. At the same time, the short text semantic feature extraction modeling method is changed to the large language model GLM for small sample texts. Then, for the obtained multi-type indicators, the feature vectors are spliced ​​using suitable different rule quantization methods. Finally, the patent value is evaluated through a classification model based on the XGBoost model.

[0142] Please note that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification. The above-mentioned embodiments only express several implementation methods of the present application, and their descriptions are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, without departing from the concept of the present application, several variations and improvements can be made, which all belong to the scope of protection of the present application. Therefore, the scope of protection of the patent in this application shall be based on the attached claims.

Claims

1. A multi-index fusion Chinese patent value assessment method, characterized in that: The method comprises: Step 1: Extracting patent text dimension features; including: extracting short text features from the abstract of the specification based on the large language model GLM, and extracting long text features from the claims based on the HBert model; Step 2: Extract the basic information indicator characteristics of the patent, including technical dimension indicators, economic dimension indicators, legal dimension indicators and enterprise dimension indicators; Step 3: Based on the patent text dimension characteristics and patent basic information indicator characteristics, use the XGBoost model to evaluate the patent value level.

2. According to the multi-index fusion Chinese patent value assessment method of claim 1, it is characterized in that: In step 1, short text features are extracted from the instruction manual abstract based on the large language model GLM; wherein: For the summary of the instruction manual, the large language model GLM model is used for feature embedding; the GLM model is based on the Transformer model. The GLM model is based on a 28-layer Transformer Encoder module, rearranges the order of normalization and residual connections, uses a single linear layer to predict the output word, and replaces the ReLU activation function with GeLUs; The Lora algorithm is used to fine-tune the GLM model. In Lora, the parameters of the original large language model are frozen, the intrinsic rank of the bypass simulation model parameters is increased, only the parameters in the bypass matrix are trained, and the parameters in the bypass matrix are superimposed on the original large language model; the text feature encoding of the instruction manual summary is performed through the fine-tuned GLM model and the fully connected layer, and the calculation formula is: GLM(W0+ΔW)=Lora(GLM(W0)) Sun A =GLM(A,W0+ΔW) V AL =Linear(V A ) Where W0 represents the original parameters of the GLM model, ΔW represents the parameters of the instruction fine-tuning bypass matrix through the Lora algorithm, A represents the input instruction summary text, V AL Represents the final feature word vector obtained as a short text feature.

3. According to the multi-index fusion Chinese patent value assessment method of claim 2, it is characterized in that: In step 1, long text features are extracted from the claims based on the HBert model; wherein: The HBert model obtains long text features in a hierarchical manner; the long text is divided into sentences, and the long text is represented as D = (sen1, sen2, ..., sen n ), D indicates long text, sen i represents the i-th sentence of a long text, which contains n sentences in total; if each sentence is split by characters, the sentence sen i Expressed as sen i =(w i1 ,w i2 ,...,w im ), w ij Indicates the sentence sen i The jth character of sen i There are m characters in total; After dividing the long text into sentences, i =(w i1 ,w i2 ,...,w im ) is used as the input of the Bert model to obtain the word vector H at the sentence granularity CLS And word vectors at word granularity H CLS Calculate the attention weight for each word vector and get the representative sentence sen by weighted summation i The word vector S i , the calculation formula is: The long text D = (sen1, sen2, ..., sen n ) is input into the Bert model, and the word vector set {S1,S2,...,S n }, concatenate a randomly initialized word vector [SCLS] before the word vector set, and replace {[SCLS], S1, S2, ..., S n } is used as input and encoded through the Transformer Encoder model to obtain a word vector H containing the granularity of long text SCLS and sentence-level word vectors H SCLS Calculate the attention weight for the word vector of each sentence, and obtain the word vector V representing the long text D by weighted summation D , the calculation formula is:

4. According to the multi-index fusion Chinese patent value assessment method of claim 3, it is characterized in that: In step 2: Technical dimension indicators include: technical efficacy matrix, number of independent claims, number of dependent claims, number of pages in the specification, number of embodiments, number of citations, number of cited documents, number of priorities, number of technical classifications, number of sub-cases, number of patent families, patent awards, and decryption of confidential patents; Economic dimension indicators include: number of applicants, type of applicants, authorization cycle, number of license filings, number of rights pledges, number of rights transfers, and customs filing status; Legal dimension indicators include: legal status, survival period / maintenance time, remaining life, litigation status, and number of invalidation requests; Enterprise dimension indicators include: company registered capital, number of insured persons, number of own risks, number of related risks, number of historical risks, number of sensitive public opinions; university level, start-up capital, number of own risks, number of related risks, number of historical risks, number of sensitive public opinions; The above indicators are divided into digital indicators, categorical indicators and matrix indicators, and the representation forms of digital indicators, categorical indicators and matrix indicators are obtained respectively. The representation vectors of all indicators are spliced ​​to obtain the basic information indicator characteristics of the patent.

5. According to the multi-index fusion Chinese patent value assessment method of claim 4, it is characterized in that: In step 3: The XGBoost model is used as the patent value evaluation model. XGBoost uses a large-scale parallel ensemble learning algorithm based on gradient boosting trees, performs a second-order Taylor expansion on the loss function, and introduces L1 and L2 regularization terms. The obtained patent text dimension features and patent basic information indicator features are concatenated to obtain the patent feature vector as the input of the XGBoost model, perform classification tasks, and output the patent value level. The calculation formula is: V=concat(V AL ,V D ,V info ) H=XGBoost(V) Among them, V AL Represents the short text feature in the patent text dimension feature, V D It represents the long text feature in the patent text dimension feature, V represents the patent feature vector, and H represents the patent value level result calculated by the model.

6. A multi-index fusion Chinese patent value assessment system, characterized in that: The system comprises a processing unit configured to perform the following steps: Extracting patent text dimension features; including: extracting short text features from the abstract of the specification based on the large language model GLM, and extracting long text features from the claims based on the HBert model; Extract the basic information indicator characteristics of patents, including technical dimension indicators, economic dimension indicators, legal dimension indicators and enterprise dimension indicators; Based on the patent text dimension characteristics and patent basic information indicator characteristics, the XGBoost model is used to evaluate the patent value level.

7. According to the multi-index fusion Chinese patent value evaluation system of claim 6, it is characterized in that: The processing unit is specifically configured to perform, based on the large language model GLM, extracting short text features from the specification abstract; wherein: For the summary of the instruction manual, the large language model GLM model is used for feature embedding; the GLM model is based on the Transformer model. The GLM model is based on a 28-layer Transformer Encoder module, rearranges the order of normalization and residual connections, uses a single linear layer to predict the output word, and replaces the ReLU activation function with GeLUs; The Lora algorithm is used to fine-tune the GLM model. In Lora, the parameters of the original large language model are frozen, the intrinsic rank of the bypass simulation model parameters is increased, only the parameters in the bypass matrix are trained, and the parameters in the bypass matrix are superimposed on the original large language model; the text feature encoding of the instruction manual summary is performed through the fine-tuned GLM model and the fully connected layer, and the calculation formula is: GLM(W0+ΔW)=Lora(GLM(W0)) Sun A =GLM(A,W0+ΔW) V AL =Linear(V A ) Where W0 represents the original parameters of the GLM model, ΔW represents the parameters of the instruction fine-tuning bypass matrix through the Lora algorithm, A represents the input instruction summary text, V AL Represents the final feature word vector obtained as a short text feature.

8. According to the multi-index fusion Chinese patent value evaluation system of claim 7, it is characterized in that: The processing unit is specifically configured to perform, based on the HBert model, extracting long text features from the claims; wherein: The HBert model obtains long text features in a hierarchical manner; the long text is divided into sentences, and the long text is represented as D = (sen1, sen2, ..., sen n ), D indicates long text, sen i represents the i-th sentence of a long text, which contains n sentences in total; if each sentence is split by characters, the sentence sen i Expressed as sen i =(w i1 ,w i2 ,...,w im ), w ij Indicates the sentence sen i The jth character of sen i There are m characters in total; After dividing the long text into sentences, i =(w i1 ,w i2 ,...,w im ) as the input of the Bert model to obtain the word vector H at the sentence granularity CLS And word vectors at word granularity H CLS Calculate the attention weight for each word vector and get the representative sentence sen by weighted summation i The word vector S i , the calculation formula is: The long text D = (sen1, sen2, ..., sen n ) is input into the Bert model, and the word vector set {S1,S2,…,S n }, concatenate a randomly initialized word vector [SCLS] before the word vector set, and replace {[SCLS], S1, S2, …, S n } is used as input and encoded through the Transformer Encoder model to obtain a word vector H containing the granularity of long text SCLS and sentence-level word vectors H SCLS Calculate the attention weight for the word vector of each sentence, and obtain the word vector V representing the long text D by weighted summation D , the calculation formula is:

9. According to the multi-index fusion Chinese patent value evaluation system of claim 8, it is characterized in that: The processing unit is specifically configured to perform: Technical dimension indicators include: technical efficacy matrix, number of independent claims, number of dependent claims, number of pages in the specification, number of embodiments, number of citations, number of cited documents, number of priorities, number of technical classifications, number of sub-cases, number of patent families, patent awards, and decryption of confidential patents; Economic dimension indicators include: number of applicants, type of applicants, authorization cycle, number of license filings, number of rights pledges, number of rights transfers, and customs filing status; Legal dimension indicators include: legal status, survival period / maintenance time, remaining life, litigation status, and number of invalidation requests; Enterprise dimension indicators include: company registered capital, number of insured persons, number of own risks, number of related risks, number of historical risks, number of sensitive public opinions; university level, start-up capital, number of own risks, number of related risks, number of historical risks, number of sensitive public opinions; The above indicators are divided into digital indicators, categorical indicators and matrix indicators, and the representation forms of digital indicators, categorical indicators and matrix indicators are obtained respectively. The representation vectors of all indicators are spliced ​​to obtain the basic information indicator characteristics of the patent.

10. A multi-index fusion Chinese patent value evaluation system according to claim 9, characterized in that: The processing unit is specifically configured to perform: The XGBoost model is used as the patent value evaluation model. XGBoost uses a large-scale parallel ensemble learning algorithm based on gradient boosting trees, performs a second-order Taylor expansion on the loss function, and introduces L1 and L2 regularization terms. The obtained patent text dimension features and patent basic information indicator features are concatenated to obtain the patent feature vector as the input of the XGBoost model, perform classification tasks, and output the patent value level. The calculation formula is: V=concat(V AL ,V D ,V info ) H=XGBoost(V) Among them, V AL Represents the short text feature in the patent text dimension feature, V D It represents the long text feature in the patent text dimension feature, V represents the patent feature vector, and H represents the patent value level result calculated by the model.

Citation Information

Patent Citations

  • Early evaluation method of patent value

    CN115186982A

  • Domain patent quality grade prediction method fusing knowledge information

    CN115204519A

  • Patent value evaluation method based on depth map and semantic learning

    CN115983877A

  • Scientific and technological achievement value evaluation method and system based on technical chain

    CN117591628A

  • Human-like value alignment method and system based on multi-dimensional feedback reinforcement learning

    CN118013016A

Cited By

  • LLM feature extraction and similarity matching-based patent value quantitative evaluation method

    CN120561714A

  • Scientific and technological achievement intelligent evaluation method and system based on multi-source data fusion

    CN121481296A