Patent label information generation method and system based on transfer learning

Through the transfer learning method, feature encoding and prediction models are constructed and pseudo-labels are generated, which solves the problems of poor flexibility and high labeling data demand in patent label information generation, and achieves efficient and accurate label generation in multi-field patent prediction.

CN120372012AActive Publication Date: 2025-07-25BEIJING ZHIGUAGUA TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510659139.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-31
Filing Date
2025-05-21
Publication Date
2025-07-25
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing patent label information generation methods have problems of poor flexibility and scalability. Rule-based methods require a large amount of manual annotation, model-based methods require a large amount of annotation data, and the general prediction model has poor adaptability in the patent field, which cannot adapt to the evolution of IPC classification and multi-target field migration.

Method used

Using a transfer learning-based method, we build source and target domain training corpus, design feature coding models and prediction models, including benchmark coding models, momentum coding models, classification prediction models, prototype prediction models and covariance prediction models, and generate pseudo-labels through semi-supervised methods to reduce labeling data needs and realize cross-domain migration.

Benefits of technology

It realizes accurate and efficient label information generation in multi-field patent prediction, reduces the economic and time cost of labeling data, and improves the adaptability and accuracy of the model in different IPC fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372012A_ABST
    Figure CN120372012A_ABST
Patent Text Reader

Abstract

The invention discloses a patent label information generation method and system based on transfer learning, the method comprises five steps of task configuration, model design, corpus construction, model training and execution prediction, the task configuration selects a patent prediction task, division is carried out based on all patent texts according to the technical field in an IPC classification system, and the patent prediction task is obtained; selecting a proper source domain and a target domain according to the patent data; according to the model design, a feature coding model and a prediction model are respectively designed according to task configuration information; corpus construction: respectively preparing training corpus data on a source domain and a target domain; model training is carried out on the selected source domain and target domain by using corresponding training corpus data; and executing prediction, applying a model interface after model training to predict unlabeled patents on the target domain, generating prediction label information and storing the prediction label information. According to the method, the economic cost and the time cost of a large amount of annotated data are effectively reduced under the condition that the accuracy of an existing model is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to the fields of natural language processing and deep learning, and particularly to a method and system for generating patent label information based on transfer learning. Background Art

[0002] The technology of natural language processing in patent text processing has become increasingly mature, which has greatly promoted the realization of intelligent and automated patent examination, application, and analysis. There are similarities in the various problems recorded in patent documents in different fields. Patents in different fields have both common basic features and specific features of the patent field. The existing methods for generating patent label information are generally implemented based on patent features combined with rules or models. The rule-based method has the problems of poor flexibility and scalability. The model-based method is generally implemented in a supervised manner and requires a large amount of manual annotation of patent data. Therefore, how to combine the natural language algorithm based on deep learning with the transfer learning ability according to the specific text characteristics of multiple patent fields, and face the contradiction between the poor unsupervised prediction effect of the patent value evaluation problem and the need for labeled data in the supervised model, and apply semi-supervised multi-objective domain transfer technology to achieve accurate and efficient patent prediction has become an urgent problem to be solved in this field.

[0003] The existing methods for generating patent label information are generally implemented based on patent features combined with rules or models. The rule-based method has the problems of poor flexibility and scalability. The model-based method is generally implemented in a supervised manner and requires a large amount of manual annotation of patent data. Since the patent text contains rich detailed description information in the legal, technical, and economic dimensions that are missing in the bibliographic items, the supervised model requires a large amount of labeled data in each field for training. However, patents cover many fields, and the annotation of patent text requires domain experts to spend a lot of time to complete. Manual annotation is time-consuming and laborious, resulting in high economic and time costs.

[0004] Currently, common prediction models are generally designed for general domains. When applied to the patent domain, due to differences in professional terms, application domains, and pragmatic scenarios, a large amount of literal, grammatical, and semantic information is contained in professional texts represented by patents, such as professional vocabulary, semantic similarity, and discourse structure relationships. At the same time, the large differences in text feature distributions caused by the diversity issues in the patent domain lead to poor multi-domain adaptability of general prediction models, and there is currently a lack of effective patent prediction models. Based on the IPC classification system, patents are divided into multi-level and multi-category professional domains. With the evolution of IPC standards accompanying technological development, the disappearance of some original IPC classification numbers, the emergence of new IPC classifications, and changes in the standard definitions and judgment rules of the same IPC classification in different periods all result in domain migration of data in the time dimension, making the original fixed classification test algorithms and models unable to adapt to this application scenario of multi-target domain migration. Summary of the Invention

[0005] The present application provides a method and system for generating patent label information based on transfer learning. By constructing source domain and target domain training corpora based on patent texts and designing a patent prediction model system based on deep learning, it effectively solves the problems of low accuracy of existing unsupervised models and lack of labeled data in supervised models, and achieves accurate and efficient multi-domain patent prediction.

[0006] In a first aspect, a method for generating patent label information based on transfer learning, the method includes five steps: task configuration, model design, corpus construction, model training, and prediction execution, which specifically include:

[0007] Task configuration: Select a patent prediction task, divide it based on the technical fields in the IPC classification system with all patent texts as the basis, and select appropriate source domain and target domain according to patent data;

[0008] Model design: Design a feature encoding model and a prediction model respectively according to the task configuration information; among them, the feature encoding model includes a reference encoding model and a momentum encoding model. The reference encoding model is used to encode labeled patents and corresponding citation patents in the source domain, and the momentum encoding model is used to keep in sync with the reference encoding model through periodic updates; the prediction model includes a classification prediction model, a prototype prediction model, and a covariance prediction model. The feature vectors generated by the feature encoding model are respectively input into the classification prediction model, the prototype prediction model, and the covariance prediction model. The classification prediction model is used to achieve classification prediction of patents, the prototype prediction model is used to predict the representative type of patents in the source domain, and the covariance prediction model is used to predict the covariance matrix of patents in the source domain;

[0009] Corpus construction, preparing training corpus data for the source domain and the target domain respectively; among them, the source domain corpus is constructed by labeled patents and corresponding patent citation data, and the target domain corpus is constructed by a small number of labeled patents and a large number of unlabeled original patent texts;

[0010] Model training, for the selected source domain and target domain, using the corresponding training corpus data to train the model;

[0011] Performing prediction, applying the model interface after model training to predict the unlabeled patents in the target domain, generating prediction label information and storing it; among them, the prediction label information at least includes the IPC classification number, technical field and legal status of the patent.

[0012] Optionally, in the task configuration, select appropriate source domain and target domain according to the patent data, specifically including:

[0013] Based on all patent texts, divide the patents according to the technical fields in the IPC classification system, select the fields with high-value patent label data as the source domain, and select the fields to which the knowledge to be extracted with little or no label data belongs as the target domain; among them, the fields are based on IPC classification and can take the third-level subclasses; the patent labels select high-value, high-disruptive, core key, green and low-carbon, key numbers, emerging industry task goals, and the fields are selected according to the task situation to limit the tobacco, communication, chip or new energy fields; the patent data selects the abstract or claims in the five patent books.

[0014] Optionally, in the model design, the encoding model designs the encoding model structure according to the configuration information of the source domain and the target domain, and selects the BERT model based on the multi-layer Transformer Encoder structure to implement. The BERT model is used as the encoding layer to implement the patent text encoding, and the CLS position vector of the output layer is taken as the feature encoding vector;

[0015] The benchmark encoding model is constructed based on the encoding model structure and is used for the embedding feature representation encoding of the citation patents corresponding to the patents in the source domain and the target domain; during initialization, the benchmark encoder is trained through the labeled dataset in the source domain;

[0016] The momentum encoding model is constructed based on the encoding model structure and is used for the embedding feature representation encoding of the patents in the source domain and the target domain; during the initialization of model training, the momentum encoder is initialized by applying the benchmark encoder, and during the model training process, the parameters of the momentum encoding model are updated using the momentum encoding model parameter update algorithm.

[0017] Optionally, the momentum encoding model parameter update algorithm is implemented by means of cyclic iteration, and the specific calculation process is through the parameter m of the momentum encoder F′ at the k-1 step k-1 , vk-1 , n k-1 , θ′ k-1 , update the parameters θ of the k-th step and the (k - 1)-th step of the reference encoder F k , θ k-1 :

[0018] m k = β1m k-1 +(1 - β1)θ k

[0019] v k = β2v k-1 +(1 - β2)(θ k - θ k-1 )

[0020]

[0021] n k = β3n k-1 +(1 - β3)[θ k +(1 - β2)(θ k - θ k-1 )] 2

[0022]

[0023] Optionally, in the model design, the classification prediction model adopts a prediction model structure, with the input being the feature vector output by the encoding model and the output being the binary classification high-value prediction probability;

[0024] The representative prediction model adopts a prediction model structure, with the input being the feature vector output by the encoding model and the output being the embedded feature vector of the high-value patent representative in the domain;

[0025] The covariance prediction model adopts a prediction model structure, with the input being the feature vector output by the encoding model and the output being the representative covariance matrix;

[0026] Optionally,, the representative of the representative prediction model is a P-dimensional feature vector calculated by the average value of the feature vectors in this type of set The metric in the embedded feature space vector of the covariance prediction model is calculated using the Bregman divergence:

[0027]

[0028] The distance calculation method between feature vectors is:

[0029]

[0030] Optionally, the dynamic mechanism of the semi-supervised embedding space domain adaptation model encodes patent and its cited patent data by applying a momentum encoder and a benchmark encoder to the source domain dataset, the labeled target dataset, and the unlabeled target dataset respectively, obtains representative embedding feature vectors by applying a representative prediction model, obtains the covariance matrix of the current instance by applying a covariance prediction model M, generates pseudo-labels for unlabeled instances, calculates the model loss, and updates the model parameters through backpropagation.

[0031] Optionally, the model training step further includes regularizing the model, specifically including:

[0032] Adding an L2 regularization term to the loss function of the model;

[0033] Setting a regularization parameter and determining its value range;

[0034] During training, imposing an L2 norm constraint on the weights of the model.

[0035] In addition, performing the prediction step further includes performing interpretive analysis on the prediction results, specifically including:

[0036] Extracting the feature importance scores related to the prediction results;

[0037] Generating a visualization chart of the feature importance scores;

[0038] Generating an interpretive report containing the feature importance scores and the prediction results.

[0039] In a second aspect, a patent label information generation system based on transfer learning, the system includes a task configuration module, a model design module, a corpus construction module, a model training module, and an execution prediction module, and specifically includes:

[0040] The task configuration module is used to select a patent prediction task, divide it based on the technical fields in the IPC classification system with all patent texts as the basis, and select appropriate source and target domains according to the patent data;

[0041] A model design module, configured to design a feature encoding model and a prediction model respectively according to task configuration information; wherein, the feature encoding model includes a baseline encoding model and a momentum encoding model, the baseline encoding model is configured to encode the labeled patents and corresponding citation patents in the source domain, and the momentum encoding model is configured to keep synchronized with the baseline encoding model by means of periodic update; the prediction model includes a classification prediction model, a prototype prediction model, and a covariance prediction model. The feature vectors generated by the feature encoding model are respectively input into the classification prediction model, the prototype prediction model, and the covariance prediction model. The classification prediction model is configured to implement the classification prediction of patents, the prototype prediction model is configured to predict the representative type of patents in the source domain, and the covariance prediction model is configured to predict the covariance matrix of patents in the source domain;

[0042] A corpus construction module, configured to prepare the training corpus data on the source domain and the target domain respectively; wherein, the source domain corpus is constructed by labeled patents and corresponding patent citation data, and the target domain corpus is constructed by a small number of labeled patents and a large number of unlabeled original patent texts;

[0043] A model training module, configured to perform model training on the selected source domain and target domain using the corresponding training corpus data;

[0044] An execution prediction module, configured to apply the model interface after model training to predict the unlabeled patents on the target domain, generate prediction label information and store it; wherein, the prediction label information at least includes the IPC classification number, technical field, and legal status of the patent.

[0045] In a third aspect, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method for generating patent label information based on transfer learning according to any one of the first aspects above is implemented.

[0046] The present invention provides a method for generating patent label information based on semi-supervised domain transfer learning. Based on the labeled patent data in the source domain and a small amount of labeled and a large amount of unlabeled patent data in the target domain, an encoding model including a general benchmark encoding model and a momentum encoding model is designed, and cross-domain transfer patent prediction is achieved by combining a classification prediction model, a representative prediction model, and a covariance prediction model. The encoding model is trained in an embedding space shared by the source and target domains based on contrast learning, and the lack of a large amount of unlabeled data in the target domain is supplemented by generating pseudo-labels for the unlabeled data in the target domain. The model gives full play to the feature representation ability of the representative in the IPC field, and corrects the deviation of the prediction result caused by the difference in the patent data distribution in different IPC fields by combining the domain dependence measure of the feature vector represented by the covariance model, effectively solving the problems of low accuracy of existing unsupervised models and lack of labeled data in supervised models, and achieving accurate and efficient multi-domain patent prediction. This method only requires labeled data in the source domain combined with a small amount of labeled data in the target domain, without a large number of labeled data in the target domain, and effectively reduces the economic cost and time cost of a large amount of labeled data while ensuring the accuracy of the existing model. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and those of ordinary skill in the art can also obtain other implementation drawings according to the provided drawings without creative efforts.

[0048] Figure 1 It is the main flow chart provided by the embodiment of the present application;

[0049] Figure 2 It is the domain transfer framework provided by the embodiment of the present application;

[0050] Figure 3 It is the overall structure diagram of the model provided by the embodiment of the present application;

[0051] Figure 4 It is the schematic diagram of the labeled and unlabeled sample spaces provided by the embodiment of the present application;

[0052] Figure 5 It is the input-output flow chart provided by the embodiment of the present application;

[0053] Figure 6 It is the overall structure diagram provided by the embodiment of the present application;

[0054] Figure 7 It is the module architecture block diagram of the patent label information generation system based on transfer learning provided by an embodiment of the present application;

[0055] Figure 8 It is the internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0056] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0057] In the description of the present application: the terms "including", "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily limit to the clearly listed steps or units, but may also include other steps or units inherent to these processes, methods, products or devices that are not clearly listed, or steps or units added by further optimized solutions based on the inventive concept.

[0058] Aiming at the problem of accurate and efficient prediction of multi-field patents, the present invention provides a method for generating patent label information based on semi-supervised domain transfer learning. A source domain and a target domain training corpus are constructed based on patent texts, and a patent prediction model system based on deep learning is designed. Among them, the source domain is the domain patent text data with labels, and the target domain is the domain patent text data with a small number of labels and a large number of unlabeled data. The model system includes a feature encoding model and a prediction model. The encoding model encodes the text into an embedded feature vector, including a general benchmark encoding model and a momentum encoding model. Among them, the benchmark encoding model is used for feature encoding of patent texts in the target domain, and the momentum encoding model is used for feature encoding of patent texts in the source domain. The encoding model is trained in an embedded space shared by the source and target domains based on contrast learning, and then pseudo-labels are generated for the unlabeled data in the target domain. The prediction model includes a classification prediction model, a representative prediction model, and a covariance prediction model. The label prediction model realizes patent label prediction based on the encoded features. The representative prediction model realizes the prediction of domain representatives, and the covariance prediction model realizes the prediction of covariance matrices, which is used to effectively alleviate the deviation of prediction results caused by the difference in the distribution of patent data in different IPC fields. Among them, the labeled patent prediction model in the source domain is trained in a supervised manner, and the unlabeled patent prediction model in the target domain is trained in a semi-supervised manner by generating pseudo-labels. After the above all models are integrated, a complete model system is formed to perform end-to-end training and inference on the data sets of the source domain and the target domain.

[0059] In one embodiment, as Figure 1As shown in the figure, a method for generating patent label information based on transfer learning is provided. This method can be applied to a server. The method includes five steps: task configuration, model design, corpus construction, model training, and prediction execution. Specifically, it includes:

[0060] For task configuration, it is necessary to first select the patent prediction task. Then, based on all patent texts, divide them according to the fields corresponding to IPC classification. Select appropriate source domains and target domains according to the patent data. The target domain includes a small number of labeled patents and a large number of unlabeled original patent texts; save and generate the above fields and task configuration information as the input for subsequent model training.

[0061] For model design, based on the above fields and task configuration information, design the model architectures for the source task and the target task respectively, including a feature encoding model and a prediction model. Among them, the feature encoding model includes a reference encoding model and a momentum encoding model, which are respectively used to encode labeled high-value patents and corresponding citation patents in the source domain. The momentum encoding model synchronizes with the reference encoding model through periodic updates. The prediction model includes a classification prediction model, a prototype prediction model, and a covariance prediction model. The feature vectors generated by the reference feature encoding model are respectively input into the classification prediction model, the prototype prediction model, and the covariance prediction model. The classification prediction model realizes the classification prediction of patents. The prototype prediction model is used to predict the representative type of patents in the source domain. The covariance prediction model is used to predict the covariance matrix of patents in the source domain; the unlabeled patents in the target domain generate pseudo-labels through the kNN method of the representative type. Design the overall loss function of the model, including patent prediction loss, representative type prediction loss, feature contrast loss, and covariance prediction loss.

[0062] For corpus construction, prepare the training data for the source domain and the target domain respectively. The source domain corpus is constructed by labeled patents and corresponding patent citation data. The target domain corpus is constructed by a small number of labeled patents and a large number of unlabeled original patent texts.

[0063] For model training, for the training corpus data selected for the source domain and the target domain respectively, select an optimizer, set training hyperparameters, train the encoding model and the prediction model, update the model parameters through the backpropagation of gradients, and save the model parameters after training.

[0064] For model application, apply the model interface trained above to predict the unlabeled patents in the target domain, generate prediction label information and store it.

[0065] Among them, the prediction label information can include IPC classification numbers, technical fields: the application fields of patents can also be described more specifically, such as specific industrial applications, product types, etc.

[0066] Patent legal status: The prediction labels may include the legal status of the patent, such as whether it is valid, under examination, expired, etc.

[0067] Patent technical effects: The prediction labels may describe the actual effects of the patented technology, such as improving efficiency, reducing costs, enhancing performance, etc.

[0068] The task configuration determines and selects appropriate source domains and target domains according to the existing high-value patent data situation.

[0069] For the selection of source and target domains, based on all patent texts, the patents are classified according to the technical fields in the IPC classification system. The fields with high-value patent label data are selected as the source domains, and the fields to which the knowledge to be extracted with few or no label data belongs are selected as the target domains.

[0070] Here, the fields are based on IPC classification, and the subclass (third level) can be taken. Patent labels can select task objectives such as high value, high disruption, core key, green and low-carbon, key figures, emerging industries, etc. The fields can be selected as restricted fields such as tobacco, communication, chips, or new energy according to the task situation. Patent data can select the abstract or claims in the five patent books.

[0071] An embodiment:

[0072] In the present invention, subclass A24F (smokers' articles) in the IPC system is selected as the source domain, and subclass C11B (production, e.g. by pressing raw materials or extracting from waste, refining or preserving fats, fatty substances such as lanolin, oils or waxes; essential oils; perfumes) is selected as the target domain. The patents in the source domain are data with high-value labels, and among the few patents in the target domain, some have high-value labels while a large number of patents have no labels. If all the patents in the target domain are without label data, some small amounts of high-value labels can be supplemented through manual assistance. Figure 4 An example figure shows the target domain and source domain of the labeled and unlabeled sample spaces.

[0073] The model design is based on the above-mentioned domain and task configuration information, and an encoding model and a prediction model are designed respectively, including a feature encoding model and a classification model. Among them, the feature encoding model includes a benchmark encoding model and a momentum encoding model, which are used to encode the labeled patents and the corresponding citation patents in the source domain respectively. The momentum encoding model is synchronized with the benchmark encoding model in a periodically updated manner. The classification model includes a classification prediction model, a prototype prediction model, and a covariance prediction model. The feature vectors generated by the feature encoding model are respectively input into the classification prediction model, the prototype prediction model, and the covariance prediction model. The classification prediction model realizes the classification prediction of patents, the prototype prediction model is used to predict the representative type of patents in the source domain, and the covariance prediction model is used to predict the covariance matrix of patents in the source domain; the unlabeled patents in the target domain generate pseudo-labels through the kNN method of the representative type. Design the overall structure of the model and the loss function of the optimization objective, including classification prediction loss, representative type prediction loss, contrast loss, covariance prediction loss, etc. Figure 5 It shows that the feature vectors are obtained by inputting the citation patents corresponding to the patents in the source domain and the target domain, and are used as the inputs of the classification prediction model, the representative type prediction model, and the covariance prediction model to obtain the outputs of each model. Figure 6 It shows the overall structure diagram of the algorithm.

[0074] Encoding model: Design the encoding model structure according to the configuration information of the source domain and the target domain, and implement it by using the BERT model based on the multi-layer Transformer Encoder structure. The BERT model is used as the encoding layer to implement patent text encoding, and the [CLS] position vector of the output layer is taken as the feature encoding vector.

[0075] Benchmark encoding model: Based on the encoding model structure, it is used for the embedding feature representation encoding of the citation patents corresponding to the patents in the source domain and the target domain. At initialization, the benchmark encoder F is trained through the labeled dataset on the source domain.

[0076] Momentum encoding model: Based on the encoding model structure, it is used for the embedding feature representation encoding of the patents in the source domain and the target domain. When initializing the model training, the benchmark encoder F is used to initialize the momentum encoder F`. During the model training process, the parameters of the momentum encoding model are updated using the following algorithm:

[0077] Algorithm 1: Parameter update algorithm for the momentum encoding model:

[0078] Input: Initialize the model parameters θ0, the number of iteration steps K, the step size η, the momentum (β1, β2, β3) ∈ [0, 1] 3 , the stability parameter ε > 0, the weight decay λ k > 0, the restart condition, the parameters θ of the benchmark encoder F at the k-th step and the (k - 1)-th step k, θ k-1 , the parameter m of the momentum encoder F′ at the (k - 1)-th step of the previous step k-1 , v k-1 , n k-1 , θ′ k-1 The target loss function f θ ,

[0079] Output: m at the k-th step of the current step of the momentum encoder F`, k , v k n k , θ′ k .

[0080] Initialization: When executed for the first time, initialize m0 = 0, v0 = 0, k = 0.

[0081] Calculation:

[0082] m k = β1m k-1 + (1 - β1)θ k

[0083] v k = β2v k-1 + (1 - β2)(θ k - θ k-1 )

[0084]

[0085] n k = β3n k-1 + (1 - β3)[θ k + (1 - β2)(θ k - θ k-1 )] 2

[0086] Return m k , v k , n k , θ′ k , Note: Here, β1 = 0.9, β2 = 0.999, β3 = 0.9999, η = 0.001, ε = 10 -8 , K = 1000.

[0087] Prediction Model: It is constructed using a two-layer fully connected neural network structure. The input is the feature vector output by the encoding model, and the output is the binary classification high-value prediction probability. The received feature vector input by the encoding model is n-dimensional. The first hidden layer contains h1 neurons, using ReLU as the activation function. The second hidden layer contains h2 neurons, also using ReLU as the activation function, mainly to increase non-linearity and accelerate training. The output layer contains 1 neuron, using the sigmoid function, and the output is restricted to between 0 and 1 to represent probability. Specifically, z1 = W1x + b1, a1 = ReLU(z1), z2 = W2x + b2, a2 = ReLU(z2), z3 = W3x + b3, y = σ(z3), where W1, W2, and W3 are weight matrices, b1, b2, and b3 are bias vectors, and σ is the sigmoid function. x is the input feature vector.

[0088] Classification Prediction Model (Classification), using the prediction model structure, with the input being the feature vector output by the encoding model and the output being the binary classification high-value prediction probability.

[0089] Prototype Prediction Model, using the prediction model structure, with the input being the feature vector output by the encoding model and the output being the embedded feature vector of the high-value patent prototype in the domain.

[0090] The prototype is a P-dimensional feature vector. It is calculated by the average value of the feature vectors in this class of sets.

[0091] Among them, is the embedded feature encoding function, implemented in the way of the encoding model, where φ is the learnable parameter of the model.

[0092] The working principle of the prototype model is to give the distance function The prototype prediction model outputs the softmax distribution of the query point x in the embedded space at a distance d in this class:

[0093]

[0094] The learning process of the model parameters is to minimize the negative log probability through gradient descent on class k:

[0095] Implemented.

[0096] Covariance Prediction Model (Covariance), using the prediction model structure, with the input being the feature vector output by the encoding model and the output being the prototype covariance matrix.

[0097] The metric embedded in the feature space vector is calculated using the Bregman divergence:

[0098]

[0099] where is a differentiable strictly Legendre convex function. Here, the Bregman divergence of the Mahalanobis distance squared is adopted:

[0100] When at that time

[0101] To simplify the calculation and improve the execution speed, Q can be implemented as the identity matrix, resulting in the Euclidean squared distance:

[0102] At this time

[0103] The covariance matrix here is the Gaussian covariance matrix. Based on the distance metric and direction in the class-specific feature vector space, a representative type is automatically learned and constructed. Among them, the clustering center is the center point of the representative type, and the Gaussian covariance matrix characterizes the confidence (uncertainty) in a specific direction, with anisotropic scaling in the feature space. Here, Q == S -1 .

[0104] The predicted output covariance matrix of the representative type network where D is the dimension of the feature vector in the embedding space.

[0105] The distance calculation method between feature vectors is:

[0106]

[0107] where is the center point, that is, the representative type of class c. M is the covariance matrix. The representative type model combined with the covariance model can learn the class and direction-dependent distance metric in the embedding space. In practical applications, to improve the model training and inference speed, the covariance matrix can be further simplified to a diagonal matrix.

[0108] Model system: The dynamic mechanism of the model system includes the synchronous update mode of the momentum encoding model and the benchmark encoding model, the calculation and update mode of the representative type feature vector, the calculation and update mode of the covariance matrix, the calculation of the loss function, and the update of model parameters.

[0109] The dynamic mechanism of the semi-supervised embedding space domain adaptation model is shown in Algorithm 2:

[0110] Algorithm 2: Dynamic mechanism of the semi-supervised embedding space domain adaptation model:

[0111] Input: Source domain dataset Ds , there is a labeled target dataset D t , there is an unlabeled target dataset D u .

[0112] Initialization: Initialize the momentum encoder F' using the baseline encoder F trained on the labeled dataset, and initialize the database Q = {}.

[0113] Loop through the following steps until the model training ends:

[0114] 1. Sample samples (x s , y s ) ∈ D with corresponding cited patents from the source domain dataset D s , sample samples (x s , y t ) ∈ D with corresponding cited patents from the labeled dataset of the target domain D t , y t ) ∈ D t , and sample samples x u with corresponding cited patents from the unlabeled dataset of the target domain D u ∈ D u .

[0115] 2. Retrieve the corresponding cited patent data of x s , x t , x u from the patent database

[0116] 3. Apply the momentum encoder F' to encode x s , x t , x u to obtain the embedded feature vectors f′ s , f′ t , f′ u , apply the baseline encoder F to to obtain the embedded feature vectors Apply the representative prediction model P to obtain the representative embedded feature vectors p′ s , p′ t , p′ u , Apply the covariance prediction model M to obtain the covariance matrix M′ of the current instance s , M′ t , M′ u , Update all the generated feature vectors, representative vectors, and covariance matrices to the database Q.

[0117] 4. Generate pseudo-labels for the unlabeled instance x u , Generate: If Then generate pseudo-labels where f is for x respectively u , Take f′ u and The distance in kNN adopts the covariance-corrected distance formula to calculate is the representative vector of class c, M is the covariance matrix is the candidate embedding feature vector in Q. Update all the generated pseudo-labels above to the database Q

[0118] 5. Calculate the model loss and update the model parameters through backpropagation

[0119] 5.1. Calculate the labeled citation patent data The cross-entropy loss L1 of the label: Apply the high-value prediction model C and input the embedding vector features of the labeled citation patent data to obtain the predicted label information Calculate the cross-entropy loss through the cross-entropy loss function where y′ are the true label and the predicted label information respectively, y′ y′ y′ s , y′ t are the labels of the citation patent data on the source domain and the target domain respectively

[0120] 5.2. Calculate the unlabeled citation patent data The cross-entropy loss L of the generated pseudo-labels u , apply the high-value prediction model C and input the embedding vector features of the unlabeled citation patent data to obtain the predicted label information Calculate the cross-entropy loss through the cross-entropy loss function where are the predicted label information and the pseudo-labels respectively the unlabeled citation patent data of the target domain generated pseudo-labels

[0121] 5.3. Apply the formula where τ is 0.9 and formula 6 Calculate the contrast loss between the unlabeled sample x u and the corresponding citation data feature vector f u and f′ u in the vector library Q: L c = L u ​​​(f u , f′ u ).

[0122] 5.4 Use a gradient reversal layer before the classifier layer to calculate the gradient, and use to update the parameters of the baseline encoder F and initialize the parameters of the momentum encoder F`, where λ u , λ c represents the regularization coefficient, whose value determines the weight of the regularization term in the total loss function, and H represents that the prototype is updated based on the entropy of the unlabeled target prediction.

[0123] 5.5 Use a gradient reversal layer before the classifier layer to calculate the gradient, and use the parameters of the high-value prediction model C where λ u , λ c represents the regularization coefficient, whose value determines the weight of the regularization term in the total loss function, and H represents that the prototype is updated based on the entropy of the unlabeled target prediction.

[0124] 5.6. Apply the formula to update the representative vectors of all categories c in all Q Apply the formula to update the representative vector of the category c predicted by the representative model where Q fc is the set of all embedded feature vectors belonging to category c in Q, Q pc is the set of all representative vectors belonging to category c in Q, N pc is the number of representative vectors in Q pc , and M is the covariance matrix. Calculate the representative prediction Euclidean distance loss:

[0125] 5.7. Apply the covariance formula to calculate the covariance matrix of all feature vectors in Q where X i , Y i represents the i-th feature vector in Q, is the expected mean of all feature vectors in Q, and N is the total number of feature vectors in Q. Apply the formula to update the covariance matrix predicted by the covariance model, where m is the covariance matrix predicted by the covariance model in the set Q M of all covariance matrices in Q, and N M is the total number of covariance matrices in Q M . Calculate the covariance matrix prediction loss: where respectively represent the elements of the i-th row and j-th column of the covariance matrix.

[0126] Sum up all the loss functions in the above steps 5.1 to 5.7:

[0127] L = L l + L u + L c + L p + L m , and use the gradient descent algorithm of backpropagation to update respectively:

[0128] The parameters of the benchmark encoder The parameters of the high-value prediction model The parameters of the representative prediction model The parameters of the covariance prediction model

[0129] 6. Update the parameters of the momentum encoder model using Algorithm 1

[0130] 7. Update all the generated feature vectors, labels, representative vectors, and covariance matrices into the database Q. Return all model parameters. The corpus construction is to prepare the training data on the source domain and the target domain respectively. The source domain corpus is constructed by labeled patents and corresponding patent citation data, and the target domain corpus is constructed by a small number of labeled patents and a large number of unlabeled original patent texts;

[0131] In the model training process, the following steps use the corresponding training corpus data to train the model. Then, train the encoder model and the predictor model, and use the backpropagation algorithm to optimize the model parameters. After training, save the trained model parameters.

[0132] In an optional embodiment of the present application, for the semi-supervised embedding space domain adaptation model dynamic mechanism, by applying the momentum encoder and the benchmark encoder on the source domain dataset, the labeled target dataset, and the unlabeled target dataset respectively, encode the patents and their citation patent data to generate embedded feature vectors. Subsequently, use the representative prediction model to extract the representative embedded feature vectors, and calculate the covariance matrix through the covariance prediction model to characterize the distribution characteristics of the data. For unlabeled instances, generate pseudo-labels based on the representative embedded feature vectors and the covariance matrix to further enrich the training data. By calculating the model loss and using backpropagation to update the model parameters, optimize the model performance, realize the knowledge transfer from the source domain to the target domain, and effectively improve the adaptability and prediction accuracy of the model to the target domain data.

[0133] In an optional embodiment of the present application, in the model training step, to prevent the model from overfitting, a step of regularizing the model is added. The specific implementation details are as follows:

[0134] Add an L2 regularization term to the loss function of the model. L2 regularization makes the weight vector of the model smoother and reduces the complexity of the model by adding the sum of squares of the weights to the loss function.

[0135] Set the regularization parameter, whose value range is [0.001, 0.1]. The regularization parameter controls the weight of the regularization term in the loss function. By adjusting the value of the regularization parameter, a balance can be found between the complexity and fitting degree of the model. Generally, a smaller value of the regularization parameter will make the model more inclined to fit the training data, while a larger value of the regularization parameter will make the model smoother and reduce the risk of overfitting. During the training process, impose an L2 norm constraint on the weights of the model. Specifically, by normalizing the weights in each iteration, ensure that the L2 norm of the weights does not exceed a preset threshold.

[0136] In an alternative embodiment of the present application, during the prediction step, in order to help users better understand the basis and logic of the prediction results, a step of performing interpretive analysis on the prediction results is added. The specific implementation details are as follows:

[0137] In the prediction model, by analyzing the weights and input features of the model, calculate the importance score of each feature for the prediction result. Specifically, methods such as SHAP (SHapley Additive exPlanations) value or LIME (Local Interpretable Model-agnostic Explanations) can be used to calculate the feature importance score. These methods can quantify the contribution of each feature to the model prediction, thereby helping users understand which features have an important impact on the prediction result.

[0138] Generate a visualization chart of the feature importance score, such as a bar chart or a heatmap, to intuitively show the impact of each feature on the prediction result. For example, a bar chart can clearly display the importance score of each feature, and a heatmap can show the interaction between features and their combined impact on the prediction result. These charts can help users quickly identify key features and their contributions to the prediction result.

[0139] Generate an interpretive report containing the feature importance score and the prediction result. The report details the importance score of each feature and its specific impact on the prediction result. In addition, the report can also include information such as the confidence of the prediction result, the source of the prediction result (such as similar cases in the training data), etc., to help users comprehensively understand the basis and logic of the prediction result.

[0140] The model application is to apply the model interface trained above to predict unlabeled patents in the target domain, generate prediction label information and store it.

[0141] In summary, it can be seen that the present invention provides a method for extracting multi-domain knowledge of patents based on feature alignment. The fields are divided based on the IPC classification of patents, and the source domain, target domain, and their corresponding combinations of labeled and unlabeled data are selectively targeted. An encoding model including a general benchmark encoding model and a momentum encoding model is designed, and a patent prediction for cross-domain migration is achieved by combining a classification prediction model, a representative prediction model, and a covariance prediction model. The encoding model and the prediction model comprehensively adopt various loss functions such as cross-entropy, contrastive loss, representative, and covariance in the embedding space shared by the source and target domains. Through the automatic generation of pseudo-labels for a large amount of unlabeled data in the target domain, the original unsupervised training method is converted into a supervised training mode. After integration, a complete model system is formed, and efficient end-to-end training and inference are achieved on the datasets of the source domain and the target domain.

[0142] In one embodiment, as Figure 7 shown, a system for generating patent label information based on transfer learning is provided. The system includes a task configuration module, a model design module, a corpus construction module, a model training module, and an execution prediction module, which specifically include:

[0143] The task configuration module is used to select a patent prediction task. Based on all patent texts, it is divided according to the technical fields in the IPC classification system, and appropriate source and target domains are selected according to the patent data;

[0144] The model design module is used to design a feature encoding model and a prediction model respectively according to the task configuration information; among them, the feature encoding model includes a benchmark encoding model and a momentum encoding model. The benchmark encoding model is used to encode the labeled patents and corresponding citation patents in the source domain, and the momentum encoding model is used to keep in sync with the benchmark encoding model through periodic updates; the prediction model includes a classification prediction model, a prototype prediction model, and a covariance prediction model. The feature vectors generated by the feature encoding model are respectively input into the classification prediction model, the prototype prediction model, and the covariance prediction model. The classification prediction model is used to achieve the classification prediction of patents, the prototype prediction model is used to predict the representative of patents in the source domain, and the covariance prediction model is used to predict the covariance matrix of patents in the source domain;

[0145] The corpus construction module is used to prepare the training corpus data on the source domain and the target domain respectively; among them, the source domain corpus is constructed by labeled patents and corresponding patent citation data, and the target domain corpus is constructed by a small amount of labeled patents and a large amount of unlabeled original patent texts;

[0146] The model training module is used to train the model using the corresponding training corpus data for the selected source domain and target domain;

[0147] An execution prediction module is used to apply the model interface after model training to predict unlabeled patents in the target domain, generate prediction label information and store it.

[0148] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structural diagram can be as Figure 8 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities, and the network interface is used to communicate with external terminals through a network. The computer device realizes the above-mentioned method for generating patent label information by loading and running a computer program.

[0149] Those skilled in the art can understand that Figure 8 the structure shown in

[0150] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0151] In one embodiment, a computer-readable storage medium is further provided, on which a computer program is stored, covering all or part of the processes in the above-mentioned embodiment method.

[0152] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in M forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Symchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0153] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

Claims

1. A method for generating patent label information based on transfer learning, characterized in that The method includes five steps: task configuration, model design, corpus construction, model training, and prediction execution, which specifically include: Task configuration: Select the patent prediction task, divide all patent texts based on the technical fields in the IPC classification system, and select appropriate source and target domains according to the patent data; Model design: Design a feature encoding model and a prediction model respectively according to the task configuration information. Among them, the feature encoding model includes a baseline encoding model and a momentum encoding model. The baseline encoding model is used to encode the labeled patents and corresponding citation patents in the source domain, and the momentum encoding model is used to keep in sync with the baseline encoding model through periodic updates; the prediction model includes a classification prediction model, a prototype prediction model, and a covariance prediction model. The feature vectors generated by the feature encoding model are respectively input into the classification prediction model, the prototype prediction model, and the covariance prediction model. The classification prediction model is used to realize the classification prediction of patents, the prototype prediction model is used to predict the representative type of patents in the source domain, and the covariance prediction model is used to predict the covariance matrix of patents in the source domain; Corpus construction: Prepare the training corpus data for the source and target domains respectively. Among them, the source domain corpus is constructed by labeled patents and corresponding patent citation data, and the target domain corpus is constructed by a small number of labeled patents and a large number of unlabeled original patent texts; Model training: Use the corresponding training corpus data to train the model for the selected source and target domains; Prediction execution: Apply the model interface after model training to predict the unlabeled patents in the target domain, generate prediction label information and store it. Among them, the prediction label information at least includes the IPC classification number, technical field, and legal status of the patent; 2. The method according to claim 1, wherein Selecting appropriate source and target domains according to the patent data in task configuration specifically includes: Based on all patent texts, divide the patents according to the technical fields in the IPC classification system, select the fields with high-value patent label data as the source domain, and select the fields to which the knowledge to be extracted with few or no label data belongs as the target domain. Among them, the fields are based on IPC classification and can take the third-level subclasses; the patent labels select high-value, high-disruptiveness, core key, green and low-carbon, key numbers, emerging industry task goals, and the fields are selected according to the task situation to limit the tobacco, communication, chip, or new energy fields; the patent data selects the abstract or claims in the five patent books; 3. The method according to claim 1, wherein In model design, the encoding model designs the encoding model structure according to the configuration information of the source and target domains, and selects the BERT model based on the multi-layer Transformer Encoder structure to implement. The BERT model is used as the encoding layer to realize patent text encoding, and the CLS position vector of the output layer is taken as the feature encoding vector; The baseline encoding model is constructed based on the encoding model structure and is used for the embedded feature representation encoding of the citation patents corresponding to the patents in the source and target domains; At initialization, the baseline encoder is trained through the labeled dataset in the source domain; The momentum encoding model is constructed based on the encoding model structure and is used for the embedded feature representation encoding of the patents in the source and target domains; When initializing model training, the momentum encoder is initialized using the baseline encoder. During the model training process, the parameters of the momentum encoding model are updated using the momentum encoding model parameter update algorithm.

4. The method according to claim 3, wherein The momentum encoding model parameter update algorithm is implemented by means of iterative updates. The specific calculation process is to update the parameters m k-1 , v k-1 ' n k-1 , θ′ k-1 of the momentum encoder F′ at the (k - 1)-th step, and update the parameters θ k , θ k-1 at the k-th step and the (k - 1)-th step of the reference encoder F: m k = β1m k-1 + (1 - β1)θ k v k = β2v k-1 + (1 - β2)(θ k - θ k-1 ) n k = β3n k-1 + (1 - β3)[θ k + (1 - β2)(θ k - θ k-1 )] 2 5. The method according to claim 1, wherein In the model design, the classification prediction model adopts the prediction model structure, with the input being the feature vector output by the encoding model and the output being the binary classification high-value prediction probability. The representative prediction model adopts the prediction model structure, with the input being the feature vector output by the encoding model and the output being the embedded feature vector of the high-value patent representative in the domain. The covariance prediction model adopts the prediction model structure, with the input being the feature vector output by the encoding model and the output being the representative covariance matrix.

6. The method according to claim 5, characterized in that The representative form of the representative prediction model is a P-dimensional feature vector It is calculated by the average value of the feature vectors in this type of set The metric embedded in the feature space vector of the covariance prediction model is calculated using the Bregman divergence: The distance calculation method between feature vectors is as follows:

7. The method according to claim 1, wherein The dynamic mechanism of the semi-supervised embedding space domain adaptation model encodes patent and its citation patent data using the momentum encoder and the baseline encoder on the source domain dataset, the labeled target dataset, and the unlabeled target dataset respectively, obtains the representative embedded feature vector using the representative prediction model, obtains the covariance matrix of the current instance using the covariance prediction model M, generates pseudo-labels for unlabeled instances, calculates the model loss, and updates the model parameters through backpropagation.

8. The method according to claim 1, characterized in that, The model training steps further include regularizing the model, specifically including: Adding an L2 regularization term to the loss function of the model; Setting the regularization parameter and determining its value range; During training, imposing an L2 norm constraint on the weights of the model; In addition, the prediction step also includes performing interpretive analysis on the prediction results, specifically including: Extracting the feature importance scores related to the prediction results; Generating a visualization chart of the feature importance scores; Generating an interpretive report containing the feature importance scores and the prediction results.

9. A patent label information generation system based on transfer learning, characterized in that The system includes a task configuration module, a model design module, a corpus construction module, a model training module, and an execution prediction module, which specifically include: The task configuration module is used to select the patent prediction task, divide based on the technical fields in the IPC classification system with all patent texts as the basis, and select appropriate source and target domains according to the patent data. The model design module is used to design the feature encoding model and the prediction model respectively according to the task configuration information; among them, the feature encoding model includes a baseline encoding model and a momentum encoding model. The baseline encoding model is used to encode the labeled patents and corresponding citation patents in the source domain, and the momentum encoding model is used to keep in sync with the baseline encoding model through periodic updates; the prediction model includes a classification prediction model, a prototype prediction model, and a covariance prediction model. The feature vectors generated by the feature encoding model are respectively input into the classification prediction model, the prototype prediction model, and the covariance prediction model. The classification prediction model is used to achieve the classification prediction of patents, the prototype prediction model is used to predict the representative of patents in the source domain, and the covariance prediction model is used to predict the covariance matrix of patents in the source domain. The corpus construction module is used to prepare the training corpus data on the source and target domains respectively; among them, the source domain corpus is constructed by labeled patents and corresponding patent citation data, and the target domain corpus is constructed by a small number of labeled patents and a large number of unlabeled original patent texts. A model training module, which is used to train a model using corresponding training corpus data for a selected source domain and target domain; An execution prediction module, which is used to apply the model interface after model training to predict unlabeled patents in the target domain, generate prediction label information and store it; wherein the prediction label information at least includes the IPC classification number, technical field and legal status of the patent.

10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Riemannian feature migration-based magnetoencephalogram signal classification method and device and medium

    CN113191206A

  • Efficient multi-modal contrast deep hash retrieval method for medical big data

    CN116881336A

  • Patent multi-domain knowledge extraction method and system based on feature alignment

    CN117172323A

  • Devices, systems, methods, and media for domain adaptation using hybrid learning

    US20230082899A1

  • Model training device, model training method, and program

    WO2024018592A1