An edge innovation patent prediction method and system

By constructing a time-series-based edge patent dataset and neural network model, combined with multi-dimensional feature evaluation, the limitations of edge innovation identification in existing technologies are overcome, enabling early and accurate identification and prediction of edge innovation patents, and supporting technology strategic decision-making.

CN119648481BActive Publication Date: 2025-10-24ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411761436.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-10-24
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify and predict emerging technologies. Traditional methods are slow to react and lack timely responses to technological trends. Furthermore, existing methods have limitations in identifying emerging technologies.

Method used

By constructing a time-series-based edge patent dataset, combining outlier and niche assessments, and using neural network models to predict the conversion time of edge patents, and combining technological novelty, market attractiveness, and cross-diversity characteristics, we can achieve multi-dimensional dynamic tracking of innovative technology development.

Benefits of technology

Accurately identify early-stage, cutting-edge innovation patents, provide scientific evidence to support technology strategy decisions, and plan ahead for future innovation directions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119648481B_ABST
    Figure CN119648481B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of technical opportunity identification, and discloses an edge innovation patent prediction method and system, which comprises the following steps: according to a patent edge innovation opportunity identification task, information of all patents in a technical field to which a target patent belongs is acquired, and a patent data set is constructed; the patent data set is divided according to a time sequence, edge evaluation is carried out, a time sequence-based edge patent data set and a center patent data set are constructed; edge innovation patents are screened out according to the time sequence, conversion time lengths are recorded, and an edge innovation patent data set is constructed; a feature system of the edge innovation opportunity is constructed, sample features of the edge innovation patent data set are acquired, a prediction model is constructed based on a neural network, and supervised training on a conversion time task is carried out; the target patent is predicted to be converted into an edge innovation patent, and conversion time is obtained. According to the time sequence characteristics of patent data, the development of technology is dynamically tracked, and technology with growth potential can be identified at an early stage.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of technical opportunity identification, and in particular to an edge innovation patent prediction method and system. BACKGROUND

[0002] With the rapid development of global technology, the application scenarios of innovative technologies are expanding, and the demand for identification and prediction of cutting-edge technologies by enterprises and research institutions is growing. Patent literature is an important record and source of technological innovation, and can reflect the dynamics of technological development. Currently, there are still great challenges in predicting edge innovation technologies. Edge technologies usually refer to technologies that have not been widely applied or lack sufficient market attention. These technologies may have great application potential in the future due to their innovative or breakthrough characteristics. However, due to their relatively hidden development path, traditional technology prediction methods, such as trend analysis based on historical data and expert review, are difficult to identify the potential of these technologies in a timely manner, leading to a slow response to technological trends, which in turn affects the formulation of innovation strategies and technology layout.

[0003] Technical opportunities are discovered by mining the development trends and relationships of existing technologies in a certain technical field, discovering the latest technology trends, and inferring the possible technical forms or development points in that field. Technical opportunities, as the basis for innovation decisions, are important factors that must be considered for any technological innovation. Technical opportunity identification methods include qualitative and quantitative analysis. The rapid development of data mining and natural language processing technologies has led to rapid progress in quantitative analysis of technical opportunities. The quantitative analysis methods for technical opportunity identification include patent map method, form analysis method, outlier detection method, semantic TRIZ method, and link prediction-based method, etc. The patent map method is based on the vectorization of patent keywords, and then performs dimensionality reduction visualization to find the blank areas in the patent map as potential technical opportunities. The core steps of the outlier detection method include identification and evaluation of outlier patents, where the evaluation step is usually based on qualitative methods such as expert experience, which is highly subjective. The link prediction method uses network structure information to predict the likelihood of connection between two nodes in the network that have not yet been connected, and finds opportunity points. Most of these methods are based on a single data or a single relationship, and have certain industry limitations.

[0004] The patent application with the publication number CN 114817567 A discloses a technical opportunity identification method, which comprises the following steps: extracting classification numbers from patent data and constructing a text set, obtaining semantic vectors and co-occurrence vectors of each classification number through a Doc2vec model, constructing classification number co-occurrence networks in different time periods respectively, dividing positive samples and negative samples to obtain a training sample set; counting all classification numbers that have appeared in all patents in the target field, and obtaining a target field node vector after weighted averaging; adding edges between the target field node and all nodes in the target field classification number set, and corresponding all nodes one by one to generate a test sample set; inputting each test sample in the test sample set into the graph neural network model trained by the training sample set to obtain the probability of each test sample generating an edge. The prior art mainly relies on classification number co-occurrence networks, and cannot comprehensively capture multi-dimensional features such as time dynamics and inventor relationship networks of patent documents; meanwhile, the prior art only predicts the probability of connection between nodes through a graph neural network, which is relatively limited in specific edge innovation opportunity identification.

[0005] The patent application with the publication number CN 117634502 A discloses a technical opportunity identification method, which comprises the following steps: collecting an initial literature data set corresponding to a to-be-analyzed technical field according to the to-be-analyzed technical field; preprocessing the initial literature data set to obtain a pre-analysis data set; screening the pre-analysis data set to determine a technical main path corresponding to the to-be-analyzed technical field, i.e., the evolution and development history of a technical theme; calculating the development maturity of the technical main path to determine the development stage of the technical main path; and determining potential technical opportunities in the to-be-analyzed technical field based on the determined development stage. The prior art mainly identifies technical opportunities by evaluating the development stage of the technical main path, which easily ignores technical opportunities located at the edge of the technical field and lacks comprehensive evaluation of the market potential of technology and edge innovation technical opportunities. SUMMARY

[0006] The purpose of the present application is to solve the problems in the prior art in the identification and prediction of patent technical opportunities. Through the composite technical opportunity prediction method, the advantages of single prediction methods are complementary, the patent opportunities are analyzed dynamically and multidimensionally combined with time changes, the potential future technical opportunities are identified, the scientific basis for technical strategic decision-making is provided, and it is helpful to layout the future technical innovation direction in advance.

[0007] An edge innovation patent prediction method, comprising:

[0008] Step 1: According to the patent edge innovation opportunity identification task published by the user and containing the target patent, the information of all patents in the technical field to which the target patent belongs is obtained and preprocessed to construct a patent data set;

[0009] Step 2: The patent dataset is divided in chronological order, and the marginality evaluation of the two dimensions of outlying and smallness is carried out respectively, and the marginal patent dataset and the center patent dataset based on time sequence are constructed;

[0010] Step 3: According to the marginal patent dataset and the center patent dataset based on time sequence, the marginal patents which have been converted into center patents are screened out in chronological order, the conversion time is recorded, and the marginal innovation patent dataset is constructed;

[0011] Step 4: A feature system of marginal innovation opportunity is constructed, the sample features of each marginal innovation patent in the marginal innovation patent dataset are obtained as model input based on the feature system, the conversion time is taken as label, and the prediction model based on neural network is constructed for the supervised training of the conversion time task;

[0012] Step 5: Based on the trained prediction model, the target patent in the patent marginal innovation opportunity identification task is predicted to be converted into a marginal innovation patent, and the conversion time of the target patent is obtained.

[0013] Further, in the step 1, the information of all patents in the technical field to which the target patent belongs is obtained and preprocessed, and a patent dataset is constructed, including:

[0014] The information of the patent includes international patent classification number, patent text, inventor co-occurrence information and disclosure date information, wherein the patent text includes patent title and abstract;

[0015] For the international patent classification number, the first 4 digits are retained to obtain the preprocessed classification number data;

[0016] For the patent text, the patent title and abstract are combined, symbols, numbers, stop words and common words are removed, and morphological reduction is performed to obtain preprocessed patent text data;

[0017] For the inventor co-occurrence information, patents with only single inventor are excluded, and the inventor information of the remaining patents is extracted to obtain preprocessed inventor information;

[0018] The preprocessed data of all patents is constructed into the patent dataset.

[0019] Further, in the step 2, the patent dataset is divided in chronological order, and the marginality evaluation of the two dimensions of outlying and smallness is carried out respectively, and the marginal patent dataset and the center patent dataset based on time sequence are constructed, including:

[0020] For the patent dataset, from the starting year to the ending year, taking the starting year as the reference, the data is divided according to the principle that each increase of one year forms a new sub-dataset, forming a series of patent datasets based on time sequence.

[0021] In the time-based patent data set, the patent text data is vectorized into a patent text vector; the method for vectorizing the patent text data includes a BERT algorithm;

[0022] The classification number data is vectorized into a classification number vector; the method for vectorizing the classification number data includes a term frequency-inverse document frequency algorithm.

[0023] The text vector and the classification number vector are fused and spliced, and the spliced vector is subjected to outlier analysis by a cross analysis method, three outlier detection algorithms of a local outlier factor, a K-nearest neighbor algorithm and an isolation forest are used for analysis, the intersection of the three analysis results is obtained, and a time-based outlier edge patent data set is obtained as an outlier evaluation result.

[0024] In the time-based patent data set, the inventors are taken as nodes, the cooperation relationship of co-inventors is taken as edges, and the actual number of cooperation between the inventors is taken as a connection weight, so as to establish an inventor cooperation relationship network graph;

[0025] The inventor cooperation relationship network graph is input into a graph convolution network to obtain a reconstructed node feature matrix, and then the reconstructed node feature matrix is subjected to link prediction by a graph autoencoder to obtain a link prediction score of all edges in the inventor cooperation relationship network graph as a potential cooperation possibility between the inventors;

[0026] The inventors are taken as nodes, edges are established between nodes with a cooperation possibility higher than a set value, and the cooperation possibility is taken as a connection weight, so as to reconstruct an inventor potential cooperation relationship network graph;

[0027] Based on the inventor cooperation relationship network graph and the inventor potential cooperation relationship network graph, the importance of each node in the two network graphs is calculated by a web page ranking algorithm, the nodes with low importance are selected, and the intersection is obtained, so as to obtain information of an inventor with strong public awareness and a patent as a first author, and a time-based small edge patent data set is obtained as a small evaluation result.

[0028] The intersection of the outlier evaluation result and the small evaluation result is defined as a time-based edge patent data set; and the remaining patents of the non-edge patent data in the time-based patent data set are defined as a time-based center patent data set.

[0029] Further, in step 3, according to the time-based edge patent data set and the center patent data set, an edge patent that has been converted into a center patent is screened out in chronological order, the conversion duration is recorded, and an edge innovation patent data set is constructed, including:

[0030] The type of each patent in the edge-based edge patent dataset and the center patent dataset is marked as an edge patent and a center patent respectively; each patent is sorted according to year from small to large to form a time sequence track of each patent containing a patent type mark;

[0031] The time sequence track of each patent is analyzed using a sliding window method to check whether there is a state transition, which indicates that the patent belongs to an edge patent in the publication year, and becomes a center patent in the subsequent years;

[0032] The patents with state transitions are extracted, and the patent name, patent publication year and transition year are recorded, all edge innovation patents are screened out and the conversion length thereof is obtained, and an edge innovation patent dataset is constructed.

[0033] Further, in step 4, the feature system of the edge innovation opportunity is constructed, including:

[0034] According to the concept and characteristics of the edge innovation patent, based on three dimensions of technical novelty, market attractiveness and cross diversity, a feature system for evaluating the edge innovation opportunity is designed, and each dimension in the feature system has a plurality of quantifiable indexes;

[0035] The quantifiable indexes of the technical novelty include the number of citations, the number of family citations, the number of cited scientific literature, the number of claims, the number of independent claims, the number of dependent claims, the number of pages of literature, and the patent index;

[0036] The quantifiable indexes of the market attractiveness include the number of citations, the number of same-family patents, the number of countries of same-family patents, and the number of emerging industry classifications;

[0037] The quantifiable indexes of the cross diversity include the number of inventors, the number of applicants or patentees, the number of IPC classifications, and the proportion of enterprise types in the applicants.

[0038] The application also provides an edge innovation patent prediction system, comprising:

[0039] The first dataset construction module is used for acquiring information of all patents in a technical field to which a target patent belongs and preprocessing the information according to a patent edge innovation opportunity identification task published by a user and containing the target patent, and constructing a patent dataset;

[0040] The second dataset construction module is used for dividing the patent dataset according to time sequence, and respectively performing edge evaluation in two dimensions of outlyingness and rarity to construct an edge patent dataset and a center patent dataset based on time sequence;

[0041] a third data set construction module for filtering out edge patents that have been converted into core patents in chronological order according to the time-based edge patent data set and the core patent data set, recording the conversion duration, and constructing an edge innovation patent data set;

[0042] a prediction model training module for constructing a feature system of edge innovation opportunities, obtaining sample features of each edge innovation patent in the edge innovation patent data set as model input based on the feature system, taking the conversion duration as a label, and performing supervised training on the prediction model based on a neural network for the conversion time task;

[0043] an edge innovation patent prediction module for predicting the conversion of target patents in the patent edge innovation opportunity identification task into edge innovation patents based on the trained prediction model, and obtaining the conversion time of the target patents.

[0044] The application solves the problem that the prior art is difficult to capture early edge innovation technologies by identifying edge innovation technology hotspots based on a time series analysis method. Compared with the prior art, the application evaluates the edge patents for outlying and smallness, filters out edge innovation patents, continuously tracks the edge innovation patents in the historical data by using the time series characteristics of the patent data, analyzes the changes in the edge innovation patent activity mode through edge innovation feature indicators, realizes multi-dimensional dynamic tracking of the development and evolution of the innovation technology, and accurately identifies edge patents with potential value in the early stage. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 For the execution steps of an edge innovation patent prediction method in an embodiment of the application;

[0046] Figure 2 For a specific process of edge innovation patent evaluation and prediction in an embodiment of the application;

[0047] Figure 3 For an edge innovation patent prediction system in an embodiment of the application. DETAILED DESCRIPTION

[0048] To make the purpose, technical scheme and beneficial effects of the application clearer, the application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the embodiments described herein are only used to explain the application and do not limit the application.

[0049] As shown in Figure 1 An edge innovation patent prediction method comprises the following steps:

[0050] S1, according to the patent edge innovation opportunity identification task published by the user and containing the target patent, obtaining the information of all patents in the technical field to which the target patent belongs and pre-processing, and constructing a patent data set;

[0051] The specific process is shown in S101-S102 in the embodiment: Figure 2

[0052] From the Incopat patent library, obtain patent information S101 in the technical field to which the target patent belongs from 2004 to 2020, including International Patent Classification (IPC) number, patent text, inventor co-occurrence information and disclosure date information, wherein the patent text includes patent title and abstract;

[0053] Preprocess the patent information S101 to construct a patent data set S102. The specific preprocessing method includes:

[0054] For the IPC number, the first four digits (IPC4) are retained to obtain the preprocessed classification number data; for the patent text, the patent title and the abstract are combined, symbols, numbers, stop words and common words are removed, and morphological reduction is performed to obtain the preprocessed patent text data; for the inventor co-occurrence information, patents with only a single inventor are removed, and the inventor information of the remaining patents is extracted to obtain the preprocessed inventor information;

[0055] The data of all patents after preprocessing is constructed into the patent data set S102.

[0056] S2, divide the patent data set according to time sequence, respectively evaluate the edge in two dimensions of outlying and smallness, and construct edge patent data set and center patent data set based on time sequence;

[0057] The specific process is shown in S201-S208 in the embodiment: Figure 2

[0058] For the patent data set S102, from the starting year to the ending year, take the starting year as the benchmark, divide according to the principle that each increase of one year forms a new sub-data set, form a series of patent data sets based on time sequence S201, the specific process is as follows:

[0059] (1) Extract the data of 2004 from the original data set to form the first patent data set;

[0060] (2) Extract the data of 2005 from the original data set, combine with the first patent data set, and form the second patent data set, that is, the second patent data set contains the data of 2004 and 2005;

[0061] ​​(3) Extract the data of 2006 from the original data set separately, and combine it with the second patent data set to form a third patent data set, that is, the third patent data set contains the data of 2004, 2005 and 2006;

[0062] (4) In an increasing annual manner, the data of each subsequent year is sequentially combined with the data of all previous years to form a new data set until all data up to 2020 is included.

[0063] Based on the above process, a series of time-based patent data sets S201 are constructed, each of which is based on the previous sub-data set and adds all the data of the next year.

[0064] In the time-based patent data set S201, the patent text data is vectorized into fixed-length patent text vectors; the method used for vectorizing the patent text data is the BERT algorithm;

[0065] The classification number data is vectorized into classification number vectors that can highlight the differences in patent technology; the method used for vectorizing the classification number data is the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm, which is specifically represented as:

[0066]

[0067] IPC j =(tfidf 1,j ,tfidf 2,j ,...,tfidf x,j )

[0068] Where tfidf i,j represents the TF-IDF value of IPC4 classification number data i in patent j, tf i,j represents the frequency of IPC4 classification number data i in patent j, N represents the total number of patents in the entire time-based patent data set S201, df i,j represents the number of patents containing IPC4 classification number data i in the time-based patent data set S201; δ is a smoothing parameter, in this embodiment, to ensure the effectiveness of the function and prevent the denominator from being zero, it is set to δ = 1; IPC j represents the classification number vector of patent j, x is the total number of IPC4 classification number data; in this way, each patent can be represented by a vector, and the dimension of this vector is equal to the number of IPC4 classification number data.

[0069] The patent text vector and the classification number vector are fused and spliced, the spliced vector S202 is subjected to outlier analysis by using a cross analysis method, three outlier detection algorithms of a Local Outlier Factor (LOF), a K-Nearest Neighbors (KNN) and an Isolation Forest (IF) are used for analysis, an intersection of three analysis results is obtained, and a time-series-based outlier edge patent dataset S203 is obtained as an outlier evaluation result.

[0070] In the time-series-based patent dataset S201, an inventor is taken as a node, a cooperation relationship of a co-inventor is taken as an edge, and an actual cooperation frequency between inventors is taken as a connection weight, and an inventor cooperation relationship network graph S204 is established.

[0071] The inventor cooperation relationship network graph is input into a graph convolution network, a node feature vector of each node is obtained after reconstruction, and a link prediction of the reconstructed node feature vector is performed through a Graph Autoencoder (GAE), a link prediction score of all edges in the inventor cooperation relationship network graph is obtained as a potential cooperation possibility between inventors, and is specifically represented as:

[0072] score uv =GAE(h u ,h v )

[0073] Wherein, h u and h v represent the node feature vectors of the nodes u and v after reconstruction through the graph convolution network, and score uv represents the link prediction score of the edge between the nodes u and v; in the embodiment, the link prediction score score uv greater than or equal to 0.6 represents a high cooperation possibility.

[0074] The inventor is taken as a node, an edge is established between nodes with a high cooperation possibility, and an inventor potential cooperation relationship network graph S205 is reconstructed by taking the cooperation possibility as a connection weight.

[0075] Based on the inventor cooperation relationship network graph S204 and the inventor potential cooperation relationship network graph S205, the importance of each node in the two network graphs is calculated through a Page-Rank algorithm, the nodes with low importance are selected, and a union set is obtained, to obtain information of an inventor with strong popularity and a patent as a first author, and a time-series-based small edge patent dataset S206 is obtained as a small evaluation result.

[0076] Wherein, the Page-Rank algorithm is specifically expressed as:

[0077]

[0078] Wherein, p represents a node to be calculated, PR(p) is the Page-Rank value of node u; d is a damping factor, representing the probability of random walk to other nodes, d = 0.85 in the embodiment; M is the total number of nodes in the network; B q represents the set of all nodes q connected to node p, PR(q) is the Page-Rank value of node q, and L(q) represents the number of edges connected to node q; after sorting all nodes in descending order of Page-Rank value, the last 10% of nodes are regarded as nodes with low importance.

[0079] The intersection of the outlying evaluation result and the rarity evaluation result is defined as a time-series-based edge patent dataset S207;

[0080] The remaining patents of the non-edge patent data in the time-series-based patent dataset S201 are defined as a time-series-based center patent dataset S208.

[0081] S3, according to the time-series-based edge patent dataset and the center patent dataset, filtering out the edge patents that have been converted into center patents in chronological order, recording the conversion duration, and constructing an edge innovation patent dataset;

[0082] Labeling the types of patents in the time-series-based edge patent dataset and the center patent dataset as edge patents and center patents respectively; sorting each patent in chronological order from far to near to form a time sequence track of each patent containing patent type labels;

[0083] Using a sliding window method to analyze the time sequence track of each patent, checking whether there is a state transition, which indicates that the patent belongs to an edge patent in the disclosure year, and becomes a center patent in the subsequent years;

[0084] Extracting patents with state transitions and recording patent names, patent disclosure years and transition years, filtering out all edge innovation patents and obtaining their conversion durations, and constructing an edge innovation patent dataset.

[0085] S4, constructing a feature system of edge innovation opportunities, obtaining sample features of each edge innovation patent in the edge innovation patent dataset as model input based on the feature system, taking the conversion duration as the label, and performing supervised training on the prediction model based on the neural network for the conversion time task;

[0086] Specifically as Figure 2 shown in S401-S406.

[0087] According to the concept and characteristics of edge innovation patents, based on the three dimensions of technical novelty S401, market attractiveness S402 and cross diversity S403, a feature system for evaluating edge innovation opportunities is designed, and each dimension in the feature system has a plurality of quantifiable indicators, as shown in Table 1:

[0088] Table 1 Feature system for evaluating edge innovation opportunities

[0089]

[0090]

[0091] Based on the feature system, all edge innovation patents in the edge innovation patent data set S301 are obtained from the Incopat patent library, and the quantifiable indicators of each year from the patent publication year to 2020 are used as multi-perspective patent data time series vectors, which can consider the development trajectory of the patent from different time nodes and multiple dimensions;

[0092] The multi-perspective patent data time series vectors are combined into a multi-perspective patent data time series matrix as the sample features S405; the sample features S405 are normalized to eliminate the scale difference between the features and used as an input data set;

[0093] A prediction model S406 based on Transformer is constructed, the input data set is divided into a training set and a test set according to a ratio of 8:2, the transformation time is used as a label, and supervised training is performed on the transformation time task, which is specifically represented as:

[0094] y=f Transformer ([X0,X1,...,X T ])

[0095] Wherein, X T represents the multi-perspective patent data time series vector of each patent in the Tth year after publication, y represents the output transformation time, and f Transformer represents a Transformer prediction model based on a multi-head attention mechanism.

[0096] S5, based on the trained prediction model, the target patent in the patent edge innovation opportunity identification task is predicted to be transformed into an edge innovation patent to obtain the transformation time of the target patent; wherein the publication time of the target patent is from 2021 to 2024.

[0097] As Figure 3 shown, the present embodiment provides an edge innovation patent prediction system, comprising:

[0098] A first data set construction module 1 is used to obtain information of all patents in a technical field to which a target patent belongs and to pre-process the information according to a patent edge innovation opportunity identification task published by a user and containing the target patent, and to construct a patent data set;

[0099] A second data set construction module 2 is used to divide the patent data set according to time sequence, to respectively perform edge evaluation in two dimensions of outlying and smallness, to construct an edge patent data set and a center patent data set based on time sequence;

[0100] A third data set construction module 3 is used to filter out edge patents that have been converted into center patents according to time sequence based on the edge patent data set and the center patent data set based on time sequence, to record conversion time length, and to construct an edge innovation patent data set;

[0101] A prediction model training module 4 is used to construct a feature system of edge innovation opportunity, to obtain sample features of each edge innovation patent in the edge innovation patent data set as model input based on the feature system, to take conversion time length as a label, and to perform supervised training on a prediction model based on a neural network for a conversion time task;

[0102] An edge innovation patent prediction module 5 is used to perform prediction on conversion of a target patent in a patent edge innovation opportunity identification task into an edge innovation patent based on the trained prediction model, and to obtain conversion time of the target patent.

[0103] The application solves the problem that the prior art is difficult to capture early edge innovation technology by identifying edge innovation technology hotspots based on a time sequence analysis method.

[0104] The preferred embodiments of the application are described above with reference to the accompanying drawings, and the scope of the right of the embodiments of the application is not limited by this. Any modification, equivalent replacement and improvement made by a person skilled in the art without departing from the scope and substance of the embodiments of the application should be within the scope of the right of the embodiments of the application.

Claims

1. An edge innovation patent prediction method, characterized in that, The application relates to a method for predicting the conversion time of a patent edge innovation opportunity, and belongs to the technical field of intellectual property management. The method comprises the following steps: Step 1: according to a patent edge innovation opportunity identification task published by a user and containing a target patent, information of all patents in a technical field to which the target patent belongs is acquired and preprocessed to construct a patent data set; Step 2: the patent data set is divided according to time sequence, and edge evaluation in two dimensions of outlyingness and rarity is respectively performed to construct a time-based edge patent data set and a center patent data set; Wherein, the edge evaluation of the time-based patent data set in the dimension of rarity comprises: In the time-based patent data set, an inventor is taken as a node, a cooperation relationship of co-inventors is taken as an edge, and an actual cooperation frequency between the inventors is taken as a connection weight, so that an inventor cooperation relationship network graph is established; The inventor cooperation relationship network graph is input into a graph convolution network to obtain a reconstructed node feature vector of each node, and then a link prediction is performed on the reconstructed node feature vector by using a graph autoencoder to obtain a link prediction score of all edges in the inventor cooperation relationship network graph, which is taken as a potential cooperation possibility between the inventors; The inventor is taken as a node, and an edge is established between nodes with high cooperation possibility, and the cooperation possibility is taken as a connection weight, so that an inventor potential cooperation relationship network graph is reestablished, and a link prediction score greater than or equal to 0.6 indicates high cooperation possibility; Based on the inventor cooperation relationship network graph and the inventor potential cooperation relationship network graph, the importance of each node in the two network graphs is respectively calculated by using a webpage ranking algorithm, the nodes with low importance are selected and taken as a union set, information of the inventors with strong rarity and the patents of which the inventors are first authors are obtained, a time-based rare edge patent data set is obtained as a rarity evaluation result, wherein, all nodes are sorted from high to low according to webpage ranking values, and the last 10% nodes are regarded as nodes with low importance; Step 3: according to the time-based edge patent data set and the center patent data set, an edge patent that has been converted into a center patent is screened out according to time sequence, a conversion duration is recorded, and an edge innovation patent data set is constructed; Step 4: a feature system of an edge innovation opportunity is constructed, sample features of each edge innovation patent in the edge innovation patent data set are obtained as model inputs based on the feature system, a conversion duration is taken as a label, a prediction model based on a neural network is constructed, and supervised training is performed on a conversion time task; 2. The method of claim 1, wherein, Step 5: based on the trained prediction model, a prediction is performed on the target patent in the patent edge innovation opportunity identification task to convert into an edge innovation patent, and a conversion time of the target patent is obtained. In step 1, the information of all patents in a technical field to which the target patent belongs is acquired and preprocessed to construct a patent data set, which comprises the following steps: The information of the patent comprises an international patent classification number, a patent text, co-inventor information and disclosure date information, wherein the patent text comprises a patent title and an abstract; For the international patent classification number, the first four data are reserved to obtain preprocessed classification number data; For the patent text, the patent title and the abstract are combined, symbols, numbers, stop words and common words are removed, and morphological reduction is performed to obtain preprocessed patent text data; For the inventor co-occurrence information, remove patents with only a single inventor, extract the inventor information of the remaining patents, and obtain pre-processed inventor information; After preprocessing all patent data, construct the patent data set.

3. The method of claim 1, wherein, In step 2, the patent data set is divided in chronological order, and the marginality of the two dimensions of outlying and smallness is evaluated respectively to construct the marginal patent data set and the central patent data set based on time sequence, including: For the patent data set, from the starting year to the ending year, take the starting year as the basis, divide it into a series of patent data sets based on time sequence according to the principle that each increase in one year forms a new sub-data set; For the patent data set, from the starting year to the ending year, take the starting year as the basis, divide it into a series of patent data sets based on time sequence according to the principle that each increase in one year forms a new sub-data set; For the patent data set, from the starting year to the ending year, take the starting year as the basis, divide it into a series of patent data sets based on time sequence according to the principle that each increase in one year forms a new sub-data set; 4. The method of claim 3, wherein, For the patent data set, from the starting year to the ending year, take the starting year as the basis, divide it into a series of patent data sets based on time sequence according to the principle that each increase in one year forms a new sub-data set; For the patent data set, from the starting year to the ending year, take the starting year as the basis, divide it into a series of patent data sets based on time sequence according to the principle that each increase in one year forms a new sub-data set; For the patent data set, from the starting year to the ending year, take the starting year as the basis, divide it into a series of patent data sets based on time sequence according to the principle that each increase in one year forms a new sub-data set; 5. The method of claim 4, wherein, For the patent data set, from the starting year to the ending year, take the starting year as the basis, divide it into a series of patent data sets based on time sequence according to the principle that each increase in one year forms a new sub-data set; 6. The method of claim 1, wherein, The method for vectorizing patent text data includes BERT algorithm; the method for vectorizing classification number data includes term frequency-inverse document frequency algorithm. In step 3, according to the marginal patent data set and the central patent data set based on time sequence, the marginal patents that have been converted into central patents are screened out in chronological order, the conversion duration is recorded, and the marginal innovation patent data set is constructed, including: For the marginal patent data set and the central patent data set based on time sequence, mark the types of patents in them as marginal patents and central patents respectively; sort each patent in chronological order from far to near to form a time sequence track of each patent containing patent type marks; Use a sliding window method to analyze the time sequence track of each patent to check if there is a state transition, which means that the patent is a marginal patent at the disclosure year, but becomes a central patent in the subsequent years; 7. The method of claim 1, wherein, Extract patents with state transitions and record patent name, patent disclosure year and transition year, screen out all marginal innovation patents and get their conversion duration, and construct a marginal innovation patent data set. In step 4, the feature system of marginal innovation opportunity is constructed, including: According to the concept and characteristics of marginal innovation patents, based on the three dimensions of technical novelty, market attractiveness and cross diversity, a feature system for evaluating marginal innovation opportunities is designed, and each dimension in the feature system has multiple quantifiable indicators; The quantifiable indicators of the novelty of the technology include citation times, family citation times, the number of cited scientific literature, the number of claims, the number of independent claims, the number of dependent claims, the number of document pages, and patent indicators; The quantifiable indicators of the market attractiveness include cited times, the number of sibling patents, the number of countries of sibling patents, and the number of emerging industry classifications; The quantifiable indicators of the cross diversity include the number of inventors, the number of applicants or patentees, the number of IPC classifications, and the proportion of enterprise types in the applicants.

8. An edge innovation patent prediction system characterized by, It comprises: A first data set construction module is configured to obtain information of all patents in a technical field to which a target patent belongs and perform preprocessing according to a patent edge innovation opportunity identification task published by a user and containing the target patent, and construct a patent data set; A second data set construction module is configured to divide the patent data set in chronological order, respectively perform edge evaluation in two dimensions of out-of-the-ordinary and rarity, and construct an edge patent data set and a center patent data set based on time sequence; The edge evaluation of the rarity of the patent data set based on time sequence comprises: In the patent data set based on time sequence, an inventor is taken as a node, a cooperation relationship of co-inventors is taken as an edge, and an actual cooperation frequency between inventors is taken as a connection weight, so as to establish an inventor cooperation relationship network graph; The inventor cooperation relationship network graph is input into a graph convolution network to obtain a reconstructed node feature vector of each node, and then a link prediction is performed on the reconstructed node feature vector by a graph autoencoder to obtain a link prediction score of all edges in the inventor cooperation relationship network graph as a potential cooperation possibility between inventors; An inventor is taken as a node, an edge is established between nodes with high cooperation possibility, and the cooperation possibility is taken as a connection weight, so as to reconstruct an inventor potential cooperation relationship network graph, and a link prediction score greater than or equal to 0.6 indicates high cooperation possibility; Based on the inventor cooperation relationship network graph and the inventor potential cooperation relationship network graph, the importance of each node in the two network graphs is calculated by a page rank algorithm respectively, the nodes with low importance are selected and taken as a union set to obtain information of inventors with strong rarity and patents of which the inventors are first authors, and an edge patent data set based on time sequence is obtained as a rarity evaluation result, wherein all nodes are sorted from high to low according to the page rank values, and the last 10% of the nodes are regarded as nodes with low importance; A third data set construction module is configured to filter out edge patents that have been converted into center patents according to the edge patent data set and the center patent data set based on time sequence, record conversion durations, and construct an edge innovation patent data set; A prediction model training module is configured to construct a feature system of an edge innovation opportunity, obtain sample features of each edge innovation patent in the edge innovation patent data set as model input based on the feature system and take the conversion duration as a label, and perform supervised training on a prediction model based on a neural network for a conversion time task. an edge innovation patent prediction module configured to make a prediction of whether a target patent in an edge innovation opportunity identification task will be converted into an edge innovation patent based on a trained prediction model, and obtain a conversion time of the target patent.

Citation Information

Patent Citations

  • Classification number co-occurrence network construction method, technical opportunity identification method and system

    CN114817567A

  • Technical opportunity recognition method and device, computer equipment and storage medium

    CN117634502A

  • Method, medium and equipment for predicting quoted quantity of achievements based on dynamic knowledge graph

    CN114817571A

  • Enterprise technology analysis method and device based on patent big data and related components

    CN115797115A