A method and device for identifying electrochemical energy storage technologies based on patent text analysis
Through the method based on patent text analysis, the characteristics of electrochemical energy storage technology are extracted and the technical topics are identified through cluster identification, and the problem of poor objectivity and timeliness of the identification of electrochemical energy storage technology in the existing technology is solved, achieving more accurate and timely technical identification.
Patent Information
- Application Number
- CN202510229385.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The qualitative analysis methods used in the prior art to identify electrochemical energy storage technology have problems of poor objectivity and timeliness, and it is difficult to provide accurate and timely evaluation results in the environment of rapid iteration of technology.
Using a method based on patent text analysis, technical features are extracted from patent text through the Dirichlet distribution theme model, combined with the number of applications of patent text in multiple patent institutions, the average marginal effect of technical features is determined, and the target characteristics and technical topics of electrochemical energy storage technology are identified through clustering technology.
It improves the objectivity and timeliness of the identification results of electrochemical energy storage technology, provides more accurate and timely technical identification information, and supports enterprises and research institutions to make scientific decisions.
Smart Images

Figure CN119739866B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method and device for identifying electrochemical energy storage technologies based on patent text analysis. Background Art
[0002] Although new energy sources such as photovoltaic and wind power can reduce carbon emissions, their intermittency and instability pose challenges to the power system. Electrochemical energy storage technologies can effectively store and release energy to ensure the stable operation of the power system. Despite the rapid development of electrochemical energy storage technologies, they still face the problem of uncertain decision-making on technical routes. Therefore, establishing an effective method for identifying electrochemical energy storage technologies can screen numerous electrochemical energy storage technologies and provide a scientific decision-making basis for enterprises and research institutions.
[0003] In the prior art, the identification methods of electrochemical energy storage technologies mainly adopt qualitative analysis methods. Such methods rely on the experience and knowledge of experts and use tools such as the Delphi method, the analytic hierarchy process, and the scenario analysis method to identify and evaluate key technologies. First, a team of domain experts is formed to design an evaluation index system, then expert opinions are collected through questionnaires or seminars, and finally conclusions are drawn by synthesizing expert scores. Due to completely relying on the individual knowledge reserves and experience judgments of experts, in an environment of rapid technological iteration, experts often have difficulty grasping the latest technological trends in a timely and comprehensive manner, resulting in poor objectivity and timeliness of evaluation results. Summary of the Invention
[0004] The present invention provides a method and device for identifying electrochemical energy storage technologies based on patent text analysis to solve the defect of poor objectivity and timeliness of the results of identifying electrochemical energy storage technologies by using qualitative analysis methods in the prior art, and to achieve the improvement of the objectivity and timeliness of the results of identifying electrochemical energy storage technologies.
[0005] The present invention provides a method for identifying electrochemical energy storage technologies based on patent text analysis, including:
[0006] Obtain a patent text database, where the patent text database includes patent texts related to electrochemical energy storage;
[0007] Through the Dirichlet distribution topic model, obtain the topic distribution corresponding to the patent text, and use the topics included in each patent text as the technical features of electrochemical energy storage;
[0008] Obtain the application numbers of the patent text in multiple patent institutions, determine the average marginal effect of each technical feature based on the topic distribution and the application numbers of the patent text, and determine the target technical feature based on the average marginal effect;
[0009] Cluster the target technical features to obtain the identification result of the electrochemical energy storage technology, where the identification result of the electrochemical energy storage technology includes the technical themes obtained by clustering the target technical features.
[0010] According to an electrochemical energy storage technology identification method based on patent text analysis provided by the present invention, the obtaining of the theme distribution corresponding to the patent text through the Dirichlet distribution topic model includes:
[0011] Sample the tag distribution of the patent text through the Dirichlet distribution topic model, where the tag is the Derwent manual code;
[0012] Sample the theme distribution for each tag through the Dirichlet distribution topic model to obtain each theme and the theme distribution corresponding to the tag of the patent text;
[0013] Sample the theme word distribution for each theme through the Dirichlet distribution topic model to obtain each theme word and the theme word distribution corresponding to the theme of the patent text.
[0014] According to an electrochemical energy storage technology identification method based on patent text analysis provided by the present invention, the determination of the average marginal effect of each technical feature based on the theme distribution of the patent text and the application quantity includes:
[0015] Construct a fitting relationship between the technical features in the patent text and the application quantity through an ordered logistic regression model;
[0016] Calculate the average marginal effect of each technical feature in each patent text based on the fitting relationship.
[0017] According to an electrochemical energy storage technology identification method based on patent text analysis provided by the present invention, the determination of the target technical features based on the average marginal effect includes:
[0018] Select the technical features with an average marginal effect lower than a preset threshold in the patent text with the lowest application quantity as the target technical features.
[0019] According to an electrochemical energy storage technology identification method based on patent text analysis provided by the present invention, the clustering of the target technical features to obtain the identification result of the electrochemical energy storage technology includes:
[0020] Obtain the similarity matrix corresponding to the target technical features, where the similarity matrix includes the similarities between the target technical features;
[0021] Based on the similarity matrix, determine the clustering result based on the affinity propagation algorithm, and use each cluster in the clustering result as a technical theme;
[0022] Determine the identification result of the electrochemical energy storage technology based on the technical theme.
[0023] According to an identification method of electrochemical energy storage technology based on patent text analysis provided by the present invention, after determining the identification result of the electrochemical energy storage technology based on the technical theme, it includes:
[0024] Conduct a technical life cycle analysis on the technical features included in the technical theme to obtain the technical maturity of each technical feature in the technical theme;
[0025] Based on the average value of the technical maturity of each technical feature included in the technical theme and each technical feature included in the technical theme, generate an evaluation result for each technical theme in the identification result of the electrochemical energy storage technology.
[0026] The present invention also provides an identification device for electrochemical energy storage technology based on patent text analysis, including:
[0027] A patent text acquisition module for acquiring a patent text database, where the patent text database includes patent texts related to electrochemical energy storage;
[0028] A text theme mining module for obtaining the theme distribution corresponding to the patent text through the Dirichlet distribution theme model, and using the themes included in each patent text as technical features of electrochemical energy storage;
[0029] A technical feature extraction module for obtaining the number of applications of the patent text in multiple patent institutions, determining the average marginal effect of each technical feature based on the theme distribution and the number of applications of the patent text, and determining target technical features based on the average marginal effect;
[0030] A technical feature clustering module for clustering the target technical features to obtain an identification result of the electrochemical energy storage technology, where the identification result of the electrochemical energy storage technology includes technical themes obtained by clustering the target technical features.
[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the identification method of electrochemical energy storage technology based on patent text analysis as described in any one of the above.
[0032] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for identifying an electrochemical energy storage technology based on patent text analysis as described in any one of the above.
[0033] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for identifying an electrochemical energy storage technology based on patent text analysis as described in any one of the above.
[0034] The method for identifying an electrochemical energy storage technology based on patent text analysis and the device provided by the present invention, the method includes: obtaining a patent text database, where the patent text database includes patent texts related to electrochemical energy storage; through the Dirichlet distribution topic model, obtaining the topic distribution corresponding to the patent texts, and taking the topics included in each patent text as the technical features of the electrochemical energy storage; obtaining the number of applications of the patent texts in multiple patent institutions, and based on the topic distribution and the number of applications of the patent texts, determining the average marginal effect of each technical feature, and determining the target technical feature based on the average marginal effect; clustering the target technical features to obtain the identification result of the electrochemical energy storage technology, where the identification result of the electrochemical energy storage technology includes the technical topics obtained by clustering the target technical features.
[0035] In this way, by analyzing the patent texts related to electrochemical energy storage, using the Dirichlet distribution topic model, mining the technical information of the patent texts and extracting technical features, and then facing the problem of inconsistent content levels of the identification results, selecting target technical features based on the patent family scale to eliminate broad or low-value technical features, and finally clustering the selected target technical features into technical topics. Since patent applications have foresight, realizing the identification of electrochemical energy storage technologies based on the analysis of a large number of patent texts can improve the objectivity and timeliness of the identification results and provide accurate reference information for decision-makers. Description of the Drawings
[0036] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 is a schematic flowchart of the method for identifying an electrochemical energy storage technology based on patent text analysis provided by the present invention.
[0038] Figure 2 is a probability diagram of the Dirichlet distribution topic model in the method for identifying an electrochemical energy storage technology based on patent text analysis provided by the present invention.
[0039] Figure 3 It is a schematic structural diagram of an electrochemical energy storage technology identification device provided by the present invention based on patent text analysis.
[0040] Figure 4 It is a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners
[0041] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] The following Figure 1 - Figure 2 describes an electrochemical energy storage technology identification method provided by the present invention based on patent text analysis. As Figure 1 shown, the electrochemical energy storage technology identification method based on patent text analysis includes the steps of:
[0043] S110. Obtain a patent text database, where the patent text database includes patent texts related to electrochemical energy storage;
[0044] S120. Through the Dirichlet distribution topic model, obtain the topic distribution corresponding to the patent text, and use the topics included in each patent text as the technical features of electrochemical energy storage;
[0045] S130. Obtain the application numbers of the patent text in multiple patent institutions, determine the average marginal effect of each technical feature based on the topic distribution and application number of the patent text, and determine the target technical feature based on the average marginal effect;
[0046] S140. Cluster the target technical features to obtain an electrochemical energy storage technology identification result, where the electrochemical energy storage technology identification result includes the technical topics obtained by clustering the target technical features.
[0047] The electrochemical energy storage technology identification method provided by the present invention analyzes the patent texts related to electrochemical energy storage, uses the Dirichlet distribution topic model to mine the technical information of the patent texts and extract technical features, and then faces the problem of inconsistent content levels in the identification results. Based on the patent family size, target technical features are selected to eliminate broad or low-value technical features, and finally the selected target technical features are clustered into technical topics. Since patent applications have forward-looking nature, realizing electrochemical energy storage technology identification based on a large amount of patent text analysis can improve the objectivity and timeliness of the identification results and provide accurate reference information for decision-makers.
[0048] By retrieving and collecting electrochemistry energy storage related patents in the patent database, a patent text database used in the method provided by the present invention is constructed. In a possible implementation, electrochemistry energy storage related patents can be retrieved in the Derwent Innovations Index to obtain the patent text database used in the method provided by the present invention. The Derwent Innovations Index is a globally used patent information and scientific and technological information institution, which can comprehensively reflect the international scientific and technological development trends. The Derwent database provides detailed classification information, including patent families, Derwent classification codes and manual codes, and this information is more refined than conventional patent classification systems (such as the International Patent Classification IPC). Collect electrochemistry energy storage related patent text data, and collect patent data since the 21st century with the search formula of "DC=(X16)". The data is preprocessed and a high-quality patent text database is formed according to fields (including patent titles, abstracts, Derwent manual codes, application details, patent family information, etc.).
[0049] Through the Dirichlet distribution topic model, technical information mining is carried out on the patent texts in the patent text database. The Derwent manual code can effectively index a patent. In the method provided by the present invention, the Derwent manual code is embedded into the Dirichlet distribution topic model to constrain the model output result and obtain directional label information. Specifically, through the Dirichlet distribution topic model, the topic distribution corresponding to the patent text is obtained, including:
[0050] Sampling the label distribution of the patent text through the Dirichlet distribution topic model, and the label is the Derwent manual code;
[0051] Sampling the topic distribution for each label through the Dirichlet distribution topic model to obtain the topic distribution corresponding to each topic and the label of the patent text;
[0052] Sampling the topic word distribution for each topic through the Dirichlet distribution topic model to obtain the topic word distribution corresponding to each topic word and the topic of the patent text.
[0053] In the method provided by the present invention, a label dimension is introduced into the document-topic-vocabulary hierarchical structure in the Dirichlet distribution topic model to form a four-level structure of document-label-topic-word.
[0054] The specific process of forming the four-level structure of document-label-topic-word is as follows:
[0055] (1) For each patent text , sampling the label distribution , where is a vector with a length of and the The first element is ;
[0056] (2)For each tag , sample the topic distribution ;
[0057] (3)For each topic , sample the topic-word distribution
[0058] (4)For the generation of each word in each document , first sample a tag from the tag distribution , then sample a latent topic from the topic distribution , and finally sample a word from the topic-word distribution .
[0059] The above process is as Figure 2 shown, and the nodes in bold represent that their values can be directly observed. Through the method provided by the present invention, the joint distribution of document-topic and the topic-word distribution can be obtained, and estimated by the Markov chain Monte Carlo sampling method.
[0060] In the above process and Figure 2 , D represents the number of patent texts in the patent text database, d represents one of the patent texts, represents the number of words in the patent text d, L represents the number of all tags in the patent text database, is one of the tags, represents all the tags owned by the document d, represents the tag distribution of the document d, represents the tag 's number of latent topics, represents the topic distribution of the document d's tag , represents the topic-word distribution of the k-th topic of the tag , represents the hyperparameter of the prior Dirichlet distribution of the topic distribution , represents the hyperparameter of the prior Dirichlet distribution of the topic-word distribution , represents the unobserved word, and z represents the latent topic of each observed word .
[0061] The method provided by the present invention uses tag information to control the extraction of only the topics associated with the tags it has in each patent text. Taking the patent classification system as tag information, it can be distinguished from traditional unsupervised text mining techniques, limit the topics mined from patent texts within a specific technical scope, so that the output topic results not only include the technical field information of patent classification tags, but also can reveal more detailed technical content that is more in line with the data sub - technical field (electrochemical energy storage technology) at a finer - grained level. Patent classification information provides a macro - technical field framework, limits the analysis scope, and reduces noise interference; text semantic information mines fine - grained technical features in patents and captures deep - level semantic associations. The combination of the two realizes multi - level technical feature extraction from macro to micro, not only improving the accuracy of technical information mining, but also enhancing the cross - field technology identification ability and adapting to complex technical ecosystems.
[0062] Taking the topics output by the model as technical features, their specific technical information is explained through the distribution of their topic words. At the same time, the method provided by the present invention represents each patent text as a document - topic distribution (that is, the proportion of each technical feature in the document, and the sum of the proportions of each topic in all documents is called the popularity of that topic), so as to obtain the technical feature vector of each patent.
[0063] Generally speaking, patents with layouts in multiple patent agencies have higher technical value. In the method provided by the present invention, the number of applications of patent texts in multiple patent agencies is obtained, and based on the topic distribution and the number of applications of patent texts, the average marginal effect of each technical feature is determined, specifically including:
[0064] Constructing a fitting relationship between the technical features in the patent text and the number of applications through an ordered logistic regression model;
[0065] Calculating the average marginal effect of each technical feature in each patent text based on the fitting relationship.
[0066] Specifically, the application numbers of the main patent application institutions with a relatively large number of patent applications in the patent text database can be obtained (such as the State Intellectual Property Office of China, the European Patent Office, the Japan Patent Office, the Korean Intellectual Property Office, and the United States Patent and Trademark Office). Taking 5 patent application institutions as an example, the application numbers of each patent in the 5 patent institutions in the patent text database are counted to obtain an integer ranging from 0 to 5: getting 0 if the patent is not applied in any of the patent institutions, getting 1 if it is applied in one patent institution, getting 2 if it is applied in two patent institutions, and so on. The obtained integers are encoded to get multiple categories. For example, 0 and 1 are encoded as category 0, 2 is encoded as category 1, 3 is encoded as category 2, and 4 and 5 are encoded as category 3. In this way, a measure index of patent technology value can be constructed as an ordered component with four categories. The higher the value, the more patent application institutions the patent is applied in, and the higher the technology value.
[0067] Based on the theme distribution of the patent text and the application number of the patent text, a fitting relationship between the technical features in the patent text and the application number is constructed. The construction of the fitting relationship can be carried out through an Ordinal Logistic Regression (OLR) model. That is, the application number index = OLR (patent technical feature vector). The relationship between the technical feature distribution of the patent text and the application number of the patent text can be expressed by the formula:
[0068] ;
[0069] where, represents the division threshold of the application number category j, represents the i-th item of the patent technical feature vector, M is the total number of technical features, represents the regression coefficient of the corresponding technical feature, Y represents the measure index of patent technology value, and P represents the probability.
[0070] According to the fitting result of the ordered logistic regression model, the average marginal effect (AME) of each technical feature can be calculated as follows:
[0071] ;
[0072] ;
[0073] where, N is the total number of patent texts, is the marginal effect calculated for the th patent text . is the average marginal effect of the k-th technical feature.
[0074] The marginal effect of a technical feature calculated for a certain patent text reflects the degree of influence of this technical feature on the number of applications of this patent text (reflecting the technical value of the patent). Taking the average of the marginal effects of a technical feature calculated for all patent texts, the average marginal effect of this technical feature is obtained.
[0075] Determining the target technical feature based on the average marginal effect includes:
[0076] Selecting the technical features whose average marginal effects in the patent text with the lowest number of applications are lower than a preset threshold as the target technical features.
[0077] Specifically, the average marginal effect of category 0 (representing the category with the lowest number of applications) can be selected as the basis for selecting the target technical features, that is, when the component of the selected technical feature increases by one unit, the change value of the probability that this patent belongs to category 0 (that is, the patent technology value measurement index is 0 or 1). Obviously, when this value of a certain technical feature is negative, the higher the component of this technical feature in the patent, the more it can reduce the probability that it belongs to category 0, that is, the probability that it belongs to high-tech value increases. Therefore, if the average marginal effect of a technical feature in the category representing the lowest technical value is negative, it indicates that it has a positive correlation with improving the technical value of the patent, and at the same time, the larger the absolute value of this value, the greater the influence of this technical feature. Finally, all technical features with an average marginal effect less than 0 (at a significance level of 0.05) in the category representing the lowest technical value are selected as the key technical features (that is, target technical features) of the electrochemical energy storage technology, and the absolute value of the average marginal effect of these features is taken as the numerical basis for calculating the technical value in the follow-up.
[0078] Quantify the technical value of patent technology by constructing a new technical value measurement index, and establish an influence model between the technical value and technical features based on an ordered logistic regression model. The method provided by the present invention innovatively combines technical value measurement and statistical modeling, and can effectively identify the key technical features that have a greater impact on the technical value, providing a scientific basis for key technology screening.
[0079] The meaning of the target technical features selected through the above steps is clear and precisely points to the technical content of high-tech value. In order to support technical management decisions, in the method provided by the present invention, the target technical features are further clustered and condensed into more representative meso-technical themes in order to refine the related technologies of electrochemical energy storage.
[0080] Clustering the target technical features to obtain the identification result of electrochemical energy storage technology, including:
[0081] Obtain the similarity matrix corresponding to the target technical features, where the similarity matrix includes the similarities between the target technical features;
[0082] Based on the similarity matrix, determine the clustering result using the affinity propagation algorithm, and take each cluster in the clustering result as a technical theme;
[0083] Determine the identification result of the electrochemical energy storage technology based on the technical theme.
[0084] First, construct a similarity matrix with the selected M target technical features, where , and is the set of the top 15 topic keywords of the th target technical feature. Each element represents the Jaccard coefficient between technical features and :
[0085] ;
[0086] Then, construct the affinity propagation algorithm (AP) based on this similarity matrix to cluster the technical features. Use an iterative message passing mechanism to propagate two pieces of information, "responsibility value" and "availability", between data points to determine the cluster center of each data point. is the responsibility value, indicating the suitability of data point considering data point as its cluster center. is the availability, indicating the suitability of data point as the cluster center of data point . Update the following iteration until convergence.
[0087] ;
[0088] ;
[0089] .
[0090] According to the final value, select the largest as the cluster center. By passing information between neighboring points and continuously iterating the responsibility matrix and the availability matrix, finally, the number of clusters and the cluster centers are adaptively obtained according to the calculation results of the two matrices, realizing the clustering of technical features into technical themes at the meso scale.
[0091] In patent theme clustering, the confirmation of the number of themes and the robustness of clustering have always been difficult problems. In the method provided by the present invention, through the affinity propagation clustering algorithm, it is not necessary to specify the number of clusters in advance. Instead, all data points are regarded as potential cluster centers, and clustering is achieved through message passing, thus avoiding the subjectivity of the hyperparameter of determining the number of clusters. At the same time, the Jaccard coefficient is used instead of the Euclidean distance metric, providing higher robustness.
[0092] Taking the Jaccard coefficient of the set of technical feature theme keywords as the similarity metric, constructing a similarity matrix, and adaptively determining the number of clusters and centers through the affinity propagation algorithm, organizing the technical features into moderately specific and strongly strategic-oriented technical themes. The method provided by the present invention does not require presetting the number of clusters, avoiding subjectivity, and at the same time has higher robustness when dealing with non-Euclidean space data, and can better adapt to the complexity and diversity of technical features. By extracting more representative meso-technical themes from the clustering results, the interpretability and practicality of decision support are improved.
[0093] Through the above clustering operation, key technical themes are finally obtained. For each key technical theme, the specific content of the top 15 theme keywords with the highest content within the theme will be combined to identify the key electrochemical energy storage technologies represented by each technical theme.
[0094] In a possible implementation, it is also possible to further obtain the technology maturity of each technical theme, and combine the technology maturity of the technical theme and the number of applications corresponding to the technical theme (reflecting the technical value of the technical feature) to evaluate each technical theme, so as to provide more reference information for technology decision-makers. That is, after determining the electrochemical energy storage technology identification result based on the technical theme, it includes:
[0095] Conduct a technology life cycle analysis on the technical features included in the technical theme to obtain the technology maturity of each technical feature in the technical theme;
[0096] Based on the average value of the technology maturity of each technical feature included in the technical theme and each technical feature included in the technical theme, generate the evaluation results of each technical theme in the electrochemical energy storage technology identification result.
[0097] Specifically, the method provided by the present invention models the identified target technical features of electrochemical energy storage through the Gompertz growth S curve in the technology life cycle analysis, expressed as:
[0098]
[0099] In the formula, is the growth saturation value, is the parameter controlling the growth rate, t is the year, indicating the popularity of the target technical features. Based on the growth data of the popularity of all target technical features, an S-curve is fitted and parameters are estimated. Take (i.e., the ratio of the situation in the most recent year to the predicted maximum limit) as the technology maturity of the technical feature. Finally, for each technology theme, the mean value of the maturity of all technical features within its cluster is used as the quantification index of its technology maturity.
[0100] , where m represents the number of technical features included under the technology theme.
[0101] This method measures the application breadth and development depth of a technical feature by measuring the technical value and technology maturity of the key technical features. The technical value and technology maturity of the technical features belonging to the same theme are averaged, and the obtained mean value is used as the technical value and technology maturity of the identified key technologies for electrochemical energy storage. After normalizing the measurement indicators of the technical value and technology maturity of each technology, they are used as the vertical axis and the horizontal axis respectively to construct a key technology evaluation matrix. According to the criteria of high and low technical value (taking the approximation of the middle value of the numerical distribution range of the technical value, i.e., 0.15) and the early, middle, and late stages of technology development (taking the technology maturity of 0.25 and 0.75), the development situation types of the key technologies are qualitatively divided. Finally, the early and middle-stage key technologies for electrochemical energy storage with high technical value are obtained.
[0102] Model the life cycle of the technology through the Gompertz growth curve, and quantify the maturity of the technology theme through the within-cluster mean value; construct a key technology evaluation matrix through the technical value and technology maturity, and finally identify the early and middle-stage key technologies with high technical value. Reveal the laws of different technologies in terms of development stage and market value, and provide clear guidance for technology management.
[0103] The following describes the electrochemical energy storage technology identification device provided by the present invention based on patent text analysis. The electrochemical energy storage technology identification device described below and the electrochemical energy storage technology identification method described above can be mutually referred to. As Figure 3 shown, the electrochemical energy storage technology identification device provided by the present invention includes:
[0104] A patent text acquisition module 310, configured to acquire a patent text database, where the patent text database includes patent texts related to electrochemical energy storage;
[0105] A text theme mining module 320, configured to obtain the theme distribution corresponding to the patent text through the Dirichlet distribution theme model, and use the themes included in each patent text as the technical features of electrochemical energy storage;
[0106] The technical feature extraction module 330 is configured to obtain the number of applications of the patent text in multiple patent institutions, determine the average marginal effect of each technical feature based on the theme distribution and the number of applications of the patent text, and determine the target technical feature based on the average marginal effect;
[0107] The technical feature clustering module 340 is configured to cluster the target technical features to obtain an identification result of the electrochemical energy storage technology, and the identification result of the electrochemical energy storage technology includes the technical themes obtained by clustering the target technical features.
[0108] Figure 4 An example of the physical structure diagram of an electronic device is shown as Figure 4 shown. The electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete mutual communication through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the method for identifying the electrochemical energy storage technology based on the patent text analysis. The method includes: obtaining a patent text database, where the patent text database includes patent texts related to electrochemical energy storage; obtaining the theme distribution corresponding to the patent text through the Dirichlet distribution topic model, and taking the themes included in each patent text as the technical features of the electrochemical energy storage; obtaining the number of applications of the patent text in multiple patent institutions, determining the average marginal effect of each technical feature based on the theme distribution and the number of applications of the patent text, and determining the target technical feature based on the average marginal effect; clustering the target technical features to obtain an identification result of the electrochemical energy storage technology, and the identification result of the electrochemical energy storage technology includes the technical themes obtained by clustering the target technical features.
[0109] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, and other various media that can store program codes.
[0110] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the electrochemical energy storage technology identification method based on patent text analysis provided by the above-mentioned various methods. The method includes: obtaining a patent text database, which includes patent texts related to electrochemical energy storage; through the Dirichlet distribution topic model, obtaining the topic distribution corresponding to the patent texts, and taking the topics included in each patent text as the technical features of electrochemical energy storage; obtaining the number of applications of the patent texts in multiple patent institutions, and based on the topic distribution and the number of applications of the patent texts, determining the average marginal effect of each technical feature, and determining the target technical feature based on the average marginal effect; clustering the target technical features to obtain the electrochemical energy storage technology identification result, and the electrochemical energy storage technology identification result includes the technical topics obtained by clustering the target technical features.
[0111] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the electrochemical energy storage technology identification method based on patent text analysis provided by the above-mentioned various methods. The method includes: obtaining a patent text database, which includes patent texts related to electrochemical energy storage; through the Dirichlet distribution topic model, obtaining the topic distribution corresponding to the patent texts, and taking the topics included in each patent text as the technical features of electrochemical energy storage; obtaining the number of applications of the patent texts in multiple patent institutions, and based on the topic distribution and the number of applications of the patent texts, determining the average marginal effect of each technical feature, and determining the target technical feature based on the average marginal effect; clustering the target technical features to obtain the electrochemical energy storage technology identification result, and the electrochemical energy storage technology identification result includes the technical topics obtained by clustering the target technical features.
[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying electrochemical energy storage technology based on patent text analysis, characterized in that: include: Obtaining a patent text database, wherein the patent text database includes patent texts related to electrochemical energy storage; The topic distribution corresponding to the patent text is obtained by using the Dirichlet distribution topic model, and the topics included in each of the patent texts are used as technical features of electrochemical energy storage; Obtain the number of applications for the patent text in multiple patent institutions, determine the average marginal effect of each of the technical features based on the subject distribution of the patent text and the number of applications, and determine the target technical feature based on the average marginal effect; Clustering the target technical features to obtain an electrochemical energy storage technology identification result, wherein the electrochemical energy storage technology identification result includes a technical theme obtained by clustering the target technical features; The topic distribution corresponding to the patent text is obtained by using the Dirichlet distribution topic model, including: The patent text is sampled and labeled using the Dirichlet distribution topic model, wherein the label is a Derwent manual code; Sampling the topic distribution of each of the labels through the Dirichlet distribution topic model to obtain the topic distribution corresponding to each topic and the label of the patent text; Sampling the subject word distribution of each of the topics through the Dirichlet distribution topic model to obtain the subject word distribution corresponding to each subject word and the subject word distribution of the topic of the patent text; The determining of the average marginal effect of each of the technical features based on the subject distribution of the patent text and the number of applications includes: The fitting relationship between the technical features in the patent text and the number of applications is constructed by an ordered logistic regression model. The relationship between the technical features in the patent text and the number of applications in the patent text can be expressed by the formula: ; in, represents the classification threshold of application quantity category j, represents the i-th item of the patent technical feature vector, M is the total number of technical features, express The corresponding regression coefficient of the technical feature, Y represents the patent technology value measurement index, the patent technology value measurement index is an ordered component, the higher the value, the more patent application institutions the patent has been applied for, and P represents the probability; Calculating the average marginal effect of each of the technical features in each of the patent texts based on the fitting relationship; Determining the target technical features based on the average marginal effect includes: The technical feature whose average marginal effect in the patent text with the lowest number of applications is lower than a preset threshold is selected as the target technical feature.
2. The electrochemical energy storage technology identification method based on patent text analysis according to claim 1 is characterized in that: The clustering of the target technology features to obtain electrochemical energy storage technology identification results includes: Obtaining a similarity matrix corresponding to the target technical features, wherein the similarity matrix includes similarities between the target technical features; Based on the similarity matrix, a clustering result is determined based on a neighbor propagation algorithm, and each cluster in the clustering result is used as a technical theme; The electrochemical energy storage technology identification result is determined based on the technical subject.
3. The electrochemical energy storage technology identification method based on patent text analysis according to claim 2 is characterized in that: After determining the electrochemical energy storage technology identification result based on the technical subject, the method includes: Performing a technology life cycle analysis on the technical features included in the technical subject to obtain the technical maturity of each of the technical features in the technical subject; Based on the average value of the technical maturity of each of the technical features included in the technical subject and each of the technical features included in the technical subject, an evaluation result of each of the technical subjects in the electrochemical energy storage technology identification result is generated.
4. An electrochemical energy storage technology identification device based on patent text analysis, characterized in that: include: A patent text acquisition module, used to acquire a patent text database, wherein the patent text database includes patent texts related to electrochemical energy storage; A text topic mining module, used to obtain the topic distribution corresponding to the patent text through the Dirichlet distribution topic model, and use the topics included in each of the patent texts as the technical features of electrochemical energy storage; A technical feature extraction module is used to obtain the number of applications of the patent text in multiple patent institutions, determine the average marginal effect of each of the technical features based on the subject distribution of the patent text and the number of applications, and determine the target technical feature based on the average marginal effect; A technical feature clustering module, used for clustering the target technical features to obtain an electrochemical energy storage technology identification result, wherein the electrochemical energy storage technology identification result includes a technical theme obtained by clustering the target technical features; The topic distribution corresponding to the patent text is obtained by using the Dirichlet distribution topic model, including: The patent text is sampled and labeled using the Dirichlet distribution topic model, wherein the label is a Derwent manual code; Sampling the topic distribution of each of the labels through the Dirichlet distribution topic model to obtain the topic distribution corresponding to each topic and the label of the patent text; Sampling the subject word distribution of each of the topics through the Dirichlet distribution topic model to obtain the subject word distribution corresponding to each subject word and the subject word distribution of the topic of the patent text; The determining of the average marginal effect of each of the technical features based on the subject distribution of the patent text and the number of applications includes: The fitting relationship between the technical features in the patent text and the number of applications is constructed by an ordered logistic regression model. The relationship between the technical features in the patent text and the number of applications in the patent text can be expressed by the formula: ; in, represents the classification threshold of application quantity category j, represents the i-th item of the patent technical feature vector, M is the total number of technical features, express The corresponding regression coefficient of the technical feature, Y represents the patent technology value measurement index, the patent technology value measurement index is an ordered component, the higher the value, the more patent application institutions the patent has been applied for, and P represents the probability; Calculating the average marginal effect of each of the technical features in each of the patent texts based on the fitting relationship; Determining the target technical features based on the average marginal effect includes: The technical feature whose average marginal effect in the patent text with the lowest number of applications is lower than a preset threshold is selected as the target technical feature.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the electrochemical energy storage technology identification method based on patent text analysis as described in any one of claims 1 to 3 is implemented.
6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the electrochemical energy storage technology identification method based on patent text analysis as described in any one of claims 1 to 3 is implemented.
7. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the electrochemical energy storage technology identification method based on patent text analysis as described in any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Large model-based science and technology text related cue word generation method and system
CN119357375A