Evolution analysis method and device for emerging technology and storage medium

By dividing time periods, preprocessing, dimensionality reduction and clustering patent data, combining cosine similarity analysis between word vectors and document vectors, the evolution path of emerging technologies is determined, which solves the problem of low accuracy in the existing technology and achieves more efficient technical evolution analysis.

CN120543328AActive Publication Date: 2025-08-26CAPITAL UNIV OF ECONOMICS & BUSINESS
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510538837.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-26
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing emerging technology evolution analysis methods have the problem of low accuracy, especially when processing patent data, which ignores the technical details and topic changes in the patent content, resulting in inaccurate analysis results.

Method used

By obtaining the patent data of the target technology, pre-processing is performed according to the time period, the patent abstract text is converted into a text vector of the first preset dimension, dimensionality reduction and clustering are performed, keywords are extracted, the cosine similarity between the word vector and the document vector is calculated, the topic name and association relationship are determined, and the evolution path of the technology is determined.

Benefits of technology

It improves the accuracy of the evolution analysis of emerging technologies, can efficiently capture the development path of technology, and solves the problem of low accuracy in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543328A_ABST
    Figure CN120543328A_ABST
Patent Text Reader

Abstract

The invention discloses an emerging technology evolution analysis method and device and a storage medium. According to the method, firstly, patent data of a target technology is obtained and divided according to time periods, so that the development states of the technology on different time nodes are observed; secondly, performing preprocessing and text vector conversion on the patent data, and extracting main technical themes in each time period through dimension reduction processing and clustering analysis; secondly, splicing the abstract text of each topic to obtain a target abstract text, and extracting keywords from the target abstract text by calculating cosine similarity between word vectors and document vectors, so as to determine a topic name of each topic. And finally, determining an incidence relation according to the cosine similarity between the topics, and determining an evolution path of the target technology. Therefore, the evolution path of the target technology can be efficiently and accurately captured, the accuracy of technology evolution analysis is improved, and the technical problem of low accuracy of technology evolution analysis in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data technology, and in particular to a method, device and storage medium for analyzing the evolution of emerging technologies. Background Art

[0002] Emerging technology evolution analysis is a method that studies the processes and patterns of a technology's emergence, maturity, diffusion, and integration with other technologies. It can help individuals, businesses, governments, and other entities better understand the trends, patterns, and impacts of technological development, enabling them to make more informed decisions. Businesses can use technology evolution analysis to predict which emerging technologies are likely to become mainstream in the future, thereby allocating R&D resources and formulating technology strategies in advance to avoid technological lag or market failure. Governments can use technology evolution analysis to formulate science and technology policies, directing resources toward strategically important emerging technology areas, and promoting industrial upgrading and economic restructuring.

[0003] Traditionally, technology evolution analysis relies heavily on expert experience and qualitative descriptions. This approach is not only time-consuming and labor-intensive, but also susceptible to subjective factors, making it difficult to fully and objectively reflect the true evolution of technology. In recent years, with the continuous advancement of big data and artificial intelligence technologies, methods for analyzing technology evolution using patent data have gradually emerged. As a key vehicle for technological innovation, patent data contains a wealth of technological information and development trends, providing new perspectives and tools for analyzing technology evolution.

[0004] However, existing methods for analyzing technological evolution based on patent data still have many shortcomings. For example, some methods simply count the number of patents, ignoring the technical details and thematic changes within the patent content, resulting in low analysis accuracy. Other methods, while attempting to analyze patent text, often suffer from low accuracy in topic extraction and keyword recognition due to the high dimensionality and noise interference of text vectors.

[0005] Therefore, how to efficiently process and analyze patent data and accurately capture the evolution path of technology has become an urgent problem to be solved in the current field of technology evolution analysis. Summary of the Invention

[0006] The embodiments of the present disclosure provide an emerging technology evolution analysis method to at least solve the technical problem of low accuracy of technology evolution analysis in the prior art.

[0007] According to one aspect of an embodiment of the present disclosure, a method for analyzing the evolution of emerging technologies is provided, comprising: obtaining patent data of a target technology for evolution analysis, dividing the patent data into S time periods according to time periods; preprocessing the patent data of each time period, converting the abstract text of each patent into a first text vector of a first preset dimension; performing dimensionality reduction processing on the first text vector to obtain a second text vector of a second preset dimension, and clustering the second text vector to obtain M topics for each time period; for each time period, concatenating the abstract texts of all patent data in each topic to obtain a target abstract text for each topic; converting each word in the target abstract text of each topic in each time period into a word vector of the first preset dimension, and converting the entire document of each patent of each topic in each time period into the word vector of the first preset dimension. A document vector of a preset dimension; calculating the cosine similarity between the word vector and the document vector, and extracting N words from the target summary text as keywords based on the N results with the highest cosine similarity, to obtain N keywords for each topic in each time period; determining the topic name of each topic in each time period based on the N keywords, calculating the cosine similarity between all topics in the two previous and next time periods based on the topic names, and determining the association relationship between the topics based on the cosine similarity between the topics; determining the evolution path of the target technology based on the association relationship; wherein, S is an integer greater than or equal to 3, M is an integer greater than or equal to 2, and N is an integer greater than or equal to 2; the first preset dimension and the second preset dimension are pre-set, and the first preset dimension is greater than the second preset dimension.

[0008] According to another aspect of the embodiments of the present disclosure, a storage medium is further provided. The storage medium includes a stored program, wherein the above method is executed by a processor when the program is running.

[0009] According to another aspect of the embodiment of the present disclosure, an emerging technology evolution analysis device is also provided, including: a patent data acquisition module, configured to obtain patent data of a target technology for evolution analysis, and divide the patent data into S time periods according to time periods; an abstract vectorization module, configured to preprocess the patent data of each time period, and convert the abstract text of each patent into a first text vector of a first preset dimension; a dimensionality reduction module, configured to perform dimensionality reduction processing on the first text vector to obtain a second text vector of a second preset dimension, and cluster the second text vector to obtain M topics for each time period; a splicing module, configured to splice the abstract texts of all patent data in each topic for each time period to obtain a target abstract text for each topic; a document vectorization module, configured to convert each word in the target abstract text of each topic in each time period into a word vector of the first preset dimension, and cluster each word of each topic in each time period The entire document of the patent is converted into a document vector of the first preset dimension; a keyword extraction module is configured to calculate the cosine similarity between the word vector and the document vector, and extract N words from the target summary text as keywords based on the N results with the highest cosine similarity, to obtain N keywords for each topic in each time period; an association determination module is configured to determine the topic name of each topic in each time period based on the N keywords, calculate the cosine similarity between all topics in the two previous and next time periods based on the topic name, and determine the association relationship between the topics based on the cosine similarity between the topics; an evolution module is configured to determine the evolution path of the target technology based on the association relationship; wherein, S is an integer greater than or equal to 3, M is an integer greater than or equal to 2, and N is an integer greater than or equal to 2; the first preset dimension and the second preset dimension are pre-set, and the first preset dimension is greater than the second preset dimension.

[0010] According to another aspect of the embodiment of the present disclosure, there is also provided an emerging technology evolution analysis device, comprising: a processor; and a memory, connected to the processor, for providing the processor with instructions for processing the following processing steps: obtaining patent data of a target technology for evolution analysis, dividing the patent data into S time periods according to time periods; preprocessing the patent data of each time period, converting the abstract text of each patent into a first text vector of a first preset dimension; performing dimensionality reduction processing on the first text vector to obtain a second text vector of a second preset dimension, and clustering the second text vector to obtain M topics for each time period; for each time period, splicing the abstract texts of all patent data in each topic to obtain a target abstract text for each topic; converting each word in the target abstract text of each topic in each time period into a word vector of the first preset dimension, and clustering the second text vector into a word vector of the first preset dimension for each time period; The method comprises the following steps: converting the entire document of each patent of each topic of each time segment into a document vector of the first preset dimension; calculating the cosine similarity between the word vector and the document vector, and extracting N words as keywords from the target summary text based on the N results with the highest cosine similarity, to obtain N keywords for each topic of each time period; determining the topic name of each topic of each time period based on the N keywords, calculating the cosine similarity between all topics of the two previous and next time periods based on the topic name, and determining the association relationship between the topics based on the cosine similarity between the topics; determining the evolution path of the target technology based on the association relationship; wherein, S is an integer greater than or equal to 3, M is an integer greater than or equal to 2, and N is an integer greater than or equal to 2; the first preset dimension and the second preset dimension are pre-set, and the first preset dimension is greater than the second preset dimension.

[0011] In the disclosed embodiment, the patent data of the target technology is first obtained and divided according to time periods in order to observe the development status of the technology at different time nodes. Then, the patent data is preprocessed and converted into text vectors, and the main technical themes in each time period are extracted through dimensionality reduction processing and cluster analysis. Furthermore, the abstract texts of each topic are spliced ​​to obtain the target abstract text, and the keywords are extracted from the target abstract text by calculating the cosine similarity between the word vector and the document vector, thereby determining the topic name of each topic. Finally, the association relationship is determined based on the cosine similarity between the topics, and the evolution path of the target technology is determined. The present application can efficiently and accurately capture the evolution path of the target technology, thereby improving the accuracy of the technology evolution analysis and solving the technical problem of low accuracy of the technology evolution analysis in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:

[0013] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to embodiment 1 of the present disclosure;

[0014] Figure 2 1 is a flow chart of the emerging technology evolution analysis method according to Example 1 of the present disclosure;

[0015] Figure 3 is a schematic diagram of the dimensionality reduction result according to the first aspect of Example 1 of the present disclosure;

[0016] Figure 4 is a schematic diagram of the clustering result according to the first aspect of Example 1 of the present disclosure;

[0017] Figure 5 is a schematic diagram of a fine-tuned BERT model according to the first aspect of Embodiment 1 of the present disclosure;

[0018] Figure 6 Schematic diagram of the technological evolution in the PBAT field according to the first aspect of Example 1 of the present disclosure;

[0019] Figure 7 is a schematic diagram of an emerging technology evolution analysis device according to Example 2 of the present disclosure;

[0020] Figure 8 2 is a schematic diagram of an emerging technology evolution analysis device according to Example 3 of the present disclosure. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.

[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0023] First, some nouns or terms that appear in the description of the embodiments of the present disclosure are subject to the following explanations:

[0024] Distilbert-base-nli-mean-tokens model: A natural language processing model developed based on the Sentence-Transformers library. It is mainly used to map sentences and paragraphs into a 768-dimensional dense vector space. It is suitable for tasks such as semantic similarity calculation, text clustering, and information retrieval.

[0025] UMAP: short for Uniform Manifold Approximation and Projection, is a nonlinear dimensionality reduction technique used to map high-dimensional data to a low-dimensional space while preserving the local and global structure of the data.

[0026] Kmeans algorithm: A clustering algorithm used to divide a data set into K clusters so that the data points within a cluster are as similar as possible, while the data points between different clusters are as dissimilar as possible.

[0027] nltk: short for Natural Language Toolkit, is a widely used Python library for natural language processing (NLP).

[0028] The NLI dataset is a dataset used for Natural Language Inference (NLI) tasks. Natural Language Inference is a natural language processing (NLP) task that aims to determine the logical relationship between two sentences (a premise and a hypothesis). It is generally divided into three types: entailment, contradiction, and neutral.

[0029] BERT, Bidirectional Encoder Representations from Transformers, is a pre-trained language model based on the Transformer architecture that aims to improve the performance of natural language processing (NLP) tasks through bidirectional contextual information.

[0030] AdamW Optimizer: The AdamW optimizer is an improved version of the Adam optimizer, incorporating the weight decay regularization method to improve model training stability and generalization. The Adam (Adaptive Moment Estimation) optimizer is an adaptive learning rate optimization algorithm that combines the advantages of momentum and RMSProp optimization methods. The core concept is to adjust the learning rate for each parameter individually while leveraging momentum to accelerate convergence and reduce oscillations.

[0031] Example 1

[0032] According to this embodiment, a method embodiment of emerging technology evolution analysis is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0033] The method embodiment provided in this embodiment can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computing device for implementing an emerging technology evolution analysis method. Figure 1 As shown, the computing device may include one or more processors (the processor may include but is not limited to a microprocessor MCU or a programmable logic device FPGA, etc.), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include: a display, a keyboard, and a cursor control device connected to the input / output interface. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0034] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computing device. As described in the embodiments of the present disclosure, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0035] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the emerging technology evolution analysis method in the embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the emerging technology evolution analysis method of the above-mentioned application. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computing device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0036] The transmission device is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by a communications provider of the computing device. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0037] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computing device.

[0038] It should be noted that, in some optional embodiments, the above Figure 1 The computing device shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computing device described above.

[0039] In the above operating environment, according to the first aspect of this embodiment, a method for analyzing the evolution of emerging technologies is provided. Figure 1 The computing device implementation shown in . Figure 2 A schematic diagram showing the process of the method is shown in FIG. Figure 2 As shown, the method includes:

[0040] S202, obtaining patent data of a target technology for evolutionary analysis, and dividing the patent data into S time periods according to time periods;

[0041] S204, pre-processing the patent data of each time period, converting the abstract text of each patent into a first text vector of a first preset dimension;

[0042] S206: Perform dimensionality reduction processing on the first text vector to obtain a second text vector of a second preset dimension, and cluster the second text vector to obtain M topics for each time period;

[0043] S208. For each time period, concatenate the abstract texts of all patent data in each topic to obtain a target abstract text for each topic;

[0044] S210: Convert each word in the target abstract text of each topic in each time period into a word vector of the first preset dimension, and convert the entire document of each patent of each topic in each time period into a document vector of the first preset dimension;

[0045] S212, calculating the cosine similarity between the word vector and the document vector, and extracting N words from the target summary text as keywords based on the N results with the highest cosine similarity, to obtain N keywords for each topic in each time period;

[0046] S214: Determine a topic name for each topic in each time period based on the N keywords, calculate the cosine similarity between all topics in the two time periods based on the topic names, and determine the association relationship between the topics based on the cosine similarity between the topics;

[0047] S216. Determine the evolution path of the target technology based on the association relationship;

[0048] Wherein, S is an integer greater than or equal to 3, M is an integer greater than or equal to 2, and N is an integer greater than or equal to 2;

[0049] The first preset dimension and the second preset dimension are preset, and the first preset dimension is greater than the second preset dimension.

[0050] In the emerging technology evolution analysis method of this embodiment, the patent data of the target technology (such as biotechnology, information technology, etc.) for evolution analysis is first obtained, and the patent data is divided into S time periods according to the time period (for example, S=3, that is, divided into 3 time periods: the embryonic stage, the rapid development stage, and the stable development stage); then the pre-trained SentenceBERT model is used to vectorize the abstract text of the patent data of the target technology, and after dimensionality reduction and clustering, multiple topics are obtained. The abstract text of each topic is spliced ​​to obtain the target abstract text, and the target abstract text is converted into a word vector. The entire document of each patent of each topic in each time period is converted into a document vector. The cosine similarity between the word vector and the document vector is calculated, and the keywords of each topic are determined based on the cosine similarity, and the topic name of each topic is determined based on the keywords of each topic. The cosine similarity between all topics in the two time periods before and after is calculated, and the correlation between the topics is determined based on the cosine similarity between the topics, so as to determine the evolution path of the target technology based on the correlation, thereby efficiently and accurately capturing the evolution path of the target technology, greatly improving the accuracy of technology evolution analysis, and solving the technical problem of low accuracy of technology evolution analysis in the existing technology.

[0051] In the embodiment of the present invention, the target technology is a technology that requires evolutionary analysis, which can be a subdivided technical field or an industry technology. For example, it can be the PBAT subdivided technology in the field of biotechnology (PBAT, Polybutylene Adipate Terephthalate, is a biodegradable thermoplastic polymer, the Chinese name is polybutylene adipate / terephthalate. It is a copolymer of butylene adipate (PBA) and butylene terephthalate (PBT), and has the characteristics of both), the printed circuit board subdivided technology in electronic information technology, etc., or it can be a technical field such as big data technology, artificial intelligence technology, and biomedical technology.

[0052] In the embodiment of the present invention, obtaining patent data of the target technology for evolution analysis refers to obtaining patent application documents in the target technology field through retrieval, screening and other processes.

[0053] In an embodiment of the present invention, dividing the patent data into S time periods by time period may include: dividing the period from the earliest application date in the patent data to the date the patent data was obtained into S time periods based on the patent application date. Alternatively, dividing the period from the emergence of the target technology to a specified date based on the history of technological development may be performed. For example, if the target technology emerged in 1885 and reached a specified date, such as 2020, the period from 1885 to 2020 may be divided into S time periods. As an alternative example, the S time periods may be divided into different stages based on the laws and current status of technological development. For example, the S time periods may be divided into three time periods: the embryonic stage, the rapid development stage, and the stable development stage. Alternatively, the S time periods may be divided into four time periods: the embryonic stage, the rapid development stage, the stable development stage, and the decline stage. As an alternative example, PBAT technology may be divided into three time periods: the embryonic stage (1900-2011), the rapid development stage (2012-2019), and the stable development stage (2020-2023).

[0054] Optionally, in step S204, the pre-processing of the patent data in each time period includes:

[0055] Using the nltk word segmentation tool to perform word segmentation on the abstract text of the patent data;

[0056] Filter out stop words, punctuation, conjunctions, prepositions, personal pronouns, and interjections;

[0057] Among them, filtering is carried out by combining manual filtering and automatic filtering.

[0058] In an embodiment of the present invention, patent data is preprocessed, also known as data cleaning. One of the most important links in emerging technology topic identification and evolution analysis is data cleaning, which is closely related to the accuracy and efficiency of subsequent topic clustering results. Therefore, data preprocessing such as information extraction, word segmentation, stop word removal, and word standardization of patent data is indispensable. The present invention utilizes the nltk word segmentation tool to perform word segmentation on the abstract text of the patent, and filters stop words, punctuation, and other words with no actual meaning such as conjunctions, prepositions, personal pronouns, and interjections. In the process of word filtering, both automatic and manual methods can be used to filter words in order to achieve better word filtering effects.

[0059] Optionally, in step S204, converting the abstract text of each patent into a first text vector of a first preset dimension includes:

[0060] Fine-tune the Distilbert-base-nli-mean-tokens model using the patent data;

[0061] Convert the summary text into a vectorized representation using a fine-tuned Distilbert-base-nli-mean-tokens model;

[0062] The Distilbert-base-nli-mean-tokens model is a SentenceBERT model that is fine-tuned on the NLI dataset using a pre-trained Distilbert-base model.

[0063] Among them, the first preset dimension is 768 dimensions.

[0064] For example, for PBAT technology, the Distilbert-base-nli-mean-tokens model was adopted and fine-tuned for PBAT domain data. The Distilbert-base-nli-mean-tokens model uses a mean pooling strategy to compute sentence representations. As a pre-trained model, its quality has been extensively evaluated for embedding sentences and paragraphs, as well as for embedding search queries.

[0065] SentenceBERT (SBERT), a sentence vector calculation model, generates sentence embedding vectors to calculate similarity between patent documents, effectively addressing the sparse semantic features of patent abstracts. SentenceBERT is an improvement to the BERT language model, primarily addressing the significant time overhead of calculating text semantic similarity with the BERT model. Sentence-BERT extends the pre-trained BERT model, using the Sentence Transformer. By loading a pre-trained model, document embeddings can be created from a set of documents.

[0066] Optionally, fine-tuning the Distilbert-base-nli-mean-tokens model using the patent data includes:

[0067] Randomly extract two pieces of data from all claims, titles, and abstracts of the patent data as positive and negative samples, respectively, and then input them into the Distilbert-base-nli-mean-tokens model to obtain the corresponding vector representations;

[0068] Calculating the cosine similarity between the vectors of the positive sample and the negative sample, and setting a label according to the cosine similarity;

[0069] The labeled patent data is input into the Distilbert-base-nli-mean-tokens model as training data.

[0070] As an optional example, in an embodiment of the present invention, the Distilbert-base-nli-mean-tokens model is fine-tuned as a pre-training model. Two pieces of data are randomly extracted from all the claims, titles, and abstracts of the original data as a positive sample and a negative sample, respectively, and then input into the model to obtain the corresponding vector representation. Next, the cosine similarity of the positive sample and the negative sample vectors is calculated, and the labels are set according to the cosine similarity. Finally, the patent data with labels is input as training data into the Distilbert-base-nli-mean-tokens model for iteration. For example, it is set as the AdamW optimizer, BCEWithLogitsLoss as the loss function, and the number of iterations is 30 to complete the fine-tuning of the pre-training model.

[0071] Among them, BCEWithLogitsLoss is a loss function for binary classification problems. It combines the Sigmoid activation function and the binary cross entropy loss to provide a more efficient and numerically stable way to calculate the loss.

[0072] Assuming that the input is the model's output Logits x and the true label y, where y∈{0,1}, the calculation formula of BCEWithLogitsLoss is:

[0073]

[0074] in, is the Sigmoid function. By combining the Sigmoid function with the loss calculation, numerical stability is improved.

[0075] Among them, Logits is the output of the last layer in the neural network, which is a numerical vector. P is the number of data, and i is the data number, ranging from 1 to P. i The i-th output, y i is the true label of the i-th image, and x is x i The set of y is y i A collection of .

[0076] For example, when the target technology is PBAT technology, the results of converting the abstract text of each patent into the first text vector of the first preset dimension are shown in Table 1 below:

[0077] Table 1: Example of vectorized representation of PBAT technical summary text

[0078]

[0079] Optionally, in step S206, performing dimensionality reduction processing on the first text vector includes:

[0080] A UMAP method is used to perform a dimensionality reduction operation on the first text vector, and after the dimensionality reduction operation, a second text vector of a second preset dimension is obtained.

[0081] The second preset dimension may be 3.

[0082] Since the dimension of the first text vector is too high, the present invention uses the UMAP algorithm to perform data dimensionality reduction on the first text vectorization matrix. The UMAP algorithm is a very effective and scalable dimensionality reduction algorithm. While retaining more global structural information of the summary text, the algorithm maps the high-dimensional probability distribution to a low-dimensional space to facilitate keyword extraction and topic clustering.

[0083] The UMAP algorithm mainly uses local manifold approximation and local fuzzy simplex set representation to construct a topological representation of high-dimensional data, which can reduce computational complexity and memory usage while preserving the characteristics of the original data to the greatest extent. Scatter plots can well represent the relationship and structure between high-dimensional embedding vectors, such as Figure 3 As shown in the figure, each point represents an embedding vector (i.e., a summary text), and the distance between points reflects their similarity in the original high-dimensional space. For example, if two points are very close in three-dimensional space, then their embedding vector representations in the high-dimensional space may be very similar. As an optional example, the present invention uses the UMAP algorithm to reduce the 768-dimensional vector processed by SentenceBERT into a three-dimensional vector, that is, the second text vector is a three-dimensional vector.

[0084] Optionally, in step S206, clustering is performed on the second text vector to obtain M topics for each time period, including:

[0085] Clustering the second text vector using the Kmeans algorithm to generate M topics for each time period;

[0086] The clustering of the second text vector using the Kmeans algorithm includes:

[0087] Randomly select K data points from the second text vector as K cluster centers;

[0088] The following three steps are iterated until the objective function converges to a square error value less than the preset error value:

[0089] Calculate the distances between the remaining data points and the centers of the K clusters, and divide the remaining data points into the clusters with the closest distances between the data points and the centers of the K clusters;

[0090] Calculate the mean of all data points in K clusters and use it as the new cluster center;

[0091] Calculate the convergence square error value of the objective function;

[0092] Wherein, the objective function is defined as:

[0093]

[0094] Where i is the cluster number, J is the objective function, K is the number of clusters, S i is the set of data points in the i-th cluster, x is a data point, μ i is the center of the ith cluster, and μ i satisfy:

[0095]

[0096] Here, K is the number of clusters.

[0097] That is, in the above steps, the second text vector is the representation of the patent data after dimensionality reduction. They are used as the input of the Kmeans algorithm. The following are the detailed steps of the clustering process:

[0098] 1) Initialize cluster centers: Randomly select K data points from the second text vector as the initial cluster centers; where K corresponds to the number of topics M expected in each time period;

[0099] 2) The algorithm enters the iterative process until the convergence condition is met (that is, the square error of the objective function convergence value is less than the preset error value):

[0100] Calculate distance and divide clusters: For each second text vector (i.e. each data point), calculate its distance from the K cluster centers. Then, divide the data point into the cluster corresponding to the cluster center with the closest distance to it;

[0101] Update cluster centers: For each cluster, calculate the mean of all data points in the cluster and use this mean as the new cluster center. This step ensures that the cluster center can represent the center position of the data points in the cluster.

[0102] Calculate the objective function: The objective function is usually defined as the sum of the squares of the distances from all data points to the center of their cluster. After each iteration, the objective function value is recalculated to evaluate the quality of the current clustering.

[0103] 3) Convergence condition: The iterative process continues until the change in the objective function value is less than the preset error value, which indicates that the cluster division has become relatively stable and the algorithm has converged;

[0104] 4) Obtain M topics for each time period: When the Kmeans algorithm converges, the data points in each time period are divided into K (i.e., M) clusters, and each cluster represents a topic.

[0105] Therefore, by clustering the second text vector using the Kmeans algorithm, M topics for each time period can be obtained. These topics reflect the technical content and trends of the patent data in different time periods.

[0106] Wherein, the objective function is defined as:

[0107]

[0108] Where i is the cluster number, J is the objective function, K is the number of clusters, S i is the set of data points in the i-th cluster, x is the data point, u i is the center of the ith cluster, and μ i satisfy:

[0109]

[0110] Here, K is the number of clusters.

[0111] The Kmeans algorithm requires that the number of clusters to be generated must be determined before execution. For example, based on the actual situation in the PBAT field, each stage is clustered into 6 clusters. The results of Kmeans clustering of some data are as follows: Figure 4 As shown, for example, patents disclosed as EP2631060A1, EP2712889A1, EP2803753A1, EP2804908A1, etc. are in the X-shaped cluster on the left.

[0112] Optionally, in step S208, for each time period, the abstract texts of all patent data in each topic are concatenated to obtain a target abstract text for each topic, including:

[0113] The abstracts of all patents in each topic are extracted and concatenated to obtain the target abstract text of each topic.

[0114] Optionally, in step S210, converting each word in the target abstract text of each topic in each time period into a word vector of the first preset dimension, and converting the entire document of each patent of each topic in each time period into a document vector of the first preset dimension includes:

[0115] The first step is data preprocessing. The input target summary text is preprocessed by removing stop words and using predicate logic to remove synonyms. The second step is vectorization. Each word in the target summary text is converted into a word vector of the first preset dimension using a fine-tuned BERT model. The entire document is converted into a document vector of the first preset dimension using a pretrained Distilbert-base-nli-mean-tokens model. The first preset dimension is 768.

[0116] Optionally, in step S212, calculating the cosine similarity between the word vector and the document vector includes:

[0117] The cosine similarity is calculated according to the following formula:

[0118]

[0119] where x 1k Indicates the kth dimension of the first calculation, x 2k represents the k-th dimension of the second computational quantity, where k is in the range [1,768], n is the number of dimensions and n=768, and cos(θ) is in the range [-1,1];

[0120] cos(θ) represents the cosine similarity between the first calculation amount and the second calculation amount.

[0121] For example, x 1k Represents word vector, x 2k Represents the document vector. The cosine similarity between the word vector and the document vector can be calculated using the above formula.

[0122] In this embodiment, combined with Figure 5 As shown, the BERT model is fine-tuned in the following way:

[0123] (1) Obtain a historical data set, where each piece of historical data in the historical data set includes all topics of a certain emerging technology in each time period, the subject name of each topic, the patent documents of all patents under each topic, and the subject abstract text corresponding to each topic;

[0124] (2) Perform word segmentation on the topic abstract text corresponding to each topic, extract all candidate words A1 to Am (such as noun phrases and verb phrases), and record the position of each candidate word in the sentence (i.e., the start and end token index);

[0125] (3) For each candidate word, the sentence in which it is located is intercepted as the context sentence, and an importance score (such as a value of 0-1) is assigned to each candidate word based on manual annotation or pseudo-labeling; wherein the importance score is used to indicate the contribution value of the candidate word to the topic name;

[0126] (4) The topic summary text data and the patent documents under the topic are used as a set of training samples. The topic summary text data is the feature information of all candidate words, which includes the context sentence, the candidate word position and the corresponding label (i.e., the importance score);

[0127] (5) Input the context sentences and candidate word positions in the topic abstract text data into the BERT model to be trained, and input the patent document under the topic into the trained Distilbert-base-nli-mean-tokens model. The BERT model outputs the word vectors X1~Xm (768 dimensions) of all candidate words, and the Distilbert-base-nli-mean-tokens model outputs the document vector (768 dimensions) of the patent document.

[0128] (6) Calculate the cosine similarity between the word vector and the document vector through the cosine similarity calculation module;

[0129] (7) Freeze all parameters of the Distilbert-base-nli-mean-tokens model and backpropagate the model parameters of the BERT model according to the following loss function:

[0130]

[0131] Among them, m is the number of candidate words, Si is the cosine similarity between the i-th candidate word and the document vector, Score the importance of the i-th candidate word;

[0132] (8) Perform iterative training according to the above steps until the loss function converges or the preset number of iterations is reached.

[0133] In this way, the word vectors output by the fine-tuned BERT model can reflect their importance in determining the subsequent topic name, thereby ensuring the accuracy of the keywords selected from the target summary text.

[0134] Optionally, in step S212, based on the N results with the highest cosine similarity, N words are extracted from the target summary text as keywords to obtain N keywords for each topic in each time period, where N is an integer greater than or equal to 2. Optionally, N is equal to 15, that is, the 15 words with the highest cosine similarity between the word vector and the document vector are obtained as keywords.

[0135] For example, in the case of PBAT technology, there are 10 keywords per topic, with multiple stages: the embryonic stage (1900-2011), the rapid development stage (2012-2019), and the stable development stage (2020-2023). Each topic is numbered, and each stage contains 6 topics. The keywords for each stage are shown in Table 2 below. The keywords for each topic are shown in the Keywords column, and the Weight column is the weight value of the keyword. The higher the weight value, the more important the keyword is in the abstract.

[0136] Table 2: Examples of keywords for each topic at each stage in the PBAT technology field

[0137]

[0138]

[0139] Optionally, in step S214, determining the theme name of each theme in each time period based on the N keywords includes:

[0140] According to the N keywords in each topic, the N keywords and related weights are input into the Large Language Model (LLM) to obtain the topic name of each topic, and the final topic name is obtained based on the keywords after expert review.

[0141] For example, taking PBAT technology as an example, the determined subject names are shown in Table 3 below.

[0142] Table 3: Theme name of each topic in each stage of PBAT technology field

[0143]

[0144]

[0145] Optionally, in step S214, calculating the cosine similarity between all topics in two time periods according to the topic names, and determining the association relationship between the topics according to the cosine similarity between the topics may include:

[0146] Starting from one time period, the cosine similarities between all topics in two adjacent time periods are calculated backwards. Then all cosine similarities are checked. If the cosine similarity is greater than a preset similarity threshold, it is determined that there is a correlation between the two topics.

[0147] As a preferred example, the similarity threshold is 0.4.

[0148] For example, if there are four time periods, namely the embryonic stage S0, the rapid development stage S1, the stable stage S2, and the decline stage S3. The embryonic stage S0 has N0 topics, the rapid development stage S1 has N1 topics, the stable stage S2 has N2 topics, and the decline stage S3 has N3 topics. Then perform the following calculations:

[0149] 1. Calculate the cosine similarity between each topic in N0 and each topic in N1, and obtain a total of N0*N1 cosine similarity values. If a cosine similarity is greater than the preset similarity threshold, it is determined that there is a correlation between the two topics for which the cosine similarity is calculated.

[0150] 2. Calculate the cosine similarity between each topic in N1 and each topic in N2, and obtain a total of N1*N2 cosine similarity values. If a cosine similarity is greater than the preset similarity threshold, it is determined that there is a correlation between the two topics for which the cosine similarity is calculated.

[0151] 3. Calculate the cosine similarity between each topic in the N2 topics and each topic in the N3 topics, and obtain a total of N2*N3 cosine similarity values. If a cosine similarity is greater than the preset similarity threshold, it is determined that there is a correlation between the two topics for which the cosine similarity is calculated.

[0152] After the above steps, we can obtain the correlation between each topic between two adjacent time periods, and use graphical methods to visualize the evolution of emerging technologies. The existence of a correlation between two topics means that there is an evolutionary relationship between the two topics.

[0153] Taking PBAT technology as an example, the evolution trend of PBAT technology is plotted using the Sankey diagram. Figure 6 As shown (only technical topics with evolutionary relationships are shown). Among them, the first stage (1900-2011) was mainly the technical manufacturing and material production of polymers, and the main content of the second stage (2012-2019) was the improvement and innovation of polymers, which had significant developments compared to the first stage, in order to obtain high-performance polymer products, and some of them were used in biodegradable film products. The third stage (2020-2023) is the further development of polymers and other materials, mainly for the improvement and application of polymers. Compared with the changes between the first and second stages, the changes between the second and third stages are relatively small, and the development speed is slower, but PBAT technology is still a hot direction.

[0154] Depend on Figure 6 The following information can also be interpreted:

[0155] The majority of the technical topics with evolving relationships are related to biodegradability and polymers, falling within the realm of polymers and their manufacturing processes, including 14 topics such as S0, M5, and L2. Five topics, such as S3, M1, and L0, primarily concern films and finished products. This suggests that these two areas are the most important technical areas within PBAT. Polymers and their manufacturing processes have been a key focus of PBAT research since 2000 and remained a popular research topic as recently as 2023. Films and their finished products have also appeared in all three phases, making them a perennial research area within the field.

[0156] The evolutionary direction between these topics in adjacent time slices is mainly inheritance and fusion, and the evolutionary intensity (i.e., the similarity between technical topics) is also relatively high, for example: S1→M0, M2→L2. This shows that the field of biodegradable polymer-related technology PBAT continues to receive widespread attention in the industry, with strong vitality and high scientific research and application value. Among them, M2 is Biodegradable Polymer-Based Polyester Composition Invention. Since PBAT is a type of polymer, this topic has influenced many subsequent topics (L0, L2, L3, L4, L5).

[0157] Furthermore, according to a second aspect of this embodiment, a storage medium is provided, wherein the storage medium includes a stored program, wherein when the program is run, a processor executes any one of the above methods.

[0158] According to the method of this embodiment, first, the patent data of the target technology is obtained and divided according to time periods in order to observe the development status of the technology at different time nodes. Then, the patent data is preprocessed and converted into text vectors, and the main technical themes in each time period are extracted through dimensionality reduction processing and cluster analysis. Furthermore, the abstract texts of each topic are spliced ​​to obtain the target abstract text, and the keywords are extracted from the target abstract text by calculating the cosine similarity between the word vector and the document vector, thereby determining the topic name of each topic. Finally, the association relationship is determined based on the cosine similarity between the topics, and the evolution path of the target technology is determined. The present application can efficiently and accurately capture the evolution path of the target technology, thereby improving the accuracy of the technology evolution analysis and solving the technical problem of low accuracy of the technology evolution analysis existing in the prior art.

[0159] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0161] Example 2

[0162] Figure 7 FIG2 shows an emerging technology evolution analysis device 700 according to this embodiment, which corresponds to the method according to the first aspect of embodiment 1. Figure 7 As shown, the apparatus 700 includes:

[0163] The patent data collection module 701 is configured to obtain patent data of a target technology for evolution analysis and divide the patent data into S time periods according to time periods;

[0164] Abstract vectorization module 702 is configured to pre-process the patent data of each time period and convert the abstract text of each patent into a first text vector of a first preset dimension;

[0165] A dimensionality reduction module 703 is configured to perform dimensionality reduction processing on the first text vector to obtain a second text vector of a second preset dimension, and cluster the second text vector to obtain M topics for each time period;

[0166] The splicing module 704 is configured to splice the abstract texts of all patent data in each topic for each time period to obtain a target abstract text for each topic;

[0167] A document vectorization module 705 is configured to convert each word in the target abstract text of each topic in each time period into a word vector of the first preset dimension, and convert the entire document of each patent of each topic in each time period into a document vector of the first preset dimension;

[0168] A keyword extraction module 706 is configured to calculate the cosine similarity between the word vector and the document vector, and extract N words from the target summary text as keywords based on the N results with the highest cosine similarity, to obtain N keywords for each topic in each time period;

[0169] The association determination module 707 is configured to determine a topic name for each topic in each time period based on the N keywords, calculate the cosine similarity between all topics in the two time periods based on the topic names, and determine the association relationship between the topics based on the cosine similarity between the topics;

[0170] An evolution module 708 is configured to determine an evolution path of the target technology according to the association relationship;

[0171] Wherein, S is an integer greater than or equal to 3, M is an integer greater than or equal to 2, and N is an integer greater than or equal to 2;

[0172] The first preset dimension and the second preset dimension are preset, and the first preset dimension is greater than the second preset dimension.

[0173] Optionally, the summary vectorization module 702 is further configured to:

[0174] Using the nltk word segmentation tool to perform word segmentation on the abstract text of the patent data;

[0175] Filter out stop words, punctuation, conjunctions, prepositions, personal pronouns, and interjections;

[0176] Among them, filtering is carried out by combining manual filtering and automatic filtering.

[0177] Optionally, the summary vectorization module 702 is further configured to:

[0178] Fine-tune the Distilbert-base-nli-mean-tokens model using the patent data;

[0179] Convert the summary text into a vectorized representation using a fine-tuned Distilbert-base-nli-mean-tokens model;

[0180] The Distilbert-base-nli-mean-tokens model is a SentenceBERT model that is fine-tuned on the NLI dataset using a pre-trained Distilbert-base model.

[0181] Optionally, fine-tuning the Distilbert-base-nli-mean-tokens model using the patent data includes:

[0182] Randomly extract two pieces of data from all claims, titles, and abstracts of the patent data as positive and negative samples, respectively, and then input them into the Distilbert-base-nli-mean-tokens model to obtain the corresponding vector representations;

[0183] Calculating the cosine similarity between the vectors of the positive sample and the negative sample, and setting a label according to the cosine similarity;

[0184] The labeled patent data is input into the Distilbert-base-nli-mean-tokens model as training data.

[0185] Optionally, the dimensionality reduction module 703 is further configured to:

[0186] The UMAP method is used to perform a dimensionality reduction operation on the first text vector.

[0187] Optionally, the dimensionality reduction module 703 is further configured to:

[0188] Clustering the second text vector using the Kmeans algorithm to generate M topics for each time period;

[0189] The clustering of the second text vector using the Kmeans algorithm includes:

[0190] Randomly select K data points from the second text vector as K cluster centers;

[0191] The following three steps are iterated until the objective function converges to a square error value less than the preset error value:

[0192] Calculate the distances between the remaining data points and the centers of the K clusters, and divide the remaining data points into the clusters with the closest distances between the data points and the centers of the K clusters;

[0193] Calculate the mean of all data points in K clusters and use it as the new cluster center;

[0194] Calculate the convergence square error value of the objective function;

[0195] Wherein, the objective function is defined as:

[0196]

[0197] Where i is the cluster number, J is the objective function, K is the number of clusters, S i is the set of data points in the i-th cluster, x is a data point, μ i is the center of the ith cluster, and μ i satisfy:

[0198]

[0199] Here, K is the number of clusters.

[0200] Optionally, the keyword extraction module 706 is further configured to calculate the cosine similarity according to the following formula:

[0201]

[0202] where x 1k Indicates the kth dimension of the first calculation, x 2k represents the k-th dimension of the second computational quantity, where k is in the range [1,768], n is the number of dimensions and n=768, and cos(θ) is in the range [-1,1];

[0203] cos(θ) represents the cosine similarity between the first calculation amount and the second calculation amount.

[0204] Optionally, the association determination module 707 is further configured to:

[0205] If the cosine similarity between two topics is greater than a preset similarity threshold, it is determined that there is an association relationship between the two topics.

[0206] It should be noted that the emerging technology evolution analysis device of Example 2 and the emerging technology evolution analysis method of Example 1 belong to the same inventive concept, solve the same technical problems, and achieve the same technical effects, and the similarities will not be repeated here.

[0207] According to an embodiment of the present device, first, the patent data of the target technology is obtained and divided according to time periods in order to observe the development status of the technology at different time nodes. Then, the patent data is preprocessed and converted into text vectors, and the main technical themes in each time period are extracted through dimensionality reduction processing and cluster analysis. Furthermore, the abstract texts of each topic are spliced ​​to obtain the target abstract text, and the keywords are extracted from the target abstract text by calculating the cosine similarity between the word vector and the document vector, thereby determining the topic name of each topic. Finally, the association relationship is determined based on the cosine similarity between the topics, and the evolution path of the target technology is determined. The present application can efficiently and accurately capture the evolution path of the target technology, thereby improving the accuracy of the technology evolution analysis and solving the technical problem of low accuracy of the technology evolution analysis in the prior art.

[0208] Example 3

[0209] Figure 8 The emerging technology evolution analysis device according to this embodiment is shown, which corresponds to the method according to the first aspect of embodiment 1. Figure 8 As shown, the device includes:

[0210] Processor 810; and

[0211] The memory 820 is connected to the processor 810 and is configured to provide the processor 810 with instructions for processing the following steps:

[0212] Obtaining patent data of a target technology for evolutionary analysis, and dividing the patent data into S time periods according to time periods;

[0213] Preprocessing the patent data of each time period, converting the abstract text of each patent into a first text vector of a first preset dimension;

[0214] Performing dimensionality reduction processing on the first text vector to obtain a second text vector of a second preset dimension, and clustering the second text vector to obtain M topics for each time period;

[0215] For each time period, the abstract texts of all patent data in each topic are spliced ​​together to obtain the target abstract text of each topic;

[0216] Convert each word in the target abstract text of each topic in each time period into a word vector of the first preset dimension, and convert the entire document of each patent of each topic in each time period into a document vector of the first preset dimension;

[0217] Calculate the cosine similarity between the word vector and the document vector, and extract N words from the target summary text as keywords based on the N results with the highest cosine similarity, to obtain N keywords for each topic in each time period;

[0218] Determine the topic name of each topic in each time period based on the N keywords, calculate the cosine similarity between all topics in the two time periods based on the topic names, and determine the association relationship between the topics based on the cosine similarity between the topics;

[0219] Determining an evolution path of the target technology based on the association relationship;

[0220] Wherein, S is an integer greater than or equal to 3, M is an integer greater than or equal to 2, and N is an integer greater than or equal to 2;

[0221] The first preset dimension and the second preset dimension are preset, and the first preset dimension is greater than the second preset dimension.

[0222] Optionally, the preprocessing of the patent data for each time period includes:

[0223] Using the nltk word segmentation tool to perform word segmentation on the abstract text of the patent data;

[0224] Filter out stop words, punctuation, conjunctions, prepositions, personal pronouns, and interjections;

[0225] Among them, filtering is carried out by combining manual filtering and automatic filtering.

[0226] Optionally, converting the abstract text of each patent into a first text vector of a first preset dimension includes:

[0227] Fine-tune the Distilbert-base-nli-mean-tokens model using the patent data;

[0228] Convert the summary text into a vectorized representation using a fine-tuned Distilbert-base-nli-mean-tokens model;

[0229] The Distilbert-base-nli-mean-tokens model is a SentenceBERT model that is fine-tuned on the NLI dataset using a pre-trained Distilbert-base model.

[0230] Optionally, fine-tuning the Distilbert-base-nli-mean-tokens model using the patent data includes:

[0231] Randomly extract two pieces of data from all claims, titles, and abstracts of the patent data as positive and negative samples, respectively, and then input them into the Distilbert-base-nli-mean-tokens model to obtain the corresponding vector representations;

[0232] Calculating the cosine similarity between the vectors of the positive sample and the negative sample, and setting a label according to the cosine similarity;

[0233] The labeled patent data is input into the Distilbert-base-nli-mean-tokens model as training data.

[0234] Optionally, performing dimensionality reduction processing on the first text vector includes:

[0235] The UMAP method is used to perform a dimensionality reduction operation on the first text vector.

[0236] Optionally, clustering the second text vector to obtain M topics for each time period includes:

[0237] Clustering the second text vector using the Kmeans algorithm to generate M topics for each time period;

[0238] The clustering of the second text vector using the Kmeans algorithm includes:

[0239] Randomly select K data points from the second text vector as K cluster centers;

[0240] The following three steps are iterated until the objective function converges to a square error value less than the preset error value:

[0241] Calculate the distances between the remaining data points and the centers of the K clusters, and divide the remaining data points into the clusters with the closest distances between the data points and the centers of the K clusters;

[0242] Calculate the mean of all data points in K clusters and use it as the new cluster center;

[0243] Calculate the convergence square error value of the objective function;

[0244] Wherein, the objective function is defined as:

[0245]

[0246] Where i is the cluster number, J is the objective function, K is the number of clusters, S i is the set of data points in the i-th cluster, x is a data point, μ iis the center of the ith cluster, and μ i satisfy:

[0247]

[0248] Here, K is the number of clusters.

[0249] Optionally, the cosine similarity is calculated according to the following formula:

[0250]

[0251] where x 1k Indicates the kth dimension of the first calculation, x 2k represents the k-th dimension of the second computational quantity, where k is in the range [1,768], n is the number of dimensions and n=768, and cos(θ) is in the range [-1,1];

[0252] cos(θ) represents the cosine similarity between the first calculation amount and the second calculation amount.

[0253] Optionally, determining the association relationship between the topics according to the cosine similarity between the topics includes:

[0254] If the cosine similarity between two topics is greater than a preset similarity threshold, it is determined that there is an association relationship between the two topics.

[0255] It should be noted that the emerging technology evolution analysis device of Example 3 and the emerging technology evolution analysis method of Example 1 belong to the same inventive concept, solve the same technical problems, and achieve the same technical effects, and the similarities will not be repeated here.

[0256] According to an embodiment of the present device, first, the patent data of the target technology is obtained and divided according to time periods in order to observe the development status of the technology at different time nodes. Then, the patent data is preprocessed and converted into text vectors, and the main technical themes in each time period are extracted through dimensionality reduction processing and cluster analysis. Furthermore, the abstract texts of each topic are spliced ​​to obtain the target abstract text, and the keywords are extracted from the target abstract text by calculating the cosine similarity between the word vector and the document vector, thereby determining the topic name of each topic. Finally, the association relationship is determined based on the cosine similarity between the topics, and the evolution path of the target technology is determined. The present application can efficiently and accurately capture the evolution path of the target technology, thereby improving the accuracy of the technology evolution analysis and solving the technical problem of low accuracy of the technology evolution analysis in the prior art.

[0257] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0258] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0259] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0260] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0261] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0262] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.

[0263] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for analyzing the evolution of emerging technologies, characterized in that: include: Obtaining patent data of a target technology for evolutionary analysis, and dividing the patent data into S time periods according to time periods; Preprocessing the patent data of each time period, converting the abstract text of each patent into a first text vector of a first preset dimension; Performing dimensionality reduction processing on the first text vector to obtain a second text vector of a second preset dimension, and clustering the second text vector to obtain M topics for each time period; For each time period, the abstract texts of all patent data in each topic are spliced ​​together to obtain the target abstract text of each topic; Convert each word in the target abstract text of each topic in each time period into a word vector of the first preset dimension, and convert the entire document of each patent of each topic in each time period into a document vector of the first preset dimension; Calculate the cosine similarity between the word vector and the document vector, and extract N words from the target summary text as keywords based on the N results with the highest cosine similarity, to obtain N keywords for each topic in each time period; Determine the topic name of each topic in each time period based on the N keywords, calculate the cosine similarity between all topics in the two time periods based on the topic names, and determine the association relationship between the topics based on the cosine similarity between the topics; Determining an evolution path of the target technology based on the association relationship; Wherein, S is an integer greater than or equal to 3, M is an integer greater than or equal to 2, and N is an integer greater than or equal to 2; The first preset dimension and the second preset dimension are preset, and the first preset dimension is greater than the second preset dimension.

2. The method according to claim 1, characterized in that The preprocessing of the patent data for each time period includes: Using the nltk word segmentation tool to perform word segmentation on the abstract text of the patent data; Filter out stop words, punctuation, conjunctions, prepositions, personal pronouns, and interjections; Among them, filtering is carried out by combining manual filtering and automatic filtering.

3. The method according to claim 1, characterized in that The step of converting the abstract text of each patent into a first text vector of a first preset dimension includes: Fine-tune the Distilbert-base-nli-mean-tokens model using the patent data; Convert the summary text into a vectorized representation using a fine-tuned Distilbert-base-nli-mean-tokens model; The Distilbert-base-nli-mean-tokens model is a SentenceBERT model that is fine-tuned on the NLI dataset using a pre-trained Distilbert-base model.

4. The method according to claim 3, characterized in that The fine-tuning of the Distilbert-base-nli-mean-tokens model using the patent data includes: Randomly extract two pieces of data from all claims, titles, and abstracts of the patent data as positive and negative samples, respectively, and then input them into the Distilbert-base-nli-mean-tokens model to obtain the corresponding vector representations; Calculating the cosine similarity between the vectors of the positive sample and the negative sample, and setting a label according to the cosine similarity; The labeled patent data is input into the Distilbert-base-nli-mean-tokens model as training data.

5. The method according to claim 1, characterized in that The clustering of the second text vector to obtain M topics for each time period includes: Clustering the second text vector using the Kmeans algorithm to generate M topics for each time period; The clustering of the second text vector using the Kmeans algorithm includes: Randomly select K data points from the second text vector as K cluster centers; The following three steps are iterated until the objective function converges to a square error value less than the preset error value: Calculate the distances between the remaining data points and the centers of the K clusters, and divide the remaining data points into the clusters with the closest distances between the data points and the centers of the K clusters; Calculate the mean of all data points in K clusters and use it as the new cluster center; Calculate the convergence square error value of the objective function; Wherein, the objective function is defined as: Where i is the cluster number, J is the objective function, K is the number of clusters, S i is the set of data points in the i-th cluster, x is a data point, μ i is the center of the ith cluster, and μ i satisfy: Here, K is the number of clusters.

6. The method according to claim 1, characterized in that The cosine similarity is calculated according to the following formula: where x 1k Indicates the kth dimension of the first calculation, x 2k represents the k-th dimension of the second computational quantity, where k is in the range [1,768], n is the number of dimensions and n=768, and cos(θ) is in the range [-1,1]; cos(θ) represents the cosine similarity between the first calculation amount and the second calculation amount.

7. The method according to claim 1, characterized in that Determining the association relationship between the topics according to the cosine similarity between the topics includes: If the cosine similarity between two topics is greater than a preset similarity threshold, it is determined that there is an association relationship between the two topics.

8. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is executed, the processor executes the method according to any one of claims 1 to 7.

9. An emerging technology evolution analysis device, characterized in that: include: a patent data acquisition module configured to acquire patent data of a target technology for evolution analysis and divide the patent data into S time periods according to time periods; An abstract vectorization module is configured to pre-process the patent data of each time period and convert the abstract text of each patent into a first text vector of a first preset dimension; a dimensionality reduction module configured to perform dimensionality reduction processing on the first text vector to obtain a second text vector of a second preset dimension, and cluster the second text vector to obtain M topics for each time period; A splicing module is configured to splice the abstract texts of all patent data in each topic for each time period to obtain a target abstract text for each topic; a document vectorization module configured to convert each word in the target abstract text of each topic in each time period into a word vector of the first preset dimension, and convert the entire document of each patent of each topic in each time period into a document vector of the first preset dimension; A keyword extraction module is configured to calculate the cosine similarity between the word vector and the document vector, and extract N words from the target summary text as keywords based on the N results with the highest cosine similarity, to obtain N keywords for each topic in each time period; an association determination module configured to determine a topic name of each topic in each time period based on the N keywords, calculate the cosine similarity between all topics in the two previous and next time periods based on the topic names, and determine the association relationship between the topics based on the cosine similarity between the topics; An evolution module, configured to determine an evolution path of the target technology according to the association relationship; Wherein, S is an integer greater than or equal to 3, M is an integer greater than or equal to 2, and N is an integer greater than or equal to 2; The first preset dimension and the second preset dimension are preset, and the first preset dimension is greater than the second preset dimension.

10. An emerging technology evolution analysis device, characterized in that: include: processor; as well as A memory, connected to the processor, configured to provide the processor with instructions for processing the following processing steps: Obtaining patent data of a target technology for evolutionary analysis, and dividing the patent data into S time periods according to time periods; Preprocessing the patent data of each time period, converting the abstract text of each patent into a first text vector of a first preset dimension; Performing dimensionality reduction processing on the first text vector to obtain a second text vector of a second preset dimension, and clustering the second text vector to obtain M topics for each time period; For each time period, the abstract texts of all patent data in each topic are spliced ​​together to obtain the target abstract text of each topic; Convert each word in the target abstract text of each topic in each time period into a word vector of the first preset dimension, and convert the entire document of each patent of each topic in each time period into a document vector of the first preset dimension; Calculate the cosine similarity between the word vector and the document vector, and extract N words from the target summary text as keywords based on the N results with the highest cosine similarity, to obtain N keywords for each topic in each time period; Determine the topic name of each topic in each time period based on the N keywords, calculate the cosine similarity between all topics in the two time periods based on the topic names, and determine the association relationship between the topics based on the cosine similarity between the topics; Determining an evolution path of the target technology based on the association relationship; Wherein, S is an integer greater than or equal to 3, M is an integer greater than or equal to 2, and N is an integer greater than or equal to 2; The first preset dimension and the second preset dimension are preset, and the first preset dimension is greater than the second preset dimension.

Citation Information

Patent Citations

  • An academic subject life cycle analysis method based on abstract clustering of sci-tech documents

    CN109344248A

  • Dynamic knowledge hotspot evolution and trend analysis method

    CN111694930A

  • BERT-based text topic extraction and time-space evolution analysis method and system

    CN118734826A

  • Quantitative scientific research project approval screening method based on patent situation analysis

    CN119149736A

  • Scientific research topic recommendation method and device based on similarity calculation

    CN119513293A