Visualization analysis method and system of scientific and technological development trends based on word vector network
By constructing a word vector network to cluster and visualize science and technology theme communities, the quality problem of keyword selection in existing technologies is solved, and an intuitive presentation and fine-grained analysis of science and technology development trends are achieved, meeting practical application needs.
Patent Information
- Application Number
- CN202510082209.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The quality of keyword selection in existing technologies is difficult to control, and corpus information is not fully mined and utilized. It is impossible to intuitively and fully reflect the trends and processes of topic changes over time. Existing visualization analysis methods cannot meet actual application needs.
By constructing a word vector network, clustering science and technology theme communities, extracting characteristic keywords, calculating community similarity, and visualizing them in a Cartesian coordinate system, we can draw an evolutionary map of science and technology development trends.
It achieves intuitive visualization of the evolutionary relationship between science and technology theme communities, reflects the changing trends and processes of themes over time, meets practical application needs, and provides fine-grained analysis of science and technology development trends.
Smart Images

Figure CN119988611B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of science and technology development trend analysis, and in particular to a method and system for visual analysis of science and technology development trends based on a word vector network. Background Art
[0002] With the gradual advancement of my country's strategy to become a national power in science and technology, and the new requirements for developing high-quality, innovative productivity, how to accurately analyze the dynamics of scientific and technological research and assess the future trajectory of scientific and technological development has become a hot topic. Scientific topic evolution analysis, as an analytical tool, can reveal the interconnected structure and evolutionary context of knowledge in the field of science and technology. It is an important tool for analyzing the metabolic process of scientific research topics in a field and thus understanding the development trends of scientific research in that field. Simultaneously, with the rise of the open science movement, a variety of well-structured, high-quality, and comprehensive open academic graphs and scientific literature databases have provided solid data support for scientific topic evolution analysis in a field. By comprehensively, rapidly, and accurately tracking the temporal development of scientific topics represented by terms within this abundant and comprehensive scientific corpus, and then analyzing the changes and updates in the knowledge structure and the systematic development of the scientific knowledge system within the field, this will help reveal the dynamic process of scientific knowledge production and dissemination. It is also of great significance for frontline researchers and science and technology decision-makers to timely grasp the trends of scientific and technological innovation and cutting-edge research hotspots, and to optimize resource allocation and deployment.
[0003] Scientific literature typically focuses on specific research topics, which can be represented by a combination of fields distributed across the document, such as the title, author keywords, and abstract (or summary). By observing the interactive combinations of characteristic keywords within a specific scientific field and mining their semantic associations, we can effectively summarize and map a specific scientific topic within that field.
[0004] Currently, scientific knowledge analysis is commonly performed using co-word networks. However, this approach suffers from difficulties in controlling the quality of keyword selection and insufficiently exploiting corpus information. Furthermore, existing visualization methods based on scientific knowledge networks typically use static charts, such as cluster diagrams, which fail to intuitively and fully reflect the trends and processes of topic evolution over time. Consequently, existing methods fail to meet practical application needs. Summary of the Invention
[0005] The present invention provides a method and system for visual analysis of scientific and technological development trends based on a word vector network, which solves the technical problems in the existing technology, such as the difficulty in controlling the quality of keyword selection, insufficient mining and utilization of corpus information, and the inability to intuitively and fully reflect the trends and processes of topic changes over time, thereby meeting actual application needs.
[0006] The present invention provides a method for visual analysis of scientific and technological development trends based on a word vector network, comprising:
[0007] Get technology keywords;
[0008] A word vector network is constructed based on the technology keywords; the vertices in the word vector network are the technology keywords, and the edges between the vertices are the semantic similarity between the word vectors of the two vertices;
[0009] Performing technology theme community clustering on the technology keyword vertices in the word vector network, and extracting the top M feature keywords with the greatest similarity to the cluster center in different technology theme community clusters;
[0010] The similarity calculation formula of science and technology theme community The community similarity cos(θ) between the science and technology theme community cluster A and the science and technology theme community cluster B is calculated; wherein, the science and technology theme community cluster A and the science and technology theme community cluster B are respectively derived from the word vector network of the adjacent time period, and the science and technology theme community cluster A and the science and technology theme community cluster B have a total of h different science and technology keywords, A i is the unique heat value of the i-th technology keyword in the technology theme community cluster A, B i is the unique heat value of the i-th technology keyword in the technology theme community cluster B;
[0011] Compare the community similarity cos(θ) with a preset threshold to determine the correlation evolution relationship between the science and technology theme community cluster A and the science and technology theme community cluster B in the adjacent time periods;
[0012] The association evolution relationship is visualized in a Cartesian coordinate system.
[0013] Specifically, the technology keyword vertices in the word vector network are clustered into technology theme communities, and the top M feature keywords with the greatest similarity to the cluster center in different technology theme community clusters are extracted, including:
[0014] By formula Calculate the maximum normalized cosine similarity s(x between the non-cluster center technology keywords and the current latest preset cluster center keywords i ); where x i is the word vector of the non-cluster center technology keyword, c j is the word vector of the latest preset cluster center keyword, and C is the set of cluster centers;
[0015] By formula Calculate the probability P(x) that the non-cluster center technology keyword is selected as the new cluster center i ), select the maximum P(x i ) The corresponding technology keywords are used as new cluster centers;
[0016] After multiple rounds of cluster center selection, K cluster centers are finally selected. The normalized cosine similarity between the technology keywords of each non-cluster center and the selected K cluster centers is calculated, and the technology keywords of each non-cluster center are arranged into the cluster C where the cluster center with the largest cosine similarity is located. i middle;
[0017] By formula Calculate the word vector c of the new cluster center of the i-th cluster i '; Among them, x i ′ belongs to the original cluster C i The word vector of the technology keywords, Indicates that the original cluster C i The sum of the word vectors of all the technology keywords in ‖C i ‖ represents the modulus of this vector sum;
[0018] The normalized cosine similarity between the new cluster center and the remaining technology keywords in the cluster to which it belongs is calculated, and the top M technology keywords with the maximum similarity to the cluster center are selected as the feature keywords.
[0019] Specifically, the association evolution relationship is visualized in a Cartesian coordinate system, including:
[0020] According to the set order of technology keyword node labels, the characteristic keywords in each technology theme community cluster are ranked according to the technology theme community cluster on the horizontal axis of the Cartesian coordinate system. After completing the ranking of the characteristic keywords, according to the different time periods to which each characteristic keyword belongs, the technology theme community cluster nodes on the corresponding ranking are drawn on the corresponding positions of the vertical axis of the Cartesian coordinate system. According to the associated evolution relationship, the directed edges between the technology theme community clusters are drawn, and the community similarities between the technology theme community clusters are drawn on the corresponding directed edges.
[0021] Specifically, in the process of drawing the science and technology theme community cluster nodes, science and technology theme community cluster nodes of different diameters are drawn according to the set node size, and the node size is determined by the total number of science and technology keywords in different science and technology theme community clusters.
[0022] Specifically, after visually presenting the association evolution relationship in a Cartesian coordinate system, the following steps are further included:
[0023] Through the community overall centripetal degree formula Calculate the community cluster C i The overall centripetality (i) in the word vector network; wherein the NorCosine function represents the calculation of the normalized cosine similarity, c i Community cluster C i The word vector of the cluster center, c j is the community cluster C i The word vectors of the remaining technology keywords in , k is the community cluster C i The total number of all technology keywords in;
[0024] The nodes of each science and technology theme community cluster and their internal characteristic keywords are arranged from left to right on the horizontal axis of the Cartesian coordinate system in descending order according to the size of the overall centripetality of each community cluster.
[0025] The present invention also provides a system for visualizing and analyzing the development of science and technology based on a word vector network, comprising:
[0026] Technology keyword acquisition module, used to obtain technology keywords;
[0027] A word vector network construction module is used to construct a word vector network based on the technology keywords; the vertices in the word vector network are the technology keywords, and the edges connecting each vertex are the semantic similarity between the word vectors of the two vertices;
[0028] A clustering module is used to cluster the technology keyword vertices in the word vector network into technology theme communities and extract the top M feature keywords with the greatest similarity to the cluster center in different technology theme community clusters;
[0029] Similarity calculation module, used to calculate the similarity formula of science and technology theme communities The community similarity cos(θ) between the science and technology theme community cluster A and the science and technology theme community cluster B is calculated; wherein, the science and technology theme community cluster A and the science and technology theme community cluster B are respectively derived from the word vector network of the adjacent time period, and the science and technology theme community cluster A and the science and technology theme community cluster B have a total of h different science and technology keywords, A i is the unique heat value of the i-th technology keyword in the technology theme community cluster A, B i is the unique heat value of the i-th technology keyword in the technology theme community cluster B;
[0030] An association evolution relationship determination module is used to compare the community similarity cos(θ) with a preset threshold to determine the association evolution relationship between the science and technology theme community cluster A and the science and technology theme community cluster B in adjacent time periods;
[0031] The visualization presentation module is used to visualize the association evolution relationship in a Cartesian coordinate system.
[0032] Specifically, the clustering module includes:
[0033] Cosine similarity calculation unit, used to calculate the cosine similarity through the formula Calculate the maximum normalized cosine similarity s(x between the non-cluster center technology keywords and the current latest preset cluster center keywords i ); where x i is the word vector of the non-cluster center technology keyword, c j is the word vector of the latest preset cluster center keyword, and C is the set of cluster centers;
[0034] Probability calculation unit, used to pass the formula Calculate the probability P(x) that the non-cluster center technology keyword is selected as the new cluster center i ), select the maximum P(x i ) The corresponding technology keywords are used as new cluster centers;
[0035] The clustering unit is used to select K cluster centers through multiple rounds of cluster center selection operations, calculate the normalized cosine similarity between the technology keywords of each non-cluster center and the selected K cluster centers, and arrange the technology keywords of each non-cluster center into the cluster C where the cluster center with the largest cosine similarity is located. i middle;
[0036] Word vector calculation unit, used to calculate the word vector through the formula Calculate the word vector c of the new cluster center of the i-th cluster i '; Among them, x i ′ belongs to the original cluster C i The word vector of the technology keywords, Indicates that the original cluster C i The sum of the word vectors of all the technology keywords in ‖C i ‖ represents the modulus of this vector sum;
[0037] The characteristic keyword acquisition unit is used to calculate the normalized cosine similarity between the new cluster center and the remaining technology keywords in the cluster to which it belongs, and select the top M technology keywords with the maximum similarity to the cluster center as the characteristic keywords.
[0038] Specifically, the visualization presentation module is specifically used to rank the characteristic keywords in each science and technology theme community cluster according to the set order of science and technology keyword node labels on the horizontal axis of the Cartesian coordinate system. After completing the ranking of the characteristic keywords, the science and technology theme community cluster nodes on the corresponding ranking are drawn on the corresponding position of the vertical axis of the Cartesian coordinate system according to the different time periods to which each characteristic keyword belongs. The directed edges between the science and technology theme community clusters are drawn according to the associated evolution relationship, and the community similarities between the science and technology theme community clusters are drawn on the corresponding directed edges.
[0039] Specifically, it also includes:
[0040] The overall centripetal degree calculation module is used to calculate the overall centripetal degree of the community through the formula Calculate the community cluster C i The overall centripetality (i) in the word vector network; wherein the NorCosine function represents the calculation of the normalized cosine similarity, c i Community cluster C i The word vector of the cluster center, C j is the community cluster C i The word vectors of the remaining technology keywords in , k is the community cluster C i The total number of all technology keywords in;
[0041] The visualization adjustment module is used to arrange the science and technology theme community cluster nodes and their internal characteristic keywords from left to right on the horizontal axis of the Cartesian coordinate system in descending order according to the size of the overall centripetality of each community cluster.
[0042] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:
[0043] First, scientific literature data within a domain is selected and matched according to rules to identify scientific keywords reflecting scientific themes. Using word embedding representation learning, the scientific keyword corpus composed of domain literature is nonlinearly transformed into a semantic hyperspace for vectorization, resulting in a domain word vector network with semantic hyperspace metrics. Second, the vertices on the word vector network (scientific keywords reflecting scientific themes) are clustered into scientific theme communities, and the top M characteristic keywords with the greatest similarity to the cluster center in each community cluster are extracted to identify scientific theme communities. Subsequently, a similarity metric is used between scientific theme communities in adjacent time periods to measure the distribution similarity of characteristic keywords between each other, thereby determining the "predecessor-successor" correlation evolution relationship between scientific themes in adjacent time periods. Finally, the correlation evolution relationship between scientific theme communities is visualized in a Cartesian coordinate system to create an evolutionary map of the scientific and technological development trends in the domain. This solves the technical problems of existing technologies, such as the difficulty in controlling keyword selection quality, insufficient utilization of corpus information, and the inability to intuitively and fully reflect the trends and processes of topic changes over time, thus meeting practical application needs.
[0044] In addition, the present invention also has the following advantages:
[0045] 1. Through a recurring matching mechanism, when using author keywords as scientific keywords, we consider including information such as titles and abstracts of scientific papers as expanded corpus fields, and do not remove stop words. Furthermore, to ensure a comprehensive and accurate analysis of the scientific and technological development trends in the field, we use camelCase naming for scientific and technological keyword phrases rather than a simple weighted summation method. We also employ the Glove method to set the VOCAB_MIN_COUNT (minimum keyword) parameter to 1 for word embedding of feature keywords. This achieves precise and comprehensive semantic modeling and representation of relevant concepts within the discipline, reflecting the semantic relevance of concepts.
[0046] 2. Because word embedding networks accurately model and reflect the semantic associations of related concepts within a discipline, they can identify scientific and technological topics through community clustering methods, determine the "predecessor-successor" relationships between topics through similarity discrimination methods, and identify the evolutionary paths of topics, thereby enabling a fine-grained understanding of scientific and technological development trends from a semantic evolutionary perspective.
[0047] 3. Since the word vector network can realize the fine-grained disclosure of the development trend of science and technology from the perspective of semantic evolution. Therefore, based on the identified field science and technology themes and their evolutionary paths, the present invention designs a theme evolution visualization algorithm under the stacked Cartesian coordinate system, and draws an evolution map of the field science and technology development trend. In this map, an indicator for calculating the overall centripetality of the science and technology theme community is proposed. By observing the science and technology themes arranged from left to right on the horizontal axis in the same period, the closeness of the correlation between different science and technology themes can be judged. Secondly, the science and technology keywords with a high correlation with the community cluster center in the same community are used as feature keywords to fully represent the science and technology theme community together, showing the super-network structure characteristics formed between science and technology theme communities. Finally, the "predecessor-successor" evolutionary relationship between science and technology theme communities and the changes in feature keywords between science and technology theme communities in adjacent time periods are further observed to identify the evolutionary development path of science and technology themes and realize the visualization of the development trend of science and technology in the field. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A flowchart of a method for visually analyzing technological development trends based on a word vector network provided by an embodiment of the present invention;
[0049] Figure 2 This is a visual presentation result in an embodiment of the present invention;
[0050] Figure 3 This is the visual presentation result after adjustment in the embodiment of the present invention;
[0051] Figure 4 The overall semantic drift of the information economy and digital economy obtained by the technology development trend visualization analysis method based on the word vector network provided by the embodiment of the present invention;
[0052] Figure 5 This is the visualization result of the topic structure evolution of the case dataset from 2016 to 2020;
[0053] Figure 6 Module diagram of the technology development trend visualization analysis system based on word vector network provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The embodiments of the present invention provide a method and system for visual analysis of scientific and technological development trends based on a word vector network, thereby solving the technical problems in the prior art of difficulty in controlling the quality of keyword selection, insufficient mining and utilization of corpus information, and inability to intuitively and fully reflect the trends and processes of topic changes over time, thereby meeting actual application needs.
[0055] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0056] like Figure 1 As shown, the method for visual analysis of scientific and technological development trends based on a word vector network provided by an embodiment of the present invention includes:
[0057] Step S110: Acquire technology keywords;
[0058] To explain this step in detail, first, scientific and technological literature data covering a complete time period in a specific subject area is screened in the basic data set, the data is continuously and uninterruptedly segmented, and the author keywords of the scientific and technological literature in the field are placed in the corpus consisting of titles and abstracts for recurrence matching to obtain scientific and technological keywords reflecting the scientific and technological themes.
[0059] Next, the Glove toolkit is used to generate word embedding vectors for N-dimensional technology keywords on each data shard. In this process, since most technology keywords are in the form of phrases / phrases, it is difficult for word vectors based on single words to fully express the relevant semantics. Therefore, this embodiment converts phrase / phrase-type technology keywords into their camel case nomenclature form to carry out learning and training of the word embedding model to obtain technology keyword word vectors that express the complete semantics of these phrases / phrases. For example: "Information System" is converted to "InformationSystem" and "SocialNetwork Analysis" is converted to "SocialNetworkAnalysis". At the same time, in order to ensure that all technology keywords in the field of scientific and technological literature can obtain corresponding word embedding vector representations, the minimum word frequency count VOCAB_MIN_COUNT for word vector training is set to 1. In addition to the minimum word frequency count, this embodiment uses the default parameter values of Glove for word vector learning and training, where the word vector dimension N is 50. These technology keywords and their N-dimensional word vectors are obtained through the above process.
[0060] After obtaining these N-dimensional technology keyword word vector sets, this embodiment uses the domain technology keywords and corresponding word vectors on the data slice to construct a domain word vector network in sequence within each time period, so as to finally obtain a time series word vector network within the field.
[0061] Step S120: constructing a word vector network based on technology keywords; the vertices in the word vector network are technology keywords, and the edges connecting each vertex are the semantic similarity between the word vectors of the two vertices;
[0062] The following is a detailed description of the word vector network in the embodiment of the present invention:
[0063] For a document set D = {d1, d2, ..., d i , d m}, where d i represents the i-th document in the text corpus. And document d i ={w i,1 , w i,2 ,...,w i,j , w i,n}, where w i,j Represents document d i For the document set D in the text corpus, its word vector network G = {V, E}, where V represents the vertex set of the network and E represents the edge set of the network. The word vector network G is an undirected weighted graph. The vertex set of the network V = d1∩d2∩...∩d i ∩...∩d m ={W1, W2, ..., W k ,...,W r}, where W k The N-dimensional word vector represents the k-th different word in the entire text corpus. The word vector is mainly obtained through the Glove toolkit. The network edge set E = {e1, e2, e 1,3 ,...,e 2,3 , e 2,4 ,...,e 2,r ..., e p,q ,...,e r-1,r}, where e p,q Represents word W p With the word W q The edge between them, this edge e p,q Further details can be obtained by numerical weight weight p,q To represent this word. In this word vector network, p,q Edge weight p,q It is based on the representation of word W in the text corpus p With the word W q The cosine similarity or Euclidean distance between the word vectors is calculated.
[0064] Specifically, let the domain word vector network in time period t be G t ={V t , E t}, and are the technology keywords in the domain word vector network in time period t, namely and set up The corresponding Glove word vector is v x =(x1, x2, ..., x N ), The corresponding Glove word vector is v y =(y1, y2, ..., y N ), then the edge between the two nodes The weight value That is, the word vector v x With v y The similarity calculation formula is as follows:
[0065]
[0066] where x i and y i They respectively represent the Glove word vectors corresponding to two scientific and technological keywords in the domain word vector network over the time period. Represents the dot product of two word vectors (the sum of the products of the values in the corresponding dimensions). and Represent the modulus (Euclidean norm) of the two word vectors respectively.
[0067] When constructing a domain time series word vector network, it is necessary to calculate the normalized cosine similarity of the corresponding word vectors between the scientific and technological keywords in the domain scientific and technological literature collection at each time period t, and then complete the entire domain word vector network G t Calculation of the edges.
[0068] Step S130: clustering the technology keyword vertices in the word vector network into technology theme communities, and extracting the top M feature keywords with the greatest similarity to the cluster center in different technology theme community clusters;
[0069] This step is explained in detail. The technology keyword vertices in the word vector network are clustered into technology theme communities, and the top M feature keywords with the greatest similarity to the cluster center in different technology theme community clusters are extracted, including:
[0070] By formula Calculate the maximum normalized cosine similarity s(x between the non-cluster center technology keywords and the current latest preset cluster center keywords i ); where x i is the word vector of the non-cluster center technology keyword, c j is the word vector of the latest preset cluster center keyword, and C is the set of cluster centers;
[0071] By formula Calculate the probability P(x) that the non-cluster center technology keyword is selected as the new cluster center i ), select the maximum P(x i ) The corresponding technology keywords are used as new cluster centers;
[0072] After multiple rounds of cluster center selection, K cluster centers are finally selected. The normalized cosine similarity between the technology keywords of each non-cluster center and the selected K cluster centers is calculated, and the technology keywords of each non-cluster center are arranged into the cluster C where the cluster center with the largest cosine similarity is located. i At the same time, these non-cluster center feature keywords are clustered. So far, all feature keywords have their corresponding clusters C i .
[0073] By formula Calculate the word vector c of the new cluster center of the i-th cluster i '; Among them, x i ′ belongs to the original cluster C i The word vector of the technology keywords, Indicates that the original cluster C i The sum of the word vectors of all the technology keywords in ‖C i ‖ represents the modulus of this vector sum;
[0074] Calculate the normalized cosine similarity between the new cluster center and the rest of the technology keywords in the cluster to which it belongs, and select the top M technology keywords with the greatest similarity to the cluster center as feature keywords to determine the label of the technology theme community.
[0075] Step S140: Calculate the similarity of the science and technology theme community using the formula The community similarity cos(θ) between technology theme community cluster A and technology theme community cluster B is calculated; among them, technology theme community cluster A and technology theme community cluster B are respectively derived from the word vector network of the adjacent time period. It is estimated that there are h different technology keywords in technology theme community cluster A and technology theme community cluster B. i is the unique heat value of 0 and 1 of the i-th technology keyword in the technology theme community cluster A, B i is the 0 and 1 unique heat value of the i-th technology keyword in the technology theme community cluster B; if the community similarity cos(θ) is larger, it means that the similarity of the technology keyword distribution between the two technology theme community clusters is higher.
[0076] Step S150: Compare the community similarity cos(θ) with a preset threshold to determine the correlation evolution relationship between the science and technology theme community cluster A and the science and technology theme community cluster B in the adjacent time periods;
[0077] To explain this step in detail, if the community similarity cos(θ) is greater than a preset threshold, it is considered that there is a correlation between them, that is, one science and technology theme community may be the predecessor or successor of another science and technology theme community. Specifically, if the community similarity between science and technology theme community cluster A in the current time period and science and technology theme community cluster B in the next time period is greater than the preset threshold, it means that science and technology theme community cluster A is the predecessor of science and technology theme community cluster B, and science and technology theme community cluster B is the successor of science and technology theme community cluster A.
[0078] Step S160: Visualize the association evolution relationship between science and technology theme communities in a Cartesian coordinate system.
[0079] This step is explained in detail, and the evolutionary relationship between science and technology theme communities is visualized in a Cartesian coordinate system, including:
[0080] According to the order of the set technology keyword node labels, the characteristic keywords in each technology theme community cluster are ranked according to the technology theme community cluster on the horizontal axis of the Cartesian coordinate system. After the ranking of the characteristic keywords is completed, the technology theme community cluster nodes on the corresponding ranking are drawn on the corresponding position of the vertical axis of the Cartesian coordinate system according to the different time periods to which each characteristic keyword belongs. According to the correlation evolution relationship between technology theme communities, the directed edges between each technology theme community cluster are drawn, and the community similarities between each technology theme community cluster are drawn on the corresponding directed edges. Specifically, if the community similarity between technology theme community cluster A in the current time period and technology theme community cluster B in the next time period is greater than the preset threshold, it means that technology theme community cluster A is the predecessor of technology theme community cluster B in evolution, and technology theme community cluster B is the successor of technology theme community cluster A in evolution. Then, the technology theme community cluster A in the previous time period points to the technology theme community cluster B in the next time period, and there is a directed edge between the two communities showing an evolutionary correlation relationship. The cosine value of the angle between the distribution vectors of the two science and technology theme communities is displayed on the directed edge to reflect the strength of the association between the two science and technology themes. The visualization results are as follows: Figure 2 shown.
[0081] In order to make the quantitative characteristics of different science and technology theme community clusters more obvious, in the process of drawing science and technology theme community cluster nodes, science and technology theme community cluster nodes of different diameters are drawn according to the set node size. The node size is determined by the total number of science and technology keywords in different science and technology theme community clusters.
[0082] In order to further describe the evolutionary trends between similar technology-themed communities, after visualizing the correlation and evolutionary relationship between technology-themed communities in a Cartesian coordinate system, the following is also included:
[0083] Through the community overall centripetal degree formula Calculate the community cluster C i The overall centripetality (i) in the word vector network; where the NorCosine function represents the calculation of the normalized cosine similarity, c i Community cluster C i The word vector of the cluster center, c j Community cluster C i The word vectors of the remaining technology keywords in , k is the community cluster C i The total number of all scientific and technological keywords in the community; the overall centrality measures the consistency within a community cluster by calculating the average normalized cosine similarity between all cluster centers within the community cluster. The calculation result reflects the cohesion within the community. A community with high centrality means that its internal nodes are more closely clustered in vector space, indicating that members within the community have high consistency in themes or concepts.
[0084] According to the overall centripetality of each community cluster, the nodes of each science and technology theme community cluster and their internal characteristic keywords are arranged from left to right on the horizontal axis of the Cartesian coordinate system in descending order. The visualization result after adjustment is as follows: Figure 3 shown.
[0085] The effectiveness of the analysis method provided in the embodiment of the present invention is verified below:
[0086] Verification 1: Comparative analysis of the conceptual semantic expression effects of the method of the embodiment of the present invention and the co-word network method
[0087] In order to reveal the differences between the domain word vector network in the embodiment of the present invention and the existing co-word network, Web of Science was selected as the scientific literature retrieval platform, and a text corpus consisting of scientific papers in the field of "information science and library science" from 2011 to 2020 was retrieved as the data set for verifying the effectiveness of the technology.
[0088] Specifically, we used scientific literature in the fields of information science and library science from 2011 to 2016 as training data, with a word embedding dimension of 50. We trained and screened word embeddings for characteristic keywords from 2016. We then used labels like "Bibliometrics," "Social Media," "Qualitative," "Knowledge Management," and "Academic Library" as examples to compare the differences between these highly co-occurring keywords and their semantic similarity, revealing the similarities and differences between the domain word embedding network and the co-word network at the level of inter-word relationships. The experimental results are shown in Tables 1-5:
[0089] Table 1 Highly co-occurring words and their semantic similarity in “Social Media”
[0090]
[0091]
[0092] Table 2. “Bibliometrics” highly co-occurring words and their semantic similarity
[0093]
[0094] Table 3 Highly co-occurring words of “Qualitative” and their semantic similarity
[0095]
[0096] Table 4. High co-occurrence words and semantic similarity of “Knowledge Management”
[0097]
[0098] Table 5. High co-occurrence words and their semantic similarity of “Academic Library”
[0099]
[0100]
[0101] The table above shows the five words that co-occur most with words such as "Social Media," "Bibliometrics," "Qualitative," "Knowledge Management," and "Academic Library," as well as the semantic similarity between the word vectors of these five highly co-occurring words and the current label word. It also identifies and calculates the most similar words between these five highly co-occurring words and the other four label words, as well as the corresponding maximum similarity.
[0102] It can be found that on the one hand, the highly co-occurring words found through the co-occurrence relationship often have a relatively high semantic similarity with the label words. Both the domain word vector network and the co-word network can reflect the knowledge structure of the scientific field to a certain extent. On the other hand, the co-occurrence relationship only considers the local co-occurrence relationship between words based on the author keyword field in the current scientific and technological literature collection, while the semantic association relationship based on the word vector network in the embodiment of the present invention represents the semantic distance between words in the global context composed of a large number of scientific and technological literature collections, and is more stable and comprehensive to a certain extent.
[0103] Verification 2: Comparison of the semantic drift process in the method of the present invention and the co-word network
[0104] In order to further reflect the role of the domain word vector network in the embodiment of the present invention in the analysis of topic evolution, and the difference from the co-word network analysis, this embodiment selects two similar but different words, Digital Economy and Information Economy, and uses the domain word vector network to carry out an empirical case of semantic drift analysis. A total of 77,749 papers with the theme of "Digital Economy", "Information Economy" and its extensions from 2013 to 2023 were selected from Web of Science. After that, the subject domain word vector network was constructed year by year, and the normalized similarity between the paper keywords that remained stable from 2013 to 2023 and information economy and digital economy was calculated. The results are as follows: Figure 4 shown.
[0105] All keywords with increasing and decreasing normalized similarities are listed to further observe the semantic drift of information economy and digital economy from a micro perspective, as shown in Table 6.
[0106] Table 6 Keywords with increasing and decreasing normalized similarity
[0107]
[0108] The keywords in Table 6 are ranked by the absolute value of the difference in normalized similarity between 2013 and 2023. A higher ranking indicates a greater degree of semantic drift. Among the keywords increasingly aligning with the digital economy, technical keywords, such as "Information And Communication Technology," are emerging as a research trend within the digital economy, focusing on the impact of digital technology on economic development. Of the keywords that appeared consistently between 2013 and 2023, only two showed a trend toward the information economy. This suggests that the information economy, while generally a long-term development, has relatively stable semantics, while the subsequent emergence of the digital economy has absorbed more digital-related connotations. This phenomenon is consistent with the earlier observations regarding the information economy.
[0109] For the co-word network, this embodiment selects the top 15 co-occurrence frequencies of information economy and digital economy between 2013 and 2023, and analyzes the co-occurrence frequency sequences of these keywords. The results are shown in Tables 7 and 8.
[0110] Table 7 Co-occurrence frequency sequence of main co-occurring words in information economy
[0111]
[0112] Table 8 Co-occurrence frequency sequence of main co-occurring keywords in digital economy
[0113]
[0114]
[0115] The results in Tables 7 and 8 show that the overall co-occurrence frequency series is very sparse. Key words co-occurring with the information economy and digital economy typically have high co-occurrence frequencies over a continuous period of time. Using co-occurrence frequency series calculated using a co-word network makes it difficult to observe continuous changes in co-occurrences, making it difficult to accurately reveal the semantic drift process.
[0116] The case analysis results of semantic drift in the information economy and digital economy show that the subject domain word vector network in the embodiment of the present invention can accurately reveal the semantic drift process that conforms to the actual situation. It is more flexible in fine-grained semantic analysis tasks and can provide a more fine-grained micro perspective to facilitate the details of the semantic drift of subject domain keywords, providing a new tool for the evolutionary analysis of subject domain research topics.
[0117] Verification 3: Evolutionary Map of Field Technology Development Trends
[0118] Relying on the multi-relationship diagram in the Cartesian coordinate system designed by self-programming research and development in this embodiment, the theme association relationship of the scientific and technological themes from 2016 to 2020 identified in the case data set is visualized, and the details of the evolution process of the scientific and technological themes at the association level are visualized. Finally, a visualization graph of the theme structure evolution analysis is obtained, as shown in the figure below. Figure 5 shown.
[0119] When drawing the visualization graph, the similarity threshold selected in this embodiment to determine whether there is a "predecessor-successor" relationship between two science and technology theme communities is 0.1, that is, when the cosine value of the angle between the characteristic keyword distribution vectors of the two previous and subsequent science and technology theme communities is greater than 0.1, it is determined that there is a relationship.
[0120] By observing the visualization results of the evolution of the science and technology theme structure from 2016 to 2020 in the above dataset, it can be found that these science and technology theme communities have a high degree of continuity between 2016 and 2020, and there are very few communities that have differentiated or merged.
[0121] Based on the multi-relationship diagram in the Cartesian coordinate system designed by the self-programming of this embodiment, it is possible to intuitively and clearly display the distribution changes of science and technology theme communities and their core feature keywords in each time period within the subject field. At the same time, it is also possible to present the "predecessor-successor" relationship between science and technology theme communities and the strength of this relationship. In addition, by observing the sorting of related elements on the visualization graph of the evolution of the theme structure, it is also possible to find changes in the overall centripetality index reflecting the closeness of the relationship between the science and technology theme and other themes in the domain word vector network in the same time period, as well as changes in the number of feature keywords in the science and technology theme in each time period, that is, the change in the theme scale.
[0122] like Figure 6 As shown, the technology development trend visualization analysis system based on the word vector network provided by the embodiment of the present invention includes:
[0123] A technology keyword acquisition module 100 is used to acquire technology keywords;
[0124] The word vector network construction module 200 is used to construct a word vector network based on technology keywords; the vertices in the word vector network are technology keywords, and the edges connecting each vertex are the semantic similarity between the word vectors of the two vertices;
[0125] Clustering module 300, for clustering the technology keyword vertices in the word vector network into technology theme communities, and extracting the top M feature keywords with the greatest similarity to the cluster center in different technology theme community clusters;
[0126] Specifically, the clustering module 300 includes:
[0127] Cosine similarity calculation unit, used to calculate the cosine similarity through the formula Calculate the maximum normalized cosine similarity s(x between the non-cluster center technology keywords and the current latest preset cluster center keywords i ); where x i is the word vector of the non-cluster center technology keyword, c j is the word vector of the latest preset cluster center keyword, and C is the set of cluster centers;
[0128] Probability calculation unit, used to pass the formula Calculate the probability P(x) that the non-cluster center technology keyword is selected as the new cluster center i ), select the maximum P(x i ) The corresponding technology keywords are used as new cluster centers;
[0129] The clustering unit is used to select K cluster centers through multiple rounds of cluster center selection operations, calculate the normalized cosine similarity between the technology keywords of each non-cluster center and the selected K cluster centers, and arrange the technology keywords of each non-cluster center into the cluster C where the cluster center with the largest cosine similarity is located. i At the same time, these non-cluster center feature keywords are clustered. So far, all feature keywords have their corresponding clusters C i .
[0130] Word vector calculation unit, used to calculate the word vector through the formula Calculate the word vector c of the new cluster center of the i-th cluster i '; Among them, x i ′ belongs to the original cluster C i The word vector of the technology keywords, Indicates that the original cluster C i The sum of the word vectors of all the technology keywords in ‖C i ‖ represents the modulus of this vector sum;
[0131] The feature keyword acquisition unit is used to calculate the normalized cosine similarity between the new cluster center and the remaining technology keywords in the cluster to which it belongs, and select the top M technology keywords with the greatest similarity to the cluster center as feature keywords to determine the label of the technology theme community.
[0132] Similarity calculation module 400 is used to calculate the similarity of science and technology theme communities through the formula The community similarity cos(θ) between technology theme community cluster A and technology theme community cluster B is calculated; among them, technology theme community cluster A and technology theme community cluster B are respectively derived from the word vector network of the adjacent time period. It is estimated that there are h different technology keywords in technology theme community cluster A and technology theme community cluster B. i is the unique heat value of 0 and 1 of the i-th technology keyword in the technology theme community cluster A, B i is the 0 and 1 unique heat value of the i-th technology keyword in the technology theme community cluster B; if the community similarity cos(θ) is larger, it means that the similarity of the technology keyword distribution between the two technology theme community clusters is higher.
[0133] The correlation evolution relationship determination module 500 is used to compare the community similarity cos(θ) with a preset threshold to determine the correlation evolution relationship between the science and technology theme community cluster A and the science and technology theme community cluster B in the adjacent time periods;
[0134] Specifically, the association evolution relationship determination module 500 is specifically used to compare the community similarity cos(θ) with a preset threshold. If the community similarity cos(θ) is greater than the preset threshold, it is considered that there is an association between them, that is, one science and technology theme community may be the predecessor or successor of another science and technology theme community. Specifically, if the community similarity between science and technology theme community cluster A in the current time period and science and technology theme community cluster B in the next time period is greater than the preset threshold, it means that science and technology theme community cluster A is the predecessor of science and technology theme community cluster B in evolution, and science and technology theme community cluster B is the successor of science and technology theme community cluster A in evolution.
[0135] The visualization presentation module 600 is used to visualize the association evolution relationship between science and technology theme communities in a Cartesian coordinate system.
[0136] Specifically, the visualization presentation module 600 is specifically used to rank the characteristic keywords in each science and technology theme community cluster according to the set order of science and technology keyword node labels on the horizontal axis of the Cartesian coordinate system. After completing the ranking of the characteristic keywords, the science and technology theme community cluster nodes on the corresponding ranking are drawn on the corresponding positions of the vertical axis of the Cartesian coordinate system according to the different time periods to which each characteristic keyword belongs. According to the correlation evolution relationship between the science and technology theme communities, the directed edges between the science and technology theme community clusters are drawn, and the community similarities between the science and technology theme community clusters are drawn on the corresponding directed edges. Specifically, if the community similarity between technology-themed community cluster A in the current time period and technology-themed community cluster B in the next time period is greater than a preset threshold, it means that technology-themed community cluster A is the evolutionary predecessor of technology-themed community cluster B, and technology-themed community cluster B is the evolutionary successor of technology-themed community cluster A. Therefore, there is a directed edge between the two communities, from technology-themed community cluster A in the previous time period to technology-themed community cluster B in the next time period, indicating an evolutionary relationship. The directed edge displays the cosine value of the angle between the technology keyword distribution vectors of the two technology-themed communities to reflect the strength of the association between the two technology themes.
[0137] In order to further describe the evolutionary trends between communities with similar technology themes, it also includes:
[0138] The overall centripetal degree calculation module is used to calculate the overall centripetal degree of the community through the formula Calculate the community cluster C i The overall centripetality (i) in the word vector network; where the NorCosine function represents the calculation of the normalized cosine similarity, c i Community cluster C i The word vector of the cluster center, c j Community cluster C i The word vectors of the remaining technology keywords in , k is the community cluster C i The total number of all scientific and technological keywords in the community; the overall centrality measures the consistency within a community cluster by calculating the average normalized cosine similarity between all cluster centers within the community cluster. The calculation result reflects the cohesion within the community. A community with high centrality means that its internal nodes are more closely clustered in vector space, indicating that members within the community have high consistency in themes or concepts.
[0139] The visualization adjustment module is used to arrange the nodes of each science and technology theme community cluster and their internal characteristic keywords from left to right on the horizontal axis of the Cartesian coordinate system in descending order according to the size of the overall centripetality of each community cluster.
[0140] The embodiment of the present invention provides a method and system for visual analysis of the development trend of science and technology based on a word vector network. By screening scientific and technological literature data to construct a domain word vector network, using a clustering algorithm to identify scientific and technological topics, measuring the correlation relationship of the evolution trend of scientific and technological topics, and designing a visualization algorithm, a domain science and technology development trend evolution map is drawn. By performing operations such as node sorting, node appearance setting, and superimposing a multi-dimensional Cartesian coordinate system on the trend evolution map, on the one hand, the "predecessor-successor" relationship between scientific and technological topics can be identified, and the changes in the domain science and technology topic community can be clearly displayed. On the other hand, through vivid visualization graphics, the evolution path identification of domain science and technology topics can be easily completed, and a multi-dimensional visualization analysis of the domain science and technology development trend can be achieved. Experimental verification shows that the word vector network in the embodiment of the present invention is superior to the existing co-word network in terms of conceptual semantic expression, semantic drift analysis, etc., and can provide a more stable and comprehensive perspective for the analysis of scientific and technological development trends, helping scientific researchers and decision-making departments to grasp cutting-edge hotspots and optimize resource allocation.
[0141] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0142] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0143] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0145] Any details not described in the embodiments of the present invention are well-known to those skilled in the art. Finally, it should be noted that the above embodiments are only intended to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or equivalents should be included in the scope of the claims of the present invention.
Claims
1. A visual analysis method for scientific and technological development trends based on word vector networks, characterized in that: include: Get technology keywords; Constructing a word vector network based on the scientific and technological keywords; The vertices in the word vector network are the technology keywords, and the edges between each vertex are the semantic similarity between the word vectors of the two vertices; Performing technology theme community clustering on the technology keyword vertices in the word vector network, and extracting the top M feature keywords with the greatest similarity to the cluster center in different technology theme community clusters; The similarity calculation formula of science and technology theme community Calculate the community similarity between technology theme community cluster A and technology theme community cluster B ; Wherein, the science and technology theme community cluster A and the science and technology theme community cluster B are respectively derived from the word vector network of the adjacent time period, and h is the total number of different science and technology keywords that exist in the science and technology theme community cluster A and the science and technology theme community cluster B. The first The unique heat value of technology keywords, The first The unique heat value of a technology keyword; The community similarity Compare with the preset threshold to determine the correlation evolution relationship between the science and technology theme community cluster A and the science and technology theme community cluster B in the adjacent time periods; Visualizing the association evolution relationship in a Cartesian coordinate system; Visualizing the association evolution relationship in a Cartesian coordinate system includes: According to the set order of technology keyword node labels, the characteristic keywords in each technology theme community cluster are ranked according to the technology theme community cluster on the horizontal axis of the Cartesian coordinate system. After completing the ranking of the characteristic keywords, according to the different time periods to which each characteristic keyword belongs, the technology theme community cluster nodes on the corresponding ranking are drawn on the corresponding positions of the vertical axis of the Cartesian coordinate system. According to the associated evolution relationship, the directed edges between the technology theme community clusters are drawn, and the community similarities between the technology theme community clusters are drawn on the corresponding directed edges.
2. The method for visual analysis of scientific and technological development trends based on a word vector network according to claim 1, characterized in that: The technology keyword vertices in the word vector network are clustered into technology theme communities, and the top M feature keywords with the greatest similarity to the cluster center in different technology theme community clusters are extracted, including: By formula Calculate the maximum normalized cosine similarity between the non-cluster center technology keywords and the current latest preset cluster center keywords ;in, is the word vector of the non-cluster center technology keyword, is the word vector of the currently latest preset cluster center keyword, is the set of cluster centers; By formula Calculate the probability that the non-cluster center technology keyword is selected as the new cluster center , select the maximum The corresponding technology keywords serve as new cluster centers; After multiple rounds of cluster center selection, K cluster centers are finally selected. The normalized cosine similarity between the technology keywords of each non-cluster center and the selected K cluster centers is calculated, and the technology keywords of each non-cluster center are arranged into the cluster where the cluster center with the largest cosine similarity is located. middle; By formula Calculate the first The word vector of the new cluster center of the cluster ;in, Belongs to the original cluster The word vector of the technology keywords, Indicates that the original cluster The vector sum is obtained by adding the word vectors of all scientific and technological keywords in , Represents the modulus of this vector sum; The normalized cosine similarity between the new cluster center and the remaining technology keywords in the cluster to which it belongs is calculated, and the top M technology keywords with the maximum similarity to the cluster center are selected as the feature keywords.
3. The method for visual analysis of scientific and technological development trends based on a word vector network according to claim 1, characterized in that: In the process of drawing the science and technology theme community cluster nodes, science and technology theme community cluster nodes of different diameters are drawn according to the set node size, and the node size is determined by the total number of science and technology keywords in different science and technology theme community clusters.
4. The method for visual analysis of scientific and technological development trends based on a word vector network according to claim 1, characterized in that: After visually presenting the association evolution relationship in a Cartesian coordinate system, the method further includes: Through the community overall centripetal degree formula Calculate the community clusters The overall centripetality in the word vector network ;in, The function represents the calculation of normalized cosine similarity. Community cluster The word vector of the cluster center, The community cluster The word vectors of the remaining technology keywords in , k is the community cluster The total number of all technology keywords in; The nodes of each science and technology theme community cluster and their internal characteristic keywords are arranged from left to right on the horizontal axis of the Cartesian coordinate system in descending order according to the size of the overall centripetality of each community cluster.
5. A visual analysis system for scientific and technological development trends based on word vector networks, characterized by: include: Technology keyword acquisition module, used to obtain technology keywords; A word vector network construction module, used to construct a word vector network based on the scientific and technological keywords; The vertices in the word vector network are the technology keywords, and the edges between each vertex are the semantic similarity between the word vectors of the two vertices; A clustering module is used to cluster the technology keyword vertices in the word vector network into technology theme communities and extract the top M feature keywords with the greatest similarity to the cluster center in different technology theme community clusters; Similarity calculation module, used to calculate the similarity formula of science and technology theme communities Calculate the community similarity between technology theme community cluster A and technology theme community cluster B ; Wherein, the science and technology theme community cluster A and the science and technology theme community cluster B are respectively derived from the word vector network of the adjacent time period, and h is the total number of different science and technology keywords that exist in the science and technology theme community cluster A and the science and technology theme community cluster B. The first The unique heat value of technology keywords, The first The unique heat value of a technology keyword; The association evolution relationship determination module is used to determine the community similarity Compare with the preset threshold to determine the correlation evolution relationship between the science and technology theme community cluster A and the science and technology theme community cluster B in the adjacent time periods; A visualization presentation module, used for visually presenting the association evolution relationship in a Cartesian coordinate system; The visualization presentation module is specifically used to rank the characteristic keywords in each science and technology theme community cluster according to the set order of science and technology keyword node labels on the horizontal axis of the Cartesian coordinate system. After completing the ranking of the characteristic keywords, the science and technology theme community cluster nodes on the corresponding ranking are drawn on the corresponding positions of the vertical axis of the Cartesian coordinate system according to the different time periods to which each characteristic keyword belongs. The directed edges between the science and technology theme community clusters are drawn according to the associated evolution relationship, and the community similarities between the science and technology theme community clusters are drawn on the corresponding directed edges.
6. The technology development trend visualization analysis system based on word vector network according to claim 5 is characterized in that: The clustering module includes: Cosine similarity calculation unit, used to calculate the cosine similarity through the formula Calculate the maximum normalized cosine similarity between the non-cluster center technology keywords and the current latest preset cluster center keywords ;in, is the word vector of the non-cluster center technology keyword, is the word vector of the currently latest preset cluster center keyword, is the set of cluster centers; Probability calculation unit, used to pass the formula Calculate the probability that the non-cluster center technology keyword is selected as the new cluster center , select the maximum The corresponding technology keywords serve as new cluster centers; The clustering unit is used to select K cluster centers through multiple rounds of cluster center selection operations, calculate the normalized cosine similarity between the technology keywords of each non-cluster center and the selected K cluster centers, and arrange the technology keywords of each non-cluster center into the cluster where the cluster center with the largest cosine similarity is located. middle; Word vector calculation unit, used to calculate the word vector through the formula Calculate the first The word vector of the new cluster center of the cluster ;in, Belongs to the original cluster The word vector of the technology keywords, Indicates that the original cluster The vector sum is obtained by adding the word vectors of all scientific and technological keywords in , Represents the modulus of this vector sum; The characteristic keyword acquisition unit is used to calculate the normalized cosine similarity between the new cluster center and the remaining technology keywords in the cluster to which it belongs, and select the top M technology keywords with the maximum similarity to the cluster center as the characteristic keywords.
7. The technology development trend visualization analysis system based on word vector network according to claim 5 is characterized in that: Also includes: The overall centripetal degree calculation module is used to calculate the overall centripetal degree of the community through the formula Calculate the community clusters The overall centripetality in the word vector network ;in, The function represents the calculation of normalized cosine similarity. Community cluster The word vector of the cluster center, The community cluster The word vectors of the remaining technology keywords in , k is the community cluster The total number of all technology keywords in; The visualization adjustment module is used to arrange the science and technology theme community cluster nodes and their internal characteristic keywords from left to right on the horizontal axis of the Cartesian coordinate system in descending order according to the size of the overall centripetality of each community cluster.
Citation Information
Patent Citations
Literature-based geoscience research hotspot extraction and visualization method and system
CN117708333A
Method and apparatus for acquiring hot topics
US20140280242A1