A method for measuring technological competitiveness based on a technological theme map
By calculating word segmentation and subject word relationship matrix of scientific and technological text data, combining clustering and layout algorithms, the density function and color intensity function are constructed, and the problem of lack of refined and fine-grained numerical measurement of existing technical competitiveness measurement methods is solved, and the competitiveness measurement of technical theme diagrams is realized, providing dual support for vision and numerical values.
Patent Information
- Application Number
- CN202210928186.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-08-03
AI Technical Summary
The lack of integration of the existing technology competitiveness measurement methods with technical topic diagrams leads to the lack of refined and fine-grained numerical measurement references for visual analysis.
By performing word segmentation processing and subject word relationship matrix calculation on scientific and technological text data, clustering algorithms and layout algorithms are used to map subject words to the spatial plane, class density functions and color intensity functions are constructed for visualization, and a competitiveness measurement model for enterprises or R&D institutions under technical topics are constructed.
The technical competitiveness measurement based on the technical theme diagram is realized, providing visual intuitive perception and refined numerical references, helping decision makers make technical decisions more accurately.
Smart Images

Figure CN115186107B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text data processing, and particularly relates to a method for measuring technical competitiveness based on a technical theme map. Background Art
[0002] With the continuous development of text mining and information visualization technologies, numerous analysis methods for scientific and technological text data such as patents, papers, technical standards, and research reports have emerged. The technical theme map integrates text mining and information visualization technologies, mines technical information from scientific and technological text data, and uses intuitive graphics to represent the technical layout, which is widely applied to analysis scenarios such as technical layout analysis, technical competition analysis, and technical structure analysis. The analysis of scientific and technological text data based on the technical theme map provides visual support for innovation decision-makers from a macroscopic perspective. However, visual perception is global, coarse-grained, and non-refined, and decision-makers often need more accurate numerical metrics to assist in decision-making during the decision-making process.
[0003] Existing technical competitiveness measurement methods separately construct competitiveness models without integrating with the technical theme map. Specifically, they can be divided into those based on the data envelopment analysis method (DEA), such as using the sequential DEA method to measure the technical gap between Chinese industrial industries and the world technology frontier mainly composed of industrial industries in OECD countries [1]. Using the common frontier theory, a parallel network DEA model is constructed to measure the innovation efficiency of Chinese high-tech manufacturing industries and the technical gap between regions from 2007 to 2015, etc. [2]. Based on the function model method, such as introducing a two-country transcendental logarithmic production model to estimate the total factor productivity gap between the manufacturing industries of China and the United States by industry [3]. Introducing the super technology function to analyze the agricultural technology gap between regions in China [4]. These methods measure technical competitiveness through a single numerical value without integrating with the technical theme map and lack a global intuitive visual perception.
[0004] Using the technical theme map for the technical analysis of scientific and technological text data is clear and easy to understand in terms of visual presentation, but it lacks refined and fine-grained decision data references. In the visual presentation of the technical theme map, the presentation of the theme density map has insufficient recognition and aesthetic degree; the visual content presentation mostly analyzes the technical competition situation from a macroscopic perspective and lacks fine-grained, refined, and accurate numerical metric references.
[0005] References:
[0006] [1] Lu Jian, Liu Jianping, Cheng Shixiong. Dynamic Measurement of the Technical Gap between Chinese and Main OECD Countries' Industrial Industries [J]. World Economy, 2014, 37(09): 25 - 52.
[0007] [2] Xiao Renqiao, Chen Zhongwei, Qian Li. Research on the Innovation Efficiency of High-Tech Manufacturing in China from the Perspective of Heterogeneous Technology [J]. Management Science, 2018, 31(01): 48-68.
[0008] [3] Huang Yongfeng, Ren Ruoen. Comparative Study on Total Factor Productivity of the Manufacturing Industries in China and the United States [J]. China Economic Quarterly, 2002(04): 161-180.
[0009] [4] Yang Guotao. A Method for Measuring the Agricultural Technology Gap among Regions in China Using a Transcendental Function [J]. Journal of Anhui Agricultural Sciences, 2007(22): 6693-6694. DOI: 10.13989 / j.cnki.0517-6611.2007.22.030. Summary of the Invention
[0010] Aiming at the deficiencies of the prior art, the present invention aims to provide a method for measuring technological competitiveness based on a technology theme map.
[0011] To achieve the above object, the present invention adopts the following technical solutions:
[0012] A method for measuring technological competitiveness based on a technology theme map, the specific process is as follows:
[0013] S1. Perform word segmentation on the scientific and technological text dataset, and calculate the membership relationship matrix between each scientific and technological text and the subject words;
[0014] S2. Based on the membership relationship matrix between each scientific and technological text and the subject words, calculate the relationship strength matrix between all subject words;
[0015] S3. According to the relationship strength matrix obtained in step S2, cluster all subject words according to the relationship strength between each subject word, and add category labels to each subject word according to the clustering result; denote the number of categories after clustering as C, that is, the subject words are divided into C groups; then, map the subject words to points in the spatial plane through different layout algorithms;
[0016] S4. Construct a plane pixel point class density function for visualization:
[0017] S4.1. Assume that the coordinates of n subject words are (x i , y i ), i = 1…n, and the average value of the two-dimensional Euclidean distance between subject words is Number i , i = 1…n, representing the number of scientific and technological texts in which the subject word i appears; after clustering, there are a total of C categories, and there are n c subject words in each category; f(Number i) is the standardized value of the subject term i; the coordinates (x, y) of the pixel point P; where the density functions α and β are non - negative numbers;
[0018] Define the density function of the pixel point and the class density function as:
[0019] Density function:
[0020] Class density function:
[0021] c represents a specific category after clustering;
[0022] S4.2. After fusing the clustering information, use Density max to represent the maximum density value, Color i to represent the RGB - mode color of the category i = 1…C; the RGB - mode color of the pixel point P(x, y) is calculated as follows:
[0023]
[0024] where Color i is the value of each channel of the RGB - mode color;
[0025] S4.3. To achieve a visualization effect similar to the contour lines of a topographic map, and at the same time the contour lines can distinguish the subject terms under the same category and also distinguish the subject terms under different categories, construct the color intensity function:
[0026] f(Density(x, y) / Density(max));
[0027] S5. Construct the competitiveness measurement model of enterprise or R & D institution i under the technical theme j:
[0028] w is the number of enterprises or R & D institutions participating in the competitiveness measurement; s is the number of technical themes covered by the technical field; n i,j is the number of literatures of enterprise or R & D institution i under the technical theme j, Position(D k ) represents the technical status of the scientific and technological text k, and its value is the color intensity calculated in step S4.3, Ability(D k ) is the quality of the scientific and technological text k; therefore, the competitiveness measurement model of enterprise or R & D institution i under the technical theme j is as follows:
[0029]
[0030] Corpration i represents enterprise or R & D institution i, Technology j represents the technical theme j.
[0031] S6. Form a competitiveness matrix among institutions and technical themes in the following form.
[0032] Furthermore, in step S1, the subordination relationship matrix is expressed as follows:
[0033]
[0034] where m represents the number of scientific and technological texts, n represents the number of subject terms, Document i represents the i-th scientific and technological text, Keyword j represents the j-th subject term, and b ij represents the number of occurrences of the j-th subject term in the i-th scientific and technological text.
[0035] Furthermore, in step S2, the relationship strength matrix is expressed as follows:
[0036]
[0037] where n represents the number of subject terms, Keyword i , Keyword j represent the i-th and j-th subject terms, and r ij represents the number of scientific and technological texts in which the i-th subject term and the j-th subject term co-occur.
[0038] Furthermore, in step S4.3, the color intensity function is specifically:
[0039]
[0040] where is the floor function, and N is the number of intensity levels.
[0041] Furthermore, the competitiveness measurement model in step S5 can be transformed into:
[0042]
[0043] where D k (x, y) is the coordinate of the scientific and technological text k in the technical theme map.
[0044] Furthermore, Ability(D k ) is expressed as the ratio of the number of citations of a scientific and technological text to the number of institutions citing the text, that is, the quality of a single scientific and technological text is evaluated by the number of citations of the scientific and technological text and the coverage of the citing institutions. Thus, the competitiveness measurement model is further transformed as follows:
[0045]
[0046] Among them, ReferencedNumber(D k ) is the number of times scientific and technological text k is cited, and CorprationNumber(D k ) is the number of institutions that have cited scientific and technological text k.
[0047] Furthermore, the competitiveness measure is normalized so that it is between 0 and 1 for easy comparison; the calculation formula is as follows:
[0048]
[0049] Furthermore, in step S6, the competitiveness matrix is expressed as follows:
[0050]
[0051] Among them, q ij is the technical competitiveness measure value of enterprise or R & D institution i under technical theme j.
[0052] The beneficial effects of the present invention are as follows: The method for measuring technical competitiveness based on technical theme map analysis provided by the present invention constructs a technical competitiveness measurement model based on the technical theme map, and measures the technical competitiveness between enterprises or R & D institutions with specific numbers, enabling decision-makers to not only have an intuitive visual perception when observing the technical theme map, but also have more direct and more refined numerical references. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a schematic diagram of the integration of the technical theme map layout algorithm in the embodiment of the present invention;
[0054] Figure 2 is a schematic diagram of the construction idea of the technical competitiveness measurement model based on the technical theme map in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The following will further describe the present invention in conjunction with the drawings. It should be noted that this embodiment is based on the present technical solution and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to this embodiment.
[0056] This embodiment provides a method for measuring technical competitiveness based on a technical theme map. The specific process is as follows:
[0057] S1. Perform word segmentation on the scientific and technological text data set, and calculate the membership relationship matrix between each scientific and technological text and the subject words. The membership relationship matrix is expressed as follows:
[0058]
[0059] Among them, m represents the number of scientific and technological texts, n represents the number of subject terms, and Document i represents the i-th scientific and technological text, and Keyword j represents the j-th subject term, and b ij represents the number of occurrences of the j-th subject term in the i-th scientific and technological text.
[0060] S2. Calculate the relationship strength matrix between all subject terms based on the membership relationship matrix between each scientific and technological text and subject terms. The relationship strength matrix is expressed as follows:
[0061]
[0062] The relationship strength matrix is specifically calculated by the product of the transpose of the membership relationship matrix and the membership relationship matrix, or calculated based on the membership relationship matrix using methods such as inverted document frequency, information entropy, and mutual information. Among them, n represents the number of subject terms, and Keyword i , Keyword j represent the i-th and j-th subject terms, and r ij represents the relationship strength between the i-th subject term and the j-th subject term.
[0063] S3. According to the relationship strength matrix obtained in step S2, cluster all subject terms according to the relationship strength between each subject term, and add category labels to each subject term according to the clustering results. Using the K-Means clustering algorithm, assume that the number of categories after clustering is C, that is, the subject terms are divided into C groups. Then, map the subject terms to points in the spatial plane through different layout algorithms. The overall layout algorithm fusion schematic diagram is as shown in Figure 1 shown.
[0064] S4. Construct a plane pixel point class density function for visualization:
[0065] S4.1. Assume that the coordinates of n subject terms are (x i , y i ), i = 1...n, and the average value of the two-dimensional Euclidean distance between subject terms is Number i , i = 1...n, represents the number of scientific and technological texts in which the subject term i appears, used to reveal the text content under the technical theme; after clustering, there are a total of C categories, and each category has n c subject terms; f(Number i ) is the normalized value of the subject term i; the coordinates of the pixel point P are (x, y). Among them, the density function α, β are non-negative numbers, and their values are different, and the theme map effects are different.
[0066] Define the density function and class density function of the pixel point as:
[0067] Density function:
[0068] Class density function:
[0069] where c represents a specific category after clustering.
[0070] S4.2. After integrating the clustering information, use Density max to represent the maximum density value, and Color i to represent the RGB mode color of category i = 1...C.
[0071] The RGB mode color of pixel point P(x, y) is calculated as follows:
[0072]
[0073] where Color i is the value of each channel of the RGB mode color.
[0074] S4.3. To achieve a visualization effect similar to the contour lines of a topographic map, and at the same time, the contour lines can distinguish the subject words under the same category and also distinguish the subject words under different categories, construct a color intensity function:
[0075] f(Density(x, y) / Density(max))
[0076] The color intensity function should be a step function to achieve the contour line effect. A simple color intensity function can be:
[0077]
[0078] where is the floor function, N is the number of intensity levels, and the number of intensity levels directly affects the drawing effect, making the visualization result different from the heat map, density map, and general topographic map form technology theme map.
[0079] S5. Construct a competitiveness measurement model for enterprise or R & D institution i under technology theme j. The construction idea is as Figure 2 shown. w is the number of enterprises or R & D institutions participating in the competitiveness measurement; s is the number of technology themes covered by the technology field, which is determined by clustering or community discovery algorithms in the technology themes; n i,j is the number of documents of enterprise or R & D institution i in technology theme j, Position(D k ) is the technical status of scientific and technological text k, and Ability(D k ) is the quality of scientific and technological text k. Corresponding to the technology theme map, Position(Dk ) The value is the color intensity calculated in step S4.3. When N = 10, its value ranges from {0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9}. Ability(D k ) is an evaluation index of the quality of a single scientific and technological text, and numerical indexes such as the number of citations of a paper and the number of patent families can be used to measure the quality of a single scientific and technological text. Therefore, the competitiveness measurement model of enterprise or R & D institution i under technical theme j is as follows:
[0080]
[0081] Corpration i represents enterprise or R & D institution i, Technology j represents technical theme j.
[0082] The above formula is transformed into the following formula:
[0083]
[0084] where D k (x, y) is the coordinate of scientific and technological text k in the technical theme graph;
[0085] Ability(D k ) is represented by the ratio of the number of citations of a scientific and technological text to the number of institutions citing the scientific and technological text, that is, the quality of a single scientific and technological text is evaluated by the number of citations of the scientific and technological text and the coverage of the citing institutions. Thus, the competitiveness measurement model is further transformed as follows:
[0086]
[0087] where, ReferencedNumber(D k ) is the number of citations of scientific and technological text k, and CorprationNumber(D k ) is the number of institutions citing scientific and technological text k.
[0088] The competitiveness measurement is normalized to be between 0 and 1 for easy comparison. The calculation formula is as follows:
[0089]
[0090] S6. Form a competitiveness matrix between institutions and technical themes in the following form.
[0091]
[0092] where, q ijIt is the measurement value of the technological competitiveness of enterprise or R & D institution i under technological theme j.
[0093] For those skilled in the art, various corresponding changes and deformations can be given according to the above technical solutions and concepts, and all these changes and deformations should be included within the protection scope of the claims of the present invention.
Claims
1. A method for measuring technological competitiveness based on a technology theme map, characterized in that, the specific process is as follows: S1. Perform word segmentation on the scientific and technological text dataset, and calculate the membership relationship matrix between each scientific and technological text and the subject words; S2. Based on the membership relationship matrix between each scientific and technological text and the subject words, calculate the relationship strength matrix between all subject words; S3. According to the relationship strength matrix obtained in step S2, apply a clustering algorithm to cluster all subject words according to the relationship strength between each subject word, and add a category label to each subject word according to the clustering result; Denote the number of categories after clustering as C, that is, the subject words are divided into C groups; Then, map the subject words to points in the spatial plane through different layout algorithms; S4. Construct a plane pixel point class density function for visualization: S4.
1. Assume that the coordinates of n subject terms are (x i , y i ), i = 1...n, and the average two-dimensional Euclidean distance between subject terms is Number i , i = 1...n, representing the number of scientific and technological texts in which subject term i appears; after clustering, there are C categories, and each category has n c subject terms; f(Number i ) is the normalized value of subject term i; the coordinates (x, y) of pixel point P; where the density functions α and β are non-negative numbers; Define the density function and class density function of the pixel points as: Density function: Class density function: c represents a specific category after clustering; S4.
2. After integrating the clustering information, use Density max to represent the maximum density value, and Color i to represent the RGB-mode color of class i = 1... C. The RGB-mode color of pixel point P(x, y) is calculated as follows: Among them, Color i is the value of each channel of the RGB mode color; S4.
3. In order to achieve a visualization effect similar to the contour lines of a topographic map, and at the same time the contour lines can distinguish the subject words under the same category and also distinguish the subject words under different categories, construct a color intensity function: Among them, is rounding down, and N is the number of strength levels; S5. Construct a competitiveness measurement model of enterprise or R & D institution i under technology theme j: w is the number of enterprises or R & D institutions participating in the competitiveness measurement; s is the number of technical topics covered by the technical field; n i,j is the number of literatures of enterprise or R & D institution i in technical topic j, Position(D k ) represents the technical status of scientific and technological text k, and its value is the color intensity calculated in step S4.
3. Ability(D k ) is the quality of scientific and technological text k. Therefore, the competitiveness measurement model of enterprise or R & D institution i under technical topic j is as follows: Corpration i represents enterprise or R & D institution i, Technology j represents technology theme j; D k (x, y) is the coordinate of scientific and technological text k in the technology theme map; S6. Form a competitiveness matrix between institutions and technology themes in the following form.
2. The method according to claim 1, characterized in that, in step S1, the membership relationship matrix is expressed as follows: Among them, m represents the number of scientific and technological texts, n represents the number of subject words, Document i represents the i-th scientific and technological text, Keyword j represents the j-th subject word, b ij represents the number of occurrences of the j-th subject word in the i-th scientific and technological text.
3. The method according to claim 1, characterized in that, in step S2, the relationship strength matrix is expressed as follows: Among them, n represents the number of subject terms, Keyword i and Keyword j represent the i-th and j-th subject terms, respectively, and r ij represents the number of scientific and technological texts in which the i-th subject term and the j-th subject term co-occur.
4. The method according to claim 1, characterized in that, Ability(D k ) It is represented by the ratio of the number of citations of a scientific and technological text to the number of institutions citing the scientific and technological text. That is, the quality of a single scientific and technological text is evaluated by the number of citations of the scientific and technological text and the coverage of the citing institutions. Thus, the competitiveness measurement model is further transformed as follows: Among them, ReferencedNumber(D k ) is the number of times scientific and technological text k is cited, and CorprationNumber(D k ) is the number of institutions that have cited scientific and technological text k.
5. The method according to claim 1, characterized in that, Perform normalization processing on the competitiveness measurement to make it between 0 and 1 for easy comparison; The calculation formula is as follows:
6. The method according to claim 1, characterized in that, in step S6, the competitiveness matrix is expressed as follows: Among them, q ij is the measurement value of the technological competitiveness of enterprise or R & D institution i under technology theme j.
Citation Information
Patent Citations
Regression analysis-based news competitiveness analysis method and visualization device
CN105373579A