A text mining-based decision-making support method and system for scientific and technological project establishment management
Through the text mining method, a multi-level and multi-dimensional scientific and technological project similarity comparison model is constructed, which solves the problem of low efficiency in the project establishment management of repeated projects and similarity identification, and realizes efficient and accurate project similarity analysis and management.
Patent Information
- Application Number
- CN202111587067.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-12-23
AI Technical Summary
In the power system, there are problems in the management of scientific and technological project project establishment, low efficiency in identifying project similarity, and unreasonable selection of review experts, resulting in low management efficiency and accuracy.
Using text mining methods, through technologies such as information extraction, layered text similarity mining and grid search, a multi-level and multi-dimensional scientific and technological project similarity comparison model is constructed, reducing manual screening, and improving the efficiency and accuracy of project similarity analysis.
It has achieved the reduction of manual screening, improved the efficiency and accuracy of project similarity analysis, reduced the risk of repeated project establishment, and improved the scientificity and efficiency of scientific and technological project establishment management.
Smart Images

Figure CN114265935B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power systems, and in particular relates to a method and system for auxiliary decision-making in the establishment and management of scientific and technological projects based on text mining. Background Art
[0002] After literature research, it was found that there is no concept of project similarity assessment or duplication checking abroad, but research on big data mining and analysis started early, and a lot of research and exploration have been carried out, accumulating rich experience and mature technology; similarity assessment or duplication checking of scientific and technological projects is essentially a text similarity calculation method, involving key information extraction technology, word segmentation technology, text similarity calculation technology, etc., and similarity assessment or duplication checking of scientific and technological projects is affected by the development of these technologies.
[0003] Many foreign scholars have conducted extensive research on text similarity calculation and have achieved many results. It can be roughly divided into two stages: the first stage is mainly based on vector calculation and semantic calculation methods; the second stage is that in recent years, with the maturity of deep learning technology, more and more scholars have begun to study the calculation of text similarity based on self-learning methods.
[0004] Although China started late in the research of text mining methods, it has carried out targeted research on the application of text mining methods in scientific and technological project management. Jiang Shaohua proposed a scientific research project management prototype system based on text mining, focusing on studying and solving problems such as segmentation and feature modeling of scientific research project texts; Zuo Chuan proposed a method based on non-segmentation technology to solve the problem of duplicate checking of scientific and technological projects. This method does not require word segmentation of the text, and uses frequent closed item sets to construct a vector space model to model the project application and calculate the similarity; Fang Yanfeng proposed using an improved TF-IDF method for duplicate checking of scientific and technological projects, considering the position and length of feature words; Wu Yan proposed a scientific and technological project classification and duplicate checking method based on hierarchical clustering, which comprehensively considered factors such as application fields, research content and technical sources when calculating the similarity of scientific and technological projects; Lin Mingcai et al. proposed an improved fuzzy clustering algorithm RM-FCM, which considered the importance of feature items of different attributes to scientific research projects when calculating project similarity; Liu Yinming et al. studied the phenomenon of duplicate project establishment in my country's scientific research from the aspects of scientific and technological novelty search practice, multiple management of regions and departments, and the number of projects supported by scientific research papers. By analyzing the application and approval process of scientific research projects, specific measures to avoid duplicate project establishment were proposed.
[0005] With the continuous deepening of power reform and the continuous development of science and technology, more and more scientific and technological research projects and scientific and technological achievements in various professional categories are being reviewed, and the problem of duplicate project establishment has become increasingly serious. From the perspective of scientific and technological project establishment management, there are mainly the following problems: First, a large amount of unstructured data of scientific and technological projects is difficult to identify, and the identification of similarities of projects to be established consumes a lot of manpower and material resources; second, the comprehensive competitiveness of the applicants of scientific and technological projects is difficult to evaluate, and there is a lack of scientific competitiveness evaluation system for applicants; third, it is difficult to accurately recommend scientific and technological project review experts, and relying on manual selection of experts from the review expert database cannot guarantee the rationality of the selection of review experts; therefore, how to use cutting-edge technologies such as big data and artificial intelligence to solve the problems of multiple and duplicate project establishment in the current scientific and technological project establishment has become a key issue in improving the level of scientific and technological project establishment management in power supply bureaus. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide a method and system for auxiliary decision-making in project establishment management of scientific and technological projects based on text mining, so as to reduce the subjective factors of manual screening and identification, and improve the efficiency and accuracy of project similarity analysis.
[0007] In order to solve the above technical problems, the present invention provides a method for assisting decision-making in the establishment and management of scientific and technological projects based on text mining, comprising:
[0008] Step S1, using information extraction technology to extract feature data from the database of pending science and technology projects and the database of historical science and technology projects, and constructing a science and technology project information database;
[0009] Step S2, performing hierarchical text similarity mining on the feature data to construct a multi-level and multi-dimensional scientific and technological project similarity comparison model;
[0010] Step S3, obtaining the similarity scores between the project to be reviewed and other projects in the feature data, and iterating and updating the weights of the feature data using a grid search method on a historical sample training set to obtain a set of optimal weights;
[0011] Step S4, calculating a comprehensive score of similarities between the project to be reviewed and other projects according to the optimal weight.
[0012] Furthermore, the characteristic data includes title, keywords, project summary, purpose and significance, research background, main research content, and expected goals.
[0013] Furthermore, the step S1 specifically includes:
[0014] Seven types of characteristic data, including title, keywords, project abstract, purpose and significance, research background, main research content, and expected goals, were extracted from the database of science and technology projects to be reviewed and the database of historical science and technology projects.
[0015] Clean the extracted feature data, remove useless characters, and process them in a unified format;
[0016] The word segmentation operation is performed by using a combination of Jieba word segmentation + power industry dictionary + stop word filtering;
[0017] Extract keywords, including research object keywords, title keywords, subject keywords and comprehensive keywords.
[0018] Furthermore, the extracting keywords further includes:
[0019] Use text topic network graph clustering to extract keywords, select the first n keywords, if the keyword exists in the historical research object keywords, then use it as the research object keyword of the project to be reviewed, otherwise select the first two words with the largest comprehensive feature value as the research object keywords of the project to be reviewed;
[0020] The textrank method is used to extract keywords from the review items, where the keyword's part of speech is one of common nouns, professional nouns, institutional groups, organization names, and work names;
[0021] The historical science and technology projects were classified by manual annotation, and the SVM model was used for multi-label classification training to obtain the classification of the subject keywords of the projects to be reviewed;
[0022] The keywords extracted by textrank and topic network graph clustering were merged 1:1 to obtain comprehensive keywords for subsequent keyword similarity comparison.
[0023] Furthermore, the step S2 includes using an improved similarity calculation method based on edit distance to calculate the similarity of the project names, which specifically includes:
[0024] Step S21, assuming there is a string s 1 and 2 , let the input string be s 1i and 2j , use the algorithm to find the longest common substring of the two input strings, the result is l s ;
[0025] Step S22, if l s The length of s is greater than 2, then 1i and 2j Do the following: remove l s , and when l s When at the beginning or end of a string, split the string into two independent strings, s 1i1 、s1i2 and 2j1 、s 2j2 Otherwise, put s 1i Merge in order into the initially empty result string s a In the 2j Merge into the result string s in order b middle;
[0026] Step S23, traverse s 1i and 2j After the split string, continue to recursively enter step S21 until all substrings are calculated; at this time, all the longest common substrings have been obtained from s 1 and 2 Removed from s, the result is stored in s a and b middle;
[0027] Step S24, a and b Calculate the edit distance and use the edit distance similarity calculation formula to calculate the similarity:
[0028]
[0029] Among them, sim(s 1 ,s 2 ) means s 1 and 2 Similarity, ED represents the edit distance, len(s 1 ) represents the string s 1 Length.
[0030] Furthermore, the step S2 includes using the Doc2vec model in deep learning to obtain a long text vector, and using it to calculate the long text similarity; the calculation of the long text similarity includes calculating the similarity of the long text at the keyword level, sentence level and paragraph level.
[0031] Furthermore, calculating the long text keyword level specifically includes:
[0032] Extracting long text keywords w through text topic network graph clustering method 1 , w 2 ,......w n , use the trained word2vec model to perform word embedding mapping and obtain the word embedding vector w corresponding to each word n =(x 1 ,x 2 ,......x m ), n is the nth word, m is the mth feature, and then cosine similarity is used to calculate w 1 =(x1 ,x 2 ,......x m ), w 2 =(y 1 ,y 2 ,......y m ) are related to:
[0033]
[0034] Keyword D for two long texts 1 =(w 11 ,w 12 ,...w 1a ),D 2 =(w 21 ,w 22 ,...w 2b ), D 1 and D 2 The word-level similarity between is calculated using the following formula:
[0035]
[0036] Among them, w 1k ,w 2l Indicates the keywords of long text 1 and long text 2, sim(w 1k ,w 2l ) represents w calculated by cosine similarity 1k ,w 2l The similarity between .
[0037] Furthermore, calculating sentence level similarity specifically includes:
[0038] The common word statistics method is used to calculate similarity, and long text 1 and long text 2 are divided into a set of sentences D according to the textrank sentence granularity. 1 =(s 11 ,s 12 ,...s 1n ), D 2 =(s 21 ,s 22 ,...s 2m ), and the importance of each sentence in the long text is as follows:
[0039] D 1 ={s 11 :w 11 ,s 12 :w 12 ,...s 1n :w 1n ), D 2={s 21 :w 21 ,s 22 :w 22 ,...s 2m :w 2m )
[0040] Among them, w 11 +w 12 +...w 1n =1,w 21 +w 22 +...w 2m = 1, use the word segmentation to segment each sentence to obtain the sentence word set s = (w 1 ,w 2 ,...w a ), use the following formula to calculate the similarity between two sentences:
[0041]
[0042] That is, the ratio of the number of common words between two sentences to the number of all words between the two sentences is used as the sentence similarity, so the paragraph similarity calculation formula corresponding to the sentence level is as follows:
[0043]
[0044] Among them, w 1k represents the weight of the k-th sentence, max(sim(s 1k ,s 2l )) represents the sentence score of sentence k in long text 1 with the highest similarity in long text 2.
[0045] Furthermore, calculating the paragraph-level similarity specifically includes: using the doc2vec model to map the long text content into a high-dimensional vector, and then calculating the paragraph-level long text similarity by the cosine formula.
[0046] Furthermore, the step S3 specifically includes:
[0047] Determine the initial weights and fluctuation ranges of the seven characteristic data, namely title, keywords, project summary, purpose and significance, research background, main research content, and expected goals, and then update the weights through grid search. The specific process is as follows:
[0048] The weights of the seven feature data are divided into 50 or more parts, with the lowest being 0 and the highest being 1;
[0049] Circularly combine the seven weights, calculate the similarity accuracy of each project group under each weight combination, and select the group of weights with the highest accuracy as the update weights, where the project to be reviewed and the historical project are one project group;
[0050] Furthermore, the step S4 specifically includes: calculating the total similarity scores of all historical projects corresponding to the project to be reviewed according to the determined weights, arranging the total similarity scores of all data in descending order, selecting the similarity score values of the first three positions of each project to be reviewed and taking the average as the high and low threshold dividing line s high , take the value of the 5% position of the total similarity score as the low and medium threshold dividing line s low ; Similarity score higher than s high That is, high similarity, in s low With s high The similarity between them is medium, and the similarity below s low That is low similarity.
[0051] The present invention also provides a scientific and technological project establishment management auxiliary decision-making system based on text mining, comprising:
[0052] An information extraction module is used to extract feature data from the database of pending science and technology projects and the database of historical science and technology projects using information extraction technology to construct a science and technology project information database;
[0053] A similarity mining module is used to perform hierarchical text similarity mining on the feature data and construct a multi-level and multi-dimensional scientific and technological project similarity comparison model;
[0054] A weight determination module is used to obtain the similarity scores of the project to be reviewed and other projects in the feature data, and to iteratively update the weights of the feature data using a grid search method on a historical sample training set to obtain a set of optimal weights;
[0055] The calculation module is used to calculate the comprehensive score of the similarity between the project to be reviewed and other projects according to the optimal weight.
[0056] The implementation of the present invention has the following beneficial effects: based on relevant text data such as science and technology project application materials, the present invention uses artificial intelligence technologies such as Word2Vec, ELMO and Doc2Vec, combined with Chinese word segmentation, entropy value and hierarchical analysis methods, to carry out similarity analysis of science and technology projects, quantitative comparative analysis of science and technology project funds and contents, evaluation of the competitiveness of applicants, precise recommendation of review experts, and analysis of the use of research results of award-winning projects. Based on the research results, the present invention develops and implements auxiliary decision-making applications for science and technology project management, assists science and technology management departments in project establishment and reward review stages, supports the innovation of the company's science and technology project establishment and reward review management model, and ensures that the quality and efficiency of science and technology project establishment and reward review management work are improved;
[0057] The present invention studies the similarity analysis model of scientific and technological projects based on the text materials related to project establishment, realizes the comprehensive analysis of the similarities between the projects to be reviewed and those under construction, other projects to be evaluated and historical projects in multiple dimensions, reduces the subjective factors of manual screening and identification, and solves the problems of low efficiency and accuracy of similarity analysis of projects manually compared by professionals in the past; performs comprehensive text similarity calculation from various content modules of scientific and technological projects, and avoids the problem of inaccurate similarity analysis caused by the single use of keyword matching search. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0059] Figure 1 The present invention is a flowchart of a method for assisting decision-making in the establishment and management of scientific and technological projects based on text mining according to an embodiment of the present invention.
[0060] Figure 2 It is a schematic diagram of the process of similarity comparison of scientific and technological projects in an embodiment of the present invention.
[0061] Figure 3a-3e Schematic diagram of feature data extraction in an embodiment of the present invention, wherein: Figure 3a Extract schematics for titles and project summaries, Figure 3b Extracting diagrams for purpose and meaning, Figure 3c Schematic diagrams were extracted for research purposes. Figure 3d This is a schematic diagram for extracting the main research content + subheadings. Figure 3e Extract schematic diagram for the intended target.
[0062] Figure 4 Schematic diagram of a text topic network in an embodiment of the present invention.
[0063] Figure 5 Schematic diagram of word2vec framework in an embodiment of the present invention.
[0064] Figure 6 Schematic diagram of the PV-DM framework in an embodiment of the present invention.
[0065] Figure 7 Schematic diagram of the PV-DBOW framework in an embodiment of the present invention.
[0066] Figure 8 This is the AUC curve diagram of the similarity and dissimilarity binary classification of the present invention. DETAILED DESCRIPTION
[0067] The following descriptions of the embodiments refer to the accompanying drawings to illustrate specific embodiments in which the present invention may be implemented.
[0068] This invention will combine the research results of predecessors, integrate theoretical research and practical application needs, and build a technology project establishment management auxiliary decision-making system based on the historical data of science and technology projects, using big data and natural language processing technology, to assist the management work of the project establishment stage of the science and technology management department, support the innovation of the company's science and technology project establishment review management model, and ensure the quality and efficiency of all aspects of science and technology project establishment management. Figure 1 As shown, the first embodiment of the present invention provides a method for assisting decision-making in the establishment and management of scientific and technological projects based on text mining, including:
[0069] Step S1, using information extraction technology to extract feature data from the database of pending science and technology projects and the database of historical science and technology projects, and constructing a science and technology project information database;
[0070] Step S2, performing hierarchical text similarity mining on the feature data to construct a multi-level and multi-dimensional scientific and technological project similarity comparison model;
[0071] Step S3, obtaining the similarity scores between the project to be reviewed and other projects in the feature data, and iterating and updating the weights of the feature data using a grid search method on a historical sample training set to obtain a set of optimal weights;
[0072] Step S4, calculating a comprehensive score of similarities between the project to be reviewed and other projects according to the optimal weight.
[0073] Specifically, please combine Figure 2 As shown, in this embodiment, the feature data includes seven types of data: title, keywords, project abstract, purpose and significance, research background, main research content, and expected goals; step S2 performs hierarchical text similarity mining on keywords and main research contents, and constructs a multi-level and multi-dimensional similarity comparison model for scientific and technological projects. The specific algorithms involved include long text similarity calculation and short text similarity calculation. The long text similarity calculation includes similarity comparisons at the long text keyword level, sentence level, and paragraph level. The short text similarity calculation uses continuous common substrings + edit distance for similarity calculation; step S3 obtains the similarity scores of the seven types of feature data, namely title, keywords, project abstract, purpose and significance, research background, main research content, and expected goals, between the project to be reviewed and other projects, and uses a grid search method on the historical sample training set to iterate the weights of the seven types of feature data to obtain a set of optimal weights; step S4 uses the set of optimal weights to calculate the comprehensive score of the similarity between the project to be reviewed and other projects.
[0074] Step S1 performs text information extraction. Since much of the input technology project data is in doc format, and doc format files cannot read information well, it is necessary to first convert doc files into docx files. Technology projects in different periods have different project structures and contents, so it is necessary to unify the comparison content. The similarity comparison of two technology projects requires comparison from all aspects, dimensions, and contents.
[0075] 1.1.1 Content Extraction
[0076] The present invention extracts seven kinds of feature data, including title, keywords, project abstract, purpose and significance, research background, main research content (including technical route) and expected goal, and performs specific similarity comparison. According to different extraction rules corresponding to the project structure of different periods, the important parts of each project are extracted and put into the information database. Because the given data belong to different years, many projects have different content structures, which makes it impossible to use the same extraction template for information extraction. Therefore, information extraction templates corresponding to various content structures are designed and combined to automatically extract information from different types of projects. The specific project extraction part is as follows: Figure 3a-3e As shown in the figure (the object in the box is the object to be extracted). The code explains how to extract text input for new projects in the future.
[0077] 1.1.2 Data Cleansing
[0078] Since there are useless characters (including spaces, carriage returns, etc.) and some messy format writing in the text data of scientific and technological projects, it will interfere with the subsequent keyword extraction and subsequent similarity calculation. Therefore, the read original document data is processed in a unified format, such as converting traditional Chinese to simplified Chinese, converting full-width to half-width, removing spaces, removing redundant and useless words, etc., to clean the text and provide high-quality data for the following tasks.
[0079] 1.1.3 Word segmentation
[0080] Considering the efficiency of word segmentation and the effect of proper nouns, the combination of Jieba word segmentation + power industry dictionary + stop word filtering is used to perform word segmentation on the total content of the scientific and technological project. At the same time, the word parts are screened, and the required parts of speech include: common nouns (n), professional nouns (nz), institutional groups (nt), organization names (ORG), and work names (nw). These parts of speech will be of great help to the keyword extraction module.
[0081] 1.1.4 Keyword Extraction
[0082] The keywords of scientific and technological projects can reflect the purpose of scientific and technological projects to a certain extent. A multi-dimensional model is constructed for keyword extraction, and the keywords are divided into the following four parts: research object keywords, title keywords, theme keywords and comprehensive keywords (keywords extracted from the full text content). For example, the 1036 scientific and technological projects given are manually screened, and the screening data is shown in Table 1, where project_name is the name of the scientific and technological project of the screened data, and the project classification column is the classification of the project obtained by information extraction. The last three columns of label content, research object and label theme are the model training sample sets obtained by manual screening.
[0083] Table 1 Example of manual keyword screening
[0084]
[0085]
[0086] The specific process includes the following:
[0087] 1.1.4.1Textrank gets keywords:
[0088] TextRank was proposed by Mihalcea and Tarau in EMNLP'04. The idea is very simple: build a network through the adjacent relationship between words, and then use PageRank to iteratively calculate the rank value of each node. The keywords can be obtained by sorting the rank values. The algorithm used by TextRank for keyword extraction is as follows:
[0089] Split the given text T into complete sentences, namely:
[0090] T=[S 1 , S 2 , …, S m ]
[0091] For each sentence Si belonging to T, perform word segmentation and part-of-speech tagging, filter out stop words, and only keep words with specified parts of speech, such as nouns, verbs, and adjectives, that is:
[0092] S i =[t i,1 ;t i,2 ,…;t i,n ]
[0093] where t i,j It is the candidate keyword after retention.
[0094] Construct a candidate keyword graph G = (V, E), where V is a node set consisting of the generated candidate keywords, and then use the co-occurrence relationship to construct an edge between any two points. An edge exists between two nodes only if their corresponding words co-occur in a window of length K, where K represents the window size, that is, at most K words co-occur.
[0095] According to the above formula, the weights of each node are iteratively propagated until convergence.
[0096] The node weights are sorted in reverse order to obtain the most important T words as candidate keywords.
[0097] The most important T words are marked in the original text. If adjacent phrases are formed, they are combined into multi-word keywords. Each sentence in the text is regarded as a node. If two sentences are similar, it is considered that there is an undirected weighted edge between the nodes corresponding to the two sentences. The method to examine the similarity of sentences is the following formula:
[0098]
[0099] where s i ,s j The numerator is the number of words in the two sentences, and the denominator is the sum of the logarithms of the number of words in the sentences. This design of the denominator can curb the advantage of longer sentences in similarity calculation.
[0100] By cleaning several parts of the extracted long text data of scientific and technological projects, we obtain relatively clean long text data, segment them and use textrank to build an overall word graph. The importance of each word is calculated by calculating the correlation between each word and other words, and then all the words are ranked according to the importance scores. The topN words are selected as textrank to obtain the important words in this project.
[0101] 1.1.4.2 Keyword Extraction Based on Text Topic Network
[0102] Compared with textrank, the keyword extraction method based on text topic network uses graph clustering related methods to extract keywords after constructing the word graph. Specifically, a text topic network G is used to represent a text D, that is, the text topic network G can represent the theme of text D. The topic network of the entire text is represented by such a language network, and the entire text D is represented by a series of topic connected subgraphs. The central high-frequency words in the connected subgraph and the relatively low-frequency words connecting the two subgraphs are the words that play a key role in G and can be used to characterize the characteristics of the text, such as Figure 4 The figure shows the text topic network diagram, where the central words b, d, g and the connecting word f are the characteristic words of G.
[0103] The text topic network is defined as follows: Text topic network G = (V, E), where V = {V i |i=1,2,...n} represents vertex combination (for example, each word after data segmentation), E={(v i ,v j )|v i ,v j ∈V} is the edge set of the text topic network. In the process of extracting keywords, we combine the clustering properties to find important words that meet the requirements. Define the text topic network node v i Degree D i =|{(v i ,v j ):(v i ,v j )∈E,v i ,v j ∈V}|, node v i The degree of aggregation
[0104] K i =|{(v j ,v k ):(v i ,v j )∈E,(v j ,v k )∈E,v i ,v j ,v k ∈V}|
[0105] Therefore, node v i The clustering coefficient can be calculated according to the following formula:
[0106]
[0107] According to graph theory, the degree of a node in a graph represents the association with the node, which is measured by the number of edges with the node. The size of the node clustering reflects the density of the nodes around the node. Combined with clustering theory, the clustering coefficient reflects the proportion of a node in the shortest path between any two nodes. Here, we define v i The clustering coefficient where g(i) jk Indicates that in the text topic network, through node v i Connect node v j and v k The number of shortest paths, g jk Then it means connecting node node v jand v k The total number of shortest paths. Based on the above theory, the comprehensive feature value of each node in the text network graph is calculated according to the following formula:
[0108]
[0109] Calculate the comprehensive feature weight corresponding to each node (i.e. word), and arrange the feature weight CF in descending order. The larger the CF value, the greater its semantic relevance to the text. Take the topN nodes as the important words of the text for downstream tasks.
[0110] 1.1.4.3 Keyword extraction implementation
[0111] In order to more comprehensively describe the main purpose of the scientific and technological project content, corresponding keywords will be extracted from four aspects: research object, title, theme and full text. However, the source of keyword texts in each dimension is different, so different methods are used to extract keywords.
[0112] 1.1.4.3.1 Keywords of research object
[0113] A partial comparison of the effects of extracting important words using TextRank and using text topic graph clustering is shown in Table 2. The first column of research objects is the manually selected keywords of research objects of scientific and technological projects. It can be seen that the method of using text topic graph clustering can better extract the keywords of research objects of scientific and technological projects.
[0114] Table 2 Comparison of keyword extraction algorithm effects
[0115]
[0116]
[0117] Text topic network graph clustering is used to extract the corresponding project keywords. The first n keywords are selected. If the keyword exists in the historical research object keywords, it will be used as the research object keyword of this project. Otherwise, the first two words with the largest comprehensive characteristic value are selected as the research object keywords of this project.
[0118] 1.1.4.3.2 Title Keywords
[0119] Title keywords are the most intuitive main information of scientific and technological projects. Usually, the general research content of a project will exist in its title, so the textrank method is used to extract keywords in scientific and technological projects. The keywords must meet certain part-of-speech requirements, that is, the part-of-speech needs to be one of common nouns (n), professional nouns (nz), institutional groups (nt), organization names (ORG), and work names (nw).
[0120] 1.1.4.3.3 Theme Keywords
[0121] The subject keywords are the research topics in Table 1. We use manual labeling to classify historical scientific and technological projects into 12 categories: lightning, wind and fire disaster protection, risk assessment, information security protection, energy saving, electricity theft, decision support, monitoring and alarm, status diagnosis, testing technology, research and development, data management, and status evaluation. We also use the SVM model for multi-label classification training to obtain the classification of the subject keywords of the projects to be reviewed.
[0122] 1.1.4.3.4 Comprehensive Keywords
[0123] After experimental comparative analysis, it was found that using textrank and topic network graph clustering to extract keywords has a better effect, so this embodiment merges the keywords extracted by the two methods 1:1 to obtain comprehensive keywords for subsequent keyword similarity comparison.
[0124] Through model training, we obtain a set of optimal weights for four levels: research object, title, theme, and comprehensive keywords (the weights obtained from the current sample data are 0.12, 0.04, 0.02, and 0.82, respectively).
[0125] 1.2 Similarity comparison
[0126] 1.2.1 Short text similarity comparison
[0127] Short text refers to text data with fewer words, such as titles of scientific and technological projects and subtitles of main research contents. Compared with long texts such as research contents, short texts contain less and more concentrated information. In addition, most of the terms in power technology projects are relatively professional. It is not appropriate to use keywords alone for comparison. Therefore, an improved similarity calculation method based on edit distance (continuous common substring + edit distance (ed)) is used to calculate the similarity of project names.
[0128] 1.2.1.1 Edit distance
[0129] Edit distance is a measure of the similarity between two strings, which indicates the minimum number of operations required to convert one string into the other. This concept was proposed by Russian scientist Vladimir Levenshtein in 1965. Edit distance is widely used in fast fuzzy matching of strings and is a good method for calculating sentence similarity.
[0130] Edit distance: refers to the minimum number of edits required to convert one substring into another. Edit operations include deletion, insertion, replacement, etc. Edit distance can be expressed as:
[0131]
[0132] Where D(str1,str2,i,j) represents the edit distance between the first i characters of string str1 and the first j characters of string str2. i Represents the ith substring of string str1. The initial value of D(str1,str2,0,0) is 0.
[0133] The above formula is a recursive definition. If there are strings s1 and s2, with lengths of m and n respectively, a matching relationship matrix of order (m+1)*(n+1) is generally used to calculate the edit distance. The element values in the matrix are:
[0134]
[0135] where d i,j It represents the value of the i-th row and j-th column in the matrix. The following is an example of a matching relationship matrix. The edit distance between "big data application" and "application big data" is calculated. The edit distance is 4, as shown in Table 3:
[0136] Table 3 Edit distance calculation matrix
[0137] big number according to answer use answer 1 2 3 3 4 use 2 2 3 4 3 big 2 3 3 4 4 number 3 2 3 4 5 according to 4 3 2 3 4
[0138] 1.2.1.2 Improved topic similarity calculation
[0139] Studying and observing the names in the science and technology project applications, we can find the following characteristics:
[0140] There are many professional words in the title and they are all combined into long words. They are not simple professional words that can be segmented. For example, "Research and Application of Equipment Visualization Monitoring Model Based on Big Data Acceleration Analysis and Three-dimensional Digitalization". The meaning of "big data acceleration analysis" and "equipment visualization detection model" has changed after being simply segmented into "big data", "acceleration", "analysis", "equipment", "visualization", "detection" and "model".
[0141] It is difficult to understand the semantics of professional titles. For example, "Research on key technologies and development models of source-end integrated energy systems" and "Research on multi-energy conversion simulation and comprehensive energy efficiency evaluation technologies of integrated energy systems" are similar in semantic understanding, but simply using the edit distance will result in a very low score.
[0142] The names of scientific and technological projects are relatively short, with the longest being about 30 characters and the shortest being only 10 characters.
[0143] Since the names of scientific and technological projects contain a large number of professional names, they are often combined together to form longer words. For two project names, if there are many repeated professional terms in the two names, then the possibility of the two projects being similar is very high. However, if the edit distance is directly used to calculate, the similarity may be very low. Based on this, it is proposed to remove all the longest continuous common substrings in the string (such as the longest common substring of "Research on Key Technologies and Development Models of Integrated Energy Systems for Source-End Bases" and "Research on Multi-Energy Conversion Simulation and Comprehensive Energy Efficiency Evaluation Technologies for Integrated Energy Systems" is "Integrated Energy Systems") before calculating the edit distance. Assume that there is a string s 1 and 2 , the calculation process of the improved algorithm is as follows:
[0144] Step S21, assuming the input string is s 1i and 2j , use the algorithm to find the longest common substring of the two input strings, the result is l s .
[0145] Step S22, if l s The length of s is greater than 2, then 1i and 2j Do the following: remove l s , and split the string into two parts (when l s At the beginning or end of a string) an independent string, s 1i1 、s 1i2 and 2j1 、s 2j2 Otherwise, put s 1i Merge into the result string s in order a (Initially empty), put s 2j Merge into the result string s in order b middle.
[0146] Step S23, traverse s 1i and 2j The segmented string continues to recursively enter step S21 until the calculation of all substrings is completed.
[0147] At this time, all the longest common substrings have been 1 and 2 Removed from s, the result is stored in s a and b middle.
[0148] Step S24,a and b Calculate the edit distance (ED), and then use the edit distance similarity calculation formula to calculate the similarity. The specific formula is as follows:
[0149]
[0150] Where sim(s 1 ,s 2 ) means s 1 and 2 Similarity, ED represents the edit distance, len(s 1 ) represents the string s 1 Length.
[0151] Some scientific and technological projects were randomly selected for project name similarity calculation using the original algorithm (single edit distance calculation) and the improved algorithm (longest common substring + edit distance), and the comparison results are shown in Table 4. It can be seen that the edit distance of the improved algorithm is relatively small, and the similarity value is higher. Compared with the original algorithm, the improved algorithm is more consistent with the real similarity value.
[0152] Table 4 Name similarity comparison results under different algorithms
[0153]
[0154]
[0155] Note: ED stands for edit distance, sim stands for similarity
[0156] Short text calculation is mainly the calculation and comparison between project titles and main research content subtitles. By splitting the main research content into full-content long text and subtitle short text, the main research content of two projects can be compared more comprehensively and specifically, especially in the subtitle comparison, which can achieve a relatively ideal effect. For example, if the main content subtitle of project A is similar to the project title or main content subtitle of project B, then A and B may have more or less similarity. This can be used as the basis for judging similar projects, and similar projects can be screened from more detailed aspects.
[0157] 1.2.2 Long text similarity comparison
[0158] 1.2.2.1 Long text similarity calculation
[0159] For the similarity calculation of unsupervised long texts, the basic direction is to vectorize the text and then determine the similarity value by calculating the distance between two item vectors. The commonly used methods are as follows:
[0160] bag of words
[0161] LDA (Latent Dilitre Allocation)
[0162] Average word vectors
[0163] Tfidf-weighting word vectors (average of word vectors with tfidf weights)
[0164] Among them, the bag-of-words model does not take into account the order of words and ignores the semantic information of words; LDA mainly calculates the topic distribution of a document or a sentence; the word vector average model first trains the word2vec / bert word vector, and simply averages all the words in the sentence paragraph. This is the most effective and simple way, but the obvious disadvantage is that it does not take into account the order of words; the word vector average with tfidf weight is to sum all the word vectors in the sentence according to the tfidf weight. It is a commonly used method for calculating long text vectors. Compared with simply averaging all long text vectors, it considers the use of tfidf weights. Therefore, the proportion of more important words in the sentence is larger, but the order of words is not considered. Compared with the above methods, the Doc2vec model not only considers the order of words but also contains semantic information. The Doc2vec model in deep learning is used to obtain long text vectors and is used to calculate long text similarity.
[0165] 1.2.2.2 Doc2vec
[0166] Doc2vec (paragraph2vec) is an unsupervised algorithm that can obtain vector representations of sentences / paragraphs / long documents. It is an extension of word2vec. The framework of word2vec is as follows: Figure 5 shown.
[0167] There are two modes for Word2vec training: CBOW and Skip-gram. Figure 6 INPUT, PROJECTION, and OUTPUT in the above code represent the input layer, hidden layer, and output layer, respectively. Taking CBOW as an example, each word is mapped into the vector space, and the word vectors of the context are concatenated or summed in a window of a specific length as features to predict the next word in the sentence. For example, the word sequence is 'carry out', 'big data', 'accelerate', 'analyze' to predict 'based on', and the objective function is:
[0168]
[0169] Where J(θ) represents the objective function we need to train, w trepresents the tth word, k represents the size of a window, k=2 means the context length is 2, and T represents the number of all words predicted in a sentence.
[0170] The prediction task is a classification problem. The last layer of the classifier uses softmax, and the calculation formula is as follows:
[0171]
[0172] Where i is the number of words in the vocabulary, y i is the predicted value of the i-th word, and ywt is the predicted value of the core word at time t to be predicted. Each word is considered as a category, y i The calculation formula is as follows:
[0173] y=b+Uh(w t-k ,...,w t+k ; W)
[0174] Among them, U, b are softmax calculation parameters, and h is the value of w t-k ,...,w t+k Each word vector is concatenated or averaged. Since each word is considered as a category in the algorithm, the number of categories is very large and the training efficiency is very low. Therefore, when normalizing Word2vec, hierarical softmax and Negative Sampling are used to speed up the calculation. Here we introduce Negative Sampling, as follows:
[0175] The core idea of Negative Sampling is to replace the central word of a word string in the corpus with another word, and construct a word string that does not exist in the corpus D as a negative sample. Under this strategy, the optimization goal becomes: maximize the probability of positive samples while minimizing the probability of negative samples. A word string (w, c) (for skip-gram, c represents the central word of w, for CBOW, c represents the context of w), uses a binomial logistic regression model to model the probability of its positive sample:
[0176]
[0177] So the likelihood function of all positive samples is:
[0178]
[0179] Similarly, the likelihood function of all negative samples is:
[0180]
[0181] It is necessary to maximize the former and minimize the latter, that is, to maximize the following formula:
[0182]
[0183] Take the log-likelihood:
[0184]
[0185] Since SGD is used, we only need to know the objective function for a positive sample (ω, c). Where NEG(ω) is the central word set of negative samples of (ω, c):
[0186]
[0187] This greatly optimizes the Word2vec normalization efficiency.
[0188] The core idea of training word vectors is to predict based on the context of each word, that is, the word pairs in the context have an impact. Similarly, the same method can be used to train Doc2vec, where Doc2vec has two modes: A distributed memory model and Paragraph Vector without word ordering: Distributed bag of words.
[0189] A distributed memory model
[0190] like Figure 5The figure shows the framework of Doc2vec PV-DM. It can be seen from the figure that in addition to the word-level vector, there is also a vector representation for each paragraph / sentence. For example, for a sentence 'the cat sat on', if you want to predict the word on in the sentence, you can not only generate corresponding features based on other words, but also generate features based on other words and sentences for prediction. Each paragraph / sentence is mapped to the vector space and can be represented by a column of the matrix. Each word is also mapped to the vector space and can be represented by a column of the matrix. Then the paragraph vector and the word vector are cascaded or averaged to obtain features to predict the next word in the sentence. The paragraph vector / sentence vector can also be considered as a word. Its function is equivalent to the memory unit of the context or the theme of this paragraph, so we generally call this training method Distributed Memory Model of Paragraph Vectors (PV-DM). During training, the context length is fixed, and the training set is generated by the sliding window method. And the paragraph / sentence vector is shared in the context. The specific Doc2vec process mainly has two steps:
[0191] Train the model to obtain word vectors, softmax parameters, and paragraph vectors / sentence vectors from known training data.
[0192] In the inference stage, for a new paragraph, its vector expression is obtained. Specifically, more columns are added to the matrix, and the above method is used for training under a fixed length, and the new D (paragraph vector matrix) is obtained using the gradient descent method, thereby obtaining the vector expression of the new paragraph.
[0193] Paragraph Vector without word ordering: Distributed bag of words (distributed bag of words model)
[0194] like Figure 7 The figure shows the framework of Doc2vec PV-DBOW. Compared with the distributed bag of words model method to train the model to obtain the paragraph vector, there is another method that ignores the input context and lets the model predict a random word in the paragraph. Here, only the paragraph vector is input, but the prediction is for all the words in the paragraph / sentence. This method is similar to skip-gram in Word2vec, called Distributed Bag of Words version of Paragraph Vector (PV-DBOW). To compare the two training methods, we use the PV-DM method for training.
[0195] 1.2.2.3 Word-level similarity
[0196] The similarity calculation of words is set for long texts, and the long text keywords w are extracted through the text topic network graph clustering method. 1 , w 2 ,......w n , use the trained word2vec model to perform word embedding mapping and obtain the word embedding vector w corresponding to each word n =(x 1 ,x 2 ,......x m ), where n is the nth word, m is the mth feature (m is 300), and then cosine similarity is used to calculate w 1 =(x 1 ,x 2 ,......x m ), w 2 =(y 1 ,y 2 ,......y m ), the cosine similarity calculation formula is as follows:
[0197]
[0198] Keyword D for two long texts 1 =(w 11 ,w 12 ,...w 1a ),D 2 =(w 21 ,w 22 ,...w 2b ), D 1 and D 2 The word-level similarity between can be calculated using the following formula:
[0199]
[0200] Among them, w 1k ,w 2l Indicates the keywords of long text 1 and long text 2, sim(w 1k ,w 2l ) represents w calculated by cosine similarity 1k ,w 2l The similarity between .
[0201] 1.2.2.4 Sentence Level Similarity
[0202] For sentence-level similarity comparison, considering the actual efficiency issue, the common word statistics method is used to calculate the similarity, and the long text 1 and the long text 2 are divided into a set of sentences D according to the textrank sentence granularity. 1 =(s 11 ,s 12 ,...s 1n ), D 2 =(s 21 ,s 22 ,...s 2m ), and get the importance of each corresponding sentence in the long text, as follows:
[0203] D 1 ={s 11 :w 11 ,s 12 :w 12 ,...s 1n :w 1n ), D 2 ={s 21 :w 21 ,s 22 :w 22 ,...s 2m :w 2m )
[0204] where w 11 +w 12 +...w 1n =1,w 21 +w 22 +...w 2m = 1, use the word segmentation to segment each sentence to obtain the sentence word set s = (w 1 ,w 2 ,...w a ), use the following formula to calculate the similarity between two sentences:
[0205]
[0206] That is, the ratio of the number of common words between two sentences to the number of all words between the two sentences is used as the sentence similarity, so the paragraph similarity calculation formula corresponding to the sentence level is as follows:
[0207]
[0208] Among them, w 1k represents the weight of the k-th sentence, max(sim(s 1k ,s 2l )) represents the sentence score of sentence k in long text 1 with the highest similarity in long text 2.
[0209] 1.2.2.5 Paragraph-level similarity
[0210] The doc2vec model is used to map long text content into high-dimensional vectors, and then the cosine formula is used to calculate the long text similarity at the paragraph level.
[0211] After model training, we get a set of optimal weights at the word, sentence, and paragraph levels (the weights obtained by the current sample data are 0.4, 0.12, and 0.48, respectively), and add them up to get the final similarity between long text 1 and long text 2. In addition, considering the importance of the "main research content" and the particularity of its structure, the similarity calculation of the "main research content" is processed separately, as follows:
[0212] The technical route content of scientific and technological projects reflects the innovativeness of their technical implementation methods, but there are many missing values, so the non-missing technical route content is merged into the main research content section.
[0213] In multiple experimental comparisons, it was found that the comprehensive keywords of the full text as the main research content keywords are more comprehensive and accurate, and have the best effect.
[0214] The sentence-level similarity method was used to calculate the similarity of the subheadings of the main research content, and by repeatedly adjusting the weights of long text words, sentences, paragraphs, and subheadings, a set of optimal results were obtained, which were 0.38, 0.1, 0.45, and 0.07, respectively.
[0215] 1.3 Weight determination
[0216] Since the extracted 7 parts have different importance, the weight value of each part and the thresholds of the three similarity levels of high, medium and low are determined according to different model algorithms.
[0217] First, based on experience, determine the initial weights of the seven parts of the scientific and technological project title, keywords, project abstract, purpose and significance, research background, main research content, and expected goals, as well as the fluctuation range of the weights of these parts. For example, the weight of the main research content is between (0.25, 0.4), and the weight of the project abstract is between (0.1, 0.25). Then, update the weights through grid search. The specific method is as follows:
[0218] 1. The weights of the seven parts (title, keywords, project summary, purpose and significance, research background, main research content, and expected goals) are divided into 50 parts (or more) with a minimum of 0 and a maximum of 1.
[0219] 2. Cycle through the seven weight combinations, calculate the topN similarity accuracy of each project group (the project to be compared and the historical project are a project group) under this weight combination, and select the group of weights with the highest accuracy as the update weights.
[0220] According to the determined weights, the total similarity scores of all historical projects corresponding to the project to be reviewed are calculated, and the total similarity scores of all data are arranged in descending order. The similarity score values of the first three positions of each project to be reviewed are selected and the average is taken as the high and low threshold dividing line s high , and set s high The value cannot be lower than 0.5; the value of the 5% position of the total similarity score is taken as the low and medium threshold dividing line s low . Similarity score higher than s high That is, high similarity, in s low With s high The similarity between them is medium, and the similarity below s low That is, low similarity. According to the training data, s high 0.377, s low It is 0.321.
[0221] 1.4 Results Evaluation
[0222] 1.4.1 TopN Test Evaluation
[0223] The present invention selects top5, top10, top15, and top20 as the research scope, selects the topN most similar projects of each project to be reviewed, and compares them with the label to be reviewed. If the real similar project (label to be reviewed) of the project to be reviewed exists in the topN similar documents of the project to be reviewed, then the comparison is correct, and the accuracy of the topN similarity of the project to be reviewed is calculated according to the following formula. Assuming that there are m projects to be reviewed:
[0224]
[0225] The specific topN test steps are as follows:
[0226] 1) The 128 training data sets are divided into training sets (109 sets) and test sets (19 sets) in a ratio of 17:3;
[0227] 2) Use grid search to determine the weights of 7 parts of the 109 training set data;
[0228] 3) Calculate the similarity scores between the 19 test sets and the other 1036 scientific and technological projects according to the determined weights, and calculate the topN accuracy of the top5, top10, top15, and top20 comparison results in turn according to the above formula;
[0229] 4) Repeat steps 1, 2, and 3 above 5 times, and take the average of the top5, top10, top15, and top20 accuracy scores obtained 5 times as the final evaluation accuracy.
[0230] According to the above steps, different combination strategies are used to calculate the similarity of various scientific and technological projects. The results are shown in Table 5.
[0231] Table 5 Statistics of topN similarity accuracy of projects
[0232]
[0233]
[0234] Note: The similarity of short texts is calculated using the improved edit distance;
[0235] The first column in Table 5 shows different strategies for similarity calculation of scientific and technological projects, among which 'full-text keywords' refers to keywords extracted from the full-text data of scientific and technological projects; 'layered keywords' refers to keywords extracted after layering the full-text data at different levels; 'keywords as a separate dimension' refers to a strategy that takes the project keyword dimension comparison out separately and establishes 7 parts of title, keywords, project summary, purpose and significance, research background, main research content, and expected goals for similarity calculation, while 'keywords not as a separate dimension' refers to a strategy that puts the project keyword dimension comparison into the main research content comparison and only forms 6 parts of title, project summary, purpose and significance, research background, main research content, and expected goals for similarity calculation. It can be seen from Table 5 that: first, the effect of the strategy of comparing project keywords as a separate dimension is generally higher than that of not taking it as a separate dimension; second, the layered keywords have a very obvious effect on the improvement, indicating that the layered keywords can capture the main content of the project well; third, it can be seen from the last few columns in the table that most of the similar projects are within the top 20, and a few projects are after the top 20.
[0236] 1.4.2 Similarity and Dissimilarity Test Evaluation
[0237] The AUC index is based on TP, FP, FN, and TN, as shown in Table 6. ROC (receiveroperating characteristic curve) is the receiver operating characteristic curve, which is used to predict the effect of a group of samples. Its horizontal axis is FP and its vertical axis is TP. AUC (Area Under Curve) is the area of the lower half of the ROC curve. The larger the area, the better the classification effect.
[0238] Table 6 Data positive example-negative example explanation table
[0239]
[0240] Randomly sample 128 similar and 128 dissimilar project groups, mark the corresponding labels as 0 (dissimilar) and 1 (similar), perform grid search on the 128 technology project training sets to obtain the corresponding 7-part weights, and calculate the similarity scores and corresponding AUC values of the 128 technology projects with other technology projects for accuracy comparison.
[0241] Specific results such as Figure 8 As shown:
[0242] Figure 8 The AUC curve for similarity and dissimilarity is shown in Figure 2. The ks curve corresponds to the difference between the true positive rate and the false positive rate. The AUC curve is the AUC value of 128 scientific and technological projects. The four coordinate values corresponding to the red dot represent the true positive rate, false positive rate, maximum difference between the true positive rate and the false positive rate (0.814) and the binary classification threshold of 0.385. At this time, the accuracy of the model is 0.955.
[0243] In summary, the multi-level and multi-dimensional scientific and technological project similarity comparison model of the present invention can accurately and effectively search for similar projects for most scientific and technological projects, providing effective assistance for project review.
[0244] Corresponding to the above-mentioned embodiment 1 of the present invention, which provides a method for assisting decision-making in the establishment and management of scientific and technological projects based on text mining, the embodiment 2 of the present invention further provides a system for assisting decision-making in the establishment and management of scientific and technological projects based on text mining, including:
[0245] An information extraction module is used to extract feature data from the database of pending science and technology projects and the database of historical science and technology projects using information extraction technology to construct a science and technology project information database;
[0246] A similarity mining module is used to perform hierarchical text similarity mining on the feature data and construct a multi-level and multi-dimensional scientific and technological project similarity comparison model;
[0247] A weight determination module is used to obtain the similarity scores of the project to be reviewed and other projects in the feature data, and to iteratively update the weights of the feature data using a grid search method on a historical sample training set to obtain a set of optimal weights;
[0248] The calculation module is used to calculate the comprehensive score of the similarity between the project to be reviewed and other projects according to the optimal weight.
[0249] For the working principle and process of this embodiment, please refer to the description of the first embodiment of the present invention, which will not be repeated here.
[0250] It can be seen from the above description that the beneficial effects brought by the present invention are: based on relevant text data such as science and technology project application materials, the present invention uses artificial intelligence technologies such as Word2Vec, ELMO and Doc2Vec, combined with Chinese word segmentation, entropy value and hierarchical analysis methods, to carry out similarity analysis of science and technology projects, quantitative comparative analysis of science and technology project funds and contents, evaluation of the competitiveness of applicants, precise recommendation of review experts, and analysis of the use of research results of award-winning projects. Based on the research results, research and development realizes the application of auxiliary decision-making in science and technology project management, assists the management work of science and technology management departments in project establishment and reward review stages, supports the innovation of the company's science and technology project establishment and reward review management model, and ensures the improvement of quality and efficiency of science and technology project establishment and reward review management work;
[0251] The present invention studies the similarity analysis model of scientific and technological projects based on the text materials related to project establishment, realizes the comprehensive analysis of the similarities between the projects to be reviewed and those under construction, other projects to be evaluated and historical projects in multiple dimensions, reduces the subjective factors of manual screening and identification, and solves the problems of low efficiency and accuracy of similarity analysis of projects manually compared by professionals in the past; performs comprehensive text similarity calculation from various content modules of scientific and technological projects, and avoids the problem of inaccurate similarity analysis caused by the single use of keyword matching search.
[0252] The above disclosure is only the preferred embodiment of the present invention, which certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A text mining-based decision-making support method for scientific and technological project establishment management. It is characterized in that include: Step S1, using information extraction technology to extract feature data from the technology project database to be reviewed and the historical technology project database, and constructing a technology project information database; Step S2, performing hierarchical text similarity mining on the feature data to construct a multi-level and multi-dimensional scientific and technological project similarity comparison model; Step S3, obtaining the similarity scores between the project to be reviewed and other projects in the feature data, and iterating and updating the weights of the feature data using a grid search method on a historical sample training set to obtain a set of optimal weights; Step S4, calculating a comprehensive score of similarity between the project to be reviewed and other projects according to the optimal weight; The step S2 includes using the Doc2vec model in deep learning to obtain a long text vector, and using it to calculate the long text similarity; the calculation of the long text similarity includes calculating the similarity of the long text keyword level, sentence level and paragraph level; Calculating the long text keyword level specifically includes: Extracting long text keywords w through text topic network graph clustering method 1 , w 2 ,......w n , use the trained word2vec model to perform word embedding mapping and obtain the word embedding vector w corresponding to each word n =(x 1 ,x 2 ,......x m ), n is the nth word, m is the mth feature, and then cosine similarity is used to calculate w 1 =(x 1 ,x 2 ,......x m ), w 2 =(y 1 ,y 2 ,......y m ) are related to: Keyword D for two long texts 1 =(w 11 ,w 12 ,...w 1a ),D 2 =(w 21 ,w 22 ,...w 2b ), D 1 and D 2 The word-level similarity between is calculated using the following formula: Among them, w 1k ,w 2l Indicates the keywords of long text 1 and long text 2, sim(w 1k ,w 2l ) represents w obtained by cosine similarity calculation 1k ,w 2l The similarity between .
2. The method according to claim 1, It is characterized in that The characteristic data include title, keywords, project summary, purpose and significance, research background, main research content, and expected goals.
3. The method according to claim 2, It is characterized in that The step S1 specifically includes: Seven types of characteristic data, including title, keywords, project abstract, purpose and significance, research background, main research content, and expected goals, were extracted from the database of science and technology projects to be reviewed and the database of historical science and technology projects. Clean the extracted feature data, remove useless characters, and process them in a unified format; The word segmentation operation is performed by using a combination of Jieba word segmentation + power industry dictionary + stop word filtering; Extract keywords, including research object keywords, title keywords, subject keywords and comprehensive keywords.
4. The method according to claim 3, It is characterized in that The extracting keywords further includes: Use text topic network graph clustering to extract keywords, select the first n keywords, if the keyword exists in the historical research object keywords, then use it as the research object keyword of the project to be reviewed, otherwise select the first two words with the largest comprehensive feature value as the research object keywords of the project to be reviewed; The textrank method is used to extract keywords from the review items, where the keyword's part of speech is one of common nouns, professional nouns, institutional groups, organization names, and work names; The historical science and technology projects were classified by manual annotation, and the SVM model was used for multi-label classification training to obtain the classification of the subject keywords of the projects to be reviewed; The keywords extracted by textrank and topic network graph clustering were merged 1:1 to obtain comprehensive keywords for subsequent keyword similarity comparison.
5. The method according to claim 1, It is characterized in that The step S2 includes using an improved similarity calculation method based on edit distance to calculate the similarity of the project names, which specifically includes: Step S21, assuming there is a string s 1 and 2 , let the input string be s 1i and 2j , use the algorithm to find the longest common substring of the two input strings, the result is l s ; Step S22, if l s The length of s is greater than 2, then 1i and 2j Do the following: remove l s , and when l s When at the beginning or end of a string, split the string into two independent strings, s 1i1 、s 1i2 and 2j1 、s 2j2 Otherwise, put s 1i Merge in order into the initially empty result string s a In the 2j Merge into the result string s in order b middle; Step S23, traverse s 1i and 2j After the split string, continue to recursively enter step S21 until all substrings are calculated; at this time, all the longest common substrings have been obtained from s 1 and 2 Removed from s, the result is stored in s a and b middle; Step S24, a and b Calculate the edit distance and use the edit distance similarity calculation formula to calculate the similarity: Among them, sim(s 1 ,s 2 ) means s 1 and 2 Similarity, ED represents the edit distance, len(s 1 ) represents the string s 1 Length.
6. The method according to claim 1, It is characterized in that Calculating sentence-level similarity specifically includes: The common word statistics method is used to calculate similarity, and long text 1 and long text 2 are divided into a set of sentences D according to the textrank sentence granularity. 1 =(s 11 ,s 12 ,...s 1n ), D 2 =(s 21 ,s 22 ,...s 2m ), and the importance of each corresponding sentence in the long text is as follows: D 1 ={s 11 :w 11 ,s 12 :w 12 ,...s 1n :w 1n ),D 2 ={s 21 :w 21 ,s 22 :w 22 ,...s 2m :w 2m ) Among them, w 11 +w 12 +...w 1n =1,w 21 +w 22 +...w 2m = 1, use the word segmentation to segment each sentence to obtain the sentence word set s = (w 1 ,w 2 ,...w a ), use the following formula to calculate the similarity between two sentences: That is, the ratio of the number of common words between two sentences to the number of all words between the two sentences is used as the sentence similarity, so the paragraph similarity calculation formula corresponding to the sentence level is as follows: Among them, w 1k represents the weight of the kth sentence, max(sim(s 1k ,s 2l )) represents the sentence score of sentence k in long text 1 with the highest similarity in long text 2.
7. The method according to claim 6, It is characterized in that Calculating paragraph-level similarity specifically includes: using the doc2vec model to map the long text content into a high-dimensional vector, and then calculating the paragraph-level long text similarity using the cosine formula.
8. The method according to claim 7, It is characterized in that The step S3 specifically includes: Determine the initial weights and fluctuation ranges of the seven characteristic data, namely title, keywords, project summary, purpose and significance, research background, main research content, and expected goals, and then update the weights through grid search. The specific process is as follows: The weights of the seven feature data are divided into 50 or more parts, with the lowest being 0 and the highest being 1; The seven weights are cyclically combined to calculate the similarity accuracy of each project group under each weight combination, and the group of weights with the highest accuracy is selected as the update weights, wherein the project to be reviewed and the historical project are one project group.
9. The method according to claim 7, It is characterized in that The step S4 specifically includes: calculating the total similarity scores of all historical projects corresponding to the project to be reviewed according to the determined weights, arranging the total similarity scores of all data in descending order, selecting the similarity score values of the first three positions of each project to be reviewed and taking the average as the high and low threshold dividing line s high , take the value of the 5% position of the total similarity score as the low and medium threshold dividing line s low ; Similarity score higher than s high That is, high similarity, in s low With s high The similarity between them is medium, and the similarity below s low That is low similarity.
10. A technology project management decision-making support system based on text mining. It is characterized in that include: An information extraction module is used to extract feature data from the database of pending science and technology projects and the database of historical science and technology projects using information extraction technology to construct a science and technology project information database; A similarity mining module is used to perform hierarchical text similarity mining on the feature data and construct a multi-level and multi-dimensional scientific and technological project similarity comparison model; A weight determination module is used to obtain the similarity scores of the project to be reviewed and other projects in the feature data, and to iteratively update the weights of the feature data using a grid search method on a historical sample training set to obtain a set of optimal weights; A calculation module, used for calculating a comprehensive score of similarity between the project to be reviewed and other projects according to the optimal weight; The similarity mining module is specifically used to obtain long text vectors using the Doc2vec model in deep learning, and to calculate the long text similarity; the long text similarity calculation includes calculating the similarity of the long text keyword level, sentence level and paragraph level; Calculating the long text keyword level specifically includes: Extracting long text keywords w through text topic network graph clustering method 1 , w 2 ,......w n , use the trained word2vec model to perform word embedding mapping and obtain the word embedding vector w corresponding to each word n =(x 1 ,x 2 ,......x m ), n is the nth word, m is the mth feature, and then cosine similarity is used to calculate w 1 =(x 1 ,x 2 ,......x m ), w 2 =(y 1 ,y 2 ,......y m ) are related to: Keyword D for two long texts 1 =(w 11 ,w 12 ,...w 1a ),D 2 =(w 21 ,w 22 ,...w 2b ), D 1 and D 2 The word-level similarity between is calculated using the following formula: Among them, w 1k ,w 2l Indicates the keywords of long text 1 and long text 2, sim(w 1k ,w 2l ) represents w obtained by cosine similarity calculation 1k ,w 2l The similarity between .
Citation Information
Patent Citations
Recommendation method based on multi-dimensional alarm information text similarity analysis
CN111159387A
Intelligent scientific research project review method and storage medium
CN112329425A