Big data-based scientific and technological project information management method and system

By using big data analytics and knowledge graph construction, the shortcomings of traditional science and technology project information management in terms of correlation and risk analysis have been addressed. This has enabled precise optimization and risk identification for new projects, thereby improving research efficiency and information acquisition effectiveness.

CN121010216BActive Publication Date: 2026-05-12SUN YAT SEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUN YAT SEN UNIV
Filing Date
2025-08-12
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Traditional methods of information management for science and technology projects are inadequate in terms of information correlation analysis, risk analysis, and information retrieval. They are unable to accurately determine the similarities and differences between new and historical projects, and cannot provide comprehensive and accurate project risk analysis and retrieval results, thus reducing research efficiency and information acquisition efficiency.

Method used

基于大数据的科技项目信息管理方法,通过计算待处理项目与基准项目的关联度偏差系数,构建知识图谱,分析项目进度、成本和资源匹配偏差,进行风险分析,并基于关键词匹配度输出项目信息推荐列表。

Benefits of technology

It enables precise optimization of new projects, comprehensive identification of potential risks, and enhances the scientific and forward-looking nature of project management. It also improves user search efficiency and the utilization efficiency of scientific research resources, ensuring the smooth progress of scientific and technological projects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010216B_ABST
    Figure CN121010216B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of information management, and particularly discloses a scientific and technological project information management method and system based on big data, which comprises the following steps: S1, obtaining to-be-processed scientific and technological project information; S2, project information correlation deviation analysis; S3, constructing a scientific and technological project knowledge graph; S4, scientific and technological project information management risk analysis; S5, scientific and technological project information matching analysis; and S6, outputting a project information recommendation list. The application judges the correlation deviation of newly-added projects and performs optimization analysis, constructs a scientific and technological project knowledge graph, analyzes the management risk of scientific and technological project information, and further recommends scientific and technological project information to users. The application realizes accurate management and efficient utilization of scientific and technological project information by means of big data and knowledge graph technology, improves the scientific nature of project information management decision-making, reduces project risks, and provides strong support for scientific and technological innovation development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information management technology, and in particular to a method and system for managing information on science and technology projects based on big data. Background Technology

[0002] In the digital transformation of science and technology project management, with the surge in the number of projects and the increasing complexity of information, leveraging big data for efficient management has become an inevitable trend. From the integration and analysis of project information to risk prediction and precise retrieval and recommendation, building an intelligent management system is crucial for improving research efficiency and ensuring project progress.

[0003] Traditional methods of managing science and technology project information primarily rely on a combination of manual operation and simple information technology tools. This approach has several limitations: First, traditional methods are significantly inadequate in analyzing information correlation. Due to a lack of in-depth analysis of the correlation between science and technology project information and historical benchmark information, it is difficult to accurately determine the similarities and differences between new and previous projects. This prevents targeted optimization of new projects based on historical experience, significantly compromising the scientific rigor and accuracy of project information management. Second, in risk analysis, traditional methods rely heavily on manual experience, lacking comprehensive analysis of multi-dimensional data on project progress, costs, and resources. This makes it difficult to comprehensively and accurately identify potential management risks and provide strong data support for project decision-making. Furthermore, in information retrieval and matching, traditional simple keyword-based search methods fail to fully consider the semantic relationships and contextual information between keywords. The search results are often inaccurate and incomplete, failing to meet diverse user search needs and unable to provide project recommendations that include information management risk analysis results, thus reducing the efficiency of users obtaining effective information. Summary of the Invention

[0004] In order to overcome the above-mentioned deficiencies of the prior art, embodiments of the present invention provide a method and system for managing science and technology project information based on big data, so as to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for managing science and technology project information based on big data, comprising the following steps:

[0006] S1: Obtain information on science and technology projects to be processed: Collect newly added science and technology project information, mark it as each science and technology project to be processed, and mark the verified historical science and technology project benchmark information in the science and technology project information management database as benchmark project information;

[0007] S2: Project Information Correlation Deviation Analysis: Calculate the correlation deviation coefficient between each science and technology project information to be processed and the benchmark project information, and make a correlation deviation judgment. Based on the correlation deviation judgment result, optimize the science and technology project information and store it in the science and technology project information database.

[0008] S3: Construct a knowledge graph of science and technology projects: Obtain science and technology project information from the science and technology project information database and preprocess it to generate a feature vector set. Based on the feature vector set, construct a knowledge graph of science and technology projects.

[0009] S4: Risk Analysis of Information Management for Science and Technology Projects: Extract project schedule deviation, project cost deviation, and project resource matching deviation coefficients from the knowledge graph of science and technology projects, analyze the management risks of information for science and technology projects, and generate information management risk analysis results for science and technology projects.

[0010] S5: Science and Technology Project Information Matching Analysis: Extract search keywords from user project search requests, match the feature vectors of search keywords with the feature vectors of keywords in the science and technology project knowledge graph, and obtain the matching degree coefficient of user search keywords;

[0011] S6: Output a list of recommended project information: Based on the matching coefficient of the user's search keywords, output the recommended results of science and technology project information corresponding to the user's search keywords. The recommended results of science and technology project information include the information management risk analysis results of the corresponding science and technology projects.

[0012] Preferably, the specific execution method for obtaining the information of the technology project to be processed is as follows:

[0013] Newly added science and technology project information is collected according to the preset collection cycle. The collected newly added science and technology project information is marked as each science and technology project information to be processed, and each science and technology project information to be processed is numbered sequentially as 1, 2, ..., i, ..., n, where i represents the number of each science and technology project information to be processed.

[0014] Preferably, the specific steps for performing the project information correlation deviation analysis are as follows:

[0015] S21: Extract feature parameters from the benchmark project information and construct the feature parameter vector A of the benchmark project information. ,in, This represents the baseline value of the k-th feature parameter in the baseline item information, where k represents the number of each feature parameter, k=1, 2, 3, ..., m, and m represents the total number of feature parameters;

[0016] S22: Extract the feature parameters from the information of each science and technology project to be processed, and construct the feature parameter vector of each science and technology project information. B i, ,in, This represents the actual value of the kth feature parameter in the i-th science and technology project information to be processed, where i represents the number of each science and technology project information to be processed.

[0017] S23: Feature parameter vector A based on benchmark project information and feature parameter vectors of each science and technology project information to be processed B i A correlation deviation analysis model was constructed to calculate the correlation deviation coefficient between each science and technology project information to be processed and the benchmark project information, and this coefficient was marked as the correlation deviation coefficient of each science and technology project information to be processed. DC i The correlation deviation coefficient is used to determine the correlation deviation of each project information to be processed. Based on the correlation deviation determination result, the science and technology project information is optimized and stored in the science and technology project information database.

[0018] Preferably, the specific content of judging the correlation deviation of each item information based on the correlation deviation coefficient is as follows:

[0019] Correlation deviation coefficient of reading information on each science and technology project to be processed DC i The correlation deviation coefficient of each science and technology project information to be processed DC i The correlation coefficient is compared with a preset threshold. If the correlation coefficient of a certain scientific and technological project information to be processed is less than or equal to the preset threshold, it is determined that the correlation deviation between the scientific and technological project information to be processed and the benchmark project information is within a reasonable range, and the scientific and technological project information to be processed is stored in the scientific and technological project information database. Otherwise, it is determined that the correlation deviation between the scientific and technological project information to be processed and the benchmark project information exceeds a reasonable range, and a project information optimization signal is issued to optimize the project information whose correlation deviation exceeds a reasonable range.

[0020] Preferably, the specific steps for performing the risk analysis of the science and technology project information management are as follows:

[0021] S41: Extracting Project Schedule Deviation: Extracting the planned time node vector of science and technology projects from the science and technology project knowledge graph. T , , This represents the planned completion date of the j-th time node, and simultaneously extracts the actual time node vector P. , This represents the actual completion date of the j-th time node, where j represents the number of each time node, j=1, 2, 3, ..., c, and c represents the total number of time nodes;

[0022] By combining the planned time node vector and the actual time node vector, the project schedule deviation is calculated. Pd ;

[0023] S42: Extracting Project Cost Deviation: Extracting the budget cost vector S of technology projects from the technology project knowledge graph. , This represents the budget amount for the h-th cost category, and simultaneously extracts the actual cost vector F. , This represents the actual expenditure amount for the h-th cost category, where h represents the number of each cost category, h = 1, 2, 3, ..., v, and v represents the total number of cost categories;

[0024] Combine the budgeted cost vector and the actual cost vector to calculate the project cost variance. Cd ;

[0025] S43: Read Project Schedule Deviation Pd Project cost deviation Cd And extract the project resource matching deviation coefficient. Mdc A management risk analysis model for science and technology project information was constructed, and the management risk index of science and technology project information was obtained. MRI Based on the management risk index, the management risk of science and technology project information is analyzed, and the information management risk analysis results corresponding to science and technology projects are generated.

[0026] Preferably, the extraction of the project resource matching deviation coefficient Mdc The content is: Extracting the resource requirement vector R of science and technology projects from the knowledge graph of science and technology projects. , This represents the planned demand for resource type g, and simultaneously extracts the actual resource allocation vector X. , Let g represent the actual allocation of resource of type g, where g represents the number of each resource type, g = 1, 2, 3, ..., y, and y represents the total number of resource types.

[0027] By combining the resource demand vector and the actual allocation vector, the project resource matching deviation can be calculated. Md The project resource matching deviation coefficient is obtained based on the project resource matching deviation analysis. Mdc .

[0028] Preferably, the specific steps for performing the science and technology project information matching and analysis are as follows:

[0029] S51: Process the user's input item search request information into text, extract search keywords, and generate a search keyword feature vector L based on the search keywords;

[0030] S52: Traverse each project u in the technology project knowledge graph, extract keywords for each project u, and generate a keyword feature vector for each project u based on the keywords. C u ;

[0031] S53: Based on the keyword feature vector L and the keyword feature vector of each item u C u Calculate the matching coefficient between the user's search keywords and each item u in the science and technology project knowledge graph. DC u Mark it as the first match coefficient of the user's search keywords. DC u ;

[0032] S54: By introducing keyword position weights to correct the first matching degree coefficient, we obtain the corrected matching degree coefficient between the user's search keywords and each item u in the science and technology project knowledge graph. It is marked as the second relevance coefficient of the user's search keywords. .

[0033] To achieve the above objectives, the present invention provides the following technical solution: a big data-based science and technology project information management system, comprising the following steps for implementing the above-mentioned big data-based science and technology project information management method:

[0034] The pending science and technology project information acquisition module is used to collect newly added science and technology project information, mark it as pending science and technology project information, and mark the verified historical science and technology project benchmark information in the science and technology project information management database as benchmark project information.

[0035] Project Information Correlation Deviation Analysis Module: This module calculates the correlation deviation coefficient between each science and technology project information to be processed and the benchmark project information, judges the correlation deviation, optimizes the science and technology project information based on the correlation deviation judgment results, and stores the information in the science and technology project information database.

[0036] The module for constructing a knowledge graph of science and technology projects is used to obtain science and technology project information from the science and technology project information database, preprocess it to generate a feature vector set, and construct a knowledge graph of science and technology projects based on the feature vector set.

[0037] The Science and Technology Project Information Management Risk Analysis Module is used to extract project schedule deviation, project cost deviation, and project resource matching deviation coefficients from the science and technology project knowledge graph, analyze the management risks of science and technology project information, and generate information management risk analysis results for the corresponding science and technology projects.

[0038] Science and Technology Project Information Matching and Analysis Module: This module is used to extract search keywords from user project search requests, match and analyze the feature vectors of the search keywords with the feature vectors of keywords in the science and technology project knowledge graph, and obtain the matching degree coefficient of the user's search keywords.

[0039] The module that outputs a project information recommendation list: Based on the matching degree coefficient of the user's search keywords, it outputs the science and technology project information recommendation results corresponding to the user's search keywords. The science and technology project information recommendation results include the information management risk analysis results of the corresponding science and technology projects.

[0040] As described above, the big data-based method and system for managing science and technology project information provided by this invention has at least the following beneficial effects:

[0041] The present invention provides a method and system for managing science and technology project information based on big data. This method acquires newly added science and technology project information to be processed, as well as verified historical benchmark information of science and technology projects in a science and technology project information management database. It then calculates the correlation deviation coefficient between the information to be processed and the benchmark information, optimizes the processing of the science and technology project information accordingly, and stores it. Subsequently, it preprocesses the information in the science and technology project information database to generate a feature vector set, thereby constructing a science and technology project knowledge graph. Next, it extracts the project progress, cost, and resource matching deviation coefficients from the knowledge graph and analyzes them to generate information management risk analysis results. Then, it extracts keywords from user search requests, matches their feature vectors with the keyword feature vectors in the knowledge graph, and obtains a matching degree coefficient. Finally, based on the matching degree coefficient, it outputs a recommended list of science and technology project information containing the information management risk analysis results. This invention utilizes correlation deviation analysis to fully leverage historical project experience for precise optimization of new projects, enhancing the scientific rigor and forward-looking nature of project management. Knowledge graph-based risk analysis comprehensively and deeply uncovers potential project risks, providing a reliable basis for project decision-making. Precise keyword matching and information recommendation not only improve user retrieval efficiency but also allow users to intuitively understand project risks, effectively reducing uncertainties in the implementation of scientific and technological projects, ensuring their smooth progress, and improving the utilization efficiency of research resources. Attached Figure Description

[0042] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0043] Figure 1 This is a flowchart illustrating the big data-based scientific and technological project information management method of the present invention.

[0044] Figure 2 This is a schematic diagram of the structure of the big data-based science and technology project information management system of the present invention. Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] Example 1

[0047] Please see Figure 1 As shown, this invention provides a method for managing science and technology project information based on big data, including the following steps:

[0048] S1: Obtain information on science and technology projects to be processed: Collect newly added science and technology project information, mark it as each science and technology project to be processed, and mark the verified historical science and technology project benchmark information in the science and technology project information management database as benchmark project information;

[0049] In this embodiment, it should be specifically explained that the execution method for obtaining the information of the technology project to be processed is as follows:

[0050] Newly added science and technology project information is collected according to the preset collection cycle. The collected newly added science and technology project information is marked as each science and technology project information to be processed, and each science and technology project information to be processed is numbered sequentially as 1, 2, ..., i, ..., n, where i represents the number of each science and technology project information to be processed.

[0051] It should be noted that, in a specific embodiment, the preset collection cycle is specifically: collecting newly added science and technology project information once a day. In this embodiment, the value of the preset collection cycle is not specifically limited, and those skilled in the art can set it freely.

[0052] S2: Project Information Correlation Deviation Analysis: Calculate the correlation deviation coefficient between each science and technology project information to be processed and the benchmark project information, and make a correlation deviation judgment. Based on the correlation deviation judgment result, optimize the science and technology project information and store it in the science and technology project information database.

[0053] In this embodiment, it should be specifically explained that the execution steps of the project information correlation deviation analysis are as follows:

[0054] S21: Extract feature parameters from the benchmark project information and construct the feature parameter vector A of the benchmark project information. ,in, This represents the baseline value of the k-th feature parameter in the baseline item information, where k represents the number of each feature parameter, k=1, 2, 3, ..., m, and m represents the total number of feature parameters;

[0055] It should be noted that the characteristic parameters include, but are not limited to, parameters that can be specifically quantified, such as R&D cycle, budget range, number of participants, equipment investment amount, number of experiments, achievement transformation cycle, and energy consumption indicators.

[0056] S22: Extract the feature parameters from the information of each science and technology project to be processed, and construct the feature parameter vector of each science and technology project information. B i , ,in, This represents the actual value of the kth feature parameter in the i-th science and technology project information to be processed, where i represents the number of each science and technology project information to be processed.

[0057] It should be specifically noted that the actual values ​​of the feature parameters in the feature parameter vectors of each technology project information to be processed correspond one-to-one with the baseline values ​​of the feature parameters in the feature parameter vectors of the baseline project information. For example, The benchmark value representing the R&D cycle in the benchmark project information is then... This represents the actual value of the R&D cycle in the i-th pending technology project information.

[0058] S23: Feature parameter vector A based on benchmark project information and feature parameter vectors of each science and technology project information to be processed B i A correlation deviation analysis model was constructed to calculate the correlation deviation coefficient between each science and technology project information to be processed and the benchmark project information, and this coefficient was marked as the correlation deviation coefficient of each science and technology project information to be processed. DC i The correlation deviation coefficient is used to determine the correlation deviation of each project information to be processed. Based on the correlation deviation determination result, the science and technology project information is optimized and stored in the science and technology project information database.

[0059] The calculation formula for the correlation deviation analysis model is as follows: ,in, DC i It represents the correlation deviation coefficient between the i-th technology project information to be processed and the benchmark project information, and is a very small positive number to avoid the denominator being zero.

[0060] In this embodiment, it should be specifically noted that the correlation deviation coefficient between each piece of information on the science and technology project to be processed and the baseline project information is used in the formula. DC i The smaller the value, the higher the correlation between the item to be processed and the benchmark item, and the smaller the deviation.

[0061] In this embodiment, it should be specifically explained that the specific content of judging the correlation deviation of each item information based on the correlation deviation coefficient is as follows:

[0062] Correlation deviation coefficient of reading information on each science and technology project to be processed DC i The correlation deviation coefficient of each science and technology project information to be processed DC i The correlation coefficient is compared with a preset threshold. If the correlation coefficient of a certain technology project information to be processed is less than or equal to the preset threshold, it is determined that the correlation deviation between the technology project information to be processed and the benchmark project information is within a reasonable range, and the technology project information to be processed is stored in the technology project information database. Otherwise, it is determined that the correlation deviation between the technology project information to be processed and the benchmark project information exceeds a reasonable range, a project information optimization signal is issued, and the project information with a correlation deviation exceeding a reasonable range is optimized.

[0063] The specific content of optimizing project information with correlation deviations exceeding a reasonable range is as follows: The correlation deviation is automatically prompted to the data entry personnel, who then supplement and improve the deviation information based on the prompts, or the system intelligently completes the information based on historical correction data. The corrected information of the technology project to be processed re-enters step S23 to calculate the correlation deviation coefficient until it meets the requirements. DC i If the correlation deviation coefficient is less than or equal to the preset threshold, the corresponding science and technology project information will be stored in the science and technology project information database.

[0064] S3: Construct a knowledge graph of science and technology projects: Obtain science and technology project information from the science and technology project information database and preprocess it to generate a feature vector set. Based on the feature vector set, construct a knowledge graph of science and technology projects.

[0065] In this embodiment, it should be specifically explained that the execution method for constructing the knowledge graph of science and technology projects is as follows:

[0066] This process involves acquiring science and technology project information from a science and technology project information database, preprocessing the information to generate standardized science and technology project information, extracting keywords from the standardized information, and generating a keyword set. Each keyword in the keyword set is then converted into a high-dimensional semantic vector, forming a semantic vector space. Hierarchical clustering analysis is performed on the vectors in the semantic vector space to identify the theme categories of the science and technology projects, generating a theme distribution matrix. Key entities, including project name, participating institutions, researchers, and funding sources, are identified from the standardized science and technology project information, and an entity set is constructed. Relationships between entities are extracted from the standardized science and technology project information, forming a set of relation triples. Finally, the theme distribution matrix, entity set, and relation triple set are fused to generate a feature vector set.

[0067] Based on knowledge in the field of science and technology projects, design the conceptual hierarchy, entity types, and relation types of the knowledge graph to form a graph pattern; map the entity information in the feature vector set to the corresponding nodes in the graph pattern to establish an entity node set; based on the relation information in the feature vector set, establish connections between the nodes in the entity node set to construct the science and technology project knowledge graph.

[0068] S4: Risk Analysis of Information Management for Science and Technology Projects: Extract project schedule deviation, project cost deviation, and project resource matching deviation coefficients from the knowledge graph of science and technology projects, analyze the management risks of information for science and technology projects, and generate information management risk analysis results for science and technology projects.

[0069] In this embodiment, it should be specifically explained that the execution steps of the risk analysis for science and technology project information management are as follows:

[0070] S41: Extracting Project Schedule Deviation: Extracting the planned time node vector of science and technology projects from the science and technology project knowledge graph. T , , This represents the planned completion date of the j-th time node, and simultaneously extracts the actual time node vector P. , This represents the actual completion date of the j-th time node, where j represents the number of each time node, j=1, 2, 3, ..., c, and c represents the total number of time nodes;

[0071] By combining the planned time node vector and the actual time node vector, the project schedule deviation is calculated. Pd The calculation formula is: ,in, Pd Indicates project schedule deviation;

[0072] S42: Extracting Project Cost Deviation: Extracting the budget cost vector S of technology projects from the technology project knowledge graph. , This represents the budget amount for the h-th cost category, and simultaneously extracts the actual cost vector F. , This represents the actual expenditure amount for the h-th cost category, where h represents the number of each cost category, h = 1, 2, 3, ..., v, and v represents the total number of cost categories;

[0073] Combine the budgeted cost vector and the actual cost vector to calculate the project cost variance. Cd The calculation formula is: ,in, Cd Indicates project cost deviation;

[0074] S43: Read Project Schedule Deviation Pd Project cost deviation Cd And extract the project resource matching deviation coefficient. Mdc A management risk analysis model for science and technology project information was constructed, and the management risk index of science and technology project information was obtained. MRI Based on the management risk index, the management risk of science and technology project information is analyzed, and the information management risk analysis results corresponding to science and technology projects are generated.

[0075] The analysis of management risk of science and technology project information based on the management risk index includes: [Analyzing the management risk index of science and technology project information]. MRI Compared with the preset management risk index threshold, if the management risk index of technology project information... MRI If the risk index is less than the preset threshold, the technology project information is judged to have no management risk; otherwise, the technology project information is judged to have management risk, an early warning is issued for the technology project information, and the corresponding information management risk analysis results are generated and sent to the mobile device of the information management personnel.

[0076] The calculation formula for the management risk analysis model is as follows: ,in, MRI An index representing the management risk of science and technology project information. These represent the preset maximum allowable project schedule deviations. 、 Maximum project cost deviation These represent the weighting coefficients for project schedule deviation, project cost deviation, and project resource matching deviation, respectively. ;

[0077] In this embodiment, it should be specifically noted that the project schedule deviation in the formula... Pd The bigger 、 Project cost deviation Cd The larger the coefficient, the greater the project resource matching deviation coefficient. Mdc The larger the value, the higher the risk index for managing science and technology project information.MRI The larger the value, the more likely there are risks in the management of the science and technology project, and the more necessary it is to manage and maintain the information about the science and technology project.

[0078] In this embodiment, it should be specifically noted that the extraction of the project resource matching deviation coefficient... Mdc The content is: Extracting the resource requirement vector R of science and technology projects from the knowledge graph of science and technology projects. , This represents the planned demand for resource type g, and simultaneously extracts the actual resource allocation vector X. , Let g represent the actual allocation of resource of type g, where g represents the number of each resource type, g = 1, 2, 3, ..., y, and y represents the total number of resource types.

[0079] By combining the resource demand vector and the actual allocation vector, the project resource matching deviation can be calculated. Md The calculation formula is: The project resource matching deviation coefficient is obtained based on the project resource matching deviation analysis. Mdc The calculation formula is: ,in, Mdc This represents the project resource matching deviation coefficient.

[0080] S5: Science and Technology Project Information Matching Analysis: Extract search keywords from user project search requests, match the feature vectors of search keywords with the feature vectors of keywords in the science and technology project knowledge graph, and obtain the matching degree coefficient of user search keywords;

[0081] In this embodiment, it should be specifically explained that the execution steps of the science and technology project information matching analysis are as follows:

[0082] S51: Process the user's input item search request information into text, extract search keywords, and generate a search keyword feature vector L based on the search keywords;

[0083] S52: Traverse each project u in the technology project knowledge graph, extract keywords for each project u, and generate a keyword feature vector for each project u based on the keywords. C u ;

[0084] S53: Based on the keyword feature vector L and the keyword feature vector of each item u C u Calculate the matching coefficient between the user's search keywords and each item u in the science and technology project knowledge graph. DC u Mark it as the first match coefficient of the user's search keywords. DCu The calculation formula is: ,in, This represents the keyword feature vector L and the keyword feature vector for each item u. C u dot product, These represent the keyword feature vector L and the keyword feature vector for each item u, respectively. C u The length of the module.

[0085] It should be noted that the first matching degree coefficient DC u The value ranges from 0 to 1, when the first matching degree coefficient DC u The closer the coefficient is to 1, the more similar the user's search keywords are to the keyword features of item u, and the higher the matching degree; when the first matching degree coefficient is... DC u The closer it is to 0, the greater the difference between the user's search keywords and the keyword characteristics of item u, and the lower the matching degree.

[0086] S54: By introducing keyword position weights to correct the first matching degree coefficient, we obtain the corrected matching degree coefficient between the user's search keywords and each item u in the science and technology project knowledge graph. It is marked as the second relevance coefficient of the user's search keywords. The calculation formula is: ,in, This represents the position weight of the z-th search keyword, where z represents the index of each search keyword, z = 1, 2, 3, ..., b, and b represents the total number of search keywords;

[0087] It should be noted that, in a specific embodiment, the keyword position weight is as follows: if the keyword appears in the title, the position weight is set to 1.5; if the keyword appears in the abstract, the position weight is set to 1.2; and if the keyword appears in the details, the position weight is set to 1.0.

[0088] S6: Output a list of recommended project information: Based on the matching coefficient of the user's search keywords, output the recommended results of science and technology project information corresponding to the user's search keywords. The recommended results of science and technology project information include the information management risk analysis results of the corresponding science and technology projects.

[0089] In this embodiment, it should be specifically explained that the execution method of the output project information recommendation list is as follows:

[0090] Read the matching coefficient between the corrected user search keywords and each item u in the science and technology project knowledge graph. ,filter Items with a matching coefficient greater than or equal to a preset threshold are grouped into an item set. For project collections The items in the middle are matched according to the second matching degree coefficient. Sort the data in descending order to generate a list of recommended science and technology projects corresponding to the user's search keywords (the higher the matching degree, the higher the ranking). The recommended science and technology project information results include the information management risk analysis results of the corresponding science and technology projects.

[0091] In this embodiment, it should be specifically noted that the above formulas are all dimensionless calculations. The formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0092] Example 2

[0093] Please see Figure 2 As shown, this invention provides a big data-based science and technology project information management system, including:

[0094] The pending science and technology project information acquisition module is used to collect newly added science and technology project information, mark it as pending science and technology project information, and mark the verified historical science and technology project benchmark information in the science and technology project information management database as benchmark project information.

[0095] Project Information Correlation Deviation Analysis Module: This module calculates the correlation deviation coefficient between each science and technology project information to be processed and the benchmark project information, judges the correlation deviation, optimizes the science and technology project information based on the correlation deviation judgment results, and stores the information in the science and technology project information database.

[0096] The module for constructing a knowledge graph of science and technology projects is used to obtain science and technology project information from the science and technology project information database, preprocess it to generate a feature vector set, and construct a knowledge graph of science and technology projects based on the feature vector set.

[0097] The Science and Technology Project Information Management Risk Analysis Module is used to extract project schedule deviation, project cost deviation, and project resource matching deviation coefficients from the science and technology project knowledge graph, analyze the management risks of science and technology project information, and generate information management risk analysis results for the corresponding science and technology projects.

[0098] Science and Technology Project Information Matching and Analysis Module: This module is used to extract search keywords from user project search requests, match and analyze the feature vectors of the search keywords with the feature vectors of keywords in the science and technology project knowledge graph, and obtain the matching degree coefficient of the user's search keywords.

[0099] The module that outputs a project information recommendation list: Based on the matching degree coefficient of the user's search keywords, it outputs the science and technology project information recommendation results corresponding to the user's search keywords. The science and technology project information recommendation results include the information management risk analysis results of the corresponding science and technology projects.

[0100] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0101] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for managing science and technology project information based on big data, characterized in that: Includes the following steps: S1: Obtain information on science and technology projects to be processed: Collect newly added science and technology project information, mark it as each science and technology project to be processed, and mark the verified historical science and technology project benchmark information in the science and technology project information management database as benchmark project information; S2: Project Information Correlation Deviation Analysis: Calculate the correlation deviation coefficient between each science and technology project information to be processed and the benchmark project information, and make a correlation deviation judgment. Based on the correlation deviation judgment result, optimize the science and technology project information and store it in the science and technology project information database. S3: Construct a knowledge graph of science and technology projects: Obtain science and technology project information from the science and technology project information database and preprocess it to generate a feature vector set. Based on the feature vector set, construct a knowledge graph of science and technology projects. S4: Risk Analysis of Information Management for Science and Technology Projects: Extract project schedule deviation, project cost deviation, and project resource matching deviation coefficients from the knowledge graph of science and technology projects, analyze the management risks of information for science and technology projects, and generate information management risk analysis results for science and technology projects. S5: Science and Technology Project Information Matching Analysis: Extract search keywords from user project search requests, match the feature vectors of search keywords with the feature vectors of keywords in the science and technology project knowledge graph, and obtain the matching degree coefficient of user search keywords; S6: Output a project information recommendation list: Based on the matching degree coefficient of the user's search keywords, output the science and technology project information recommendation results corresponding to the user's search keywords. The science and technology project information recommendation results include the information management risk analysis results of the corresponding science and technology projects. The specific execution method for obtaining information on the science and technology projects to be processed is as follows: Newly added science and technology project information is collected according to the preset collection cycle. The collected newly added science and technology project information is marked as each science and technology project information to be processed, and each science and technology project information to be processed is numbered sequentially as 1, 2, ..., i, ..., n, where i represents the number of each science and technology project information to be processed. The specific steps for performing the project information correlation deviation analysis are as follows: S21: Extract feature parameters from the benchmark project information and construct the feature parameter vector A of the benchmark project information. ,in, This represents the baseline value of the k-th feature parameter in the baseline item information, where k represents the number of each feature parameter, k=1, 2, 3, ..., m, and m represents the total number of feature parameters; S22: Extract the feature parameters from the information of each science and technology project to be processed, and construct the feature parameter vector of each science and technology project information. B i , ,in, This represents the actual value of the kth feature parameter in the i-th science and technology project information to be processed, where i represents the number of each science and technology project information to be processed. S23: Feature parameter vector A based on benchmark project information and feature parameter vectors of each science and technology project information to be processed B i A correlation deviation analysis model was constructed to calculate the correlation deviation coefficient between each science and technology project information to be processed and the benchmark project information, and this coefficient was marked as the correlation deviation coefficient of each science and technology project information to be processed. DC i The correlation deviation coefficient is used to determine the correlation deviation of each project information to be processed. Based on the correlation deviation determination result, the science and technology project information is optimized and stored in the science and technology project information database. The specific details of judging the correlation deviation of each item's information based on the correlation deviation coefficient are as follows: Correlation deviation coefficient of reading information on each science and technology project to be processed DC i The correlation deviation coefficient of each science and technology project information to be processed DC i The correlation coefficient is compared with a preset threshold. If the correlation coefficient of a certain technology project information to be processed is less than or equal to the preset threshold, it is determined that the correlation deviation between the technology project information to be processed and the benchmark project information is within a reasonable range, and the technology project information to be processed is stored in the technology project information database. Otherwise, it is determined that the correlation deviation between the technology project information to be processed and the benchmark project information exceeds a reasonable range, a project information optimization signal is issued, and the project information with a correlation deviation exceeding a reasonable range is optimized. The specific steps for performing the risk analysis of information management for the aforementioned science and technology projects are as follows: S41: Extracting Project Schedule Deviation: Extracting the planned time node vector of science and technology projects from the science and technology project knowledge graph. T , , This represents the planned completion date of the j-th time node, and simultaneously extracts the actual time node vector P. , This represents the actual completion date of the j-th time node, where j represents the number of each time node, j=1, 2, 3, ..., c, and c represents the total number of time nodes; By combining the planned time node vector and the actual time node vector, the project schedule deviation can be calculated. Pd ; S42: Extracting Project Cost Deviation: Extracting the budget cost vector S of technology projects from the technology project knowledge graph. , This represents the budget amount for the h-th cost category, and simultaneously extracts the actual cost vector F. , This represents the actual expenditure amount for the h-th cost category, where h represents the number of each cost category, h = 1, 2, 3, ..., v, and v represents the total number of cost categories; Combine the budgeted cost vector and the actual cost vector to calculate the project cost variance. Cd ; S43: Read Project Schedule Deviation Pd Project cost deviation Cd And extract the project resource matching deviation coefficient. Mdc A management risk analysis model for science and technology project information was constructed, and the management risk index of science and technology project information was obtained. MRI Based on the management risk index, the management risk of science and technology project information is analyzed, and the information management risk analysis results corresponding to science and technology projects are generated.

2. The method for managing science and technology project information based on big data according to claim 1, characterized in that: The extraction of the project resource matching deviation coefficient Mdc The content is: Extracting the resource requirement vector R of science and technology projects from the knowledge graph of science and technology projects. , This represents the planned demand for resource type g, and simultaneously extracts the actual resource allocation vector X. , Let g represent the actual allocation of resource of type g, where g represents the number of each resource type, g = 1, 2, 3, ..., y, and y represents the total number of resource types. By combining the resource demand vector and the actual allocation vector, the project resource matching deviation can be calculated. Md The project resource matching deviation coefficient is obtained based on the project resource matching deviation analysis. Mdc .

3. The method for managing science and technology project information based on big data according to claim 1, characterized in that: The specific steps for performing the matching and analysis of the science and technology project information are as follows: S51: Process the user's input item search request information into text, extract search keywords, and generate a search keyword feature vector L based on the search keywords; S52: Traverse each project u in the technology project knowledge graph, extract keywords for each project u, and generate a keyword feature vector for each project u based on the keywords. C u ; S53: Based on the keyword feature vector L and the keyword feature vector of each item u C u Calculate the matching coefficient between the user's search keywords and each item u in the science and technology project knowledge graph. DC u Mark it as the first match coefficient of the user's search keywords. DC u ; S54: By introducing keyword position weights to correct the first matching degree coefficient, we obtain the corrected matching degree coefficient between the user's search keywords and each item u in the science and technology project knowledge graph. It is marked as the second relevance coefficient of the user's search keywords. .

4. A big data-based science and technology project information management system, used to implement the big data-based science and technology project information management method according to any one of claims 1-3, characterized in that, include: The pending science and technology project information acquisition module is used to collect newly added science and technology project information, mark it as pending science and technology project information, and mark the verified historical science and technology project benchmark information in the science and technology project information management database as benchmark project information. Project Information Correlation Deviation Analysis Module: This module calculates the correlation deviation coefficient between each science and technology project information to be processed and the benchmark project information, judges the correlation deviation, optimizes the science and technology project information based on the correlation deviation judgment results, and stores the information in the science and technology project information database. The module for constructing a knowledge graph of science and technology projects is used to obtain science and technology project information from the science and technology project information database, preprocess it to generate a feature vector set, and construct a knowledge graph of science and technology projects based on the feature vector set. The Science and Technology Project Information Management Risk Analysis Module is used to extract project schedule deviation, project cost deviation, and project resource matching deviation coefficients from the science and technology project knowledge graph, analyze the management risks of science and technology project information, and generate information management risk analysis results for the corresponding science and technology projects. Science and Technology Project Information Matching and Analysis Module: This module is used to extract search keywords from user project search requests, match and analyze the feature vectors of the search keywords with the feature vectors of keywords in the science and technology project knowledge graph, and obtain the matching degree coefficient of the user's search keywords. The module that outputs a project information recommendation list: Based on the matching degree coefficient of the user's search keywords, it outputs the science and technology project information recommendation results corresponding to the user's search keywords. The science and technology project information recommendation results include the information management risk analysis results of the corresponding science and technology projects.