Scientific and technological information management method and system based on big data

By analyzing the historical management data and vulnerability parameters of the scientific and technological information files, determining the monitoring frequency and recovery strategy, the information integrity problem is solved and the effectiveness and reliability of scientific and technological information management is improved.

CN120374057AInactive Publication Date: 2025-07-25CHINA METEOROLOGICAL ADMINISTRATION WEATHER MODIFICATION CENT
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510842000.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the scientific and technological information management methods based on big data focus on information classification management, ignoring the integrity of information during the circulation process, resulting in loss of file content, errors in formats and garbled codes, affecting the effectiveness and reliability of management.

Method used

By receiving scientific and technological information files, analyzing their historical management data and vulnerability parameters, determining monitoring frequency and recovery requirements tags, using different recovery strategies for file recovery, and determining the recovery effect to adjust the strategy.

Benefits of technology

The integrity monitoring of the target files is achieved, the recovery efficiency and management reliability is improved, and the effectiveness and reliability of file recovery is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374057A_ABST
    Figure CN120374057A_ABST
Patent Text Reader

Abstract

The invention discloses a science and technology information management method and system based on big data, and belongs to the technical field of science and technology information processing, and the method comprises the following steps: analyzing the management coefficient of a target file, determining the monitoring frequency of the target file, and analyzing the vulnerability coefficient of the target file; determining a recovery demand label of the target file in combination with the management coefficient of the target file, analyzing a first execution strategy of the target file and executing a recovery process when the file recovery demand label of the target file is waiting recovery or emergency recovery, receiving a recovery completion signal of the target file, and judging the recovery effect of the target file. And determining a second execution strategy of the target file. By monitoring the target file, storage and calling of the target file are monitored comprehensively, and the problem that in the prior art, integrity monitoring of the target file is blank is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of science and technology information processing, and in particular, to a science and technology information management method and system based on big data. Background Art

[0002] With the development of technology, information management technology has been widely applied to multiple fields such as data analysis of documents, intelligent decision-making, and pattern recognition, including aspects such as document storage, classification tags, and document archiving. In the prior art, more emphasis is placed on constructing a classification system for documents, constructing models for classification according to document content, and ignoring the integrity monitoring of documents. During the actual operation process, the accuracy of information processing, storage, and invocation in the system and the integrity of documents have a direct impact on the overall performance. Science and technology information may be affected by various dynamic factors during the circulation process, including the content coverage phenomenon when a document is viewed and cited, the associated fracture effect caused by multi-version iteration, and the fault phenomenon of missing key information in cross-team collaboration. These phenomena may lead to changes such as the loss of the core content of the document, the breakage of the logical chain, the distortion of knowledge reuse, and the weakening of decision support.

[0003] For example, a method for constructing a patent technology classification system in the field of air traffic management announced in the invention patent announcement with the announcement number: CN119250071B belongs to the technical field of patent analysis and mining. The method includes step 1, mining key technology classification vocabulary in the field of air traffic management; step 2, calculating the metric value of the technology classification vocabulary; step 3, constructing an initial patent technology classification system in the field of air traffic management; step 4, constructing an expert set, and experts in the expert set evaluate the technology classification vocabulary to obtain a first evaluation matrix; step 5, reducing the expert set and calculating a third evaluation matrix; step 6, calculating the importance degree and consistency degree of the technology classification vocabulary; step 7, reducing the technology classification vocabulary set to obtain the final patent technology classification system.

[0004] For example, a literature analysis method and management system based on big data announced in the invention patent announcement with the announcement number: CN116821349B includes obtaining the text data of the literature to be tested, preprocessing the text data to obtain first data and second data, the first data representing the text relationship of the preprocessed text data, the second data representing the citation times and the cited frequencies, the text relationship representing the relationship information of the theme of the text data, calculating a comprehensive score according to the first data and the second data, constructing a classification model according to the comprehensive score, inputting the first data and the second data into the classification model to obtain a classification, and classifying and managing the text to be tested according to the classification. However, in the process of implementing the technical solution of the invention in the embodiments of the present application, it is found that the above technologies have at least the following technical problems: In the prior art, the scientific and technological information management methods based on big data mostly focus on the classified management of information, ignoring the integrity problem of information during the transfer process. During the process of opening and viewing a file, problems such as file content loss, format error, and file garbled characters may occur, thus affecting the effectiveness and reliability of scientific and technological information management. Summary of the Invention

[0005] In a first aspect of the present invention, a scientific and technological information management method based on big data is provided, including the following steps: receiving a scientific and technological information file, denoted as a target file, obtaining the historical management data of the target file, analyzing the management coefficient of the target file, and determining the monitoring frequency of the target file.

[0006] Perform vulnerability monitoring on the target file based on the monitoring frequency of the target file, analyze the vulnerability coefficient of the target file, and determine the recovery requirement label of the target file in combination with the management coefficient of the target file.

[0007] When the file recovery requirement label of the target file is waiting for recovery or urgent recovery, analyze the first execution strategy of the target file based on the vulnerability coefficient of the target file. When the file recovery requirement label of the target file is urgent recovery, directly execute the recovery process based on the first execution strategy of the target file.

[0008] When the file recovery requirement label of the target file is waiting for recovery, obtain the queue waiting duration of each waiting-for-recovery file, determine the recovery priority index of each waiting-for-recovery file in combination with the vulnerability coefficient of each waiting-for-recovery file, and thus execute the recovery process of each waiting-for-recovery file in combination with the first execution strategy of each waiting-for-recovery file.

[0009] Receive the target file recovery completion signal, perform the recovery effect determination of the target file, and thus determine the second execution strategy of the target file.

[0010] Further, the management coefficient of the target file is specifically analyzed as follows: The historical management data of the target file includes the historical average number of citations, historical update frequency, and historical recovery times of the target file.

[0011] Analyze the management coefficient of the target file based on the historical management data of the target file.

[0012] The management coefficient of the target file is the quantitative data of the influence degree of the historical average number of citations, historical update frequency, and historical recovery times on the activity status of the same type of files corresponding to the target file. The specific analysis process is as follows: Compare the historical average number of citations, historical update frequency, and historical recovery times with the corresponding reference values, and perform coupling processing on the results of each comparison process in combination with the corresponding importance ratio, so as to obtain the management coefficient of the target file.

[0013] Further, for the monitoring frequency of the target file, the specific analysis process is as follows: Extract the first verification factor of the management coefficient and the second verification factor of the management coefficient of the target file preset in the database; if the management coefficient of the target file is greater than or equal to the first verification factor of the management coefficient, record the monitoring frequency determination result as high-frequency monitoring.

[0014] If the management coefficient of the target file is less than the first verification factor of the management coefficient and greater than the second verification factor of the management coefficient, record the monitoring frequency determination result as medium-frequency monitoring.

[0015] If the management coefficient of the target file is less than or equal to the second verification factor of the management coefficient, record the monitoring frequency determination result as low-frequency monitoring.

[0016] Further, for the vulnerability coefficient of the target file, the specific analysis process is as follows: During the preset monitoring period, collect the characterization parameters of the target file, including the hash difference rate, semantic difference degree, and reference anomaly number of the target file.

[0017] Analyze the vulnerability coefficient of the target file based on the characterization parameters of the target file.

[0018] The vulnerability coefficient of the target file is the quantitative data of the influence degree of the hash difference rate, semantic difference degree, and reference anomaly number on the vulnerability coefficient of the target file. The specific analysis process is: Compare the hash difference rate, semantic difference degree, and reference anomaly number with the corresponding reference values, and perform a coupling process on the results of each comparison process combined with the corresponding importance scores, so as to obtain the vulnerability coefficient of the target file.

[0019] Further, to determine the recovery requirement label of the target file, the specific analysis process is as follows: Analyze the recovery requirement value of the target file based on the vulnerability coefficient of the target file and the management coefficient of the target file.

[0020] The recovery requirement value of the target file is obtained by performing a coupling process on the vulnerability coefficient of the target file and the management coefficient of the target file combined with the corresponding weights, so as to obtain the recovery requirement value of the target file.

[0021] Extract the first recovery requirement verification factor and the second recovery requirement verification factor preset in the database.

[0022] If the recovery requirement value of the target file is greater than or equal to the first recovery requirement verification factor, determine the recovery requirement label of the target file as emergency recovery.

[0023] If the recovery requirement value of the target file is less than the first recovery requirement verification factor and greater than the second recovery requirement verification factor of the target file, determine the recovery requirement label of the target file as waiting for recovery.

[0024] If the recovery requirement value of the target file is less than or equal to the second recovery requirement verification factor, the recovery requirement label of the target file is determined as no need to recover.

[0025] Further, directly execute the recovery process based on the first execution policy of the target file. The specific analysis process is as follows: Extract the preset vulnerability coefficient verification factor in the database.

[0026] If the vulnerability coefficient of the target file is greater than the vulnerability coefficient verification factor, record the first execution policy of the target file as full - coverage recovery.

[0027] If the vulnerability coefficient of the target file is less than or equal to the vulnerability coefficient verification factor, record the first execution policy of the target file as location - based recovery.

[0028] When the file recovery requirement label of the target file is emergency recovery, directly execute the recovery process.

[0029] Further, execute the recovery process of each waiting - to - recover file in combination with the first execution policy of each waiting - to - recover file. The specific analysis process is as follows: When the file recovery requirement label of the target file is waiting to recover, transfer the target file to the recovery waiting queue.

[0030] The recovery waiting queue includes the first recovery waiting queue and the second recovery waiting queue.

[0031] Extract the preset third recovery requirement verification factor in the database.

[0032] If the recovery requirement value of the target file is less than the third recovery requirement verification factor, transfer the target file to the first recovery waiting queue.

[0033] If the recovery requirement value of the target file is greater than or equal to the third recovery requirement verification factor, transfer the target file to the second recovery waiting queue.

[0034] Obtain the recovery priority indicators of each waiting - to - recover file in the corresponding recovery waiting queue of the target file.

[0035] Sort the recovery priority indicators of each waiting - to - recover file from largest to smallest, and use the sorting order as the recovery priority of each waiting - to - recover file.

[0036] Obtain the first execution policy of each waiting - to - recover file, and execute the recovery process of each waiting - to - recover file in combination with the recovery priority of each waiting - to - recover file.

[0037] Further, for the recovery priority indicators of each waiting - to - recover file, the specific analysis process is as follows: During the preset monitoring period, collect the queue waiting duration of each waiting - to - recover file.

[0038] Obtain the vulnerability coefficient of each waiting - to - recover file.

[0039] Analyze the recovery priority indicators of each file waiting to be recovered based on the queue waiting duration of each file waiting to be recovered and the vulnerability coefficient of each file waiting to be recovered.

[0040] The specific analysis process of the recovery priority indicators of each file waiting to be recovered is as follows: Compare the queue waiting duration of each file waiting to be recovered with the corresponding reference value, and correct the comparison results based on the corresponding vulnerability coefficient, so as to obtain the recovery priority indicators of each file waiting to be recovered.

[0041] Furthermore, the second execution strategy of the target file is analyzed as follows: Re-obtain the recovery requirement value of the target file, denoted as the first recovery requirement value.

[0042] If the first recovery requirement value is greater than the first recovery requirement verification factor of the target file, it is marked as an invalid recovery, and the second execution strategy of the target file is recorded as generating a warning message.

[0043] If the first recovery requirement value is less than or equal to the first recovery requirement verification factor of the target file and greater than the second recovery requirement verification factor of the target file, it is marked as a partial recovery, and the second execution strategy of the target file is recorded as performing a secondary recovery.

[0044] If the first recovery requirement value is less than or equal to the second recovery requirement verification factor of the target file, it is marked as a complete recovery, and the second execution strategy of the target file is recorded as continuing to monitor.

[0045] The second aspect of the present invention provides a science and technology information management system based on big data, including: a management module for target files, a recovery module for target files, a damage discrimination module for target files, and a recovery effect verification module for target files.

[0046] Among them, the management module for target files is used to obtain the historical management data of the target file, analyze the management coefficient of the target file, and determine the monitoring frequency of the target file.

[0047] The recovery module for target files is used to monitor the vulnerability of the target file, analyze the vulnerability coefficient of the target file, and determine the recovery requirement label of the target file in combination with the management coefficient of the target file.

[0048] The damage discrimination module for files waiting to be recovered is used to discriminate the recovery requirement labels of each file waiting to be recovered, analyze and execute the recovery process of each file waiting to be recovered in combination with the first execution strategy of each file waiting to be recovered.

[0049] The recovery effect verification module for target files is used to receive the target file recovery completion signal, determine the recovery effect of the target file, determine the second execution strategy of the target file, and complete the verification of the recovery effect of the target file.

[0050] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. A method and system for managing scientific and technological information based on big data provided by the present invention realizes the integrity monitoring function of the target file by providing vulnerability monitoring of the target file, thereby determining the recovery requirement label of the target file, and effectively solves the problem of the blank of the integrity monitoring of the target file in the prior art.

[0051] 2. By discriminating the recovery label of the target file, the present invention can divide the target file into different categories according to the recovery requirements of the target file, and further dynamically adjust the recovery strategy for the target file according to the recovery requirement label of the target file, thereby improving the recovery efficiency of the target file.

[0052] 3. By determining the second execution strategy of the target file through the determination of the recovery effect of the target file, for the partially recovered files, after re-acquiring each parameter, the recovery process is performed again, and a warning prompt message is given for the files that cannot be repaired effectively, thereby ensuring the effectiveness of the target file recovery and improving the reliability of scientific and technological information management. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flowchart of a method for managing scientific and technological information based on big data provided by an embodiment of the present application.

[0054] Figure 2 It is a schematic structural diagram of a system for managing scientific and technological information based on big data provided by an embodiment of the present application.

[0055] Figure 3 It is a flowchart of the judgment and execution of the recovery strategy involved in an embodiment of the present application.

[0056] Figure 4 It is a flowchart of the verification of the recovery effect involved in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] It should be noted that the scientific and technological information-related files involved in the embodiments of the present application take the relevant documents with the content of artificial weather modification as an example.

[0059] Refer to Figure 1As shown in the figure, the first aspect of the present invention provides a technology information management method based on big data, including the following steps: Receive technology information - related files, denoted as target files, obtain the historical management data of the target files, analyze the management coefficient of the target files, and determine the monitoring frequency of the target files.

[0060] In this embodiment, the specific analysis of the management coefficient of the target files is as follows: The historical management data of the target files includes the historical average citation times, historical update frequency, and historical recovery times of the target files.

[0061] It should be noted that the historical management data of the target files refers to the historical management data of files of the same type as the target files stored in the current system.

[0062] It should also be noted that the historical average citation times, historical update frequency, and historical recovery times of the target files can all be collected from the system program logs.

[0063] It should be supplemented that when the historical average citation times of the target files increase, it usually means that the file content has importance or practicality, which prompts the increase of the file's update frequency; at the same time, due to high - frequency citation, the historical recovery times of the target files also tend to increase. If the historical update frequency increases, it usually indicates that the historical citation times of the target files are relatively high, thus causing an increase in the recovery times. There is a positive - correlation relationship among the three. When any one of the parameters fluctuates significantly, the other two parameters will also show a synchronous growth trend.

[0064] Extract the reference historical average citation times, reference historical update frequency, and reference historical recovery times of the target files stored in the database.

[0065] Extract the importance ratios of the historical average citation times, historical update frequency, and historical recovery times of the target files preset in the database.

[0066] It should be noted that the importance ratio of the historical average citation times of the target file, the importance ratio of the historical update frequency of the target file, and the importance ratio of the historical recovery times of the target file all have a value range between 0 and 1, and the sum of the importance ratio of the historical average citation times of the target file, the importance ratio of the historical update frequency of the target file, and the importance ratio of the historical recovery times of the target file is 1. When using, the preset values can be directly extracted from the database. The specific extraction method is, for example, to construct a one-to-one mapping set between the historical average citation times of the target file, the historical update frequency of the target file, and the historical recovery times of the target file and the corresponding importance ratios of the historical average citation times of the target file, the importance ratios of the historical update frequency of the target file, and the importance ratios of the historical recovery times of the target file. When using, the obtained historical average citation times of the target file, the historical update frequency of the target file, and the historical recovery times of the target file are respectively input into the corresponding mapping sets, so as to extract the importance ratio of the historical average citation times of the target file, the importance ratio of the historical update frequency of the target file, and the importance ratio of the historical recovery times of the target file.

[0067] Analyze the management coefficient of the target file based on the historical management data of the target file.

[0068] The management coefficient of the target file is the quantitative data of the influence degree of the historical average citation times, the historical update frequency, and the historical recovery times on the activity status of the corresponding same-type files of the target file. The specific analysis process is as follows: compare the historical average citation times, the historical update frequency, and the historical recovery times with the corresponding reference values, and perform coupling processing on the results of each comparison processing in combination with the corresponding importance ratios, so as to obtain the management coefficient of the target file.

[0069] In a specific embodiment, the management coefficient of the target file is specifically represented as: , Among them, is the management coefficient of the target file, is the historical average citation times of the target file, is the historical update frequency of the target file, is the historical recovery times of the target file, is the reference historical average citation times, is the reference historical update frequency, is the reference historical recovery times, is the importance ratio of the historical average citation times, is the importance ratio of the historical update frequency, is the importance ratio of the historical recovery times.

[0070] In this embodiment, the monitoring frequency of the target file is analyzed as follows: Extract the first verification factor of the management coefficient and the second verification factor of the management coefficient of the preset target file in the database.

[0071] If the management coefficient of the target file is greater than or equal to the first verification factor of the management coefficient, record the monitoring frequency determination result as high-frequency monitoring.

[0072] It should be noted that if the management coefficient of the target file is greater than or equal to the first verification factor of the management coefficient, it indicates that the target file exhibits greater instability, usually caused by a high number of operations on the target file. Therefore, it is determined as high-frequency monitoring.

[0073] If the management coefficient of the target file is less than the first verification factor of the management coefficient and greater than the second verification factor of the management coefficient, record the monitoring frequency determination result as medium-frequency monitoring.

[0074] It should be noted that if the management coefficient of the target file is less than the first verification factor of the management coefficient and greater than the second verification factor of the management coefficient, it indicates that the target file is relatively stable, usually caused by the number of operations on the target file being within an acceptable range. Therefore, it is determined as medium-frequency monitoring.

[0075] If the management coefficient of the target file is less than or equal to the second verification factor of the management coefficient, record the monitoring frequency determination result as low-frequency monitoring.

[0076] It should be noted that if the management coefficient of the target file is less than the first verification factor of the management coefficient and greater than the second verification factor of the management coefficient, it indicates that the target file is very stable, usually caused by a small number of operations on the target file. Therefore, it is determined as low-frequency monitoring.

[0077] It should be noted that high-frequency monitoring, medium-frequency monitoring, and low-frequency monitoring are all preset values in the database.

[0078] It should be added that the first verification factor of the management coefficient is greater than the second verification factor of the management coefficient.

[0079] Perform vulnerability monitoring on the target file based on the monitoring frequency of the target file, analyze the vulnerability coefficient of the target file, and determine the recovery requirement label of the target file in combination with the management coefficient of the target file.

[0080] Refer to Figure 3 As shown, it is the flowchart of the recovery strategy judgment and execution involved in the embodiment of the present application. Determine the recovery execution strategy through the vulnerability monitoring result, and perform recovery effect verification after the recovery operation ends.

[0081] In this embodiment, the vulnerability coefficient of the target file is analyzed as follows: During a preset monitoring period, collect the characterization parameters of the target file, including the hash difference rate, semantic difference degree, and number of reference anomalies of the target file.

[0082] It should be noted that the hash difference rate, semantic difference degree, and number of reference anomalies of the target file are collected from the system program logs.

[0083] It should be added that when the hash difference rate of the target file increases, it usually means that the target file has been changed, and the target file is likely to be damaged, resulting in a synchronous increase in the semantic difference rate. At the same time, due to the change of the target file information, the reference is likely to become invalid, and the number of reference anomalies also tends to increase. If the semantic difference rate increases abnormally, it usually indicates that the content of the target file has changed, causing fluctuations in the hash difference rate and reference validity. There is a positive correlation among the three. When any parameter fluctuates significantly, the other two parameters will also show a synchronous growth trend.

[0084] Extract the reference hash difference rate, reference semantic difference degree, and reference number of reference anomalies stored in the database.

[0085] Extract the importance ratio of the hash difference rate, the importance ratio of the semantic difference degree, and the importance ratio of the number of reference anomalies preset in the database.

[0086] It should be noted that the value ranges of the importance ratio of the hash difference rate, the importance ratio of the semantic difference degree, and the importance ratio of the number of reference anomalies are all between 0 and 1, and the sum of the importance ratio of the hash difference rate, the importance ratio of the semantic difference degree, and the importance ratio of the number of reference anomalies is 1. When using, the preset values can be directly extracted from the database. The specific extraction method is, for example, to construct a one-to-one mapping set between the hash difference rate, semantic difference degree, and number of reference anomalies and their corresponding importance ratios of the hash difference rate, importance ratios of the semantic difference degree, and importance ratios of the number of reference anomalies. When using, input the obtained hash difference rate, semantic difference degree, and number of reference anomalies into the corresponding mapping sets respectively, so as to extract the importance ratio of the hash difference rate, the importance ratio of the semantic difference degree, and the importance ratio of the number of reference anomalies.

[0087] Analyze the vulnerability coefficient of the target file based on the characterization parameters of the target file.

[0088] The vulnerability coefficient of the target file is the quantitative data of the influence degree of the hash difference rate, semantic difference degree, and number of reference anomalies on the vulnerability coefficient of the target file. The specific analysis process is as follows: compare the hash difference rate, semantic difference degree, and number of reference anomalies with the corresponding reference values, and couple the results of each comparison process with the corresponding importance scores to obtain the vulnerability coefficient of the target file.

[0089] In a specific embodiment, the vulnerability coefficient of the target file is specifically represented as: , Among them, is the vulnerability coefficient of the target file, is the hash difference rate, is the semantic difference degree, is the number of reference anomalies, is the reference hash difference rate, is the reference semantic difference degree, is the reference number of reference anomalies, is the importance proportion of the hash difference rate, is the importance proportion of the semantic difference degree, is the importance proportion of the number of reference anomalies.

[0090] In this embodiment, to determine the recovery requirement label of the target file, the specific analysis process is as follows: Analyze the recovery requirement value of the target file based on the vulnerability coefficient of the target file and the management coefficient of the target file.

[0091] In this embodiment, to analyze the recovery requirement value of the target file, the specific analysis process is as follows: Extract the vulnerability coefficient of the target file and the management coefficient of the target file, and perform coupling processing in combination with the corresponding weights to obtain the recovery requirement value of the target file.

[0092] Re-extract the importance proportion of the vulnerability coefficient of the target file and the importance proportion of the management coefficient of the target file.

[0093] It should be noted that the value ranges of the importance proportion of the vulnerability coefficient and the importance proportion of the management coefficient of the target file are both between 0 and 1, and the sum of the importance proportion of the vulnerability coefficient and the importance proportion of the management coefficient of the target file is 1.

[0094] In a specific embodiment, the recovery requirement value of the target file is specifically represented as: , Among them, is the recovery requirement value of the target file, is the vulnerability coefficient of the target file, is the management coefficient of the target file, is the importance proportion of the vulnerability coefficient of the target file, is the importance proportion of the management coefficient of the target file.

[0095] If the recovery requirement value of the target file is greater than or equal to the first recovery requirement verification factor, the recovery requirement label of the target file is determined to be emergency recovery.

[0096] It should be noted that if the recovery requirement value of the target file is greater than or equal to the first recovery requirement verification factor, it indicates that the damage degree of the target file is high or the target file belongs to a core technology file, and a recovery operation needs to be performed immediately. Therefore, it is determined as an emergency recovery.

[0097] If the recovery requirement value of the target file is less than the first recovery requirement verification factor and greater than the second recovery requirement verification factor of the target file, the recovery requirement label of the target file is determined as waiting for recovery.

[0098] It should be noted that if the recovery requirement value of the target file is less than the first recovery requirement verification factor and greater than the second recovery requirement verification factor of the target file, it indicates that the damage degree of the target file is low or the target file does not belong to a core technology file, and a recovery operation does not need to be performed immediately. Thus, it enters the waiting queue to wait for recovery. Therefore, it is determined as waiting for recovery.

[0099] If the recovery requirement value of the target file is less than or equal to the second recovery requirement verification factor, the recovery requirement label of the target file is determined as no need for recovery.

[0100] It should be noted that if the recovery requirement value of the target file is less than the second recovery requirement verification factor, it indicates that the integrity of the target file has not changed and no recovery operation is required. Continue to perform the monitoring operation of the target file. Therefore, it is determined as no need for recovery.

[0101] It should be added that the first recovery requirement verification factor is greater than the second recovery requirement verification factor of the target file.

[0102] When the file recovery requirement label of the target file is waiting for recovery or emergency recovery, analyze the first execution strategy of the target file based on the vulnerability coefficient of the target file. When the file recovery requirement label of the target file is emergency recovery, directly execute the recovery process based on the first execution strategy of the target file.

[0103] In this embodiment, directly execute the recovery process based on the first execution strategy of the target file. The specific analysis process is as follows: Extract the preset vulnerability coefficient verification factor in the database.

[0104] If the vulnerability coefficient of the target file is greater than the vulnerability coefficient verification factor, record the first execution strategy of the target file as full - text overwrite recovery.

[0105] It should be noted that if the vulnerability coefficient of the target file is greater than the vulnerability coefficient verification factor, it indicates that the damage degree of the target file is high. In order to quickly recover the damaged file and improve the recovery efficiency, it is necessary to perform the recovery of the target file in the way of full - text overwrite. Therefore, record the first execution strategy of the target file as full - text overwrite recovery.

[0106] It should be added that the data covering the whole text is sourced from the system backup library.

[0107] If the vulnerability coefficient of the target file is less than or equal to the vulnerability coefficient verification factor, record the first execution strategy of the target file as location recovery.

[0108] It should be noted that if the vulnerability coefficient of the target file is less than or equal to the vulnerability coefficient verification factor, it indicates that the damage degree of the target file is low. To reduce the waiting time for the target file to be restored, directly adopt the method of locating and restoring according to the original text to restore the target file. Therefore, record the first execution strategy of the target file as location recovery.

[0109] When the file recovery requirement label of the target file is emergency recovery, directly execute the recovery process.

[0110] When the file recovery requirement label of the target file is waiting for recovery, obtain the queue waiting duration of each file waiting for recovery, and determine the recovery priority index of each file waiting for recovery in combination with the vulnerability coefficient of each file waiting for recovery. Then, execute the recovery process of each file waiting for recovery in combination with the first execution strategy of each file waiting for recovery.

[0111] In this embodiment, execute the recovery process of each file waiting for recovery in combination with the first execution strategy of each file waiting for recovery. The specific analysis process is as follows: When the file recovery requirement label of the target file is waiting for recovery, transfer the target file to the recovery waiting queue.

[0112] The recovery waiting queue includes the first recovery waiting queue and the second recovery waiting queue.

[0113] Extract the preset third recovery requirement verification factor in the database.

[0114] If the recovery requirement value of the target file is less than the third recovery requirement verification factor, transfer the target file to the first recovery waiting queue.

[0115] It should be noted that if the recovery requirement value of the target file is less than the third recovery requirement verification factor, it indicates that the recovery requirement of the target file is relatively low and can continue to wait for the recovery operation. If it is monitored that there is a target file for emergency recovery, immediately stop the recovery process of the first recovery waiting queue and switch to perform the recovery operation for the target file for emergency recovery. If the first recovery waiting queue is currently recovering the target file for emergency recovery, the subsequent target files for emergency recovery continue to wait.

[0116] If the recovery requirement value of the target file is greater than or equal to the third recovery requirement verification factor, transfer the target file to the second recovery waiting queue.

[0117] It should be noted that if the recovery requirement value of the target file is greater than or equal to the third recovery requirement verification factor, it indicates that the recovery requirement of the target file is at an intermediate value. Therefore, the second recovery waiting queue will not execute the interrupted recovery process operation and always maintains the service of repairing the target file waiting to be repaired.

[0118] Obtain the recovery priority indicators of each waiting-to-recover file in the recovery waiting queue corresponding to the target file.

[0119] Refer to Figure 4 As shown, it is a flowchart for verifying the recovery effect involved in the embodiments of the present application. The recovery effect is determined based on the size relationship among the first recovery requirement value, the first recovery requirement verification factor of the target file, and the second recovery requirement verification factor of the target file, and the execution strategy is determined based on the recovery effect.

[0120] In this embodiment, the specific analysis process of the recovery priority indicators of each waiting-to-recover file is as follows: During a preset monitoring period, collect the queue waiting durations of each waiting-to-recover file.

[0121] Obtain the vulnerability coefficients of each waiting-to-recover file.

[0122] Extract the reference queue waiting durations of each waiting-to-recover file preset in the database and the reference vulnerability coefficients of each waiting-to-recover file.

[0123] Extract the queue waiting duration importance scores of each waiting-to-recover file preset in the database and the vulnerability coefficient importance scores of each waiting-to-recover file.

[0124] It should be noted that the value ranges of the queue waiting duration importance scores of each waiting-to-recover file and the vulnerability coefficient importance scores of each waiting-to-recover file are both between 0 and 1, and the sum of the queue waiting duration importance scores of each waiting-to-recover file and the vulnerability coefficient importance scores of each waiting-to-recover file is 1. When using, the pre-set values can be directly extracted from the database. The specific extraction method is, for example, to construct a one-to-one mapping set between the queue waiting durations of each waiting-to-recover file and the vulnerability coefficients of each waiting-to-recover file and the corresponding queue waiting duration importance scores of each waiting-to-recover file and the vulnerability coefficient importance scores of each waiting-to-recover file. When using, the obtained queue waiting durations of each waiting-to-recover file and the vulnerability coefficients of each waiting-to-recover file are respectively input into the corresponding mapping sets, so as to extract the queue waiting duration importance scores of each waiting-to-recover file and the vulnerability coefficient importance scores of each waiting-to-recover file.

[0125] The specific analysis process of the recovery priority index for each waiting-to-be-recovered file is as follows: The queue waiting duration of each waiting-to-be-recovered file is combined with the corresponding reference value for comparison processing, and each comparison result is corrected based on the corresponding vulnerability coefficient, so as to obtain the recovery priority index for each waiting-to-be-recovered file.

[0126] In a specific embodiment, the recovery priority index for each waiting-to-be-recovered file is specifically represented as follows: , where, is the recovery priority index of the i-th waiting-to-be-recovered file, is the vulnerability coefficient of the i-th waiting-to-be-recovered file, is the queue waiting duration of the i-th waiting-to-be-recovered file, is the reference queue waiting duration, is the importance score of the vulnerability coefficient, is the importance score of the queue waiting duration, i is the number of the waiting-to-be-recovered file, , and n is the number of waiting-to-be-recovered files.

[0127] Sort the recovery priority indexes of each waiting-to-be-recovered file from largest to smallest, and use the sorting order as the recovery priority of each waiting-to-be-recovered file.

[0128] Obtain the first execution strategy for each waiting-to-be-recovered file, and execute the recovery process of each waiting-to-be-recovered file in combination with the recovery priority of each waiting-to-be-recovered file.

[0129] Receive the signal that the target file recovery is completed, perform the determination of the recovery effect of the target file, and thus determine the second execution strategy of the target file.

[0130] It should be added that after the target file completes the recovery process, a signal indicating that the target file recovery is completed is generated.

[0131] In this embodiment, the specific analysis process of the second execution strategy of the target file is as follows: Re-obtain the recovery requirement value of the target file, denoted as the first recovery requirement value.

[0132] If the first recovery requirement value is less than or equal to the second recovery requirement verification factor of the target file, it is marked as fully recovered, and the second execution strategy of the target file is recorded as continuous monitoring.

[0133] It should be noted that if the first recovery requirement value is less than or equal to the second recovery requirement verification factor of the target file, it means that the recovery degree of the target file is good and all recoveries are successful. Therefore, it is marked as fully recovered, and the monitoring operation of the target file continues to be executed.

[0134] If the first recovery requirement value is less than or equal to the first recovery requirement verification factor of the target file and greater than the second recovery requirement verification factor of the target file, it is marked as partial recovery, and the second execution policy of the target file is recorded as performing secondary recovery.

[0135] It should be noted that if the first recovery requirement value is less than or equal to the first recovery requirement verification factor of the target file and greater than the second recovery requirement verification factor of the target file, it indicates that only a part of the target file has been successfully recovered and not fully recovered. Therefore, it is marked as partial recovery, and based on this, a secondary recovery operation is performed.

[0136] It should be added that the secondary recovery operation means re-obtaining all the parameters of the target file and re-entering the recovery process of the target file. If it is still marked as partial recovery after secondary recovery, the mark is modified to invalid repair and a warning prompt message is generated.

[0137] If the first recovery requirement value is greater than the first recovery requirement verification factor of the target file, it is marked as invalid recovery, and the second execution policy of the target file is recorded as generating a warning prompt message.

[0138] It should be noted that if the first recovery requirement value is greater than the first recovery requirement verification factor of the target file, it indicates that the target file has basically not been recovered or not recovered at all, and the recovered target file has basically not changed compared with the target file before recovery. Therefore, it is marked as invalid recovery, and at the same time, a warning prompt message is generated.

[0139] Refer to Figure 2 As shown, the second aspect of the present invention provides a science and technology information management system based on big data, including: A management module for the target file, which is used to obtain the historical management data of the target file, analyze the management coefficient of the target file, and determine the monitoring frequency of the target file.

[0140] A recovery module for the target file, which is used to monitor the vulnerability of the target file, analyze the vulnerability coefficient of the target file, and determine the recovery requirement label of the target file in combination with the management coefficient of the target file.

[0141] A damage discrimination module for the files waiting for recovery, which is used to discriminate the recovery requirement labels of the files waiting for recovery, analyze and execute the recovery process of the files waiting for recovery in combination with the first execution policy of each file waiting for recovery.

[0142] A verification module for the recovery effect of the target file, which is used to receive the target file recovery completion signal, determine the recovery effect of the target file, determine the second execution policy of the target file, and complete the verification of the recovery effect of the target file.

[0143] Those skilled in the art will appreciate that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0144] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0145] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realize the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0147] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0148] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A scientific and technological information management method based on big data, characterized in that, It includes the following steps: Receive science and technology information - type files, denoted as target files, obtain the historical management data of the target files, analyze the management coefficient of the target files, and determine the monitoring frequency of the target files; Conduct vulnerability monitoring on the target files based on the monitoring frequency of the target files, analyze the vulnerability coefficient of the target files, and determine the recovery requirement label of the target files in combination with the management coefficient of the target files; When the file recovery requirement label of the target file is waiting for recovery or urgent recovery, analyze the first execution strategy of the target file based on the vulnerability coefficient of the target file. When the file recovery requirement label of the target file is urgent recovery, directly execute the recovery process based on the first execution strategy of the target file; When the file recovery requirement label of the target file is waiting for recovery, obtain the queue waiting duration of each waiting - for - recovery file, determine the recovery priority index of each waiting - for - recovery file in combination with the vulnerability coefficient of each waiting - for - recovery file, and thus execute the recovery process of each waiting - for - recovery file in combination with the first execution strategy of each waiting - for - recovery file; Receive the target file recovery completion signal, conduct the determination of the recovery effect of the target file, and thus determine the second execution strategy of the target file.

2. The method for managing scientific and technological information based on big data according to claim 1, wherein: The management coefficient of the target file is specifically analyzed as follows: The historical management data of the target file includes the historical average number of citations, historical update frequency, and historical recovery times of the target file; Analyze the management coefficient of the target file based on the historical management data of the target file; The management coefficient of the target file is the quantitative data of the influence degree of the historical average number of citations, historical update frequency, and historical recovery times on the activity status of the corresponding type of files of the target file. The specific analysis process is as follows: Compare the historical average number of citations, historical update frequency, and historical recovery times with the corresponding reference values, and perform coupling processing on the results of each comparison in combination with the corresponding importance ratios to obtain the management coefficient of the target file.

3. The method for managing scientific and technological information based on big data according to claim 2, characterized in that: The monitoring frequency of the target file is specifically analyzed as follows: Extract the first verification factor and the second verification factor of the management coefficient of the target file preset in the database; If the management coefficient of the target file is greater than or equal to the first verification factor of the management coefficient, record the monitoring frequency determination result as high - frequency monitoring; If the management coefficient of the target file is less than the first verification factor of the management coefficient and greater than the second verification factor of the management coefficient, record the monitoring frequency determination result as medium - frequency monitoring; If the management coefficient of the target file is less than or equal to the second verification factor of the management coefficient, record the monitoring frequency determination result as low - frequency monitoring.

4. The method for managing scientific and technological information based on big data according to claim 1, wherein: The vulnerability coefficient of the target file is specifically analyzed as follows: During the preset monitoring period, collect the characterization parameters of the target file, including the hash difference rate, semantic difference degree, and reference anomaly number of the target file; Analyze the vulnerability coefficient of the target file based on the characterization parameters of the target file; The vulnerability coefficient of the target file is the quantified data of the influence degrees of the hash difference rate, semantic difference degree, and reference anomaly number on the vulnerability coefficient of the target file. The specific analysis process is as follows: The hash difference rate, semantic difference degree, and reference anomaly number are compared with the corresponding reference values, and the results of each comparison process are coupled with the corresponding importance scores to obtain the vulnerability coefficient of the target file.

5. The method for managing scientific and technological information based on big data according to claim 1, characterized in that: The specific analysis process for determining the recovery requirement label of the target file is as follows: Analyze the recovery requirement value of the target file based on the vulnerability coefficient of the target file and the management coefficient of the target file; The recovery requirement value of the target file is obtained by coupling the vulnerability coefficient of the target file and the management coefficient of the target file with the corresponding weights, so as to obtain the recovery requirement value of the target file; Extract the first recovery requirement verification factor and the second recovery requirement verification factor preset in the database; If the recovery requirement value of the target file is greater than or equal to the first recovery requirement verification factor, the recovery requirement label of the target file is determined to be emergency recovery; If the recovery requirement value of the target file is less than the first recovery requirement verification factor and greater than the second recovery requirement verification factor of the target file, the recovery requirement label of the target file is determined to be waiting for recovery; If the recovery requirement value of the target file is less than or equal to the second recovery requirement verification factor, the recovery requirement label of the target file is determined to be no recovery required.

6. The method for managing scientific and technological information based on big data according to claim 4, wherein: The specific analysis process for directly executing the recovery process based on the first execution strategy of the target file is as follows: Extract the vulnerability coefficient verification factor preset in the database; If the vulnerability coefficient of the target file is greater than the vulnerability coefficient verification factor, record the first execution strategy of the target file as full-text overwrite recovery; If the vulnerability coefficient of the target file is less than or equal to the vulnerability coefficient verification factor, record the first execution strategy of the target file as location recovery; When the file recovery requirement label of the target file is emergency recovery, directly execute the recovery process.

7. The method for managing scientific and technological information based on big data according to claim 1, characterized in that: The specific analysis process for executing the recovery processes of the waiting-for-recovery files in combination with the first execution strategies of the waiting-for-recovery files is as follows: When the file recovery requirement label of the target file is waiting for recovery, transfer the target file to the recovery waiting queue; The recovery waiting queue includes a first recovery waiting queue and a second recovery waiting queue; Extract the third recovery requirement verification factor preset in the database; If the recovery requirement value of the target file is less than the third recovery requirement verification factor, transfer the target file to the first recovery waiting queue; If the recovery requirement value of the target file is greater than or equal to the third recovery requirement verification factor, transfer the target file to the second recovery waiting queue; Obtain the recovery priority indicators of the waiting-for-recovery files corresponding to the target file in the recovery waiting queue; Sort the recovery priority indicators of the waiting-for-recovery files from largest to smallest, and use the sorting order as the recovery priority of the waiting-for-recovery files; Obtain the first execution strategies of the waiting-for-recovery files, and execute the recovery processes of the waiting-for-recovery files in combination with the recovery priorities of the waiting-for-recovery files.

8. The method for managing scientific and technological information based on big data according to claim 7, characterized in that: The specific analysis process for the recovery priority indicators of the waiting-for-recovery files is as follows: During a preset monitoring period, collect the queue waiting durations of each file waiting to be restored; Obtain the vulnerability coefficients of each file waiting to be restored; Based on the queue waiting durations of each file waiting to be restored and the vulnerability coefficients of each file waiting to be restored, analyze the restoration priority indicators of each file waiting to be restored; The specific analysis process of the restoration priority indicators of each file waiting to be restored is as follows: Combine the queue waiting durations of each file waiting to be restored with corresponding reference values for comparison processing, and correct each comparison result based on the corresponding vulnerability coefficient, so as to obtain the restoration priority indicators of each file waiting to be restored.

9. The method for managing scientific and technological information based on big data according to claim 1, characterized in that: The specific analysis process of the second execution strategy of the target file is as follows: Re-obtain the restoration requirement value of the target file, denoted as the first restoration requirement value; If the first restoration requirement value is greater than the first restoration requirement verification factor of the target file, mark it as an invalid restoration, and record the second execution strategy of the target file as generating a warning message; If the first restoration requirement value is less than or equal to the first restoration requirement verification factor of the target file and greater than the second restoration requirement verification factor of the target file, mark it as a partial restoration, and record the second execution strategy of the target file as performing a secondary restoration; If the first restoration requirement value is less than or equal to the second restoration requirement verification factor of the target file, mark it as a complete restoration, and record the second execution strategy of the target file as continuing to monitor.

10. A system applying the big data-based scientific and technological information management method according to any one of claims 1-9, characterized in that, Including: The management module of the target file, the restoration module of the target file, the damage discrimination module of the target file, the restoration effect verification module of the target file; Among them, the management module of the target file is used to obtain the historical management data of the target file, analyze the management coefficient of the target file, and determine the monitoring frequency of the target file; The restoration module of the target file is used to perform vulnerability monitoring on the target file, analyze the vulnerability coefficient of the target file, and determine the restoration requirement label of the target file in combination with the management coefficient of the target file; The damage discrimination module of the file waiting to be restored is used to discriminate the restoration requirement labels of each file waiting to be restored, analyze and execute the restoration process of each file waiting to be restored in combination with the first execution strategy of each file waiting to be restored; The restoration effect verification module of the target file is used to receive the target file restoration completion signal, determine the restoration effect of the target file, determine the second execution strategy of the target file, and complete the verification of the target file restoration effect.

Citation Information

Patent Citations

  • A document analysis method and management system based on big data

    CN116821349B

  • A method for constructing a patent technology classification system in the field of air traffic management

    CN119250071B

  • Science and technology project document key information extraction method based on document elements

    CN118349674A

  • Security control method for consistency of source code version and product version based on signature

    CN118627135A

  • Terminal data automatic backup and recovery method and system

    CN119668939A