Efficient Management System, Method, and Medium for Diversified New Archival Information Resources Driven by Big Data
Through the big data-driven archive information resource management system, the problem of inconvenience in integration and retrieval in traditional archive management methods is solved, efficient classification integration and convenient information retrieval are achieved, and the effect of archive management is improved.
Patent Information
- Application Number
- CN202411731402.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Traditional archive management methods cannot effectively integrate archive data from different sources and types, and cannot adapt to changes in business needs, resulting in inefficient management and inconvenient information retrieval.
A diversified new archive information resource management system driven by big data is adopted to receive and pre-process archive information resources through data reception and pre-processing modules, and a key information annotation and analysis module is used to extract data characteristics, classify, integrate and store, and build a search mechanism.
It realizes efficient classification, integration and management of diversified archive information resources, improves processing efficiency and reliability, and ensures the convenience of information retrieval and utilization.
Smart Images

Figure CN119690923B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to an efficient management system, method, and medium for diversified new archive information resources driven by big data. Background Art
[0002] At present, with the rapid development of information technology and the explosive growth of information volume, the types and quantities of archive information resources have become extremely large and diversified;
[0003] Traditional archive management methods face many challenges, such as low data processing efficiency, inconvenient information retrieval and utilization, etc. At the same time, due to the diversification of the types and quantities of archive information resources, the existing technologies cannot effectively integrate archive data from different sources and different types, and cannot adapt to the changing business requirements and technical environments, thus greatly reducing the management effect of diversified archive information resources;
[0004] Therefore, in order to overcome the above defects, the present invention provides an efficient management system, method, and medium for diversified new archive information resources driven by big data. Summary of the Invention
[0005] The present invention provides an efficient management system, method, and medium for diversified new archive information resources driven by big data, which are used to receive and preprocess different categories of archive information resources, so as to facilitate the integration and processing of different types of archive information resources. At the same time, key information is marked on the archive information resources, which provides convenience for the management of diversified new archive information resources. Secondly, the archive information resources are analyzed based on the key information marking results to accurately and effectively determine data characteristics such as the association relationship and change law between the archive information resources, so as to effectively classify and integrate the archive information resources according to the obtained data characteristics, improve the processing efficiency and reliability of diversified new archive information resources. Finally, the archive information resources are stored according to the classification and integration results, and a retrieval mechanism is constructed to quickly locate the required archive information resources according to the user's needs through the retrieval mechanism and effectively manage the utilization methods, improve the convenience of information retrieval and utilization, and ensure the management effect of the archive information resources.
[0006] The present invention provides an efficient management system for diversified new archive information resources driven by big data, including:
[0007] A data receiving and preprocessing module, which is used to receive different categories of archive information resources based on a data interface, and preprocess and mark key information on the archive information resources;
[0008] A data analysis module, which is used to analyze the preprocessed archival information resources based on the key information annotation results, extract the data characteristics of the archival information resources, and classify and integrate the archival information resources based on the data characteristics;
[0009] A data management module, which is used to store the archival information resources based on the classification and integration results, construct a retrieval mechanism based on the storage results, and open up multi-dimensional utilization methods for the stored archival information resources.
[0010] Preferably, a diversified new type of archival information resource efficient management system driven by big data, the data reception and preprocessing module includes:
[0011] A data source docking unit, which is used for:
[0012] Determining the allowable management resource types based on management requirements, and determining the allowable docking sources based on the allowable management resource types;
[0013] Performing interface docking with the allowable docking sources based on the data interface, and extracting the basic attributes of each allowable docking source based on the interface docking results;
[0014] A data reception unit, which is used for:
[0015] Analyzing the basic attributes to obtain the data categories corresponding to each allowable docking source, and determining the different requirements for the basic communication parameters of the data interface based on the different data categories;
[0016] Unifying and debugging the basic communication parameters of different data categories based on the different requirements, and receiving the archival information resources of different allowable docking sources based on the debugging results.
[0017] Preferably, a diversified new type of archival information resource efficient management system driven by big data, the data reception and preprocessing module includes:
[0018] A result acquisition unit, which is used for:
[0019] Obtaining the obtained archival information resources, and determining the data service structures of different categories of archival information resources based on the data sources of the archival information resources;
[0020] Determining the key fields and corresponding structure conversion rules in the archival information resources based on the data service structures, and performing structure assignment on the key fields based on the structure conversion rules to obtain structured archival information resources;
[0021] A data preprocessing unit, which is used for:
[0022] Visualizing the structured archival information resources in a visualization coordinate system, and determining the local outliers in the structured archival information resources based on the visualization display;
[0023] Remove data points for local outliers, fit the structured archival information resources based on the data point removal results, and perform data standardization processing on the fitting results based on a preset standardization interval to complete the preprocessing of the archival information resources.
[0024] Preferably, a diversified new type of archival information resource efficient management system driven by big data, the data reception and preprocessing module, includes:
[0025] The result analysis unit is used to obtain the preprocessed archival information resources and determine the business characteristics of each category of archival information resources based on the data sources of the archival information resources;
[0026] The key information determination unit is used for:
[0027] Determine the information indicators for extracting key information from each category of archival information sources based on the business characteristics. At the same time, segment the archival information sources based on the natural language system, and perform part-of-speech tagging on each target word based on the segmentation results;
[0028] Extract the semantics of each target word, perform context association on the semantics of each target word based on the part-of-speech tagging results, and extract key information from the association results based on the information indicators;
[0029] The annotation unit is used to perform differential annotation on the key information of different categories of archival information resources based on a preset annotation strategy.
[0030] Preferably, a diversified new type of archival information resource efficient management system driven by big data, the data analysis module, includes:
[0031] The data analysis unit is used for:
[0032] Obtain the key information annotation results, extract the visual symbol representations of the key information annotation results, and perform preliminary classification on the preprocessed archival information resources based on the visual symbol representations;
[0033] Determine the data distribution characteristics of each category of archival information resources based on the preliminary classification results, and select key features from the key information based on the data distribution characteristics;
[0034] Perform vector conversion on the key features to obtain key feature vectors, and perform correlation analysis on the key features based on the key feature vectors to obtain the preliminary correlations between the key features;
[0035] The data feature determination unit is used for:
[0036] Cluster the key features based on the preliminary correlations to obtain relevant data groups, and analogously display the relevant data groups in a rectangular coordinate system;
[0037] Determine the personalized change trends of each key feature in the relevant data group based on the analogy display results, and successively perform independent variable locking analysis on each key feature based on the personalized change trends to obtain the multi-dimensional relative change relationships between each key feature and the remaining key features;
[0038] Obtain the multi-dimensional association rules between the key features in the relevant data group based on the multi-dimensional relative change relationships, and determine the confidence interval based on the business analysis requirements;
[0039] Extract the frequent item sets of each key feature under the multi-dimensional association rules, and determine the relative size relationship between the multi-dimensional association rules and the confidence interval based on the frequent item sets;
[0040] Determine the association rules that meet the confidence interval as strong association rules, and obtain the data characteristics of the archival information resources based on the strong association rules;
[0041] The data classification and integration unit is used to determine the classification criteria for the archival information resources based on the data characteristics, and classify and integrate the archival information resources based on the classification criteria.
[0042] Preferably, a diversified new type of archival information resource efficient management system driven by big data, the data classification and integration unit includes:
[0043] The result acquisition subunit is used to:
[0044] Obtain the classification criteria for the archival information resources, and classify the archival information resources based on the classification criteria;
[0045] Perform content verification on each category of archival information resources based on the classification results, remove duplicates from each category of archival information resources based on the content verification results, and merge the archival information resources with strong association rules;
[0046] The visualization subunit is used to generate the association topology structure path of each category of archival information resources and the association path between different categories of archival information resources based on the merging results, and generate a topology structure diagram of the archival information resources based on the association topology structure path and the association path.
[0047] Preferably, a diversified new type of archival information resource efficient management system driven by big data, the data management module includes:
[0048] The data specification unit is used to:
[0049] Obtain the basic parameters of the preset storage space. At the same time, obtain the basic structure parameters of each archival information resource, and compare the differences in the basic structure parameters of each archival information resource to obtain the structure difference parameters;
[0050] Determine the compatibility characteristics between each archival information resource and the preset storage space based on the basic parameters and structural difference parameters of the preset storage space, and determine the unified data storage format for each archival information resource based on the compatibility characteristics;
[0051] A data storage unit, which is used to perform format conversion on the classified and integrated archival information resources based on the unified data storage format, and perform spatial partition storage of the classified and integrated archival information resources in the preset storage space based on the format conversion result.
[0052] Preferably, a diversified new type of archival information resource efficient management system driven by big data, the data management module includes:
[0053] A retrieval mechanism construction unit, which is used for:
[0054] Determine the storage addresses of each archival information resource based on the storage result. At the same time, extract the data characteristics of each archival information resource, and generate retrievable indexes for each archival information resource based on the storage addresses and data characteristics;
[0055] Determine the allowable data retrieval conditions based on the preset data retrieval service, and map the allowable data retrieval conditions with the retrievable indexes;
[0056] Invoke the standard data retrieval engine based on the big data, and perform personalized parameter configuration on the script file of the standard data retrieval engine based on the mapping result to obtain the retrieval mechanism;
[0057] A data management unit, which is used for:
[0058] Obtain the retrieval requirement, analyze the retrieval requirement based on the retrieval mechanism, obtain the retrieval target for the stored archival information resources, and extract the encryption level of the retrieval target;
[0059] Determine the allowable multi-dimensional utilization methods of the retrieval target based on the encryption level, and open the utilization permissions of the retrieval target based on the allowable multi-dimensional utilization methods.
[0060] The present invention provides a method for efficiently managing diversified new type of archival information resources driven by big data, including:
[0061] Step 1: Receive different types of archival information resources based on the data interface, and perform preprocessing and key information annotation on the archival information resources;
[0062] Step 2: Analyze the preprocessed archival information resources based on the key information annotation results, extract the data characteristics of the archival information resources, and classify and integrate the archival information resources based on the data characteristics;
[0063] Step 3: Store the archival information resources based on the classification integration result, construct a retrieval mechanism based on the storage result, and enable multi-dimensional utilization methods for the stored archival information resources.
[0064] The present invention provides a data-driven diverse new type of efficient management medium for archival information resources, on which a computer program is stored:
[0065] When the computer program is executed by a processor, it implements the steps of any one of the data-driven diverse new type of efficient management methods for archival information resources.
[0066] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0067] By receiving and preprocessing archival information resources of different categories, it is convenient to integrate and process archival information resources of different types. At the same time, key information is marked on the archival information resources, which provides convenience for the management of diverse new type of archival information resources. Secondly, by analyzing the archival information resources based on the key information marking results, it is possible to accurately and effectively determine data characteristics such as the association relationship and change law between archival information resources, so as to effectively classify and integrate the archival information resources according to the obtained data characteristics, improving the processing efficiency and reliability of diverse new type of archival information resources. Finally, the archival information resources are stored according to the classification integration result, and a retrieval mechanism is constructed to quickly locate the required archival information resources according to user needs through the retrieval mechanism and effectively manage the utilization methods, improving the convenience of information retrieval and utilization and ensuring the management effect of archival information resources.
[0068] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures specifically pointed out in this application document.
[0069] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0070] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0071] Figure 1 is a structural diagram of a big data-driven diverse new type of efficient management system for archival information resources in an embodiment of the present invention;
[0072] Figure 2This is the structural diagram of the data reception and preprocessing module in a diversified new type of archive information resource efficient management system driven by big data according to an embodiment of the present invention;
[0073] Figure 3 This is the flowchart of a diversified new type of archive information resource efficient management method driven by big data according to an embodiment of the present invention. Specific embodiments
[0074] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.
[0075] Embodiment 1:
[0076] This embodiment provides a diversified new type of archive information resource efficient management system driven by big data, as Figure 1 shown, including:
[0077] A data reception and preprocessing module, configured to receive different types of archive information resources based on a data interface, and perform preprocessing and key information annotation on the archive information resources;
[0078] A data analysis module, configured to analyze the preprocessed archive information resources based on the key information annotation results, extract the data features of the archive information resources, and classify and integrate the archive information resources based on the data features;
[0079] A data management module, configured to store the archive information resources based on the classification and integration results, construct a retrieval mechanism based on the storage results, and open multi-dimensional utilization modes for the stored archive information resources based on the construction results.
[0080] In this embodiment, different types of archive information resources refer to data in different formats or types, such as text, images, audio, etc.
[0081] In this embodiment, preprocessing refers to operations such as removing abnormal data and performing standardization processing on the archive information resources.
[0082] In this embodiment, key information annotation refers to circling the main content in the archive information resources in a prominent manner to facilitate highlighting of the key information.
[0083] In this embodiment, data features refer to the correlation relationships or potential change rules existing between the archive information resources, etc., to facilitate classification, integration, and management of the archive information resources.
[0084] In this embodiment, the retrieval mechanism is a tool for parsing user requirements and retrieving the required archive information resources from the stored archive information resources.
[0085] In this embodiment, the multi-dimensional utilization method refers to the permission to utilize the stored archival information resources, which can be, for example, online browsing and downloading, etc., aiming to effectively manage different categories of archival information resources.
[0086] The working principle and beneficial effects of the above technical solution are as follows: By receiving and preprocessing different categories of archival information resources, it is convenient to integrate and process different types of archival information resources. At the same time, key information is marked on the archival information resources, which provides convenience for the management of diverse new archival information resources. Secondly, through the analysis of the archival information resources based on the key information marking results, it is possible to accurately and effectively determine data characteristics such as the correlation relationship and change law between archival information resources, so as to effectively classify and integrate the archival information resources according to the obtained data characteristics, improving the processing efficiency and reliability of diverse new archival information resources. Finally, the archival information resources are stored according to the classification and integration results, and a retrieval mechanism is constructed to quickly locate the required archival information resources according to the user's needs through the retrieval mechanism and effectively manage the utilization method, improving the convenience of information retrieval and utilization and ensuring the management effect of archival information resources.
[0087] Embodiment 2:
[0088] Based on Embodiment 1, this embodiment provides an efficient management system for diverse new archival information resources driven by big data. The data reception and preprocessing module includes:
[0089] The data source docking unit is used for:
[0090] Determining the allowable management resource types based on management requirements and determining the allowable docking sources based on the allowable management resource types;
[0091] Performing interface docking with the allowable docking sources based on the data interface and extracting the basic attributes of each allowable docking source based on the interface docking results;
[0092] The data reception unit is used for:
[0093] Analyzing the basic attributes to obtain the data categories corresponding to each allowable docking source and determining the different requirements for the basic communication parameters of the data interface based on different data categories;
[0094] Unifying and debugging the basic communication parameters of different data categories based on the different requirements and receiving the archival information resources of different allowable docking sources based on the debugging results.
[0095] In this embodiment, the management requirements are known in advance and are used to characterize the types capable of resource management, that is, the allowable resource management types.
[0096] In this embodiment, the allowable docking source refers to the data source or terminal source corresponding to the allowable resource management type.
[0097] In this embodiment, the basic attributes refer to the configuration of each allowable docking source and information such as the generated resource categories and compositions.
[0098] In this embodiment, the basic communication parameters refer to the data communication format and data communication rate of the data interface, etc.
[0099] In this embodiment, the differential requirements refer to the different parameters in the basic communication parameter requirements of different data categories for the data interface.
[0100] In this embodiment, the parameter unified debugging refers to making the basic communication parameters of different data categories compatible and unified according to the differential requirements, so as to facilitate the effective reception of the file information resources generated by different allowable docking sources through the data interface.
[0101] The working principle and beneficial effects of the above technical solution are: by analyzing the management requirements, effectively determining the manageable resource types and allowable docking sources, so as to facilitate the docking of the allowable docking sources with the data interface. Secondly, after the interface docking, the basic attributes of each allowable docking source are extracted, and the parameter unified debugging of the basic communication parameters of the data interface is realized according to the basic attributes, so as to facilitate ensuring the reliable and effective reception of the file information resources of different allowable docking sources and guaranteeing the accuracy of resource reception.
[0102] Embodiment 3:
[0103] Based on Embodiment 1, this embodiment provides a highly efficient management system for diversified new file information resources driven by big data. The data reception and preprocessing module includes:
[0104] Result acquisition unit, used for:
[0105] Obtaining the obtained file information resources, and determining the data service structures of different categories of file information resources based on the data sources of the file information resources;
[0106] Determining the key fields and corresponding structure conversion rules in the file information resources based on the data service structures, and performing structure assignment on the key fields based on the structure conversion rules to obtain structured file information resources;
[0107] Data preprocessing unit, used for:
[0108] Visualize the structured archival information resources in a visualization coordinate system, and determine the local outliers in the structured archival information resources based on the visualization;
[0109] Remove the data points of the local outliers, fit the structured archival information resources based on the result of the data point removal, and perform data standardization processing on the fitting result based on a preset standardization interval to complete the preprocessing of the archival information resources.
[0110] In this embodiment, the data service structure refers to the data composition of different categories of archival information resources, which facilitates the structure conversion of different categories of archival information resources, that is, converting unstructured archival information resources into structured archival information resources.
[0111] In this embodiment, the structure conversion rule refers to the specific scheme or strategy for converting archival information resources.
[0112] In this embodiment, structure assignment refers to reassigning the key fields in the archival information resources through the structure conversion rule, so as to convert the archival information resources into structured archival information resources, that is, archival information resources with unified structure and format definitions.
[0113] In this embodiment, the local outlier refers to a data sample in the structured archival information resources that deviates from the average data value level.
[0114] In this embodiment, fitting the structured archival information resources based on the result of the data point removal means that after removing the data points of the local outliers, smoothing the remaining data points to ensure the continuity of the data.
[0115] In this embodiment, the preset standardization interval is set in advance, for example, it can be [0, 1], etc. The purpose is to convert different categories of structured archival information resources to a unified value range, so as to facilitate the management of different categories of archival information resources.
[0116] The working principle and beneficial effects of the above technical solution are as follows: By determining the data service structure of different categories of archival information resources, converting different categories of archival information resources into structured archival information resources according to the data service structure, which facilitates the effective processing of different categories of archival information resources. Secondly, visualizing the structured archival information resources in a visualization coordinate system, effectively determining and removing the local outliers according to the visualization result, ensuring the reliability of the finally obtained archival information resources. Finally, fitting and performing data standardization processing on the archival information resources after removing the local outliers, realizing accurate and effective preprocessing of the archival information resources, providing convenience for the efficient management of diverse new archival information resources.
[0117] Example 4:
[0118] Based on Example 1, this example provides an efficient management system for diverse new archive information resources driven by big data, as Figure 2 shown, the data reception and preprocessing module includes:
[0119] A result analysis unit, which is used to obtain the preprocessed archive information resources and determine the business characteristics of each category of archive information resources based on the data sources of the archive information resources;
[0120] A key information determination unit, which is used for:
[0121] Determining the information indicators for extracting key information from each category of archive information sources based on the business characteristics. At the same time, segmenting the archive information sources based on the natural language system, and performing part-of-speech tagging on each target word based on the segmentation results;
[0122] Extracting the semantics of each target word, associating the semantics of each target word contextually based on the part-of-speech tagging results, and extracting key information from the association results based on the information indicators;
[0123] A tagging unit, which is used to perform differential tagging on the key information of different categories of archive information resources based on a preset tagging strategy.
[0124] In this example, the business characteristics refer to the business features corresponding to each category of archive information resources, including business types, etc.
[0125] In this example, the information indicators refer to the basis for determining the extraction of key information from each category of archive information sources according to the business characteristics. For example, they can be identity information, business category information, etc.
[0126] In this example, the natural language system is preset and is used to represent the structure of the language and the corresponding parts of speech of different words in the text, etc.
[0127] In this example, the target words refer to multiple different words obtained by segmenting the archive information sources based on the natural language system.
[0128] In this example, part-of-speech tagging refers to annotating the parts of speech of different words. For example, it can be noun or verb annotation, etc.
[0129] In this example, the preset tagging strategy is preset and is used to represent the method or strategy for tagging different categories of archive information resources. Among them, differential tagging means using different standard methods to tag the key information of different categories of archive information resources, so as to facilitate the distinction of the tagging results.
[0130] The working principle and beneficial effects of the above technical solution are as follows: By determining the business characteristics of each category of archival information resources, the information indicators for key information extraction from various categories of archival information sources are determined according to the business characteristics, providing a reference basis for key information extraction. Secondly, the archival information sources are segmented and part-of-speech tagged. At the same time, the semantics of each target word are extracted, and the semantics of each target word are contextually associated according to the part-of-speech tagging results, thereby realizing the effective extraction of key information. Finally, a preset annotation strategy is used to differentially annotate the key information of different categories of archival information resources, facilitating the distinction of the key information of different categories of archival information resources.
[0131] Embodiment 5:
[0132] Based on Embodiment 1, this embodiment provides a data-driven diversified and efficient management system for new archival information resources. The data analysis module includes:
[0133] The data analysis unit is used for:
[0134] Obtain the key information annotation results, extract the visual symbol representation of the key information annotation results, and preliminarily classify the preprocessed archival information resources based on the visual symbol representation;
[0135] Based on the preliminary classification results, determine the data distribution characteristics of each category of archival information resources, and select key features from the key information based on the data distribution characteristics;
[0136] Perform vector transformation on the key features to obtain key feature vectors, and perform correlation analysis on the key features based on the key feature vectors to obtain the preliminary correlation between the key features;
[0137] The data feature determination unit is used for:
[0138] Cluster the key features based on the preliminary correlation to obtain relevant data groups, and analogously display the relevant data groups in a rectangular coordinate system;
[0139] Based on the analog display results, determine the personalized change trends of the key features in the relevant data groups, and sequentially perform independent variable locking analysis on the key features based on the personalized change trends to obtain the multi-dimensional relative change relationships between each key feature and the remaining key features;
[0140] Based on the multi-dimensional relative change relationships, obtain the multi-dimensional association rules between the key features in the relevant data groups, and determine the confidence interval based on the business analysis requirements;
[0141] Extract the frequent item sets of each key feature under the multi-dimensional association rules, and determine the relative size relationship between the multi-dimensional association rules and the confidence interval based on the frequent item sets;
[0142] Determine the association rules that meet the confidence interval as strong association rules, and obtain the data characteristics of the archival information resources based on the strong association rules;
[0143] A data classification and integration unit, configured to determine the classification criteria for the archival information resources based on the data characteristics, and classify and integrate the archival information resources based on the classification criteria.
[0144] In this embodiment, the visual symbol representation refers to the specific symbol form corresponding to the key information annotation result, etc.
[0145] In this embodiment, the preliminary classification refers to a general classification of the preprocessed archival information resources according to the visual symbol representation.
[0146] In this embodiment, the data distribution characteristics refer to the value distribution range and other situations of each category of archival information resources after preliminary classification.
[0147] In this embodiment, the key features refer to the key features extracted from the key information of each archival information resource.
[0148] In this embodiment, the key feature vector refers to the result of converting the key features into a vector form, that is, converting the text into a vector form.
[0149] In this embodiment, the preliminary correlation refers to the degree of correlation existing between each key feature obtained after performing a correlation analysis on the key features according to the key feature vector. The correlation analysis can be determined by calculating the cosine value between the key feature vectors.
[0150] In this embodiment, the relevant data group refers to clustering each key feature according to the obtained preliminary correlation, that is, classifying the data with correlation into one category.
[0151] In this embodiment, the analogical display refers to displaying the relevant data groups in the same coordinate system, so as to facilitate determining the association relationship and limiting relationship between different data, etc.
[0152] In this embodiment, the personalized change trend refers to the change situation corresponding to each key feature.
[0153] In this embodiment, the independent variable locking and analysis refers to performing a single variable analysis on each key feature, that is, when analyzing the change situation of one variable, fixing the values of other variables, so as to facilitate determining the change situation of the current variable.
[0154] In this embodiment, the multi-dimensional relative change relationship refers to the mutual influence relationship between each key feature and the remaining key features.
[0155] In this embodiment, the multi-dimensional association rule refers to multiple different association relationships existing among key features.
[0156] In this embodiment, the confidence interval is determined according to the requirements of business analysis, can be adjusted, and is used as a reference basis for screening multi-dimensional association rules.
[0157] In this embodiment, the frequent item set refers to the association relationships that appear more frequently in the multi-dimensional association rules.
[0158] In this embodiment, determining the relative magnitude relationship between the multi-dimensional association rule and the confidence interval based on the frequent item set is used to determine whether different association rules meet the confidence interval, so as to screen out strong association rules from the multi-dimensional association rules.
[0159] The working principle and beneficial effects of the above technical solution are as follows: By initially classifying the preprocessed archival information resources according to the visual symbol representation of the key information annotation results, and determining the key features from the key information of each category of archival information resources after the initial classification, it provides convenience for classifying the archival information resources. Secondly, converting the key features into key feature vectors, realizing the determination of the initial correlation between key features according to the key feature vectors, realizing the further clustering of key features according to the initial correlation, and also facilitating the analysis of the potential association relationships between key features according to the clustering results. Finally, analogously displaying the relevant data in a rectangular coordinate system, realizing the accurate and effective determination of strong association rules between key features according to the analog display results, realizing the further determination of the classification criteria for archival information resources according to the strong association rules, and ultimately realizing the classification and integration of archival information resources according to the classification criteria, improving the reliability, accuracy, and effectiveness of the management of diverse new archival information resources.
[0160] Embodiment 6:
[0161] Based on Embodiment 5, this embodiment provides an efficient management system for data-driven diverse new archival information resources. The data classification and integration unit includes:
[0162] A result acquisition subunit, used for:
[0163] Obtaining the classification criteria for archival information resources, and classifying the archival information resources based on the classification criteria;
[0164] Verifying the content of each category of archival information resources based on the classification results, removing duplicates from each category of archival information resources based on the content verification results, and merging the archival information resources with strong association rules;
[0165] A visualization subunit, configured to generate an associated topological structure path of the archival information resources of each category and an association path between the archival information resources of different categories based on the merging result, and generate a topological structure diagram of the archival information resources based on the associated topological structure path and the association path.
[0166] In this embodiment, content verification refers to verifying the normativity of the content of the archival information resources of each category.
[0167] In this embodiment, the associated topological structure path refers to the association relationship existing between the archival information resources of each category.
[0168] In this embodiment, the association path refers to the association relationship between the archival information resources of different categories.
[0169] The working principle and beneficial effects of the above technical solution are as follows: By classifying the archival information resources according to the classification criteria and verifying the content of the archival information resources of each category according to the classification result, the reliability of the final archival information resources is ensured. Finally, after the verification is passed, the associated topological structure path of the archival information resources of each category and the association path between the archival information resources of different categories are determined, and the topological structure diagram of the archival information resources is generated according to the associated topological structure path and the association path, so as to facilitate a quick and effective understanding of the association situation between different archival information resources.
[0170] Embodiment 7:
[0171] Based on Embodiment 1, this embodiment provides a data-driven diversified new type of archival information resource efficient management system. The data management module includes:
[0172] A data specification unit, configured to:
[0173] Obtain the basic parameters of the preset storage space. At the same time, obtain the basic structure parameters of each archival information resource, and compare the differences in the basic structure parameters of each archival information resource to obtain a structure difference parameter;
[0174] Determine the compatibility characteristics between each archival information resource and the preset storage space based on the basic parameters of the preset storage space and the structure difference parameter, and determine the unified data storage format for each archival information resource based on the compatibility characteristics;
[0175] A data storage unit, configured to perform format conversion on the classified and integrated archival information resources based on the unified data storage format, and perform spatial partitioning and storage of the classified and integrated archival information resources in the preset storage space based on the format conversion result.
[0176] In this embodiment, the preset storage space is set in advance and is an area for storing archival information resources.
[0177] In this embodiment, the basic parameters refer to the size of the preset storage space and the requirements for the data format, etc.
[0178] In this embodiment, the basic structure parameters refer to the structural conditions and data format conditions of each archival information resource.
[0179] In this embodiment, the structural difference parameters refer to the structural differences existing between the basic structure parameters of each archival information resource.
[0180] In this embodiment, the compatibility feature refers to the requirements that the preset storage space can effectively meet when storing different archival information resources, that is, environmental configuration parameters, etc.
[0181] In this embodiment, space division storage means dividing the preset storage space according to the categories of archival information resources, that is, storing each category of archival information resources in the same area.
[0182] The working principle and beneficial effects of the above technical solution are as follows: By determining the basic parameters of the preset storage space and the basic structure parameters of each archival information resource, effective parameter configuration of the preset storage space is achieved according to the basic parameters and basic structure parameters, format conversion and storage of various categories of archival information resources are achieved according to a unified data storage format, ensuring the reliability of the storage management of diverse new archival information resources.
[0183] Embodiment 8:
[0184] Based on Embodiment 1, this embodiment provides an efficient management system for data-driven diverse new archival information resources. The data management module includes:
[0185] A retrieval mechanism construction unit, which is used for:
[0186] Determining the storage addresses of each archival information resource based on the storage results, and at the same time, extracting the data characteristics of each archival information resource, and generating retrievable indicators for each archival information resource based on the storage addresses and data characteristics;
[0187] Determining the allowable data retrieval conditions based on the preset data retrieval service, and mapping the allowable data retrieval conditions with the retrievable indicators;
[0188] Invoking a standard data retrieval engine based on big data, and performing personalized parameter configuration on the script file of the standard data retrieval engine based on the mapping result to obtain a retrieval mechanism;
[0189] A data management unit, which is used for:
[0190] Obtain the retrieval requirements, analyze the retrieval requirements based on the retrieval mechanism to obtain the retrieval target of the stored archival information resources, and extract the encryption level of the retrieval target;
[0191] Determine the allowable multi-dimensional utilization methods of the retrieval target based on the encryption level, and open the utilization permissions of the retrieval target based on the allowable multi-dimensional utilization methods.
[0192] In this embodiment, the storage address refers to the specific storage address information corresponding to each archival information resource in the storage space.
[0193] In this embodiment, the data characteristics refer to the core content and business characteristics of each archival information resource, etc.
[0194] In this embodiment, the retrievable index refers to the label for positioning each archival information resource, that is, the corresponding archival information resource can be quickly located through the retrievable index.
[0195] In this embodiment, the preset data retrieval service is set in advance.
[0196] In this embodiment, the allowable data retrieval conditions refer to the specific requirements and events that can perform retrieval.
[0197] In this embodiment, the standard data retrieval engine is retrieved through big data and is a tool for performing retrieval.
[0198] In this embodiment, the personalized parameter configuration of the script file of the standard data retrieval engine based on the mapping result refers to configuring the background operation data of the standard data retrieval engine through the mapping result of the allowable data retrieval conditions and the retrievable index, ensuring that the standard data retrieval engine can effectively perform the corresponding retrieval tasks when working, and also facilitating the quick positioning of the archival information resources to be retrieved.
[0199] In this embodiment, the retrieval target refers to the archival information resources that need to be accessed after analyzing the retrieval requirements.
[0200] In this embodiment, the allowable multi-dimensional utilization methods include online browsing and downloading, etc.
[0201] The working principle and beneficial effects of the above technical solution are as follows: By determining the storage addresses of each archival information resource and the data characteristics of each archival information resource, accurate and effective generation of retrievable indicators for each archival information resource is achieved based on the storage address and data characteristics. Secondly, the allowable data retrieval conditions are determined, and the allowable data retrieval conditions are mapped to the retrievable indicators, and personalized parameter configuration is performed on the script file of the standard data retrieval engine according to the mapping result, so as to accurately and effectively formulate the retrieval mechanism. Finally, the retrieval requirements are analyzed through the retrieval mechanism to quickly and effectively determine the retrieval target. At the same time, the encryption level of the retrieval target is determined, so as to determine the allowable multi-dimensional utilization mode of the retrieval target according to the encryption level, and effectively manage the utilization mode of the retrieval target according to the allowable multi-dimensional utilization mode.
[0202] Embodiment 9:
[0203] This embodiment provides a method for efficiently managing diverse new archival information resources driven by big data, as Figure 3 shown, including:
[0204] Step 1: Receive archival information resources of different categories based on a data interface, and perform preprocessing and key information annotation on the archival information resources;
[0205] Step 2: Analyze the preprocessed archival information resources based on the key information annotation results, extract the data characteristics of the archival information resources, and classify and integrate the archival information resources based on the data characteristics;
[0206] Step 3: Store the archival information resources based on the classification and integration results, construct a retrieval mechanism based on the storage results, and open up multi-dimensional utilization modes for the stored archival information resources based on the construction results.
[0207] The working principle and beneficial effects of the above technical solution are as follows: By receiving and preprocessing archival information resources of different categories, it is convenient to integrate and process different types of archival information resources. At the same time, key information annotation is performed on the archival information resources, which provides convenience for managing diverse new archival information resources. Secondly, by analyzing the archival information resources based on the key information annotation results, accurate and effective determination of data characteristics such as the correlation relationship and change law between archival information resources is achieved, so as to effectively classify and integrate the archival information resources according to the obtained data characteristics, improving the processing efficiency and reliability of diverse new archival information resources. Finally, the archival information resources are stored according to the classification and integration results, and a retrieval mechanism is constructed, so as to quickly locate the required archival information resources according to the user's needs through the retrieval mechanism and effectively manage the utilization mode, improving the convenience of information retrieval and utilization and ensuring the management effect of archival information resources.
[0208] Example 10:
[0209] This embodiment provides a data-driven diversified new type of efficient management medium for archival information resources, on which a computer program is stored:
[0210] When the computer program is executed by a processor, it implements the steps of the data-driven diversified new type of efficient management method for archival information resources as described in any one of Examples 1 to 8.
[0211] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A highly efficient management system for diverse new archival information resources driven by big data, characterized in that, Including: A data receiving and preprocessing module, which is used to receive archive information resources of different categories based on a data interface, and preprocess the archive information resources and annotate key information; A data analysis module, which is used to analyze the preprocessed archive information resources based on the key information annotation results, extract the data characteristics of the archive information resources, and classify and integrate the archive information resources based on the data characteristics; A data management module, which is used to store the archive information resources based on the classification and integration results, construct a retrieval mechanism based on the storage results, and enable multi-dimensional utilization methods for the stored archive information resources based on the construction results; Among them, the data analysis module includes: A data analysis unit, which is used to: Obtain the key information annotation results, extract the visual symbol representations of the key information annotation results, and preliminarily classify the preprocessed archive information resources based on the visual symbol representations; Determine the data distribution characteristics of each category of archive information resources based on the preliminary classification results, and select key features from the key information based on the data distribution characteristics; Perform vector conversion on the key features to obtain key feature vectors, and perform correlation analysis on the key features based on the key feature vectors to obtain the preliminary correlations between the key features; A data feature determination unit, which is used to: Cluster the key features based on the preliminary correlations to obtain relevant data groups, and analogously display the relevant data groups in a rectangular coordinate system; Determine the personalized change trends of the key features in the relevant data groups based on the analog display results, and sequentially perform independent variable locking analysis on the key features based on the personalized change trends to obtain the multi-dimensional relative change relationships between each key feature and the remaining key features; Obtain the multi-dimensional association rules between the key features in the relevant data groups based on the multi-dimensional relative change relationships, and determine the confidence interval based on the business analysis requirements; Extract the frequent item sets of the key features under the multi-dimensional association rules, and determine the relative size relationship between the multi-dimensional association rules and the confidence interval based on the frequent item sets; Determine the association rules that meet the confidence interval as strong association rules, and obtain the data characteristics of the archive information resources based on the strong association rules; A data classification and integration unit, which is used to determine the classification criteria for the archive information resources based on the data characteristics, and classify and integrate the archive information resources based on the classification criteria.
2. The efficient management system for diverse new archive information resources driven by big data according to claim 1, wherein The data receiving and preprocessing module includes: A data source docking unit, which is used to: Determine the allowable management resource types based on management requirements, and determine the allowable docking sources based on the allowable management resource types; Perform interface docking with the allowable docking sources based on the data interface, and extract the basic attributes of each allowable docking source based on the interface docking results; A data receiving unit, which is used to: Analyze the basic attributes to obtain the data categories corresponding to each allowable docking source, and determine the different requirements for the basic communication parameters of the data interface based on the different data categories; Unify and debug the basic communication parameters of different data categories based on the different requirements, and receive the archive information resources of different allowable docking sources based on the debugging results.
3. The efficient management system for diversified new archive information resources driven by big data according to claim 1, characterized in that, The data receiving and preprocessing module includes: A result obtaining unit, which is used to: Obtain the acquired archival information resources, and determine the data service structures of different categories of archival information resources based on the data sources of the archival information resources; Determine the key fields in the archival information resources and the corresponding structure conversion rules based on the data service structures, and perform structure assignment on the key fields based on the structure conversion rules to obtain structured archival information resources; The data preprocessing unit is used for: Visually display the structured archival information resources in a visual coordinate system, and determine the local outliers in the structured archival information resources based on the visual display; Remove the data points of the local outliers, fit the structured archival information resources based on the data point removal results, and perform data standardization processing on the fitting results based on a preset standardization interval to complete the preprocessing of the archival information resources.
4. A highly efficient management system for diverse new archival information resources driven by big data according to claim 1, characterized in that, The data reception and preprocessing module includes: The result analysis unit is used to obtain the preprocessed archival information resources and determine the business characteristics of each category of archival information resources based on the data sources of the archival information resources; The key information determination unit is used for: Determine the information indicators for extracting key information from each category of archival information sources based on the business characteristics. At the same time, segment the archival information sources based on the natural language system, and perform part-of-speech tagging on each target word based on the segmentation results; Extract the semantics of each target word, associate the semantics of each target word based on the part-of-speech tagging results, and extract key information from the association results based on the information indicators; The annotation unit is used to perform differential annotation on the key information of different categories of archival information resources based on a preset annotation strategy.
5. A data-driven diversified new type of efficient management system for archival information resources according to claim 1, characterized in that, The data classification and integration unit includes: The result acquisition subunit is used for: Obtain the classification criteria for the archival information resources, and classify the archival information resources based on the classification criteria; Verify the content of each category of archival information resources based on the classification results, remove duplicates from each category of archival information resources based on the content verification results, and merge the archival information resources with strong association rules; The visualization subunit is used to generate the associated topological structure paths of each category of archival information resources and the association paths between different categories of archival information resources based on the merging results, and generate a topological structure diagram of the archival information resources based on the associated topological structure paths and association paths.
6. The data-driven diversified new type of archival information resource efficient management system according to claim 1, characterized in that, The data management module includes: The data specification unit is used for: Obtain the basic parameters of the preset storage space. At the same time, obtain the basic structure parameters of each archival information resource, and compare the differences in the basic structure parameters of each archival information resource to obtain the structure difference parameters; Determine the compatibility characteristics between each archival information resource and the preset storage space based on the basic parameters of the preset storage space and the structure difference parameters, and determine the unified data storage format for each archival information resource based on the compatibility characteristics; The data storage unit is used to perform format conversion on the classified and integrated archival information resources based on the unified data storage format, and perform spatial partition storage of the classified and integrated archival information resources in the preset storage space based on the format conversion results.
7. A data-driven diversified new type of efficient management system for archival information resources according to claim 1, characterized in that, The data management module includes: The retrieval mechanism construction unit is used for: Determine the storage addresses of each archival information resource based on the stored results. At the same time, extract the data characteristics of each archival information resource, and generate retrievable indicators for each archival information resource based on the storage addresses and data characteristics; Determine the allowable data retrieval conditions based on the preset data retrieval service, and map the allowable data retrieval conditions to the retrievable indicators; Invoke the standard data retrieval engine based on big data, and perform personalized parameter configuration on the script file of the standard data retrieval engine based on the mapping result to obtain the retrieval mechanism; The data management unit is used for: Obtain the retrieval requirement, analyze the retrieval requirement based on the retrieval mechanism to obtain the retrieval target of the stored archival information resource, and extract the encryption level of the retrieval target; Determine the allowable multi-dimensional utilization method of the retrieval target based on the encryption level, and open the utilization permission of the retrieval target based on the allowable multi-dimensional utilization method.
8. A big data-driven efficient management method for diverse new archival information resources of a data-driven diverse new archival information resource efficient management system according to claim 1, characterized in that, Including: Step 1: Receive archival information resources of different categories based on the data interface, and perform preprocessing and key information annotation on the archival information resources; Step 2: Analyze the preprocessed archival information resources based on the key information annotation results, extract the data characteristics of the archival information resources, and classify and integrate the archival information resources based on the data characteristics; Step 3: Store the archival information resources based on the classification and integration results, construct a retrieval mechanism based on the storage results, and open a multi-dimensional utilization method for the stored archival information resources based on the construction results.
9. A data-driven diversified new type of archival information resource efficient management medium, on which a computer program is stored, characterized in that: When the computer program is executed by a processor, it realizes the steps of the data-driven diversified new type of archival information resource efficient management method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data source management method and architecture of data connection platform
CN116226458A
Archive data storage system based on big data
CN117725283A