Government affair integrated data acquisition and storage management system

Through the integrated government data collection and storage management system, the problems of unstandard data processing, unreasonable resource allocation and insufficient attention to historical document data in traditional government data management are solved, and the efficient, accurate collection and long-term storage of data are achieved, and the system's response speed and adaptability are improved.

CN120407529AActive Publication Date: 2025-08-01SHANDONG SHUIFA ZIGUANG BIG DATA CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510911895.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-01
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

There are problems in traditional government data management systems with lack of standardization of data processing processes, unreasonable allocation of storage resources, lack of focus on historical document data, and imperfect dynamic management and maintenance mechanisms, resulting in low data integrity and accuracy and insufficient system response speed and flexibility.

Method used

The integrated government data acquisition and storage management system is adopted, and through precise time and order marking, the data acquisition sequence and storage resource allocation are optimized using entropy weight method and linear planning algorithm, and the genetic algorithm and supervisory learning model are used for dynamic management to ensure efficient collection and long-term preservation of historical document data.

Benefits of technology

It improves the integrity and traceability of government data, reduces duplicate work, saves costs, ensures the security and convenient access of historical document data, and realizes the forward-looking and flexible system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407529A_ABST
    Figure CN120407529A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information, in particular to a government affair integrated data acquisition and storage management system. Comprising the following parts: a data acquisition and primary processing module for acquiring processing time and a processing sequence of N pieces of government affair data, and marking the processing time of the N pieces of government affair data according to the processing sequence of the N pieces of government affair data; and for the historical literature data, processing time and a processing sequence of the historical literature data are recorded in detail. According to the invention, through accurate time and sequence marking, the accuracy and traceability of government affair data are improved; the standardized data integrity evaluation optimizes the acquisition process and storage resource allocation, reduces the repeated work and saves the cost; long-term storage and convenient access of historical literature data are particularly concerned, and safety and effective utilization of the historical literature data are ensured; by predicting a future data access mode, flexible and prospective management is realized, and the response speed and adaptability of the system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and particularly to an integrated government affairs data collection and storage management system. Background Art

[0002] In recent years, with the rapid development of information technology and the advancement of government digital transformation, the scale and complexity of government affairs data have been increasing continuously. The traditional data collection and storage management methods are difficult to meet the needs of the efficient operation of modern governments. Especially when dealing with historical document data, these data not only carry important historical information but also have irreplaceable value for policy-making, public services, and social research. However, due to the unique nature of historical document data, such as large data volume, diverse formats, high preservation requirements, etc., its collection, storage, and management face many challenges. In actual operation, government agencies need to process a large amount of government affairs data, including daily administrative records, regulatory documents, and statistical reports. The collection and storage of these data must ensure efficiency, accuracy, and traceability to support government decision-making and service provision. At the same time, as a part of government affairs data, historical document data, due to its importance and particularity, requires a specially designed method to ensure its long-term preservation and effective utilization.

[0003] Currently, many government affairs data management systems have the following problems: The data processing process lacks standardization, resulting in chaotic data collection order, which affects the integrity and accuracy of the data. The storage resource allocation is unreasonable, causing low utilization rate of storage media and increasing the operation cost. There is a lack of focused attention on historical document data, and the needs for historical document data cannot be effectively guaranteed. The dynamic management and maintenance mechanism is not perfect, and the system is difficult to adapt to the constantly changing data access patterns, reducing the response speed and flexibility of the system.

[0004] To solve the above problems, there is an urgent need for an integrated data collection and storage management system. Summary of the Invention

[0005] In order to overcome the disadvantages of low data processing accuracy and poor practicability in existing government affairs management systems, the present invention provides an integrated government affairs data collection and storage management system.

[0006] The technical implementation solution of the present invention is as follows: An integrated government affairs data collection and storage management system includes the following parts: Data Acquisition and Preliminary Processing Module: Obtain the processing time and processing order of N government affairs data, and mark the processing time of N government affairs data according to the processing order of N government affairs data; for historical literature data, record in detail the processing time and processing order of historical literature data; determine the acquisition order of historical literature data according to the processing order of N government affairs data; the acquisition order of the historical literature data is to separately mark the acquisition order of historical literature in government affairs data; the historical literature data belongs to a part of the government affairs data. Data Evaluation and Resource Allocation Module: Use the government affairs data acquisition integrity formula based on the entropy weight method to evaluate the integrity of government affairs data, and obtain the evaluation result of government affairs data acquisition integrity; and adjust the acquisition order of government affairs data according to the evaluation result of government affairs data acquisition integrity. Based on the adjusted acquisition order of government affairs data, determine the acquisition and update order of historical literature data; according to the evaluation result of government affairs data acquisition integrity, reasonably allocate storage media and storage space for N government affairs data, and optimize the storage space using the linear programming algorithm. Storage Management and Optimization Module: Mark the storage media and remaining space of all government affairs data; specifically mark the storage situation of historical literature data; optimize the storage media and storage space of historical literature data using the genetic algorithm according to the storage situation of other government affairs data. Dynamic Management and Maintenance Module: Implement dynamic management of government affairs data using the long-term validity formula of government affairs data according to the acquisition order, storage media, and storage space of government affairs data, and use the supervised learning model to predict the future government affairs data access pattern and adjust the government affairs data distribution.

[0007] Preferably, the Data Acquisition and Preliminary Processing Module: Obtain the processing time and processing order of N government affairs data, and mark the processing time of N government affairs data according to the processing order of N government affairs data; for historical literature data, record in detail the processing time and processing order of historical literature data, including: Obtain the processing time and processing order of N government affairs data; take the processing time as the first factor and the processing order as the second factor, multiply them after normalizing the first factor and the second factor to obtain the comprehensive scoring result of N government affairs data. Based on the comprehensive scoring result, determine the acquisition order of government affairs data according to the principle that the sum of adjacent two items is the smallest. Obtain the finally determined acquisition order of government affairs data, separately mark the historical literature data, and record in detail the processing time and processing order of historical literature data in the acquisition order of government affairs data.

[0008] Preferably, the determination of the acquisition order of historical literature data according to the processing order of N government affairs data includes: Mapping storage space for the historical document data based on the processing time and processing order of the historical document data; said mapping storage space is to allocate virtual storage space for the historical document data; Based on the initial collection order of the historical document data, mapping the storage space validity for the historical document data in the initial collection order; the mapping storage space validity is to allocate virtual storage time to the historical document data; A linear programming algorithm is used to optimize the allocated virtual storage space and virtual storage time.

[0009] Preferably, the data evaluation and resource allocation module: uses the government data collection integrity formula based on the entropy weight method to evaluate the integrity of the government data, and obtains the government data collection integrity evaluation result; and adjusts the collection order of the government data according to the government data collection integrity evaluation result, and determines the collection and update order of the historical document data based on the adjusted government data collection order, including: the government data collection integrity formula is as follows, , in, The collection completeness of government data, range: [0,1]; is the importance weight factor of historical document data; To adjust the parameters; is the normalized processing time function; The sequence number of the government data currently being processed; The total number of times government data needs to be processed; For the The time it takes to complete the processing of government data; is the standard deviation of the processing completion time for all government data; The maximum number of remaining processing times.

[0010] Preferably, the method of reasonably allocating storage media and storage space for N government data based on the government data collection integrity assessment result and optimizing the storage space using a linear programming algorithm includes: Based on the government data collection integrity assessment results, the government data are sorted from low to high in terms of collection integrity and storage media and storage space are allocated; Obtain virtual storage space and virtual storage time of historical document data; Extracting government data that complies with virtual storage space and virtual storage time; The estimated storage time and occupied space of government data that conforms to the virtual storage space and virtual storage time are obtained as an alternative storage space for historical document data.

[0011] Preferably, the storage medium and remaining space for marking all government data include: If the alternative storage space for historical document data is less than or equal to the predefined alternative storage space threshold, then integrate the remaining storage space of other government affairs data; If the alternative storage space for historical document data is greater than the predefined alternative storage space threshold, then directly allocate the storage medium for the alternative storage space.

[0012] Preferably, the storage situation of the specially marked historical document data is described; according to the storage situation of other government affairs data, optimize the storage medium and storage space of historical document data, including: Based on the storage medium of the alternative storage space, preferentially allocate the storage medium with the longest storage aging time for historical document data, and synchronously allocate the storage medium with the shortest storage aging time as the temporary storage medium.

[0013] Preferably, the dynamic management and maintenance module: according to the collection order, storage medium and storage space of government affairs data, implement dynamic management of government affairs data using the long-term validity formula of government affairs data, including: The long-term validity formula of government affairs data is as follows, , where, is the long-term validity of government affairs data, range: [0,1]; is the collection integrity of government affairs data calculated through the government affairs data collection integrity formula, range: [0,1]; is a regulation parameter; is the time difference in the process of collecting government affairs data.

[0014] Preferably, the implementation of dynamic management includes: Obtain the collection order, storage medium and storage space of government affairs data; Taking the collection integrity of historical document data as the center point of the collection sequence, use the K-means clustering algorithm for clustering to obtain the collection sequence clustering result; Taking the long-term validity of historical document data as the center point of the storage sequence, use the K-means clustering algorithm for clustering to obtain the storage sequence clustering result; Use the Euclidean distance formula to calculate the similarity; Perform superposition adjustment according to the collection sequence clustering result and the storage sequence clustering result to obtain the superposition adjustment result; Determine the dynamic management method based on the superposition adjustment result.

[0015] Preferably, the determination of the dynamic management method based on the superposition adjustment result includes: Based on the clustering results of the collected sequences and the clustering results of the stored sequences, superimpose adjacent clustering points. If the superimposed result is less than the set threshold, trigger the dynamic adjustment mechanism. If the superimposed result is greater than or equal to the set threshold, maintain the current configuration.

[0016] Beneficial effects: Through precise time and sequence marking, the present invention ensures that all government affairs data, especially historical literature data, can be efficiently and accurately collected and recorded, enhancing data integrity and traceability; adjusts the data collection order through a standardized data integrity assessment method and reasonably allocates storage resources according to the assessment results, reducing duplicate work and saving time and costs; pays special attention to the long-term preservation and fast retrieval requirements of historical literature data, and guarantees the security and convenient access of these important data through mapping storage space and effectiveness allocation; realizes the forward-looking and flexibility of data distribution by predicting future data access patterns, improving the response speed and adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic structural diagram of the government affairs integration data collection and storage management system of the present invention; Figure 2 is a schematic flowchart of implementing dynamic management of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] A government affairs integration data collection and storage management system, as Figure 1 and Figure 2 shown, includes the following parts: Data collection and preliminary processing module: Obtain the processing time and processing order of N government affairs data, and mark the processing time of N government affairs data according to the processing order of N government affairs data; for historical literature data, record the processing time and processing order of historical literature data in detail; determine the collection order of historical literature data according to the processing order of N government affairs data; the collection order of the historical literature data is to separately mark the collection order of historical literature in government affairs data; the historical literature data belongs to a part of the government affairs data; Data Evaluation and Resource Allocation Module: Evaluate the integrity of government affairs data using the government affairs data collection integrity formula based on the entropy weight method to obtain the evaluation result of the government affairs data collection integrity; and adjust the collection order of government affairs data according to the evaluation result of the government affairs data collection integrity. Based on the adjusted collection order of government affairs data, determine the collection and update order of historical literature data; according to the evaluation result of the government affairs data collection integrity, reasonably allocate storage media and storage space for N pieces of government affairs data, and optimize the storage space using the linear programming algorithm. Storage Management and Optimization Module: Mark the storage media and remaining space of all government affairs data; specifically mark the storage status of historical literature data; optimize the storage media and storage space of historical literature data using the genetic algorithm according to the storage status of other government affairs data. Dynamic Management and Maintenance Module: Implement dynamic management of government affairs data using the long-term validity formula of government affairs data according to the collection order, storage media, and storage space of government affairs data, and use the supervised learning model to predict the future government affairs data access pattern and adjust the government affairs data distribution.

[0020] Data Collection and Preliminary Processing Module: Obtain the processing time and processing order of N pieces of government affairs data, and mark the processing time of N pieces of government affairs data according to the processing order of N pieces of government affairs data; for historical literature data, record in detail the processing time and processing order of historical literature data, including: Obtain the processing time and processing order of N pieces of government affairs data; take the processing time as the first factor and the processing order as the second factor, multiply them after normalization, and obtain the comprehensive scoring result of N pieces of government affairs data. Based on the comprehensive scoring result, determine the collection order of government affairs data according to the principle that the sum of adjacent two items is the smallest. Obtain the finally determined collection order of government affairs data, mark historical literature data separately, and record in detail the processing time and processing order of historical literature data in the government affairs data collection order.

[0021] Further explanation: Automatically capture the processing time (actual time consumption) and processing order (execution sequence) of government affairs data, and collect them in real time through log monitoring or time series database. Historical literature data synchronously records its independent processing time series.

[0022] Multi-factor Comprehensive Scoring: Multiply the processing time (efficiency) and processing order (priority) after normalization to generate a comprehensive score (for example: short time and high order result in a high score).

[0023] Dynamic Sorting Optimization: Rearrange the collection order according to the principle of minimizing the sum of scores of adjacent government affairs data to reduce the overall processing delay (similar to the shortest job first algorithm).

[0024] Historical data marking: Individually mark historical literature data in the global sequence and record its precise time series to ensure long-term traceability.

[0025] Example, assume 5 government affairs data (including 2 historical literature data H1, H2): Original data: Processing time: 10s, 8s, 15s, 20s, 5s; Processing order: 2, 1, 3, 5, 42, 1, 3, 5, 4. Normalization and scoring: Time normalization: 0.5, 0.4, 0.75, 1, 0.25; Order normalization: 0.4, 0.2, 0.6, 1, 0.8; Comprehensive score: Time × Order → 0.2, 0.08, 0.45, 1, 0.2. Optimize the acquisition order: Sort according to the principle of the smallest adjacent sum: Data 2(0.08) → Data 1(0.2) → Data 5(0.2) → Data 3(0.45) → Data 4(1). Historical data marking: If H1 = Data 3 and H2 = Data 4, then mark their positions and processing times (15s, 20s).

[0026] Determine the acquisition order of historical literature data according to the processing order of N government affairs data, including: Based on the processing time and processing order of historical literature data, map storage spaces for historical literature data; The mapped storage space is to allocate virtual storage spaces for historical literature data; Based on the initial acquisition order of historical literature data, map the validity of storage spaces for historical literature data in the initial acquisition order; The validity of the mapped storage space is to allocate virtual storage times for historical literature data; Use the linear programming algorithm to optimize the allocated virtual storage space and virtual storage time.

[0027] Further explanation: Storage space mapping: Allocate virtual storage blocks (logical addresses) according to the processing time series of historical literature data to achieve physical storage decoupling.

[0028] Validity mapping: Allocate virtual storage time windows according to the initial acquisition order (such as the validity period [from January 1, 2025 to December 31, 2030]) to define the life cycle.

[0029] Linear programming optimization: Take the minimum storage cost as the objective function, virtual space and time as decision variables, and hardware capacity as constraint conditions to solve the optimal allocation.

[0030] Example, 5 government affairs data (including H1, H2): Virtual mapping: H1 is allocated virtual block B1 (capacity 50GB), and the validity period is up to 2030; H2 is allocated B2 (80GB), and the validity period is up to 2035. Conflict problem: B1 shares the physical disk with other government affairs data block C3 (40GB), but the validity period of C3 is only up to 2028.

[0031] Linear programming algorithm model construction: variables: Represents government data On storage media Capacity allocated on the server (GB); Goal: Minimize total cost ,in It is a storage medium Unit capacity cost (yuan / GB); Constraints: Media capacity limit: ( For medium Maximum capacity, GB); data capacity requirements: ( For data Total demand, GB); fixed allocation for historical data: ( exist Media fixed occupancy 50GB); other data limits: ( exist Media usage ≤ 40GB); non-negative constraints: ; Solve and execute: Solve the model to obtain the optimal solution. If the result requires and , then from Migrate to .

[0032] Data evaluation and resource allocation module: Use the government data collection integrity formula based on the entropy weight method to evaluate the integrity of government data and obtain the government data collection integrity evaluation result; and adjust the collection order of government data according to the government data collection integrity evaluation result, and determine the collection and update order of historical document data based on the adjusted government data collection order, including: The government data collection integrity formula is as follows, , in, The collection completeness of government data, range: [0,1]; is the importance weight factor of historical document data; is the adjustment parameter; is the normalized processing time function; The sequence number of the government data currently being processed; The total number of times government data needs to be processed; For the The time it takes to complete the processing of government data; is the standard deviation of the processing completion time for all government data; The maximum number of remaining processing times.

[0033] For further illustration, the steps to adjust the collection order of government affairs data are as follows: Obtain the evaluation results of the integrity of government affairs data collection; Sort the government affairs data in ascending order of integrity; Adjust the collection order of government affairs data based on the sorting results.

[0034] The steps to determine the collection and update order of historical literature data are as follows: Obtain the virtual storage space and virtual storage time of historical literature data; Extract the government affairs data in the government affairs data that meet the virtual storage space and virtual storage time of historical literature data; Obtain the estimated storage time and estimated storage occupancy space of the government affairs data that meet the requirements; Use the estimated storage time and estimated storage occupancy space as the alternative storage space for historical literature data; Determine the collection and update order of historical literature data based on the alternative storage space.

[0035] Objectively quantify the weight of historical literature data through the entropy weight method ( ), combined with temporal characteristics (processing order , time-consuming ) and global volatility ( ), and dynamically evaluate the integrity of data collection ( Rising = rising integrity).

[0036] Formula decomposition: Numerator: : The weight of historical data calculated by the entropy weight method; : Processing efficiency factor ( The shorter → the larger the value); Denominator: : Penalty for remaining tasks (the more unprocessed data → The lower); : Global time-consuming volatility (the greater the volatility → The lower); Explanation of the entropy weight method, function: Objectively determine weight, avoid subjective deviation. Steps: Data matrix: Construct m government affairs data × n evaluation indicators (such as historical value, update frequency), Normalization: Perform positive / negative processing on the indicators, Entropy value calculation: ( = proportion of the indicator), Weight assignment: (entropy value decreases → weight increases), is the number of government affairs data, is the proportion of the th data in the th indicator.

[0037] Example, calculation by the entropy weight method : The entropy value of historical literature indicators is lower ( = 0.2, =0.3 vs ordinary data e = 0.8) → =1.5, integrity Calculation (taking H1 as an example): =3, =5, =15s, =11.6s, =5.5s, =0.618, let =0.1 → 1 - exp(-0.1×0.618) = 0.06, =0.09 / 11 = 0.008, Sequence adjustment: The W value of H1 is the lowest → The priority is promoted to the first place in the update sequence.

[0038] Indicates the relationship between the current processing time and the processing order, and the calculation formula is: , Where Is the average value of all processing completion times, Is the standard deviation, standardizing the processing time to ensure that the time differences of different processing tasks do not cause result distortion; Is from 1 to Of integers, indicating the processing order; Represents a positive integer, indicating the total amount of data to be processed; Represents the specific processing completion time, and the unit is selected according to the actual situation (such as seconds, minutes); Measures the variability of the processing time and reflects the dispersion degree of the processing time, Is the completion time of each processing task, Is the average completion time, Is the maximum value of the remaining processing times, preventing the denominator from being zero. When = When, the denominator is not zero, ensuring the rationality of the formula. If - >0, then use ; Otherwise use 1.

[0039] According to the evaluation results of the integrity of the government affairs data collection, reasonably allocate storage media and storage space for N pieces of government affairs data, including: Based on the evaluation results of the integrity of the government affairs data collection, sort the government affairs data from low to high in terms of collection integrity and allocate storage media and storage space; Obtain the virtual storage space and virtual storage time of the historical literature data; Extract the government affairs data that meet the virtual storage space and virtual storage time; The estimated storage time and occupied space of government data that conforms to the virtual storage space and virtual storage time are obtained as an alternative storage space for historical document data.

[0040] Further explanation is that integrity-driven sorting: by data integrity assessment results Sort values from low to high ( Decrease = Increased data quality risk), prioritize allocating high-reliability storage resources (such as SSDs) to low-integrity data.

[0041] Historical document storage matching: Utilize the preset virtual storage parameters (space / time) of historical documents to select storage blocks with matching spatiotemporal attributes in other government data as a candidate pool to ensure the long-term storage compatibility of historical data.

[0042] Example, 5 government data integrity rankings: H1( =0.008)→Data 2( =0.2)→Data1( =0.3)→Data 5( =0.5)→H2( =0.7), historical document H1 requirement: 50GB of virtual space, valid until 2030. Resource allocation process: Allocation by integrity: H1 (lowest integrity) is allocated to a high-speed SSD (100GB), and H2 is allocated to a regular hard drive (200GB). Alternative storage matching: Filter other data blocks that meet H1 requirements (≥50GB and valid ≥2030) → discover the storage block of data 5 (80GB, valid until 2035). Mark the data 5 block as the alternative storage pool for H1, and automatically switch when the primary SSD fails.

[0043] Storage Management and Optimization Module: Marks the storage media and remaining space for all government data, including: If the alternative storage space for historical document data is less than or equal to the pre-selected storage space threshold, the remaining storage space for other government data will be integrated; If the candidate storage space for historical document data is greater than the pre-selected storage space threshold, the storage medium with the selected storage space is directly allocated.

[0044] Further explanation is how to set the pre-device storage space threshold: Basic threshold setting: Determine the basic value based on the minimum storage capacity requirement of historical document data to ensure that its basic storage requirements are met; redundancy expansion: Add the system fault tolerance redundancy ratio (such as 10%-20%) to the basic value to form the final threshold to cope with storage fluctuations or sudden demands.

[0045] Alternative storage space verification: Verify whether the alternative storage space for historical document data meets the requirements (capacity ≥ requirement). Trigger cross-block resource integration when it is insufficient.

[0046] Dynamic resource expansion: Insufficient alternatives: Scan the remaining space of non-alternative storage blocks (such as fragmented space released by other government affairs data) and combine for expansion. Sufficient alternatives: Directly enable the storage media in the alternative pool to avoid resource waste.

[0047] Example: H1 requires 50GB, and only 30GB is available in its alternative pool (data block 5) → insufficient. Remaining space of other government affairs data: Data 2 (20GB) + Data 1 (15GB). Operation process: Extract fragmented space: Combine the remaining space of Data 2 (20GB) and Data 1 (15GB) → Generate a new 35GB block, combined with the 30GB in the alternative pool → Total 65GB (>50GB requirement). Allocation and execution: Primary storage: SSD 100GB (originally allocated), Auxiliary storage: 65GB hybrid fragmented block (emergency backup).

[0048] Specially mark the storage situation of historical document data; According to the storage situation of other government affairs data, use genetic algorithm to optimize the storage media and storage space of historical document data, including: Based on the storage media of the alternative storage space, preferentially allocate the storage media with the longest storage time limit for historical document data, and synchronously allocate the storage media with the shortest storage time limit as the temporary storage media.

[0049] For further explanation, dual-track storage strategy: Primary storage: Select the medium with the longest storage time limit in other government affairs data (such as optical disc / tape) to ensure long-term preservation of historical documents.

[0050] Temporary storage medium: Synchronously occupy the block with the shortest storage time limit (such as memory / temporary SSD) to achieve fast response for high-frequency access.

[0051] Genetic algorithm optimization: Take minimizing storage cost and minimizing access latency as the objective functions, and iteratively solve the optimal combination of primary storage and temporary storage media through crossover and mutation.

[0052] Example: Storage time limit of other government affairs data: Data 5 (optical disc, valid until 2035) → longest, Data 1 (SSD, valid until 2028) → shortest. H1 requires long-term preservation and supports frequent retrieval; Genetic Algorithm Operations: Initialize the population: Solution 1: Main memory = Data 5 optical discs, Temporary storage = Data 1 SSD; Solution 2: Main memory = Data 2 magnetic tapes, Temporary storage = Data 5 SSD. Iterative optimization: Calculate fitness (cost + latency weighted): Score of Solution 1 > Score of Solution 2. Crossover and mutation: Replace the temporary storage of Solution 1 with Data 3 memory (expiring in 2026). Output the optimal solution: Main storage: Data 5 optical discs (expiring in 2035), Temporary storage medium: Data 3 memory (expiring in 2026, high-frequency access buffer).

[0053] Dynamic Management and Maintenance Module: According to the collection order, storage medium, and storage space of government affairs data, dynamically manage government affairs data using the long-term validity formula for government affairs data, including: The long-term validity formula for government affairs data is as follows , where is the long-term validity of government affairs data, range: [0, 1]; is the collection integrity of government affairs data calculated through the government affairs data collection integrity formula, range: [0, 1]; is a tuning parameter; is the time difference during the collection process of government affairs data.

[0054] Further explanation is that is used to control the non-linear impact of time on validity, usually a positive number. In practical applications the value of should be selected according to industry standards or rules of thumb in a specific field. For example, selecting a small value reflects the long-term value of historical literature data, or selecting a large value emphasizes the importance of fresh data in a real-time decision support system; The unit is selected according to the actual situation (such as seconds, minutes) to measure the timeliness of data; As time increases, will tend to 1, meaning the validity of the data will gradually decrease; If is large, it means the data quickly loses its value over time; If is small, it means the value of the data changes slowly over time.

[0055] Example, H1 integrity = 0.008 (low), storage = 2 years, Historical literature parameter: = 0.01 (low decay), Validity calculation: = 0.008 * (1 - 0.9802) ≈ 0.00016, Dynamic response: Monitoring: < 0.001 threshold alarm, Maintenance action: Re-collect H1 data to improve (→ =0.05), migrate to a more stable storage medium ( down to 0.005), after the update : ≈0.0005.

[0056] Implement dynamic management, including: Obtain the collection order, storage media and storage space of government data; Taking the collection integrity of historical document data as the center point of the collection sequence, the K-means clustering algorithm is used to cluster and obtain the collection sequence clustering results; Taking the long-term validity of historical document data as the center point of the storage sequence, the K-means clustering algorithm is used to cluster and obtain the storage sequence clustering results; Similarity is calculated using the Euclidean distance formula; Performing superposition adjustment according to the acquisition sequence clustering result and the storage sequence clustering result to obtain a superposition adjustment result; A dynamic management method is determined based on the superimposed adjustment result.

[0057] Further explanation is that dual sequence collaborative management: sequence construction: collection sequence: government data processing time sequence flow (including integrity ), storage sequence: storage medium distribution status (including long-term validity ), historical data dual-center clustering: collection sequence based on historical documents As the cluster center (to ensure the priority of key data collection), the storage sequence is based on historical documents For cluster centers (to ensure long-term storage stability), dynamic adjustment mechanism: Euclidean distance quantifies the collection / storage cluster similarity, and the superposition result trigger strategy: reorganize the data distribution when the difference is too large.

[0058] Example, acquisition sequence: storage sequence: (H1 is located on the CD), clustering and adjustment: acquisition clustering (based on H1 =0.008 as the center): Cluster 1 (high priority): H1, data 2 (distance from the center < 0.1), storage clustering (based on H1's =0.95 as the center): Cluster 1 (long-term storage): CD, SSD (distance < 0.05), overlay analysis: data 2 in cluster 1 originally stored in memory ( =0.6), the distance from the storage center =8.2 (Euclidean distance), trigger action: migrate data 2 to the CD ( =0.3).

[0059] Determining a dynamic management method based on the superimposed adjustment result includes: Based on the clustering results of the acquisition sequence and the storage sequence, superimpose adjacent clustering points. If the superimposed result is less than the set threshold, trigger the dynamic adjustment mechanism; if the superimposed result is greater than or equal to the set threshold, maintain the current configuration. Use a supervised learning model to predict future government data access patterns and adjust the distribution of government data.

[0060] For further explanation, the dynamic management trigger mechanism: clustering superposition analysis: calculate the sum of the Euclidean distances between the clustering centers of the acquisition sequence and the storage sequence (the superimposed result ). (Threshold) indicates a mismatch between the data distribution and the acquisition strategy → trigger adjustment. Maintain the status quo, prediction intervention: The supervised learning model (such as LSTM) predicts future patterns based on historical access logs to suppress unnecessary adjustments.

[0061] The method for setting the threshold is to determine the basic threshold according to the minimum storage capacity requirement of historical literature data; superimpose the system fault tolerance redundancy ratio (10%-20%) on the basic value to form the final threshold.

[0062] Example, the superimposed result : The acquisition cluster center ( = 0.008) and the storage cluster center ( = 0.95), the distance = 8.5, the set threshold = 5 → > , operation process: threshold determination: = 8.5 > = 5 → do not trigger adjustment, prediction intervention: The LSTM model analyzes the historical access trend and predicts that the access volume of H1 will surge by 50% in the next 3 months. Active optimization: Migrate H1 from optical disc to SSD ( drops to 2.1, still > . but improves the response speed).

[0063] The above has introduced this application in detail. Specific examples are used in this article to elaborate on the principle and implementation method of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. An integrated government affairs data collection, storage and management system, characterized in that, It includes the following parts: Data acquisition and preliminary processing module: Obtain the processing time and processing order of N government affairs data, and mark the processing time of N government affairs data according to the processing order of N government affairs data; For historical literature data, record the processing time and processing order of historical literature data in detail; Determine the acquisition order of historical literature data according to the processing order of N government affairs data; The acquisition order of the historical literature data is to separately mark the acquisition order of historical literature in government affairs data; The historical literature data belongs to a part of the government affairs data; Data evaluation and resource allocation module: Use the government affairs data acquisition integrity formula based on the entropy weight method to evaluate the integrity of government affairs data, and obtain the government affairs data acquisition integrity evaluation result; And adjust the acquisition order of government affairs data according to the government affairs data acquisition integrity evaluation result. Based on the adjusted acquisition order of government affairs data, determine the acquisition update order of historical literature data; According to the government affairs data acquisition integrity evaluation result, reasonably allocate storage media and storage space for N government affairs data, and optimize the storage space using the linear programming algorithm; Storage management and optimization module: Mark the storage media and remaining space of all government affairs data; Especially mark the storage situation of historical literature data; According to the storage situation of other government affairs data, use the genetic algorithm to optimize the storage media and storage space of historical literature data; Dynamic management and maintenance module: According to the acquisition order, storage media and storage space of government affairs data, implement dynamic management of government affairs data using the long-term validity formula of government affairs data, and use the supervised learning model to predict the future government affairs data access pattern and adjust the government affairs data distribution.

2. A government affairs integration data collection and storage management system according to claim 1, characterized in that, The data acquisition and preliminary processing module: Obtain the processing time and processing order of N government affairs data, and mark the processing time of N government affairs data according to the processing order of N government affairs data; For historical literature data, record the processing time and processing order of historical literature data in detail, including: Obtain the processing time and processing order of N government affairs data; Take the processing time as the first factor and the processing order as the second factor, multiply them after normalizing the first factor and the second factor to obtain the comprehensive scoring result of N government affairs data; Based on the comprehensive scoring result, determine the acquisition order of government affairs data according to the principle that the sum of adjacent two items is the smallest; Obtain the finally determined acquisition order of government affairs data, separately mark the historical literature data, and record the processing time and processing order of historical literature data in the government affairs data acquisition order in detail.

3. A government affairs integration data collection and storage management system according to claim 1, characterized in that, The determination of the acquisition order of historical literature data according to the processing order of N government affairs data includes: Based on the processing time and processing order of historical literature data, map storage space for historical literature data; The mapped storage space is to allocate virtual storage space for historical literature data; Based on the initial acquisition order of historical literature data, map the storage space validity of historical literature data in the initial acquisition order; The mapped storage space validity is to allocate virtual storage time for historical literature data; Use the linear programming algorithm to optimize the allocated virtual storage space and virtual storage time.

4. A government affairs integration data collection and storage management system according to claim 1, characterized in that, The data evaluation and resource allocation module: evaluates the integrity of government affairs data using the government affairs data collection integrity formula based on the entropy weight method, and obtains the evaluation result of the government affairs data collection integrity; And adjusts the collection order of government affairs data according to the evaluation result of the government affairs data collection integrity. Based on the adjusted collection order of government affairs data, determines the collection and update order of historical literature data, including: The government affairs data collection integrity formula is as follows. , Among them, is the acquisition integrity of government affairs data, range: [0, 1]; is the importance weight factor of historical literature data; is the adjustment parameter; is the standardized processing time function; is the serial number of the government affairs data being currently processed; is the total number of times the government affairs data needs to be processed; is the time when the th government affairs data is processed; is the standard deviation of the completion times of all government affairs data; is the maximum value of the remaining number of processing times.

5. A government affairs integration data collection and storage management system according to claim 3, characterized in that, According to the evaluation result of the government affairs data collection integrity, reasonably allocates storage media and storage space for N pieces of government affairs data, and optimizes the storage space using the linear programming algorithm, including: Based on the evaluation result of the government affairs data collection integrity, sorts the government affairs data from low to high in terms of collection integrity and allocates storage media and storage space; Obtains the virtual storage space and virtual storage time of historical literature data; Extracts the government affairs data that meets the virtual storage space and virtual storage time; Obtains the estimated storage time and occupied space of the government affairs data that meets the virtual storage space and virtual storage time, as the alternative storage space for historical literature data.

6. An integrated government affairs data collection, storage and management system according to claim 5, characterized in that The storage management and optimization module: marks the storage media and remaining space of all government affairs data, including: If the alternative storage space for historical literature data is less than or equal to the preset alternative storage space threshold, integrates the remaining storage space of other government affairs data; If the alternative storage space for historical literature data is greater than the preset alternative storage space threshold, directly allocates the storage media of the alternative storage space.

7. An integrated government affairs data collection, storage and management system according to claim 5, characterized in that, Specially marks the storage situation of historical literature data; According to the storage situation of other government affairs data, optimizes the storage media and storage space of historical literature data using the genetic algorithm, including: Based on the storage media of the alternative storage space, preferentially allocates the storage media with the longest storage time limit for historical literature data, and synchronously allocates the storage media with the shortest storage time limit as the temporary storage media.

8. A government affairs integration data collection and storage management system according to claim 1, characterized in that, The dynamic management and maintenance module: implements dynamic management of government affairs data using the government affairs data long-term validity formula according to the collection order, storage media, and storage space of government affairs data, including: The government affairs data long-term validity formula is as follows. , Among them, is the long-term validity of government affairs data, range: [0, 1]; is the collection integrity of government affairs data calculated through the government affairs data collection integrity formula, range: [0, 1]; is a regulation parameter; is the time difference in the process of government affairs data collection.

9. A government affairs integration data collection and storage management system according to claim 8, characterized in that, Implementing dynamic management includes: Obtains the collection order, storage media, and storage space of government affairs data; Uses the historical literature data collection integrity as the center point of the collection sequence, and clusters using the K-means clustering algorithm to obtain the collection sequence clustering result; Uses the long-term validity of historical literature data as the center point of the storage sequence, and clusters using the K-means clustering algorithm to obtain the storage sequence clustering result; Calculates the similarity using the Euclidean distance formula; Performs superposition adjustment according to the collection sequence clustering result and the storage sequence clustering result to obtain the superposition adjustment result; Determines the dynamic management method based on the superposition adjustment result.

10. A government affairs integrated data collection and storage management system according to claim 9, characterized in that, Determining the dynamic management method based on the superposition adjustment result includes: Based on the collection sequence clustering result and the storage sequence clustering result, superimposes adjacent clustering points. If the superposition result is less than the set threshold, triggers the dynamic adjustment mechanism. If the superposition result is greater than or equal to the set threshold, maintains the current configuration.

Citation Information

Patent Citations

  • Data acquisition system based on bond transaction and data acquisition method thereof

    CN107123047A

  • High-reliability time sequence data transmission and storage system based on cloud edge collaboration

    CN117591496A

  • Nuclear power metadata storage management method and system

    CN117891789A

  • Intelligent power plant data storage management method and system

    CN119025833A

  • Government affair data operation and maintenance intelligent management method

    CN119620950A