An integrated government affairs data collection and storage management system

Through the integrated government data collection and storage management system, the problems of non-standard data processing, irrational resource allocation and insufficient attention to historical document data in traditional government data management have been solved, efficient and accurate data collection and recording have been achieved, and the system's response speed and adaptability have been improved.

CN120407529BActive Publication Date: 2025-09-12SHANDONG SHUIFA ZIGUANG BIG DATA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510911895.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-12
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Traditional government data management systems have problems such as lack of standardization in data processing processes, irrational allocation of storage resources, lack of focus on historical document data, and imperfect dynamic management mechanisms, resulting in poor data integrity and traceability, and low system response speed and flexibility.

Method used

An integrated government data collection and storage management system is adopted, which uses precise time and sequence marking, entropy weight method and linear programming algorithm to optimize data collection sequence and storage resource allocation, and supervised learning model to predict data access patterns and implement dynamic management.

Benefits of technology

Ensure the efficient and accurate collection and recording of government data, especially historical document data, enhance data integrity and traceability, reduce duplication of work, save costs, ensure the security and convenient access to important data, and improve system response speed and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407529B_ABST
    Figure CN120407529B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of information technology, and in particular to an integrated government data collection and storage management system. It includes the following parts: a data collection and preliminary processing module: obtaining the processing time and processing sequence of N government data, and marking the processing time of the N government data according to the processing sequence of the N government data; for historical document data, recording the processing time and processing sequence of the historical document data in detail. The present invention improves the accuracy and traceability of government data through precise time and sequence marking; standardized data integrity assessment optimizes the collection process and storage resource allocation, reduces duplication of work and saves costs; special attention is paid to the long-term preservation and convenient access of historical document data to ensure its security and effective use; by predicting future data access patterns, flexible and forward-looking management is achieved, and the response speed and adaptability of the system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a government affairs integrated data collection and storage management system. Background Art

[0002] In recent years, with the rapid development of information technology and the advancement of government digital transformation, the scale and complexity of government data have continued to grow. Traditional data collection, storage, and management methods are no longer able to meet the needs of efficient modern government operations, especially when processing historical documentary data. This data not only carries important historical information but also holds irreplaceable value for policymaking, public services, and social research. However, due to the unique characteristics of historical documentary data, such as large data volumes, diverse formats, and strict preservation requirements, its collection, storage, and management present numerous challenges. In practice, government agencies must process massive amounts of government data, including daily administrative records, regulatory documents, and statistical reports. The collection and storage of this data must be efficient, accurate, and traceable to support government decision-making and service delivery. Furthermore, as a component of government data, historical documentary data, due to its importance and unique nature, requires a specifically designed approach to ensure its long-term preservation and effective utilization.

[0003] Currently, many government data management systems face the following issues: Lack of standardization in data processing processes leads to disorganized data collection, compromising data integrity and accuracy. Irrational allocation of storage resources results in low storage media utilization, increasing operating costs. Lack of focus on historical document data leads to a failure to effectively meet demand for historical document data. Inadequate dynamic management and maintenance mechanisms make it difficult for systems to adapt to changing data access patterns, reducing their responsiveness and flexibility.

[0004] In order to solve the above problems, an integrated data acquisition and storage management system is urgently needed. Summary of the Invention

[0005] In order to overcome the shortcomings of low data processing accuracy and poor practicality in existing government affairs management systems, the present invention provides an integrated government affairs data collection and storage management system.

[0006] The technical implementation scheme of the present invention is: an integrated government affairs data collection and storage management system, including the following parts:

[0007] Data collection and preliminary processing module: obtains the processing time and processing order of N government data, and marks the processing time of the N government data according to the processing order of the N government data; for historical document data, records the processing time and processing order of the historical document data in detail; determines the collection order of the historical document data based on the processing order of the N government data; the collection order of the historical document data is to separately mark the collection order of the historical documents in the government data; the historical document data is part of the government data;

[0008] Data evaluation and resource allocation module: Use the government data collection integrity formula based on the entropy weight method to evaluate the integrity of government data and obtain the government data collection integrity evaluation results; adjust the collection order of government data based on the government data collection integrity evaluation results, and determine the collection and update order of historical document data based on the adjusted government data collection order; and reasonably allocate storage media and storage space for N government data based on the government data collection integrity evaluation results, and optimize the storage space using a linear programming algorithm;

[0009] Storage management and optimization module: Marks the storage media and remaining space of all government data; specifically marks the storage status of historical document data; uses genetic algorithms to optimize the storage media and storage space of historical document data based on the storage status of other government data;

[0010] Dynamic management and maintenance module: Based on the collection order, storage media and storage space of government data, the long-term validity formula of government data is used to implement dynamic management of government data, and the supervised learning model is used to predict future government data access patterns and adjust the distribution of government data.

[0011] Preferably, the data collection and preliminary processing module: obtains the processing time and processing order of N government data, and marks the processing time of the N government data according to the processing order of the N government data; for historical document data, records the processing time and processing order of the historical document data in detail, including:

[0012] Obtain the processing time and processing order of N government data; take the processing time as the first factor and the processing order as the second factor, normalize the first factor and the second factor, and then multiply them to obtain the comprehensive scoring result of the N government data;

[0013] Based on the comprehensive scoring results, the order of collecting government data is determined according to the principle of minimizing the sum of two adjacent items;

[0014] Obtain the finalized order of government data collection, mark the historical document data separately, and record in detail the processing time and order of the historical document data in the government data collection order.

[0015] Preferably, determining the collection order of historical document data according to the processing order of N government data includes:

[0016] Mapping storage space for the historical document data based on the processing time and processing order of the historical document data; said mapping storage space is to allocate virtual storage space for the historical document data;

[0017] Based on the initial collection order of the historical document data, mapping the storage space validity for the historical document data in the initial collection order; the mapping storage space validity is to allocate virtual storage time to the historical document data;

[0018] A linear programming algorithm is used to optimize the allocated virtual storage space and virtual storage time.

[0019] Preferably, the data evaluation and resource allocation module: uses the government data collection integrity formula based on the entropy weight method to evaluate the integrity of the government data, and obtains the government data collection integrity evaluation result; and adjusts the collection order of the government data according to the government data collection integrity evaluation result, and determines the collection and update order of the historical document data based on the adjusted government data collection order, including: the government data collection integrity formula is as follows,

[0020] ,

[0021] in, The collection completeness of government data, range: [0,1]; is the importance weight factor of historical document data; To adjust the parameters; is the normalized processing time function; The sequence number of the government data currently being processed; The total number of times government data needs to be processed; For the The time it takes to complete the processing of government data; is the standard deviation of the processing completion time for all government data; The maximum number of remaining processing times.

[0022] Preferably, the method of reasonably allocating storage media and storage space for N government data based on the government data collection integrity assessment result and optimizing the storage space using a linear programming algorithm includes:

[0023] Based on the government data collection integrity assessment results, the government data are sorted from low to high in terms of collection integrity and storage media and storage space are allocated;

[0024] Obtain virtual storage space and virtual storage time of historical document data;

[0025] Extracting government data that complies with virtual storage space and virtual storage time;

[0026] The estimated storage time and occupied space of government data that conforms to the virtual storage space and virtual storage time are obtained as an alternative storage space for historical document data.

[0027] Preferably, the storage medium and remaining space for marking all government data include:

[0028] If the alternative storage space for historical document data is less than or equal to the pre-selected storage space threshold, the remaining storage space for other government data will be integrated;

[0029] If the candidate storage space for historical document data is greater than the pre-selected storage space threshold, the storage medium with the selected storage space is directly allocated.

[0030] Preferably, the storage status of the historical document data is specially marked; and according to the storage status of other government data, the storage medium and storage space of the historical document data are optimized, including:

[0031] Based on the storage media of the candidate storage space, the storage medium with the longest storage validity time is preferentially allocated to the historical document data, and the storage medium with the shortest storage validity time is simultaneously allocated as a temporary storage medium.

[0032] Preferably, the dynamic management and maintenance module: implements dynamic management of government data using the government data long-term validity formula according to the collection order, storage medium and storage space of government data, including: the government data long-term validity formula is as follows,

[0033] ,

[0034] in, The long-term validity of government data, range: [0,1]; is the collection integrity of government data calculated by the government data collection integrity formula, ranging from [0,1]; is a tuning parameter; It is the time difference in the process of collecting government data.

[0035] Preferably, the implementation of dynamic management includes:

[0036] Obtain the collection order, storage media and storage space of government data;

[0037] Taking the collection integrity of historical document data as the center point of the collection sequence, the K-means clustering algorithm is used to cluster and obtain the collection sequence clustering results;

[0038] Taking the long-term validity of historical document data as the center point of the storage sequence, the K-means clustering algorithm is used to cluster and obtain the storage sequence clustering results;

[0039] Similarity is calculated using the Euclidean distance formula;

[0040] Performing superposition adjustment according to the acquisition sequence clustering result and the storage sequence clustering result to obtain a superposition adjustment result;

[0041] A dynamic management method is determined based on the superimposed adjustment result.

[0042] Preferably, determining a dynamic management method based on the superimposed adjustment result includes:

[0043] Based on the clustering results of the acquisition sequence and the storage sequence, adjacent cluster points are superimposed. If the superimposed result is less than a set threshold, a dynamic adjustment mechanism is triggered. If the superimposed result is greater than or equal to the set threshold, the current configuration is maintained.

[0044] Beneficial effects: The present invention ensures that all government data, especially historical document data, can be collected and recorded efficiently and accurately through precise time and sequence marking, thereby enhancing the integrity and traceability of the data; adjusts the data collection sequence through standardized data integrity assessment methods, and reasonably allocates storage resources based on the assessment results, thereby reducing duplication of work, saving time and costs; pays special attention to the long-term preservation and rapid retrieval requirements of historical document data, and ensures the security and convenient access to these important data by mapping storage space and validity allocation; achieves foresight and flexibility in data distribution by predicting future data access patterns, and improves the response speed and adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a schematic diagram of the structure of the government affairs integrated data collection and storage management system of the present invention;

[0046] Figure 2 This is a flow chart of the implementation of dynamic management in the present invention. DETAILED DESCRIPTION

[0047] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0048] An integrated government data collection and storage management system, such as Figure 1 and Figure 2 As shown, it includes the following parts:

[0049] Data collection and preliminary processing module: obtains the processing time and processing order of N government data, and marks the processing time of the N government data according to the processing order of the N government data; for historical document data, records the processing time and processing order of the historical document data in detail; determines the collection order of the historical document data based on the processing order of the N government data; the collection order of the historical document data is to separately mark the collection order of the historical documents in the government data; the historical document data is part of the government data;

[0050] Data evaluation and resource allocation module: Use the government data collection integrity formula based on the entropy weight method to evaluate the integrity of government data and obtain the government data collection integrity evaluation results; adjust the collection order of government data based on the government data collection integrity evaluation results, and determine the collection and update order of historical document data based on the adjusted government data collection order; and reasonably allocate storage media and storage space for N government data based on the government data collection integrity evaluation results, and optimize the storage space using a linear programming algorithm;

[0051] Storage management and optimization module: Marks the storage media and remaining space of all government data; specifically marks the storage status of historical document data; uses genetic algorithms to optimize the storage media and storage space of historical document data based on the storage status of other government data;

[0052] Dynamic management and maintenance module: Based on the collection order, storage media and storage space of government data, the long-term validity formula of government data is used to implement dynamic management of government data, and the supervised learning model is used to predict future government data access patterns and adjust the distribution of government data.

[0053] Data collection and preliminary processing module: obtains the processing time and processing order of N government data, and marks the processing time of N government data according to the processing order of N government data; for historical document data, records the processing time and processing order of historical document data in detail, including:

[0054] Obtain the processing time and processing order of N government data; take the processing time as the first factor and the processing order as the second factor, normalize the first factor and the second factor, and then multiply them to obtain the comprehensive scoring result of the N government data;

[0055] Based on the comprehensive scoring results, the order of collecting government data is determined according to the principle of minimizing the sum of two adjacent items;

[0056] Obtain the finalized order of government data collection, mark the historical document data separately, and record in detail the processing time and order of the historical document data in the government data collection order.

[0057] Further explanation: The processing time (actual time consumption) and processing sequence (execution sequence) of government data are automatically captured in real time through log monitoring or time series databases. Historical document data is also recorded with its independent processing sequence.

[0058] Multi-factor comprehensive score: Processing time (efficiency) and processing order (priority) are normalized and multiplied to generate a comprehensive score (for example, a shorter time and a higher order will result in a higher score).

[0059] Dynamic sorting optimization: Rearrange the collection order according to the principle of minimizing the sum of adjacent government data scores to reduce overall processing delay (similar to the shortest job first algorithm).

[0060] Historical data tagging: historical document data is individually marked in the global sequence and its precise time sequence is recorded to ensure long-term traceability.

[0061] For example, assume 5 government data (including 2 historical document data H1 and H2): original data: processing time: 10s, 8s, 15s, 20s, 5s; processing order: 2, 1, 3, 5, 42, 1, 3, 5, 4. Normalization and scoring: time normalization: 0.5, 0.4, 0.75, 1, 0.25; order normalization: 0.4, 0.2, 0.6, 1, 0.8; comprehensive score: time × order → 0.2, 0.08, 0.45, 1, 0.2. Optimize the collection order: sort by the principle of proximity and minimum: data 2 (0.08) → data 1 (0.2) → data 5 (0.2) → data 3 (0.45) → data 4 (1). Historical data marking: if H1 = data 3 and H2 = data 4, mark their position and processing time (15s, 20s).

[0062] The order of collecting historical document data is determined based on the processing order of N government data, including:

[0063] Mapping storage space for the historical document data based on the processing time and processing order of the historical document data; said mapping storage space is to allocate virtual storage space for the historical document data;

[0064] Based on the initial collection order of the historical document data, mapping the storage space validity for the historical document data in the initial collection order; the mapping storage space validity is to allocate virtual storage time to the historical document data;

[0065] A linear programming algorithm is used to optimize the allocated virtual storage space and virtual storage time.

[0066] A further explanation is that storage space mapping: virtual storage blocks (logical addresses) are allocated according to the processing sequence of historical document data to achieve physical storage decoupling.

[0067] Validity mapping: Assign virtual storage time windows (such as validity period [2025-01-01 to 2030-12-31]) according to the initial collection order to define the life cycle.

[0068] Linear programming optimization: Minimizing storage cost is the objective function, virtual space and time are the decision variables, and hardware capacity is the constraint condition to solve the optimal allocation.

[0069] For example, consider five government data blocks (including H1 and H2): Virtual mapping: H1 is assigned virtual block B1 (50GB), valid until 2030; H2 is assigned B2 (80GB), valid until 2035. Conflict: B1 shares the physical disk with another government data block, C3 (40GB), but C3 is valid only until 2028.

[0070] Linear programming algorithm model construction: variables: Represents government data On storage media Capacity allocated on the server (GB); Goal: Minimize total cost ,in It is a storage medium Unit capacity cost (yuan / GB); Constraints: Media capacity limit: ( For medium Maximum capacity, GB); data capacity requirements: ( For data Total demand, GB); fixed allocation for historical data: ( exist Media fixed occupancy 50GB); other data limits: ( exist Media usage ≤ 40GB); non-negative constraints: ; Solve and execute: Solve the model to obtain the optimal solution. If the result requires and , then from Migrate to .

[0071] Data evaluation and resource allocation module: Use the government data collection integrity formula based on the entropy weight method to evaluate the integrity of government data and obtain the government data collection integrity evaluation result; and adjust the collection order of government data according to the government data collection integrity evaluation result, and determine the collection and update order of historical document data based on the adjusted government data collection order, including: The government data collection integrity formula is as follows,

[0072] ,

[0073] in, The collection completeness of government data, range: [0,1]; is the importance weight factor of historical document data; To adjust the parameters; is the normalized processing time function; The sequence number of the government data currently being processed; The total number of times government data needs to be processed; For the The time it takes to complete the processing of government data; is the standard deviation of the processing completion time for all government data; The maximum number of remaining processing times.

[0074] Further explanation is that the steps for adjusting the order of government data collection are: obtaining the results of the government data collection integrity assessment; sorting the government data from low to high according to the integrity of the government data; and adjusting the order of government data collection based on the sorting results.

[0075] Steps for determining the order of collecting and updating historical document data: obtaining the virtual storage space and virtual storage time of historical document data; extracting government data from government data that meets the virtual storage space and virtual storage time of historical document data; obtaining the estimated storage time and estimated storage space occupied by government data that meets the requirements; using the estimated storage time and estimated storage space occupied as alternative storage space for historical document data; and determining the order of collecting and updating historical document data based on the alternative storage space.

[0076] Objectively quantify the weight of historical document data through the entropy weight method ( ), combined with temporal features (processing order ,time consuming ) and global volatility ( ), dynamically evaluate data collection integrity ( Rising = rising integrity).

[0077] Formula breakdown: Numerator: : Historical data weight calculated by entropy weight method; : Processing efficiency factor ( Shorter → larger value); Denominator: : Penalty for remaining tasks (the more unprocessed data → the lower); :Global time consumption volatility (the greater the volatility → the lower);

[0078] Entropy weight method description, function: objective determination Weights are used to avoid subjective bias. Steps: Data matrix: Construct m government data × n evaluation indicators (such as historical value, update frequency), Normalization: Process the indicators positively / negatively, Entropy calculation:

[0079] ( =Indicator ratio), weight distribution: (Entropy Down → Weight rise), is the amount of government data, For the The data in The proportion of indicators.

[0080] Example, entropy weight calculation : The entropy value of historical document indicators is lower ( =0.2, =0.3 vs normal data e=0.8) → =1.5, completeness Calculation (taking H1 as an example): =3, =5, =15s, =11.6s, =5.5s, =0.618, assuming =0.1→1-exp(-0.1×0.618)=0.06, =0.09 / 11=0.008, order adjustment: H1 has the lowest W value → its priority is increased to the first place in the update sequence.

[0081] It represents the relationship between the current processing time and the processing order. The calculation formula is:

[0082] ,

[0083] in is the average of all processing completion times, is the standard deviation, which normalizes the processing time to ensure that the time differences between different processing tasks do not distort the results; From 1 to An integer indicating the processing order; Represents a positive integer, indicating the total amount of data that needs to be processed; Indicates the specific processing completion time. The unit is selected according to the actual situation (such as seconds or minutes); Measures the variability of processing time and reflects the degree of discreteness of processing time. is the completion time of each processing task, is the average completion time, To handle the maximum value of the remaining times and prevent the denominator from being zero, when = When the denominator is not zero, to ensure the rationality of the formula, if - >0, then use ; otherwise use 1.

[0084] Based on the government data collection integrity assessment results, storage media and storage space are reasonably allocated for N government data, including:

[0085] Based on the government data collection integrity assessment results, the government data are sorted from low to high in terms of collection integrity and storage media and storage space are allocated;

[0086] Obtain virtual storage space and virtual storage time of historical document data;

[0087] Extracting government data that complies with virtual storage space and virtual storage time;

[0088] The estimated storage time and occupied space of government data that conforms to the virtual storage space and virtual storage time are obtained as an alternative storage space for historical document data.

[0089] Further explanation is that integrity-driven sorting: by data integrity assessment results Sort values ​​from low to high ( Decrease = Increased data quality risk), prioritize allocating high-reliability storage resources (such as SSDs) to low-integrity data.

[0090] Historical document storage matching: Utilize the preset virtual storage parameters (space / time) of historical documents to select storage blocks with matching spatiotemporal attributes in other government data as a candidate pool to ensure the long-term storage compatibility of historical data.

[0091] Example, 5 government data integrity rankings: H1( =0.008)→Data 2( =0.2)→Data1( =0.3)→Data 5( =0.5)→H2( =0.7), historical document H1 requirement: 50GB of virtual space, valid until 2030. Resource allocation process: Allocation by integrity: H1 (lowest integrity) is allocated to a high-speed SSD (100GB), and H2 is allocated to a regular hard drive (200GB). Alternative storage matching: Filter other data blocks that meet H1 requirements (≥50GB and valid ≥2030) → discover the storage block of data 5 (80GB, valid until 2035). Mark the data 5 block as the alternative storage pool for H1, and automatically switch when the primary SSD fails.

[0092] Storage Management and Optimization Module: Marks the storage media and remaining space for all government data, including:

[0093] If the alternative storage space for historical document data is less than or equal to the pre-selected storage space threshold, the remaining storage space for other government data will be integrated;

[0094] If the candidate storage space for historical document data is greater than the pre-selected storage space threshold, the storage medium with the selected storage space is directly allocated.

[0095] Further explanation is how to set the pre-device storage space threshold:

[0096] Basic threshold setting: Determine the basic value based on the minimum storage capacity requirement of historical document data to ensure that its basic storage requirements are met; redundancy expansion: Add the system fault tolerance redundancy ratio (such as 10%-20%) to the basic value to form the final threshold to cope with storage fluctuations or sudden demands.

[0097] Alternative storage space verification: Check whether the alternative storage space for historical document data meets the requirements (capacity ≥ requirements). If insufficient, cross-block resource integration will be triggered.

[0098] Dynamic resource expansion: Insufficient alternatives: Scan the remaining space in non-alternative storage blocks (such as fragmented space released by other government data) and combine them for capacity expansion. Sufficient alternatives: Directly activate storage media in the alternative pool to avoid resource waste.

[0099] For example, H1 requires 50GB, but only 30GB is available in its alternative pool (data block 5), which is insufficient. The remaining space for other government data is: data 2 (20GB) + data 1 (15GB). Operation process: Extract fragmented space: Merge the remaining space of data 2 (20GB) and data 1 (15GB) → Generate a new 35GB block, combined with 30GB in the alternative pool → Total 65GB (>50GB requirement). Allocation execution: Primary storage: SSD 100GB (original allocation).

[0100] Auxiliary storage: 65GB mixed fragment blocks (emergency backup).

[0101] Specially mark the storage status of historical document data; based on the storage status of other government data, use genetic algorithms to optimize the storage media and storage space of historical document data, including:

[0102] Based on the storage media of the candidate storage space, the storage medium with the longest storage validity time is preferentially allocated to the historical document data, and the storage medium with the shortest storage validity time is simultaneously allocated as a temporary storage medium.

[0103] A further explanation is the dual-track storage strategy: Primary storage: select the medium with the longest storage period among other government data (such as CD / tape) to ensure the long-term preservation of historical documents.

[0104] Temporary storage media: Synchronously occupies the blocks with the shortest storage validity period (such as memory / temporary SSD) to achieve fast response to high-frequency access.

[0105] Genetic algorithm optimization: With the objective function of minimizing storage cost and access latency, the optimal combination of main memory and temporary storage media is solved through crossover and mutation iteration.

[0106] For example, the storage period of other government data is: Data 5 (CD-ROM, valid until 2035) → the longest, Data 1 (SSD, valid until 2028) → the shortest. H1 requires long-term storage and supports frequent retrieval.

[0107] Genetic algorithm operation: Initialize the population: Plan 1: main memory = Data 5 CD, temporary storage = Data 1 SSD, Plan 2: main memory = Data 2 tape, temporary storage = Data 5 SSD, iterative optimization: calculate fitness (cost + delay weighted): Plan 1 score > Plan 2, crossover mutation: replace Plan 1 temporary storage with Data 3 memory (expiration date 2026), output optimal solution: main storage: Data 5 CD (expiration date 2035), temporary storage medium: Data 3 memory (expiration date 2026, high-frequency access buffer).

[0108] Dynamic management and maintenance module: Based on the collection order, storage media and storage space of government data, the government data is dynamically managed using the government data long-term validity formula, including: The government data long-term validity formula is as follows:

[0109] ,

[0110] in, The long-term validity of government data, range: [0,1]; is the collection integrity of government data calculated by the government data collection integrity formula, ranging from [0,1]; is a tuning parameter; It is the time difference in the process of collecting government data.

[0111] Further explanation is, Used to control the nonlinear effect of time on effectiveness, usually a positive number. In practical applications, The value should be chosen based on industry standards or rules of thumb for the specific field, such as choosing a small The value reflects the long-term value of historical document data, or choose a large The value emphasizes the importance of fresh data in real-time decision support systems; The unit is selected according to the actual situation (such as seconds, minutes) to measure the timeliness of the data; The increase, will tend to 1, which means that the validity of the data will gradually decrease; if If the value is large, it means that the data loses its value rapidly over time; if Small means that the value of the data changes slowly over time.

[0112] Example, H1 completeness =0.008 (low), storage =2 years, historical document parameters: =0.01 (low attenuation), effectiveness calculation: =0.008*(1-0.9802)≈0.00016, dynamic response: Monitoring: <0.001 threshold alarm, maintenance action: re-collect H1 data to improve (→ =0.05), migrate to a more stable storage medium ( down to 0.005), after the update : ≈0.0005.

[0113] Implement dynamic management, including:

[0114] Obtain the collection order, storage media and storage space of government data;

[0115] Taking the collection integrity of historical document data as the center point of the collection sequence, the K-means clustering algorithm is used to cluster and obtain the collection sequence clustering results;

[0116] Taking the long-term validity of historical document data as the center point of the storage sequence, the K-means clustering algorithm is used to cluster and obtain the storage sequence clustering results;

[0117] Similarity is calculated using the Euclidean distance formula;

[0118] Performing superposition adjustment according to the acquisition sequence clustering result and the storage sequence clustering result to obtain a superposition adjustment result;

[0119] A dynamic management method is determined based on the superimposed adjustment result.

[0120] Further explanation is that dual sequence collaborative management: sequence construction: collection sequence: government data processing time sequence flow (including integrity ), storage sequence: storage medium distribution status (including long-term validity ), historical data dual-center clustering: collection sequence based on historical documents As the cluster center (to ensure the priority of key data collection), the storage sequence is based on historical documents For cluster centers (to ensure long-term storage stability), dynamic adjustment mechanism: Euclidean distance quantifies the collection / storage cluster similarity, and the superposition result trigger strategy: reorganize the data distribution when the difference is too large.

[0121] Example, acquisition sequence: storage sequence: (H1 is located on the CD), clustering and adjustment: acquisition clustering (based on H1 =0.008 as the center): Cluster 1 (high priority): H1, data 2 (distance from the center < 0.1), storage clustering (based on H1's =0.95 as the center): Cluster 1 (long-term storage): CD, SSD (distance < 0.05), overlay analysis: data 2 in cluster 1 originally stored in memory ( =0.6), the distance from the storage center =8.2 (Euclidean distance), trigger action: migrate data 2 to the CD ( =0.3).

[0122] Determining a dynamic management method based on the superimposed adjustment result includes:

[0123] Based on the clustering result of the acquisition sequence and the clustering result of the storage sequence, adjacent cluster points are superimposed. If the superposition result is less than a set threshold, a dynamic adjustment mechanism is triggered. If the superposition result is greater than or equal to the set threshold, the current configuration is maintained.

[0124] Use supervised learning models to predict future government data access patterns and adjust government data distribution.

[0125] Further explanation is that the dynamic management trigger mechanism is: Cluster overlay analysis: Calculate the sum of the Euclidean distances between the cluster centers of the collected sequence and the stored sequence (overlay results ), (Threshold) indicates a mismatch between data distribution and acquisition strategy → triggers adjustment, Maintain the status quo and predict intervention: Supervised learning models (such as LSTM) predict future patterns based on historical access logs, suppressing unnecessary adjustments.

[0126] The method for setting the threshold is to determine the basic threshold based on the minimum storage capacity requirement of historical document data; the system fault tolerance redundancy ratio (10%-20%) is added to the basic value to form the final threshold.

[0127] Example, overlay results : Collection cluster center ( =0.008) and the storage cluster center ( =0.95) distance = 8.5, set threshold =5→ > , Operation process: Threshold determination: =8.5> =5→No adjustment is triggered, predictive intervention: LSTM model analyzes historical access trends and predicts that H1 access will surge by 50% in the next three months. Active optimization: migrate H1 from optical disk to SSD ( Down to 2.1, still> But improve the response speed).

[0128] The above is a detailed introduction to the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for those skilled in the art, based on the idea of ​​the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An integrated government affairs data collection and storage management system, characterized by: Includes the following sections: Data collection and preliminary processing module: obtains the processing time and processing order of N government data, and marks the processing time of N government data according to the processing order of N government data; For historical document data, the processing time and processing order of the historical document data are recorded in detail; the collection order of the historical document data is determined according to the processing order of N government data; the collection order of the historical document data is to separately mark the collection order of the historical documents in the government data; The historical document data is part of the government data; Data evaluation and resource allocation module: Use the government data collection integrity formula based on the entropy weight method to evaluate the integrity of government data and obtain the government data collection integrity evaluation results; and adjusting the collection order of government data according to the government data collection integrity assessment result, and determining the collection and update order of historical document data based on the adjusted government data collection order; According to the government data collection integrity assessment results, reasonably allocate storage media and storage space for N government data, and optimize the storage space using a linear programming algorithm; Storage management and optimization module: Marks the storage media and remaining space of all government data; especially marks the storage status of historical document data; Based on the storage situation of other government data, use genetic algorithms to optimize the storage media and storage space of historical document data; Dynamic management and maintenance module: This module dynamically manages government data based on its collection sequence, storage media, and storage space, using the long-term validity formula for government data. It also uses supervised learning models to predict future government data access patterns and adjust the data distribution. The data collection and preliminary processing module: obtains the processing time and processing order of N government data, and marks the processing time of the N government data according to the processing order of the N government data; For historical document data, the processing time and processing sequence of the historical document data are recorded in detail, including: Obtain the processing time and processing order of N government data; take the processing time as the first factor and the processing order as the second factor, normalize the first factor and the second factor, and then multiply them to obtain the comprehensive scoring result of the N government data; Based on the comprehensive scoring results, the order of collecting government data is determined according to the principle of minimizing the sum of two adjacent items; Obtain the finalized order of government data collection, mark historical document data separately, and record in detail the processing time and order of historical document data in the government data collection sequence; Determining the order of collecting historical document data according to the processing order of N government data includes: Mapping storage space for the historical document data based on the processing time and processing order of the historical document data; said mapping storage space is to allocate virtual storage space for the historical document data; Based on the initial collection order of the historical document data, mapping the storage space validity for the historical document data in the initial collection order; the mapping storage space validity is to allocate virtual storage time to the historical document data; A linear programming algorithm is used to optimize the allocated virtual storage space and virtual storage time.

2. The government affairs integrated data collection and storage management system according to claim 1 is characterized in that: The data evaluation and resource allocation module: uses the government data collection integrity formula based on the entropy weight method to evaluate the integrity of government data and obtain the government data collection integrity evaluation result; The collection order of government data is adjusted according to the government data collection integrity assessment result, and the collection and update order of historical document data is determined based on the adjusted government data collection order, including: the government data collection integrity formula is as follows, ,in, The collection completeness of government data, range: [0,1]; is the importance weight factor of historical document data; To adjust the parameters; is the normalized processing time function; The sequence number of the government data currently being processed; The total number of times government data needs to be processed; For the The time it takes to complete the processing of government data; is the standard deviation of the processing completion time for all government data; The maximum number of remaining processing times.

3. The integrated government affairs data collection and storage management system according to claim 1 is characterized in that: The method of rationally allocating storage media and storage space for N government data based on the government data collection integrity assessment results and optimizing the storage space using a linear programming algorithm includes: Based on the government data collection integrity assessment results, the government data are sorted from low to high in terms of collection integrity and storage media and storage space are allocated; Obtain virtual storage space and virtual storage time of historical document data; Extracting government data that complies with virtual storage space and virtual storage time; The estimated storage time and occupied space of government data that conforms to the virtual storage space and virtual storage time are obtained as an alternative storage space for historical document data.

4. The government affairs integrated data collection and storage management system according to claim 3 is characterized in that: The storage management and optimization module marks the storage media and remaining space of all government data, including: If the alternative storage space for historical document data is less than or equal to the pre-selected storage space threshold, the remaining storage space for other government data will be integrated; If the candidate storage space for historical document data is greater than the pre-selected storage space threshold, the storage medium with the selected storage space is directly allocated.

5. The government affairs integrated data collection and storage management system according to claim 3 is characterized in that: The method of optimizing the storage medium and storage space of historical document data using a genetic algorithm based on the storage conditions of other government data includes: Based on the storage media of the candidate storage space, the storage medium with the longest storage validity time is preferentially allocated to the historical document data, and the storage medium with the shortest storage validity time is simultaneously allocated as a temporary storage medium.

6. The integrated government affairs data collection and storage management system according to claim 1 is characterized in that: The dynamic management and maintenance module: implements dynamic management of government data based on the collection order, storage medium and storage space of government data using the government data long-term validity formula, including: the government data long-term validity formula is as follows: ,in, The long-term validity of government data, range: [0,1]; is the collection integrity of government data calculated by the government data collection integrity formula, ranging from [0,1]; is a tuning parameter; It is the time difference in the process of collecting government data.

7. The integrated government affairs data collection and storage management system according to claim 6 is characterized in that: The implementation of dynamic management includes: Obtain the collection order, storage media and storage space of government data; Taking the collection integrity of historical document data as the center point of the collection sequence, the K-means clustering algorithm is used to cluster and obtain the collection sequence clustering results; Taking the long-term validity of historical document data as the center point of the storage sequence, the K-means clustering algorithm is used to cluster and obtain the storage sequence clustering results; Similarity is calculated using the Euclidean distance formula; Performing superposition adjustment according to the acquisition sequence clustering result and the storage sequence clustering result to obtain a superposition adjustment result; A dynamic management method is determined based on the superimposed adjustment result.

8. The government affairs integrated data collection and storage management system according to claim 7 is characterized in that: The determining of a dynamic management method based on the superimposed adjustment result includes: Based on the clustering results of the acquisition sequence and the storage sequence, adjacent cluster points are superimposed. If the superimposed result is less than a set threshold, a dynamic adjustment mechanism is triggered. If the superimposed result is greater than or equal to the set threshold, the current configuration is maintained.

Citation Information

Patent Citations

  • Government affair data operation and maintenance intelligent management method

    CN119620950A

  • Wind power plant data optimization storage management system and method

    CN120144058A