An enterprise data optimization management system and method based on big data

By setting up multi-level file entry ports and customized saving strategies in the enterprise file processing system, the problems of file loss and redundancy are solved, and more efficient file management and memory utilization are achieved.

CN119003469BActive Publication Date: 2025-06-20IND & INFORMATION TECH (BEIJING) IND DEV RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411208362.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-06-20
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

During the process of enterprise file processing, files may be lost due to failure to save in time, and due to high repetitive files, files may be redundant, which is not conducive to memory management.

Method used

The file is saved through the first-level, second-level and third-level file entry ports, and the time the file handler opens and closes the file each time, and calculates the processing cycle. According to the threshold interval and acquisition interval, a customized saving strategy is implemented to save files during processing to avoid file redundancy.

Benefits of technology

It realizes timely storage of files, avoids file loss, reduces file redundancy, optimizes memory management, and improves the efficiency of file processing flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119003469B_ABST
    Figure CN119003469B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of enterprise data optimization, and discloses an enterprise data optimization management system and method based on big data, including: executing a processing cycle calculation strategy to obtain a processing cycle; when a first-level file handler processes a file: within the processing cycle; setting a threshold interval for determining whether the first-level file handler accidentally opens the file; according to the threshold interval, executing a preprocessing strategy to obtain an execution interval set; according to the execution interval set and the collection interval, executing a customized saving strategy to save the file during the processing; if the first-level file handler chooses to submit, executing a submission strategy to process the file received by the second-level file handler. By tracking the change in the file size, the storage space can be effectively managed to avoid unnecessary data accumulation. By reasonably setting the threshold, it can help determine when additional storage or backup is needed, ensure data security, make reasonable use of the storage space, and reduce the accumulation of invalid or outdated information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of enterprise data optimization, and specifically provides an enterprise data optimization management system and method based on big data. Background Art

[0002] Enterprise document processing and management is an important part of an enterprise information system, which is related to the effective transmission and storage of information. With the expansion of enterprise scale and the complexity of business processes, manually managing a large number of documents becomes cumbersome and error-prone. Therefore, modern enterprises increasingly rely on technical means to optimize the document management process. Document management has functions such as file version control, permission management, automatic archiving, search and retrieval. Version control allows multiple users to edit and review files at different time points while retaining historical version records for easy backtracking and recovery. By setting different access permissions, the security and controllability of sensitive information are ensured. Users can access specific files or perform specific operations according to their authorization levels, effectively preventing data leakage. In summary, by leveraging modern information technology, enterprises can achieve more efficient and secure document processing and management, thereby enhancing work efficiency and enterprise competitiveness.

[0003] In the existing process of processing enterprise documents, there is a risk of file loss if not saved in time. At the same time, considering the memory capacity, excessive saving of highly repetitive files causes file redundancy, which is not conducive to memory management.

[0004] This solution proposes to obtain a set of execution time points according to the collection interval and the set of execution intervals, judge the size of the file at each time point, and execute different strategies according to the judgment results. Summary of the Invention

[0005] The present invention provides an enterprise data optimization management system and method based on big data to help solve the problems mentioned in the above background art.

[0006] The present invention provides the following technical solution: An enterprise data optimization management method based on big data, including:

[0007] Through the primary file entry port: Save the files processed by the primary file handler.

[0008] Through the secondary file entry port: Save the files submitted by the primary file handler for review by the secondary file handler after completion.

[0009] Through the tertiary file entry port: Save the files submitted after the review by the secondary file handler is completed.

[0010] For each file at the primary file entry port;

[0011] Record the time when the primary file handler opens and closes the file each time to obtain a set of processing times;

[0012] Execute the processing cycle calculation strategy to obtain the processing cycle;

[0013] When the primary file handler processes the file:

[0014] Within the processing cycle;

[0015] Set a threshold interval for determining whether the primary file handler accidentally opens the file;

[0016] According to the threshold interval, execute the preprocessing strategy to obtain a set of execution intervals;

[0017] Set the acquisition interval for automatically saving the file;

[0018] According to the set of execution intervals and the acquisition interval, execute the customized saving strategy to save the file during the processing, obtain file copies, and form a set of file copies;

[0019] Each time the primary file handler closes the file, ask the primary file handler whether to submit the file processed this time to the secondary file handler;

[0020] If the primary file handler chooses to submit, execute the submission strategy to process the file received by the secondary file handler.

[0021] Optionally, the execution of the processing cycle calculation strategy includes:

[0022] All files in the tertiary file input port are files that have been processed;

[0023] Obtain the total number of files that the primary file handler has completed processing, denoted as the number of completions;

[0024] When the number of completions ≥ 1:

[0025] Obtain all the files that the primary file handler has completed processing;

[0026] For each file:

[0027] Record the time point when the file is first opened as the processing start point;

[0028] Record the time point when the file is last closed as the processing end point;

[0029] Calculate the interval between the processing start point and the processing end point for each file, and denote them as the first interval, the second interval... the rth interval respectively;

[0030] Calculate the processing cycle;

[0031] Calculate the processing cycle through the following formula:

[0032] (First interval + second interval +... + r-th interval) ÷ r = processing cycle.

[0033] Optionally, the execution of the processing cycle calculation strategy includes:

[0034] When the completion count = 0:

[0035] Obtain the processing cycles of all other first-level file handlers, denoted as the first cycle, the second cycle... the t-th cycle respectively;

[0036] Calculate the processing cycle;

[0037] Calculate the processing cycle through the following formula:

[0038] (First cycle + second cycle +... + t-th cycle) ÷ t = the processing cycle for the first processing by the first-level file handler.

[0039] Optionally, the execution of the preprocessing strategy according to the threshold interval to obtain the execution interval set includes:

[0040] In the processing time set:

[0041] Calculate the interval between each file opening and the corresponding file closing, denoted as the processing interval;

[0042] Obtain all the processing intervals and compare them with the threshold interval:

[0043] Obtain all the processing intervals greater than the threshold interval to form the execution interval set.

[0044] Optionally, the execution of the customized saving strategy according to the execution interval set and the collection interval to save the files during the processing to obtain the file copy set includes:

[0045] Obtain any element in the execution interval set, denoted as the marked interval;

[0046] Execute the execution interval division strategy to obtain the execution time point set;

[0047] Obtain the size of the file corresponding to each time point in the execution time point set, denoted as the file storage data, and form the file storage data set in the order of the appearance time from early to late;

[0048] Set the threshold data for determining whether to save the file, and the threshold data is greater than 0;

[0049] Traverse each element in the file storage data set in the order of the element appearance time:

[0050] S1. Obtain an element in the file storage data set, denoted as the first marked data;

[0051] S2. Obtain the next element adjacent to the first marker data, denoted as the second marker data;

[0052] S3. Calculate the second marker data - the first marker data, and denote the result as the judgment data;

[0053] S4. When the judgment data is greater than or equal to the threshold data, save the file at the time point corresponding to the second marker data to obtain a file copy;

[0054] S5. When the judgment data is greater than zero and less than the threshold data:

[0055] Obtain the next element adjacent to the second marker data in the file storage dataset, replace the current second marker data, and repeat S3 - S5;

[0056] S6. When the judgment data is less than zero, execute a backtracking optimization strategy on all file storage data before the second marker data.

[0057] Optionally, the executing the backtracking optimization strategy on all file storage data before the second marker data includes:

[0058] Obtain all file storage data before the second marker data to form a backtracking dataset;

[0059] Subtract the value of each element in the backtracking dataset from the value of the second marker data, and the results form a backtracking difference set;

[0060] Traverse all elements in the backtracking difference set, and obtain the largest element less than zero, denoted as the backtracking positioning difference;

[0061] Denote the time point when the file storage data corresponding to the backtracking positioning difference appears as the backtracking positioning start point;

[0062] Denote the time point corresponding to the second marker data as the backtracking positioning end point;

[0063] Then, delete all file copies saved from the backtracking positioning start point to the backtracking positioning end point from the file copy set;

[0064] Save the file copy corresponding to the second marker data.

[0065] Optionally, the executing the interval division strategy to obtain the execution time point set includes:

[0066] Obtain the start time point of the marker interval;

[0067] Obtain the end time point of the marker interval;

[0068] Denote the time point 1 acquisition interval away from the start moment as the first time point;

[0069] Record the time point that is two acquisition intervals away from the start time as the second time point;

[0070] ……

[0071] Record the time point that is u acquisition intervals away from the start time as the u-th time point, where the u-th time point is less than or equal to the end time point;

[0072] The start time point, the end time point, and all the time points between the start time point and the end time point together form the execution time point set.

[0073] Optionally, the execution submission policy includes:

[0074] The secondary file handler receives the file sent by the primary file handler, denoted as the received file:

[0075] For all the files in the secondary file input port:

[0076] If there is a file with the same file name as the received file, denote this file as the first draft file;

[0077] Replace the first draft file with the received file and store it in the secondary file input port;

[0078] If there is no file with the same file name as the received file, traverse all the files in the secondary file input port and compare them with the received file:

[0079] Set the threshold repeatability for determining whether a file is saved as the first version in the secondary file input port;

[0080] If there is a file with a repeatability greater than the threshold repeatability with the received file, denote this file as the second draft file;

[0081] Replace the second draft file with the received file and store it in the secondary file input port;

[0082] If there is no file with a repeatability greater than the threshold repeatability with the received file, store the received file in the secondary file input port.

[0083] A system for an enterprise data optimization management method based on big data, including:

[0084] Data acquisition module: primary file input port, secondary file input port, and tertiary file input port;

[0085] For each file in the primary file input port, record the time when the primary file handler opens and closes the file each time to obtain the processing time set;

[0086] Data processing module: execute the processing cycle calculation policy to obtain the processing cycle; according to the threshold interval, execute the preprocessing policy to obtain the execution interval set;

[0087] Execution module: According to the execution interval set and the collection interval, execute the customized saving strategy to obtain a set of file copies; when the first-level file handler submits a file to the second-level file handler, execute the submission strategy to process the file received by the second-level file handler.

[0088] The present invention has the following beneficial effects:

[0089] 1. A method for optimizing enterprise data management based on big data. When the completion times of the first-level file handler are greater than or equal to 1, obtain the processing start point and processing end point of each file, calculate the interval of each file processing, calculate the mean value, and obtain the processing cycle of the first-level file handler. By recording the start and end times of each file, the actual time required to process each file can be accurately calculated. Calculating the processing cycle helps to evaluate the efficiency of the file processing process, so as to find out the improvement points.

[0090] 2. A method for optimizing enterprise data management based on big data. When the completion times of the first-level file handler are 0, it means that the first-level file handler is a new person. Obtain the processing cycles of all other first-level file handlers, calculate the mean value, and obtain the processing cycle of the new person, providing an expected processing cycle for the new first-level file handler to help set reasonable goals and expectations.

[0091] 3. A method for optimizing enterprise data management based on big data. Calculate the processing interval of the file and compare it with the threshold interval, which can avoid redundant files caused by executing the customized saving strategy at the corresponding processing interval due to accidentally opening the file, saving memory.

[0092] 4. A method for optimizing enterprise data management based on big data. Obtain the marked interval from the execution interval set, execute the execution interval division strategy to obtain the set of execution time points, obtain the file size at each time point, and compare them. When the second marked data is greater than the first marked data, it means that the content of the file has increased. If it is determined that the data is greater than the threshold data, save the file corresponding to the second marked data; if it is determined that the data is greater than 0 and less than the threshold data, obtain the new second marked data and repeat the determination. By tracking the change of the file size, the storage space can be effectively managed to avoid unnecessary data accumulation. Reasonably setting the threshold can help determine when additional storage or backup is needed to ensure data security.

[0093] 5. An enterprise data optimization management method based on big data. When it is determined that the data is less than the threshold data, it indicates that the first-level file handler has deleted the file. Calculate the value obtained by subtracting the value of each element in the backtracking dataset from the value of the second marker data, and obtain the largest element less than zero in the backtracking difference set to get the backtracking positioning difference. The file storage data corresponding to the backtracking positioning difference is the size of the file that is closest to and greater than the second marker data, representing the highest degree of duplication between the file corresponding to the backtracking positioning start point and the file corresponding to the second marker data. All file copies saved from the backtracking positioning start point to the backtracking positioning end point are redundant. By deleting the useless file copies, the use of storage space and management complexity are reduced. Redundant data is avoided from being retained in the file processing system, thereby improving the operating efficiency of the system. Through accurate time point positioning, it is ensured that only relevant data is deleted while necessary historical records are retained.

[0094] 6. An enterprise data optimization management method based on big data. Obtain the start time point and end time point of the marking interval, and obtain multiple time points by taking different numbers of acquisition intervals starting from the start time point. The start time point, end time point, and all time points between the start time point and the end time point together form the execution time point set, and then subsequent customized saving strategies are executed. The systematic generation of time points helps ensure the coherence and integrity of data acquisition. By setting time points, the efficiency of data processing and analysis can be optimized.

[0095] 7. An enterprise data optimization management method based on big data. When there is a file in the secondary file entry port with the same file name as the received file, replace the first draft file with the received file and store it in the secondary file entry port; otherwise, if there is a file with a duplication degree greater than the threshold duplication degree with the received file, it indicates that the received file is not the first version in the secondary file entry port. Replace the second draft file with the received file and store it in the secondary file entry port; if there is no file with a duplication degree greater than the threshold duplication degree with the received file, it indicates that the received file is the first version in the secondary file entry port. Store the received file in the secondary file entry port to update the file in a timely manner to ensure the accuracy and currency of information. By checking the file duplication degree, storage resources are effectively managed, and unnecessary data duplication is avoided. Storage space is reasonably utilized, and the accumulation of invalid or outdated information is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] Figure 1 It is a schematic diagram of the system of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0097] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0098] Embodiment 1: Through the primary file entry port: Save the files processed by the primary file handler.

[0099] Through the secondary file entry port: Save the files submitted by the primary file handler for review by the secondary file handler after completion of processing.

[0100] Through the tertiary file entry port: Save the files submitted after the review by the secondary file handler is completed.

[0101] For each file at the primary file entry port;

[0102] Record the time when the primary file handler opens and closes the file each time to obtain a processing time set.

[0103] Execute a processing cycle calculation strategy to obtain a processing cycle.

[0104] When the primary file handler processes a file:

[0105] Within the processing cycle;

[0106] Set a threshold interval for determining whether the primary file handler accidentally opens a file.

[0107] According to the threshold interval, execute a preprocessing strategy to obtain an execution interval set.

[0108] Set a collection interval for automatically saving files.

[0109] According to the execution interval set and the collection interval, execute a customized saving strategy to save the files during the processing to obtain file copies, which form a file copy set.

[0110] Each time the primary file handler closes a file, ask the primary file handler whether to submit the file processed this time to the secondary file handler.

[0111] If the primary file handler chooses to submit, execute a submission strategy to process the files received by the secondary file handler.

[0112] The execution of the processing cycle calculation strategy includes:

[0113] All the files in the tertiary file entry port are files that have been completed processing;

[0114] Obtain the total number of files that the first-level file handler has completed processing, denoted as the number of completions;

[0115] When the number of completions ≥ 1:

[0116] Obtain all the files that the first-level file handler has completed processing;

[0117] For each file:

[0118] Record the time point when the file is first opened as the starting point of processing;

[0119] Record the time point when the file is finally closed as the ending point of processing;

[0120] Calculate the interval between the starting point and the ending point of processing for each file, denoted as the first interval, the second interval... the rth interval respectively;

[0121] Calculate the processing cycle;

[0122] Calculate the processing cycle through the following formula:

[0123] (The first interval + the second interval +... + the rth interval) ÷ r = the processing cycle.

[0124] In this embodiment, r = 4;

[0125] The first interval = 4 days, the second interval = 3 days, the third interval = 5 days, the fourth interval = 4;

[0126] Calculate (the first interval + the second interval +... + the rth interval) ÷ r = (4 + 3 + 5 + 4) ÷ 4 = 4 days;

[0127] When the number of completions of the first-level file handler is greater than or equal to 1, obtain the starting point and the ending point of processing for each file, calculate the interval of processing for each file, calculate the average value, and obtain the processing cycle of the first-level file handler. By recording the start and end times of each file, the actual time required to process each file can be accurately calculated. Calculating the processing cycle helps to evaluate the efficiency of the file processing process, so as to find out the improvement points.

[0128] The execution of the processing cycle calculation strategy includes:

[0129] When the number of completions = 0:

[0130] Obtain the processing cycles of all other first-level file handlers, denoted as the first cycle, the second cycle... the tth cycle respectively;

[0131] Calculate the processing cycle;

[0132] Calculate the processing cycle through the following formula:

[0133] (The first cycle + the second cycle +... + the t-th cycle) ÷ t = the processing cycle for the first time by the first-level document handler.

[0134] In this embodiment, t = 5;

[0135] The first cycle = 4 days, the second cycle = 4 days, the third cycle = 6 days, the fourth cycle = 6 days, and the fifth cycle = 5 days;

[0136] Calculate (the first cycle + the second cycle +... + the t-th cycle) ÷ t = (4 + 4 + 6 + 6 + 5) ÷ 5 = 5 days;

[0137] When the completion times of the first-level document handler are 0, it indicates that the first-level document handler is a new person. Obtain the processing cycles of all other first-level document handlers, calculate the average value, and obtain the processing cycle of the new person, providing an expected processing cycle for the new first-level document handler to help set reasonable goals and expectations.

[0138] The execution of the preprocessing strategy according to the threshold interval to obtain the execution interval set includes:

[0139] In the processing time set:

[0140] Calculate the interval between each time the file is opened and the corresponding time it is closed, denoted as the processing interval;

[0141] Obtain all the processing intervals and compare them with the threshold interval:

[0142] Obtain all the processing intervals greater than the threshold interval to form the execution interval set.

[0143] Calculating the processing interval of the file and comparing it with the threshold interval can avoid the execution of the customized saving strategy at the corresponding processing interval due to accidentally opening the file, resulting in file redundancy and saving memory.

[0144] The execution of the customized saving strategy according to the execution interval set and the acquisition interval to save the files during the processing to obtain the file copy set includes:

[0145] Obtain any element in the execution interval set, denoted as the marked interval;

[0146] Execute the execution interval division strategy to obtain the execution time point set;

[0147] Obtain the size of the file corresponding to each time point in the execution time point set, denoted as the file storage data, and form the file storage data set in the order of the appearance time from early to late;

[0148] Set the threshold data for judging whether to save the file, and the threshold data is greater than 0;

[0149] Traverse each element in the file storage dataset in the chronological order of element appearance:

[0150] S1. Obtain an element in the file storage dataset, denoted as the first marked data;

[0151] S2. Obtain the next element adjacent to the first marked data, denoted as the second marked data;

[0152] S3. Calculate the second marked data - the first marked data, and denote the result as the judgment data;

[0153] S4. When the judgment data is greater than or equal to the threshold data, save the file at the time point corresponding to the second marked data to obtain a file copy;

[0154] S5. When the judgment data is greater than zero and less than the threshold data:

[0155] Obtain the next element adjacent to the second marked data in the file storage dataset, replace the current second marked data, and repeat S3 - S5;

[0156] S6. When the judgment data is less than zero, perform a backtracking optimization strategy on all file storage data before the second marked data.

[0157] In this embodiment, the files stored in the current file copy set are respectively the file copy at the first time point, the file copy at the second time point, and the file copy at the third time point;

[0158] The second marked data is the file storage data of the file copy at the fourth time point;

[0159] The first marked data is the file storage data of the file copy at the third time point;

[0160] Calculate the judgment data, and the judgment data is less than 0;

[0161] Obtain the marked interval from the execution interval set, perform the execution interval division strategy to obtain the execution time point set, obtain the file size of each time point, and compare. When the second marked data is greater than the first marked data, it indicates an increase in the file content. If the judgment data is greater than the threshold data, save the file corresponding to the second marked data; if the judgment data is greater than 0 and less than the threshold data, obtain the new second marked data and repeat the judgment. By tracking the change of the file size, the storage space can be effectively managed to avoid unnecessary data accumulation. Reasonably setting the threshold can help determine when additional storage or backup is needed to ensure data security.

[0162] Performing the backtracking optimization strategy on all file storage data before the second marked data includes:

[0163] Obtain all file storage data before the second marker data to form a backtracking dataset;

[0164] Subtract the value of each element in the backtracking dataset from the value of the second marker data, and the results form a backtracking difference set;

[0165] Traverse all elements in the backtracking difference set, and obtain the largest element less than zero, denoted as the backtracking positioning difference;

[0166] Record the time point when the file storage data corresponding to the backtracking positioning difference appears as the backtracking positioning start point;

[0167] Record the time point corresponding to the second marker data as the backtracking positioning end point;

[0168] Then, delete all file copies saved from the backtracking positioning start point to the backtracking positioning end point from the file copy set;

[0169] Save the file copy corresponding to the second marker data.

[0170] Obtain the file storage data of the file copy at the first time point, which is equal to 10, the file storage data of the file copy at the second time point, which is equal to 20, and the file storage data of the file copy at the third time point, which is equal to 30, to form a backtracking dataset;

[0171] The second marker data is equal to 18;

[0172] Subtract the value of each element in the backtracking dataset from the value of the second marker data, and respectively obtain the difference corresponding to the first time = 18 - 10 = 8; the difference corresponding to the second time = 18 - 20 = -2; the difference corresponding to the third time = 18 - 30 = -12;

[0173] The largest element less than zero is the difference corresponding to the second time, then the backtracking positioning start point is the second time, and the backtracking positioning end point is the fourth time;

[0174] The file copies saved from the backtracking positioning start point to the backtracking positioning end point are the file copies at the third time point;

[0175] Delete the file copies at the third time point from the file copy set;

[0176] Save the file copies at the fourth time point;

[0177] The elements in the file copy set are the file copies at the first time point, the file copies at the second time point, and the file copies at the fourth time point;

[0178] When it is determined that the data is less than the threshold data, it indicates that the first-level file handler has deleted some files. Calculate the difference between the value of the second marker data and the value of each element in the backtracking dataset, and obtain the largest element less than zero in the backtracking difference set, which is the backtracking positioning difference. The file storage data corresponding to the backtracking positioning difference is the size of the file that is closest to and greater than the second marker data, representing the highest degree of duplication between the file corresponding to the backtracking positioning start point and the file corresponding to the second marker data. All file copies saved from the backtracking positioning start point to the backtracking positioning end point are redundant. By deleting the useless file copies, the use of storage space and the management complexity can be reduced. Avoid retaining redundant data in the file processing system, thereby improving the operating efficiency of the system. Through accurate time point positioning, ensure that only relevant data is deleted and necessary historical records are retained.

[0179] The execution interval division strategy obtains an execution time point set, including:

[0180] Obtain the start time point of the marking interval;

[0181] Obtain the end time point of the marking interval;

[0182] Record the time point at a distance of 1 acquisition interval from the start time as the first time point;

[0183] Record the time point at a distance of 2 acquisition intervals from the start time as the second time point;

[0184] ……

[0185] Record the time point at a distance of u acquisition intervals from the start time as the u-th time point, where the u-th time point is less than or equal to the end time point;

[0186] The start time point, the end time point, and all time points between the start time point and the end time point together form the execution time point set.

[0187] Obtain the start time point and the end time point of the marking interval, obtain multiple time points by starting from the start time point and spacing different numbers of acquisition intervals, and jointly form the execution time point set with the start time point, the end time point, and all time points between the start time point and the end time point. Then, execute the subsequent customized saving strategy. The systematic generation of time points helps ensure the coherence and integrity of data acquisition. By setting time points, the efficiency of data processing and analysis can be optimized.

[0188] The execution submission strategy includes:

[0189] The second-level file handler receives the file sent by the first-level file handler, denoted as the received file:

[0190] For all files in the second-level file entry port:

[0191] If there is a file with the same file name as the received file, record this file as the first draft file;

[0192] Replace the first draft file with the received file and store it in the secondary file input port;

[0193] If there is no file with the same file name as the received file, traverse all the files in the secondary file input port and compare them with the received file:

[0194] Set the threshold repeatability for determining whether a file is saved as the first version in the secondary file input port;

[0195] If there is a file with a repeatability greater than the threshold repeatability with the received file, record this file as the second draft file;

[0196] Replace the second draft file with the received file and store it in the secondary file input port;

[0197] If there is no file with a repeatability greater than the threshold repeatability with the received file, store the received file in the secondary file input port.

[0198] When there is a file with the same file name as the received file in the secondary file input port, replace the first draft file with the received file and store it in the secondary file input port; otherwise, if there is a file with a repeatability greater than the threshold repeatability with the received file, it means that the received file is not the first version in the secondary file input port, replace the second draft file with the received file and store it in the secondary file input port; if there is no file with a repeatability greater than the threshold repeatability with the received file, it means that the received file is the first version in the secondary file input port, store the received file in the secondary file input port, and update the file in time to ensure the accuracy and currency of the information. By checking the file repeatability, effectively manage the storage resources, avoid unnecessary data duplication. Reasonably utilize the storage space and reduce the accumulation of invalid or obsolete information.

[0199] Example 2, refer to Figure 1 , a system for an enterprise data optimization management method based on big data, includes:

[0200] Data acquisition module: a primary file input port, a secondary file input port, and a tertiary file input port;

[0201] For each file in the primary file input port, record the time when the primary file handler opens and closes the file each time to obtain a set of processing times;

[0202] Data processing module: execute the processing cycle calculation strategy to obtain the processing cycle; according to the threshold interval, execute the preprocessing strategy to obtain the execution interval set;

[0203] Execution module: According to the execution interval set and the acquisition interval, execute the customized saving strategy to obtain a set of file copies; when the primary file handler submits a file to the secondary file handler, execute the submission strategy to process the file received by the secondary file handler.

[0204] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0205] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for optimizing enterprise data management based on big data, comprising: Through the primary file input port: save the files processed by the primary file processor; Through the secondary file input port: save the files that the primary file processor has processed and submitted to the secondary file processor for review; Through the third-level file entry port: save the files submitted by the second-level file processor after review; For each file in the primary file entry port; Record the time when the primary file processor opens and closes the file each time, and obtain the processing time set; Execute the processing cycle calculation strategy to obtain the processing cycle; When a primary document processor processes a document: During the processing cycle; Set the threshold interval for judging whether the primary file handler has opened the file by mistake; According to the threshold interval, the preprocessing strategy is executed to obtain the execution interval set; Set the collection interval for automatic file saving; According to the execution interval set and the collection interval, a customized saving strategy is executed to save the files in the process, obtain the file copies, and form a file copy set; Each time the primary file processor closes a file, the primary file processor is asked whether to submit the file processed this time to the secondary file processor; If the primary file processor chooses to submit, the submission strategy is executed to process the file received by the secondary file processor; The step of executing a customized saving strategy according to the execution interval set and the collection interval, saving the files in the processing process, and obtaining the file copy set includes: Get any element in the execution interval set and record it as the marked interval; Execute the interval partitioning strategy to obtain the execution time point set; Get the size of the file corresponding to each time point in the execution time point set, record it as file storage data, and form a file storage data set in the order of appearance time from early to late; Set the threshold data for judging whether to save the file, the threshold data is greater than 0; Traverse each element in the file storage data set in the order of the time when the element appears: S1, obtaining an element in the file storage data set, recorded as the first marked data; S2, obtaining the next element adjacent to the first labeled data, and recording it as the second labeled data; S3, calculating the second labeled data minus the first labeled data, and recording the result as judgment data; S4. When the judgment data is greater than or equal to the threshold data, the file is saved at the time point corresponding to the second mark data to obtain a file copy; S5. When the judgment data is greater than zero and less than the threshold data: Obtain the next element adjacent to the second labeled data in the file storage data set, replace the current second labeled data, and repeat S3-S5; S6. When the judgment data is less than zero, a backtracking optimization strategy is executed on all file storage data before the second marking data, specifically including: Acquire all file storage data before the second marked data to form a backtracking data set; Subtract the value of each element in the backtracking data set from the value of the second labeled data, and the results form a backtracking difference set; Traverse all elements in the backtracking difference set and obtain the largest element less than zero, which is recorded as the backtracking location difference; The time point at which the file storage data corresponding to the backtracking positioning difference appears is recorded as the backtracking positioning starting point; The time point corresponding to the second marked data is recorded as the end point of the backtracking positioning; Then, all the file copies saved from the backtracking positioning start point to the backtracking positioning end point are deleted from the file copy set; A copy of the file corresponding to the second markup data is saved.

2. The enterprise data optimization management method based on big data according to claim 1 is characterized by: The execution processing cycle calculation strategy includes: The files in the third-level file input port are all files that have been processed; Get the total number of files that have been processed by the first-level file processor, which is recorded as the number of completions; When the number of completions ≧ 1: Obtain all the documents that have been processed by the first-level document processor; For each file: The time point when the file is first opened is recorded as the processing starting point; The last time the file was closed is recorded as the end point of processing; Calculate the interval between the processing start point and the processing end point of each file, and record them as the first interval, the second interval, ... the rth interval respectively; Calculate processing cycles; The processing cycle is calculated by the following formula: (first interval + second interval + ... + rth interval) ÷ ​​r = processing period.

3. The enterprise data optimization management method based on big data according to claim 2 is characterized by: The execution processing cycle calculation strategy includes: When the number of completions = 0: Obtain the processing cycles of all other first-level file processors, which are recorded as the first cycle, the second cycle, and so on the tth cycle; Calculate processing cycles; The processing cycle is calculated by the following formula: (First cycle + second cycle + … + tth cycle) ÷ t = the processing cycle of the first processing by the first-level file processor.

4. The enterprise data optimization management method based on big data according to claim 1 is characterized by: The step of executing the preprocessing strategy according to the threshold interval to obtain the execution interval set includes: In the processing time set: Calculate the interval between each file opening and the corresponding file closing, and record it as the processing interval; Get all processing intervals and compare them with the threshold interval: Get all processing intervals greater than the threshold interval to form an execution interval set.

5. The enterprise data optimization management method based on big data according to claim 1 is characterized by: The execution interval division strategy is used to obtain a set of execution time points, including: Get the start time of the marking interval; Get the end time point of the marking interval; The time point that is one sampling interval away from the start time is recorded as the first time point; The time point that is 2 sampling intervals away from the start time is recorded as the second time point; The time point u collection intervals away from the start time is recorded as the u-th time point, where the u-th time point is less than or equal to the end time point; The start time point, the end time point, and all time points between the start time point and the end time point together constitute an execution time point set.

6. The enterprise data optimization management method based on big data according to claim 1 is characterized by: The execution submission strategy includes: The secondary file processor receives the file sent by the primary file processor, which is recorded as received file: For all files in the secondary file entry port: If there is a file with the same file name as the received file, then the file is recorded as the first draft file; Replace the first draft file with the received file and store it in the secondary file input port; If there is no file with the same file name as the received file, traverse all files in the secondary file input port and compare them with the received file: Set the threshold duplication for judging whether a file is saved as the first version in the secondary file input port; If there is a file whose duplication degree with the received file is greater than the threshold duplication degree, the file is recorded as the second draft file; Replace the second draft file with the received file and store it in the secondary file input port; If there is no file whose duplication degree with the received file is greater than the threshold duplication degree, the received file will be stored in the secondary file input port.

7. A system for implementing the enterprise data optimization management method based on big data as claimed in claim 1, comprising: Data collection module: primary file entry port, secondary file entry port and tertiary file entry port; For each file of the primary file input port, record the time when the primary file processor opens and closes the file each time to obtain a processing time set; Data processing module: executes the processing cycle calculation strategy to obtain the processing cycle; executes the preprocessing strategy according to the threshold interval to obtain the execution interval set; Execution module: Executes the customized preservation strategy according to the execution interval set and the collection interval to obtain the file copy set; when the primary file processor submits the file to the secondary file processor, the submission strategy is executed to process the file received by the secondary file processor.

Citation Information

Patent Citations

  • Smart versioning for files

    CN111936987A

  • Electronic approval method and device, equipment and medium

    CN114066425A