A distributed budget archive collaborative management system based on big data
By designing a distributed budget archive collaborative management system based on big data, using feature analysis, label setting and archive optimization modules, problems such as inconsistent formats and inconsistent templates in the processing of massive archive files are solved, and efficient and energy-saving archive file optimization is achieved.
Patent Information
- Application Number
- CN202510102093.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
When processing massive archive files, the existing technology has problems such as content errors, inconsistent formats, and inconsistent templates, resulting in high computing power consumption and low efficiency of analysis and processing.
Design a distributed budget archive collaborative management system based on big data, including archive acquisition module, feature analysis module, label setting module and archive optimization module. The feature analysis module extracts the data expression characteristics of the archive file, and the tag setting module sets the presentation label for the archive file, and performs targeted processing based on the tag, including information clustering and content verification.
When facing a large number of archive files, the system can effectively identify and optimize archive files, improve the consistency of archive files, save computing power, and improve processing efficiency.
Smart Images

Figure CN119537670B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed archive management, and in particular to a distributed budget archive collaborative management system based on big data. Background Art
[0002] With the development of computer technology and Internet technology, computer office has penetrated into all walks of life. For example, in the office process, computers are needed to process various archival files. These files are mostly in the form of tables. However, in actual situations, archival files compiled by multiple parties often have problems such as inconsistent formats and templates, which brings difficulties and challenges to archive sorting. Therefore, relevant technologies for processing archival files are valued by people.
[0003] For example, Chinese patent publication number: CN114817676A, discloses an archive management system, which includes an electronic archive centralized management system, which includes an electronic file pre-archiving system, an archive electronic file management system and an electronic archive long-term preservation system; the electronic file pre-archiving system includes a data receiving module, an electronic file conversion module, a metadata capture module, a metadata conversion module, a metadata supplement module, an electronic file supplement scanning module, an archive data comparison module, a receiving list generation module and an archive data submission module; the archive electronic file management system includes a user management module, an archive receiving module, a receiving detection module, an online archiving module, a data utilization module and a data sharing module; the electronic archive long-term preservation system includes an archive storage module, a management module, an archive sharing module and an archive retrieval module.
[0004] However, there are still the following problems in the prior art:
[0005] In actual situations, archive files may have problems such as content errors, inconsistent formats, inconsistent templates, etc. that need to be optimized. However, when dealing with massive archive files, the computing power consumption for analyzing and processing the archive files is high and the efficiency is low. Summary of the invention
[0006] To this end, the present invention provides a distributed budget archive collaborative management system based on big data, which is used to overcome the problems in the prior art that archive files may have content errors, inconsistent formats, inconsistent templates, etc., which need to be optimized and processed. However, when dealing with massive archive files, the computing power consumption for analyzing and processing the archive files is high and the efficiency is low.
[0007] To achieve the above objectives, the present invention provides a distributed budget archive collaborative management system based on big data, comprising:
[0008] The file acquisition module is used to acquire the archive files uploaded by the user;
[0009] A feature analysis module, which is connected to the archive acquisition module, and is used to call a number of the archive files, perform image analysis and text analysis on the archive files, and obtain data representation features of the archive files, including texture distribution features and text structure features;
[0010] a label setting module, connected to the feature analysis module, for comparing the data representation form characteristics of the archive file with the data representation form characteristics of the standard archive template, determining the data representation difference representation value, and setting a representation label for each archive file according to the data representation difference representation value;
[0011] The archive optimization module is connected to the archive acquisition module and the label setting module respectively, and is used to process the archive files according to the labels of each archive file, including:
[0012] Extracting information from archive files for clustering, building associations between various types of information and various types of standard information in standard archive templates one by one, and filling the extracted information into the standard archive template according to the associations to obtain an optimized archive file;
[0013] or, performing content verification on the information in the archive file, and marking the information in the archive file that fails the verification;
[0014] The content verification includes verifying the difference between the data structure of each type of information and the data structure of a standard file template, and the standard file template is pre-constructed.
[0015] Furthermore, the feature analysis module performs image analysis and text analysis on the archive file, including:
[0016] To obtain the image data converted from the archive file and extract the key texture in the image data;
[0017] It is used to extract the interval characteristics of key textures in different reference directions and the closed area formed by key textures to obtain texture distribution characteristics;
[0018] To obtain the text data of the archive file, extract the data structure and part of speech category of the text data, and obtain the text structure features;
[0019] The reference direction includes a horizontal reference direction and a vertical reference direction, and the key texture includes all textures presented in the non-text part of the archive file.
[0020] Furthermore, the label setting module compares the data representation characteristics of the archive file with the data representation characteristics of the standard archive template, including:
[0021] for comparing the texture distribution features and text structure features of the archive file and the standard archive template in each coordinate area under the same two-dimensional coordinate system;
[0022] It is used to record the horizontal interval mean difference ratio, vertical interval mean difference ratio and closed area mean difference ratio in each coordinate area obtained by comparison;
[0023] To determine consistency parameters based on the comparison results of text structure features;
[0024] The comparison result includes whether the data structures are the same and whether the part-of-speech categories are the same.
[0025] Furthermore, the label setting module determines the data performance difference characterization value including:
[0026] The weighted sum of the horizontal interval mean difference ratio, the vertical interval mean difference ratio, the closed area mean difference ratio and the consistency parameter of each coordinate area is used to obtain the regional data performance difference representation value;
[0027] The data performance difference characterization values corresponding to each coordinate region are averaged to obtain the data performance difference characterization value.
[0028] Furthermore, the label setting module sets a presentation label for each archive file, including:
[0029] If the data representation difference characterization value corresponding to the archive file is greater than or equal to a preset data representation difference characterization threshold, setting a difference representation label for the archive file;
[0030] If the data representation difference characterization value corresponding to the archive file is less than a preset data representation difference characterization threshold, a non-difference representation label is set for the archive file.
[0031] Furthermore, the archive optimization module processes the archive files according to the labels of the archive files, including:
[0032] If the archive file is a differential expression label, the archive optimization module extracts information from the archive file for clustering, builds associations between various types of information and various types of standard information in the standard archive template one by one, and fills the extracted information into the standard archive template according to the association to obtain an optimized archive file;
[0033] If the archive file has a non-differential presentation label, the archive optimization module verifies the content of the information in the archive file to determine whether to change the presentation label of the archive file.
[0034] Furthermore, the archive optimization module extracts information from the archive file for clustering, including:
[0035] for extracting certain information from the archive file and clustering each of the information according to the clustering condition;
[0036] The clustering conditions include that the data structure and part-of-speech category of the single-category information are the same and each piece of information is located in the same reference direction.
[0037] Furthermore, the archive optimization module constructs the association relationship between various types of information and various types of standard information in the standard archive template one by one, including:
[0038] It is used to compare the single-category information of archive files one by one with the various standard information of standard archive templates to solve the correlation ranking;
[0039] Establish an association relationship between the single-category standard information corresponding to the maximum association degree and the single-category information of the archive file.
[0040] Furthermore, the archive optimization module fills the extracted information into the standard archive template according to the association relationship to obtain an optimized archive file including:
[0041] A single type of standard information used to determine the association between various types of information in the archive file;
[0042] It is used to determine the position area of the single-category standard information in the standard archive template, and fill the single-category information of the archive file into the corresponding position area to obtain an optimized archive file.
[0043] Furthermore, the content verification performed by the archive optimization module includes:
[0044] Extracting information from archive files and standard archive templates in a predetermined extraction method for comparison, and calculating the correlation between the information;
[0045] If the correlation degree corresponding to the information is greater than or equal to a predetermined correlation degree threshold, it is determined that the information verification has passed;
[0046] If the correlation degree is less than a predetermined correlation degree threshold, it is determined that the information verification fails.
[0047] Compared with the prior art, the present invention is provided with an archive acquisition module, a feature analysis module, a label setting module, and an archive optimization module. The feature analysis module is used to identify the data expression form characteristics of the archive file with less computing power. The label setting module is used to trigger the data expression difference representation value from the data standard form characteristics, and a label is set for the archive file. Then, the archive optimization module is used to perform targeted classification processing on the archive file according to the label of the archive file. In particular, for the archive files with differential expression labels, considering that the differences of the archive files are strong, the archive files are optimized by completely extracting and reconstructing the information. For the non-differential expression labels, since the standard type of the archive files is strong, processing can be recommended. Therefore, when facing massive archive files, the present invention can perform targeted processing, save computing power, and improve the consistency of the archive files.
[0048] In particular, the feature parsing module of the present invention calls the archive file to obtain the data representation form characteristics of the archive file. In actual situations, there may be differences in the archive files aggregated by multiple parties, such as inconsistent formats, inconsistent templates or incorrect filling. Therefore, it is considered to analyze the archive files from two dimensions: image and text. Especially in the image dimension, for archive files with high differences, there are obvious differences in the representation form. Therefore, it is considered to analyze the differences between the archive files and the standard archive template from the texture distribution characteristics. For computers, it is possible to identify the differences in archive files faster and more intuitively. The texture distribution characteristics include the interval characteristics of key textures in different reference directions and the area of the closed area formed by the key textures. From the computer perspective, the parameters are quantifiable, easy to compare, and can reflect the formal differences, format differences or template differences of the archive files. For the text dimension, the data structure and part of speech categories are easy to extract quickly and do not occupy computing power. Therefore, the present invention relatively simply extracts the data representation form characteristics, calculates the data representation difference representation value, and provides data support for the subsequent setting of labels for archive files. In addition, when facing massive archive files, it can be processed in a targeted manner, save computing power, and improve the consistency of archive files.
[0049] In particular, the present invention processes archive files in a targeted manner based on the labels of the archive files. For archive files with different expression labels, it is characterized that there are large differences in the archive files observed in terms of expression forms. Therefore, the present invention adopts an information clustering method for such archive files to extract various types of information. Based on the method of building association relationships, it is possible to find the standard information corresponding to each type of information in the standard archive template, and then fill in the corresponding information and reconstruct the archive files to avoid the impact of large differences in archive files on the subsequent processes. Furthermore, when faced with massive archive files, it is possible to carry out targeted processing, save computing power, and improve the consistency of archive files.
[0050] In particular, for archive files with non-differential expression labels, the present invention adopts a content verification method. Since the format of such archive files is relatively uniform and the template is consistent, a predetermined extraction method can be used for rapid verification to mark information with potential problems. Therefore, when faced with massive archive files, targeted processing can be performed to save computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A schematic diagram of the structure of a distributed budget archive collaborative management system based on big data according to an embodiment of the invention;
[0052] Figure 2 A logic block diagram of setting a presentation tag for each archive file according to an embodiment of the invention;
[0053] Figure 3 A logic block diagram of processing archive files according to the labels of each archive file according to an embodiment of the invention;
[0054] Figure 4 A logical block diagram of content verification of information in an archive file according to an embodiment of the present invention. DETAILED DESCRIPTION
[0055] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0056] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0057] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the term "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0058] See also Figures 1 to 4 As shown, Figure 1 This is a schematic diagram of the structure of a distributed budget archive collaborative management system based on big data according to an embodiment of the invention. Figure 2 A logic block diagram of setting a presentation tag for each archive file according to an embodiment of the invention, Figure 3 This is a logic block diagram of processing archive files according to the labels of each archive file according to an embodiment of the invention. Figure 4The logical block diagram of the content verification of information in the archive file of the embodiment of the invention is as follows: the distributed budget archive collaborative management system based on big data of the embodiment of the invention includes:
[0059] The file acquisition module is used to acquire the archive files uploaded by the user;
[0060] A feature analysis module, which is connected to the archive acquisition module, and is used to call a number of the archive files, perform image analysis and text analysis on the archive files, and obtain data representation features of the archive files, including texture distribution features and text structure features;
[0061] a label setting module, connected to the feature analysis module, for comparing the data representation form characteristics of the archive file with the data representation form characteristics of the standard archive template, determining the data representation difference representation value, and setting a representation label for each archive file according to the data representation difference representation value;
[0062] The archive optimization module is connected to the archive acquisition module and the label setting module respectively, and is used to process the archive files according to the labels of each archive file, including:
[0063] Extracting information from archive files for clustering, building associations between various types of information and various types of standard information in standard archive templates one by one, and filling the extracted information into the standard archive template according to the associations to obtain an optimized archive file;
[0064] or, performing content verification on the information in the archive file, and marking the information in the archive file that fails the verification;
[0065] The content verification includes verifying the difference between the data structure of each type of information and the data structure of a standard file template, and the standard file template is pre-constructed.
[0066] Specifically, there is no limitation on the specific structures of the archive acquisition module, feature analysis module, label setting module and archive optimization module, which can be composed of logical components, including field programmable components, computers or microprocessors in computers.
[0067] Specifically, there is no limitation on the form of archival documents. It is understandable that archival documents are mostly presented in a table format, which records the required content.
[0068] Specifically, there is no limitation on the specific form of the standard file template. Technical personnel in this field can pre-construct a standard file template for reference based on actual conditions. The standard file template can be presented in a tabular form. To facilitate subsequent data comparison, preferably, the standard file template needs to be filled with correct standard information, which will not be repeated here.
[0069] Specifically, the feature analysis module performs image analysis and text analysis on the archive file, including:
[0070] To obtain the image data converted from the archive file and extract the key texture in the image data;
[0071] It is used to extract the interval characteristics of key textures in different reference directions and the area of closed regions formed by key textures to obtain texture distribution characteristics. It can be understood that in archive files, key textures are mostly distributed horizontally or vertically, and then the interval characteristics of key textures in different reference directions can be known. Taking the vertically distributed texture as an example, the interval characteristics of the horizontal reference direction can be obtained, which will not be repeated here.
[0072] To obtain the text data of the archive file, extract the data structure and part of speech category of the text data, and obtain the text structure features;
[0073] The reference direction includes a horizontal reference direction and a vertical reference direction, and the key texture includes all textures presented in the non-text part of the archive file.
[0074] Specifically, there is no limitation on the method of converting the archive file into image data, and any document processing tool can be used to convert the archive file into image data, such as image data in PDF format, which will not be elaborated here.
[0075] Specifically, there is no limitation on the method of extracting key textures from image data, and an image segmentation algorithm may be used to identify prominent texture features in the image data, which will not be described in detail.
[0076] Specifically, commonly, the key textures identified include the edges of the table frame and the line segments filled in the table, which will not be described in detail.
[0077] Specifically, data structures include letters, numbers, text, and symbols.
[0078] Specifically, there is no limitation on the method of obtaining the part-of-speech categories. The part-of-speech categories can be analyzed using existing natural language tools to perform word segmentation on the information to obtain the part-of-speech category of each word segmentation, which will not be elaborated here.
[0079] Specifically, the label setting module compares the data representation characteristics of the archive file with the data representation characteristics of the standard archive template, including:
[0080] for comparing the texture distribution features and text structure features of the archive file and the standard archive template in each coordinate area under the same two-dimensional coordinate system;
[0081] It is used to record the horizontal interval mean difference ratio, vertical interval mean difference ratio and closed area mean difference ratio in each coordinate area obtained by comparison;
[0082] To determine consistency parameters based on the comparison results of text structure features;
[0083] The comparison result includes whether the data structures are the same and whether the part-of-speech categories are the same.
[0084] Specifically, the difference ratio of two data is calculated as the ratio of the absolute value of the difference between the two data to the mean of the two data.
[0085] Specifically, in implementation, if the data structures are the same and the part-of-speech categories are the same, the consistency parameter value is 0; if the data structures are different and / or the part-of-speech categories are different, the consistency parameter value is 1.
[0086] Specifically, the label setting module determines the data performance difference characterization value including:
[0087] The weighted sum of the horizontal interval mean difference ratio, the vertical interval mean difference ratio, the closed area mean difference ratio and the consistency parameter of each coordinate area is used to obtain the regional data performance difference representation value;
[0088] The data performance difference characterization values corresponding to each coordinate region are averaged to obtain the data performance difference characterization value.
[0089] In implementation, the weights of the horizontal interval mean difference ratio, the vertical interval mean difference ratio, the closed area mean difference ratio and the consistency parameter in the weighted summation are 0.15, 0.15, 0.3 and 0.4 respectively.
[0090] The feature parsing module of the present invention calls the archive file to obtain the data representation form characteristics of the archive file. In actual situations, there may be differences in the archive files aggregated by multiple parties, such as inconsistent formats, inconsistent templates or incorrect filling. Therefore, it is considered to analyze the archive files from two dimensions, image and text. Especially in the image dimension, for archive files with high differences, there are obvious differences in the representation form. Therefore, it is considered to analyze the difference between the archive file and the standard archive template from the texture distribution feature. For the computer, based on this, the difference of the archive file can be identified faster and more intuitively. The texture distribution feature includes the interval characteristics of the key texture in different reference directions and the area of the closed area formed by the key texture. From the computer perspective, the parameters can be quantified, which is convenient for comparison, and can reflect the form difference, format difference or template difference of the archive file. For the text dimension, the data structure and part of speech category are easy to extract quickly and do not occupy computing power. Therefore, the present invention is relatively simple to extract the data representation form characteristics, calculate the data representation difference representation value, and provide data support for the subsequent setting of labels for the archive file. Then, when facing massive archive files, it can be targeted for processing, save computing power, and improve the consistency of the archive file.
[0091] Specifically, the label setting module sets a presentation label for each archive file, including:
[0092] If the data representation difference characterization value corresponding to the archive file is greater than or equal to a preset data representation difference characterization threshold, setting a difference representation label for the archive file;
[0093] If the data representation difference characterization value corresponding to the archive file is less than a preset data representation difference characterization threshold, a non-difference representation label is set for the archive file.
[0094] Specifically, the purpose of setting the data performance difference characterization threshold is to reflect the situation where there is a high difference between the archive file and the standard archive file. The data performance difference characterization threshold is predetermined, where:
[0095] A number of archive files that meet the standards are screened out in advance from the archive files, the data performance difference characterization value corresponding to each archive file is determined, and the mean is solved. The data performance difference characterization threshold is set as the product of the mean and the precision coefficient, and the precision coefficient is selected in the interval [0.45, 0.65].
[0096] Specifically, the archive optimization module processes the archive files according to the labels of the archive files, including:
[0097] If the archive file is a differential expression label, the archive optimization module extracts information from the archive file for clustering, builds associations between various types of information and various types of standard information in the standard archive template one by one, and fills the extracted information into the standard archive template according to the association to obtain an optimized archive file;
[0098] If the archive file has a non-differential presentation label, the archive optimization module verifies the content of the information in the archive file to determine whether to change the presentation label of the archive file.
[0099] Specifically, the archive optimization module extracts information from the archive file for clustering, including:
[0100] for extracting certain information from the archive file and clustering each of the information according to the clustering condition;
[0101] The clustering conditions include that the data structure and part-of-speech category of the single-category information are the same and each piece of information is located in the same reference direction.
[0102] It is understandable that the purpose of clustering is to identify single-category information in archive files. Single-category information is often described in the same way and has the same semantic orientation. In tabular archive files, it is often located in a single column or row. Based on this, clustering conditions are set.
[0103] Specifically, the archive optimization module constructs the association relationship between various types of information and various types of standard information in the standard archive template one by one, including:
[0104] It is used to compare the single-category information of archive files one by one with the various standard information of standard archive templates to solve the correlation ranking;
[0105] Establish an association relationship between the single-category standard information corresponding to the maximum association degree and the single-category information of the archive file.
[0106] Specifically, the relevance is a semantic relevance, which can be calculated using a variety of methods. For example, the cosine similarity algorithm is used to vectorize the keywords and then calculate the cosine similarity. Of course, when calculating the relevance of information, a single type of information may contain multiple keywords. It is necessary to solve the mean cosine similarity between the keywords in the two types of information and determine the mean cosine similarity as the relevance. This will not be repeated here.
[0107] Specifically, in implementation, it is mostly applicable to text-based archival files. For information in the form of letters and numbers, since letters and information do not have semantic association, the association can be directly set to 0, and this type of information can be marked to facilitate subsequent manual processing.
[0108] Specifically, the archive optimization module fills the extracted information into the standard archive template according to the association relationship to obtain an optimized archive file including:
[0109] A single type of standard information used to determine the association between various types of information in the archive file;
[0110] It is used to determine the position area of the single-category standard information in the standard archive template, and fill the single-category information of the archive file into the corresponding position area to obtain an optimized archive file.
[0111] The present invention processes archive files in a targeted manner based on the labels of the archive files. For archive files with labels for different expressions, it is characterized that there are large differences in the archive files when observing the expression forms. Therefore, the present invention adopts an information clustering method for such archive files to extract various types of information. Based on the method of building association relationships, it is possible to find the standard information corresponding to each type of information in the standard archive template, and then fill in the corresponding information and reconstruct the archive files to avoid the impact of large differences in archive files on the subsequent processes. Furthermore, when faced with massive archive files, it is possible to process them in a targeted manner, save computing power, and improve the consistency of the archive files.
[0112] Specifically, the content verification performed by the archive optimization module includes:
[0113] Extracting information from archive files and standard archive templates in a predetermined extraction method for comparison, and calculating the correlation between the information;
[0114] If the correlation degree corresponding to the information is greater than or equal to a predetermined correlation degree threshold, it is determined that the information verification has passed;
[0115] If the correlation degree is less than a predetermined correlation degree threshold, it is determined that the information verification fails.
[0116] Specifically, the purpose of setting the correlation threshold is to characterize the situation where the difference between the two types of information is small. The correlation threshold is pre-set, where:
[0117] A number of archive files that meet the standards are selected from the uploaded archive files as samples, and information of the same category is extracted for comparison. The mean correlation is calculated, and the correlation threshold is set to the product of the mean correlation and the offset coefficient. The offset coefficient is selected in the interval [0.85, 0.95].
[0118] For archive files with non-differential expression labels, the present invention adopts a content verification method. Since the formats of such archive files are relatively uniform and the templates are consistent, a predetermined extraction method can be used for rapid verification and to mark information with potential problems. Therefore, when faced with massive archive files, targeted processing can be performed to save computing power.
[0119] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A distributed budget archive collaborative management system based on big data, characterized in that: include: The file acquisition module is used to acquire the archive files uploaded by the user; A feature analysis module, which is connected to the archive acquisition module, and is used to call a number of the archive files, perform image analysis and text analysis on the archive files, and obtain data representation features of the archive files, including texture distribution features and text structure features; a label setting module, connected to the feature analysis module, for comparing the data representation form characteristics of the archive file with the data representation form characteristics of the standard archive template, determining the data representation difference representation value, and setting a representation label for each archive file according to the data representation difference representation value; The archive optimization module is connected to the archive acquisition module and the label setting module respectively, and is used to process the archive files according to the labels of each archive file, including: Extracting information from archive files for clustering, building associations between various types of information and various types of standard information in standard archive templates one by one, and filling the extracted information into the standard archive template according to the associations to obtain an optimized archive file; or, performing content verification on the information in the archive file, and marking the information in the archive file that fails the verification; Wherein, the content verification includes verifying the difference between the data structure of each type of information and the data structure of the standard file template, and the standard file template is pre-built; The file optimization module constructs the association relationship between various types of information and various types of standard information in the standard file template one by one, including: It is used to compare the single-category information of archive files one by one with the various standard information of standard archive templates to solve the correlation ranking; Establishing an association relationship between the single-category standard information corresponding to the maximum association degree and the single-category information of the archive file; The archive optimization module fills the extracted information into the standard archive template according to the association relationship to obtain an optimized archive file including: A single type of standard information used to determine the association between various types of information in the archive file; To determine the location area of the single-category standard information in the standard archive template, and fill the single-category information of the archive file into the corresponding location area to obtain an optimized archive file; The label setting module compares the data representation characteristics of the archive file with the data representation characteristics of the standard archive template, including: for comparing the texture distribution features and text structure features of the archive file and the standard archive template in each coordinate area under the same two-dimensional coordinate system; It is used to record the horizontal interval mean difference ratio, vertical interval mean difference ratio and closed area mean difference ratio in each coordinate area obtained by comparison; To determine consistency parameters based on the comparison results of text structure features; Wherein, the comparison results include whether the data structures are the same and whether the part of speech categories are the same; The label setting module determines the data performance difference characterization value including: The weighted sum of the horizontal interval mean difference ratio, the vertical interval mean difference ratio, the closed area mean difference ratio and the consistency parameter of each coordinate area is used to obtain the regional data performance difference representation value; To average the data performance difference characterization values corresponding to each coordinate region to obtain the data performance difference characterization value; The label setting module sets a presentation label for each archive file, including: If the data representation difference characterization value corresponding to the archive file is greater than or equal to a preset data representation difference characterization threshold, setting a difference representation label for the archive file; If the data representation difference characterization value corresponding to the archive file is less than a preset data representation difference characterization threshold, a non-difference representation label is set for the archive file.
2. The distributed budget archive collaborative management system based on big data according to claim 1 is characterized in that: The feature analysis module performs image analysis and text analysis on the archive file, including: To obtain the image data converted from the archive file and extract the key texture in the image data; It is used to extract the interval characteristics of key textures in different reference directions and the closed area formed by key textures to obtain texture distribution characteristics; To obtain the text data of the archive file, extract the data structure and part of speech category of the text data, and obtain the text structure features; The reference direction includes a horizontal reference direction and a vertical reference direction, and the key texture includes all textures presented in the non-text part of the archive file.
3. The distributed budget archive collaborative management system based on big data according to claim 1 is characterized in that: The archive optimization module processes the archive files according to the labels of each archive file, including: If the archive file is a differential expression label, the archive optimization module extracts information from the archive file for clustering, builds associations between various types of information and various types of standard information in the standard archive template one by one, and fills the extracted information into the standard archive template according to the association to obtain an optimized archive file; If the archive file has a non-differential presentation label, the archive optimization module verifies the content of the information in the archive file to determine whether to change the presentation label of the archive file.
4. The distributed budget archive collaborative management system based on big data according to claim 1 is characterized in that: The archive optimization module extracts information from archive files for clustering, including: for extracting certain information from the archive file and clustering each of the information according to the clustering condition; The clustering conditions include that the data structure and part-of-speech category of the single-category information are the same and each piece of information is located in the same reference direction.
5. The distributed budget archive collaborative management system based on big data according to claim 1 is characterized in that: The content verification performed by the archive optimization module includes: Extracting information from archive files and standard archive templates in a predetermined extraction method for comparison, and calculating the correlation between the information; If the correlation degree corresponding to the information is greater than or equal to a predetermined correlation degree threshold, it is determined that the information verification has passed; If the correlation degree is less than a predetermined correlation degree threshold, it is determined that the information verification fails.
Citation Information
Patent Citations
Archive management system
CN114817676A
Project standard test document automatic generation method and device
CN116011400A
Integrity detection system, method and equipment for electronic file and medium
CN117390696A
Document classification method and device, electronic equipment and medium
CN117612182A