Automatic file data cleaning method of file system

By comprehensively evaluating file complexity and matching to identify key files and performing data cleaning based on benchmark files, the problem of incomplete or redundant conversion of PDF-format contract files into Excel files in existing technologies is solved, achieving more efficient data cleaning results.

CN120610951AActive Publication Date: 2025-09-09JINAN KEJIN INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511081882.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-09-09
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

Existing intelligent cleaning algorithms cannot accurately identify key data and convert it into the target format when extracting data from contract documents in PDF format, resulting in incomplete or redundant information in the converted Excel files.

Method used

By obtaining the data type complexity and data volume of the contract documents to be processed, combined with the matching degree between the file name and the required information, key files are identified and data cleaning is performed based on the benchmark files, including preprocessing, word segmentation, semantic expansion and association analysis, to determine the type complexity of the file type and select the most matching benchmark file for cleaning.

Benefits of technology

It improves the accuracy of data cleaning, reduces the probability of incorrect cleaning of important data and the probability of missing redundant data, and makes the conversion of PDF format contract files to Excel files more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610951A_ABST
    Figure CN120610951A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic file integration, in particular to an automatic file data cleaning method for a file system, which comprises the following steps: acquiring a type identifier, demand information, a plurality of file types and a plurality of contract files to be processed included in each file type, obtaining corresponding file complexity according to the data type complexity and the data size of each to-be-processed contract file; obtaining a corresponding demand matching degree according to the file complexity, the file name and the matching degree between the file content and the demand information of each to-be-processed contract file; determining a key file corresponding to each file type according to the demand matching degree of each to-be-processed contract file; and identifying a reference file in the plurality of key files, and performing data cleaning on the plurality of to-be-processed contract files based on the reference file. According to the method and the device, a plurality of to-be-processed contract files can obtain a better data cleaning effect, so that the contract in the PDF format is accurately converted into the file in the EXCEL format.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automated file integration, and in particular to an automated file data cleaning method for a file system. Background Art

[0002] With the rapid development of digitalization, the importance of file systems as a core vehicle for data storage and management has become increasingly prominent. From the massive business files generated by daily business operations to the vast experimental data files accumulated by scientific research institutions, the data stored in file systems is experiencing explosive growth. With the advancement of artificial intelligence, various file integration software programs have emerged. In the process of managing multiple file resources, data cleansing (the process of extracting key data from the original file and converting it into the target format, while also filtering out redundant data) is often required.

[0003] At present, when extracting data from contract files in PDF format, existing intelligent algorithms cannot accurately identify the key data of contract files and convert them effectively. In particular, when batch converting multiple contract files, they cannot identify the data associations between multiple contract files. This makes the Excel files converted from contract files in PDF format incomplete or contain redundant information.

[0004] That is to say, the data cleaning effect of the intelligent cleaning algorithm provided by the existing technology is poor. Summary of the Invention

[0005] In order to solve the technical problem that the intelligent cleaning algorithms provided by the prior art have poor data cleaning effects, the present invention aims to provide an automated file data cleaning method for a file system. The technical solutions adopted are as follows: In a first aspect, an embodiment of the present invention provides a method for automatically cleaning file data in a file system, the method comprising: Acquire a type identifier, requirement information, multiple file types, and multiple contract files to be processed included in each file type, where the type identifier is used to indicate a target file type among the multiple file types; Obtain the file complexity of each contract document to be processed based on the data type complexity and data volume of each contract document to be processed; Obtaining the requirement matching degree of each contract document to be processed based on the file complexity of each contract document to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information; Determine the key file corresponding to each file type based on the demand matching degree of each pending contract file, wherein the key file is the pending contract file with the highest demand matching degree among the multiple pending contract files of the corresponding file type; Identifying a benchmark file among the multiple key files, wherein the benchmark file is: a key file corresponding to the target file type, and / or a key file corresponding to a file type with the highest type complexity among the multiple file types, where the type complexity is calculated based on the complexity of multiple files associated with the corresponding file type; Perform data cleansing on multiple pending contract documents based on benchmark files.

[0006] In one embodiment, the file complexity of each contract file to be processed is obtained based on the data type complexity and data volume of each contract file to be processed, including: Determine the data type complexity of each contract file to be processed as the ratio of the first number to the second number corresponding to each contract file to be processed, wherein the first number is the number of data types included in the corresponding contract file to be processed, and the second number is the total number of data types included in the file type to which the corresponding contract file to be processed belongs; The product of the data type complexity and the data volume of each contract document to be processed is determined as the original complexity of each contract document to be processed; The original complexity of each contract document to be processed is normalized to obtain the file complexity of each contract document to be processed.

[0007] In one embodiment, the step of obtaining a first degree of matching between the file name of each to-be-processed contract document and the requirement information includes: Perform word segmentation on the file name of each pending contract document to obtain a title phrase of each pending contract document; and perform word segmentation on the demand information to obtain a demand phrase; The intersection-and-union ratio between the title phrase and the requirement phrase of each to-be-processed contract document is determined as the first matching degree between the file name of each to-be-processed contract document and the requirement information.

[0008] In one embodiment, the step of obtaining the second degree of matching between the content of each to-be-processed contract document and the requirement information includes: Perform word segmentation on the content of each pending contract document to obtain content phrases of each pending contract document; and perform word segmentation on the demand information to obtain demand phrases; Performing semantic expansion on the content phrases of each pending contract document to obtain a hypernym of each pending contract document; and performing semantic expansion on the requirement phrases to obtain a hyponym, wherein a plurality of first participles in the content phrases of each pending contract document correspond one-to-one to a plurality of second participles in the corresponding hypernym, the second participles being hypernyms of the corresponding first participles, a plurality of third participles included in the requirement phrases correspond one-to-one to a plurality of fourth participles included in the hyponym, the fourth participles being hyponyms of the corresponding third participles; An association analysis is performed on the hypernym and hyponym of each contract document to be processed to obtain a second matching degree between the file content of each contract document to be processed and the requirement information.

[0009] In one embodiment, association analysis is performed on the hypernyms and hyponyms of each pending contract document to obtain a second matching degree between the content of each pending contract document and the requirement information, including: Taking the intersection of the hypernym and hyponym of each pending contract document to obtain the intersection phrase of each pending contract document; The ratio of the number of segment words included in the intersection phrase of each contract document to be processed to the number of segment words included in the hypernym phrase is determined as the second matching degree between the file content of each contract document to be processed and the demand information.

[0010] In one embodiment, the requirement matching degree of each contract document to be processed is obtained based on the file complexity of each contract document to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information, including: The product of the file complexity of each contract file to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information is determined as the requirement matching degree of each contract file to be processed.

[0011] In one embodiment, the step of obtaining the type complexity of each file type includes: Calculating the average of multiple file complexities of multiple contract files to be processed included in each file type to obtain the average complexity corresponding to each file type; According to the complexity mean corresponding to each file type, the type complexity of each file type is obtained.

[0012] In one embodiment, the type complexity of each file type is obtained based on the complexity mean corresponding to each file type, including: Determine the ratio of the number of pending contract documents included in each file type to the total number of pending contract documents included in multiple file types as the type coefficient of each file type; The product of the complexity mean corresponding to each file type and the type coefficient of each file is determined as the type complexity of each file type.

[0013] In one embodiment, a reference file is identified among a plurality of key files, including: determining the target file type as a primary type, and determining other file types except the primary type among the plurality of file types as secondary types; The difference between the type complexity of the main type and the type complexity of each sub-type is determined as the type difference value of each sub-type; Sum the type difference values ​​of each sub-type to obtain the difference sum value; When the difference sum value is greater than a preset difference threshold, the key file corresponding to the target file type is determined as the benchmark file, wherein the difference threshold is a negative number; When the difference sum value is less than or equal to the difference threshold, the key file corresponding to the subtype with the highest type complexity among the multiple subtypes is determined as the benchmark file.

[0014] In one embodiment, the step of obtaining multiple contract documents to be processed includes: Among multiple original files, content recognition is performed on each original file to obtain the original content of each original file, where the original file is a file input into the file system by the user; Preprocessing the original content of each original file to obtain preprocessed content of each original file, wherein the preprocessing includes at least one of missing value processing, outlier processing, and duplicate value processing; The file name of each original file and the corresponding pre-processed content are spliced ​​to obtain multiple contract files to be processed, wherein the file name of the contract file to be processed is the file name of the corresponding original file, and the file content of the contract file to be processed is the pre-processed content of the corresponding original file.

[0015] In a second aspect, another embodiment of the present invention provides an automated file data cleaning system for a file system, the system comprising: An acquisition module is used to acquire a type identifier, requirement information, multiple file types, and multiple contract files to be processed included in each file type, wherein the type identifier is used to indicate a target file type among the multiple file types; The first processing module is used to obtain the file complexity of each contract file to be processed according to the data type complexity and data volume of each contract file to be processed; A second processing module is configured to obtain a requirement matching degree for each contract document to be processed based on the file complexity of each contract document to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information; A key file determination module is used to determine the key file corresponding to each file type based on the requirement matching degree of each pending contract file, wherein the key file is the pending contract file with the highest requirement matching degree among the multiple pending contract files included in the corresponding file type; a reference file identification module for identifying a reference file among multiple key files, wherein the reference file is: a key file corresponding to the target file type, and / or a key file corresponding to a file type with the highest type complexity among multiple file types, where the type complexity is calculated based on the complexity of multiple files associated with the corresponding file type; The data cleaning module is used to perform data cleaning on multiple contract files to be processed based on the benchmark file.

[0016] In a third aspect, another embodiment of the present invention further provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the method of the first aspect when executed by the processor.

[0017] In a fourth aspect, another embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned first aspect method are implemented.

[0018] The present invention has the following beneficial effects: The present invention first comprehensively determines the file complexity of each contract file to be processed based on the data type complexity and data volume of the contract file to be processed. Then, based on the file complexity of each contract file to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information, the requirement matching degree of the contract file to be processed is comprehensively determined from multiple aspects such as the file itself and the degree of adaptation between the file and the requirement information, so as to determine the file with the highest requirement matching degree in each file type as the key file. Then, based on the file complexity of each contract file to be processed, the type complexity of each file type is determined, so as to identify the most important key file from a number of key files. That is, by comprehensively analyzing the correlation between multiple contract files to be processed and the correlation between multiple contract files to be processed and the requirement information, a benchmark file that matches the user's requirements and has a strong correlation with other contract files to be processed is selected. Finally, data cleaning of multiple contract files to be processed is completed based on the benchmark file, which can reduce the probability of incorrectly cleaning important data in the file and the probability of omitting cleaning of redundant data in the file, so that multiple contract files to be processed can achieve better data cleaning results, that is, the Excel file converted from the PDF format contract file is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 A schematic flow chart of an automated file data cleaning method for a file system provided by one embodiment of the present invention; Figure 2 A schematic structural diagram of an automatic file data cleaning system for a file system provided by one embodiment of the present invention; Figure 3 The present invention provides a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0021] To further illustrate the technical means and effectiveness of the present invention in achieving its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of an automated file data cleaning method for a file system proposed in accordance with the present invention. In the following description, references to different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0022] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0023] The following describes in detail a specific solution of an automatic file data cleaning method for a file system provided by the present invention in conjunction with the accompanying drawings.

[0024] The present invention proposes an automatic file data cleaning method for a file system. Figure 1 , which shows a schematic flow chart of an automatic file data cleaning method for a file system provided by an embodiment of the present invention, the method comprising the following steps: Step S1: Obtain type identification, requirement information, multiple file types, and multiple contract files to be processed included in each file type.

[0025] The type identifier is used to indicate the target file type among multiple file types.

[0026] In the application, the user inputs several files that need to be integrated (that is, multiple contract files to be processed included in each of the aforementioned file types) into the file system (also known as the file integration system), and selects the file type to be generated (indicated by the aforementioned type identifier) ​​and the required information of the integration process (through text description).

[0027] It should be noted that the contract files to be processed described in the present invention are all contract files in PDF format, and the files that have undergone data cleaning are all files in EXCEL format.

[0028] After receiving the plurality of files to be integrated input by the user, the file system may classify the received plurality of files to be integrated to determine the aforementioned plurality of file types.

[0029] In the application, different pending contract files can be distinguished based on the identity of the user who entered the pending contract files in the file system. For example, the pending contract files entered by user A correspond to file type A, the pending contract files entered by user B correspond to file type B, and so on. The number of multiple file types is determined by the number of users who can enter the pending contract files.

[0030] Alternatively, different pending contract documents can be distinguished based on the contract subject amount associated with the pending contract documents. For example, pending contract documents with a contract subject amount of less than or equal to RMB 500,000 correspond to the first file type, pending contract documents with a contract subject amount between RMB 500,000 and RMB 5 million correspond to the second file type, and pending contract documents with a contract subject amount greater than or equal to RMB 5 million correspond to the third file type.

[0031] It should be noted that the distinction between multiple file types can be adaptively set according to actual needs (such as the identity of the user entering the file, the contract amount, the industry to which the enterprise or organization involved in the contract belongs, etc.), and the present invention does not limit this.

[0032] The above type identifier may be a character string, such as "file type 1", "file type 2" and "file type 3".

[0033] The above-mentioned demand information is used to describe the user's requirements for the integration / data cleaning of multiple pending contract documents. For example, the demand information may be "extract the date data and key indicators (such as sales volume) under the corresponding dates from all pending contract documents, and summarize them in order by time."

[0034] In some embodiments, the step of obtaining multiple contract documents to be processed includes: Among multiple original files, content recognition is performed on each original file to obtain the original content of each original file, where the original file is a file input into the file system by the user; Preprocessing the original content of each original file to obtain preprocessed content of each original file, wherein the preprocessing includes at least one of missing value processing, outlier processing, and duplicate value processing; The file name of each original file and the corresponding pre-processed content are spliced ​​to obtain multiple contract files to be processed, wherein the file name of the contract file to be processed is the file name of the corresponding original file, and the file content of the contract file to be processed is the pre-processed content of the corresponding original file.

[0035] After the file system receives multiple original files input by the user, based on the above settings, the integrity and accuracy of the file content of the contract files to be processed used in subsequent processes can be guaranteed.

[0036] Exemplarily, the missing value processing may include: compensating the missing value based on a fixed preset value, or dynamically compensating the missing value by combining the non-missing data around the missing value with an interpolation algorithm; The above-mentioned abnormal value processing can be: setting the abnormal value to 0 or NULL; The above-mentioned duplicate value processing may be: merging the duplicate values.

[0037] Step S2: Obtain the file complexity of each contract file to be processed based on the data type complexity and data volume of each contract file to be processed.

[0038] The data type complexity is used to indicate the richness of the data types involved in the content of the contract document to be processed. In the present invention, the more data types a contract document to be processed includes, the greater the possibility that the data type complexity of the contract document to be processed is high.

[0039] Exemplarily, the data types involved in the file content of the contract document to be processed may include: numerical type (such as integers, floating point numbers), text type (such as strings, characters), date and time type (such as dates, timestamps), Boolean type (such as true / false values), and categorical type (such as enumeration values, labels).

[0040] Data volume refers to the amount of storage space occupied by the pending contract file, i.e., the file size of the pending contract file. For example, the file size of the pending contract file is 5MB (data volume).

[0041] In this step, the file complexity of the contract file to be processed is comprehensively evaluated from two aspects: the complexity of the data type involved in the contract file to be processed and the amount of data. This can comprehensively and accurately quantify the complexity of the contract file to be processed, that is, comprehensively and accurately quantify the probability that the contract file to be processed contains important data required by the user. The higher the file complexity of the contract file to be processed, the greater the probability that the contract file to be processed contains important data required by the user.

[0042] Specifically, based on the data type complexity and data volume of each contract file to be processed, the file complexity of each contract file to be processed is obtained, including: Determine the data type complexity of each contract file to be processed as the ratio of the first number to the second number corresponding to each contract file to be processed, wherein the first number is the number of data types included in the corresponding contract file to be processed, and the second number is the total number of data types included in the file type to which the corresponding contract file to be processed belongs; The product of the data type complexity and the data volume of each contract document to be processed is determined as the original complexity of each contract document to be processed; The original complexity of each contract document to be processed is normalized to obtain the file complexity of each contract document to be processed.

[0043] For example, the formula for expressing the complexity of the nth contract document to be processed under the mth file type (file complexity) may be:

[0044] In the above formula, Indicates the complexity of the nth contract document to be processed under the mth file type. Represents a normalization function (such as using Min-Max normalization or Z-Score normalization), Indicates the number of data types contained in the nth pending contract file under the mth file type (that is, the first number corresponding to the pending contract file), Indicates the total number of data types contained in all files under the mth file type (that is, the second number corresponding to the contract file to be processed), Indicates the data volume of the nth contract file to be processed under the mth file type, It can be understood as the aforementioned data type complexity.

[0045] Based on the above settings, the probability that the contract file to be processed contains important data required by the user is quantified by the ratio of the data types contained in the contract file to be processed and the total number of data types contained in the file type to which the contract file to be processed belongs. The total number of data types contained in the file type to which the contract file to be processed belongs is selected as a reference instead of the total number of data types contained in all contract files to be processed. This can accurately reflect the richness of the data types involved in the contract file to be processed, thereby making the calculated data type complexity more accurate and reliable.

[0046] The above-mentioned normalization measures are intended to reduce the impact of extreme values ​​(such as data volumes that are too large or too small) so that the file complexity can accurately represent the richness of the information content contained in the corresponding contract documents to be processed (that is, the complexity of the contract documents to be processed).

[0047] Step S3: Obtain the requirement matching degree of each contract file to be processed based on the file complexity of each contract file to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information.

[0048] The above-mentioned demand matching degree can be understood as: the degree of matching between the contract document to be processed and the user demand reflected by the demand information.

[0049] In this step, the demand matching degree of the contract document to be processed is calculated by comprehensively considering the file complexity of the contract document to be processed, the first matching degree between the file name of the contract document to be processed and the demand information, and the second matching degree between the file content of the contract document to be processed and the demand information. The evaluation is performed from multiple aspects, including the richness of the information content contained in the contract document to be processed and the matching degree between the contract document to be processed and the demand information, so as to more accurately achieve a quantitative representation of the matching degree between the contract document to be processed and the user needs reflected by the demand information.

[0050] Among them, the file name and file content of the contract document to be processed are regarded as different dimensional parts, and the matching degree between the demand information of the two is calculated separately. This can avoid the calculation deviation caused by single-dimensional processing, and comprehensively and accurately analyze the matching degree between the overall contract document to be processed and the demand information, thereby making the final calculated demand matching degree more accurate and reliable.

[0051] Among them, the first matching degree can be understood as the degree of correlation between the semantics of the file name of the contract document to be processed and the semantics of the requirement information. Similarly, the second matching degree can be understood as the degree of correlation between the semantics of the file content of the contract document to be processed and the semantics of the requirement information.

[0052] Exemplarily, the process of obtaining the first matching degree may be: Pre-process the file name and requirement information of the contract document to be processed respectively to remove the file extension included in the file name and punctuation marks and stop words included in the requirement information; Use a pre-trained semantic model (such as Sentence-BERT or Universal Sentence Encoder) to encode the pre-processed file name and requirement information into fixed-dimensional vectors. Finally, calculate the cosine similarity or dot product similarity between the two vectors and use it as the first match.

[0053] Exemplarily, the process of obtaining the second matching degree may be: Pre-process the content and requirement information of the contract documents to be processed respectively to clean up punctuation marks, stop words, etc. included in the content and requirement information of the contract documents to be processed; Perform keyword extraction (e.g., using the TF-IDF algorithm) on the pre-processed file content and demand information to obtain a content keyword set and a demand keyword set, respectively; Then, the intersection-and-union ratio of the content keyword set and the demand keyword set is calculated and determined as the second matching degree.

[0054] Step S4: Determine the key files corresponding to each file type based on the requirement matching degree of each contract file to be processed.

[0055] Among them, the key file is the pending contract file with the highest demand matching degree among the multiple pending contract files included in the corresponding file type.

[0056] Step S5: Identify a reference file among multiple key files.

[0057] Among them, the benchmark file is: the key file corresponding to the target file type, and / or the key file corresponding to the file type with the highest type complexity among multiple file types, and the type complexity is calculated based on the complexity of multiple files associated with the corresponding file type.

[0058] The aforementioned multiple key files correspond one-to-one to the aforementioned multiple file types.

[0059] In one example, the average of multiple file complexities of multiple to-be-processed contract documents included in each file type may be determined as the type complexity of each file type.

[0060] The above-mentioned type complexity is used to quantitatively represent: the overall complexity of the multiple contract documents to be processed included in the corresponding file type, that is, the comprehensive probability that the multiple contract documents to be processed included in the corresponding file type contain important data required by the user.

[0061] By comparing the difference between the type complexity of the target file type and the type complexity of other file types, a more reliable benchmark file can be determined.

[0062] For example, the number of file types other than the target file type among the multiple file types may be counted and defined as the first number; and the number of file types other than the target file type among the multiple other file types whose type complexity is higher than that of the target file type may be counted and defined as the second number; Calculate the ratio of the second number to the first number, and compare the ratio with a set ratio threshold (such as 0.3 or 0.35). If the ratio is greater than the set ratio threshold, select the key file corresponding to the file type with the highest type complexity among several other file types as the benchmark file; if the ratio is less than or equal to the set ratio threshold, select the key file corresponding to the target file type as the benchmark file.

[0063] Step S6: Perform data cleaning on the multiple contract documents to be processed based on the benchmark file.

[0064] For example, the process of performing data cleansing on multiple contract documents to be processed based on the benchmark file may be: First, parse the benchmark file to obtain its field specifications (such as field name, data type, constraints (such as non-null, uniqueness, value range, etc.)) and relationships (such as foreign key constraints). At the same time, extract field descriptions and annotation information to form a field mapping table as the data cleaning standard. It then reads all pending contract files except the baseline file and converts them into the target intermediate format (such as a DataFrame) based on the field mapping table. It automatically handles format differences during this process (for example, converting date fields to "YYYY-MM-DD" format and removing thousandths signs from numeric fields). After completing the format conversion, data cleansing operations are performed on multiple pending contract files one by one according to the constraint rules of the benchmark file (referring to field specifications and association relationships) to convert multiple pending contract files into cleaned files in EXCEL format. Each time a data cleansing operation is performed, a log is generated to record abnormal data and data cleansing details.

[0065] During the data cleaning process, duplicate values ​​are removed based on the primary key field, retaining the first occurrence or the latest record by timestamp. Missing values ​​are filled in according to the default values ​​configured in the benchmark file (e.g., 0 for numeric values ​​and "unknown" for text values) or completed through other data tables based on associated fields (e.g., associating customer information through order ID). Outliers are automatically corrected or marked based on the preset range (e.g., age 0-120 years old). For association verification, the integrity of foreign key references is checked (e.g., the department ID in the employee table must exist in the department table).

[0066] The present invention first comprehensively determines the file complexity of each contract file to be processed based on the data type complexity and data volume of the contract file to be processed. Then, based on the file complexity of each contract file to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information, the requirement matching degree of the contract file to be processed is comprehensively determined from multiple aspects, such as the file itself and the degree of adaptation between the file and the requirement information. Therefore, the file with the highest requirement matching degree in each file type is determined as the key file. Then, based on the file complexity of each contract file to be processed, the type complexity of each file type is determined. Therefore, the most important key file is identified from a number of key files. That is, by comprehensively considering the correlations between multiple contract files to be processed and the correlations between multiple contract files to be processed and the requirement information, a benchmark file that matches the user's requirements and has a strong correlation with other contract files to be processed is selected. Finally, data cleaning of multiple contract files to be processed is completed based on the benchmark file. This can reduce the probability of incorrectly cleaning important data in the file and the probability of omitting cleaning of redundant data in the file, so that multiple contract files to be processed can achieve better data cleaning results.

[0067] In one embodiment, the step of obtaining a first degree of matching between the file name of each to-be-processed contract document and the requirement information includes: Perform word segmentation on the file name of each pending contract document to obtain a title phrase of each pending contract document; and perform word segmentation on the demand information to obtain a demand phrase; The intersection-and-union ratio between the title phrase and the requirement phrase of each to-be-processed contract document is determined as the first matching degree between the file name of each to-be-processed contract document and the requirement information.

[0068] Exemplarily, the above word segmentation processing can be completed by Jieba word segmentation tool.

[0069] The intersection-over-union ratio between the title phrase and the requirement phrase of each pending contract document can also be understood as the Jaccard correlation coefficient between the title phrase and the requirement phrase of each pending contract document.

[0070] For example, if the title phrase of a pending contract document is set to [a1, a2, a3, a4] and the requirement phrase is [a2, a3, a5], then the intersection between the title phrase and the requirement phrase is [a2, a3], and the union is [a1, a2, a3, a4, a5], and the intersection-union ratio is 0.4, that is, the first matching degree between the pending contract document and the requirement information is 0.4.

[0071] In this embodiment, the first matching degree between the file name of each contract document to be processed and the requirement information is conveniently determined by segmenting the content and calculating the intersection-union ratio between different phrases obtained by segmentation, thereby improving the data cleaning efficiency of the method of the present invention.

[0072] In one embodiment, the step of obtaining the second degree of matching between the content of each to-be-processed contract document and the requirement information includes: Perform word segmentation on the content of each pending contract document to obtain content phrases of each pending contract document; and perform word segmentation on the demand information to obtain demand phrases; Performing semantic expansion on the content phrases of each pending contract document to obtain a hypernym of each pending contract document; and performing semantic expansion on the requirement phrases to obtain a hyponym, wherein a plurality of first participles in the content phrases of each pending contract document correspond one-to-one to a plurality of second participles in the corresponding hypernym, the second participles being hypernyms of the corresponding first participles, a plurality of third participles included in the requirement phrases correspond one-to-one to a plurality of fourth participles included in the hyponym, the fourth participles being hyponyms of the corresponding third participles; An association analysis is performed on the hypernym and hyponym of each contract document to be processed to obtain a second matching degree between the file content of each contract document to be processed and the requirement information.

[0073] In the present invention, semantic expansion is used to find the hypernyms / hyponyms of the input word in the vocabulary based on the pre-established semantic mapping. Among them, hypernyms usually represent words with broader and more general concepts, which include the concepts represented by the corresponding hyponyms, for example: "animal" is the hypernym of "dog", and "fruit" is the hypernym of "apple". Hyponyms usually represent words with more specific and more particular concepts, which belong to a subcategory of the concept referred to by the corresponding hypernym, for example: "rose" is the hyponym of "flower", and "golden retriever" is the hyponym of "dog". The aforementioned semantic mapping is used to describe the hierarchical relationship between the upper and lower levels of each word in the vocabulary.

[0074] Among them, the one-to-one correspondence between the multiple first participles in the content phrase of each pending contract document and the multiple second participles in the corresponding superordinate phrase is: any first participle among the multiple first participles in the content phrase of each pending contract document can find a corresponding superordinate phrase in the corresponding superordinate phrase.

[0075] Similarly, the one-to-one correspondence between the multiple third participles included in the demand phrase and the multiple fourth participles included in the hyponym phrase is as follows: any third participle among the multiple third participles included in the demand phrase can find corresponding hyponyms in the corresponding hyponym phrase.

[0076] In this embodiment, the hypernyms of each phrase in the content phrase of the contract document to be processed are obtained to form the hypernym of the contract document to be processed, and the hyponyms of each phrase in the requirement phrase of the requirement information are obtained to form the hyponym. The second matching degree between the file content of the contract document to be processed and the requirement information is determined by performing an association analysis on the two, so as to utilize the hypernym and hyponym semantic relationship between the participles to conveniently complete the accurate quantification of the matching degree between the file content of the contract document to be processed and the requirement information.

[0077] In the above setting, by obtaining the hyponyms of each phrase in the demand phrase of the demand information, a reasonable expansion of the content of the demand information can be achieved, so as to fully reflect the semantic correlation between the demand information and the file content of the contract document to be processed in conjunction with the hyponyms of the contract document to be processed, so that the calculated second matching degree is more accurate and reliable.

[0078] The second matching degree between the content of each contract document to be processed and the requirement information is obtained by performing association analysis on the hypernym and hyponym of each contract document to be processed, including: Taking the intersection of the hypernym and hyponym of each pending contract document to obtain the intersection phrase of each pending contract document; The ratio of the number of segment words included in the intersection phrase of each contract document to be processed to the number of segment words included in the hypernym phrase is determined as the second matching degree between the file content of each contract document to be processed and the demand information.

[0079] For example, the word segmentation array (i.e., the aforementioned content phrase) for the nth pending contract document under the mth file type can be denoted as word segmentation array n. WordNet is used to obtain the hypernyms of each word in word segmentation array n to form a hypernym for the nth pending contract document under the mth file type. Assuming there are R word segmentations in word segmentation array n, the hypernym for the nth pending contract document under the mth file type will include R' hypernyms, where R' is less than or equal to R.

[0080] Identify whether R' hypernyms appear in the hyponym group corresponding to the demand information. If so, the word segmentation correlation between the content of the nth pending contract document under the mth file type and the demand information is added by 1, and then the word segmentation correlation between the content of the nth pending contract document under the mth file type and the demand information can be obtained. .

[0081] That is, the mathematical formula for the matching degree between the file content of the nth contract file to be processed under the mth file type and the required information (also known as the second matching degree) can be expressed as:

[0082] Where, Indicates the matching degree between the content of the nth contract document to be processed under the mth file type and the requirement information. It represents the word segmentation correlation between the file content and requirement information of the nth contract file to be processed under the mth file type, and R represents the number of word segments included in the word segmentation array n.

[0083] In one embodiment, the requirement matching degree of each contract document to be processed is obtained based on the file complexity of each contract document to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information, including: The product of the file complexity of each contract file to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information is determined as the requirement matching degree of each contract file to be processed.

[0084] For example, if the requirement matching degree of the nth contract document to be processed under the mth file type is set to ,but The mathematical formula can be expressed as:

[0085] in, Indicates the second matching degree between the file content of the nth contract file to be processed under the mth file type and the requirement information, Indicates the file complexity of the nth contract file to be processed under the mth file type, Indicates the first matching degree between the file name of the nth contract file to be processed under the mth file type and the requirement information.

[0086] In one embodiment, the step of obtaining the type complexity of each file type includes: Calculating the average of multiple file complexities of multiple contract files to be processed included in each file type to obtain the average complexity corresponding to each file type; According to the complexity mean corresponding to each file type, the type complexity of each file type is obtained.

[0087] Among them, according to the complexity mean corresponding to each file type, the type complexity of each file type is obtained, including: Determine the ratio of the number of pending contract documents included in each file type to the total number of pending contract documents included in multiple file types as the type coefficient of each file type; The product of the complexity mean corresponding to each file type and the type coefficient of each file is determined as the type complexity of each file type.

[0088] For example, if you set represents the type complexity of the mth file type among the aforementioned multiple file types, then The calculation formula can be expressed as:

[0089] in, Indicates the number of pending contract documents included in the mth file type. Indicates the total number of contract documents to be processed including multiple file types, M is the number of multiple file types, Indicates the file complexity of the nth contract file to be processed under the mth file type, It can be understood as the type coefficient of the mth file type. It can be understood as the average complexity of multiple files corresponding to the mth file type.

[0090] In this embodiment, the importance of each file type among multiple file types is represented based on the proportion of the number of pending contract files included in each file type among all pending contract files (the higher the proportion, the higher the importance of the corresponding file type). The importance of the multiple pending contract files included in each file type is represented based on the average of the multiple file complexities of the multiple pending contract files included in each file type. By comprehensively analyzing the importance of each file type in both horizontal (among multiple file types) and vertical (including multiple pending contract files) dimensions, the importance of each file type (i.e., type complexity) is comprehensively and accurately evaluated, making the determined type complexity more accurate.

[0091] In one embodiment, a reference file is identified among a plurality of key files, including: determining the target file type as a primary type, and determining other file types except the primary type among the plurality of file types as secondary types; The difference between the type complexity of the main type and the type complexity of each sub-type is determined as the type difference value of each sub-type; Sum the type difference values ​​of each sub-type to obtain the difference sum value; When the difference sum value is greater than a preset difference threshold, the key file corresponding to the target file type is determined as the benchmark file, wherein the difference threshold is a negative number; When the difference sum value is less than or equal to the difference threshold, the key file corresponding to the subtype with the highest type complexity among the multiple subtypes is determined as the benchmark file.

[0092] The above differences and values ​​are used to indicate the degree of difference between the type complexity of the target file type and the type complexity of other file types.

[0093] For example, if the difference and value are set to D, the calculation formula of D can be expressed as:

[0094] in, Indicates the type complexity of the main type, represents the type complexity of the jth subtype among M file types, where M is the number of the aforementioned multiple file types. In this example, the difference threshold can be set to -0.25.

[0095] When the difference sum is greater than the preset difference threshold (that is, the difference sum is a positive value or a small negative number), it means that among multiple file types, the file type with the highest complexity is the main type (or the main type ranks first in complexity). In other words, the main type is not only the file type specified by the user, but also the main type includes multiple pending contract files (the number of pending contract files, the amount of data of each pending contract file, and the number of data types of each pending contract file) with the richness of content that ranks first or ranks first among multiple file types. Therefore, determining the key file corresponding to the main type as the benchmark file can more effectively guide the data cleaning work of the remaining multiple pending contract files.

[0096] When the difference sum is less than or equal to the difference threshold (i.e., the difference sum is a large negative number), it means that although the main type is the file type specified by the user, the comprehensive content richness of the multiple contract files to be processed included therein is relatively low among the multiple file types. Therefore, if the key file corresponding to the main type is mechanically determined as the benchmark file, it is easy to cause large errors in the data cleaning work of the remaining multiple contract files to be processed. In this case, choosing the key file corresponding to the sub-type with the highest type complexity among the multiple sub-types as the benchmark file can effectively overcome the interference problem caused by inappropriate file types specified by the user, so that the data cleaning work of the remaining multiple contract files to be processed can be better performed, that is, multiple PDF format contract files to be processed can be converted into EXCEL file format more accurately.

[0097] The present invention proposes an automatic file data cleaning system for a file system, see Figure 2 , which shows a schematic structural diagram of an automatic file data cleaning system 200 for a file system provided by an embodiment of the present invention, the system includes: An acquisition module 201 is configured to acquire a type identifier, requirement information, multiple file types, and multiple contract files to be processed included in each file type, wherein the type identifier is used to indicate a target file type among the multiple file types; The first processing module 202 is configured to obtain the file complexity of each contract file to be processed based on the data type complexity and data volume of each contract file to be processed; The second processing module 203 is configured to obtain a requirement matching degree for each contract document to be processed based on the file complexity of each contract document to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information; The key file determination module 204 is configured to determine the key file corresponding to each file type based on the requirement matching degree of each pending contract file, wherein the key file is the pending contract file with the highest requirement matching degree among the multiple pending contract files of the corresponding file type; A reference file identification module 205 is configured to identify a reference file from among the multiple key files, wherein the reference file is: a key file corresponding to the target file type, and / or a key file corresponding to a file type with the highest type complexity among the multiple file types, where the type complexity is calculated based on the complexity of multiple files associated with the corresponding file type; The data cleaning module 206 is used to perform data cleaning on the plurality of contract files to be processed based on the reference file.

[0098] It should be noted that the system provided in the above embodiment is merely an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the automated file data cleaning system for a file system provided in the above embodiment and the automated file data cleaning method for a file system are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0099] The embodiment of the present invention also provides an electronic device. Figure 3 , the electronic device may include a processor 301, a memory 302, and a program 3021 stored in the memory 302 and executable on the processor 301.

[0100] When the program 3021 is executed by the processor 301, it can achieve Figure 1 Any steps in the corresponding method embodiments and achieving the same beneficial effects will not be repeated here.

[0101] Those skilled in the art will appreciate that all or part of the steps of implementing the above-described embodiment method can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0102] The embodiment of the present invention further provides a readable storage medium, on which a computer program is stored, which can realize the above-mentioned Figure 1 Any steps in the corresponding method embodiments can achieve the same technical effects and will not be described again here to avoid repetition.

[0103] The computer-readable storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component.

[0104] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0105] The program code contained on the storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0106] Computer program code for performing the operations of the present invention may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or terminal. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0107] An embodiment of the present invention further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement an automatic file data cleaning method for a file system provided in the above embodiment.

[0108] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0109] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. A method for automatically cleaning file data in a file system, characterized in that: Methods include: Acquire a type identifier, requirement information, multiple file types, and multiple contract files to be processed included in each file type, where the type identifier is used to indicate a target file type among the multiple file types; Obtain the file complexity of each contract document to be processed based on the data type complexity and data volume of each contract document to be processed; Obtaining the requirement matching degree of each contract document to be processed based on the file complexity of each contract document to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information; Determine the key file corresponding to each file type based on the demand matching degree of each pending contract file, wherein the key file is the pending contract file with the highest demand matching degree among the multiple pending contract files of the corresponding file type; Identifying a benchmark file among the multiple key files, wherein the benchmark file is: a key file corresponding to the target file type, and / or a key file corresponding to a file type with the highest type complexity among the multiple file types, where the type complexity is calculated based on the complexity of multiple files associated with the corresponding file type; Perform data cleansing on multiple pending contract documents based on benchmark files.

2. The automatic file data cleaning method of the file system according to claim 1, characterized in that: According to the data type complexity and data volume of each contract document to be processed, the file complexity of each contract document to be processed is obtained, including: Determine the data type complexity of each contract file to be processed as the ratio of the first number to the second number corresponding to each contract file to be processed, wherein the first number is the number of data types included in the corresponding contract file to be processed, and the second number is the total number of data types included in the file type to which the corresponding contract file to be processed belongs; The product of the data type complexity and the data volume of each contract document to be processed is determined as the original complexity of each contract document to be processed; The original complexity of each contract document to be processed is normalized to obtain the file complexity of each contract document to be processed.

3. The automatic file data cleaning method of the file system according to claim 1, characterized in that: The step of obtaining the first matching degree between the file name of each contract document to be processed and the requirement information includes: Perform word segmentation on the file name of each pending contract document to obtain a title phrase of each pending contract document; and perform word segmentation on the demand information to obtain a demand phrase; The intersection-and-union ratio between the title phrase and the requirement phrase of each to-be-processed contract document is determined as the first matching degree between the file name of each to-be-processed contract document and the requirement information.

4. The automatic file data cleaning method of a file system according to claim 1, characterized in that: The step of obtaining the second matching degree between the document content of each to-be-processed contract document and the requirement information includes: Perform word segmentation on the content of each pending contract document to obtain content phrases of each pending contract document; and perform word segmentation on the demand information to obtain demand phrases; Performing semantic expansion on the content phrases of each pending contract document to obtain a hypernym of each pending contract document; and performing semantic expansion on the requirement phrases to obtain a hyponym, wherein a plurality of first participles in the content phrases of each pending contract document correspond one-to-one to a plurality of second participles in the corresponding hypernym, the second participles being hypernyms of the corresponding first participles, a plurality of third participles included in the requirement phrases correspond one-to-one to a plurality of fourth participles included in the hyponym, the fourth participles being hyponyms of the corresponding third participles; An association analysis is performed on the hypernym and hyponym of each contract document to be processed to obtain a second matching degree between the file content of each contract document to be processed and the requirement information.

5. The automatic file data cleaning method of the file system according to claim 4, characterized in that: Perform association analysis on the hypernyms and hyponyms of each pending contract document to obtain a second matching degree between the content of each pending contract document and the required information, including: Taking the intersection of the hypernym and hyponym of each pending contract document to obtain the intersection phrase of each pending contract document; The ratio of the number of segment words included in the intersection phrase of each contract document to be processed to the number of segment words included in the hypernym phrase is determined as the second matching degree between the file content of each contract document to be processed and the demand information.

6. The automatic file data cleaning method of a file system according to claim 1, characterized in that: According to the file complexity of each contract file to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information, the requirement matching degree of each contract file to be processed is obtained, including: The product of the file complexity of each contract file to be processed, the first matching degree between the file name and the requirement information, and the second matching degree between the file content and the requirement information is determined as the requirement matching degree of each contract file to be processed.

7. The automatic file data cleaning method of a file system according to claim 1, characterized in that: The steps for obtaining the type complexity of each file type include: Calculating the average of multiple file complexities of multiple contract files to be processed included in each file type to obtain the average complexity corresponding to each file type; According to the complexity mean corresponding to each file type, the type complexity of each file type is obtained.

8. The automatic file data cleaning method of the file system according to claim 7, characterized in that: According to the complexity mean corresponding to each file type, the type complexity of each file type is obtained, including: Determine the ratio of the number of pending contract documents included in each file type to the total number of pending contract documents included in multiple file types as the type coefficient of each file type; The product of the complexity mean corresponding to each file type and the type coefficient of each file is determined as the type complexity of each file type.

9. The automatic file data cleaning method of a file system according to claim 1, characterized in that: Identify the benchmark file among several key files, including: determining the target file type as a primary type, and determining other file types except the primary type among the plurality of file types as secondary types; The difference between the type complexity of the main type and the type complexity of each sub-type is determined as the type difference value of each sub-type; Sum the type difference values ​​of each sub-type to obtain the difference sum value; When the difference sum value is greater than a preset difference threshold, the key file corresponding to the target file type is determined as the benchmark file, wherein the difference threshold is a negative number; When the difference sum value is less than or equal to the difference threshold, the key file corresponding to the subtype with the highest type complexity among the multiple subtypes is determined as the benchmark file.

10. The automatic file data cleaning method of a file system according to claim 1, characterized in that: There are multiple steps to obtain the contract documents to be processed, including: Among multiple original files, content recognition is performed on each original file to obtain the original content of each original file, where the original file is a file input into the file system by the user; Preprocessing the original content of each original file to obtain preprocessed content of each original file, wherein the preprocessing includes at least one of missing value processing, outlier processing, and duplicate value processing; The file name of each original file and the corresponding pre-processed content are spliced ​​to obtain multiple contract files to be processed, wherein the file name of the contract file to be processed is the file name of the corresponding original file, and the file content of the contract file to be processed is the pre-processed content of the corresponding original file.

Citation Information

Patent Citations

  • Intelligent semantic recognition method based on HSE

    CN111291562A

  • File processing method and device, electronic equipment and storage medium

    CN115269832A

  • Contract data verification method and device and server

    CN117036115A

  • File duplicate checking method and device, equipment, storage medium and program product

    CN118733717A

  • Dictionary file generation method and device, equipment and storage medium

    CN119292715A