Report file similarity detection method and device and medium

By extracting and sorting report header fields and combining common and intelligent algorithms to calculate similarity, the problem of low efficiency and accuracy in similarity detection of grassroots reports is solved, and fast and accurate report file similarity detection and data combing are achieved.

CN120745596APending Publication Date: 2025-10-03INSPUR ZHUOSHU BIG DATA IND DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511026894.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the existing technology, the similarity detection efficiency and accuracy of grassroots reports are low, resulting in a large amount of manpower and material resources invested and low detection efficiency.

Method used

By extracting the header fields of the Excel format report, sorting them using UTF-8 encoding, and combining common algorithms and artificial intelligence algorithms to calculate the similarity, the similarity between the report file and the target file list is determined and stored in the target or duplicate file list.

Benefits of technology

It achieves efficient and accurate report file similarity detection, quickly identifies duplicate reports, provides a reliable data foundation, lays the foundation for subsequent data analysis and integration operations, and supports data quality monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745596A_ABST
    Figure CN120745596A_ABST
Patent Text Reader

Abstract

The invention discloses a report file similarity detection method and device and a medium. The method comprises the steps that a to-be-detected report file in an excel format is read; extracting header fields of the report file to be detected, and sorting the header fields according to the UTF-8 code; according to the sorted header fields, performing similarity calculation on the to-be-detected report file and a target report file in the target file list; the target file list does not have the same target report file; judging whether a target report file with the similarity greater than or equal to a preset similarity threshold exists or not; if not, storing the to-be-detected report file to a target file list; and if yes, storing the to-be-detected report file to a duplicate file list. And similarity detection can be performed on the report file more efficiently and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a report file similarity detection method, device, and medium. Background Art

[0002] Grassroots reporting forms are the foundation for higher-level authorities to understand public sentiment, formulate policies, and carry out their work. Their importance cannot be underestimated. However, the numerous forms and frequent requests for materials force officials to "circle around forms" and become "stuck in the documents," leading to repeated submission of the same information and multiple reports.

[0003] There are usually thousands of reports of various forms, which need to be filled out periodically, and tens of thousands of reports need to be submitted every year. Therefore, if you want to streamline the reports, you must first sort out the types of existing reports. However, at present, there are thousands of reports of various forms in every place. For example, some titles are horizontal, some titles are vertical, some titles are only one line, and some titles are three or five lines. In addition, the reports that collect data for the same thing have different formats and numbers of items in different towns and streets, and different districts and counties; for the reports that collect data for the same thing, this year's report format and last year's report format and number of items are not exactly the same, and the report names are also different. For example, the statistical table of wheat planting area in District A in 2024 and the statistical table of wheat planting area in District A in 2025 are actually the same report. In addition, the specific expressions of the same item are not exactly the same, such as "mobile phone", "mobile phone number", "telephone number", "contact information", etc.

[0004] Because reports for the same event vary in time and space, and data titles have semantic differences, similarity detection currently relies solely on file names to determine similarity, resulting in low accuracy. Manual identification requires significant manpower and resources, and identifying, screening, and merging similar items from thousands of reports also results in low detection efficiency. Summary of the Invention

[0005] The embodiments of the present application provide a report file similarity detection method, device, and medium for solving the problem of low efficiency and accuracy in report file similarity detection.

[0006] The embodiments of this application adopt the following technical solutions:

[0007] On the one hand, an embodiment of the present application provides a report file similarity detection method, the method comprising: reading a report file to be detected in Excel format; extracting header fields of the report file to be detected, and sorting the header fields according to UTF-8 encoding; calculating the similarity between the report file to be detected and a target report file in a target file list based on the sorted header fields; the target file list does not have the same target report file; judging whether there is a target report file with a similarity greater than or equal to a preset similarity threshold; if not, storing the report file to be detected in the target file list; if so, storing the report file to be detected in a duplicate file list.

[0008] In one example, the similarity calculation between the report file to be detected and the target report file in the target file list is performed based on the sorted header fields, specifically including: connecting the sorted header fields into a string; comparing the string of the report file to be detected with the string of the target report file in the target file list to obtain the similarity between the report file to be detected and the target report file.

[0009] In one example, the value range of the similarity is greater than 0 and less than or equal to 1, and the character string of the report file to be detected is compared with the character string of the target report file in the target file list to obtain the similarity between the report file to be detected and the target report file, specifically including: comparing the character string of the report file to be detected with the character string of the target report file in the target file list character by character; when they are exactly the same, determining that the similarity between the report file to be detected and the target report file is 1; when they are not exactly the same, calculating the similarity between the report file to be detected and the target report file according to the similarity intelligent model.

[0010] In one example, the similarity between the report file to be detected and the target report file is calculated based on the similarity intelligent model, specifically including: generating a report word vector to be detected for each field of the report file to be detected, and generating a target report word vector for each field of the target report file; performing weighted averaging on multiple report word vectors to be detected to obtain a report sentence vector to be detected; and performing weighted averaging on multiple target report word vectors to obtain a target report sentence vector; and obtaining the similarity between the report file to be detected and the target report file based on the cosine similarity between the report sentence vector to be detected and the target report sentence vector.

[0011] In one example, obtaining the similarity between the report file to be detected and the target report file based on the cosine similarity between the report sentence vector to be detected and the target report sentence vector specifically includes: calculating the cosine similarity between the report sentence vector to be detected and the target report sentence vector; matching the cosine similarity in a similarity mapping relationship table to obtain the similarity between the report file to be detected and the target report file.

[0012] In one example, storing the report file to be detected in a duplicate file list specifically includes:

[0013] Determine the storage information of the report file to be detected; the storage information includes file name, file storage path, header field, file name of the target report file, and similarity between the target report file and the target report file; based on the storage information, store the report file to be detected in a duplicate file list.

[0014] In one example, the method further includes: when there are multiple target report files with a similarity greater than or equal to a preset similarity threshold, determining multiple storage information of the report file to be detected; and storing the report file to be detected multiple times in a duplicate file list based on the multiple storage information.

[0015] In one example, after storing the report file to be detected in the duplicate file list, the method further includes: receiving a judgment result of the client; when the judgment result is a duplicate file, marking the report file to be detected as a duplicate file; when the judgment result is a non-duplicate file, storing the report file to be detected in the target report file.

[0016] On the other hand, an embodiment of the present application provides a report file similarity detection device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a report file similarity detection method described in any one of the above items.

[0017] On the other hand, an embodiment of the present application provides a non-volatile computer storage medium for report file similarity detection, which stores computer-executable instructions. The computer-executable instructions can execute any of the above-mentioned report file similarity detection methods.

[0018] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:

[0019] Compare based on table header fields (rather than the entire data), focusing on the core structural features of the report to avoid structural similarities being masked by differences in data content (such as data from different periods in the same report).

[0020] By extracting header fields and calculating the similarity of the sorted header fields, it is possible to automatically identify reports to be detected that are highly similar to those in the target file list, and to more quickly compare duplicate reports (such as reports with the same fields but in a different order), thus preventing duplicate files from being stored in the target file list.

[0021] The target file list stores only unique reports, ensuring its uniqueness as a benchmark dataset and providing a reliable data foundation for subsequent data analysis, report integration, and other operations based on the list. Categorized storage of duplicate file lists facilitates tracing the source of duplicate data and supports data quality monitoring (such as identifying abnormal data from repeated reports).

[0022] To sum up, report file similarity detection can be performed more efficiently and accurately, thereby quickly completing grassroots report sorting. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solution of the present application, some embodiments of the present application will be described in detail below with reference to the accompanying drawings, in which:

[0024] Figure 1 A flowchart of a report file similarity detection method provided in an embodiment of the present application;

[0025] Figure 2 A schematic diagram of the structure of a report file similarity detection device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0027] Some embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0028] Figure 1This is a flow chart of a report file similarity detection method provided in an embodiment of the present application. This method can be applied to various business areas, such as internet finance, e-commerce, instant messaging, gaming, and government affairs. The process can be executed by computing devices in the corresponding fields, and certain input parameters or intermediate results in the process can be manually adjusted to help improve accuracy.

[0029] The analysis method involved in the embodiments of the present application can be implemented by a terminal device or a server, and the present application does not impose any special restrictions on this. For ease of understanding and description, the following embodiments are described in detail using a server as an example.

[0030] It should be noted that the server can be a single device or a system composed of multiple devices, that is, a distributed server, and this application does not make any specific restrictions on this.

[0031] Figure 1 The process in includes the following steps:

[0032] S101: Reading a report file to be tested in Excel format.

[0033] Among them, all report files are read from the directory specified by the user.

[0034] S102: Extracting header fields of the report file to be detected, and sorting the header fields according to UTF-8 encoding.

[0035] It should be noted that sorting can allow the header fields to be presented according to unified coding rules, which facilitates data processing, comparison or standardized management.

[0036] S103: performing similarity calculation between the report file to be detected and the target report file in the target file list according to the sorted header fields; the target file list does not contain the same target report file.

[0037] In some embodiments of the present application, all files in the target list are result files of the base report combing, and the report file information stored in the target file list includes file name, file path, and title item (header field). The similarity value range is greater than 0 and less than or equal to 1, and the similarity is calculated as follows:

[0038] First, the sorted header fields are concatenated into a string. The header fields are concatenated in order to obtain a string.

[0039] Then, the character string of the report file to be detected is compared with the character string of the target report file in the target file list to obtain the similarity between the report file to be detected and the target report file.

[0040] Among them, the comparison process can be combined with ordinary algorithm comparison and artificial intelligence algorithm comparison.

[0041] Based on this, the character string of the report file to be detected is compared character by character with the character string of the target report file in the target file list.

[0042] When they are exactly the same, the similarity between the report file to be detected and the target report file is determined to be 1.

[0043] That is, when they are exactly the same, it can be said that the report file to be detected and the target report file in the target file list are duplicates.

[0044] When they are not completely identical, the similarity between the report file to be detected and the target report file is calculated based on the sorted header fields and the similarity intelligent model.

[0045] Based on this, artificial intelligence algorithms can choose cosine similarity, Word2Vec, BERT-based similarity calculation methods, etc. When using Word2Vec, the similarity intelligent model calculation process can be as follows:

[0046] First, a to-be-detected report word vector is generated for each field of the to-be-detected report file, and a target report word vector is generated for each field of the target report file.

[0047] It should be noted that word vectors can be generated by using a pre-trained word vector model. When the fields are synonyms, the word vectors are usually the same or similar.

[0048] Then, a weighted average is performed on the word vectors of multiple reports to be detected to obtain the sentence vector of the report to be detected. A weighted average is also performed on the word vectors of multiple target reports to obtain the sentence vector of the target report.

[0049] Finally, the similarity between the report file to be detected and the target report file is obtained based on the cosine similarity between the sentence vector of the report to be detected and the sentence vector of the target report.

[0050] The cosine similarity between the sentence vector of the report to be detected and the sentence vector of the target report is calculated. Then, the cosine similarity is matched in the similarity mapping relationship table to obtain the similarity between the report file to be detected and the target report file.

[0051] It should be noted that the sentence vector solution can achieve the following effects:

[0052] Dimensionality reduction aggregation: compress multiple word vectors into a single sentence vector.

[0053] Semantic Fusion: Weighted averaging preserves core semantic features and can retain similarities between synonyms even if they are not in the same ranking position.

[0054] It should be noted that the character string of the report file to be detected and the character string of the target report file may also be input into the TF-IDF function to output the similarity between the report file to be detected and the target report file.

[0055] S104: Determine whether there is a target report file with a similarity greater than or equal to a preset similarity threshold.

[0056] The report file to be detected is compared one by one with each target report file in the target file list.

[0057] It should be noted that when the similarity is greater than or equal to a preset similarity threshold, the report file to be detected and the target report file are determined to be duplicate files.

[0058] S105: If no such file exists, the report file to be detected is stored in the target file list.

[0059] The stored content includes file name, file path, and header fields.

[0060] S106: If there is a duplicate file, store the report file to be detected in a duplicate file list.

[0061] It should be noted that all report files in the duplicate file list are duplicate report files to be discarded in the grassroots report sorting work. The report file information stored in the duplicate file list includes file name, file path, title item (header field), file name of the target report file, and similarity between the target report file and the target report file.

[0062] In some embodiments of the present application, if the report file to be detected is similar to multiple target report files, the report file to be detected needs to be stored multiple times in the duplicate file list to facilitate further manual judgment.

[0063] Based on this, when there are multiple target report files with similarity greater than or equal to a preset similarity threshold, multiple storage information of the report file to be detected is determined. Then, based on the multiple storage information, the report file to be detected is stored multiple times in the duplicate file list.

[0064] It should be noted that after the report files in the duplicate file list are manually judged and confirmed, they will be marked as duplicate files and wait for regular deletion. In other words, the unmarked report files are files waiting for manual judgment.

[0065] Based on this, the client receives the judgment result. If the judgment result is a duplicate file, the report file to be detected is marked as a duplicate file. Then, if the judgment result is a non-duplicate file, the report file to be detected is stored in the target report file.

[0066] It should be noted that although the embodiments of this application are based on Figure 1 To introduce and explain step S101 to step S106 in sequence, but this does not mean that step S101 to step S106 must be executed in a strict order. Figure 1 The order shown in FIG1 is to introduce and explain steps S101 to S106 in order to facilitate those skilled in the art to understand the technical solutions of the embodiments of the present application. In other words, in the embodiments of the present application, the order of steps S101 to S106 can be adjusted appropriately according to actual needs.

[0067] It's important to clarify that, among the myriad of diverse reports, the types of all report files must be identified from a business-essential perspective. In other words, reports developed for the same data collection purpose should be considered the same type, even if they differ slightly in terms of space, time, and semantics.

[0068] Based on this, the comparison is based on the header fields (rather than the full data), focusing on the core structural features of the report to avoid masking structural similarities due to differences in data content (such as different period data of the same type of report).

[0069] By extracting header fields and calculating the similarity of the sorted header fields, it is possible to automatically identify reports to be detected that are highly similar to those in the target file list, and to more quickly compare duplicate reports (such as reports with the same fields but in a different order), thus preventing duplicate files from being stored in the target file list.

[0070] The target file list stores only unique reports, ensuring its uniqueness as a benchmark dataset and providing a reliable data foundation for subsequent data analysis, report integration, and other operations based on the list. Categorized storage of duplicate file lists facilitates tracing the source of duplicate data and supports data quality monitoring (such as identifying abnormal data from repeated reports).

[0071] To sum up, report file similarity detection can be performed more efficiently and accurately, thereby quickly completing grassroots report sorting.

[0072] Furthermore, when calculating similarity, combining conventional algorithms with intelligent algorithms, leveraging the fuzzy matching capabilities of AI algorithms, can resolve semantic issues such as similar characters with different meanings and slight differences in the number of items. In other words, conventional algorithms offer speed while AI algorithms offer fuzzy intelligence. This combination ensures both high efficiency and intelligent accuracy.

[0073] Furthermore, the duplicate file list stores information such as file name, file path, title item, duplicate file name, similarity, etc., which can be used for further manual comparison and verification, and final selection and rejection decisions are made to ensure that the report sorting work is rigorous and no meaningful report is missed.

[0074] Based on the same idea, some embodiments of the present application also provide devices and non-volatile computer storage media corresponding to the above methods.

[0075] Figure 2 A schematic diagram of a report file similarity detection device provided in an embodiment of the present application includes:

[0076] at least one processor; and,

[0077] a memory communicatively connected to the at least one processor; wherein,

[0078] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the above-mentioned report file similarity detection methods.

[0079] Some embodiments of the present application provide a non-volatile computer storage medium for report file similarity detection, which stores computer-executable instructions. The computer-executable instructions can execute any of the above-mentioned report file similarity detection methods.

[0080] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device and medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.

[0081] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0082] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0083] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0084] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0086] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0087] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0088] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0089] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0090] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the technical principles of the present application should fall within the scope of protection of the present application.

Claims

1. A report file similarity detection method, characterized in that: The method comprises: Read the report file to be tested in Excel format; Extract the header fields of the report file to be tested, and sort the header fields according to UTF-8 encoding; Calculating similarity between the report file to be detected and the target report file in the target file list according to the sorted header fields; the target file list does not contain the same target report file; Determine whether there is a target report file with a similarity greater than or equal to a preset similarity threshold; If the target file does not exist, the report file to be detected is stored in the target file list; If there is a duplicate file, the report file to be detected is stored in a duplicate file list.

2. The method according to claim 1, characterized in that The calculating of similarity between the report file to be detected and the target report file in the target file list according to the sorted header fields specifically includes: Concatenate the sorted header fields into a string; The character string of the report file to be detected is compared with the character string of the target report file in the target file list to obtain the similarity between the report file to be detected and the target report file.

3. The method according to claim 2, characterized in that The similarity value range is greater than 0 and less than or equal to 1. The character string of the report file to be detected is compared with the character string of the target report file in the target file list to obtain the similarity between the report file to be detected and the target report file, specifically including: Compare the character string of the report file to be detected with the character string of the target report file in the target file list character by character; When they are exactly the same, the similarity between the report file to be detected and the target report file is determined to be 1; When they are not completely identical, the similarity between the report file to be detected and the target report file is calculated based on the similarity intelligent model.

4. The method according to claim 3, characterized in that The calculating the similarity between the report file to be detected and the target report file according to the similarity intelligent model specifically includes: Generating a to-be-detected report word vector for each field of the to-be-detected report file, and generating a target report word vector for each field of the target report file; Perform weighted averaging on multiple to-be-detected report word vectors to obtain a to-be-detected report sentence vector; and perform weighted averaging on multiple target report word vectors to obtain a target report sentence vector; The similarity between the report file to be detected and the target report file is obtained according to the cosine similarity between the report sentence vector to be detected and the target report sentence vector.

5. The method according to claim 4, characterized in that The obtaining of the similarity between the report file to be detected and the target report file based on the cosine similarity between the report sentence vector to be detected and the target report sentence vector specifically includes: Calculating the cosine similarity between the report sentence vector to be detected and the target report sentence vector; In the similarity mapping relationship table, the cosine similarity is matched to obtain the similarity between the report file to be detected and the target report file.

6. The method according to claim 1, characterized in that The storing of the report file to be detected into the duplicate file list specifically includes: Determine the storage information of the report file to be detected; the storage information includes the file name, file storage path, header field, file name of the target report file, and similarity between the target report file and the target report file; According to the storage information, the report file to be detected is stored in a duplicate file list.

7. The method according to claim 6, characterized in that The method further comprises: When there are multiple target report files with a similarity greater than or equal to a preset similarity threshold, determining multiple storage information of the report file to be detected; According to the multiple storage information, the report file to be detected is stored multiple times in a duplicate file list.

8. The method according to claim 1, characterized in that After storing the report file to be detected in the duplicate file list, the method further includes: Receive the judgment result of the client; When the result of the judgment is that the file is a duplicate file, the report file to be detected is marked as a duplicate file; When the judgment result is that the file is non-duplicate, the report file to be detected is stored in the target report file.

9. A report file similarity detection device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the report file similarity detection method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium for detecting similarity between report files, storing computer executable instructions, characterized in that: The computer executable instructions can execute the report file similarity detection method described in any one of claims 1 to 8.