Data comparison method and device, equipment, medium and program product

By receiving the file group to be compared, comparison parameters, and hidden row selection information, dynamically identifying the table header, and automatically comparing table data, this solves the problem of high manual intervention and low accuracy in Excel BOM difference comparison, and achieves efficient and accurate data difference comparison.

CN121882017APending Publication Date: 2026-04-17HUZHOU LUXSHARE PRECISION INDUSTRY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUZHOU LUXSHARE PRECISION INDUSTRY CO LTD
Filing Date
2025-12-25
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, comparing BOM differences in Excel spreadsheets requires a large amount of manual intervention, resulting in low accuracy and efficiency, and making it difficult to meet the need for accurate display of differences during the material comparison process.

Method used

By receiving the file group to be compared, comparison parameter information, and hidden row selection information, the system dynamically identifies the table header, automatically compares the table data, generates a difference comparison report, improves the accuracy and flexibility of the table header identification, and achieves data difference comparison without human intervention.

Benefits of technology

It improves the accuracy and efficiency of table data comparison, ensures the accuracy of the initial data range, and enhances the flexibility and accuracy of table data comparison.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121882017A_ABST
    Figure CN121882017A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data comparison method and device, equipment, a medium and a program product. Comprising the following steps: receiving a to-be-compared file group, comparison parameter information and hidden row selection information; reading a to-be-compared file group based on the hidden row selection information, and determining a to-be-compared table in the to-be-compared file group; performing dynamic header identification on the to-be-compared tables according to the comparison parameter information, and determining at least one target comparison table group; performing data difference comparison on each target comparison table group according to the comparison parameter information, and determining a difference comparison result; and generating a difference comparison report according to the difference comparison result. The accuracy and flexibility of header identification are improved, so that the matching of the target comparison table group determined on the basis of the accurately and flexibly determined header is more accurate, and the accuracy and the determination flexibility of the table data for data difference comparison are improved. The comparison of the data differences among the target comparison table groups is realized through an automatic mode without artificial participation according to the comparison parameter information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data comparison method, apparatus, device, medium, and program product. Background Technology

[0002] Currently, major industries are fully implementing digitalization in data processing, which has significantly improved the efficiency and quality of industry development. For design, production, and manufacturing companies, the Bill of Materials (BOM) is an essential document, and most companies choose to present the BOM in spreadsheet form (Excel).

[0003] However, because the Bill of Materials (BOM) often undergoes continuous upgrades and changes at different stages of a product's development, and different configuration requirements within the same product series may result in different BOM details, it is often necessary to compare BOMs from different stages or configurations when outputting the BOM to confirm whether the upgrades, modifications, or configuration differences for the product are correct.

[0004] Currently, comparing the differences between two Excel spreadsheets often involves opening both files and comparing them manually using various formulas or by eye. This process is time-consuming, especially for those who are not proficient in using formulas and functions. The accuracy and intuitiveness of the comparison are often unsatisfactory, making it difficult to meet the need for accurate display of differences during material comparison. Summary of the Invention

[0005] This invention provides a data comparison method, apparatus, device, medium, and program product. When comparing tabular files, it comprehensively considers the possible hidden rows and table headers with non-fixed positions in the table, improving the accuracy and flexibility of the tabular data used for data difference comparison. It achieves data difference comparison between different tables in an automated manner without human intervention, further improving the efficiency and accuracy of tabular data comparison.

[0006] In a first aspect, embodiments of the present invention provide a data comparison method, including:

[0007] Receive the file group to be compared, comparison parameter information, and hidden line selection information;

[0008] Read the file group to be compared based on the hidden line selection information, and determine the comparison table in the file group to be compared;

[0009] Based on the comparison parameter information, the header of each comparison table is dynamically identified to determine at least one target comparison table group.

[0010] Based on the comparison parameter information, the data differences of each target comparison table group are compared to determine the comparison results.

[0011] A difference comparison report is generated based on the difference comparison results.

[0012] Secondly, embodiments of the present invention also provide a data comparison device, comprising:

[0013] The information receiving module is used to receive the file group to be compared, comparison parameter information, and hidden line selection information;

[0014] The comparison table determination module is used to read the file group to be compared based on the hidden line selection information and determine the comparison table in the file group to be compared.

[0015] The target group determination module is used to dynamically identify the header of each comparison table based on the comparison parameter information, and determine at least one target comparison table group.

[0016] The comparison result determination module is used to perform data difference comparison on each target comparison table group based on the comparison parameter information and determine the difference comparison result.

[0017] The comparison report generation module is used to generate a comparison report based on the comparison results.

[0018] Thirdly, embodiments of the present invention also provide a data comparison device, comprising:

[0019] At least one processor; and a memory communicatively connected to the at least one processor;

[0020] The memory stores a computer program that can be executed by at least one processor, which enables the at least one processor to implement the data comparison method of any embodiment of the present invention.

[0021] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the data comparison method of any embodiment of the present invention.

[0022] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program, which, when executed by a processor, is used to perform the data comparison method of any embodiment of the present invention.

[0023] This invention provides a data comparison method, apparatus, device, medium, and program product. The method involves receiving a group of files to be compared, comparison parameter information, and hidden line selection information; reading the group of files to be compared based on the hidden line selection information to determine the comparison tables within the group; dynamically identifying the headers of each comparison table based on the comparison parameter information to determine at least one target comparison table group; performing data difference comparison on each target comparison table group based on the comparison parameter information to determine the difference comparison results; and generating a difference comparison report based on the difference comparison results. By employing the above technical solution, and by receiving comparison parameter information including at least hidden line selection information, it is clear whether the hidden line content in the received files to be compared requires data comparison processing, ensuring the accuracy of the initial data range required for data comparison in the files to be compared. Meanwhile, when determining the comparison tables to be compared, the first row is not used as the default header. Instead, dynamic header identification is performed for each comparison table in each file based on the comparison parameter information. This improves the accuracy and flexibility of header identification, resulting in more accurate matching of the target comparison table groups determined based on accurately and flexibly determined headers. This enhances the accuracy and flexibility of the table data used for data difference comparison. Furthermore, the comparison of data differences between target comparison table groups is automated without human intervention based on the comparison parameter information, further improving the efficiency and accuracy of table data comparison.

[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart of a data comparison method provided in Embodiment 1 of the present invention;

[0027] Figure 2 This is a flowchart of a data comparison method provided in Embodiment 2 of the present invention;

[0028] Figure 3 This is a schematic diagram of a process provided in Embodiment 2 of the present invention to update the data in the comparison table to be compared based on the matching result, and to determine the updated comparison table to be compared as the target comparison table;

[0029] Figure 4 This is a schematic diagram of the structure of a data comparison device provided in Embodiment 3 of the present invention;

[0030] Figure 5 This is a schematic diagram of the structure of a data comparison device provided in Embodiment 4 of the present invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] Example 1

[0034] Figure 1 This is a flowchart illustrating a data comparison method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations requiring efficient and accurate data comparison of tabular documents with reduced manual intervention. The method can be executed by a data comparison device, which can be implemented using software and / or hardware and can be configured within a data comparison equipment. Optionally, the data comparison device can be a laptop, desktop computer, or smart tablet, etc., and this embodiment of the present invention does not impose any limitations on this.

[0035] like Figure 1 As shown in the figure, an embodiment of the present invention provides a data comparison method, which specifically includes the following steps:

[0036] S101. Receive the file group to be compared, comparison parameter information, and hidden line selection information.

[0037] In this embodiment, the file group to be compared can be specifically understood as two tabular files input by the user, which have the same fields but may have different stored data, and need to be compared. It can be understood that the files to be compared can be Excel files, and each file to be compared can contain multiple sheets, where each sheet can serve as a comparison sheet.

[0038] Optionally, the user can import the comparison files into the comparison file group by specifying the paths of two Excel files (referred to as File A and File B in the following description) using the "Browse" button included in the pre-built data comparison interface. Optionally, the pre-built data comparison interface can also provide a "path switching" function to facilitate swapping the positions of the two files. It is understood that, generally, the first imported file, File A, can be used as the source file for comparison; that is, the data changes in File B are compared based on File A. The method for determining the source file can be adaptively specified according to different needs, and this embodiment of the invention does not limit this.

[0039] In this embodiment, the comparison parameter information can be understood as information provided by the user to clarify the data comparison requirements, and / or to clarify input, output, and selection requirements. The hidden line selection information can be understood as information provided by the user to clarify whether it is necessary to compare data in temporarily hidden lines in the file to be compared that are not displayed on the interface during the data comparison process.

[0040] Optionally, the comparison parameter information may also include user-provided comparison field information for selecting the fields to be compared, table matching method information for matching tables in different comparison files, and information for indicating the output method of the data comparison results, etc., and the embodiments of the present invention do not impose any limitations on this.

[0041] Optionally, the files to be compared and the comparison parameter information can be selected and entered through an input interface on a pre-built data comparison interface, or they can be entered through a pop-up interface. This embodiment of the invention does not limit this.

[0042] Optionally, the information in the comparison parameters that indicates the output method of the data comparison results can be specified by the user in the area for setting the output directory on the pre-built data comparison interface, allowing the user to specify the storage directory for the comparison result-related files. For example, the storage directory can default to the current program running directory, or the user can specify the storage directory for the comparison result-related files through the "Browse" button included on the data comparison interface.

[0043] Optionally, the comparison parameter information, which indicates the output method of the data comparison results, may also include information indicating the output of the original file after comparison. For example, a "Export the original file after comparison" checkbox can be set in the pre-built data comparison interface. After the user checks this option, while generating the difference comparison report based on the difference comparison results, two additional copies of the original file with highlighted marks for the difference items will also be generated to visually indicate the difference positions in the original data.

[0044] Optionally, the comparison field information may include one or more fields selected by the user for comparison in this data comparison. These fields can be understood as header fields in the files to be compared. In other words, the received comparison field information clarifies which header data in each table of the file group to be compared needs to be compared. For example, the comparison field information can be obtained by clicking on a field selection area included in a pre-built data comparison interface. Users can select which fields (columns) to use for this data comparison from a preset or custom field list by checking checkboxes. The field selection area also provides an "Add Field" input box, allowing users to add new field names and click the "Add" button to add the new field to the field selection area. Additionally, the field selection area provides options to delete custom fields and restore the default field list, allowing users to flexibly adjust the fields used for data comparison. It is understood that the adjustments and configurations made in the field selection area will be automatically saved for later use.

[0045] Specifically, when users need to compare data from different table files, they can select and input the two table files to be compared as a group of files through a pre-designed interface. At the same time, they can input comparison parameter information, including at least the hidden row selection information, at the corresponding location on the interface according to their needs.

[0046] S102. Read the file group to be compared based on the hidden line selection information, and determine the comparison table in the file group to be compared.

[0047] In this embodiment, the comparison table can be specifically understood as a sheet existing in the comparison file.

[0048] Specifically, based on the hidden row selection information, it is determined whether hidden rows in the files to be compared need to be considered in this data comparison. When reading each data row in each table of the files to be compared, it is first determined whether the data row is hidden. If the hidden row selection information determines that the data comparison needs to read hidden rows, then the data row can be read into memory for subsequent processing; if the hidden row selection information determines that the data comparison does not need to read hidden rows, then the data row can be skipped and not read into memory. After all data rows in the file group to be compared have been read, multiple comparison tables corresponding to the file group to be compared can be determined from memory.

[0049] Optionally, when reading the comparison file, the file format of the comparison file can also be detected. For example, if the comparison file is in .xls format (Excel 97-2003), the system's Excel component can be automatically and silently called to convert it into a temporary .xlsx file to solve the compatibility problem between different format versions.

[0050] Optionally, when reading the file to be compared, each sheet in the file can be read one by one. When reading each sheet, all merged cells can be automatically detected and unmerged. Then, the value of the top left cell of the original merged area is filled into the entire area to ensure that the data is complete and not lost.

[0051] S103. Based on the comparison parameter information, perform dynamic header identification on each comparison table to determine at least one target comparison table group.

[0052] In this embodiment, dynamic header recognition can be understood as scanning the sheet and determining the header based on the data information contained in different rows of the sheet, rather than directly defaulting to the header of the first row of the sheet. The target comparison table group can be understood as a set of tables consisting of two comparison tables that meet the matching requirements in the user-input comparison parameters.

[0053] Specifically, based on the comparison parameter information, the scope of each comparison table in the comparison file that can be used for data comparison, the fields that need to be compared, and the matching rules between comparison tables in different comparison files are determined. Then, based on the determined scope and fields of each comparison table, dynamic header identification is performed in each comparison table to determine the corresponding header rows. Based on the determined header rows, the data range that can be used for data comparison in each comparison table is determined. After clarifying the scope of each comparison table, the comparison tables are matched according to the matching rules between different comparison files. Two tables that meet the matching rules and whose data ranges are completely defined are identified as a target comparison table group. It can be understood that processing each comparison table can yield one or more target comparison table groups.

[0054] S104. Based on the comparison parameter information, perform data difference comparison on each target comparison table group to determine the difference comparison results.

[0055] In this embodiment, the difference comparison result can be specifically understood as the result of comparing the data differences based on the comparison parameter information for the valid data contained in each target comparison table group, which is used to characterize the differences that may exist in different comparison tables, such as data deletion, data addition, and inconsistent data occurrence frequency.

[0056] Specifically, based on the information related to the selection of comparison fields contained in the comparison parameter information, the fields that need to be compared in each target comparison table are determined. Then, for each target comparison table group, the data differences of the same fields are compared row by row to determine the possible data differences such as deletion, addition, and inconsistency, and the difference comparison results corresponding to each target comparison table group are obtained.

[0057] S105. Generate a difference comparison report based on the difference comparison results.

[0058] In this embodiment, the difference comparison report can be specifically understood as a report displayed to the user, showing the differences existing in the group of files to be compared, or in different target comparison tables within the group of files to be compared. Optionally, the difference comparison report can be presented in the form of a new Excel file, where each sheet can correspond to the difference comparison results of a target comparison table group, or all difference comparison results of the same group of files to be compared can be presented through a single sheet. This embodiment of the invention does not impose any limitations on this.

[0059] Specifically, based on the difference comparison results of different target comparison table groups, the data rows with differences in each target comparison table group, the fields corresponding to the differences, and the columns of the fields corresponding to the differences in the comparison table can be identified. Based on this, the above information is integrated into the pre-built result report template to generate the corresponding difference comparison report.

[0060] For example, when generating a difference comparison report based on the difference comparison results, a new Excel file can be created as the difference comparison report, and each sheet in the difference comparison report can correspond to the difference comparison results of a set of target comparison tables. The sheet name can be generated according to the comparison mode (e.g., comparison result_1 or the source table name of the target comparison table group). A description line can be added to the top of each sheet to indicate the source table name and the comparison fields actually used for data comparison. The main body of the difference comparison report can be a table listing only all the difference items in the difference comparison results. Optionally, to facilitate reading and identification and improve readability, the background color can be automatically filled according to the type of difference item when generating the difference comparison report. For example, gray can be used to represent rows of "deleted" data differences; green to represent rows of "added" data differences; and yellow to represent rows of "inconsistent counts" data differences, etc. This embodiment of the invention does not limit this.

[0061] The technical solution of this embodiment involves receiving a group of files to be compared, comparison parameter information, and hidden line selection information; reading the group of files to be compared based on the hidden line selection information to determine the comparison tables within the group; dynamically identifying the headers of each comparison table based on the comparison parameter information to determine at least one target comparison table group; performing data difference comparison on each target comparison table group based on the comparison parameter information to determine the difference comparison results; and generating a difference comparison report based on the difference comparison results. By adopting the above technical solution, and by receiving comparison parameter information including at least hidden line selection information, it is clear whether the hidden line content contained in the received comparison files needs data comparison processing, thus ensuring the accuracy of the initial data range required for data comparison in the comparison files. Meanwhile, when determining the comparison tables to be compared, the first row is not used as the default header. Instead, dynamic header identification is performed for each comparison table in each file based on the comparison parameter information. This improves the accuracy and flexibility of header identification, resulting in more accurate matching of the target comparison table groups determined based on accurately and flexibly determined headers. This enhances the accuracy and flexibility of the table data used for data difference comparison. Furthermore, the comparison of data differences between target comparison table groups is automated without human intervention based on the comparison parameter information, further improving the efficiency and accuracy of table data comparison.

[0062] Example 2

[0063] Figure 2 This is a flowchart of a data comparison method provided in Embodiment 2 of the present invention. The technical solution of the present invention is further optimized based on the above-mentioned optional technical solutions. First, based on the hidden row selection information in the comparison parameter information, the information in the file group to be compared is read, and the comparison table to be compared is determined from the file group to be compared to determine the data range required for this data comparison. Then, based on the comparison field information in the comparison parameter information, fuzzy matching is performed on each data row within a preset range of each comparison table to select the most appropriate data row as the header row in the preset range of each comparison table, so that each comparison table can complete the update of the comparison data range based on the header row, and obtain the updated target comparison table that can be used for this data comparison. Based on the table matching method information in the comparison parameter information, each target comparison table is matched to more accurately determine the target of this data comparison. Simultaneously, after updating each comparison table, identical data rows in the resulting target comparison tables are grouped to identify duplicate rows and their counts, which serve as the basis for in-depth comparison in subsequent data difference comparisons. Then, each target comparison table group is merged based on comparison field information, and the merged table is analyzed row by row. Differences are compared at two levels: those directly reflected in the shallow data rows and those reflected in the deeper differences through varying counts of duplicates. A difference comparison report is generated based on the comparison results of each target comparison table group. This improves the accuracy and flexibility of the tabular data used for data difference comparison, and further enhances the efficiency and accuracy of tabular data comparison by automating the comparison of data differences between target comparison table groups without human intervention, based on comparison parameters.

[0064] like Figure 2 As shown, the data comparison method provided in Embodiment 2 of the present invention specifically includes the following steps:

[0065] S201. Receive the file group to be compared, comparison parameter information, and hidden line selection information.

[0066] S202. Read the file group to be compared based on the hidden line selection information, and determine the comparison table in the file group to be compared.

[0067] S203. For each comparison table, perform fuzzy matching between the comparison field information in the comparison parameter information and each data row in the comparison table within the preset range, update the data in the comparison table according to the matching results, and determine the updated comparison table as the target comparison table.

[0068] In this embodiment, the preset range can be understood as a range of data rows in the comparison table that can be used to determine the table header, set in advance according to the actual situation. For example, the preset range can be the first 50 rows of the comparison table, or it can be set arbitrarily according to the actual situation. This embodiment of the invention does not impose any restrictions on this. Fuzzy matching can be understood as a matching algorithm that can still identify similar items or similarity when there are some differences between the data in the data row used for comparison and the comparison field information. The matching result can be understood as the result information used to determine the degree of matching between the data row and the comparison field information based on the similarity between each data row and the comparison field information.

[0069] Specifically, for each comparison table, data rows within a preset range are used as the basis for determining the header row. Each data row is then fuzzily matched with the comparison field information in the comparison parameter information to obtain the matching results for each data row. These matching results are then analyzed to determine the header row in the comparison table with the highest matching degree to the comparison parameter information. This header row is then used as the basis to determine the data available for comparison in the comparison table. Based on the determined data, the comparison table is updated so that the data contained in the updated target comparison table are all the data required for comparison in this case.

[0070] Optional, Figure 3 This is a schematic diagram of a process provided in Embodiment 2 of the present invention, which updates the data in the comparison table to be compared based on the matching results and determines the updated comparison table to be compared as the target comparison table. Figure 3 As shown, the specific steps include the following:

[0071] S2031. Based on the matching results, determine the data row with the most matching fields and the one that is sorted first as the header row.

[0072] Specifically, to ensure that the data to be compared in the comparison table is completely and effectively preserved, when determining the header row based on the matching results, the number of matching fields in the matching results corresponding to each data row can be compared first, and the data row with the most matching fields can be used as the header row. If there are multiple data rows with the most matching fields, to retain as much valid data as possible in the comparison table, the row with the most matching fields that appears first in the comparison table can be used as the header row.

[0073] S2032. Perform data cleaning on the header row and all data below it, and update the cleaned data to the comparison table to determine the target comparison table.

[0074] Specifically, all data in the header row and below in the comparison table are converted into a structured data table for subsequent data comparison processing. At the same time, the data in each cell of the data table can be cleaned to standardize it. All the processed data is then updated in the comparison table, and the updated comparison table is determined as the target comparison table.

[0075] For example, data cleaning operations may include: removing invisible characters and extra spaces; converting null values, None, nan, etc. into the string "N / A" to ensure that null values ​​can also participate in the comparison; converting pure numeric strings into floating-point numbers and retaining 4 decimal places to unify the numerical format, etc., and the embodiments of the present invention do not limit this.

[0076] Optionally, to support the generation of the original file copy for highlighting the differences as described above, when reading and updating the comparison table, the row number of each row of data in the original comparison file can be recorded for subsequent highlighting processing in the original comparison file.

[0077] Optionally, to facilitate the identification of duplicate entries within the same comparison table during the data comparison process, the target comparison table can be traversed and filtered after the comparison table is updated to identify the rows containing duplicate entries. This can be achieved in the following way:

[0078] After updating the data in the comparison table based on the matching results and determining the updated comparison table as the target comparison table, the updated target comparison table is grouped according to the comparison field information. Multiple data rows with identical content are determined as a data group; the number of data rows in the data group is determined as the number of duplicate items corresponding to the data content in the data group.

[0079] For example, generating the target comparison table is equivalent to traversing each data row in the table. During this process, it can be determined whether there are data rows with completely identical content. If so, multiple data rows with identical content can be grouped into a data group, and an internal counter starting from 1 can be added to it. Adding a data row with the same content increments the internal counter by one. Finally, after traversing all data rows in the target comparison table, the counter number corresponding to each data group, that is, the number of data rows in each data group, can be determined as the number of duplicate items corresponding to the data content in that data group. For example, if "Material A" appears three times, the internal counter in its corresponding data group will count to 3, indicating that there are three data rows containing "Material A", and these three data rows will be marked as 1, 2, and 3 respectively.

[0080] S204. Based on the table matching method information in the comparison parameter information, determine the matching relationship between each target comparison table, and determine two target comparison tables that match each other as a target comparison table group.

[0081] In this embodiment, the table matching method information can be specifically understood as the information provided by the user in the comparison parameter information, which clarifies how to match the comparison tables in different comparison files during the data comparison process. For example, the table matching method information may include sequential comparison and name-based matching: when using sequential comparison, the first sheet of File A can be matched with the first sheet of File B as a target comparison table group, the second sheet of File A can be matched with the second sheet of File B as a target comparison table group, and so on; when using name-based matching, sheets with the same name in the comparison files can be found as target comparison table groups, and sheets that exist only in one of the comparison files can be marked as either newly added or missing on one side.

[0082] Specifically, based on the table matching method information in the comparison parameter information, the user's preferred matching method for each table in the comparison file group is determined. Based on the determined matching method, the comparison tables in each comparison file are sorted, or the table names in each comparison file are matched, thus obtaining the matching relationship between each target comparison table. In this way, two target comparison tables that successfully match each other can be identified as a target comparison table group.

[0083] S205. For each target comparison table group, based on the comparison field information in the comparison parameter information, merge the target comparison table groups using the comparison field information as the connection key to determine the table to be compared and merged.

[0084] Specifically, for each target comparison table group, the comparison field information in the comparison parameter information can be determined as the valid comparison field in the target comparison table group that needs to be matched. Then, each valid comparison field can be used as a join key to perform an outer join on the data in the comparison tables belonging to different files to be compared, and the merged table obtained after the join is determined as the merged table to be compared.

[0085] Optionally, the generation and processing of the table to be compared and merged can be achieved through the merge function of the pandas library, or any other feasible table merging function can be used. This embodiment of the invention does not limit this.

[0086] S206. Perform row-by-row analysis on the merged tables to be compared, and determine the comparison results based on the data source of each row in the merged tables.

[0087] Specifically, each data row in the table to be compared and merged is analyzed row by row. Data from different data sources that belong to the same data row and have the same comparison fields are compared. The comparison results corresponding to each comparison field are combined to determine the difference comparison result corresponding to that data row.

[0088] Optionally, the difference comparison results can be determined based on the data source of each merged table's data rows, which may include the following situations:

[0089] 1) If the data source of the merged table data row is the first target comparison table belonging to the first file to be compared in the target comparison table group, then the difference comparison result of the merged table data row is determined to be deleted.

[0090] 2) If the data source of the merged table data row is the second target comparison table belonging to the second file to be compared in the target comparison table group, then the difference comparison result of the merged table data row is determined as an addition.

[0091] 3) If the data source of the merged table data row includes two target comparison tables in the target comparison table group, then compare the number of duplicate items in the two target comparison tables corresponding to the data content in the merged table data row. If the number of duplicate items is different, then the difference comparison result of the merged table data row is determined to be inconsistent.

[0092] For example, the first file to be compared can be understood as File A in the above example, and the second file to be compared can be understood as File B in the above example. The expansion of determining the difference comparison results based on the data source of each merged table's data rows can be explained as follows:

[0093] 1) If the data in the merged table data row comes only from the first target comparison table of the first file to be compared in the target comparison table group, it can be considered that the data in the merged table data row only exists in File A, which means that the data is missing from File B relative to File A. In this case, the difference comparison result of the merged table data row can be determined as "deleted".

[0094] 2) If the data source of the merged table data row is the second target comparison table belonging to the second file to be compared in the target comparison table group, it can be considered that the data in the merged table data row only exists in File B, which means that the data is extra in File B relative to File A. In this case, the difference comparison result of the merged table data row can be determined as "new".

[0095] 3) If the data source of the merged table data row includes both target comparison tables in the target comparison table group, then the data corresponding to the merged table data row can be considered to require in-depth comparison. In this case, based on the data content corresponding to the merged table data row, the number of duplicate items corresponding to the data content in the two target comparison tables can be determined respectively. If the number of duplicate items for the data content in the two target comparison tables is the same, then the data content corresponding to the merged table data row can be considered to have no difference in the two target comparison tables, and there is no need to record it in the subsequent difference comparison report. If the number of duplicate items for the data content in the two target comparison tables is different, then the data content corresponding to the merged table data row can be considered to appear at different times in the two target comparison tables. In this case, the difference comparison result of the merged table data row can be determined as "inconsistent frequency".

[0096] Optionally, when determining the difference comparison results based on the data source of each merged table data row, to facilitate the use of the highlighting function, the row number and columns involved in the original file of each merged table data row with differences can be recorded, and the records can be stored in a difference location dictionary. This allows for accurate highlighting of the original file by directly extracting the corresponding information from the difference location dictionary if highlighting is required later.

[0097] S207. Generate a difference comparison report based on the difference comparison results.

[0098] Optionally, during the execution of the above data comparison method, a log display area can be set on the pre-built data comparison interface to display all key steps, difference comparisons (such as missing fields), error messages, and the final result path in real time. Simultaneously, a progress bar area can be set on the pre-built data comparison interface. This progress bar will continuously scroll during the execution of the data comparison method and stop upon completion. All buttons on the data comparison interface will be disabled during the execution of the data comparison method and will return to normal after completion.

[0099] Optionally, after the data comparison method is completed, the system will automatically clean up the temporary files generated during the operation to keep the user's system clean.

[0100] The technical solution of this embodiment first reads the information in the comparison file group based on the hidden row selection information in the comparison parameter information, and determines the comparison table of the data range required for this data comparison from the comparison file group. Then, based on the comparison field information in the comparison parameter information, fuzzy matching is performed on each data row within a preset range in each comparison table to select the most appropriate data row as the header row. This allows each comparison table to be updated based on the comparison data range of the header row, resulting in an updated target comparison table that can be used for this data comparison. Based on the table matching method information in the comparison parameter information, each target comparison table is matched to more accurately determine the target of this data comparison. At the same time, after updating each comparison table, the identical data rows in each target comparison table are grouped to determine the data rows with duplicates and the number of duplicates in the duplicate data rows. This information is used as the basis for in-depth comparison in subsequent data difference comparison. Then, each target comparison table group is merged based on the comparison field information, and the merged table is analyzed row by row. Differences are directly reflected in the shallow data rows, and differences are reflected in the deeper levels through different numbers of duplicate items. This completes the comparison of data differences between the two target comparison tables corresponding to the merged table. A difference comparison report is generated based on the difference comparison results of each target comparison table group. This improves the accuracy and flexibility of the tabular data used for data difference comparison, and further enhances the efficiency and accuracy of tabular data comparison by automating the comparison of data differences between target comparison table groups without human intervention based on comparison parameter information.

[0101] Example 3

[0102] Figure 4 This is a schematic diagram of the structure of a data comparison device provided in Embodiment 3 of the present invention, as shown below. Figure 4 As shown, the data comparison device includes an information receiving module 31, a comparison table determination module 32, a target group determination module 33, a comparison result determination module 34, and a comparison report generation module 35.

[0103] The system includes: an information receiving module 31 for receiving a group of files to be compared, comparison parameter information, and hidden line selection information; a comparison table determination module 32 for reading the group of files to be compared based on the hidden line selection information and determining the comparison tables in the group; a target group determination module 33 for dynamically identifying the headers of each comparison table based on the comparison parameter information and determining at least one target comparison table group; a comparison result determination module 34 for performing data difference comparison on each target comparison table group based on the comparison parameter information and determining the difference comparison result; and a comparison report generation module 35 for generating a difference comparison report based on the difference comparison result.

[0104] The technical solution of this invention, by receiving comparison parameter information including at least hidden row selection information, clarifies whether the hidden row content in the received comparison file needs data comparison processing, ensuring the accuracy of the initial data range required for data comparison in the comparison file. Simultaneously, when determining the comparison table to be compared, the first row is not defaulted to as the table header; instead, dynamic header identification is performed for each comparison table in each comparison file based on the comparison parameter information. This improves the accuracy and flexibility of header identification, thereby making the target comparison table group determined based on the accurately and flexibly determined header more accurate in matching, and improving the accuracy and flexibility of the table data used for data difference comparison. Furthermore, by using comparison parameter information to automate the comparison of data differences between target comparison table groups without human intervention, the efficiency and accuracy of table data comparison are further improved.

[0105] Optionally, the target group determination module 33 is specifically used for:

[0106] For each comparison table, the comparison field information in the comparison parameter information is fuzzy matched with each data row in the preset range of the comparison table. The data in the comparison table is updated according to the matching results, and the updated comparison table is determined as the target comparison table.

[0107] Based on the table matching method information in the comparison parameter information, the matching relationship between each target comparison table is determined, and two target comparison tables that match each other are identified as a target comparison table group.

[0108] Optionally, the data in the comparison table to be compared is updated based on the matching results, and the updated comparison table to be compared is determined as the target comparison table, including:

[0109] Based on the matching results, the data row with the most matching fields and the highest sort order is determined as the header row;

[0110] Perform data cleaning on the header row and all data below it, and update the cleaned data in the comparison table to determine the target comparison table.

[0111] Optionally, the data comparison device further includes: a duplicate item determination module, specifically used for:

[0112] After updating the data in the comparison table based on the matching results and determining the updated comparison table as the target comparison table, the target comparison table is grouped according to the comparison field information. Multiple data rows with identical content are determined as a data group; the number of data rows in the data group is determined as the number of duplicate items corresponding to the data content in the data group.

[0113] Optionally, the comparison result determination module 34 is specifically used for:

[0114] For each target comparison table group, the target comparison table groups are merged based on the comparison field information in the comparison parameter information, using the comparison field information as the connection key, to determine the table to be compared and merged.

[0115] The merged tables to be compared are analyzed row by row, and the differences are determined based on the data source of each row in the merged tables.

[0116] Optionally, the difference comparison results can be determined based on the data source of each merged table's data rows, including:

[0117] If the data source of the merged table data row is the first target comparison table in the target comparison table group that belongs to the first file to be compared, then the difference comparison result of the merged table data row is determined to be deleted;

[0118] If the data source of the merged table data row is the second target comparison table belonging to the second file to be compared in the target comparison table group, then the difference comparison result of the merged table data row is determined as an addition;

[0119] If the data source of the merged table data row includes two target comparison tables in the target comparison table group, then the number of duplicate items corresponding to the data content in the merged table data row is compared between the two target comparison tables. If the number of duplicate items is different, then the difference comparison result of the merged table data row is determined to be inconsistent.

[0120] The data comparison device provided in this embodiment of the invention can execute the data comparison method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0121] Example 4

[0122] Figure 5 This is a schematic diagram of a data comparison device provided in Embodiment 4 of the present invention. The data comparison device 40 can represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The data comparison device 40 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0123] like Figure 5As shown, the data comparison device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded from storage unit 48 into the RAM 43. The RAM 43 can also store various programs and data required for the operation of the data comparison device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0124] Multiple components in the data comparison device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a disk, optical disk, etc.; and a communication unit 49, such as a network card, modem, wireless transceiver, etc. The communication unit 49 allows the data comparison device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0125] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as data comparison methods.

[0126] In some embodiments, the data comparison method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on the data comparison device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the data comparison method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the data comparison method by any other suitable means (e.g., by means of firmware).

[0127] Optionally, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the data comparison method provided in any embodiment of the present invention.

[0128] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0129] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0130] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0131] To provide user interaction, the systems and techniques described herein can be implemented on a data comparison device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the data comparison device. Other types of devices can also be used to provide user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0132] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0133] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0134] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data comparison method, characterized in that, include: Receive the file group to be compared, comparison parameter information, and hidden line selection information; Based on the hidden line selection information, the file group to be compared is read, and the comparison table in the file group to be compared is determined; Based on the comparison parameter information, dynamic header identification is performed on each of the comparison tables to be compared to determine at least one target comparison table group; Based on the comparison parameter information, data differences are compared for each of the target comparison table groups to determine the comparison results; A difference comparison report is generated based on the difference comparison results.

2. The data comparison method according to claim 1, characterized in that, The step of dynamically identifying the header of each of the comparison tables based on the comparison parameter information to determine at least one target comparison table group includes: For each of the comparison tables, the comparison field information in the comparison parameter information is fuzzily matched with each data row in the comparison table within a preset range. The data in the comparison table is updated according to the matching results, and the updated comparison table is determined as the target comparison table. Based on the table matching method information in the comparison parameter information, the matching relationship between each target comparison table is determined, and two target comparison tables that match each other are identified as a target comparison table group.

3. The data comparison method according to claim 2, characterized in that, The step of updating the data in the comparison table based on the matching results and determining the updated comparison table as the target comparison table includes: Based on the matching results, the data row with the most matching fields and the highest sort order is determined as the header row; Data cleaning is performed on the header row and all data below it, and the cleaned data is updated in the comparison table to determine the target comparison table.

4. The data comparison method according to claim 2, characterized in that, After updating the data in the comparison table based on the matching results and determining the updated comparison table as the target comparison table, the method further includes: The target comparison table is grouped according to the comparison field information, and multiple data rows with identical content are identified as a data group. The number of data rows in the data group is determined as the number of duplicate items corresponding to the data content in the data group.

5. The data comparison method according to claim 1, characterized in that, The step of performing data difference comparison on each of the target comparison table groups based on the comparison parameter information and determining the difference comparison result includes: For each target comparison table group, the target comparison table groups are merged according to the comparison field information in the comparison parameter information, using the comparison field information as the connection key, to determine the comparison and merging table to be compared; The merged tables to be compared are analyzed row by row, and the comparison results are determined based on the data source of each row in the merged tables.

6. The data comparison method according to claim 5, characterized in that, The determination of the difference comparison results based on the data source of each merged table data row includes: If the data source of the merged table data row is the first target comparison table belonging to the first file to be compared in the target comparison table group, then the difference comparison result of the merged table data row is determined to be deleted; If the data source of the merged table data row is the second target comparison table belonging to the second file to be compared in the target comparison table group, then the difference comparison result of the merged table data row is determined as an addition; If the data source of the merged table data row includes two target comparison tables in the target comparison table group, then the number of duplicate items corresponding to the data content in the merged table data row is compared between the two target comparison tables. If the number of duplicate items is different, then the difference comparison result of the merged table data row is determined to be inconsistent.

7. A data comparison device, characterized in that, include: The information receiving module is used to receive the file group to be compared, comparison parameter information, and hidden line selection information; The comparison table determination module is used to read the file group to be compared based on the hidden line selection information and determine the comparison table in the file group to be compared. The target group determination module is used to dynamically identify the header of each of the comparison tables based on the comparison parameter information, and determine at least one target comparison table group. The comparison result determination module is used to perform data difference comparison on each of the target comparison table groups based on the comparison parameter information, and determine the difference comparison result; The comparison report generation module is used to generate a comparison report based on the comparison results.

8. A data comparison device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data comparison method of any one of claims 1-6.

9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the data comparison method as claimed in any one of claims 1-6.

10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the data comparison method as claimed in any one of claims 1-6.