Data comparison method and device, electronic equipment and storage medium
Through comparative exception processing, the problem of comparison exceptions caused by different data volumes of data sets is solved, the data set comparison is realized, the range of comparable data sets is expanded, and the user experience is improved.
Patent Information
- Application Number
- CN202510341890.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
When performing data comparison, the prior art requires that the data volume of the two sets of data sets is the same, but in actual scenarios, due to data loss or abnormality, the data volume of the data set is different, and the comparison cannot be performed.
By obtaining the first data set and the second data set selected by the user, we judge whether there are data comparison exceptions, and perform comparison exception processing, including data filling or data removal, to ensure that the data volume of the data set is the same before comparing.
It realizes that data comparison can be performed even if data is lost in the data set, expanding the range of comparable data sets and improving the user experience.
Smart Images

Figure CN120217012A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology. Specifically, this application relates to a data comparison method, apparatus, electronic device, and storage medium. Background Art
[0002] Currently, in various current technical fields, data comparison is a key analysis link in various applications. Through data comparison, pattern differences between data can be discovered, and the pattern differences can provide important bases for subsequent technical decisions.
[0003] In the solutions of the prior art, when performing data comparison, the terminals for data comparison often require that the data volumes of the two data sets to be compared are the same. However, in actual operation scenarios, in many cases, some data in the data sets may be lost or there are obvious abnormal data, resulting in different data volumes of the two data sets to be compared. When the above situation occurs, the existing data comparison terminals cannot perform the data comparison of these two data sets and prompt the user that the data volumes of the two data sets need to be adjusted to be the same before the data comparison can be performed.
[0004] Therefore, a solution is needed to enable the data comparison terminal to compare a data set with other data sets even if there is a data set with lost data. Summary of the Invention
[0005] The purpose of this application aims to solve at least one of the above technical defects. The technical solutions provided by the embodiments of this application are as follows: In a first aspect, an embodiment of this application provides a data comparison method, including: Obtaining a first data set and a second data set selected by a user; According to a preset comparison anomaly condition, determining whether there is a data comparison anomaly between the first data set and the second data set, and when it is determined that there is a data comparison anomaly, performing a comparison anomaly process on the first data set and the second data set to obtain a first comparison data set and a second comparison data set; Performing a data comparison on the first comparison data set and the second comparison data set to obtain a data comparison result and display it.
[0006] In a second aspect, an embodiment of this application provides a data comparison apparatus, including: A data set acquisition module, configured to obtain a first data set and a second data set selected by a user; A data set anomaly judgment module, configured to judge whether there is a data comparison anomaly between a first data set and a second data set according to a preset comparison anomaly condition, and when it is judged that there is a data comparison anomaly, perform comparison anomaly processing on the first data set and the second data set to obtain a first comparison data set and a second comparison data set; A data set display module, configured to perform data comparison on the first comparison data set and the second comparison data set to obtain a data comparison result and display it.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory; The processor executes the computer program to implement the method provided in the embodiment of the first aspect or any optional embodiment of the first aspect.
[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method provided in the embodiment of the first aspect or any optional embodiment of the first aspect is implemented.
[0009] The beneficial effects brought by the technical solution provided by the embodiment of the present application are: In the solution provided by the present application, since the data volumes of the data included in the two data sets of the user are different and cannot be compared, it will cause a comparison anomaly in the system. Therefore, before data comparison, the data set with missing data can be filled with data, or the corresponding part of the data can be removed from the complete data set, so that the data volumes of the data sets to be compared are the same, so that even if the data sets selected by the user have data loss, two or more data sets can be compared, expanding the range of data sets that can be compared and improving the user experience. Description of the Drawings
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments of the present application.
[0011] Figure 1 It is a schematic flowchart of a data comparison method provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of a data comparison method in an example provided by an embodiment of the present application; Figure 3 It is a structural block diagram of a data comparison device provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0012] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0013] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology, etc. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include a wireless connection or a wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0014] To make the purpose, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0015] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps, etc. in different embodiments, they will not be described repeatedly.
[0016] Figure 1 A flowchart of a data comparison method is provided for an embodiment of the present application. The execution subject of this method can be a data comparison terminal (such as a computer, a mobile phone, etc.). As Figure 1 shown, this method may include: Step S101, obtaining a first data set and a second data set selected by the user.
[0017] In the embodiments of the present application, the user can be the user of the data comparison terminal. The data of the first data set and the second data set can be graphic data from a graph, or tabular data from a table, etc. The embodiments of the present application do not make limitations here.
[0018] Specifically, in the embodiments of the present application, when a user needs to compare data sets, the original data to be compared can be directly input into the terminal. Since the data sources of the original data to be compared (for example, one piece of data comes from a graph and the other comes from a table) or the formats are different, when the terminal receives the original data, it can first preprocess the original data. The preprocessing operations can include sorting, cleaning, formatting, or standardization, etc., to uniformly convert each original data into the same format and distinguish them according to the data sources to form corresponding data sets respectively.
[0019] Step S102: According to the preset comparison anomaly conditions, determine whether there is a data comparison anomaly between the first data set and the second data set. When it is determined that there is a data comparison anomaly, perform a comparison anomaly process on the first data set and the second data set to obtain a first comparison data set and a second comparison data set.
[0020] In the embodiments of the present application, the data comparison anomaly can be that there is data missing in the data set, or there are data spikes (i.e., data with obvious outliers) in the data set, etc. The embodiments of the present application do not make limitations here.
[0021] Specifically, after obtaining the two data sets, it is necessary to analyze the two data sets to determine whether they meet the preset comparison anomaly conditions. If they do not meet, the comparison process of the two data sets can be directly started; if they meet, further processing of the two data sets is required before starting the comparison process.
[0022] It should be noted that in the embodiments of the present application, different anomaly handling methods will be adopted for different comparison anomalies. For example, if the data comparison anomaly is caused by data missing in the data set, then the data set with missing data can be filled or some data in the data set without missing data can be removed to ensure that the two data sets contain the same amount of data; if the data comparison anomaly is caused by the presence of data spikes in the data set, then the corresponding outliers can be removed or the corresponding normal values can be calculated and used as replacements, etc. The embodiments of the present application do not make limitations here.
[0023] Step S103: Compare the first comparison data set and the second comparison data set to obtain a data comparison result and display it.
[0024] Specifically, after performing the corresponding comparison anomaly process on the first data set and the second data set, the comparison process is started. After the comparison is completed, the data comparison result of the first comparison data set and the second comparison data set obtained after the anomaly process is displayed.
[0025] In the solution provided by this application, since the amounts of data contained in the two data sets of the user are different and cannot be compared, it will cause a comparison anomaly in the system. Therefore, before data comparison, the data set with missing data can be filled with data, or the corresponding part of the data can be removed from the complete data set, so that the amounts of data contained in the data sets to be compared are the same. Thus, even if there is data loss in the data sets selected by the user, two or more data sets can be compared, expanding the range of data sets that can be compared and improving the user experience.
[0026] Based on the above embodiments, as an alternative embodiment, if the data comparison anomaly indicates that there is missing data in the first data set relative to the second data set; Perform data comparison anomaly processing on the first data set and the second data set to obtain a first comparison data set and a second comparison data set, which specifically includes: Judge whether the first data set meets the preset data filling condition. If it meets, fill the missing data of the first data set to obtain the first comparison data set, and use the second data set as the second comparison data set; If it does not meet, remove the missing data of the first data set from the second data set to obtain the second comparison data set, and use the first data set as the first comparison data set.
[0027] In the embodiments of this application, the preset data filling condition is used to judge whether the first data set needs to be filled. Since in actual operation, sometimes the missing data in the first data set has little impact on the comparison result, and data filling sometimes consumes more computing resources, so this condition can be used to judge whether data filling is required first.
[0028] Specifically, before data comparison, judge whether the first data set needs to be filled. If the judgment result indicates that the first data set needs to be filled, then the missing part of the data in the first data set can be filled first, and then the comparison process can be started. After the comparison is completed, the comparison result of the filled first data set and the second data set is displayed; if the judgment result indicates that the first data set does not need to be filled (that is, the missing data has little or no impact on the comparison result), then the data corresponding to the missing part of the first data set in the second data set is removed to ensure that there is no missing data in the first data set relative to the second data set after the data is removed.
[0029] It should be noted that in the embodiments of the present application, "data missing" in the first data set may refer to a data set with a part of data missing relative to the second data set. For example, for two sales data sets, one is the sales data from January to October, and the other is the sales data from January to August and October. By analysis, it can be seen that the sales data from January to August and October is missing the sales data for September compared to the sales data from January to October. At this time, the sales data from January to August and October can be determined as the first data set, and the sales data from January to October can be determined as the second data set.
[0030] Based on the above various embodiments, as an optional embodiment, the missing data in the first data set is filled to obtain a first comparison data set, which specifically includes: Determine the dimension information corresponding to the data included in the first data set, and based on the preset correspondence between the dimension information and the filling method, determine the filling method for the first data set; Call the filling method to fill the missing data in the first data set to obtain a first comparison data set.
[0031] In the embodiments of the present application, the filling method may be the interpolation method, or the fitting method or other filling methods. The user can set it according to their actual needs, and the embodiments of the present application do not limit it here. The data type can characterize the change trend of the data, such as the data changes smoothly or the data changes complexly, etc.
[0032] Specifically, in the embodiments of the present application, since data filling needs to ensure that the filled data is as accurate as possible, and different filling methods will directly affect the accuracy of data filling. For example, when some period of sales data is missing, the missing period of sales data can be predicted and filled according to the change situation between the sales times of other periods. For the missing meteorological data in a certain area, it is necessary to predict and fill according to the influence degree of other adjacent areas on the missing area. Therefore, the most suitable filling method needs to be selected for different data. For this, the embodiments of the present application can pre-find the most suitable filling method corresponding to each dimension information through a large number of experiments in advance, and establish a correspondence between each data type and its corresponding most suitable filling method and store it, so that the terminal can directly select the most suitable filling method according to the correspondence after determining the data type. Specifically, when it is determined that data filling is required for the first data set, first, the dimension information corresponding to the data included in the first data set can be determined, and then according to the preset correspondence between the dimension information and the filling method, the corresponding filling method is determined for the first data set, and then the filling method is called to fill the missing part of the data in the first data set. After the filling is completed, a first comparison data set can be obtained.
[0033] Based on the above various embodiments, as an alternative embodiment, the dimension information corresponding to the data included in the first data set is spatial type dimension information; Call a filling method to fill the missing data in the first data set to obtain a first comparison data set, which specifically includes: Select non-missing data from the first data set that has a spatial impact on the missing data based on the dimension information; For each non-missing data, determine the weight of the non-missing data based on the magnitude of the spatial impact of the non-missing data on the missing data; Perform weighted summation on each non-missing data based on the weights of each non-missing data to obtain a filling value for the missing data, and fill the filling value into the missing data in the first data set to obtain a first comparison data set.
[0034] In the embodiments of the present application, the spatial type dimension information may indicate that the data has a certain structure in space or geometry, such as geographical location, material distribution location, etc. The magnitude of the spatial impact may indicate the spatial correlation relationship between the actual spatial position corresponding to the non-missing data and the actual spatial position corresponding to the missing data. For example, if the actual spatial position corresponding to the non-missing data is closer to the actual spatial position corresponding to the missing data, then the magnitude of their spatial impact is larger.
[0035] Specifically, for the first data set with dimension information being spatial type dimension information (such as meteorological data in different regions), the inverse distance weighted interpolation method, natural neighbor interpolation method, or Kriging interpolation method can be used for data filling. Specifically, several non-missing data with spatial geometric correlation with the missing part of the data can be selected according to the spatial geometric relationship of each data in the first data set, and then the weight corresponding to each non-missing data is set according to the magnitude of the spatial impact of the determined non-missing data on the missing data (such as the distance in the actual space, etc.), and then weighted summation is performed on each non-missing data according to each weight, and the result of the weighted summation is used as the data value that the missing data needs to be filled.
[0036] It should be noted that the inverse distance weighted interpolation method sets corresponding weights for each non-missing data according to the positions of the missing data in the first dataset and other non-missing data in the first dataset. The magnitude of the weight is inversely proportional to the positional difference between the non-missing data and the missing data in the first dataset. This method only considers the influence relationship caused by the distance of each data and is applicable to application scenarios where the spatial geometric relationship of each data in the first dataset is relatively simple. When determining the spatial geometric correlation relationship between non-missing data and missing data, the natural domain interpolation method constructs a Voronoi diagram to divide the corresponding regions for each data, and then determines the influence of non-missing data on missing data based on the division results, thereby determining the magnitude of the weight of each non-missing data. It is more applicable to scenarios where there are interactions only between some of the data in the first dataset. The Kriging interpolation method establishes a corresponding semi-variogram based on the spatial relationships between all non-missing data and the missing data (this function reflects the spatial correlation between non-missing data and missing data and represents the degree of change of two data in space), and then determines the magnitude of the weight of each non-missing data according to the spatial distance between each non-missing data and the missing data and the semi-variogram. It is more applicable to scenarios where there are interactions between all the data in the first dataset.
[0037] Based on the above various embodiments, as an alternative embodiment, the dimension information corresponding to the data included in the first dataset is not the spatial type dimension information; Call a filling method to fill the missing data in the first dataset to obtain a first comparison dataset, which specifically includes: Determine whether there is a linear relationship between the non-missing data; If there is a linear relationship, obtain the change rate of each non-missing data, determine the filling value of the missing data based on the change rate and each non-missing data, and fill the filling value into the missing data; If there is no linear relationship, determine a non-linear function about the first dataset based on the non-missing data, obtain the position of the missing data in the first dataset, determine the function value corresponding to the position in the non-linear function, and use the function value as the filling value to fill the missing data to obtain the first comparison dataset.
[0038] In the embodiments of the present application, the existence of a linear relationship means that the data in the first dataset continuously changes at a stable change rate, and the non-existence of a linear relationship means that the data in the first dataset does not continuously change at a stable change rate.
[0039] Specifically, for the first data set whose dimension information is not spatial type dimension information (such as sales data at different times), linear interpolation method or spline interpolation method can be used for data filling; among them, if it is found that there is a linear relationship between the data in the first data set, the linear interpolation method can be used, that is, by calculating the change rate between the non-missing data in the first data set, and then predicting and filling the missing data in the first data set according to the change rate. The characteristic of the linear interpolation method is that the calculation is simple, but for the case of non-linear change of data, the prediction calculation accuracy is relatively low.
[0040] For the first data set where there is no linear relationship between the data, two filling methods, spline interpolation method and triangular interpolation method, can be used. The filling principle of these two methods is to first represent the data in the first data set in the form of data points on the coordinate axis, then connect the data points in turn through a smooth curve, and then predict the position corresponding to the missing data on the coordinate axis (that is, the position of the missing data in the first data set) according to the drawn smooth curve (i.e., non-linear function). The predicted value is the function value of the non-linear function corresponding to the smooth curve at this position, and this function value is used as the filling value of the missing data for filling. The characteristics of these two interpolation methods are that the calculation accuracy is relatively high, but the calculation amount is relatively large.
[0041] It should be noted that the various interpolation methods described above are not the only interpolation methods that can be implemented in the embodiments of the present application. According to actual needs, other interpolation methods can also be used for data filling, and the method of data filling through other interpolation methods should also be regarded as the protection scope of the present invention.
[0042] Optionally, in the embodiments of the present application, data filling can also be performed only by fitting, that is, fitting operations are performed on the existing data in the first data set to obtain the corresponding fitting function, and then the missing data is simulated and calculated and filled according to the fitting function.
[0043] Based on the above various embodiments, as an optional embodiment, determining whether the first data set meets the preset data filling condition specifically includes: Performing a fitting operation on the second data set to obtain the first fitting function corresponding to the second data set; Removing the data corresponding to the missing data part of the first data set from the second data set to obtain a third data set; Performing a fitting operation on the third data set to obtain the second fitting function corresponding to the third data set; Determining the first fitting error between the first fitting function and the second data set, and determining the second fitting error between the second fitting function and the third data set; Determine the difference between the first fitting error and the second fitting error. If the difference is greater than a preset threshold, it is determined that the first data set meets the preset data filling condition; if the difference is not greater than the preset threshold, it is determined that the first data set does not meet the preset data filling condition.
[0044] In an embodiment of the present application, the fitting operation may be an operation of representing the relationship between the data in the data set by an approximate function. The fitting function is the approximate function obtained in the fitting operation that can represent the relationship between the data. The fitting error may characterize the actual error between the fitting function and the relationship between the data in the data set.
[0045] Specifically, when determining whether the first data set needs to be filled with data, the main basis is whether the influence of the corresponding part of the data in the second data set on the comparison result is large. If the influence is large, the first data set needs to be filled; if the influence is small, there is no need to fill. Specifically, first perform a fitting operation on the existing data in the second data set to obtain the corresponding first fitting function, and then remove the data corresponding to the missing part in the second data set (that is, the data whose influence on the comparison result needs to be judged) to obtain the third data set; then calculate the fitting errors of the second data set and the third data set respectively with the corresponding fitting functions. Since the data in the second data set is complete, the first fitting error corresponding to this data set can reflect the influence of each data on the relationship between the data in the second data set. In the third data set, since some data has been removed, the second fitting error corresponding to this data set can only reflect the influence of the data that has not been removed on the relationship between the data in the second data set. It can be understood that if a data has a large influence on the relationship between the data in the data set, then this influence will be significantly reflected in the fitting error. Therefore, when the difference between the first fitting error and the second fitting error is large, it means that the data removed from the second data set has a large influence on the relationship between the data in this data set, indicating that this part of the data is an indispensable data in the second data set. Correspondingly, the missing part of the data in the first data set needs to be filled; otherwise, it means that the data removed from the second data set has a small influence on the relationship between the data in this data set, and there is no need to fill. Specifically, the size of the difference can be judged by setting a preset threshold, and the preset threshold can be determined in advance through a large number of calculation experiments.
[0046] Optionally, in an embodiment of the present application, the two data sets input by the user may also have missing data from each other (it can also be said that the data volumes of the data included in the two data sets are different). For example, data set A contains sales data from January to October, and data set B contains sales data from February to November. For these two data sets, the following operations can be taken: 1. First, take their common part, "sales data from February to October", as the first data set, and then perform a fitting operation to obtain the first fitting function of data set A and the first fitting function of data set B respectively; 2. Exclude the sales data of January in data set A to obtain the third data set corresponding to data set A. Exclude the sales data of October in data set B to obtain the third data set corresponding to data set B. Then, perform a fitting operation on the third data set of data set A and the third data set of data set B respectively to obtain the second fitting function corresponding to data set A and the second fitting function corresponding to data set B; 3. Determine the difference between the first fitting function and the second fitting function of data set A and the difference between the first fitting function and the second fitting function of data set B; 4. Compare the difference corresponding to data set A and the difference corresponding to data set B with a preset threshold respectively. If the difference corresponding to data set A or the difference corresponding to data set B is greater than the preset threshold, it is determined that the "sales data of January" or "sales data of November" needs to be filled; if the difference corresponding to data set A or the difference corresponding to data set B is not greater than the preset threshold, it is determined that the "sales data of January" or "sales data of November" does not need to be filled; 5. Fill the data that needs to be filled according to the judgment result, and do not fill the data that does not need to be filled and exclude the corresponding data in the other data set.
[0047] Optionally, in the embodiment of the present application, if the terminal itself cannot or does not expect to carry the above complex calculation process for determining whether to perform data filling, the terminal can also be set to the state of "default data filling", so that as long as the terminal detects the first data set with missing data, it fills the missing data in the first data set.
[0048] Based on the above various embodiments, as an optional embodiment, determine the first fitting error between the first fitting function and the second data set, and determine the second fitting error between the second fitting function and the third data set, specifically including: Determine each first fitting value corresponding to each data included in the second data set based on the first fitting function, and determine each second fitting value corresponding to each data included in the third data set based on the second fitting function; Determine each first data value difference between each data included in the second data set and the corresponding first fitting value, and determine each second data value difference between each data included in the third data set and the corresponding second fitting value; Take the mean value of each first data value difference as the first fitting error, and take the mean value of each second data value difference as the second fitting error.
[0049] In an embodiment of the present application, the first fitting value may be the data value corresponding to a position in the first dataset that conforms to the first fitting function; the second fitting value may be the data value corresponding to a position in the second dataset that conforms to the second fitting function.
[0050] Specifically, in an embodiment of the present application, when determining the fitting error, it can be obtained by comparing and calculating each fitting value obtained by the fitting function with the corresponding true value in the corresponding data value. Specifically, the following formula can be referred to: MSE = ) 2 Where MSE represents the fitting error, n represents the number of data in the data value, i represents the position of each data in the data value, represents the true data value of the i-th data in the dataset, represents the fitting value of the i-th data in the dataset, ([[]] ) 2 can be used to represent the difference between the true data value and the fitting value.
[0051] Through the above calculation method, the first fitting error corresponding to the first dataset and the second fitting error corresponding to the second dataset can be calculated.
[0052] It should be noted that the fitting error calculation method provided in the embodiments of the present application is not the only one. Users can also select other methods for calculation according to actual needs (for example, directly using the absolute value of the difference between the fitting value and the true data value as the difference between the first data value or the second data value, etc.). The embodiments of the present application do not make limitations here.
[0053] Based on the above various embodiments, as an optional embodiment, the first comparison dataset and the second comparison dataset are compared to obtain a data comparison result and display it, specifically including: Obtain the display requirements of the user for the data comparison result; Based on the display requirements, determine each statistical type to be displayed and the display form for each dataset; For each statistical type, respectively obtain each data to be statistically analyzed that needs to be statistically analyzed from the first comparison dataset and the second comparison dataset, and determine the calculation method corresponding to the statistical type. Process each data to be statistically analyzed based on the calculation method to obtain the statistical result corresponding to the statistical type; Based on the display form, display the first comparison dataset and the second comparison dataset, and display the statistical results of each statistical type.
[0054] In the embodiments of the present application, the display requirements can be set in advance by the user or set on the spot by the user. The embodiments of the present application do not limit this. The display form can be a graphical form such as a line chart, a bar chart, a pie chart, or other table forms, etc. The embodiments of the present application do not limit this. The statistical type can be the statistical data that the user needs to calculate, such as the average value, the year-on-year value, the month-on-month value, etc. The embodiments of the present application do not limit this. The statistical result is the specific value calculated corresponding to the statistical type.
[0055] Specifically, after completing the filling of the first data set to obtain the first comparison data set or completing the data exclusion of the second data set to obtain the second comparison data set, the comparison process of the first comparison data set and the second comparison data set is started. First, it is necessary to clarify the user's display requirements for the comparison results. The display requirements include the display method of the data in the first comparison data set and the second comparison data set and the statistical type that needs to be calculated for the data in the first comparison data set and the second comparison data set. In this process, first determine the to-be-statistical data to be statistically analyzed according to the statistical type. For example, if the statistical type is the average value, all the data in the first comparison data set are used as the to-be-statistical data for average value calculation, or only a part of the data is used as the to-be-statistical data for average value calculation according to the requirements; if the statistical type is year-on-year, the data in the first comparison data set that needs to be compared year-on-year and the corresponding data in the second comparison data set that needs to be compared year-on-year are jointly used as the to-be-statistical data for calculation. In actual operation, the embodiments of the present application can simultaneously receive multiple statistical type calculation requests initiated by the user, or receive a custom statistical type calculation method set by the user himself / herself, and calculate the user's custom statistical type. After the calculation of all statistical types is completed, the statistical results corresponding to all statistical types are displayed, and at the same time, the data in the first comparison data set and the data in the second comparison data set are displayed in the display form required by the user.
[0056] Based on the above various embodiments, as an optional embodiment, obtaining the user's display requirements for the data comparison results specifically includes: Obtain the user's identification information, and match the display template corresponding to the identification information from the preset database based on the identification information; Analyze the display template to determine the display requirements; Among them, the display template is generated by the user in advance through a preset template editing interface and stored in the preset database.
[0057] In the embodiments of the present application, the identification information is used to represent the user's identity information, which may be the user's username, ID (Identity document), etc. The embodiments of the present application do not make any limitations here. The display template may be the template framework of the comparison result display interface, which can be preset by the user. Each user can set multiple display templates, and each display template is stored corresponding to the user's identification information in the preset database. The preset database can be used to record the user identification information and the display templates set by the user. The template editing interface may be an interface for the user to directly operate and edit the display template.
[0058] Specifically, when obtaining the user's display requirements for the comparison results, it is possible to first detect in the preset database whether the user has a preset display template. If the display template corresponding to the user cannot be detected in the database, a prompt of "Please set the display template" can be sent to the user. When the user confirms "Set", the preset template editing interface can be displayed to the user for editing; if the display template corresponding to the user can be detected in the database, all the templates corresponding to the user are displayed to the user in the form of thumbnails for selection. At the same time, a control button of "New display template" can also be provided so that the user can immediately create the display template they need. After the user selects one of the display templates, analyze the display template to determine the user's display requirements. The display requirements include the display form of the data in the dataset and the statistical types that need to be calculated.
[0059] Optionally, the embodiments of the present application also provide the user with the function of "predicting and displaying comparison results". This function can display the future predicted comparison results after the original data in the first comparison dataset and the second comparison dataset to the user. For example, when only the "sales data from January to October" is displayed in the first comparison dataset and the second comparison dataset, through this function, the "sales data from November to December" can be further displayed. When implementing the above function, fitting operations can be performed on the first comparison dataset and the second comparison dataset, and future data predictions and displays are respectively performed according to the fitting functions obtained from the fitting operations.
[0060] Next, in combination with Figure 2 , the process of the data comparison method provided by the embodiments of the present application will be introduced, as shown in Figure 2As shown, firstly, the data to be compared input by the user is obtained. After receiving the data to be compared, each data to be compared is preprocessed (including sorting, cleaning, standardization, formatting and other operations) to generate two data sets to be compared. Next, the two data sets to be compared are analyzed to determine a first data set with missing data and a second data set with complete data. Then, it is judged whether the missing data in the first data set needs to be filled: if necessary, a suitable filling method is selected to fill the missing data in the first data set; if not, the corresponding part of the data in the second data set is removed, and then the comparison result of the two data sets is displayed to the user according to the user's display requirements.
[0061] Figure 3 A structural block diagram of a data comparison device provided in an embodiment of the present application, such as Figure 3 As shown, the Internet technology data comparison device 300 may include: a data set acquisition module 301, a data set anomaly judgment module 302 and a data set display module 303, wherein: The data set acquisition module 301 is used to acquire a first data set and a second data set selected by a user; The data set abnormality judgment module 302 is used to judge whether there is a data comparison abnormality between the first data set and the second data set according to a preset comparison abnormality condition, and when it is judged that there is a data comparison abnormality, perform comparison abnormality processing on the first data set and the second data set to obtain a first comparison data set and a second comparison data set; The data set display module 303 is used to perform data comparison on the first comparison data set and the second comparison data set to obtain and display the data comparison result.
[0062] The solution provided by the present application cannot be compared due to the different amounts of data contained in two data sets of the user, which will cause a comparison anomaly in the system. Therefore, before the data comparison, the data set with missing data can be filled with data, or the corresponding part of the data can be removed from the complete data set, so that the data sets to be compared contain the same amount of data, so that even if the data set selected by the user has data missing, two or more data sets can be compared, thereby expanding the range of data sets that can be compared and improving the user experience.
[0063] Based on the above embodiments, as an optional embodiment, if the data comparison anomaly indicates that the first data set is missing data relative to the second data set; The data set anomaly judgment module is used to: Performing a comparison and abnormality processing on the first data set and the second data set to obtain a first comparison data set and a second comparison data set specifically includes: Determine whether the first data set meets the preset data filling condition. If it meets, fill the missing data in the first data set to obtain the first comparison data set, and use the second data set as the second comparison data set; If it does not meet, remove the missing data of the first data set from the second data set to obtain the second comparison data set, and use the first data set as the first comparison data set.
[0064] Based on the above embodiments, as an alternative embodiment, the data set anomaly judgment module is further configured to: Determine the dimension information corresponding to the data included in the first data set, and determine the filling method for the first data set based on the corresponding relationship between the preset dimension information and the filling method; Call the filling method to fill the missing data in the first data set to obtain the first comparison data set.
[0065] Based on the above embodiments, as an alternative embodiment, the dimension information corresponding to the data included in the first data set is spatial type dimension information; The data set anomaly judgment module can also be used to: Select non-missing data from the first data set that has a spatial impact on the missing data based on the dimension information; For each non-missing data, determine the weight of the non-missing data based on the spatial impact size of the non-missing data on the missing data; Perform weighted summation on the non-missing data based on the weights of the non-missing data to obtain the filling value of the missing data, and fill the filling value into the missing data in the first data set to obtain the first comparison data set.
[0066] Based on the above embodiments, as an alternative embodiment, the dimension information corresponding to the data included in the first data set is not spatial type dimension information; The data set anomaly judgment module can also be used to: Judge whether there is a linear relationship between the non-missing data; If there is a linear relationship, obtain the change rate of each non-missing data, and determine the filling value of the missing data based on the change rate and each non-missing data, and fill the filling value into the missing data; If there is no linear relationship, determine a non-linear function about the first data set based on the non-missing data, obtain the position of the missing data in the first data set, determine the function value corresponding to the position in the non-linear function, and use the function value as the filling value to fill the missing data in the first data set to obtain the first comparison data set.
[0067] Based on the above embodiments, as an alternative embodiment, the data set anomaly judgment module is specifically configured to: Perform a fitting operation on the second data set to obtain a first fitting function corresponding to the second data set; Remove the data corresponding to the missing data part of the first data set from the second data set to obtain a third data set; Perform a fitting operation on the third data set to obtain a second fitting function corresponding to the third data set; Determine a first fitting error between the first fitting function and the second data set, and determine a second fitting error between the second fitting function and the third data set; Determine the difference between the first fitting error and the second fitting error. If the difference is greater than a preset threshold, it is determined that the first data set meets the preset data filling condition; if the difference is not greater than the preset threshold, it is determined that the first data set does not meet the preset data filling condition.
[0068] Based on the above various embodiments, as an optional embodiment, the data set display module is specifically used for: Obtain the user's display requirements for the data comparison result; Based on the display requirements, determine each statistical type to be displayed and the display form for each data set; For each statistical type, respectively obtain each data to be statistically analyzed from the first comparison data set and the second comparison data set, and determine the calculation method corresponding to the statistical type. Process each data to be statistically analyzed based on the calculation method to obtain the statistical result corresponding to the statistical type; Based on the display form, display the first comparison data set and the second comparison data set, and display the statistical results of each statistical type.
[0069] Based on the above various embodiments, as an optional embodiment, the data set display module is further used for: Obtain the user's identification information, and match the display template corresponding to the identification information from the preset database based on the identification information; Analyze the display template to determine the display requirements; Wherein, the display template is generated by the user in advance through a preset template editing interface and stored in the preset database.
[0070] Next, refer to Figure 4 , which shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present application (for example, a terminal device or a server that executes the method shown in Figure 1 ). The electronic device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc.Figure 4 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0071] The electronic device includes: a memory and a processor. The memory is used to store programs for executing the methods described in the above respective method embodiments; the processor is configured to execute the programs stored in the memory. Herein, the processor may be referred to as the processing device 401 described below, and the memory may include at least one of the read-only memory (ROM) 402, random access memory (RAM) 403, and storage device 408 described below, as specifically shown below: As Figure 4 shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 402 or the programs loaded from the storage device 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, ROM 402, and RAM 403 are connected to each other via a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0072] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0073] Specifically, according to the embodiments of the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network via the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above functions defined in the methods of the embodiments of the present application are executed.
[0074] It should be noted that the above-mentioned computer-readable storage medium in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0075] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of the communication network include a local area network ("LAN"), a wide area network ("WAN"), the Internet (for example, the Internet), and a peer-to-peer network (for example, an ad hoc peer-to-peer network), as well as any currently known or future-developed network.
[0076] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device.
[0077] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to: Obtain the first data set and the second data set selected by the user; according to the preset comparison anomaly condition, determine whether there is a data comparison anomaly between the first data set and the second data set, and when it is determined that there is a data comparison anomaly, perform comparison anomaly processing on the first data set and the second data set to obtain the first comparison data set and the second comparison data set; compare the first comparison data set and the second comparison data set to obtain a data comparison result and display it.
[0078] Computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0079] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0080] The modules or units involved in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module or unit does not, in some cases, constitute a limitation on the unit itself. For example, the first constraint acquisition module can also be described as "the module for acquiring the first constraint".
[0081] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0082] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or Flash Memory), an optical fiber, a portable Compact Disc Read Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0083] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order limitation for the execution of these steps, and they can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts of the accompanying drawings can include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a portion of other steps or sub-steps or stages of other steps.
[0084] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A data comparison method, characterized in that: include: Acquire a first data set and a second data set selected by a user; According to a preset abnormal comparison condition, determining whether there is a data comparison abnormality between the first data set and the second data set, and when it is determined that there is a data comparison abnormality, performing abnormal comparison processing on the first data set and the second data set to obtain a first comparison data set and a second comparison data set; The first comparison data set and the second comparison data set are compared to obtain a data comparison result and displayed.
2. The method according to claim 1, characterized in that If the data comparison anomaly indicates that the first data set is missing data relative to the second data set; The performing comparison and abnormality processing on the first data set and the second data set to obtain a first comparison data set and a second comparison data set includes: Determine whether the first data set meets a preset data filling condition, and if so, fill in the missing data of the first data set to obtain the first comparison data set, and use the second data set as the second comparison data set; If not, the missing data of the first data set are removed from the second data set to obtain the second comparison data set, and the first data set is used as the first comparison data set.
3. The method according to claim 2, characterized in that The filling of missing data of the first data set to obtain the first comparison data set includes: Determine dimension information corresponding to the data included in the first data set, and determine a filling method for the first data set based on a preset correspondence between the dimension information and the filling method; The filling method is called to fill in the missing data of the first data set to obtain the first comparison data set.
4. The method according to claim 3, characterized in that The dimension information corresponding to the data contained in the first data set is space type dimension information; The calling the filling method to fill the missing data of the first data set to obtain the first comparison data set includes: Selecting non-missing data having a spatial impact on the missing data from the first data set based on the dimensional information; For each non-missing data, determining the weight of the non-missing data based on the spatial impact of the non-missing data on the missing data; A weighted sum is performed on each non-missing data based on the weight of each non-missing data to obtain a filling value of the missing data, and the filling value is filled into the missing data in the first data set to obtain the first comparison data set.
5. The method according to claim 3, characterized in that The dimension information corresponding to the data contained in the first data set is not spatial type dimension information; The calling the filling method to fill the missing data of the first data set to obtain the first comparison data set includes: Determine whether there is a linear relationship between the non-missing data; If there is a linear relationship, obtaining a change rate of each non-missing data, determining a filling value of the missing data based on the change rate and each non-missing data, and filling the missing data with the filling value; If there is no linear relationship, a nonlinear function about the first data set is determined based on each non-missing data, the position of the missing data in the first data set is obtained, the function value corresponding to the position in the nonlinear function is determined, and the function value is used as the filling value to fill the missing data in the first data set to obtain the first comparison data set.
6. The method according to claim 2, characterized in that The determining whether the first data set meets a preset data filling condition includes: Performing a fitting operation on the second data set to obtain a first fitting function corresponding to the second data set; Eliminate data corresponding to the missing data portion of the first data set from the second data set to obtain a third data set; Performing a fitting operation on the third data set to obtain a second fitting function corresponding to the third data set; determining a first fitting error between the first fitting function and the second data set, and determining a second fitting error between the second fitting function and the third data set; Determine a difference between the first fitting error and the second fitting error, and if the difference is greater than a preset threshold, determine that the first data set meets the preset data filling condition; if the difference is not greater than the preset threshold, determine that the first data set does not meet the preset data filling condition.
7. The method according to claim 1, characterized in that The step of performing data comparison on the first comparison data set and the second comparison data set to obtain and display a data comparison result includes: Obtaining the user's demand for displaying the data comparison result; Determine the statistical types to be displayed and the display format of each data set based on the display requirements; For each statistical type, respectively obtaining each statistical data to be counted from the first comparison data set and the second comparison data set, and determining a calculation method corresponding to the statistical type, and processing each statistical data to be counted based on the calculation method to obtain a statistical result corresponding to the statistical type; The first comparison data set and the second comparison data set are displayed based on the display form, and statistical results of each statistical type are displayed.
8. The method according to claim 7, characterized in that The obtaining of the user's demand for displaying the data comparison result includes: Acquire identification information of the user, and match a display template corresponding to the identification information from a preset database based on the identification information; Analyze the display template to determine the display requirements; The display template is generated in advance by the user through a preset template editing interface and stored in the preset database.
9. A data comparison device, characterized in that: include: A data set acquisition module, used to acquire a first data set and a second data set selected by a user; a data set anomaly judgment module, configured to judge whether there is a data comparison anomaly between the first data set and the second data set according to a preset comparison anomaly condition, and when it is judged that there is a data comparison anomaly, perform comparison anomaly processing on the first data set and the second data set to obtain a first comparison data set and a second comparison data set; The data set display module is used to perform data comparison on the first comparison data set and the second comparison data set to obtain and display the data comparison result.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.