Method and device for analyzing influence factors of data problem, equipment and medium

By performing statistical analysis of the distribution characteristics and confidence level screening on the problem dataset, the influencing factors of the data problem are automatically analyzed, which solves the problems of low efficiency and inaccuracy in existing technologies and achieves efficient and accurate discovery of influencing factors.

CN121412291APending Publication Date: 2026-01-27吕国晖
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511593544.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

In existing technologies, the methods for analyzing factors affecting data problems are inefficient and inaccurate, heavily relying on the prior knowledge of domain experts, making it difficult to efficiently and accurately identify the factors affecting problematic data.

Method used

By acquiring the problem data set, analyzing the distribution characteristics of statistical data fields, filtering fields that meet the candidate criteria, and using confidence levels to perform secondary filtering of candidate influencing factors, the final influencing factors of the data problem are automatically discovered.

Benefits of technology

It enables efficient and accurate analysis of the influencing factors of data problems without relying on prior knowledge. It can automatically discover single and combined influencing factors, has a scientific confidence assessment system, is applicable to various industries and data structures, and lowers the application threshold.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121412291A_ABST
    Figure CN121412291A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for analyzing influence factors of data problems, equipment and a medium. Relates to the field of data analysis, and the method comprises the steps: obtaining a problem data set, the problem data set comprises at least one problem data set, carrying out the statistics of the attributes of data fields in a to-be-analyzed problem data set, and obtaining the distribution characteristics of the data fields corresponding to the to-be-analyzed problem data set, and screening fields meeting candidate conditions from the distribution characteristics of the data fields to obtain a candidate influence factor set of the to-be-analyzed problem data set, and screening candidate influence factors meeting confidence conditions from the candidate influence factor set to obtain final influence factors of the to-be-analyzed problem data set. According to the technical scheme provided by the embodiment of the invention, the influence factors of the problem data can be relatively efficiently and accurately analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis, and more specifically, to a method, apparatus, device, and medium for analyzing the influencing factors of data problems. Background Technology

[0002] Data distortion is inevitable at any stage of the process of data collection, processing, transmission, storage, display, synchronization, manipulation, circulation, or migration. This results in data problems where the original authenticity and accuracy are lost, making it impossible to accurately reflect objective facts.

[0003] Only by understanding the factors influencing problematic data can we fundamentally prevent its recurrence. Therefore, analyzing these factors has significant practical and production value. While data consistency checks can identify problematic data and correct them through overwriting, understanding why these data problems occur is a very labor-intensive and time-consuming task. Furthermore, current methods for analyzing the underlying causes of data problems heavily rely on the prior knowledge of domain experts. This leads to traditional methods for analyzing the influencing factors of data problems often being inefficient and inaccurate.

[0004] Therefore, how to provide a relatively efficient and accurate method for analyzing the influencing factors of problem data has become a problem that needs to be solved. Summary of the Invention

[0005] The purpose of one embodiment of this application is to provide a method, apparatus, device, and medium for analyzing the influencing factors of data problems. The technical solution of the embodiment of this application can analyze the influencing factors of problem data relatively efficiently and accurately.

[0006] In a first aspect, embodiments of this application provide a method for analyzing the influencing factors of data problems, comprising: acquiring a problem data set, the problem data set including at least one problem dataset, wherein the data in each problem dataset corresponds to the same type of data problem, different problem datasets correspond to different data problems, and each problem dataset includes at least one problem data; statistically analyzing the attributes of data fields in the problem dataset to be analyzed to obtain the distribution characteristics of the data fields corresponding to the problem dataset to be analyzed, wherein the problem dataset to be analyzed is any one of the at least one problem datasets; filtering fields that meet candidate conditions from the distribution characteristics of the data fields to obtain a set of candidate influencing factors for the problem dataset to be analyzed, the set of candidate influencing factors including at least one candidate influencing factor; and filtering candidate influencing factors that meet confidence level conditions from the set of candidate influencing factors to obtain the final influencing factors for the problem dataset to be analyzed.

[0007] This application embodiment first screens candidate influencing factors, and then uses confidence levels to perform a second screening of the candidate influencing factors to obtain the final influencing factors. Through the analysis of the data itself and the second screening, this application embodiment can find influencing factors relatively objectively and accurately without relying on personal experience. Compared with the prior art, it achieves efficient and accurate analysis of influencing factors of problem data.

[0008] In one implementation, the attributes of the data fields in the dataset to be analyzed include: time type, numeric type, or text type.

[0009] In one implementation, the candidate condition is met when the corresponding value of the field covers a proportion of data in the dataset to be analyzed that is higher than the candidate condition threshold or reaches 100%.

[0010] In one implementation, the corresponding value of the field includes one value of a field, multiple values ​​of a field, or one or more values ​​of multiple fields.

[0011] In one implementation, the confidence level condition includes one or more of the following conditions: internal confidence level is higher than an internal confidence level threshold or reaches 100%, random confidence level is higher than a random confidence level threshold or reaches 100%, and global confidence level is higher than a global confidence level threshold or reaches 100%. The internal confidence level is 1 - p_internal, where p_internal represents the proportion of candidate influencing factors in other problem datasets besides the problem dataset to be analyzed in the problem dataset set. The random confidence level is 1 - p_random, where p_random represents the proportion of candidate influencing factors in random datasets among other datasets besides the problem dataset to be analyzed in the overall dataset set. The global confidence level is 1 - p_global, where p_global represents the proportion of candidate influencing factors in other datasets besides the problem dataset to be analyzed in the overall dataset set.

[0012] In one implementation, the problem data set includes the problem data itself. If there are no fields that meet the candidate conditions or no candidate influencing factors that meet the confidence level conditions, the method further includes: after further reorganizing or refining the problem data classification, performing attribute statistics on the data fields, filtering the candidate influencing factor set, and filtering the final influencing factors again; and / or, performing statistics on the data fields of the associated data of the problem dataset to be analyzed and the attributes of the data fields in the problem dataset to be analyzed together, updating the statistical results to the distribution characteristics of the data fields corresponding to the problem dataset to be analyzed, and then performing the filtering of the candidate influencing factor set and filtering the final influencing factors again.

[0013] In one implementation, the step of statistically analyzing the attributes of data fields in the dataset to be analyzed to obtain the distribution characteristics of the data fields corresponding to the dataset to be analyzed includes: statistically analyzing the data fields of the associated data of the dataset to be analyzed and the attributes of the data fields in the dataset to be analyzed together to obtain the distribution characteristics of the data fields corresponding to the dataset to be analyzed.

[0014] Secondly, one embodiment of this application provides an apparatus for analyzing the influencing factors of data problems, comprising: an acquisition unit, configured to acquire a problem data set, the problem data set including at least one problem dataset, wherein the data in each problem dataset corresponds to the same type of data problem, different problem datasets correspond to different data problems, and each problem dataset includes at least one problem data; a statistics unit, configured to perform statistics on the attributes of data fields in the problem dataset to be analyzed, to obtain the distribution characteristics of the data fields corresponding to the problem dataset to be analyzed, wherein the problem dataset to be analyzed is any one of the at least one problem datasets; a first filtering unit, configured to filter fields that meet candidate conditions from the distribution characteristics of the data fields, to obtain a set of candidate influencing factors for the problem dataset to be analyzed, the set of candidate influencing factors including at least one candidate influencing factor; and a second filtering unit, configured to filter candidate influencing factors that meet confidence level conditions from the set of candidate influencing factors to obtain the final influencing factors for the problem dataset to be analyzed.

[0015] Thirdly, one embodiment of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the methods described in the first aspect and any embodiment of the first aspect.

[0016] Fourthly, one embodiment of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can implement the method as described in the first aspect and any embodiment of the first aspect.

[0017] Fifthly, one embodiment of this application provides a computer program product, the computer program product including a computer program, wherein the computer program, when executed by a processor, can implement the method as described in the first aspect and any embodiment of the first aspect. Attached Figure Description

[0018] To more clearly illustrate the technical solution of one embodiment of this application, the accompanying drawings used in one embodiment of this application will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart of a method for analyzing the influencing factors of data problems provided in one embodiment of this application; Figure 2 A schematic diagram of an apparatus for analyzing factors influencing data problems, provided as an embodiment of this application; Figure 3 This is a schematic diagram of an electronic device provided for one embodiment of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0021] Before describing the embodiments of this application, some concepts involved in the embodiments of this application will be explained as follows.

[0022] Data problems refer to data distortion or anomalies that occur during data generation or processing (such as synchronization, processing, transfer, or migration). These issues manifest as inconsistencies between the data content and the expected or source data, thus affecting the accuracy, completeness, and reliability of the data. For example, during user information migration, the age field for some users was incorrectly assigned null or abnormal values ​​(such as negative numbers).

[0023] Influencing factors: These refer to the root causes or conditions leading to data problems, typically involving technical, process, environmental, or human factors. Analyzing these influencing factors allows us to pinpoint the root cause of the problem and develop targeted improvement measures. For example: During data synchronization, operator error resulted in some order prices not being updated correctly (process / human factor). Here, the influencing factor refers to the operator, not the order prices (data problem).

[0024] Problematic data refers to data itself that has issues. This application's embodiments primarily involve analyzing problematic data to identify the factors contributing to these issues.

[0025] For example, an e-commerce platform might suddenly see a large number of abnormal order records with an "order amount" of 0. While operations and maintenance personnel can quickly fix these erroneous data, the more critical question is: are these abnormal orders caused by users in a specific region, within a specific time period, through a specific version of the app client, or by a newly launched marketing campaign? This application aims to analyze this type of problematic data to identify the influencing factors that produced it.

[0026] Related data: refers to data that has a certain relationship with another data. For example, if a data is a user's sales data, its related data may include the user's basic information, the user's attention information, etc.

[0027] Overall data: This includes problematic data and data that does not exist.

[0028] Next, we will describe in detail the current status of the influencing factors of the problem being analyzed.

[0029] Analyzing the influencing factors of problematic data has significant practical and production value. While data consistency comparisons can identify which data content is problematic and correct it through overwriting, understanding why these data problems occur is a very labor-intensive and time-consuming challenge. Furthermore, current technologies for analyzing the underlying causes of data problems heavily rely on the prior knowledge of domain experts. This leads to traditional methods for analyzing the influencing factors of data problems often being inefficient and inaccurate.

[0030] One approach is rule-based problem analysis, which involves pre-defining a series of data quality rules (e.g., a field cannot be empty, a value must be within a specified range), and the system automatically validates the data and alerts for problems. However, this approach lacks flexibility and adaptability: the rule base requires manual maintenance, and when business logic or data patterns change, the rules must be manually updated, resulting in high maintenance costs. Furthermore, it cannot detect anomalies outside the rules, and while it primarily identifies problems, it cannot pinpoint their causes (i.e., influencing factors). This approach also suffers from inefficiency and inaccuracy.

[0031] Another approach is to use pre-labeled problem data and their causes from historical data to train a classification or prediction model to determine the possible causes of newly emerging data problems. However, this approach relies on high-quality labeled data: the success of supervised learning depends on having a large number of accurate "(problem data, cause)" labeled samples. Furthermore, the model typically can only identify problem types that have appeared in the training data. For new business scenarios or novel data problems, the model's predictive performance will be significantly reduced, requiring the collection of new data and retraining. Moreover, influencing factors are uncertain and diverse; the same data problem may be caused by different influencing factors in different batches of runs. This approach also suffers from inefficiency and inaccuracy.

[0032] In summary, existing methods for analyzing factors affecting data issues generally suffer from low efficiency and inaccuracy.

[0033] To address the problems in the prior art, this application provides a method for analyzing the influencing factors of data problems, which can achieve efficient and accurate analysis of the influencing factors of problematic data.

[0034] Specifically, this application embodiment first screens candidate influencing factors, and then uses confidence levels to perform a second screening of the candidate influencing factors to obtain the final influencing factors. Through the analysis of the data itself and the second screening, this application embodiment can find influencing factors relatively objectively and accurately without relying on personal experience. Compared with the prior art, it achieves efficient and accurate analysis of influencing factors of problem data.

[0035] Furthermore, the method in this application embodiment does not require supervision or pre-determining of rules, and can automatically discover factors affecting data problems.

[0036] The following description, for ease of understanding and illustration, and not as a limitation, uses relevant figures to illustrate the method for analyzing the influencing factors of data problems in this application.

[0037] The following is in conjunction with the appendix Figure 1 This application provides an exemplary embodiment of a method for analyzing the influencing factors of data problems. For example... Figure 1 The method shown can be performed by a device for analyzing the influencing factors of a data problem, and the method includes: 110, retrieve the problem data set.

[0038] The problem data set includes at least one problem dataset, wherein the data in each problem dataset corresponds to the same type of data problem, and different problem datasets correspond to different data problems. Each problem dataset includes at least one problem data entry. It should be understood that in this embodiment, one data entry can represent a row of data in the dataset.

[0039] It should be understood that the problem data set here can be obtained from a third party, or it can be a set of problem data with data issues obtained after filtering through overall data analysis.

[0040] For example, suppose the user processes a dataset D, and after comparison or remediation, problematic datasets, i.e., problematic datasets, are identified as E. Based on the characteristics of E (e.g., null values, zero values, or all values ​​in a certain field being incorrect), the user classifies these problematic datasets into multiple problematic datasets E1, E2, E3, etc. For one of the categories Ei, the user needs to determine what factors caused this type of error, i.e., the key influencing factors of this type of problem, where i is an integer representing the dataset label.

[0041] It should be understood that the types and forms of data problems in this application embodiment are not limited, and can be determined according to the actual data problems in the problem data. In addition, the types of problem data, the methods of classifying problem data, and the methods of discovering problem data from the overall data set are not limited in this application embodiment. Any method can be used to obtain the problem data set. The focus of this application is on how to analyze the problem data set after obtaining the problem data set to obtain the influencing factors corresponding to each problem dataset.

[0042] 120. Statistical analysis is performed on the attributes of the data fields in the dataset to be analyzed to obtain the distribution characteristics of the data fields corresponding to the dataset to be analyzed.

[0043] The problem dataset to be analyzed is any one of the at least one problem datasets.

[0044] It should be understood that the dataset to be analyzed can be any dataset Ei. The embodiments of this application can perform corresponding analysis on each dataset, or analyze part of the dataset to determine its influencing factors. The embodiments of this application are not limited to this.

[0045] Specifically, in this application embodiment, the data in the problem dataset can be split according to fields, and then the fields can be statistically analyzed.

[0046] For example, as one implementation, the attributes of the data fields in the dataset to be analyzed include: time type, numeric type, or text type.

[0047] For example, in this embodiment of the application, different statistical strategies can be adopted for fields of different data types on the Ei dataset: For time types, the time type can be date / time (Datetime), and multi-level drill-down statistics can be performed, such as analyzing by "year, month, day, hour, minute, second" layer by layer to form multi-level distribution statistics results, so as to find out whether the problem is concentrated at a certain precise point in time or time period.

[0048] For numeric data, data can be binned and distributed across ranges to identify whether problems are concentrated in a particular price, age, or measurement range.

[0049] For text types, which can be Category / Text, the frequency of different text values ​​can be counted, and the Top-N (e.g., Top-10) values ​​and their proportions can be found.

[0050] It should be understood that step 120 aims to list the distribution characteristics of all relevant fields in the data as comprehensively as possible. In step 120, for interval-based distribution characteristics, a standard uniform boundary can be used for segmented statistics initially, followed by further refinement based on coverage in the next step. For example, this step only analyzes the distribution of year / month / day for date fields, but the influencing factors may range from one day to another, requiring adjustment through influence range analysis in the next step.

[0051] 130. Select fields that meet the candidate conditions from the distribution characteristics of the data fields to obtain the set of candidate influencing factors for the dataset of the problem to be analyzed.

[0052] The set of candidate influencing factors includes at least one candidate influencing factor.

[0053] Optionally, as one implementation, the corresponding value of the field includes one value of a field, multiple values ​​of a field, or one or more values ​​of multiple fields.

[0054] Optionally, as one implementation, the candidate condition is satisfied when the corresponding value of the field covers a proportion of data in the dataset of the problem to be analyzed that is higher than the candidate condition threshold or reaches 100%.

[0055] It should be understood that, in this embodiment of the application, the proportion of data in the dataset to be analyzed covered by the corresponding value of a field can be measured by the ratio of the amount of data (number of data entries) corresponding to the field's value to the total amount of data in the dataset to be analyzed. That is, the proportion of all such data (i.e., the number of data entries that have the field and whose value is the corresponding value mentioned above) to the total amount of data in the dataset to be analyzed.

[0056] Similarly, the confidence level calculation below is also based on the ratio of the amount of data corresponding to the field's value to the total amount of data. The total amount of data corresponding to different confidence levels is different, which can be found in the definition below, and will not be repeated here.

[0057] Specifically, after completing step 120 of the distribution statistics, this embodiment of the application can automatically find one or more fields (or combinations of fields) whose specific values ​​of one or more fields or a certain range have a high proportion (here referred to as the range of influence) in the problem dataset Ei, for example, higher than the candidate condition threshold or reaching 100%. The higher the proportion, the more likely it is that an influencing factor explaining these differences in the data will be found later. A "field-value" pair that meets the above conditions with a high proportion can be regarded as a candidate influencing factor C (hereinafter also referred to as candidate factor C). It should be understood that the candidate condition threshold here can be a pre-set value, for example, 95%, 90%, or 80%, and this embodiment of the application is not limited to this. In practical applications, the data proportion can be directly set to 100% to be identified as a candidate influencing factor, or it can be set to be higher than the candidate condition threshold to be identified as a candidate influencing factor.

[0058] For example, embodiments of this application may determine whether only a subset of fields meet the candidate criteria based on the statistical results. For instance, fields may be filtered sequentially according to their frequency of occurrence from highest to lowest, or only a subset of fields with a high statistical count may be filtered, or fields with a statistical count reaching a certain condition, such as those exceeding a certain frequency or a certain frequency proportion may be filtered. Alternatively, embodiments of this application may filter all fields that appear in the statistics; however, embodiments of this application are not limited to these methods.

[0059] Optionally, in this embodiment, after selecting a candidate influencing factor, subsequent steps can be performed directly. After completing step 140, the process can be looped back to select the next candidate influencing factor. This method enables rapid identification when there are few influencing factors, improving the efficiency of influencing factor identification. Alternatively, in step 120 of this embodiment, all fields that meet the requirements can be selected first before proceeding to subsequent steps.

[0060] If the requirement is that the influence range must reach 100% to determine candidate influencing factors, that is, to find the influencing factors of all differences in category Ei, then it is necessary to find the value range / list of one or more fields with a distribution percentage of 100%. This can be filtered from high to low using the distribution statistics obtained in step 120.

[0061] For example, if a certain value or range of values ​​in a field has reached 100% coverage, then that value or range of values ​​in the field can be used as a candidate influencing factor C.

[0062] For example, if multiple values ​​(called a value list) or multiple range combinations of a field can form 100% coverage of Ei, then the value list or range combination of that field can be used as a candidate influencing factor C.

[0063] For example, if a combination of specific values / value lists / value ranges of multiple fields can achieve 100% coverage of Ei, then the values / value lists / value ranges of these fields can be used as candidate influencing factors C.

[0064] Optionally, as another embodiment, if no field meets the candidate criteria or no candidate influencing factors meet the confidence level criteria, the method may further include: after further reorganizing or refining the classification of the problem data, performing attribute statistics on the data fields and filtering the set of candidate influencing factors again. In other words, in this embodiment, the classification of the problem datasets in the problem data set can be readjusted. For example, two or more problem datasets can be merged into one problem dataset, or one problem dataset can be split into multiple problem datasets, or the problem datasets can be reorganized. The reorganization can be flexible, and this embodiment does not limit this. Since no influencing factors were found previously, the previous classification may have been unreasonable or inaccurate. Adjusting the classification of the problem data before determining the influencing factors can increase the probability of finding them.

[0065] For example, the classification of E might need to be adjusted to find better coverage and explanatory factors. For instance, initially, "null value differences" and "zero value differences" might be analyzed as a single difference Ei. If it turns out that a distribution feature field with 100% coverage cannot be found, then the two types of differences can be separated into two categories, and the analysis can be repeated.

[0066] Alternatively, if it is acceptable to have a similar but not necessarily 100% impact range, i.e., to find the factors influencing most of the differences in classification Ei, then the above steps can be used to screen candidate influencing factors based on a self-defined threshold, i.e., the candidate condition threshold. The specific process is similar to the 100% coverage case described above, and will not be repeated here.

[0067] For example, if the condition "Source Department = 'Marketing Department A'" covers 98% of the data in Ei (exceeding the candidate condition threshold), then C=('Source Department', '=', 'Marketing Department A') is a strong candidate factor, that is, it can be used as a candidate influencing factor.

[0068] If multiple candidate factors C1, C2, ... with high proportions are found, they are all included in the candidate influencing factor set and then screened through subsequent confidence level calculations.

[0069] 140. From the set of candidate influencing factors, select the candidate influencing factors that meet the confidence level conditions to obtain the final influencing factors of the dataset of the problem to be analyzed.

[0070] Optionally, as one implementation, the confidence condition includes one or more of the following conditions: The internal confidence level is higher than the internal confidence level threshold or reaches 100%, the random confidence level is higher than the random confidence level threshold or reaches 100%, and the global confidence level is higher than the global confidence level threshold or reaches 100%.

[0071] Wherein, the internal confidence level is 1 - p_internal, where p_internal represents the proportion of candidate influencing factors in other problem datasets besides the problem dataset to be analyzed in the problem dataset set; the random confidence level is 1 - p_random, where p_random represents the proportion of candidate influencing factors in random datasets (including "correct data" without problems) in other datasets besides the problem dataset to be analyzed in the overall dataset set; and the global confidence level is 1 - p_global, where p_global represents the proportion of candidate influencing factors in other datasets (including "correct data" without problems) in other datasets besides the problem dataset to be analyzed in the overall dataset set.

[0072] Specifically, the aforementioned steps yielded a list of candidate factors, C. However, some of these candidate factors are highly likely to be spurious. For example, consider a dataset about men's clothing purchases. Analyzing a specific variance dataset Ei, it was found that the "gender" field of this dataset completely covers the variance data in Ei. Therefore, the gender field becomes a candidate influencing factor. However, it's clear that the gender field has no explanatory power for the variance here, as all gender fields in this dataset are "male." Using this field does not help in understanding the cause of the variance and is therefore not a true influencing factor for the variance classification Ei.

[0073] To evaluate the true explanatory power of candidate influencing factor C and avoid spurious correlations, this application introduces three core confidence indices. The core idea behind these indices is that a good influencing factor should be highly concentrated in the problem data and as sparse as possible in the non-problem data.

[0074] It should be understood that in practical applications, the confidence level conditions may include any one or more of the above conditions, and may be flexibly adjusted according to different confidence level requirements. The embodiments of this application are not limited thereto.

[0075] For example, internal confidence is defined as measuring how well a candidate factor C distinguishes between the current problem category Ei and other problem categories.

[0076] The internal confidence level can be calculated as follows: Let E' = E - Ei (i.e., the portion of the total problem data excluding the current category Ei). Calculate the proportion p_internal of C appearing in E'. The internal confidence level is defined as 1 - p_internal.

[0077] It should be understood that a higher internal confidence level indicates that factor C is more specifically targeted at problems like Ei, rather than a feature that is universally present in all problem data. Internal confidence levels can be used to select one or more of the most explanatory factors from the candidate set (i.e., the candidate influence set) as the final influencing factors.

[0078] Random confidence is defined as follows: it measures the ability of a candidate factor C to distinguish between the problem data and randomly sampled normal data. This method is suitable for big data scenarios and can be quickly verified.

[0079] The dataset can be calculated as follows: After excluding the problematic data Ei from the overall dataset D, the resulting dataset D' = D - Ei. Randomly select N data points (e.g., N=10000) from D' to form a sample set S. Calculate the proportion p_random of C appearing in S. The random confidence level is defined as 1 - p_random.

[0080] It should be understood that a higher random confidence score indicates that factor C is concentrated in the problem data but rare in the rest of the data, further increasing its likelihood as a root cause factor. The random confidence score can be used to select one or more of the most explanatory factors from the candidate set as the final influencing factors.

[0081] Global confidence is defined as follows: it measures the ability of candidate factor C to distinguish between problematic data and all normal data, and is the most stringent evaluation metric.

[0082] The global confidence score can be calculated as follows: Calculate the proportion p_global of C occurrences in the entire normal dataset D' = D - Ei. The global confidence score is defined as 1 - p_global.

[0083] It should be understood that a higher global confidence score indicates stronger explanatory power and greater reliability of the factor. When computational resources permit, this can be used as the final basis for judgment.

[0084] Optionally, as an implementation approach, calculating various confidence levels can be significantly resource-intensive in scenarios with large datasets. In practical applications, calculations can be performed in the priority order of internal confidence level > random confidence level > global confidence level to achieve a balance between efficiency and accuracy, quickly identifying the most likely influencing factors. This allows for the selection of the most reliable explanatory factor from multiple potential candidate influencing factors as the final influencing factor. If the data volume is very large, random confidence level may serve as an ideal verification standard in many scenarios. Specifically, various combinations of the above conditions can be used to filter the final influencing factors according to actual needs.

[0085] It should be understood that the internal confidence threshold can be 95%, 90%, 85%, or 80%, etc. The embodiments of this application are not limited to this. In practical applications, when the confidence condition includes the size of the internal confidence, the internal confidence can be higher than the internal confidence threshold, or the internal confidence can be equal to 100%. The embodiments of this application are not limited to this.

[0086] It should be understood that the random confidence threshold can be 95%, 90%, 85%, or 80%, etc. The embodiments of this application are not limited to this. In practical applications, when the confidence condition includes the size of the random confidence, the random confidence can be higher than the random confidence threshold, or the random confidence can be equal to 100%. The embodiments of this application are not limited to this.

[0087] It should be understood that the global confidence threshold can be 95%, 90%, 85%, or 80%, etc. The embodiments of this application are not limited to this. In practical applications, when the confidence condition includes the global confidence level, the global confidence level can be higher than the global confidence threshold, or the global confidence level can be equal to 100%. The embodiments of this application are not limited to this.

[0088] It should also be understood that when the confidence level condition includes multiple confidence level values, different confidence levels can be higher than the corresponding confidence level threshold, or reach 100%, or partially higher than the corresponding confidence level threshold, or partially reach 100%. The embodiments of this application are not limited to this.

[0089] Optionally, as another embodiment, the problem data set includes the problem data itself. If there are no fields that meet the candidate conditions or no candidate influencing factors that meet the confidence level conditions, the method further includes: After further reorganizing or refining the problem data classification, the actions of re-performing attribute statistics of data fields, filtering candidate influencing factor sets, and filtering final influencing factors are carried out; and / or, The data fields of the associated data in the dataset to be analyzed and the attributes of the data fields in the dataset to be analyzed are statistically analyzed together. The statistical results are updated to reflect the distribution characteristics of the data fields corresponding to the dataset to be analyzed. Then, the process of filtering the set of candidate influencing factors and filtering the final influencing factors is repeated.

[0090] Specifically, in this embodiment, if no field meeting the candidate conditions is found after step 130, or if no candidate influencing factors meeting the confidence level conditions are found after step 140, this embodiment can further reorganize or refine the classification of the problem data, and then re-execute steps 120, 130, and 140. Alternatively, the data fields of the associated data of the problem dataset to be analyzed and the attributes of the data fields in the problem dataset to be analyzed can be statistically analyzed together, and the statistical results can be updated to the distribution characteristics of the data fields corresponding to the problem dataset to be analyzed, and then steps 130 and 140 can be re-executed.

[0091] The above describes initially performing statistics only on the fields of the problem data. Only when no influencing factors can be found will adjustments be made to the problem data classification, or statistics will be performed on the fields together with the problem data and related data. This process is repeated until an influencing factor is found or a termination condition is met, such as satisfying a certain number of iterations (e.g., N iterations, where N is an integer greater than or equal to 2). Alternatively, this embodiment can consider the possibility that influencing factors may exist in related data from the outset, i.e., statistics will be performed on the fields of related data from the beginning. Correspondingly, as another embodiment, the statistical analysis of the attributes of the data fields in the problem dataset to be analyzed to obtain the distribution characteristics of the data fields corresponding to the problem dataset can include: performing statistics on the data fields of the related data of the problem dataset to be analyzed and the attributes of the data fields in the problem dataset to be analyzed together to obtain the distribution characteristics of the data fields corresponding to the problem dataset to be analyzed.

[0092] For example, if the problem with a dataset might be caused by the influence of other related datasets, the related datasets should be linked and integrated with dataset D into a complete dataset before analysis, replacing D. Simultaneously, when performing data field statistics, the fields of both the problem data and the related data should be statistically analyzed. Then, the analysis should proceed using the steps described above. For instance, if the problem data originates in the user expense details table, but the influencing factors might appear in the user attribute table, these two tables need to be linked first (i.e., both need field data statistics) before continuing the aforementioned analysis process.

[0093] Compared with the prior art, the embodiments of this application can achieve the following beneficial effects: Completely unsupervised and requiring no prior knowledge: This application's embodiments do not rely on any predefined rule base, expert experience, or labeled data, and can automatically infer and discover patterns that lead to problems directly from raw data. This makes the method highly universal and applicable to any industry and any type of structured data, greatly lowering the application threshold.

[0094] Automatic discovery of single and combined influencing factors: This application's embodiments, through comprehensive distribution statistics on single-field and multi-field combinations, can discover data problems caused by complex conditions, solving the pain point of traditional manual investigation being unable to discover combined factors. For example, it can discover deep-seated problems such as "order discount calculation errors only occur for VIP users who place orders through the iOS client on weekends."

[0095] A scientific confidence assessment system: This application proposes three confidence indices—internal, random, and global—to scientifically and quantitatively evaluate the explanatory power of candidate factors from different perspectives. Depending on actual needs, any combination of these three confidence indices can effectively eliminate accidental coincidences and spurious correlations, ensuring the high credibility of the final output influencing factors and providing a reliable basis for subsequent decision-making.

[0096] Efficient Analysis Strategy: In embodiments of this application, when all three confidence levels need to be calculated, the confidence levels are calculated in stages (first internal, then random, and finally global), enabling efficient analysis on large datasets. Furthermore, users can selectively perform calculations based on timeliness requirements, finding the optimal balance between analysis depth and resource consumption.

[0097] It should be understood that the embodiments of this application have broad application value in multiple key stages of the data lifecycle, and can be applied to various scenarios to solve specific problems in corresponding scenarios. Several application scenarios are listed below as examples.

[0098] Data Governance and Quality Monitoring: When the data quality monitoring platform detects data anomalies, it can automatically call this method to perform root cause analysis, helping the data governance team quickly locate the source of the problem (such as a faulty ETL task or an incorrect data entry standard in a department), thereby formulating precise remediation and prevention measures.

[0099] Data Migration and Integration: In data migration projects, inconsistencies between the source and target data frequently arise. This method can quickly analyze the key factors causing these inconsistencies (such as errors in handling specific data types or incompatibility with specific timestamp formats), significantly shortening the project debugging cycle and ensuring migration quality.

[0100] Business operations analysis: Business analysts can use this method to analyze abnormal fluctuations in business metrics. For example, analyzing "why the user churn rate of a certain product line suddenly increased last week" might reveal that it was caused by "compatibility issues of the new version of the app with specific mobile phone models," thereby driving product optimization.

[0101] A / B test result analysis: In A / B testing, if there are unexpected and significant differences in the core indicators between the experimental group and the control group, this method can be used to analyze whether there are certain "contamination" factors (such as abnormal user traffic from specific channels) that interfere with the fairness of the experimental results.

[0102] Please refer to Figure 2 , Figure 2 A block diagram of an apparatus for analyzing the influencing factors of data problems, according to an embodiment of this application, is shown. Figure 2 The device 200 shown can be a device for performing influencing factor analysis or a functional module within a device. It should be understood that the device 200 corresponds to the above method and is capable of performing the various steps involved in the above method embodiments. The specific functions of the device 200 can be found in the description above. To avoid repetition, detailed descriptions are appropriately omitted here.

[0103] Figure 2 The illustrated device 200 includes at least one software function module that can be stored in a memory or embedded in the device in the form of software or firmware. Figure 2The apparatus 200 shown includes: an acquisition unit 210, used to acquire a problem data set, the problem data set including at least one problem dataset, wherein the data in each problem dataset corresponds to the same type of data problem, different problem datasets correspond to different data problems, and each problem dataset includes at least one problem data; a statistics unit 220, used to perform statistics on the attributes of data fields in the problem dataset to be analyzed, to obtain the distribution characteristics of the data fields corresponding to the problem dataset to be analyzed, wherein the problem dataset to be analyzed is any one of the at least one problem datasets; a first filtering unit 230, used to filter fields that meet candidate conditions from the distribution characteristics of the data fields, to obtain a set of candidate influencing factors for the problem dataset to be analyzed, the set of candidate influencing factors including at least one candidate influencing factor; and a second filtering unit 240, used to filter candidate influencing factors that meet confidence conditions from the set of candidate influencing factors to obtain the final influencing factors for the problem dataset to be analyzed.

[0104] In one implementation, the attributes of the data fields in the dataset to be analyzed include: time type, numeric type, or text type.

[0105] In one implementation, the candidate condition is met when the corresponding value of the field covers a proportion of data in the dataset to be analyzed that is higher than the candidate condition threshold or reaches 100%.

[0106] In one implementation, the corresponding value of the field includes one value of a field, multiple values ​​of a field, or one or more values ​​of multiple fields.

[0107] In one implementation, the confidence level condition includes one or more of the following conditions: internal confidence level is higher than an internal confidence level threshold or reaches 100%, random confidence level is higher than a random confidence level threshold or reaches 100%, and global confidence level is higher than a global confidence level threshold or reaches 100%. The internal confidence level is 1 - p_internal, where p_internal represents the proportion of candidate influencing factors in other problem datasets besides the problem dataset to be analyzed in the problem dataset set. The random confidence level is 1 - p_random, where p_random represents the proportion of candidate influencing factors in random datasets among other datasets besides the problem dataset to be analyzed in the overall dataset set. The global confidence level is 1 - p_global, where p_global represents the proportion of candidate influencing factors in other datasets besides the problem dataset to be analyzed in the overall dataset set.

[0108] In one embodiment, the problem data set includes the problem data itself. If there are no fields that meet the candidate conditions or no candidate influencing factors that meet the confidence level conditions, the device further includes: a processing unit, used to further refine the classification of the fields and then re-perform the actions of statistical analysis of the data fields, filtering the set of candidate influencing factors, and filtering the final influencing factors; and / or, after updating the data associated with the problem data to the problem data set, to re-perform the actions of statistical analysis of the data fields, filtering the set of candidate influencing factors, and filtering the final influencing factors.

[0109] In one implementation, the problem dataset includes the problem data itself and data associated with the problem data.

[0110] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the aforementioned method, and will not be elaborated further here.

[0111] like Figure 3 As shown, one embodiment of this application provides an electronic device 300, which includes a memory 310, a processor 320, and a computer program stored in the memory 310 and executable on the processor 320. When the processor 320 reads the program from the memory 310 via a bus 330 and executes the program, it can implement the methods as described in any of the above embodiments. Optionally, Figure 3 The illustrated device may also include a transceiver that can be used for sending and / or receiving data. Optionally, Figure 3 The device shown may also include a data collector, which can be used to collect data. Optional electronic device 300 may also include other input devices or other hardware devices, etc., and the embodiments of this application are not limited thereto.

[0112] Processor 320 can process digital signals and may include various computing architectures. For example, it may be a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements multiple instruction set combinations. In some examples, processor 320 may be a microprocessor.

[0113] The memory 310 can be used to store instructions executed by the processor 320 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all of the functions of one or more modules described in the embodiments of this application. The processor 320 of this disclosure embodiment can be used to execute the instructions in the memory 310 to implement the above-described methods. The memory 310 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.

[0114] One embodiment of this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the methods described in the above embodiments.

[0115] An embodiment of this application also provides a computer program product, which includes a computer program, wherein the computer program, when executed by a processor, can implement the methods provided in the above embodiments.

[0116] It should be noted that the processor in the embodiments of the present invention (e.g., Figure 3 The processor in the above method embodiments can be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by software instructions. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The software module can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0117] It can be understood that the memory in the embodiments of the present invention (e.g., Figure 3The memory in the memory can be volatile or non-volatile, or it can include both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0118] It should be understood that the transceiver unit or transceiver in the embodiments of the present invention may also be referred to as a communication unit.

[0119] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0120] It should be understood that, in the embodiments of the present invention, "B corresponding to A" means that B is associated with A, and B can be determined based on A. However, it should also be understood that determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information.

[0121] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0122] In summary, the above description is merely a preferred embodiment of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for analyzing the influencing factors of data problems, characterized in that, include: Obtain a problem data set, which includes at least one problem dataset, wherein the data in each problem dataset corresponds to the same type of data problem, and different problem datasets correspond to different data problems, and each problem dataset includes at least one problem data. The attributes of the data fields in the dataset to be analyzed are statistically analyzed to obtain the distribution characteristics of the data fields corresponding to the dataset to be analyzed, wherein the dataset to be analyzed is any one of the at least one datasets of problems; Fields that meet the candidate conditions are selected from the distribution characteristics of the data fields to obtain a set of candidate influencing factors for the dataset of the problem to be analyzed. The set of candidate influencing factors includes at least one candidate influencing factor. The final influencing factors of the dataset to be analyzed are obtained by selecting candidate influencing factors that meet the confidence criteria from the set of candidate influencing factors.

2. The method according to claim 1, characterized in that, The attributes of the data fields in the dataset to be analyzed include: time type, numeric type, or text type.

3. The method according to claim 1 or 2, characterized in that, If the corresponding value of a field covers a percentage of the data in the dataset to be analyzed that is higher than the candidate condition threshold or reaches 100%, then the candidate condition is satisfied.

4. The method according to claim 3, characterized in that, The corresponding values ​​of the field include one value of a field, multiple values ​​of a field, or one or more values ​​of multiple fields.

5. The method according to claim 1 or 2, characterized in that, The confidence level conditions include one or more of the following conditions: The internal confidence level is higher than the internal confidence threshold or reaches 100%. The random confidence level is higher than the random confidence level threshold or reaches 100%. The global confidence level is higher than the global confidence threshold or reaches 100%. Wherein, the internal confidence level is 1-p_internal, where p_internal represents the proportion of candidate influencing factors in other problem datasets besides the problem dataset to be analyzed in the problem dataset set; the random confidence level is 1-p_random, where p_random represents the proportion of candidate influencing factors in random datasets among other datasets besides the problem dataset to be analyzed in the overall dataset set; and the global confidence level is 1-p_global, where p_global represents the proportion of candidate influencing factors in other datasets besides the problem dataset to be analyzed in the overall dataset set.

6. The method according to claim 1 or 2, characterized in that, If no field meets the candidate criteria or no candidate influencing factor meets the confidence level criteria, the method further includes: After further reorganizing or refining the problem data classification, the actions of re-performing attribute statistics of data fields, filtering candidate influencing factor sets, and filtering final influencing factors are carried out; and / or, The data fields of the associated data in the dataset to be analyzed and the attributes of the data fields in the dataset to be analyzed are statistically analyzed together. The statistical results are updated to reflect the distribution characteristics of the data fields corresponding to the dataset to be analyzed. Then, the process of filtering the set of candidate influencing factors and filtering the final influencing factors is repeated.

7. The method according to claim 1 or 2, characterized in that, The step of statistically analyzing the attributes of the data fields in the dataset to be analyzed, to obtain the distribution characteristics of the data fields corresponding to the dataset to be analyzed, includes: The data fields of the associated data in the dataset to be analyzed and the attributes of the data fields in the dataset to be analyzed are statistically analyzed together to obtain the distribution characteristics of the data fields corresponding to the dataset to be analyzed.

8. An apparatus for analyzing the influencing factors of data problems, characterized in that, include: The acquisition unit is used to acquire a problem data set, which includes at least one problem dataset. Each problem dataset contains data corresponding to the same type of data problem. Different problem datasets correspond to different data problems. Each problem dataset includes at least one problem data. The statistical unit is used to perform statistics on the attributes of the data fields in the dataset of the problem to be analyzed, so as to obtain the distribution characteristics of the data fields corresponding to the dataset of the problem to be analyzed, wherein the dataset of the problem to be analyzed is any one of the at least one dataset of problems; The first filtering unit is used to filter fields that meet the candidate conditions from the distribution characteristics of the data fields to obtain a set of candidate influencing factors for the dataset of the problem to be analyzed, wherein the set of candidate influencing factors includes at least one candidate influencing factor. The second screening unit is used to screen candidate influencing factors that meet the confidence level conditions from the set of candidate influencing factors to obtain the final influencing factors of the dataset of the problem to be analyzed.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor, wherein the computer program is executed by the processor to perform the method as claimed in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, performs the method as described in any one of claims 1-7.