Method and apparatus for determining data anomaly, and device and medium
By acquiring the historical status of key fields in the dataset, the system automatically identifies data anomalies and determines the causal fields, thus solving the accuracy and efficiency problems of data anomaly identification in existing technologies and achieving efficient dataset management.
Patent Information
- Application Number
- PCT/SG2024/050402
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2025-12-26
AI Technical Summary
Existing technologies struggle to accurately and efficiently identify data anomalies in datasets, leading to inaccurate execution of analysis tasks.
A method and apparatus are provided to determine data anomalies by acquiring the historical status of key fields and automatically identify the fields that cause the anomalies, thereby reducing the complexity of manual management and improving the efficiency of dataset management.
It enables accurate identification and cause analysis of data anomalies, reduces the complexity of manual management, and improves the efficiency and accuracy of dataset management.
Smart Images

Figure SG2024050402_26122025_PF_FP_ABST
Abstract
Description
[0001]Methods, Apparatus, Devices, and Media for Determining Data Anomalies The present disclosure generally relates to dataset management, and particularly to methods, apparatus, devices, and computer-readable storage media for determining data anomalies in datasets. Background Art Datasets can be used to store various types of data, such as application-related data. Multiple users can install applications on their respective client devices, generating a large amount of data as users use the data. The dataset can then include numerous fields from multiple data sources. Analytical tasks can be performed on the data in the dataset, such as determining the relationships between certain fields, etc. However, anomalies may occur in the dataset, preventing accurate execution of analytical tasks. Typically, dataset administrators need to manually discover and handle anomalies to determine the source of the anomalies. Therefore, it is desirable to determine data anomalies in the dataset in a more accurate and effective manner. Summary of the Invention In a first aspect of this disclosure, a method for determining data anomalies is provided. In this method, key fields in a dataset are determined, the dataset including multiple data sources, and each of the multiple data sources includes at least one field. The historical state of the key fields within a historical time period is obtained. In response to determining that changes in data in a key field, indicating a historical state, satisfy anomaly conditions, a data anomaly is determined in the key field, indicating that an anomaly occurred in the key field within a historical time period. At least one cause field associated with the data anomaly is identified in the dataset, and an anomaly in the at least one cause field causes a data anomaly in the key field. In a second aspect of this disclosure, an apparatus for determining a data anomaly is provided.The apparatus includes: a field determination module configured to determine a key field in a dataset, the dataset including multiple data sources, and each of the multiple data sources including at least one field; a status acquisition module configured to acquire the historical status of the key field within a historical time period; an anomaly determination module configured to determine that a data anomaly exists in the key field in response to a determination that a change in data in the key field satisfies an anomaly condition, the data anomaly indicating that an anomaly occurred in the key field within the historical time period; and a cause determination module configured to determine at least one cause field in the dataset associated with the data anomaly, the data anomaly of the at least one cause field causing the data anomaly of the key field. In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to the first aspect of this disclosure when executed by the at least one processing unit. In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, the computer program causing the processor to implement the method according to the first aspect of this disclosure when executed by a processor. In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure. It should be understood that the description in this content section is not intended to limit the key or essential features of the implementations of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Brief Description of the Drawings In the following detailed description, the above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent.In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 shows a block diagram of an application environment for determining data anomalies; Figure 2 shows a block diagram of a method for determining data anomalies according to some implementations of the present disclosure; Figure 3 shows a block diagram of a module for determining data anomalies according to some implementations of the present disclosure; Figure 4 shows a block diagram of a method for determining key fields according to some implementations of the present disclosure; Figure 5 shows a block diagram of a page for providing information related to data anomalies according to some implementations of the present disclosure; Figure 6 shows a block diagram of the mapping relationship between various fields according to some implementations of the present disclosure; Figure 7 shows a flowchart of a method for determining data anomalies according to some implementations of the present disclosure; Figure 8 shows a block diagram of an apparatus for determining data anomalies according to some implementations of the present disclosure; and Figure 9 shows a block diagram of an apparatus capable of implementing multiple implementations of the present disclosure. Detailed Description Implementations of the present disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the accompanying drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure. In the description of implementations of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationship between various data. For example, the aforementioned relationship can be obtained based on various currently known and / or future-developed technical solutions. It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.It is understood that before using the technical solutions disclosed in the embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained. For example, in response to receiving a user's active request, a prompt message can be sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the operation of the technical solutions of this disclosure, based on the prompt message. As an optional but non-limiting implementation, the method of sending a prompt message to the user in response to receiving a user's active request can be, for example, a pop-up window, in which the prompt message can be presented in text form. Furthermore, the pop-up window can also include a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device. It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation of this disclosure; other methods that comply with relevant laws and regulations can also be applied to the implementation of this disclosure. The term "in response to" as used herein refers to the state in which a corresponding event occurs or a condition is met. It will be understood that the timing of subsequent actions performed in response to an event or condition is not necessarily strongly correlated with the time when the event or condition occurs. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; in other cases, they may be performed some time after the event or condition has occurred. The example environment currently presents several dataset management technologies that can utilize datasets to store various application-related data. Multiple users can install applications on their respective client devices, generating a large amount of data as users use the application. Datasets can include numerous fields from multiple data sources.Referring to Figure 1, which describes an application environment according to some implementations of this disclosure, Figure 1 shows a block diagram 100 of an application environment for determining data anomalies. As shown in Figure 1, dataset 110 may include multiple data sources 120, ..., and 130. Each data source may include one or more fields. For example, data source 120 may include fields 122, ..., and 124, and data source 130 may include fields 132, ..., and 134, etc. Analysis tasks can be performed on the data in the dataset, such as determining the relationships between certain fields, etc. However, anomalies may occur in the data in the dataset, which may prevent accurate execution of analysis tasks. Typically, the dataset administrator needs to manually discover and handle anomalies to determine the cause of the data anomalies. Therefore, it is desirable to determine data anomalies in the dataset in a more accurate and effective manner. To at least partially address the shortcomings of the prior art, according to one implementation of this disclosure, a method for determining data anomalies is proposed. Referring to Figure 2, which describes an overview of one implementation of this disclosure, Figure 2 shows a block diagram 200o for determining data anomalies according to some implementations of this disclosure. As shown in Figure 2, dataset 110 may include multiple data sources 120, ..., 130, and each of the multiple data sources may each include at least one field. A key field in dataset 110 can be determined; a key field 210 can be represented using a slashed box, where field 122 is the key field. Dataset 110 can be continuously updated over time, and the historical state of the key field within a historical time period can be obtained. Furthermore, this historical state can be analyzed to determine whether a data anomaly exists in the key field. In response to determining that the historical state indicates that changes in the data in the key field meet anomaly conditions, it can be determined that a data anomaly exists in the key field, indicating that an anomaly occurred in the key field within a historical time period. Specific anomaly conditions can be specified according to the specific application environment; for example, a threshold amplitude of data fluctuation, a threshold time period length of data fluctuation, etc., can be specified. Furthermore, at least one cause field associated with the data anomaly can be determined in the dataset; the data anomaly of at least one cause field causes the data anomaly of the key field.For example, a grid wireframe can be used to represent the cause field 220, where field 132 is the cause field. Here, the cause field is the reason for the data anomaly in the key field; that is, the data anomaly in the cause field leads to the data anomaly in the key field. Using the implementation of this disclosure, data in each field of the dataset can be dynamically managed, and the anomaly in the key field of interest to the user can be automatically determined, along with the specific reason for the anomaly. This reduces the complexity of manual management and improves the efficiency of dataset management. The detailed process for determining data anomalies has been described in summary according to some implementations of this disclosure. Further information regarding determining data anomalies is described below with reference to Figure 3. Figure 3 shows a block diagram 300 of a module for determining data anomalies according to some implementations of this disclosure. As shown in Figure 3, the data management module 310 can be used to acquire data of interest from the dataset and perform further management. Specifically, the data acquisition module 312 can extract one or more important data indicators (e.g., key fields) from the existing dataset and then provide alerts by observing fluctuations in these data indicators. Since each key field can cover fluctuations in most business anomaly scenarios, various anomalies can be analyzed during the business process. The fluctuation alarm module 314 can present alarms in real time based on fluctuations and present them in a visual manner. This allows the recipient to clearly observe the key fields and time periods in which the fluctuation alarms occurred. The data output module 316 can take the determined key field names and occurrence times as the output of the data management module 310 and input them into the anomaly diagnosis module 320. Furthermore, the anomaly diagnosis module 320 can be responsible for diagnosis-related work and provide diagnostic results. Specifically, after receiving the key field names and anomaly time periods from upstream, the dimension decomposition module 322 can automatically decompose the key fields into multiple fields through enumeration. Furthermore, the cause localization module 324 can find one or more cause fields from the decomposed multiple fields. Specifically, a data script can be run to analyze which specific dimension the data fluctuations during the alarm time period occurred in. Through the pre-obtained dimension tracing table, the relevant fields of the data source are found.Finally, the results providing module 326 can provide the diagnostic results and fluctuation alarms together to relevant personnel, such as the dataset administrator or the person initiating a specific task within the dataset. Further details of each module are described below with reference to the accompanying drawings. According to some implementations of this disclosure, the dataset can store various types of data. For example, in an application environment that manages application data, the application provider can publish the application, and multiple users can download and install the application on their respective client devices. Media items can be provided to various client devices via the application, and users can interact with these media items, thereby generating various types of events. In this application environment, the dataset is used to store data associated with client devices among multiple client devices. Multiple fields may include a first set of attributes of the client device (e.g., the device's region, operating system type, operating system version number, etc.), a second set of attributes of the application installed on the client device (e.g., application name, identifier, version number, etc.), a third set of attributes of data items sent to the client device via the application (data item identifier, type, source, provider, etc.), and a fourth set of events associated with the data items (e.g., click events, comment events, forwarding events, conversion events, etc.). In this way, multiple attributes involved during application runtime can be fully recorded, thereby improving the efficiency of application management. According to some implementations of this disclosure, in the process of determining key fields in the dataset, candidate fields to be observed in the dataset can be determined based on user requirements from among the multiple fields included in the dataset. Specifically, based on business needs and experience, one or more key fields can be defined and collected. These fields have a direct and profound impact on the business and can therefore serve as a data foundation to measure whether various data collected during application operation exhibit anomalies. For example, candidate fields may include user conversion rate, click volume, resource consumption, etc. According to some implementation methods of this disclosure, among multiple fields in the dataset, upstream fields affecting candidate fields can be identified and designated as key fields. Specifically, to save data storage costs and alleviate computational pressure, while ensuring that the key fields of interest are clear and easy to understand, more fundamental fields are sought among the upstream fields to reduce the absolute number of key fields.It is possible to determine the dependencies between various fields. For example, it is possible to obtain the relevant calculation formulas for each field, and then use the fields of the variables involved in the formula as upstream fields. See Figure 4 for further details. Figure 4 shows a block diagram 400o for determining key fields according to some implementations of this disclosure. As shown in Figure 4, field 431 indicates "reach rate" and is calculated based on "clicks" in field 410. Therefore, field 410 is an upstream field of field 431. Field 410 is an upstream field of fields 432, 433, and 434. That is, assuming the candidate fields indicate reach rate, conversion rate, and click-through rate, the key field can be determined as field 410, i.e., clicks. Similarly, field 420 (activation count) is an upstream field of fields 432, 435, and 436. That is, assuming the candidate fields indicate conversion rate and cycle value, the key field can be determined as field 420, i.e., activation count. In this way, the number of fields being detected can be reduced, thereby reducing various related resource overheads. According to some implementations of this disclosure, in the process of determining key fields in a dataset, The task to be executed in the dataset can be obtained, which is to determine the correlation between multiple events. Furthermore, key fields related to the task can be identified from the multiple fields included in the dataset. For ease of description, more details on identifying key fields are described using the execution of a judgment task as an example. In this case, for the judgment task, key fields may further include device identification information, etc. Since the above information directly affects the judgment result, identifying key fields based on the above information can more accurately identify abnormal fields that may lead to data judgment anomalies, thereby improving the accuracy of subsequent data judgments. According to some implementation methods of this disclosure, in the process of obtaining the historical state of key fields within a historical time period, an automated script can be used to extract the historical state from the multiple fields included in the dataset. Specifically, an executable script can be pre-built to extract the historical state of the key fields of interest from the dataset. For example, the name of the key field and the time period to be processed (i.e., the period of interest, such as the past 1 day, 2 days, 1 week, etc.) can be specified to obtain the corresponding data. For example, the above script can be executed periodically, and the extraction results can be stored in a specified data table.According to some implementation methods of this disclosure, these key fields can be managed in real time through a set of predetermined rules and threshold settings. The threshold settings can be automatically and dynamically changed over time. The management system will perform data comparison and analysis after the end of each data capture cycle. When the data fluctuation of a certain indicator exceeds the preset threshold, the system will automatically trigger an alarm mechanism. According to some implementation methods of this disclosure, the abnormal condition can specify at least one of the following: a threshold for determining the magnitude of data abnormality, or a threshold for determining the duration of data abnormality. A magnitude threshold can be pre-specified, which can represent the fluctuation range of data. When the data fluctuation exceeds this range, it is considered that a data abnormality has occurred; and when the data fluctuation is within this range, it is considered that no data abnormality has occurred. An absolute quantity or a relative quantity can be used to specify the magnitude threshold. Assuming that the key field is the number of clicks, the magnitude threshold can be represented as N (a positive integer). In this case, when the number of clicks fluctuates upward or downward by more than N, it is considered that an abnormality has occurred. Alternatively and / or additionally, the amplitude threshold can be expressed as M% (M being a number less than 100). In this case, if the click volume fluctuates upward or downward by more than M%, it is considered an anomaly. Alternatively and / or additionally, a duration threshold can be pre-specified. This duration threshold represents the time range of data fluctuations. If the data fluctuation involves a long period and exceeds this range, it is considered an anomaly; and if the data fluctuates only within a short period and is within this range, it is considered not an anomaly. The duration threshold can be specified using absolute or relative quantities. For example, the threshold can be specified as 10 minutes, 30 minutes, 1% of the monitoring period, or other values, etc. This method facilitates adjusting the anomaly judgment conditions for whether data anomalies occur, allowing for flexible adjustment based on the specific application environment. Specifically, assuming historical data indicates a significant increase in user clicks on weekends and holidays, the amplitude threshold for weekends and holidays can be appropriately increased, etc. If data anomalies are confirmed, alert information can be provided. This alert information may include the names of key fields, the time period in which the data anomalies occurred, and the magnitude of the fluctuations relative to historical data.According to some implementations of this disclosure, an exception page associated with data anomalies is provided. The exception page includes filtering parameters for presenting data anomalies, including at least one of the following: the time range of the data anomaly, the application involved in the data anomaly, the region involved in the data anomaly, the type of data item involved in the data anomaly, and the source involved in the data anomaly. Furthermore, in response to receiving an interaction with the filtering parameters, the exception page is updated. See Figure 5 for further details regarding the provision of anomaly information. Figure 5 shows a block diagram of a page 500 for providing information related to data anomalies according to some implementations of this disclosure. As shown in Figure 5, information related to multiple key fields can be presented on page 500. For example, control 522 can correspond to the key field "number of users," and the user can press control 522 to view information related to the number of users. Similarly, controls 524, 526, and 528, etc., can each correspond to multiple other key fields. Assuming the user selects control 522, the fluctuation curve 520 of the user number data can be displayed, and the historical average of the user number data 530 can also be displayed. According to some implementations of this disclosure, potential anomalous fields affected by a key field can be identified in the dataset; and potential anomalous data associated with these fields can be provided. As shown in Figure 5, page 500 can further display other information about the user number, such as user number realization rate (i.e., the ratio between the current user number and the expected user number), resource tracking (i.e., the ratio between currently used resources and planned resources), date tracking (i.e., the ratio between the time period presented on the current page and the expected observation time period), etc. In this way, information about other fields that may be affected by the key field can be automatically provided, allowing the user to have a comprehensive understanding of the data in the dataset. Page 500 can further include multiple filtering parameters; for example, the user can press control 510 to select the time range of data anomalies, such as displaying data anomalies by quarter, month, day, etc.Users can press control 512 to display data anomalies related to a specific application, press control 514 to display data anomalies related to a specific region (e.g., city A, city B, etc.), press control 516 to display anomalies related to a specific type of data item presented in the application, press control 518 to select the source of the anomaly data, and so on. This allows for the specification of the desired anomaly data from multiple perspectives, enabling users to obtain more information. It should be understood that the specific content of page 500 is merely illustrative, and page 500 may present more, less, or different information. The various steps performed by the data management module 310 have been described; further details of the anomaly diagnosis module 320 will be described below. According to some implementations of this disclosure, once a data anomaly in a key field is detected, the anomaly diagnosis module 320 can be invoked to determine at least one cause field associated with the data anomaly in the dataset. Specifically, an automated script can be used to determine at least one cause field, describing the mapping relationship between the key field and the at least one cause field. For example, the anomaly diagnosis module 320 can receive the names of key fields and information about the time period in which the anomaly occurred, and then call the dimension decomposition module 322 to automatically decompose the relevant dimensions. The anomaly diagnosis module 320 can call a predefined data script to execute the corresponding process. For example, the dimension decomposition module can automatically analyze various dimensional fields in the dataset based on the key fields. Specifically, for the dimension of client devices, it can enumerate the region where the device is located, the type of operating system of the device, the version number of the operating system of the device, etc.; for the dimension of applications installed on client devices, it can enumerate the name, identifier, version number of the application, etc.; for the dimension of data items published to client devices through applications, it can enumerate the identifier, type, source, provider of the data item, etc.; and for the event dimension associated with the data item, it can enumerate click events, comment events, forwarding events, conversion events, etc. In this way, the key fields that cause data anomalies can be automatically obtained as more refined dimensional fields, which makes it easier to find the fields that cause the data anomalies in the dataset. Furthermore, the field causing the data anomaly can be located from the individual fields after decomposition.It should be understood that a dataset may include multiple data sources, and the names of fields in different data sources may be different. For example, in one data source, a field name may be represented as "DT_ID"; however, in another data source, a field containing the same content may be represented as "AF_DT_ID". In this case, the field name alone cannot confirm that the two fields correspond to the same data item; a mapping relationship needs to be established between the fields. See Figure 6 for more information, which shows a block diagram of the mapping relationship 600 between the fields according to some implementations of this disclosure. As shown in Figure 6, dimension field 610 represents the decomposed dimension field (the "data item" field "DT_ID" obtained by decomposing the key field), source table name 620 represents the name of another data source determined through the tracing process, and source field 630 represents the field name of the "data item" in the data table "APP EVENT LOG" as "AF_DT_ID". In this way, a mapping relationship can be established between the field "DT_ID" in one data source and the field "AF_DT_ID" in another data source. In other words, although the names of the two fields are different, the content stored in both fields is a "data item". It should be understood that although Figure 6 only schematically shows an example of a mapping relationship between fields, alternatively and / or additionally, mapping relationship 600 may include more rows, and each row may describe a mapping relationship. For example, another mapping relationship may represent: the field "APP ID" in one data source and the field "AF_APP ID" in another data source "APP EVENT LOG". The ID corresponds to this. In this way, it's possible to quickly find the cause field using automated scripts. Furthermore, it's possible to determine whether the data in each found cause field is abnormal. In response to determining that the data in at least one cause field is abnormal, an anomaly status associated with the key field and the cause field can be provided. Specifically, if an anomaly is found in the data of the cause field (e.g., the data exceeds the normal threshold range, or the duration of the anomaly exceeds the allowable threshold duration), then it can be determined that the cause field itself is also abnormal.At this point, anomaly states can be provided to users, namely, the anomaly states of key fields and the anomaly states of various cause fields found through the tracing process. This provides richer information to dataset administrators, supporting their subsequent actions. According to some implementations of this disclosure, corresponding alert conditions can be specified for different fields. Here, alert conditions can include at least one of the following: the duration of the anomaly meets a threshold duration, and the magnitude of the anomaly meets a threshold magnitude change. Furthermore, in response to determining that the anomaly of each field meets the alert conditions, anomaly states associated with the key and cause fields can be provided. This allows for a more flexible and effective presentation of anomaly alerts, facilitating dataset administrators to discover the relationships between various anomaly fields and improving the accuracy of attribution tasks. Using the example implementations of this disclosure, upstream fields can be determined through specific calculation formulas for each key field, thereby reducing the number of fields to be processed. Alert thresholds can be dynamically set, allowing for the definition of fluctuations according to individual needs, effectively reducing the number of alerts and improving anomaly detection efficiency. Furthermore, the visualization page can provide optional dynamic indicators, thereby effectively reducing the complexity of manual operations. By establishing a dimensional source table, direct source tracing of abnormal dimensions can be achieved, thereby supporting dataset administrators to fully grasp all anomalies in the dataset. Example process Figure 7 shows a flowchart of a method 700 for determining data anomalies according to some implementations of this disclosure. At box 710, a key field in the dataset is determined, the dataset including multiple data sources, and each of the multiple data sources includes at least one field. At box 720, the historical state of the key field within a historical time period is obtained. At box 730, in response to determining that the historical state indicates that changes in the data in the key field meet anomaly conditions, a data anomaly is determined in the key field, indicating that an anomaly occurred in the key field within the historical time period. At box 740, at least one cause field associated with the data anomaly is determined in the dataset, the data anomaly of the at least one cause field causing the data anomaly of the key field.According to some implementations of this disclosure, determining key fields in a dataset includes: determining candidate fields to be observed in the dataset based on user needs among multiple fields included in the dataset; determining upstream fields that influence the candidate fields among multiple fields; and identifying the upstream fields as key fields. According to some implementations of this disclosure, obtaining the historical state of key fields within a historical time period includes: extracting historical states from multiple fields included in the dataset using an automated script. According to some implementations of this disclosure, the anomaly condition specifies at least one of the following: a threshold for determining the magnitude of a data anomaly, or a threshold for determining the duration of a data anomaly. According to some implementations of this disclosure, the method further includes: providing an anomaly page associated with the data anomaly, the anomaly page including data anomaly filtering parameters, the filtering parameters including at least one of the following: the time range of the data anomaly, the application involved in the data anomaly, the region involved in the data anomaly, the data item type involved in the data anomaly, and the source involved in the data anomaly; and updating the anomaly page in response to receiving an interaction regarding the filtering parameters. According to some implementations of this disclosure, determining at least one cause field associated with data anomalies in a dataset includes: using an automated script to determine at least one cause field, the automated script describing the mapping relationship between a key field and at least one cause field; and the method further includes: in response to determining that data in the cause field of the at least one cause field is abnormal, providing an anomaly state associated with the key field and the cause field. According to some implementations of this disclosure, providing an anomaly state associated with the key field and the cause field includes: in response to determining at least one of the following, providing an anomaly state: the duration of the anomaly meets a threshold duration, and the magnitude of the anomaly meets a threshold magnitude change. According to some implementations of this disclosure, the method further includes: determining potential anomaly fields in the dataset affected by the key field; and providing potential anomaly data associated with the potential anomaly fields.According to some implementations of this disclosure, the dataset is used to store data associated with client devices among multiple client devices. Multiple fields include first multiple attributes of the client devices, second multiple attributes of applications installed on the client devices, third multiple attributes of data items published to the client devices via the applications, and fourth multiple events associated with the data items. According to some implementations of this disclosure, determining key fields in the dataset further includes: obtaining a task to be performed on the dataset, the task being to determine the relationships between the fourth multiple events; and determining the key fields associated with the task among the multiple fields included in the dataset. Example Apparatus and Device Figure 8 shows a block diagram of an apparatus 800 for determining data anomalies according to some implementations of this disclosure. The apparatus 800 includes: a field determination module 810, configured to determine key fields in a dataset, the dataset including multiple data sources, and each of the multiple data sources including at least one field; a status acquisition module 820, configured to acquire the historical status of the key fields within a historical time period; an anomaly determination module 830, configured to determine that a data anomaly exists in the key field in response to a determination that a change in data in the key field indicates that an anomaly occurs in the key field within a historical time period; and a cause determination module 840, configured to determine at least one cause field in the dataset associated with the data anomaly, wherein the data anomaly of the at least one cause field causes the data anomaly of the key field. According to some implementations of this disclosure, the field determination module is further configured to: determine candidate fields to be observed in the dataset based on user requirements among the multiple fields included in the dataset; determine upstream fields affecting the candidate fields among the multiple fields; and determine the upstream fields as key fields. According to some implementations of this disclosure, the status acquisition module is further configured to: extract historical status from the multiple fields included in the dataset using automated scripts. According to some implementations of this disclosure, the abnormal condition specifies at least one of the following: a threshold for determining the magnitude of the data abnormality, or a threshold for determining the duration of the data abnormality.According to some implementations of this disclosure, the apparatus further includes: a page providing module configured to provide an exception page associated with data anomalies, the exception page including data anomaly filtering parameters, the filtering parameters including at least one of the following: the time range of the data anomaly, the application involved in the data anomaly, the region involved in the data anomaly, the data item type involved in the data anomaly, and the source involved in the data anomaly; and a page updating module configured to update the exception page in response to receiving an interaction with the filtering parameters. According to some implementations of this disclosure, the cause determination module is further configured to: determine at least one cause field using an automated script, the automated script describing the mapping relationship between a key field and at least one cause field. According to some implementations of this disclosure, the apparatus further includes: a providing module configured to provide an anomaly status associated with the key field and the cause field in response to determining that data in the cause field of at least one cause field is anomaly. According to some implementations of this disclosure, the providing module is further configured to: provide an anomaly status in response to determining at least one of the following: the duration of the anomaly meets a threshold duration, and the magnitude change of the anomaly meets a threshold magnitude change. According to some implementations of this disclosure, the apparatus further includes: a potential anomaly determination module configured to determine potential anomaly fields affected by key fields in the dataset; and a potential anomaly providing module configured to provide potential anomaly data associated with the potential anomaly fields. According to some implementations of this disclosure, the dataset is used to store data associated with client devices among a plurality of client devices, the plurality of fields including first plurality of attributes of the client devices, second plurality of attributes of applications installed on the client devices, third plurality of attributes of data items published to the client devices via the applications, and fourth plurality of events associated with the data items. According to some implementations of this disclosure, the field determination module is further configured to: obtain a task to be performed in the dataset, the task being to determine the relationships between the fourth plurality of events; and determine the key field associated with the task among the plurality of fields included in the dataset. Figure 9 shows a block diagram of an apparatus 900 capable of implementing various implementations of this disclosure.It should be understood that the computing device 900 shown in Figure 9 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementation described herein. The computing device 900 shown in Figure 9 can be used to implement the methods described above. As shown in Figure 9, the computing device 900 is in the form of a general-purpose computing device. Components of the computing device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage devices 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 may be a physical or virtual processor and is capable of performing various processes according to a program stored in memory 920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 900. The computing device 900 typically includes multiple computer storage media. Such media can be any available media accessible to the computing device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 920 may be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 930 may be a removable or non-removable medium and may include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 900. Computing device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9, disk drives for reading from or writing to removable, non-volatile disks (e.g., “floppy disks”) and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each driver may be connected to a bus (not shown) via one or more data media interfaces. The memory 920 may include a computer program product 925 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.The communication unit 940 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 900 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 900 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node. The input device 950 can be one or more input devices, such as a mouse, keyboard, trackball, etc. The output device 960 can be one or more output devices, such as a monitor, speakers, printer, etc. The computing device 900 can also communicate as needed with one or more external devices (not shown) via the communication unit 940. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with the computing device 900, or with any device that enables the computing device 900 to communicate with one or more other computing devices (e.g., network interface card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown). According to an implementation of this disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the method described above. According to an implementation of this disclosure, a computer program product is provided, on which a computer program is stored, which, when executed by a processor, implements the method described above. Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of a flowchart illustration and / or block diagram, and combinations of blocks in a flowchart illustration and / or block diagram, can be implemented by computer-readable program instructions.These computer-readable program instructions can be provided to the processing unit of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, these instructions create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. The computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions, which execute on the computer, other programmable data processing apparatus, or other device, to implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram. The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. Various implementations of this disclosure have been described above; the foregoing description is exemplary, not exhaustive, and is not limited to the disclosed implementations.Without departing from the scope and spirit of the implementations described, many modifications and changes will be apparent to those skilled in the art. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the implementations disclosed herein.
Claims
Claims 1. A method for determining data anomalies, comprising: determining a key field in a data set, the data set including a plurality of data sources, and each of the plurality of data sources including at least one field respectively; obtaining a historical state of the key field in a historical time period; in response to determining that the historical state indicates that a change in data in the key field satisfies an exception condition, determining that the key field has a data exception, the data exception indicating that the key field has an anomaly in the historical time period; and determining at least one cause field in the data set that is associated with the data exception, the data exception of the at least one cause field causing the data exception of the key field.
2. The method of claim 1, wherein determining the key fields in the dataset comprises: determining, among a plurality of fields included in the data set, a candidate field to be observed in the data set based on a user demand; determining, among the plurality of fields, an upstream field that affects the candidate field; and determining the upstream field as the key field.
3. The method of claim 1, wherein obtaining the historical states of the key fields over the historical time period comprises: extracting the historical state from the plurality of fields included in the data set using an automated script.
4. The method of claim 1, wherein the exception condition specifies at least one of: a magnitude threshold for determining the data exception, or a duration threshold for determining the data exception.
5. The method of claim 1, further comprising: providing an exception page associated with the data exception, the exception page including a filter parameter for presenting the data exception, the filter parameter including at least one of: a time range of the data exception, an application involved in the data exception, a region involved in the data exception, a data item type involved in the data exception, and a source involved in the data exception; and in response to receiving an interaction with respect to the filter parameter, updating the exception page.
6. The method of claim 1, wherein determining in the dataset the at least one cause field associated with the data anomaly comprises: determining the at least one cause field using an automated script, the automated script describing a mapping relationship between the key field and the at least one cause field; and the method further including: in response to determining that data in a cause field of the at least one cause field has an anomaly, providing an exception state associated with the key field and the cause field.
7. The method of claim 6, wherein providing an abnormal state associated with the keyword segment and the reason field comprises: in response to determining at least one of: a duration length of the exception satisfies a threshold time length, a magnitude change of the exception satisfies a threshold magnitude change.
8. The method of claim 1, further comprising: determining a potential exception field in the data set that is affected by the key field; and providing potential exception data associated with the potential exception field.
9. The method of claim 1, wherein the dataset is used to store data associated with client devices among a plurality of client devices, the plurality of fields including a first plurality of attributes of the client device, a second plurality of attributes of an application installed on the client device, a third plurality of attributes of data items published to the client device via the application, and a fourth plurality of events associated with the data items.
10. The method of claim 9, wherein determining the key fields in the dataset further comprises: Obtain the task to be performed in the dataset, the task being to determine the correlation between the fourth plurality of events; And among the multiple fields included in the dataset, identify the key fields associated with the task.
11. An apparatus for determining data anomalies, comprising: The field determination module is configured to determine key fields in a dataset. The set includes multiple data sources, and each of the multiple data sources includes at least one field; a status acquisition module is configured to acquire the historical status of the key field within a historical time period; an anomaly determination module is configured to determine that the key field has data anomalies in response to determining that the historical status indicates that the changes in the data in the key field meet the anomaly conditions, and the data anomalies indicate that the key field has anomalies within the historical time period. And a cause determination module, configured to determine in the dataset at least one cause field associated with the data anomaly, wherein the data anomaly of the at least one cause field causes the data anomaly of the key field.
12. An electronic device, comprising: At least one processing unit; And at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.
13. A computer-readable storage medium having a computer program stored thereon, the computer program causing the processor to implement the method according to any one of claims 1 to 10 when executed by a processor.
14. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Anomaly detection method and device of photovoltaic power generation tracking system, and storage medium
CN113708490A
User-level KQI anomaly detection using markov chain model
US20180285320A1