Data exception attribution method and device, equipment, storage medium and program product
By converting data anomaly detection results into domain-specific language descriptions and combining time and semantic analysis, the problem of data anomaly detection and attribution separation is solved, and accurate conversion and automated attribution from the data layer to the semantic layer are realized, improving analysis efficiency and accuracy.
Patent Information
- Application Number
- CN202510496107.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, data anomaly detection and attribution analysis are separated, making it difficult to correlate with other abnormal indicators, attribution analysis is poor, and lacks business semantic understanding, resulting in high response delays and false positive rates.
By obtaining target data, identifying exception points and transforming them into domain-specific language descriptions, analyzing related events, and generating attribution results in combination with time and semantic correlation analysis, it realizes accurate conversion and automated attribution from the data layer to the semantic layer.
The abnormal positioning time is shortened, from hourly to minutely, the false alarm rate is reduced, the accuracy and efficiency of attribution analysis is improved, the cognitive burden is reduced, and the abnormal detection and attribution is automated.
Smart Images

Figure CN120407257A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data anomaly detection, and specifically to a method, apparatus, device, storage medium, and program product for attributing data anomalies. Background Art
[0002] As the businesses involved by major enterprises become more extensive, in order to ensure the stable operation of the business and the timeliness of anomaly handling, business monitoring and business system operation and maintenance are becoming increasingly important. Currently, the anomaly detection and attribution analysis generated in business monitoring and system operation and maintenance are carried out independently, making it difficult to associate with other anomaly indicators, and the attribution analysis is mostly a shallow numerical conclusion, which affects the interpretability of the attribution analysis and results in poor attribution analysis effects. Summary of the Invention
[0003] In view of this, the present disclosure provides a method, apparatus, device, storage medium, and program product for attributing data anomalies to solve the problem of poor attribution analysis effects of data anomalies.
[0004] In a first aspect, the present disclosure provides a method for attributing data anomalies, including: obtaining target data to be detected; identifying anomaly points of the target data to generate an anomaly detection result; converting the anomaly detection result into a domain-specific language description; parsing the domain-specific language description to determine relevant events corresponding to the anomaly points; and analyzing the correlation between the relevant events and the anomaly points to generate an attribution result corresponding to the anomaly points.
[0005] In a second aspect, the present disclosure provides an apparatus for attributing data anomalies, including: a data acquisition module for obtaining target data to be detected; an anomaly identification module for identifying anomaly points of the target data to generate an anomaly detection result; a structure conversion module for converting the anomaly detection result into a domain-specific language description; a parsing module for parsing the domain-specific language description to determine relevant events corresponding to the anomaly points; and an attribution module for analyzing the correlation between the relevant events and the anomaly points to generate an attribution result corresponding to the anomaly points.
[0006] In a third aspect, the present disclosure provides a computer device, including: a memory and a processor, which are communicatively connected to each other, and the memory stores computer instructions, and the processor executes the computer instructions to execute the method for attributing data anomalies according to the first aspect or any corresponding implementation thereof.
[0007] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method for attributing data anomalies according to the first aspect or any corresponding implementation thereof.
[0008] Fifth aspect, the present disclosure provides a computer program product, including computer instructions for causing a computer to execute the method for attributing data anomalies according to the first aspect or any corresponding embodiment thereof as described above.
[0009] The method, apparatus, device, storage medium and program product for attributing data anomalies provided by the embodiments of the present disclosure, after identifying the anomaly points in the target data, can map the anomaly detection result to a structured semantic description by converting the anomaly detection result into a domain-specific language description, realizing an accurate conversion from the "data layer" to the "semantic layer", and solving the understanding gap between the anomaly data and the business semantics. By parsing the domain-specific language description to determine the relevant events corresponding to the anomaly points, a deep association between the anomaly points and other relevant events is realized, so as to accurately determine the main events that generate the anomaly points in combination with the correlation analysis between the relevant events and the anomaly points, and obtain the corresponding attribution results. Thereby, the automation from anomaly detection to the generation of attribution results is realized, the problem of the separation between anomaly detection and attribution is overcome, and the attribution analysis effect of data anomalies is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the related art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure;
[0012] Figure 2 is a flowchart of the method for attributing data anomalies according to an embodiment of the present disclosure;
[0013] Figure 3 is a schematic diagram of uploading target data according to an embodiment of the present disclosure; [[ID=2,2]]
[0014] Figure 4 is a flowchart of another method for attributing data anomalies according to an embodiment of the present disclosure;
[0015] Figure 5 is a flowchart of yet another method for attributing data anomalies according to an embodiment of the present disclosure;
[0016] Figure 6 is a schematic diagram of displaying the attribution result of CPI anomaly analysis according to an embodiment of the present disclosure;
[0017] Figure 7 is a block diagram of the structure of the apparatus for attributing data anomalies according to an embodiment of the present disclosure;
[0018] Figure 8 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. Detailed implementation manners
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0020] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0021] For example, when receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as a computer device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0022] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the computer device.
[0023] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0024] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations, and related provisions.
[0025] Currently, the core pain points faced by enterprises in business monitoring and system operation and maintenance are mainly as follows: (1) Fragmentation of the analysis process, that is, anomaly detection and root cause analysis belong to independent systems, and manual connection of data, algorithms, and business knowledge is required, resulting in a response delay of hours. For example, during large-scale e-commerce promotions, when there are abnormal fluctuations in the order volume, cross-team collaboration is needed to conduct investigations, and the golden processing window is missed. (2) Shallow causal association, that is, the anomaly attribution methods in related technologies only attribute based on temporal proximity, lacking in-depth reasoning about the event impact path. When the crash rate of a mobile application suddenly increases, it is difficult to distinguish whether it is caused by a version release error, a cloud service failure, or a surge in users. (3) Isolation of multi-source knowledge islands, that is, internal and external event data (system logs / market dynamics, etc.) are stored separately and cannot be associated with metric anomalies in real time. When the transaction failure rate of a certain financial App is abnormal, operation and maintenance personnel need to manually compare data from multiple platforms such as version release records and third-party payment interface statuses. (4) Weak explanation ability, that is, in related technologies, only pure numerical conclusions are output, lacking business semantic transformation. For example, "The P99 of the API response time has increased by 200 ms" cannot be directly associated with the business impact of "The checkout abandonment rate has increased by 1.2%". (5) Fragmented result presentation, that is, the attribution conclusion is separated from the original chart, increasing the cognitive burden. When the operation and maintenance supervisor views the server load report, he / she needs to repeatedly switch between the monitoring chart, the event list, and the analysis document to verify the hypothesis.
[0026] To address the above problems, currently, the following solutions are generally adopted: (1) Independent anomaly detection systems. For example, algorithms such as Isolation Forest and LSTM are used to identify data anomalies, but there is a lack of automatic docking with the event knowledge base. For example, an e-commerce platform uses Prometheus + AlertManager to monitor the order volume. After detecting an anomaly, manual login to Jira is required to query the concurrent change events, resulting in 30% of the anomalies not being able to find associated events within 1 hour. (2) Rule-based event matching. By presetting a time window (such as ±2 hours) to match operation and maintenance events, but lacking semantic reasoning ability. For example, a bank system forcibly associates the server expansion event with the abnormal transaction delay. In fact, the delay is caused by third-party API rate limiting, resulting in a 42% false alarm attribution. (3) Manual knowledge integration dashboard. Operation and maintenance personnel manually add event markers beside the Grafana chart, but the update is lagging and depends on personal experience. For example, the SRE team of a cloud service provider needs to view Datadog metrics, the Slack operation and maintenance channel, and the customer support system simultaneously. The time consumed for cross-system switching accounts for 65% of the total analysis time respectively. (4) Numerical report output. The anomaly analysis results are presented in numerical tables, lacking business semantic transformation. For example, the monitoring system of a logistics enterprise shows that "the sorting efficiency has decreased by 15%", but it does not associate with the traffic control event caused by heavy rain, resulting in the operation team misjudging it as a device failure.
[0027] The above solutions have problems such as mechanized timing correlation (ignoring the hysteresis of event impact in fixed time window matching), oversimplified causal determination (evaluating relevance only with Pearson correlation coefficient and unable to identify the chain effect of multiple events), passive knowledge application (relying on manual memory of historical event patterns and unable to identify new event types in a timely manner, such as the first cloud service provider failure), and unstructured conclusion expression (storing analysis results in non-standard text, hindering subsequent reuse).
[0028] Based on this, the technical solution of the present disclosure adopts a closed-loop mapping architecture of "anomaly point - DSL description - related events", which overcomes the problem of the separation between detection and attribution in existing anomaly analysis and realizes end-to-end automation from anomaly recognition to cause explanation. Through the above closed-loop mapping architecture, the anomaly location time is shortened from the hour level to the minute level. For example, the time-consuming for attributing abnormal e-commerce order volume is reduced from an average of 3 hours to 5 minutes, and the response speed is increased by 36 times. Combining dynamic time window matching with semantic knowledge graph, the false alarm rate is reduced by 52% (such as the misjudgment cases of bank transaction delay are reduced to 8%), the recognition accuracy of complex event chains reaches 89%, and the depth of causal reasoning is greatly enhanced. Realize the real-time association of internal and external events (system logs / market dynamics, etc.) with metric anomalies, and the cross-system data integration efficiency is increased by 70% (such as the data source switching required for abnormal analysis of financial App transactions is reduced by 83%), realizing the integration of multi-source knowledge. Through DSL semantic mapping technology, the business impact of "API response time P99 increases by 200ms" is automatically associated with "checkout conversion rate increases by 1.2%", and the accuracy of explanation generation is increased to 92%. The visualization injection technology enables the attribution conclusion to be integrated with the original chart by 95%, and the information retrieval path of operation and maintenance personnel is shortened by 60% (such as the analysis time of server load reports is reduced from 45 minutes to 18 minutes), reducing the cognitive burden. The structured DSL description improves the machine learning reuse rate of historical anomaly events (from 12% to 68%), and speeds up the recognition of new event types (such as the recognition time of the first cloud service failure is reduced from 24 hours to 4.8 hours).
[0029] As an optional application scenario of the embodiment of the present disclosure, as Figure 1 shown, this application scenario includes a computer device 102 corresponding to a data detection system 101. In Figure 1 , only a limited number of components and exemplary connection relationships are shown. It should be understood that this is for the purpose of convenience of description and easy illustration and is not intended to limit the scope of the present disclosure, and there may be other different components. For example, display components and input components, etc. In an exemplary rather than restrictive manner, the diagnostic results of the target page can be displayed on the display component, and the problem description and problem type can be adjusted through the input component.
[0030] According to an embodiment of the present disclosure, a technician can upload target data to be detected through an interaction page provided by the data detection system 101. Correspondingly, the data detection system 101 can perform operations such as anomaly identification, DSL conversion, event retrieval query generation, related event retrieval, attribution analysis, and visualization injection on the target data, so as to display the target data and the attribution result on the same interface.
[0031] Among them, the data detection system 101 is deployed in the computer device 102. The computer device 102 represents a device with computing resources or computing capabilities, and can be a device with computing capabilities. For example, the computer device can be provided with a processor and a memory, etc., and can also be equipped with a dedicated accelerator (such as a graphics processing unit (GPU), etc.). In addition, the computer device can store and maintain data.
[0032] Examples of the computer device 102 can include supercomputers, personal computers, laptop computers, in-vehicle computing devices, mobile devices (such as smartphones, tablets, etc.), or a combination of any one or more of the above devices. It should be understood that the computer devices described herein are only exemplary and not restrictive. For example, other different types of computer devices can also be used.
[0033] According to an embodiment of the present disclosure, an embodiment of a method for attributing data anomalies is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0034] In this embodiment, a method for attributing data anomalies is provided, which can be used in computer devices such as computers and tablets. Figure 2 is a flowchart of a method for attributing data anomalies according to an embodiment of the present disclosure, as Figure 2 shown, the process includes the following steps:
[0035] Step S201, obtain target data to be detected.
[0036] The target data is data generated during business operation, and specifically can include business monitoring data, system operation and maintenance data, etc. The target data can be represented by an SVG trend chart or in the form of a table. The form of the target data is not specifically limited here.
[0037] In a specific example, a data anomaly detection system is deployed in the computer device. The data anomaly detection system provides an interaction page, and a technician can upload the target data to be detected through the interaction page. Correspondingly, the computer device can obtain the target data uploaded by the user, such asFigure 3 as shown
[0038] In another specific example, the data anomaly detection system can be communicatively connected to the business system, and the business system can regularly send the data generated during its operation to the data anomaly detection system. Correspondingly, the data anomaly detection system in the computer device can receive this data as the target data for anomaly detection. The specific method for obtaining the target data is not specifically limited herein, and those skilled in the art can determine it according to actual needs.
[0039] Step S202: Identify the anomaly points of the target data and generate an anomaly detection result.
[0040] An anomaly point is a data point with anomalies in the target data, and this anomaly point is generated when an anomaly event occurs. The anomaly detection result is used to characterize the relevant anomaly information corresponding to the anomaly point, including the confidence level of the anomaly point and the anomaly characteristics of the anomaly point, etc. Specifically, multiple anomaly detection algorithms are integrated in the data anomaly detection system. After obtaining the target data, the target data is parsed to extract the time series information corresponding to the target data. Multiple anomaly detection algorithms are used to detect the anomaly points in the time series information, and the corresponding anomaly detection results are obtained.
[0041] In some specific examples, preprocessing such as data cleaning, handling missing values, standardization, or normalization is performed on the target data to obtain preprocessed data, and the time series information corresponding to the preprocessed data is extracted. Subsequently, an integrated detection method (such as the Bagging algorithm, Boosting algorithm, etc.) can be used to detect the anomaly points in the time series information; a distance-based method (such as K-nearest neighbor, isolation forest, etc.) can be used to detect the anomaly points in the time series information; a density-based method (such as local outlier factor, density-based spatial clustering, etc.) can be used to detect the anomaly points in the time series information; statistical values of each data point can also be calculated based on a statistical method, and the anomaly points in the time series information are detected using the statistical values; or the anomaly points in the time series information can be detected based on a pre-trained model (such as a neural network model, support vector machine, etc.). Of course, other methods can also be used, which are not specifically limited herein, as long as the anomaly points in the target data can be identified.
[0042] It should be noted that by combining multiple algorithms, the advantages of each can be mutually complemented to improve the accuracy and robustness of anomaly point detection.
[0043] Step S203: Convert the anomaly detection result into a domain-specific language description.
[0044] A Domain-Specific Language (DSL) is used to describe outliers by using a specific language in the business domain where the outliers are located. That is, the DSL description is adapted to the business domain where the outliers are located, solving the problem of the mismatch between abnormal data and business semantics.
[0045] Extract the features of the outliers in the anomaly detection results to obtain the corresponding anomaly features. Using the grammar rules of DSL (such as the grammar rules defined by Backus-Naur Form BNF), map the numerical anomaly detection results to structured semantic descriptions to achieve the conversion of anomaly detection results from the "data layer" to the "semantic layer".
[0046] Among them, the grammar rules of DSL satisfy the following key principles: structured expression, that is, using a JSON-compatible hierarchical structure; semantic completeness, that is, ensuring that all key features of the outliers can be expressed; business scalability, that is, supporting semantic extensions in different business domains; machine parseability, that is, facilitating subsequent automated processing.
[0047] Step S204, parse the domain-specific language description to determine the relevant events corresponding to the outliers.
[0048] Parse the semantics represented by the domain-specific language description, extract the key information carried by the domain-specific language description, use the key information to convert the domain-specific language description into event queries for different dimensions, and conduct targeted queries for each dimension in different time windows, achieving multi-dimensional and multi-time window event retrieval queries, and completing the conversion from anomaly features to event retrieval queries.
[0049] Using the event retrieval queries implemented above, conduct event queries in multiple event sources related to the outliers, and determine all potential events related to the abnormal occurrence time period. All these potential events are the relevant events corresponding to the outliers.
[0050] Step S205, analyze the correlation between the relevant events and the outliers, and generate the attribution result corresponding to the outliers.
[0051] The attribution result is used to characterize the cause of the anomaly and the event analysis associated with the outlier. When all relevant events associated with the outlier are determined, conduct time correlation analysis and semantic correlation analysis between each relevant event and the outlier. Combining time correlation and semantic correlation, determine the target event that causes the anomaly from the relevant events, and evaluate the determined target event, and generate the attribution result corresponding to the outlier in combination with the evaluation result.
[0052] The method for attributing data anomalies provided in this embodiment can, after identifying anomaly points in target data, map the anomaly detection results to structured semantic descriptions by converting the anomaly detection results into domain-specific language descriptions, achieving an accurate conversion from the "data layer" to the "semantic layer" and solving the understanding gap between anomaly data and business semantics. By parsing the domain-specific language descriptions to determine relevant events corresponding to the anomaly points, a deep association between the anomaly points and other relevant events is achieved, so as to accurately determine the main events that generate the anomaly points in combination with the correlation analysis between the relevant events and the anomaly points, and obtain the corresponding attribution results. Thus, the automation from anomaly detection to the generation of attribution results is realized, overcoming the problem of the separation between anomaly detection and attribution, and improving the attribution analysis effect of data anomalies.
[0053] In this embodiment, a method for attributing data anomalies is provided, which can be used in computer devices such as computers, tablets, etc. Figure 4 is a flowchart of the method for attributing data anomalies according to an embodiment of the present disclosure, as Figure 4 shown, and the process includes the following steps:
[0054] Step S301, obtain target data to be detected. For details, please refer to the relevant descriptions of the corresponding steps in the above-mentioned embodiments, and details will not be repeated here.
[0055] Step S302, identify anomaly points in the target data and generate anomaly detection results. For details, please refer to the relevant descriptions of the corresponding steps in the above-mentioned embodiments, and details will not be repeated here.
[0056] Step S303, convert the anomaly detection results into domain-specific language descriptions.
[0057] Specifically, the above step S303 includes:
[0058] Step S3031, extract anomaly key features from the anomaly detection results.
[0059] Anomaly key features are the features that anomaly points have compared with normal data points, specifically including basic information (such as index name, timestamp corresponding to the anomaly point, etc.), statistical features (such as Z-score value, IQR deviation, etc.), pattern features (such as mutation type, duration, etc.), and related indicators (i.e., other indicators that are simultaneously abnormal). Specifically, by parsing each data field in the anomaly detection results, the field information containing anomaly key features can be determined, and then the corresponding anomaly key features can be extracted from the determined field information.
[0060] Step S3032, perform feature conversion on the anomaly key features to obtain anomaly information corresponding to the anomaly key features.
[0061] Anomaly information is semantic information formed for the key features of an anomaly. Specifically, by performing feature transformation on the extracted key features of the anomaly according to pre-set grammar rules, the key features of the anomaly can be transformed into anomaly information described using DSL.
[0062] In some alternative embodiments, the anomaly information includes anomaly description information, context information, and associated metrics. Among them, the anomaly description information represents the core anomaly description of the anomaly point, including basic information such as the metric name (METRIC), time point (TIMESTAMP), anomaly type (TYPE), severity (MAGNITUDE), etc.; the context information includes a text (CONTEXT) object for storing context such as the data before and after the anomaly point, expected range, rate of change, etc.; the associated metrics are used to describe other metrics or potential factors related to the anomaly of the anomaly point, and the associated metrics can be represented in the form of an associated array (RELATED array).
[0063] Correspondingly, step S3032 described above includes:
[0064] Step a1: Perform semantic mapping on the numerical features in the key features of the anomaly using a pre-set mapping rule to obtain the anomaly description information corresponding to the numerical features.
[0065] Step a2: Perform feature fusion on the non-numerical features in the key features of the anomaly using a pre-set sliding window to determine the context information corresponding to the non-numerical features.
[0066] Step a3: Analyze the anomaly metrics in the key features of the anomaly to determine the associated metrics related to the anomaly metrics.
[0067] The pre-set mapping rule is a rule pre-set for semantic mapping of numerical features, and this pre-set mapping rule is one of the rules in the grammar rules of DSL. Specifically, as shown in Table 1, different semantic mapping conditions can be set for different numerical features, and the corresponding semantic descriptions (i.e., anomaly description information) can be obtained.
[0068] Table 1 Numerical Feature Mapping
[0069] Numerical feature Semantic mapping condition Semantic description Relative change rate <-20% "severe drop" Relative change rate -20%~-10% "moderate drop" Z-score <-3.0 "severe anomaly" Z-score -3.0~-2.0 "moderate anomaly" Duration <1 hour "transient" Duration > 24 hours "persistent"
[0070] The preset sliding window is a time window preset for the generation time of outliers, such as 7 days before and after the occurrence time of the outlier. Feature fusion is performed on the historical data, expected range, seasonal features, etc. associated with the abnormal key features. The coefficient of variation CV is used to calculate the historical volatility corresponding to the outlier. The seasonal adjustment value corresponding to the outlier is determined by the STL decomposition method, and the expected value interval corresponding to the outlier is calculated based on the historical data of the same period. The calculated historical volatility, seasonal adjustment, and expected value interval are used as the context information corresponding to the outlier.
[0071] Since the occurrence of an outlier may cause other indicators to become abnormal, at this time, other indicators that are abnormal simultaneously can be detected at the occurrence time of the outlier. Correlation analysis is performed between the abnormal indicator corresponding to the outlier and the other detected indicators to determine the index correlation degree between the abnormal indicator and the other indicators. Other indicators whose index correlation degree exceeds the preset correlation degree (such as 0.7) are determined as the associated indicators associated with the abnormal indicator.
[0072] In the above embodiments, by performing feature transformation on the abnormal key features, abnormal description information, context information, and associated indicators in terms of semantics for the abnormal key features are obtained, ensuring that the subsequent generated DSL description can cover all key features of the outlier, improving the generation accuracy of the DSL description, and ensuring that the DSL description can accurately represent the relevant abnormal information of the outlier.
[0073] Step S3033, assemble the abnormal information according to the syntax assembly rules described in the domain-specific language to obtain the domain-specific language description corresponding to the abnormal detection result.
[0074] As described above, the DSL has corresponding syntax rules, and the syntax rules also include syntax assembly rules, which are preset rules for assembling various abnormal information. Assemble the above-obtained abnormal information according to the syntax assembly rules of the DSL to obtain a complete DSL description. Specifically, when assembling the abnormal information, first fill in the core abnormal description, that is, the abnormal description information generated above; secondly, add the context information, that is, splice the context information generated above into the abnormal description information; thirdly, associate the relevant indicators, that is, associate the generated associated indicators with the abnormal indicator corresponding to the outlier to realize the correlation analysis of the abnormal event.
[0075] In some optional embodiments, after generating the DSL description, the compliance verification of the DSL description can be performed to verify whether the currently generated DSL description conforms to the syntax rules and business constraints. If the verification fails, the DSL description is adjusted to make it conform to the syntax rules and business constraints.
[0076] Step S304: Parse the domain-specific language description and determine the relevant events corresponding to the anomaly points.
[0077] Specifically, the above step S304 includes:
[0078] Step S3041: Use the pre-trained retrieval conversion model to convert the domain-specific language description into an event retrieval query.
[0079] The retrieval conversion model is a pre-trained model with DSL understanding ability and specific domain knowledge understanding ability. Specifically, this retrieval conversion model can be trained based on the large language model architecture and fine-tuned with business domain data to enhance the model's understanding of industry terms and knowledge in a specific domain on the basis of its semantic understanding ability.
[0080] Input the domain-specific language description into the retrieval conversion model, and convert the DSL description into an event retrieval query with multiple dimensions and multiple time windows through the retrieval conversion model to retrieve events from multiple event sources.
[0081] In some alternative embodiments, the above step S3041 includes:
[0082] Step b1: Parse the structure of the domain-specific language description, extract the key description information, and construct the target context information corresponding to the key description information.
[0083] Step b2: Use the key description information and the target context information to construct event query information from multiple dimensions.
[0084] Step b3: Obtain the preset time window, and use the target context information, the event query information, and the preset time window to construct the prompt description information.
[0085] Step b4: Use the prompt description information to guide the retrieval conversion model to convert the domain-specific language description into an event retrieval query.
[0086] Parse the syntax structure of the DSL description, extract the corresponding key description information from it, and perform semantic parsing of the context information according to the key description information to obtain the target context information related to the key description information.
[0087] According to the target context information and the key description information, construct the event query information related to the anomaly point, such as the direct impact event query information, the competitor association query information, the technology association event query information, the industry background query information, the historical pattern query information, etc.
[0088] The preset time window is a predefined critical time window, such as the core time window (24 hours before and after the anomaly point, i.e., ±24h), the pre-time window (3 to 5 days before the anomaly, i.e., -5 days to -1 day), the continuation time window (1 to 3 days after the anomaly, i.e., +1 day to +3 days), etc.
[0089] Use the target context information, event query information, and preset time window to construct prompt description information to guide the retrieval conversion model to convert the DSL description into an event retrieval query. Specifically, the target context information, event query information for each dimension, and preset time window can be used to construct prompt description information to guide the retrieval conversion model to generate corresponding event retrieval queries for each dimension according to the DSL description. It is also possible to use the target context information, multi-dimensional event query information, and preset time window to construct prompt description information to guide the retrieval conversion model to generate multi-dimensional event retrieval queries simultaneously according to the DSL description. No specific limitation is made here, and those skilled in the art can determine it according to the actual business scenario.
[0090] In a specific example, for the scenario of the daily active users of a mobile application decreasing, the multi-dimensional query examples generated are as follows:
[0091] Direct impact event query:
[0092] "Technical events or system changes that affected the daily active users of the mobile application around March 15, 2023";
[0093] time_range: 2023-03-14T00:00:00Z to 2023-03-16T00:00:00Z.
[0094] Competitor association query:
[0095] "New features, promotional activities, or marketing events launched by Competitor X or Competitor Y from March 13 to 17, 2023";
[0096] time_range: 2023-03-13T00:00:00Z to 2023-03-17T00:00:00Z.
[0097] Technical association event query:
[0098] "Mobile operating system (iOS / Android) update or network service interruption events from March 13 to 16, 2023";
[0099] time_range: 2023-03-13T00:00:00Z to 2023-03-16T00:00:00Z.
[0100] Historical mode query:
[0101] "Typical event cases in history that led to a sudden drop of more than 20% in the daily active users of mobile applications";
[0102] time_range: any.
[0103] In the above embodiments, by constructing an event retrieval query for DSL description from multiple dimensions and multiple time windows, the conversion from abnormal key features to event retrieval queries is achieved, so as to retrieve relevant events from multiple dimensions and multiple time windows, ensuring that all potential events related to the abnormal points can be retrieved, improving the depth of anomaly detection, and facilitating the determination of the root cause of the anomaly.
[0104] Step S3042: Based on the event retrieval query, retrieve the events that occurred within the target time window and determine the relevant events corresponding to the abnormal points.
[0105] Using the above-constructed multi-dimensional and multi-time window event retrieval query, retrieve all the events that occurred within a specific time window for multiple event sources corresponding to the abnormal points, obtain multiple candidate events, calculate the correlation between each candidate event and the abnormal point, and determine the relevant events associated with the abnormal point therefrom, so as to make the causal relationship between the abnormal point and its relevant events more comprehensive and accurate.
[0106] In some alternative embodiments, the above step S3042 includes:
[0107] Step c1: Obtain multiple event sources corresponding to the abnormal points.
[0108] Step c2: Use the event retrieval query to retrieve events from multiple event sources and obtain the relevant events within the target time window.
[0109] Since the reasons for generating abnormal points can usually be traced back to certain events that occurred within a specific time window, and there will be a temporal sequence and semantic association between these events and the anomaly. Moreover, the combination of multiple events may also jointly cause the generation of abnormal points. Therefore, it is necessary to retrieve events from multiple event sources.
[0110] The event sources include various events generated during the business operation process. Specifically, the multiple event sources may include internal event sources, external event sources, and historical event sources. Among them, the internal event sources include system release records, configuration change logs, and operation activity records, etc.; the external event sources include industry news, market dynamics, competitor activities, social media, etc.; the historical event sources include the mapping database of historical anomalies and events. Multiple event sources are built into the data detection system, and when retrieving relevant events for abnormal points, the multiple event sources corresponding to the abnormal points can be obtained.
[0111] The target time window is the key time window corresponding to the anomaly point, that is, the events occurring within the target time window are potential events leading to the anomaly. Specifically, the target time window can be one or more of the core time window, the lead time window, and the continuation time window described above, or can be other time windows set according to the business scenario, which is not specifically limited here.
[0112] Specifically, when performing event searches on multiple event sources, corresponding search parameters are constructed for the event retrieval queries, and each event retrieval query is retrieved in each target time window according to the search parameters to obtain at least one relevant event associated with the anomaly point.
[0113] In the above embodiment, by retrieving relevant events associated with the anomaly point in multiple event sources, the integration of multi-source knowledge is achieved, ensuring the comprehensive retrieval of relevant events and improving the retrieval accuracy of relevant events.
[0114] Step S305, analyze the correlation between the relevant events and the anomaly point, and generate an attribution result corresponding to the anomaly point. For details, please refer to the relevant description of the corresponding steps in the above embodiment, which will not be repeated here.
[0115] The anomaly attribution method for data provided in this embodiment extracts the anomaly key features in the anomaly detection result to convert the anomaly key features into anomaly information at the semantic level, and then assembles the anomaly information according to the grammar assembly rules described in the domain-specific language to obtain the corresponding DSL description, thereby realizing the conversion of the anomaly key features from the data layer to the business domain semantic layer, realizing the structured description of the anomaly point, which is beneficial to the reuse of anomaly attribution and anomaly recognition of subsequent relevant anomaly events. By retrieving the conversion model to convert the domain-specific language description into an event retrieval query to retrieve the events occurring within the target time window, the comprehensiveness of the retrieved relevant events is ensured, facilitating the real-time anomaly association of the anomaly point with the relevant event metrics in each event source, which is beneficial to improving the accuracy of anomaly positioning and avoiding the influence of event retrieval isolation on the attribution accuracy.
[0116] In this embodiment, an anomaly attribution method for data is provided, which can be used in computer devices such as computers and tablets. Figure 5 is a flowchart of the anomaly attribution method for data according to an embodiment of the present disclosure, as Figure 5 shown, the process includes the following steps:
[0117] Step S401, obtain the target data to be detected. For details, please refer to the relevant description of the corresponding steps in the above embodiment, which will not be repeated here.
[0118] Step S402: Identify the abnormal points of the target data and generate an anomaly detection result. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be elaborated here.
[0119] Step S403: Convert the anomaly detection result into a domain-specific language description. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be elaborated here.
[0120] Step S404: Parse the domain-specific language description and determine the relevant events corresponding to the abnormal points. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be elaborated here.
[0121] Step S405: Analyze the correlation between the relevant events and the abnormal points and generate an attribution result corresponding to the abnormal points.
[0122] Specifically, the above Step S405 includes:
[0123] Step S4051: Conduct a temporal correlation analysis on the abnormal points and the relevant events to determine the temporal correlation degree.
[0124] Temporal correlation represents the temporal relevance between the occurrence of an anomaly and its relevant events; the temporal correlation degree is used to characterize the temporal correlation between the abnormal points and the relevant events. By analyzing the temporal correlation between the abnormal points and their relevant events, the abnormal events that may lead to the generation of the abnormal points are determined. Specifically, the occurrence time of the abnormal points can be matched with the occurrence time of the relevant events in chronological order, the influence of different types of relevant events on the generation of the abnormal points can be evaluated, different weights can be assigned to different chronological orders and different types of relevant events, and the temporal correlation degree between the abnormal points and their relevant events can be determined by using a weighted method.
[0125] In some alternative embodiments, the above Step S4051 includes:
[0126] Step d1: Obtain the abnormal occurrence time corresponding to the abnormal points and the event occurrence time of the relevant events, and determine the time difference between the abnormal occurrence time and the event occurrence time.
[0127] Step d2: Based on the preset time difference range where the time difference is located, determine the temporal correlation weights of different types of relevant events.
[0128] Step d3: According to the temporal correlation weights and the time difference, determine the temporal correlation degree between the relevant events and the abnormal points.
[0129] The generation of abnormal points corresponds to corresponding timestamps. When an abnormality is detected, the abnormal points will be recorded, and the timestamp of the occurrence of the abnormality will be recorded synchronously. The time represented by this timestamp is the time of the occurrence of the abnormality. Since the occurrence of related events is related to the generation of abnormal points, the event occurrence time of related events can be collected to perform a temporal correlation analysis on related events and abnormal points.
[0130] Specifically, calculate the time difference between the abnormal occurrence time T1 of the abnormal point and the event occurrence time T2 of the related event, that is, time_diff = T1 - T2.
[0131] The preset time difference range is a pre-set time difference range. Different time difference ranges have different degrees of influence on the generation of abnormalities. Within the same time difference range, different types of related events also have different degrees of influence on the generation of abnormalities. Therefore, by combining the preset time difference range where the time difference is located, different time-related weights can be set for different types of related events. Subsequently, the event correlation degree of the abnormal point and its related events is calculated according to the time-related weight and the time difference.
[0132] In a specific example, if the preset time difference range where the time difference is located is (0, 24h), that is, the related event occurs within 24h before the abnormality. At this time, obtain the time-related weight lag_weight of each related event within the time difference range (0, 24h) that affects the generation of the abnormal point. Subsequently, combine the time difference and the time-related weight lag_weight to calculate the time correlation degree temporal_score. The specific calculation method is as follows:
[0133] temporal_score = math.exp(-abs(time_diff) / (24 * lag_weight)).
[0134] Among them, abs(time_diff) represents calculating the absolute value of the time difference time_diff; 24 * lag_weight is used to convert the time difference into a proportion within one day; math.exp(-…) is the natural exponential function exp. The negative sign and dividing by 24 * lag_weight indicate that as the time difference increases, the time correlation degree will decrease exponentially.
[0135] If the preset time difference range where the time difference is located is (-12h, 12h), that is, the related event occurs within 12h after the abnormality. The related events occurring within this time difference may be influencing measures taken for the abnormal point. At this time, the time correlation degree temporal_score can be calculated in combination with the time difference. The specific calculation method is as follows:
[0136] temporal_score = C * math.exp(-abs(time_diff) / 12).
[0137] Among them, abs(time_diff) represents the absolute value of the calculated time difference time_diff; math.exp(-…) is the natural exponential function; C is a coefficient, such as 0.5.
[0138] If the preset time difference range where the time difference is located is relatively large, it means that the time is far apart. At this time, the temporal relevance temporal_score can be calculated in a linearly decaying manner. Taking 1 week as an example, its specific calculation method is as follows:
[0139] temporal_score = max(0, 1 - abs(time_diff) / (7 * 24)).
[0140] Among them, abs(time_diff) represents the absolute value of the calculated time difference time_diff; max() is to take the maximum value.
[0141] In the above-mentioned implementation manner, through the temporal relevance analysis of the abnormal points and related events, in order to determine the related events that cause the abnormal points from the temporal relevance, the correlation analysis between the related events and the abnormal points in time is realized.
[0142] Step S4052, perform semantic relevance analysis on the abnormal points and related events to determine the semantic relevance.
[0143] Semantic relevance represents the relevance in semantic description between the abnormal points and their related events; semantic relevance is used to characterize the semantic relevance between the abnormal points and related events. By analyzing the relevance in semantic description of the abnormal points and their related events, the abnormal events that may cause the abnormal points are determined. Specifically, the abnormal description of the abnormal points can be subject-matched with the event description of the related events, the possible paths of the event impact indicators can be analyzed, and matched with the known abnormalities in the historical database. Combining the subject-matching results, path analysis results and historical matching results, the semantic relevance between the abnormal points and their related events is determined.
[0144] In some alternative implementation manners, the above-mentioned step S4052 includes:
[0145] Step e1, perform semantic matching on the first semantics corresponding to the abnormal points and the second semantics corresponding to the related events to determine the semantic relevance.
[0146] Step e2, analyze the index impact path of the related events according to the preset knowledge graph to determine the path relevance.
[0147] Step e3: Retrieve historical abnormal events, match the relevant events with the historical abnormal events, and determine the relevance of the historical abnormal events.
[0148] The first semantic representation is the abnormal description for the abnormal point; the second semantic representation is the event description for the relevant event. Extract the key features in the first semantic and the second semantic respectively, and perform semantic matching on the key features corresponding to the first semantic and the key features corresponding to the second semantic to obtain the semantic relevance between the abnormal point and its relevant events. In a specific example, the key features corresponding to the first semantic and the key features corresponding to the second semantic can be mapped to the same vector space, the key features corresponding to the first semantic and the key features corresponding to the second semantic are converted into feature vectors, and the vector distance between the feature vectors is calculated. This vector distance represents the semantic relevance. Of course, other methods can also be used for calculation, which is not specifically limited here.
[0149] The preset knowledge graph is a knowledge graph preset for the current business domain, and this knowledge graph includes different types of relevant events and the metrics they are associated with. Therefore, the relevant events can be matched with the preset knowledge graph, and the possible paths of the event affecting the metrics can be analyzed according to the knowledge graph. Analyze the metric impact path of the relevant events according to the preset knowledge graph to determine the path relevance.
[0150] For the historical abnormal events that have occurred, there is a corresponding case library, and each historical abnormal event that has occurred corresponds to a corresponding abnormal attribution result. Match the relevant events associated with the current abnormal point with the historical abnormal events in the historical case library to determine whether the occurrence of the current abnormal point occurred within the historical time. Thus, by combining the matching results of the relevant events and the historical abnormal events, the relevance of the historical abnormal events can be determined.
[0151] According to the influence degrees of the semantic relevance, path relevance, and historical abnormal event relevance on the abnormal point, set corresponding weights. Then, perform weighted summation on the semantic relevance, path relevance, and historical abnormal event relevance to obtain the corresponding semantic correlation degree.
[0152] In a specific example, if the weight corresponding to the semantic relevance is 0.4, the weight corresponding to the path relevance is 0.3, and the weight corresponding to the historical abnormal event relevance is 0.3, the calculation method of the semantic correlation degree is as follows:
[0153] semantic_score = 0.4 * keyword_match_score + 0.3 * path_score + 0.3 * historical_score.
[0154] Among them, semantic_score represents semantic relevance; keyword_match_score represents semantic correlation; path_score represents path correlation; historical_score represents historical anomaly event correlation.
[0155] In the above embodiments, by performing semantic correlation analysis on anomaly points and related events, relevant events that cause the generation of anomaly points are determined from the deep semantic relevance, realizing the association analysis of relevant events and anomaly points in terms of semantics.
[0156] Step S4053, generate an attribution result based on the fusion result of time correlation and semantic correlation.
[0157] Combine the time correlation and semantic correlation for weighted processing to obtain a weighted processing result. According to the weighted processing result, determine the most relevant event that causes the generation of the anomaly point from multiple relevant events, and perform anomaly attribution according to the most relevant event to generate an attribution result.
[0158] In some optional embodiments, the above step S405 further includes:
[0159] Step f1, obtain anomaly evaluation factors corresponding to the attribution result, where the anomaly evaluation factors include at least one of multi-source verification information, rule matching degree, and historical statistical correlation.
[0160] Step f2, use the anomaly evaluation factors to perform attribution evaluation on the anomaly point to generate an attribution evaluation result.
[0161] Multi-source verification information is used to represent how many independent sources have verified the determined attribution result; the rule matching degree is used to represent the matching degree with the causal rules in the domain knowledge base; the historical statistical correlation is used to represent the statistical correlation between the relevant events and index changes in the attribution result and the historical data.
[0162] To further evaluate the evidence strength of the attribution result, place the attribution result generated for the anomaly point into multiple event sources for anomaly verification, and determine the number of anomaly verification event sources passed by the attribution result. At the same time, match the attribution result with the causal rules in the current business domain knowledge base to determine the knowledge matching degree between the two. Among them, the knowledge matching degree can be calculated through vector distance and can also be calculated through cosine similarity, which is not specifically limited here. Further, perform mathematical statistics on the relevant events and their index changes represented in the attribution result and the situations of such events and index changes occurring in the historical data, and determine the historical statistical correlation of the attribution result in combination with the mathematical statistical results.
[0163] According to the influencing degrees of abnormal evaluation factors such as multi-source verification information, rule matching degree, and historical statistical correlation on the attribution result, corresponding weights are set. Then, the abnormal evaluation factors such as multi-source verification information, rule matching degree, and historical statistical correlation are weighted and summed to obtain the corresponding weighted result, and this weighted result represents the attribution evaluation result.
[0164] In a specific example, if the weight corresponding to the multi-source verification information is 0.4, the weight corresponding to the rule matching degree is 0.4, and the weight corresponding to the historical statistical correlation is 0.2, the determination method of the attribution evaluation result is specifically as follows:
[0165] evidence_score = 0.4 * source_score + 0.4 * rule_match_score + 0.2 * stat_score.
[0166] Among them, evidence_score is the weighted result of the abnormal evaluation factor and is used to represent the attribution evaluation result; source_score represents the multi-source verification information; rule_match_score represents the rule matching degree; stat_score represents the historical statistical correlation.
[0167] In a specific example of attribution analysis for the scenario of the daily active users of a mobile application decreasing, the specific attribution analysis is as follows:
[0168] (1) Retrieve relevant events corresponding to the abnormal points:
[0169] Event A: "Release of v2.5.0 version" (internal event, 03-14 23:30)
[0170] Event B: "Android system security update" (external event, 03-15 02:15)
[0171] Event C: "Competitor A launches a time-limited offer" (competitor event, 03-15 09:00)
[0172] Event D: "Server expansion" (internal event, 03-15 14:30)
[0173] (2) Time correlation analysis:
[0174] Event A: 0.92 (occurred 4.5 hours before the anomaly)
[0175] Event B: 0.85 (occurred 2 hours before the anomaly)
[0176] Event C: 0.60 (time is close but after the user active peak period)
[0177] Event D: 0.20 (Occurs after an anomaly and may be a response measure)
[0178] (3) Semantic relevance analysis:
[0179] Event A: 0.88 (Highly relevant to version release and application performance)
[0180] Event B: 0.75 (System updates may affect application compatibility)
[0181] Event C: 0.45 (Competitor activities usually have an indirect impact)
[0182] Event D: 0.15 (Server expansion usually does not lead to a decrease in daily active users)
[0183] (4) Attribution result evaluation:
[0184] Event A: 0.90 (Confirmed by multiple internal logs and there are historical cases)
[0185] Event B: 0.80 (There are official announcements and user feedback)
[0186] Event C: 0.40 (Only market monitoring data)
[0187] Event D: 0.30 (Only internal records)
[0188] (5) Comprehensive attribution result:
[0189] Main reason: "The incompatibility between application version v2.5.0 and Android security updates leads to an increase in the crash rate" (87% confidence)
[0190] Auxiliary evidence: The crash rate increased by 150% during the same period, and there were a large number of crash complaints on social media.
[0191] In the above implementation manner, by evaluating the attribution result, it is beneficial to improve the evidence strength of the attribution result and make the attribution result more explanatory.
[0192] Step S406, use the pre-trained annotation generation model to generate the annotation content corresponding to the attribution result.
[0193] The annotation generation model is a pre-trained model with semantic understanding ability, and this annotation generation model can be trained based on the model architecture of a large language model. The annotation content represents the natural language explanation generated for the attribution result. Input the attribution result into the annotation generation model, and through the semantic understanding of the attribution result by the annotation generation model, generate the corresponding annotation content.
[0194] Specifically, the above step S406 may include:
[0195] Step A1: Obtain the pre - constructed prompt description information.
[0196] The prompt description information is preset for annotation generation and is used to guide the annotation generation model to output annotation content that meets the requirements. Specifically, technicians can construct a prompt template according to actual needs, fill in the corresponding input parameters, output parameters, and prompt words in the prompt template, and then generate the corresponding prompt description information. Correspondingly, when annotation content generation is required, the prompt description information can be directly obtained.
[0197] Step A2: Use the prompt description information to guide the annotation generation model to compress and transform the attribution result, and generate at least one level of annotation content.
[0198] Call the annotation generation model using the obtained prompt description information, so that the annotation generation model compresses and transforms the attribution result according to the prompt description information to obtain multiple levels of annotation information to meet the requirements of different business scenarios. Among them, the multiple levels of annotation information can include brief prompt information (e.g., within 20 words), basic explanation information (such as 50 - 100 words), detailed analysis, etc., which are not specifically limited here.
[0199] In the above - mentioned embodiment, by generating annotation content for the attribution result, semantic transformation in the business domain is achieved, and the interpretability of the attribution result is improved compared with pure numerical conclusions.
[0200] In some alternative embodiments, step S406 may further include:
[0201] Step B1: In response to a display trigger operation for the annotation content, display the brief information of the annotation content.
[0202] The brief information is used to represent the brief prompt of the annotation content, and the brief information has corresponding character limitations, such as not exceeding 20 words. The display trigger operation is an operation triggered by the user to display the annotation content, such as hover display.
[0203] Specifically, when the user triggers a display for the mark corresponding to the annotation content, the annotation content will be displayed in the form of brief information. For example, when the mouse is hovered over the mark of the annotation content, the brief information of the annotation content will be displayed.
[0204] Step B2: In response to a click trigger operation for the annotation content, display the explanatory information of the annotation content.
[0205] The explanatory information is used to represent the basic summary information of the annotation content, and the explanatory information also has corresponding character limitations, such as 50 - 100 words. The click trigger operation is an operation triggered by the user to view the annotation content, such as directly displaying after clicking.
[0206] Specifically, when the user clicks on the marker corresponding to the annotation content, the annotation content will be directly displayed in the form of explanatory information. For example, by clicking on the marker of the annotation content with the mouse, the explanatory information for the annotation content will be directly displayed.
[0207] Step B3, in response to a viewing trigger operation for the annotation content, display the detailed information of the annotation content.
[0208] The detailed information is used to represent all the attribution details of the annotation content, and the length of the characters of the detailed information is not limited. The viewing trigger operation is an operation triggered by the user to view the details of the annotation content. For example, a viewing control for the detailed information is set, and after the user clicks on the viewing control, the attribution result details are directly displayed.
[0209] Specifically, when the user clicks on the marker corresponding to the annotation content, the annotation content will be directly displayed in the form of explanatory information, and at the same time, the corresponding viewing control will be displayed. If the user wants to view the attribution details, clicking on the viewing control will display the attribution details for the annotation content.
[0210] In the above embodiments, by supporting the generation of annotation content at different levels, it is beneficial to improve the user experience of viewing the annotation content.
[0211] In some alternative embodiments, the display of the annotation content also supports expand / collapse interaction animations to achieve flexible display of the annotation content.
[0212] In some alternative embodiments, the visual display of the annotation content also supports animation display. For example, when viewing the annotation content, the corresponding evidence chain and related events are automatically played to achieve interactive exploration between the evidence chain and the related events, enhancing the user experience.
[0213] Step S407, analyze the data structure of the target data to determine the annotation position of the annotation content in the target data.
[0214] The annotation position is the display position of the annotation content; the data structure represents each data element that makes up the target data. Since different types of target data have different data structures (for example, target data in the form of an SVG diagram has a corresponding diagram structure, and tabular data has a corresponding table layout, etc.), it is necessary to parse each data element that makes up the target data, locate its display position in the target data in combination with the occurrence time of the anomaly point, and calculate the annotation position for displaying the annotation content in combination with the element position of each data element in the target data to avoid obscuring the data elements.
[0215] Specifically, the above step S407 includes:
[0216] Step C1: Analyze the data structure of the target data to determine the element positions of each data element in the target data.
[0217] Parse the data structure of the target data to identify the data elements that make up the target data, such as trend path elements, point elements, axis elements, cell elements, etc. Construct a coordinate system for the target data to determine the coordinate positions of each data element in this coordinate system, and this coordinate position is the element position.
[0218] Taking an SVG document as an example, by parsing the SVG document, identify the axis elements in the SVG document, and parse the time scale and value scale to determine the positions of each data point element and the path position of the path formed between two points.
[0219] Step C2: Based on the abnormal time of the abnormal point, locate the target position where the abnormal point is located.
[0220] The target position is the position where the abnormal point where the abnormality occurs is located. Combining the timestamp of the abnormal point can determine the time when the abnormality occurs. Thus, by parsing the time scale position in the target data in the constructed coordinate system, the target position of the abnormal point can be located.
[0221] Step C3: Parse the element position and the target position to determine the annotation position of the annotation content. The annotation position is associated with the target position and does not coincide with the element position.
[0222] Combining the coordinate information of the element position and the coordinate information of the target position, calculate the coordinate area that will not obscure the element position and is closest to the target position in the way of linear programming. The position surrounded by this coordinate area can be used as the annotation position for displaying the annotation content.
[0223] In the above embodiment, by combining the element positions corresponding to each data element and the target position corresponding to the abnormal point, the annotation position is determined, ensuring that the annotation position will not obscure each data element and can be relatively close to the position of the abnormal point, ensuring the best annotation position.
[0224] Step S408: Inject the annotation content into the annotation position for visual display.
[0225] Create an annotation element at the annotation position and associate the annotation element with the annotation content, so as to inject the annotation content into the annotation position corresponding to the abnormal point through the annotation element, realizing the visual display of the annotation content in the target data. As Figure 6 shown, the visual display of the attribution result generated by the data detection system for the abnormal point of CPI.
[0226] Specifically, the above step S408 includes:
[0227] Step D1, create annotation elements corresponding to the annotation content. Among them, the annotation elements include at least one of marker elements, connection line elements, annotation box elements, and text interaction elements.
[0228] The marker element is a marker for abnormal points, such as a circle, a triangle, etc.; the annotation box element is used to display the annotation content, such as a rounded rectangle; the connection line element is a connection line between the marker element and the annotation box element, such as a Bezier curve; the text interaction element represents expanding or collapsing the annotation content.
[0229] Specifically, after generating the annotation content, the data detection system can create corresponding annotation elements for the annotation content to generate corresponding annotation elements at the annotation position.
[0230] Step D2, associate the annotation content with the annotation elements, and inject the annotation elements into the annotation positions corresponding to the target data to visually display the annotation content through the annotation elements.
[0231] Associate the annotation content with each of the above-created annotation elements, and then the annotation content can be injected into the corresponding annotation elements. Subsequently, inject the created annotation elements into the annotation positions to display the annotation content at the annotation positions.
[0232] In the above embodiments, by creating annotation elements, the annotation content is injected into the annotation positions, realizing the visual display of the annotation content in the target data, avoiding the fragmentation of the presentation of the attribution results, avoiding the repeated switching between the target data and the attribution results, and realizing the synchronous presentation of the target data and the attribution results.
[0233] An example of the implementation of SVG injection and interaction in the scenario of the daily active users of a mobile application decreasing is as follows:
[0234] (1) Abnormal point marking: Add a red circle with a radius of 8px at the data point position on X month X day; set the marker transparency to 80%, corresponding to the abnormal confidence level; add a pulse animation effect to the marker to highlight it.
[0235] (2) Annotation content generation: Brief information: "The daily active users decreased by 21.9%"; Explanation information: "The new version of the application is incompatible with the Android system update, resulting in a 150% increase in the crash rate, seriously affecting user activity"; Detailed analysis: Include the complete event timeline, evidence chain, and recommended measures.
[0236] (3) Interaction behavior implementation: When the mouse hovers over the abnormal point, the brief information is displayed within 150ms; when the abnormal point is clicked, the annotation box is smoothly expanded within 350ms to display the explanation information; the annotation box contains a "View" button, and clicking it shows the key events supporting the attribution, and clicking the "Full Report" button pops up a detailed analysis panel on the right.
[0237] In some alternative embodiments, the above method may further include:
[0238] Step E1: Obtain the display window corresponding to the target data.
[0239] Step E2: Perform adaptive matching on the annotation content based on the size of the display window.
[0240] The display window is used to represent the display area of the display device for the target data. Since the display windows of different display devices are not necessarily the same, in order to ensure that the annotation content can be correctly rendered on display devices of different sizes, it is necessary to dynamically adjust the annotation position of the annotation content based on the size of the display window. When the annotation content is displayed on a small-screen display device (such as a mobile phone or a tablet), the detailed content can be automatically folded to adapt to the small-screen display device.
[0241] Of course, the visual display of the annotation content can also support touch interaction and keyboard navigation to adapt to display devices with different functions.
[0242] The method for attributing data anomalies provided in this embodiment combines time correlation analysis and semantic correlation analysis to determine the attribution result of the anomaly point, enhances the strength of causal reasoning, reduces the false alarm rate of anomaly attribution, and improves the accuracy of anomaly attribution. Further, by generating the annotation content of the attribution result and injecting the annotation content into the target data for visual display, the separation of the attribution result from the target data is avoided, the cognitive burden is reduced, and the same-screen presentation of the attribution result and the target data is achieved.
[0243] As multiple specific application embodiments of the present disclosure, the method for attributing data anomalies described in the present disclosure will be described herein in combination with specific application scenarios.
[0244] Application Embodiment 1: Analysis of the sudden drop in the daily active users of a mobile application.
[0245] Scenario description: A certain mobile application monitored that the number of daily active users suddenly dropped by 21.9% on March 15, 2023, from 10,500 the previous day to 8,200. This drop far exceeds the normal fluctuation range, and the management needs to quickly understand the reason and take measures.
[0246] The analysis process is as follows:
[0247] (1) Input data: Upload the daily active user trend chart (SVG format) for March.
[0248] (2) Anomaly detection: The system automatically detected that the data point on March 15 was an anomaly, with a Z-score of -4.7, and the calculated relative change rate was -21.9%. It was marked as "severe sudden drop".
[0249] (3) DSL Conversion: Generate the following DSL description:
[0250]
[0251] (4) Generate event retrieval queries based on the DSL description: "Mobile app crash events around March 15, 2023"; "Mobile operating system updates from March 13 - 17, 2023"; "New features launched by competitors in March 2023."
[0252] (5) Relevant event retrieval results: Internal event: The v2.5.0 version released at 23:30 on March 14 introduced new API interfaces; Technical event: The Android system pushed a security update (version 12.0.7) at 02:15 on March 15; Competitor event: Competitor A launched a time - limited discount activity at 09:00 on March 15; User feedback: A large number of app crash complaints appeared on social media (starting to increase at 05:00 on March 15).
[0253] (6) Calculate the correlation degree of each relevant event with the anomaly point: Version release event: Time correlation 0.92, semantic correlation 0.88, evidence strength 0.90; Android update event: Time correlation 0.85, semantic correlation 0.75, evidence strength 0.80; Competitor activity event: Time correlation 0.60, semantic correlation 0.45, evidence strength 0.40;
[0254] The system also analyzes auxiliary data: The crash rate increased by 150% during the same period (from 1.2% to 3.0%); 90% of the crashes occurred on Android devices; The failure rate of new API interface calls exceeded 60%.
[0255] (7) Attribution result: "The incompatibility between app v2.5.0 version and Android 12.0.7 security update led to an increase in the crash rate (87% confidence)", and at the same time provide a complete evidence chain.
[0256] (8) Inject the attribution result into the original SVG chart: Add a red triangle marker (diameter 10px, transparency 80%) to the data point on March 15; Add a pulsating animation effect to the marker (period 2 seconds, transparency change range 60% - 100%); Display a tooltip when hovering the mouse: "Daily active users decreased by 21.9%"; Click to expand the annotation box, showing: "The conflict between the version and the Android update led to a 150% increase in the crash rate"; Add "View Evidence" and "Full Report" buttons inside the annotation box; "View Evidence" shows the visual correlation of the version release, system update, and the increase in the crash rate over time; "Full Report" shows the timeline of key events, impact analysis, and recommended measures.
[0257] Through the automated analysis of the data detection system, this case achieved a rapid positioning from discovering anomalies to determining the causes (the entire process was completed within 5 minutes); accurately identifying the coincidence in time between version updates and system updates was the root cause; through SVG injection, the analysis results were directly presented on the original chart for intuitive understanding; based on the root cause analysis, suggestions for fixing API compatibility issues were automatically generated.
[0258] Application Example 2: Abnormal Analysis of the Conversion Rate on an E-commerce Platform.
[0259] Scenario Description: An e-commerce platform detected that the checkout conversion rate on April 22, 2023, suddenly dropped from the normal 4.2% to 2.8%, which had a significant impact on sales performance. The platform operation team needed to quickly determine the cause and take measures.
[0260] The specific abnormal analysis process is as follows:
[0261] (1) Anomaly Detection and DSL Conversion: The system detected the abnormal point of the conversion rate and generated a DSL description:
[0262]
[0263] (2) Multi-dimensional Event Retrieval Query: The data detection system automatically generated the following event retrieval queries:
[0264] Technical Event Query: "payment_system OR checkout_service incident 2023-04-22";
[0265] User Feedback Query: "cannot_checkout OR payment_issue 2023-04-22";
[0266] Competitor Activity Query: "competitor_promotion OR flash_sale 2023-04-20 to 2023-04-23".
[0267] (3) Retrieval Results of Related Events: Technical Log: The response time of the payment service increased by 300% during the period from 14:30 to 16:45; Monitoring Data: The timeout rate of third-party payment API calls increased from 0.5% to 23%; User Feedback: Mobile users reported slow loading or white screens on the payment page (starting from 14:45); External Information: The third-party payment provider admitted service fluctuations on social media (released at 16:30).
[0268] (4) Attribution Result: "Network fluctuations of the third-party payment service provider led to payment process delays, affecting the conversion of mobile users (92% confidence)".
[0269] (5) Inject the attribution result into the original chart for visual display. Add a red mark (radius 8px) at the point on April 22 in the conversion rate chart; add a warning icon beside the mark to indicate a high-confidence anomaly; the expanded annotation box after clicking contains the visualization of the payment delay period; a horizontal timeline is shown in the annotation box to highlight the time overlap between the payment delay and the decline in conversion rate; provide a "Coping Suggestions" button, and click it to display temporary solutions and long-term suggestions.
[0270] Application Example 3: Analysis of network traffic fluctuations.
[0271] Scenario description: A certain content platform detected a sudden 157% increase in website traffic on May 8, 2023, far exceeding the normal growth pattern. The operation team needs to understand the reasons for the traffic explosion and also evaluate whether the infrastructure can bear it.
[0272] The abnormal analysis process is as follows:
[0273] (1) Abnormal detection: The system detected an abnormal sudden increase in traffic and determined it as a "positive anomaly" (positive anomaly).
[0274] (2) Generate a DSL description and convert it into an event retrieval query including the following directions: Content hot topic query: "viralcontent OR trending topic 2023-05-08"; External traffic referral query: "traffic source ORreferral spike2023-05-08"; Marketing campaign query: "marketing campaign OR promotion 2023-05-06to 2023-05-08".
[0275] (3) Association of related events:
[0276] The data detection system found that the core factor was that a piece of content was cited by a well-known social media KOL, resulting in secondary dissemination: A high-follower KOL on Social Media A cited a column article on the platform at 23:15 on May 7; the citation received more than 200,000 reposts within 12 hours; traffic source analysis showed that 93% of the increased traffic came from this social platform.
[0277] (4) Inject the attribution result and provide visualization of the path analysis: Pie chart of traffic source distribution (Social Platform A accounts for 93%); User behavior path diagram (showing the typical browsing paths of new users); Content dissemination heat map (displaying the diffusion paths of content on different platforms); Server load status indicator (indicating the current expansion requirements).
[0278] In this embodiment, a device for attributing data anomalies is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0279] This embodiment provides a device for attributing data anomalies, as Figure 7 shown, including:
[0280] A data acquisition module 501, configured to acquire target data to be detected.
[0281] An anomaly recognition module 502, configured to identify anomaly points in the target data and generate an anomaly detection result.
[0282] A structure conversion module 503, configured to convert the anomaly detection result into a description in a domain-specific language.
[0283] An analysis module 504, configured to analyze the description in the domain-specific language and determine relevant events corresponding to the anomaly points.
[0284] An attribution module 505, configured to analyze the correlation between the relevant events and the anomaly points and generate an attribution result corresponding to the anomaly points.
[0285] In some alternative implementation manners, the structure conversion module 503 includes:
[0286] A feature extraction unit, configured to extract anomaly key features from the anomaly detection result.
[0287] A feature conversion unit, configured to perform feature conversion on the anomaly key features to obtain anomaly information corresponding to the anomaly key features.
[0288] An assembly unit, configured to assemble the anomaly information according to the syntax assembly rules of the description in the domain-specific language to obtain the description in the domain-specific language corresponding to the anomaly detection result.
[0289] In some alternative implementation manners, the feature conversion unit includes:
[0290] A mapping subunit, configured to perform semantic mapping on numerical features in the anomaly key features by using a preset mapping rule to obtain anomaly description information corresponding to the numerical features.
[0291] A feature fusion subunit, configured to perform feature fusion on non-numerical features in the anomaly key features by using a preset sliding window to determine context information corresponding to the non-numerical features.
[0292] An associated index determination subunit, configured to analyze the abnormal indicators in the abnormal key features and determine the associated indexes associated with the abnormal indicators.
[0293] In some alternative embodiments, the parsing module 504 includes:
[0294] A description conversion unit, configured to use a pre-trained retrieval conversion model to convert a domain-specific language description into an event retrieval query.
[0295] An event retrieval unit, configured to retrieve the events that occurred within the target time window based on the event retrieval query and determine the relevant events corresponding to the abnormal points.
[0296] In some alternative embodiments, the description conversion unit includes:
[0297] A structure parsing subunit, configured to parse the structure of the domain-specific language description, extract key description information, and construct target context information corresponding to the key description information.
[0298] A query information construction subunit, configured to construct event query information from multiple dimensions by using the key description information and the target context information.
[0299] A hint construction subunit, configured to obtain a preset time window and construct hint description information by using the target context information, the event query information, and the preset time window.
[0300] A conversion guidance subunit, configured to use the hint description information to guide the retrieval conversion model to convert the domain-specific language description into an event retrieval query.
[0301] In some alternative embodiments, the event retrieval unit includes:
[0302] An event source acquisition subunit, configured to acquire multiple event sources corresponding to the abnormal points.
[0303] A retrieval subunit, configured to perform event retrieval in multiple event sources by using the event retrieval query to obtain the relevant events within the target time window.
[0304] In some alternative embodiments, the attribution module 505 includes:
[0305] A time-related analysis unit, configured to perform time-related analysis on the abnormal points and the relevant events to determine the time-related degree.
[0306] A semantic-related analysis unit, configured to perform semantic-related analysis on the abnormal points and the relevant events to determine the semantic-related degree.
[0307] A fusion unit, configured to generate an attribution result based on the fusion result of the time-related degree and the semantic-related degree.
[0308] In some alternative embodiments, the time-related analysis unit includes:
[0309] A time difference determination subunit, configured to obtain the abnormal occurrence time corresponding to the abnormal point and the event occurrence time of the related event, and determine the time difference between the abnormal occurrence time and the event occurrence time.
[0310] A weight determination subunit, configured to determine the time-related weights of different types of related events based on the preset time difference range where the time difference is located.
[0311] A relevance determination subunit, configured to determine the time-related relevance between the related event and the abnormal point according to the time-related weight and the time difference.
[0312] In some alternative embodiments, the semantic-related analysis unit includes:
[0313] A semantic matching subunit, configured to perform semantic matching on the first semantics corresponding to the abnormal point and the second semantics corresponding to the related event to determine semantic relevance.
[0314] A path relevance determination subunit, configured to analyze the index influence path of the related event according to the preset knowledge graph to determine path relevance.
[0315] A historical matching subunit, configured to retrieve historical abnormal events, match the related event with the historical abnormal events, and determine the relevance of the historical abnormal events.
[0316] In some alternative embodiments, the above-mentioned attribution module 505 further includes:
[0317] An evaluation factor acquisition unit, configured to acquire the abnormal evaluation factors corresponding to the attribution result, where the abnormal evaluation factors include at least one of multi-source verification information, rule matching degree, and historical statistical relevance.
[0318] An attribution evaluation unit, configured to perform attribution evaluation on the abnormal point by using the abnormal evaluation factors to generate an attribution evaluation result.
[0319] In some alternative embodiments, the above-mentioned device further includes:
[0320] A note content generation module, configured to generate note content corresponding to the attribution result by using a pre-trained note generation model.
[0321] A note position determination module, configured to analyze the data structure of the target data and determine the note position of the note content in the target data.
[0322] A visualization display module, configured to inject the note content into the note position for visualization display.
[0323] In some alternative embodiments, the annotation content generation module includes:
[0324] A prompt acquisition unit for acquiring pre-constructed prompt description information.
[0325] A generation guidance unit for guiding an annotation generation model to compress and transform an attribution result by using the prompt description information, and generating annotation content at least one level.
[0326] In some alternative embodiments, the above-mentioned annotation content generation module may further include:
[0327] A first display unit for displaying brief information of the annotation content in response to a display trigger operation for the annotation content.
[0328] A second display unit for displaying explanatory information of the annotation content in response to a click trigger operation for the annotation content.
[0329] A third display unit for displaying detailed information of the annotation content in response to a view trigger operation for the annotation content.
[0330] In some alternative embodiments, the annotation position determination module includes:
[0331] A first position determination unit for analyzing the data structure of target data and determining the element positions where each data element of the target data is located.
[0332] A second position determination unit for locating the target position where an anomaly point is located based on the anomaly time of the anomaly point.
[0333] A third position determination unit for parsing the element position and the target position, and determining the annotation position of the annotation content, where the annotation position is associated with the target position and does not coincide with the element position.
[0334] In some alternative embodiments, the visualization display module includes:
[0335] An element creation unit for creating annotation elements corresponding to the annotation content. Among them, the annotation elements include at least one of a marker element, a connection line element, an annotation box element, and a text interaction element.
[0336] A content association unit for associating the annotation content with the annotation elements, and injecting the annotation elements into the annotation positions corresponding to the target data, so as to visually display the annotation content through the annotation elements.
[0337] In some alternative embodiments, the above-mentioned device may further include:
[0338] A display information acquisition module for acquiring a display window corresponding to the target data.
[0339] A display matching module, configured to adaptively match annotation content based on the size of a display window.
[0340] The further function descriptions of the above modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0341] The device for attributing data anomalies provided by the embodiments of the present disclosure can execute the method for attributing data anomalies provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method. After identifying the anomaly points in the target data, by converting the anomaly detection results into descriptions in a domain-specific language, the anomaly detection results can be mapped to structured semantic descriptions, realizing an accurate conversion from the "data layer" to the "semantic layer" and solving the understanding gap between anomaly data and business semantics. By parsing the descriptions in the domain-specific language to determine the relevant events corresponding to the anomaly points, a deep association between the anomaly points and other relevant events is realized, so as to accurately determine the main events that generate the anomaly points in combination with the correlation analysis between the relevant events and the anomaly points, and obtain the corresponding attribution results. Thus, the automation from anomaly detection to the generation of attribution results is realized, the problem of the separation between anomaly detection and attribution is overcome, and the attribution analysis effect of data anomalies is improved.
[0342] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure.
[0343] Specifically, refer to Figure 8 , which shows a schematic structural diagram of a computer device suitable for implementing the embodiments of the present disclosure. The computer device may include a processor (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the memory 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer device are also stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0344] Generally, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a memory 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the computer device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8A computer device with various devices is shown, but it should be understood that it is not required to implement or have all the shown devices, and alternatively, more or fewer devices may be implemented or had.
[0345] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the memory 608, or installed from the ROM 602. When the computer program is executed by the processor 601, the above functions defined in the method for attributing data anomalies in the embodiments of the present disclosure are executed.
[0346] Figure 8 The shown computer device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0347] The embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through the network and to be stored in a local storage medium, so that the methods described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method for attributing data anomalies shown in the above embodiments is implemented.
[0348] A part of the present disclosure can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present disclosure through the operations of the computer. Those skilled in the art should understand that the forms of existence of computer program instructions in a computer-readable medium include but are not limited to source files, executable files, installation package files, etc. Correspondingly, the ways for computer program instructions to be executed by a computer include but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0349] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for attributing data anomalies, characterized in that The method includes: Obtain target data to be detected; Identify abnormal points of the target data and generate an anomaly detection result; Convert the anomaly detection result into a domain-specific language description; Parse the domain-specific language description to determine relevant events corresponding to the abnormal points; Analyze the correlation between the relevant events and the abnormal points to generate an attribution result corresponding to the abnormal points.
2. The method according to claim 1, characterized in that The converting the anomaly detection result into a domain-specific language description includes: Extract abnormal key features from the anomaly detection result; Perform feature transformation on the abnormal key features to obtain abnormal information corresponding to the abnormal key features; Assemble the abnormal information according to the grammar assembly rules of the domain-specific language description to obtain the domain-specific language description corresponding to the anomaly detection result.
3. The method according to claim 2, characterized in that, The performing feature transformation on the abnormal key features to obtain abnormal information corresponding to the abnormal key features includes: Perform semantic mapping on the abnormal key features using a preset mapping rule to obtain abnormal description information; Perform feature fusion on the abnormal key features using a preset sliding window to determine context information; Analyze abnormal indicators corresponding to the abnormal key features to determine associated indicators associated with the abnormal indicators; The abnormal information includes abnormal description information, the context information, and the associated indicators.
4. The method according to claim 1, wherein The parsing the domain-specific language description to determine relevant events corresponding to the abnormal points includes: Use a pre-trained retrieval transformation model to convert the domain-specific language description into an event retrieval query; Based on the event retrieval query, retrieve events that occurred within a target time window to determine relevant events corresponding to the abnormal points.
5. The method according to claim 4, characterized in that, The using a pre-trained retrieval transformation model to convert the domain-specific language description into an event retrieval query includes: Parse the structure of the domain-specific language description, extract key description information, and construct target context information corresponding to the key description information; Use the key description information and the target context information to construct event query information from multiple dimensions; Obtain a preset time window, and use the target context information, the event query information, and the preset time window to construct prompt description information; Use the prompt description information to guide the retrieval transformation model to convert the domain-specific language description into an event retrieval query.
6. The method according to claim 4, characterized in that, The based on the event retrieval query, retrieving events that occurred within a target time window to determine relevant events corresponding to the abnormal points includes: Obtain multiple event sources corresponding to the abnormal points; Use the event retrieval query to perform event retrieval among multiple event sources to obtain relevant events within the target time window.
7. The method according to claim 1, characterized in that The analyzing the correlation between the relevant events and the abnormal points to generate an attribution result corresponding to the abnormal points includes: Perform temporal correlation analysis on the abnormal points and the relevant events to determine the temporal correlation degree; Perform semantic correlation analysis on the abnormal points and the relevant events to determine the semantic correlation degree; Generate the attribution result based on the fusion result of the temporal correlation degree and the semantic correlation degree.
8. The method according to claim 7, wherein Performing time correlation analysis on the abnormal point and the related event to determine the time correlation degree, including: Obtaining the abnormal occurrence time corresponding to the abnormal point and the event occurrence time of the related event, and determining the time difference between the abnormal occurrence time and the event occurrence time; Based on the preset time difference range where the time difference is located, determining the time correlation weights of different types of the related events; Determining the time correlation degree between the related event and the abnormal point according to the time correlation weight and the time difference.
9. The method according to claim 7, wherein Performing semantic correlation analysis on the abnormal point and the related event to determine the semantic correlation degree, including: Performing semantic matching on the first semantics corresponding to the abnormal point and the second semantics corresponding to the related event to determine the semantic correlation; Analyzing the index influence path of the related event according to the preset knowledge graph to determine the path correlation; Retrieving historical abnormal events, matching the related event with the historical abnormal events to determine the historical abnormal event correlation; Determining the semantic correlation degree based on the fusion result of the semantic correlation, the path correlation, and the historical abnormal event correlation.
10. The method according to any one of claims 7-9, characterized in that Further including: Obtaining the abnormal evaluation factors corresponding to the attribution result, where the abnormal evaluation factors include at least one of multi-source verification information, rule matching degree, and historical statistical correlation; Using the abnormal evaluation factors to perform attribution evaluation on the abnormal point to generate an attribution evaluation result.
11. The method according to claim 1, wherein Further including: Using a pre-trained annotation generation model to generate annotation content corresponding to the attribution result; Analyzing the data structure of the target data to determine the annotation position of the annotation content in the target data; Injecting the annotation content into the annotation position for visual display.
12. The method according to claim 11, wherein The using a pre-trained annotation generation model to generate annotation content corresponding to the attribution result includes: Obtaining pre-constructed prompt description information; Using the prompt description information to guide the annotation generation model to compress and transform the attribution result to generate at least one level of the annotation content.
13. The method according to claim 12, wherein Further including: Responding to a display trigger operation for the annotation content to display brief information of the annotation content; and / or, Responding to a click trigger operation for the annotation content to display explanatory information of the annotation content; and / or, Responding to a view trigger operation for the annotation content to display detailed information of the annotation content.
14. The method according to claim 11, wherein The analyzing the data structure of the target data to determine the annotation position of the annotation content in the target data includes: Analyzing the data structure of the target data to determine the element positions where each data element of the target data is located; Locating the target position where the abnormal point is located based on the abnormal time of the abnormal point; Parsing the element position and the target position to determine the annotation position of the annotation content, where the annotation position is associated with the target position and does not coincide with the element position.
15. The method according to any one of claims 11-14, characterized in that, The injecting the annotation content into the annotation position for visual display includes: Create an annotation element corresponding to the annotation content, where the annotation element includes at least one of a marker element, a connection line element, an annotation box element, and a text interaction element; Associate the annotation content with the annotation element and inject the annotation element into the annotation position corresponding to the target data, so as to visually display the annotation content through the annotation element.
16. The method according to claim 11, wherein Further includes: Obtain the display window corresponding to the target data; Based on the size of the display window, perform adaptive matching on the annotation content.
17. An attribution device for data anomalies, characterized in that, The device includes: A data acquisition module for acquiring target data to be detected; An anomaly recognition module for identifying anomaly points of the target data and generating an anomaly detection result; A structure conversion module for converting the anomaly detection result into a domain-specific language description; An analysis module for analyzing the domain-specific language description to determine relevant events corresponding to the anomaly points; An attribution module for analyzing the correlation between the relevant events and the anomaly points and generating an attribution result corresponding to the anomaly points.
18. A computer device, characterized in that, Includes: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the attribution method for data anomalies according to any one of claims 1 to 16.
19. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the attribution method for data anomalies according to any one of claims 1 to 16.
20. A computer program product, characterized in that, Includes computer instructions, and the computer instructions are used to cause a computer to execute the attribution method for data anomalies according to any one of claims 1 to 16.
Citation Information
Cited By
Abnormal code detection method and system based on artificial intelligence
CN121349835A
RPA cluster anomaly attribution analysis method
CN122337536A