Abnormal data processing method and device, equipment, medium and product

By using natural language text to identify target information in enterprise operation and management, acquiring heterogeneous data and generating attribution factor tokens for causal reasoning, the problem of low analysis efficiency, limited accuracy and lack of transparency in existing technologies is solved, and efficient and accurate location of abnormal indicator causes is achieved.

CN121523943APending Publication Date: 2026-02-13INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511664666.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing technologies in enterprise operation and management suffer from low efficiency, limited accuracy, difficulty in utilizing internal enterprise knowledge, poor user interaction experience, and lack of intelligent multi-round exploration capabilities, resulting in inaccurate analysis results and opaque analysis processes.

Method used

By identifying target information from the natural language text of business anomalies, acquiring heterogeneous data, generating attribution factor tokens and labeling event impacts, and using a large model for causal reasoning, the anomaly analysis results are obtained.

Benefits of technology

It achieves efficient, accurate, reliable and interpretable localization of the causes of indicator anomalies, solves the problems of insufficient causal reasoning ability and opaque analysis process, and improves analysis efficiency and interpretability of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121523943A_ABST
    Figure CN121523943A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal data processing method, device and equipment, a medium and a product, relates to the technical field of artificial intelligence and large models, and can be applied to the field of financial science and technology. The method comprises the following steps: determining target information of a data item according to a natural language text of a business exception problem; wherein the target information comprises a target intention and a target abnormal item; wherein the target intention comprises a target keyword, a time range and an exception description; acquiring heterogeneous data from a business internal query platform according to the target information; wherein the heterogeneous data comprises index dimension data and information data; generating an attribution factor token according to the heterogeneous data, and performing event influence labeling on the attribution factor token to obtain an influence vector; and based on a large model, performing causal reasoning according to the attribution factor token and the influence vector to obtain an anomaly analysis result. According to the technical scheme, efficient, accurate, reliable and explainable index abnormity reason positioning is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and large model technology, which can be used in the field of financial technology, and particularly relates to an abnormal data processing method, device, equipment, medium and product. BACKGROUND

[0002] In enterprise operation management, monitoring and abnormal analysis of key business indicators are crucial. Traditional indicator abnormal reason analysis methods mainly include: 1) manual analysis method: business analysts or operation and maintenance personnel observe the indicator monitoring dashboard, find the indicator abnormality, manually log in to each related system (such as database, log system, internal information platform, etc.), query the indicator historical data, related dimension data, recent changes, related events or news, and make comprehensive judgments based on personal experience and knowledge to locate the abnormal reason. 2) Rule-based alarm system: preset fixed threshold or rule, when the indicator triggers the rule, the system automatically alarms. Some advanced systems may associate some preset troubleshooting steps or knowledge base links, but the analysis process still needs manual intervention, and it is difficult to cover all complex scenarios. 3) General data analysis tool combined with AI capability: some business intelligence tools or data analysis platforms start to integrate AI capabilities, such as allowing users to query data through natural language or providing some basic anomaly detection algorithms. But these general tools often lack deep integration with enterprise internal specific knowledge base, information platform and indicator system, and it is difficult to conduct targeted, context-aware reason inference. SUMMARY

[0003] The present application provides an abnormal data processing method, device, equipment, medium and product to realize efficient, accurate, reliable and explainable indicator abnormal reason positioning.

[0004] According to an aspect of the present application, an abnormal data processing method is provided, which comprises:

[0005] According to the natural language text of the business abnormal problem, the target information of the data item is determined; wherein the target information includes a target intention and a target abnormal item; wherein the target intention includes a target keyword, a time range and an abnormal description;

[0006] According to the target information, heterogeneous data is obtained from the internal business query platform; wherein the heterogeneous data includes indicator dimension data and information data;

[0007] According to the heterogeneous data, an attribution factor token is generated, and event impact annotation is performed on the attribution factor token to obtain an impact vector;

[0008] Based on a large model, causal reasoning is performed according to the attribution factor token and the impact vector to obtain an abnormal analysis result.

[0009] According to another aspect of the present application, there is provided an abnormal data processing apparatus, comprising:

[0010] a target information determination module configured to determine target information of the data item according to the natural language text of the business abnormal problem; wherein the target information comprises a target intention and a target abnormal item; wherein the target intention comprises a target keyword, a time range and an abnormal description;

[0011] a heterogeneous data acquisition module configured to acquire heterogeneous data from a business internal query platform according to the target information; wherein the heterogeneous data comprises index dimension data and information data;

[0012] an attribution factor token generation module configured to generate attribution factor tokens according to the heterogeneous data, and perform event impact labeling on the attribution factor tokens to obtain an impact vector;

[0013] an abnormal analysis result determination module configured to perform causal reasoning based on a large model according to the attribution factor tokens and the impact vector to obtain an abnormal analysis result.

[0014] According to another aspect of the present application, there is provided an electronic device, comprising:

[0015] at least one processor; and a memory communicatively connected to the at least one processor; wherein

[0016] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the abnormal data processing method according to any one of the embodiments of the present application.

[0017] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for enabling a processor to implement the abnormal data processing method according to any one of the embodiments of the present application when executed by the processor.

[0018] According to another aspect of the present application, there is provided a computer program product comprising a computer program for implementing the abnormal data processing method according to any one of the embodiments of the present application when executed by a processor.

[0019]

[0020] The technical scheme of the embodiment of the present application determines target information of a data item according to a natural language text of a business abnormal problem; wherein the target information includes a target intention and a target abnormal item; wherein the target intention includes a target keyword, a time range and an abnormal description; acquires heterogeneous data from a business internal query platform according to the target information; wherein the heterogeneous data includes index dimension data and information data; generates an attribution factor token according to the heterogeneous data, and performs event impact labeling on the attribution factor token to obtain an impact vector; performs causal reasoning based on a large model according to the attribution factor token and the impact vector to obtain an abnormal analysis result. The above technical scheme solves the technical problems that a general large language model is prone to fact confusion, insufficient causal relationship reasoning ability, non-transparent (black box) analysis process and difficulty in utilizing historical analysis experience of an enterprise when directly processing complex and heterogeneous enterprise internal data, and realizes efficient, accurate, reliable and interpretable index abnormal reason positioning.

[0021] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained according to these drawings without creative labor for those skilled in the art.

[0023] Figure 1 is a flowchart of an abnormal data processing method according to an embodiment of the present application;

[0024] Figure 2 is a flowchart of an abnormal data processing method according to an embodiment of the present application;

[0025] Figure 3 is a structural schematic diagram of an abnormal data processing device according to an embodiment of the present application;

[0026] Figure 4 is a structural schematic diagram of an electronic device for implementing an abnormal data processing method according to an embodiment of the present application. DETAILED DESCRIPTION

[0027] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application, so that those skilled in the art can better understand the technical solutions of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should belong to the protection scope of the present application.

[0028] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0029] In addition, it should also be noted that in the technical solutions of the present application, the collected information is information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards of relevant countries and regions, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal.

[0030] The prior art has the following disadvantages in practical application: 1) low analysis efficiency: the manual analysis method highly depends on the experience and efficiency of the analyst, needs to query multiple scattered systems, is time-consuming and laborious, and is difficult to respond quickly. 2) analysis accuracy is limited: the accuracy of manual analysis is limited by personal knowledge and experience, is easily affected by subjective factors, and key information may be missed. The rule-based system is helpless for unexpected abnormal patterns. 3) enterprise internal knowledge is not fully utilized: general tools are difficult to effectively utilize valuable contextual information such as internal accumulated index definitions, business terms, historical events and specific information, resulting in that the analysis result may not be accurate or lack depth. 4) poor user interaction experience: users usually need to have certain technical background to operate complex query systems or understand raw data, and the threshold for non-technical background business personnel is high. 5) lack of intelligent multi-round exploration capability: when the preliminary analysis result fails to completely answer the user's question, the existing system mostly does not support convenient and context-based follow-up questions and in-depth exploration.

[0031] Figure 1 This is a flowchart of an abnormal data processing method according to an embodiment of the present invention. This embodiment is applicable to situations in enterprise operation and maintenance management where key business indicators are monitored and abnormal analysis is performed. The method can be executed by an abnormal data processing device, which can be implemented in hardware and / or software. This device can be configured in an electronic device that carries abnormal data processing functions, such as a server. Figure 1 As shown, the method includes:

[0032] S110. Determine the target information of the data item based on the natural language text of the business anomaly issue.

[0033] In this embodiment, "business anomaly" refers to anomalies in business-related metrics. "Natural language text" refers to text directly input by the user to describe the business anomaly. "Data item" refers to a business metric item. "Target information" refers to detailed information about the data item; optionally, target information includes target intent and target anomaly item; where target intent includes target keywords, time range, and anomaly description; target intent refers to the intent determined based on the natural language text. "Target anomaly item" refers to a business metric item that exhibits anomalies. It should be noted that metrics refer to quantifiable data used to measure enterprise operational status, business performance, system performance, etc., such as daily active users, sales revenue, CPU utilization, etc.

[0034] An alternative approach involves determining the target information of a data item based on the natural language text of a business anomaly issue, including: responding to the natural language text input by the user regarding the business anomaly issue on the front-end page; performing intent recognition on the natural language text to obtain the target intent; and performing precise matching from the enterprise's internal knowledge base based on the target keywords in the target intent to obtain the target anomaly item.

[0035] Specifically, users can input natural language text regarding business anomalies through the input interface on the computer's front-end page. Then, an intent recognition model can identify the intent from the natural language text to obtain the target intent. This intent recognition model can be a small model or a large language model; it should be noted that a small model refers to a neural network model with fewer network parameters and a simpler network structure, while a large model refers to a neural network model with more network parameters and a more complex network structure. Then, the target keywords in the target intent are precisely matched against the enterprise's internal knowledge base to obtain the target anomaly. If a match fails, a notification of matching failure is sent to the user on the front end.

[0036] Understandably, by performing intent recognition on the natural speech input by users and combining it with the enterprise's internal knowledge base, the target intent and target anomalies can be obtained. This enables the identification of intent and the locking of abnormal indicators, laying the foundation for subsequent anomaly identification.

[0037] It should be noted that an enterprise's internal private knowledge base refers to an internal database or knowledge system that stores enterprise-specific indicator definitions, business terminology, indicator metadata, etc.

[0038] S120. Obtain heterogeneous data from the internal business query platform based on the target information.

[0039] In this embodiment, the internal business query platform refers to a platform used within the enterprise for data querying. Heterogeneous data refers to data of different dimensions and data types; optionally, heterogeneous data includes indicator-level data and information data. Indicator-level data refers to detailed data corresponding to indicators, such as time series data, breakdown data from different channels / regions / production areas, etc.; information data refers to textual information data related to indicators, such as market activities, system changes, industry news, etc. It should be noted that the internal business query platform includes an indicator query platform and an information query platform. The indicator query platform refers to an internal system platform used by the enterprise to query, display, and monitor various indicator data; the information query platform refers to an internal or externally integrated system platform used to query business-related news, announcements, industry trends, market events, internal fault reports, and other information.

[0040] An optional approach is to obtain heterogeneous data from an internal business query platform based on target information, including: obtaining indicator dimension data from the internal business query platform based on target anomalies and time ranges; and obtaining information data within a time range from the internal business query platform based on target anomalies and their associated entities.

[0041] Specifically, the system can use the target anomaly and time range as query conditions to obtain the corresponding indicator dimension data from the indicator query platform of the internal business query platform. At the same time, it can use the target anomaly and its associated entities as indexes and the time range as the query time condition to obtain information data from the information query platform of the internal business query platform.

[0042] It is understandable that parallel acquisition of heterogeneous data can improve the efficiency of heterogeneous data acquisition.

[0043] S130. Generate attribution factor tokens based on heterogeneous data, and label the attribution factor tokens with event impact to obtain the impact vector.

[0044] In this embodiment, the attribution factor token is a standardized intermediate data structure for encapsulating a potential cause event that may lead to fluctuations in indicators; optionally, the attribution factor token includes structured objects of event description, timestamp, impact type, source, and confidence, etc. metadata as the basic unit for causal reasoning by a large model. The impact vector is a key attribute in the attribution factor token, which is used to predict or label which indicators an event may affect and in which direction (positive / negative). For example, the impact vector of a "market promotion activity" event may be labeled as { "daily active users": "+", "new user registration volume": "+", "average user payment": "-"}.

[0045] Specifically, the heterogeneous data is analyzed and processed to convert into structured attribution factor tokens. At the same time, an impact vector is labeled for each attribution factor token; for example, based on rules / knowledge base, for known event types, a preset impact vector can be matched from the knowledge base. For example, the knowledge base defines that an "App new version release" event will usually bring { "App crash rate": "+", "new function usage rate": "+"} impact. For example, based on historical data learning, the actual fluctuations of various indicators after the occurrence of similar events in history can be analyzed to learn and generate an impact vector. For example, the event description can also be submitted to a large language model to predict the possible impact vector according to common sense and business background. By labeling the impact vector for the attribution factor, the accuracy of subsequent indicator anomaly analysis can be improved.

[0046] S140, based on a large model, causal reasoning is performed according to the attribution factor token and the impact vector to obtain an anomaly analysis result.

[0047] In this embodiment, the anomaly analysis result refers to the result of analyzing the cause of the unexpected fluctuations or deviation from the normal range of indicators; optionally, the anomaly analysis result is in the form of a report.

[0048] Specifically, the attribution factor token and the impact vector can be combined into a data set, a target prompt word is generated for the data set, the target prompt word is input into the large model, and causal reasoning is performed by the large model to obtain the anomaly analysis result.

[0049] The technical scheme of the embodiment of the application determines target information of a data item according to a natural language text of a business abnormal problem, wherein the target information includes a target intention and a target abnormal item; the target intention includes a target keyword, a time range and an abnormal description; acquires heterogeneous data from a business internal query platform according to the target information; the heterogeneous data includes index dimension data and information data; generates an attribution factor token according to the heterogeneous data, and performs event impact labeling on the attribution factor token to obtain an impact vector; and performs causal reasoning based on a large model according to the attribution factor token and the impact vector to obtain an abnormal analysis result. The above technical scheme solves technical problems such as fact confusion, insufficient causal relationship reasoning capability, non-transparent (black box) analysis process and difficulty in utilizing historical analysis experience of an enterprise when a general large language model directly processes complex and heterogeneous enterprise internal data, and realizes efficient, accurate, reliable and interpretable index abnormal reason positioning.

[0050] Figure 2 is a flowchart of an abnormal data processing method according to an embodiment of the application, and the embodiment optimizes the steps of “generating an attribution factor token according to heterogeneous data” and “performing causal reasoning based on a large model according to the attribution factor token and the impact vector to obtain an abnormal analysis result” based on the above embodiment, and provides an optional implementation scheme. As shown in Figure 2 , the method comprises:

[0051] S210, determining target information of a data item according to a natural language text of a business abnormal problem.

[0052] The target information includes a target intention and a target abnormal item; the target intention includes a target keyword, a time range and an abnormal description.

[0053] S220, acquiring heterogeneous data from a business internal query platform according to the target information.

[0054] The heterogeneous data includes index dimension data and information data.

[0055] S230, generating an attribution factor token according to the heterogeneous data, and performing event impact labeling on the attribution factor token to obtain an impact vector.

[0056] S240, performing causal reasoning based on a large model according to the attribution factor token and the impact vector to obtain an abnormal analysis result.

[0057] An optional way of generating an attribution factor token according to heterogeneous data comprises: performing factor extraction on the heterogeneous data to obtain an abnormal event; tokenizing and packaging the abnormal event to obtain an attribution factor token; and the attribution factor token includes at least one field of a unique identifier, a factor type, a timestamp, an event description, a data source and a data pointer.

[0058] Specifically, one or more preprocessing modules can be used to extract factors from heterogeneous data to obtain abnormal events. For each identified abnormal event, a standardized attribution factor token is created, which includes at least the following fields:

[0059] Factor_ID: Unique identifier.

[0060] Factor_Type: Factor type, enumerated values such as INTERNAL_DATA_SHIFT (internal data change), MARKETING_CAMPAIGN (marketing campaign), SYSTEM_RELEASE (system release), EXTERNAL_NEWS (external news), etc.

[0061] Timestamp: Timestamp, i.e. the exact time point or time period when the event occurs or takes effect.

[0062] Description: Event description, i.e. a natural language short description of the event (e.g. "Channel A user inflow volume down 50%," "Double 11 preheat activity online").

[0063] Source: Data source (e.g. "Indicator platform - channel dimension," "Information platform - market department announcement").

[0064] Data_Evidence: Data pointer, i.e. a pointer or key data segment pointing to the original data.

[0065] It can be understood that by converting unstructured data into structured data through attribution factor tokens, subsequent large models can perform anomaly analysis, improving anomaly analysis efficiency and accuracy.

[0066] An optional way is to perform causal reasoning based on a large model according to attribution factor tokens and influence vectors to obtain abnormal analysis results, including: constructing structured reasoning prompt words according to attribution factor tokens and influence vectors; performing reasoning on the structured reasoning prompt words through the large model to obtain abnormal analysis results.

[0067] Specifically, according to the attribution factor token and the influence vector, the attribution factor token list is constructed, and the structured reasoning prompt word is constructed based on the prompt word engineering according to the attribution factor token list. The constructed prompt word is a highly structured prompt word, which is no longer a messy original text, but an attribution factor token list. Exemplarily, the structured reasoning prompt word format is as follows:

[0068] Then, the large model performs causal chain reasoning based on structured reasoning prompts and prompts with explicit reasoning instructions, rather than simply summarizing text. The model is guided to match consistency between time and impact. The large model generates a logically clear and evidence-based analysis report, i.e., the anomaly analysis results.

[0069] Understandably, transforming fuzzy correlation analysis problems into structured causal reasoning problems effectively reduces the "illusion" of large models, making their analytical logic closer to the thinking patterns of human experts.

[0070] The technical solution of this invention determines the target information of data items based on the natural language text of business anomalies. The target information includes the target intent and the target anomaly. The target intent includes target keywords, a time range, and an anomaly description. Heterogeneous data is obtained from an internal business query platform based on the target information. This heterogeneous data includes indicator dimension data and information data. Attribution factor tokens are generated based on the heterogeneous data, and event impact annotations are applied to these tokens to obtain an impact vector. Based on a large model, causal inference is performed using the attribution factor tokens and the impact vector to obtain anomaly analysis results. This technical solution solves the technical challenges of general-purpose large language models when directly processing complex and heterogeneous internal enterprise data, such as factual confusion, insufficient causal reasoning ability, opaque analysis process (black box), and difficulty in utilizing historical analysis experience accumulated by the enterprise. It achieves efficient, accurate, reliable, and interpretable location of the causes of indicator anomalies.

[0071] Based on the above embodiments, as an optional aspect of the present invention, it further includes: visualizing the anomaly analysis results; responding to follow-up questions about the anomaly analysis results; locating the attribution factor token based on the follow-up questions; and performing data querying and analysis based on the attribution factor token to conduct multi-round dialogue.

[0072] Specifically, the system displays anomaly analysis results and responds to user inquiries about these results, providing follow-up questions such as, "Please provide a detailed explanation of the situation with Channel A." Based on these questions, the system locates the attribution factor token with ID AFT001 and performs deeper data queries and analysis around the token's Data_Evidence, significantly improving the focus and accuracy of multi-turn dialogues.

[0073] Understandably, enhancing the interpretability of the analysis process, with each conclusion in the anomaly analysis results supported by a clear attribution factor token that can be traced back to the original data source and event, makes the entire analysis process no longer a black box and increases the business's trust in the results.

[0074] The attribution factor tokens and their influence vectors proposed in this invention can be stored and optimized. Each successful analysis can be used to calibrate and enrich the influence vector knowledge base. This forms a virtuous cycle of knowledge accumulation, enabling the system to "become smarter with use," which is something that general solutions cannot achieve. This enables the accumulation and reuse of enterprise knowledge.

[0075] The technical solution of this invention decomposes complex analysis tasks into two stages: "factor extraction" and "causal inference." The factor extraction stage can be processed in parallel by lighter, more specialized models or rules. The main analysis LLM only needs to process standardized attribution factor tokens, reducing the complexity and cost of its processing context and making it easier to access new data sources, i.e., only requiring the development of new attribution factor token generators. This improves system processing efficiency and scalability.

[0076] The closest existing alternative is the "single-stage end-to-end analysis method," which involves indiscriminately piecing together all raw text and data—including user questions, metric data, and related information—into a single large prompt word and directly inputting it into a large language model, hoping it will output the analysis results in one go. This approach has significant disadvantages: 1) Low reliability: The massive, unstructured input information easily confuses the large model, making it difficult to accurately capture the temporal and logical connections between key events, leading to unstable results and potential illusions. 2) Lack of interpretability: The analysis process is a complete black box; it's impossible to know what specific information the model is based on to reach its conclusions, making it difficult to gain user trust. 3) Inability to retain knowledge: Each analysis is a one-off event, making it impossible to structure, store, and reuse the causal relationships extracted during the analysis process (e.g., which metric is affected by a system change). 4) High cost and low efficiency: Processing extremely long, unstructured contexts consumes enormous computational resources from the large model and has a slow response time. Therefore, although the above-mentioned alternatives exist, they cannot solve the core technical problem addressed by the present invention, nor can they bring the beneficial effects of the present invention in terms of reliability, interpretability, and knowledge accumulation.

[0077] Figure 3 This is a schematic diagram of an anomaly data processing device according to an embodiment of the present invention. This embodiment is applicable to situations in enterprise operation and maintenance management involving the monitoring and anomaly analysis of key business indicators. The anomaly data processing device can be implemented in hardware and / or software, and can be configured in an electronic device that carries the anomaly data processing function, such as a server. Figure 3 As shown, the device includes:

[0078] The target information determination module 310 is configured to determine target information of the data item according to the natural language text of the business exception problem; wherein the target information comprises a target intention and a target exception item; wherein the target intention comprises a target keyword, a time range and an exception description;

[0079] The heterogeneous data acquisition module 320 is configured to acquire heterogeneous data from a business internal query platform according to the target information; wherein the heterogeneous data comprises index dimension data and information data;

[0080] The attribution factor token generation module 330 is configured to generate an attribution factor token according to the heterogeneous data, and perform event impact labeling on the attribution factor token to obtain an impact vector;

[0081] The abnormal analysis result determination module 340 is configured to perform causal reasoning based on a large model according to the attribution factor token and the impact vector to obtain an abnormal analysis result.

[0082] The technical scheme of the embodiment of the application determines target information of the data item according to the natural language text of the business exception problem; wherein the target information comprises a target intention and a target exception item; wherein the target intention comprises a target keyword, a time range and an exception description; acquires heterogeneous data from a business internal query platform according to the target information; wherein the heterogeneous data comprises index dimension data and information data; generates an attribution factor token according to the heterogeneous data, and performs event impact labeling on the attribution factor token to obtain an impact vector; and performs causal reasoning based on a large model according to the attribution factor token and the impact vector to obtain an abnormal analysis result. The above technical scheme solves the technical problems that a general large language model is prone to fact confusion, causal relationship reasoning deficiency, non-transparent (black box) analysis process and difficulty in utilizing historical analysis experience of an enterprise when directly processing complex and heterogeneous enterprise internal data, and realizes efficient, accurate, reliable and interpretable index exception reason positioning.

[0083] Optionally, the target information determination module 310 is configured to:

[0084] respond to the natural language text input by a user on a front-end page for the business exception problem;

[0085] perform intention recognition on the natural language text to obtain the target intention;

[0086] perform accurate matching on the target keyword in the target intention from an enterprise internal knowledge base to obtain the target exception item.

[0087] Optionally, the heterogeneous data acquisition module 320 is configured to:

[0088] acquire the index dimension data from the business internal query platform according to the target exception item and the time range;

[0089] According to the target abnormal item and the associated entity of the target abnormal item, information data in a time range is acquired from a business internal query platform.

[0090] Optionally, the attribution factor token generation module 330 is configured to:

[0091] Factor extraction is performed on the heterogeneous data to obtain an abnormal event.

[0092] The abnormal event is tokenized and encapsulated to obtain an attribution factor token; the attribution factor token includes at least one of a unique identifier, a factor type, a timestamp, an event description, a data source, and a data pointer.

[0093] Optionally, the abnormal analysis result determination module 340 is configured to:

[0094] According to the attribution factor token and the influence vector, a structured reasoning prompt word is constructed.

[0095] The structured reasoning prompt word is reasoned by a large model to obtain an abnormal analysis result.

[0096] Optionally, the device further includes a post-processing module configured to:

[0097] The abnormal analysis result is visually processed.

[0098] In response to a follow-up content of the abnormal analysis result.

[0099] According to the follow-up content, the attribution factor token is located, and data query and analysis are performed according to the attribution factor token to perform multiple rounds of dialogue.

[0100] The abnormal data processing device provided in the embodiments of the present application can execute the abnormal data processing method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0101] According to the embodiments of the present application, the present application further provides an electronic device, a readable storage medium, and a computer program product.

[0102] Figure 4 FIG. 1 is a structural schematic diagram of an electronic device for implementing the abnormal data processing method according to an embodiment of the present application. Figure 4A structural diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0103] As shown, Figure 4 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., connected in communication with the at least one processor 11, where the memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0104] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, speakers, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0105] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the abnormal data processing method.

[0106] In some embodiments, the abnormal data processing method can be implemented as a computer program tangibly embodied in a computer readable storage medium, e.g., storage unit 18. In some embodiments, parts or all of the computer program can be loaded and / or installed onto electronic device 10 via, e.g., ROM 12 and / or communication unit 19. When the computer program is loaded onto RAM 13 and executed by processor 11, one or more steps of the above-described abnormal data processing method can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the abnormal data processing method by way of other any suitable means, e.g., by way of firmware.

[0107] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, specially designed application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0108] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor of the machine, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0109] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0110] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0111] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0112] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business extensibility in traditional physical host and virtual private service.

[0113] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.

[0114] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An abnormal data processing method characterized by, The method comprises the following steps: According to the natural language text of the business exception problem, the target information of the data item is determined; wherein, the target information includes target intent and target exception item; wherein, the target intent includes target keyword, time range and exception description; According to the target information, heterogeneous data is obtained from the business internal query platform; wherein, the heterogeneous data includes index dimension data and information data; According to the heterogeneous data, the attribution factor token is generated, and the event impact annotation is performed on the attribution factor token to obtain the impact vector; Based on the large model, the cause and effect reasoning is performed according to the attribution factor token and the impact vector to obtain the abnormal analysis result.

2. The method of claim 1, wherein, According to the natural language text of the business exception problem, the target information of the data item is determined, including: In response to the natural language text input by the user on the front-end page of the business exception problem; The target intent is obtained by performing intent recognition on the natural language text; According to the target keyword in the target intent, the target exception item is obtained by accurately matching the enterprise internal knowledge base.

3. The method of claim 1, wherein, According to the target information, heterogeneous data is obtained from the business internal query platform, including: According to the target exception item and the time range, the index dimension data is obtained from the business internal query platform; According to the target exception item and the associated entity of the target exception item, the information data within the time range is obtained from the business internal query platform.

4. The method of claim 1, wherein, According to the heterogeneous data, the attribution factor token is generated, including: The factor extraction is performed on the heterogeneous data to obtain the abnormal event; Tokenization encapsulation is performed on the abnormal event to obtain the attribution factor token; wherein, the attribution factor token includes at least one field of unique identifier, factor type, timestamp, event description, data source and data pointer.

5. The method of claim 1, wherein, Based on the large model, the cause and effect reasoning is performed according to the attribution factor token and the impact vector to obtain the abnormal analysis result, including: According to the attribution factor token and the impact vector, the structured reasoning prompt word is constructed; The large model is used to perform reasoning on the structured reasoning prompt word to obtain the abnormal analysis result.

6. The method according to any one of claims 1-5, characterized in that, Further comprising: The abnormal analysis result is visualized; In response to the follow-up content of the abnormal analysis result; According to the follow-up content, the attribution factor token is located, and data query and analysis are performed according to the attribution factor token to perform multiple rounds of dialogue.

7. An abnormal data processing device characterized by comprising: The method comprises the following steps: The target information determination module is used to determine the target information of the data item according to the natural language text of the business exception problem; wherein, the target information includes target intent and target exception item; wherein, the target intent includes target keyword, time range and exception description; The heterogeneous data acquisition module is used to obtain heterogeneous data from the business internal query platform according to the target information; wherein, the heterogeneous data includes index dimension data and information data; The attribution factor token generation module is used to generate the attribution factor token according to the heterogeneous data, and perform event impact annotation on the attribution factor token to obtain the impact vector; The abnormal analysis result determination module is used to perform cause and effect reasoning based on the large model according to the attribution factor token and the impact vector to obtain the abnormal analysis result.

8. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the abnormal data processing method in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the abnormal data processing method in any one of claims 1-6 when executed.

10. A computer program product, characterised in that, The computer program product comprises a computer program which, when executed by a processor, implements the abnormal data processing method according to any one of claims 1-6.