A data processing method and device, electronic equipment and storage medium

By generating anomaly analysis and processing prompts, abnormal objects on the big data platform are automatically processed, solving the problem of low operation and maintenance efficiency in existing technologies and achieving efficient operation and maintenance.

CN122431927APending Publication Date: 2026-07-21HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
Filing Date
2026-04-10
Publication Date
2026-07-21

Smart Images

  • Figure CN122431927A_ABST
    Figure CN122431927A_ABST
Patent Text Reader

Abstract

The application discloses a data processing method and device, electronic equipment and storage medium, comprising: generating a first prompt word of an abnormal analysis task based on an abnormal analysis guide scheme input for the abnormal analysis task, the abnormal analysis guide scheme being used for indicating an analysis logic of an abnormal object in a target data platform based on abnormal object data; obtaining abnormal object data of the target data platform based on the first prompt word, and executing the abnormal analysis task based on the first prompt word and the abnormal object data to determine a target abnormal object of the target data platform currently occurring an abnormality; generating a second prompt word of an abnormal processing task based on an abnormal processing guide scheme of the target abnormal object, the abnormal processing guide scheme being used for indicating a processing logic of the identified target abnormal object; the application meets the efficient operation and maintenance demand of the big data platform, realizes the automatic operation and maintenance processing of the big data platform, and effectively improves the operation and maintenance efficiency of the big data platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically to a data processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of computer technology, various applications, in order to achieve comprehensive awareness of user behavior, typically connect their access points to big data platforms. As the core carrier for storing, computing, and analyzing massive amounts of data, the stable and efficient operation of big data platforms directly determines the accuracy of business decisions and the continuity of business operations. Currently, the core functions of big data platforms encompass the full lifecycle management of big data tasks, including mainstream distributed computing frameworks. This specifically includes providing development environments for big data tasks, submitting and deploying them, monitoring their runtime status, handling faults, and scheduling resources.

[0003] To ensure the reliable operation of big data platforms and various big data tasks, an operation and maintenance monitoring system is an indispensable core component. Under current technology, the operation and maintenance monitoring chain of big data platforms typically involves acquiring various monitoring data, reviewing the data according to preset rules, and triggering alarms. After alarm detection, abnormal data is handled manually or automatically. However, existing data monitoring methods rely on fixed preset rules and thresholds to determine abnormal data used by big data tasks and require manual review and subsequent processing. This makes it difficult to meet the high-efficiency operation and maintenance needs of big data platforms, and the lack of an automated data operation and maintenance process for big data platforms results in low operation and maintenance efficiency. Summary of the Invention

[0004] This application provides a data processing method, apparatus, electronic device, and storage medium. It constructs a first prompt word for an anomaly analysis task based on an anomaly analysis guidance scheme, and executes the anomaly analysis task using the first prompt word to identify the target anomaly object currently experiencing an anomaly on the target data platform. Then, it generates a second prompt word for an anomaly handling task based on an anomaly handling guidance scheme for the target anomaly object, and executes the anomaly handling task using the second prompt word to process the anomaly of the target anomaly object. This provides an automated processing flow for the determination and handling of anomaly objects in a target data platform, meeting the high-efficiency operation and maintenance needs of big data platforms, realizing automated data operation and maintenance processing of big data platforms, and effectively improving the operation and maintenance efficiency of big data platforms.

[0005] In a first aspect, embodiments of this application provide a data processing method, including: Based on the anomaly analysis guidance scheme for the anomaly analysis task input, a first prompt word for the anomaly analysis task is generated, wherein the anomaly analysis guidance scheme is used to indicate: the analysis logic of anomaly objects in the target data platform based on anomaly object data; Based on the first prompt word, obtain the abnormal object data of the target data platform, and perform the abnormal analysis task based on the first prompt word and the abnormal object data to determine the target abnormal object that is currently experiencing an anomaly in the target data platform; A second prompt word for the exception handling task is generated based on the exception handling guidance scheme for the target exception object. The exception handling guidance scheme is used to indicate the processing logic for the identified target exception object. The exception handling task is executed based on the second prompt word to handle the exception of the target exception object.

[0006] Secondly, embodiments of this application provide a data processing apparatus, comprising: The first generation unit is used to generate a first prompt word for the anomaly analysis task based on the anomaly analysis guidance scheme input for the anomaly analysis task, wherein the anomaly analysis guidance scheme is used to indicate: the analysis logic of anomaly objects in the target data platform based on anomaly object data; The first execution unit is configured to obtain abnormal object data of the target data platform based on the first prompt word, and execute the abnormal analysis task based on the first prompt word and the abnormal object data to determine the target abnormal object that is currently experiencing an anomaly in the target data platform; The second generation unit is used to generate a second prompt word for the exception handling task based on the exception handling guidance scheme of the target exception object. The exception handling guidance scheme is used to indicate the processing logic for the identified target exception object. The second execution unit is used to execute the exception handling task based on the second prompt word in order to handle the exception of the target exception object.

[0007] Thirdly, embodiments of this application also provide an electronic device, including a memory storing multiple instructions; a processor loading instructions from the memory to execute the steps of any of the data processing methods provided in embodiments of this application.

[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute the steps of any of the data processing methods provided in embodiments of this application.

[0009] Fifthly, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in any of the data processing methods provided in embodiments of this application.

[0010] The solution adopted in this application embodiment can construct a first prompt word for an anomaly analysis task based on an anomaly analysis guidance scheme, and execute the anomaly analysis task using the first prompt word to identify the target anomaly object currently experiencing an anomaly on the target data platform; then, based on the anomaly handling guidance scheme for the target anomaly object, a second prompt word for anomaly handling task is generated, and the anomaly handling task is executed using the second prompt word to handle the anomaly of the target anomaly object; this provides an automated processing flow for the determination and handling of anomaly objects in the target data platform, meeting the high-efficiency operation and maintenance needs of big data platforms, realizing automated data operation and maintenance processing of big data platforms, and effectively improving the operation and maintenance efficiency of big data platforms. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of a data processing system provided in the embodiments of this application; Figure 2 This is a schematic flowchart of one embodiment of the data processing method provided in this application. Figure 3 This is a flowchart of a big data platform provided in the embodiments of this application; Figure 4 This is a schematic flowchart of another embodiment of the data processing method provided in this application; Figure 5 This is a schematic diagram of the pre-plan script provided in the embodiments of this application; Figure 6 This is a schematic diagram of the data processing method provided in the embodiments of this application, illustrating the available data. Figure 7 This is a schematic diagram of the structure of the data processing device provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. At the same time, in the description of the embodiments of this application, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0014] This application provides a data processing method, apparatus, electronic device, and computer-readable storage medium. Specifically, this embodiment will be described from the perspective of a data processing apparatus, which can be integrated into an electronic device. That is, the data processing method of this application embodiment can be executed by an electronic device. Optionally, the electronic device may include a terminal device. The terminal device may be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC), etc.

[0015] The data processing method provided in this application can be applied to interactive systems, such as terminal devices and servers. The terminal can be a device that includes both receiving and transmitting hardware, i.e., a device with receiving and transmitting hardware capable of performing bidirectional communication over a bidirectional communication link. The terminal device and the server can communicate bidirectionally via a network.

[0016] Optionally, the server can be a standalone server, or a server network or server cluster, including but not limited to computers, network hosts, single network servers, multiple network server sets, or cloud servers composed of multiple servers. Cloud servers consist of a large number of computers or network servers based on cloud computing.

[0017] In one embodiment of this disclosure, the data processing method can run on a local terminal device or a server. When the game interaction method runs on a server, the method can be implemented and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and a client device.

[0018] Please see Figure 1 , Figure 1This is a schematic diagram of a data processing system provided in an embodiment of this application. The system may include at least one terminal, at least one server, at least one database, and a network. A user's terminal can connect to different servers via the network. The terminal is any device with computing hardware capable of supporting and executing software products corresponding to model generation. Furthermore, when the system includes multiple terminals, multiple servers, and multiple networks, different terminals can connect to each other through different networks and servers. The network can be a wireless network or a wired network, such as a wireless local area network (WLAN), local area network (LAN), cellular network, 2G network, 3G network, 4G network, 5G network, etc. Additionally, different terminals can also connect to other terminals or servers using their own Bluetooth networks or hotspot networks. For example, multiple users can connect online through different terminals via appropriate networks and synchronize with each other to support multi-user access. Furthermore, the system may include multiple databases coupled to different servers, and can continuously store information related to the operating environment in the databases while different users are using the system online.

[0019] The following detailed description is provided in conjunction with the accompanying drawings. In this embodiment, the execution subject is a terminal device as an example. It should be noted that the order of description in the following embodiments is not intended to limit the preferred order of the embodiments. Although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than that shown in the accompanying drawings.

[0020] Please see Figure 2 , Figure 2 This is a flowchart illustrating a data processing method provided in an embodiment of this application. The specific flow of the data processing method can be summarized in steps 101 to 104, wherein: Step 101: Based on the anomaly analysis guidance scheme input for the anomaly analysis task, generate the first prompt word for the anomaly analysis task, wherein the anomaly analysis guidance scheme is used to indicate: the analysis logic of anomaly objects in the target data platform based on anomaly object data.

[0021] The anomaly analysis guidance scheme can be a first thought chain. For example, a specific thought chain description includes: 1) Obtaining alarm information, monitoring information, inspection anomalies, dynamic rules, cluster load, etc., within the period; 2) Obtaining relevant logs of abnormal tasks or clusters, and some basic information of the tasks (priority and / or tags, etc.); 3) Classifying and counting the anomaly types of log and alarm issues; 4) If possible, obtaining the possible anomaly types corresponding to various anomaly logs through tools; 5) Outputting in standard JSON format, including anomaly type, occurrence count, whether it is a cluster anomaly, severity level, log and monitoring data, judgment basis, etc. The detailed description of the thought chain also needs to provide sufficient descriptive examples. For example, the data volume should not differ too much from yesterday's year-on-year, otherwise it is a traffic anomaly; another example is that failover counts that have not occurred consecutively within 3 periods are not considered anomalies. This example can be dynamically updated based on historical experience summarized manually.

[0022] Among them, the operation and maintenance expert intelligent agent of the AI ​​large language model can generate the first prompt word of the anomaly analysis task based on the anomaly analysis guidance scheme input for the anomaly analysis task.

[0023] Step 102: Obtain abnormal object data of the target data platform based on the first prompt word, and perform the abnormal analysis task based on the first prompt word and the abnormal object data to determine the target abnormal object currently experiencing an anomaly in the target data platform.

[0024] Specifically, the step "obtaining abnormal object data of the target data platform based on the first prompt word" includes: Based on the first prompt word, the task abnormality data and / or cluster abnormality data within a specified period are obtained through the first interface corresponding to the first agent of the target large model.

[0025] Specifically, data can be obtained through interface tools of the Model Context Protocol (MCP). There are two ways for large AI models to acquire data: one is to aggregate all data and submit it to the large AI model for analysis; the other is for the AI ​​model's agent to call various MCP interfaces to obtain information. Due to the intelligence of the agent, allowing the agent to autonomously decide which data is needed is more efficient than acquiring all data at once. Currently provided data interfaces include... Figure 6 As shown, Figure 6The MCP TOOL interface showcases the data types available and details of the retrieved data. All MCP interfaces return results based on the input start and end times, thus supporting AI dynamic semantics. Simply tell the large model using prompts that you need to query data from the most recent hour and compare it with the data from the previous hour to check volatility. The large model will then use the tools' functions to retrieve data from the current and previous periods. The large model's agent is capable of self-analysis and iteration; for example, for year-on-year and month-on-month data comparisons, the AI ​​large model will automatically input historical timestamps to obtain historical statistics.

[0026] Furthermore, the step "performing the anomaly analysis task based on the first prompt word and the anomaly object data to determine the target anomaly object currently experiencing an anomaly on the target data platform" includes: Based on the first prompt word, the task anomaly data, and / or the cluster anomaly data, the anomaly analysis task is executed to determine the anomaly task or anomaly platform currently experiencing anomalies in the target data platform, which is then identified as the target anomaly object.

[0027] Exception objects can include exception tasks and exception platforms.

[0028] In one embodiment, the step "execute the anomaly analysis task based on the first prompt word, the task anomaly data, and / or the cluster anomaly data to determine the anomaly task or anomaly platform currently experiencing anomalies in the target data platform, as the target anomaly object" includes: Based on the first prompt word, the task anomaly data, and / or the cluster anomaly data, the anomaly analysis task is executed to determine whether there is a platform anomaly. If a platform anomaly exists, the abnormal platform corresponding to the platform anomaly is identified as the target anomaly object, and the anomaly type and anomaly level of the abnormal platform are determined based on the first preset rule.

[0029] The first preset rule is a set of classification standards and judgment criteria for determining the type and level of anomalies on an abnormal platform. Specifically, the first preset rule may include anomaly type determination rules and anomaly level determination rules. The anomaly type determination rules may include resource anomaly determination rules, service anomaly determination rules, storage anomaly determination rules, and configuration anomaly determination rules, etc. The specific anomaly type can be determined by judging whether the above anomaly type determination rules are met based on the task anomaly data and / or cluster anomaly data. For example, if the data corresponding to CPU or memory meets the resource anomaly determination rules, the anomaly type is determined to be a resource anomaly type; if there is an interface timeout, service crash, or process exit, it meets the service anomaly determination rules and is determined to be a service anomaly type; if there is a read / write failure or mount anomaly, it meets the storage anomaly determination rules and is determined to be a storage anomaly type; if there is a permission anomaly or access interception, it meets the configuration anomaly determination rules and is determined to be a configuration anomaly type.

[0030] The rules for determining the level of anomalies include multiple anomaly levels and corresponding level determination rules for each anomaly type. Specifically, these can include very severe level (level 0), severe level (level 1), moderate level (level 2), and minor level (level 3). Each anomaly type determination rule includes at least one indicator threshold corresponding to the anomaly type. For example, for the resource anomaly type, the indicator threshold can be 95% and the duration can be greater than 10 minutes to be classified as level 0, the indicator threshold can be 75% and the duration can be greater than 10 minutes to be classified as level 1, the indicator threshold can be 65% and the duration can be greater than 10 minutes to be classified as level 2, and the indicator threshold can be 55% and the duration can be greater than 10 minutes to be classified as level 3. In this case, when the CPU utilization is detected to exceed 95% and the duration is greater than 10 minutes, the abnormal platform can be determined to be a resource anomaly type and the anomaly level is level 0.

[0031] In one specific embodiment, when the agent of the large model determines that a large number of tasks in the real-time cluster are starting to alarm with Checkpoint execution failures, and a large number of offline tasks are also failing, the alarm triggers analysis logic. Since the above situations occur simultaneously and on a large scale, the large model will analyze and determine that the above situations have triggered a platform-level anomaly, and identify the abnormal platform as the target anomaly object. Next, the agent can obtain audit information to rule out platform-level anomalies caused by updates to the big data platform. At the same time, it can check the cluster monitoring to determine whether the cluster machines or network are normal, and search the logs to check the key error logs. It can also determine whether there are sudden increases or decreases in traffic in the monitoring data. Finally, based on the comprehensive analysis of the above data, it can rule out big data platform changes, cluster anomalies, etc. If a large number of HDFS router connection anomalies are found from the key anomaly logs, the agent can determine that this is a platform-level anomaly caused by an external HDFS router, and can contact HDFS operations and maintenance for resolution.

[0032] Specifically, monitoring data, including upstream and downstream data in the task (such as data quality and completeness), can be used to comprehensively judge whether the overall operation is normal. For some traffic data, such as the app logs missing a certain version of the system log (such as the client version of system 14 being abnormal), or a sudden drop in traffic (such as data collection, gateway and data center), or multiple tasks failing frequently (such as task processing clusters), it is necessary to determine whether there is an anomaly at the platform level and handle it.

[0033] Optionally, the method further includes: If there is no platform anomaly, the anomaly analysis task will continue to be executed based on the first prompt word, the task anomaly data and / or the cluster anomaly data to determine whether there is a task anomaly. If so, the abnormal task corresponding to the task abnormality problem is determined as the target abnormal object, and the abnormality problem type and abnormality problem level of the abnormal task are determined based on the second preset rule.

[0034] The second preset rule is a standardized and automatically executable set of judgment criteria for classifying abnormal tasks into abnormal problem types and determining their abnormal problem levels.

[0035] Specifically, the second preset rule may include rules for determining the type of anomaly and rules for determining the level of anomaly. The rules for determining the type of anomaly may include rules for determining task execution anomalies, rules for determining task resource usage anomalies, rules for determining task data anomalies, etc. Based on the abnormal task data and / or abnormal cluster data, it is determined whether the above rules for determining the type of anomaly are met, thus identifying the specific type of anomaly. For example, if there is execution failure or execution timeout, it meets the rules for determining task execution anomalies, and the anomaly type can be determined as a task execution anomaly. If a sudden increase in CPU or memory on the task side is detected, it meets the rules for determining task resource usage anomalies, and the anomaly type can be determined as a task resource usage anomaly. If invalid input data, incorrect data format, empty data, or duplicate data is detected, it meets the rules for determining task data anomalies, and the anomaly type can be determined as a task data anomaly.

[0036] The rules for determining the level of anomalies include multiple anomaly levels and corresponding level determination rules for each anomaly type. Specifically, these can include extremely severe level (level 0), severe level (level 1), moderate level (level 2), and minor level (level 3). Each anomaly type determination rule includes at least one indicator threshold corresponding to the anomaly type. For example, for the task execution anomaly type, the indicator threshold can be that the task fails more than 10 times consecutively and cannot recover on its own, which is classified as level 0; the indicator threshold can be that the task fails more than 8 times consecutively and cannot recover on its own, which is classified as level 1; the indicator threshold can be that the task fails more than 6 times consecutively and cannot recover on its own, which is classified as level 2; and the indicator threshold can be that the task fails more than 2 times consecutively, which is classified as level 3.

[0037] For example, if there is no platform anomaly, the anomaly analysis task continues to be executed based on the first prompt word, the task anomaly data, and / or the cluster anomaly data to determine whether there is a task anomaly. Specifically, the key anomalies of the big data task and the alarm information of the past hour can be viewed, combined with the input data volume and output data volume of the past hour, and compared with the previous data to determine whether the big data task may have anomalies.

[0038] Specifically, task anomalies are mainly detected through monitoring data. For example, polling monitoring may reveal that the output traffic has dropped to zero, whereas there was output before. Therefore, all task-related information can be fed into the AI ​​model. The model can then determine that this zeroing is a problem (e.g., the input has not changed, but the output has dropped to zero). After searching for various information, it may be due to severe backpressure monitoring and the logs may contain timeouts. This could indicate an anomaly in the downstream data storage service, which will trigger a task anomaly alert.

[0039] Optionally, the timing for judging task anomalies includes the following two types: 1) Task alarms: The alarm conditions are configured by the user. The built-in alarm conditions are mainly failure, offline, timeout, real-time failure, etc. Other alarm conditions can also be configured by the user, such as traffic monitoring, lag, write latency, read latency, memory GC count, CKP count and time, etc., with a lot of free choices. 2) Rotation training: Every 30 minutes, a comprehensive judgment is made on the traffic data, number of alarms, and number of abnormal logs to see if it is necessary to trigger the platform's abnormal judgment. This standard is mainly defined by the AI ​​big model itself. The AI ​​big model can make abnormal judgments based on the input prompt words. The prompt words can be: check the key abnormality (RootException) of the task, alarm information in the past hour, combined with the input data volume in the past hour, output data volume, and compared with the previous data to determine whether the task may have abnormalities.

[0040] Step 103: Generate a second prompt word for the exception handling task based on the exception handling guidance scheme for the target exception object. The exception handling guidance scheme is used to indicate the processing logic for the identified target exception object.

[0041] The exception handling guidance scheme can be a second thought chain. For example, a specific thought chain description includes: 1) Obtaining a solution description through tools based on the exception information; 2) Determining whether the solution is suitable for solving the problem; if so, outputting the script ID; 3) Outputting the people who need to be notified of the problem; if there are contacts in the solution, outputting them; otherwise, specifying the platform duty officer (xxxx); step 4: Standard format output (exception, whether to execute the script, script ID, notify contacts). Downstream, the code program is called to start the session and execute the standard script according to the results output by the AI ​​large model.

[0042] Step 104: Execute the exception handling task based on the second prompt word to handle the exception of the target exception object.

[0043] In one embodiment, the exception handling guidance scheme includes specifying an exception problem level, a third preset rule, and an exception problem handling scheme; the step "executing the exception handling task based on the second prompt word to handle the exception of the target exception object" includes: Based on the aforementioned anomaly level and the third preset rule, determine whether the anomaly handling scheme can be used to process the target anomaly object; If so, the anomaly of the target anomaly object is processed based on the anomaly handling scheme.

[0044] In this embodiment, an AI big data model can determine whether an anomaly handling solution can be used to handle the target anomaly based on the anomaly level and a third preset rule. If so, the AI ​​big data model handles the anomaly of the target anomaly using the anomaly handling solution; otherwise, a contact person is notified to handle the issue manually. The third preset rule is a standardized judgment criterion used by the AI ​​big data model after obtaining an anomaly handling solution for the target anomaly, with the anomaly level as the core constraint, to automatically determine whether the anomaly handling solution can be executed or not. Specifically, the third preset rule can include multiple execution judgment logics. The first execution judgment logic is that for anomaly handling solutions with an anomaly level lower than or equal to the first preset level (e.g., level 1), the anomaly of the target anomaly can be handled by the AI ​​big data model. The second execution judgment logic is that for anomaly handling solutions with an anomaly level higher than the first preset level (e.g., level 1), the anomaly of the target anomaly cannot be handled by the AI ​​big data model, and a contact person needs to be notified for manual handling. For example, when the anomaly level is level 2, the anomaly of the target anomaly object can be handled by an AI large model based on the anomaly handling scheme.

[0045] In another embodiment, the system can automatically determine whether to execute or not execute the abnormal problem handling solution based on the feasibility of execution, risk controllability, business compliance, and solution effectiveness. Execution feasibility assesses whether the necessary resources, permissions, environment, and dependencies for implementing the anomaly handling solution are available. Risk controllability assesses whether implementing the anomaly handling solution will cause secondary failures, expand the scope of the anomaly, or lead to business interruption. Business compliance assesses whether implementing the anomaly handling solution conforms to business strategies, operational standards, or data security requirements. Solution effectiveness assesses whether the implemented anomaly handling solution matches the anomaly type and level, and whether there are obvious logical errors, irrelevance to the root cause, or invalid historical verification. When assessing the execution feasibility, risk controllability, business compliance, and solution effectiveness based on the anomaly handling solution, if at least one of the following conditions exists, the AI ​​large model cannot be used to handle the anomaly of the target anomaly object based on the anomaly handling solution. Specific conditions may include: lack of necessary resources, permissions, environment, and dependencies for the anomaly handling solution; execution of the anomaly handling solution will cause secondary failures, expand the scope of the anomaly, or lead to business interruption; non-compliance with business strategies, operational standards, or data security requirements; mismatch between the implemented anomaly handling solution and the anomaly type and level; obvious logical errors, irrelevance to the root cause, or invalid historical verification.

[0046] Optionally, the exception handling guidance scheme further includes specifying the communication channel and the target user account associated with the target exception object, and the method further includes: If not, an exception message is generated based on the target exception object, and the exception message is sent to the target user account through the designated communication channel to notify the user of the target user account that the target exception object has an exception.

[0047] Specifically, in this embodiment of the application, one or more intelligent agents are defined as analysis expert roles through an AI big model. Based on the information such as the type, frequency, and description of the problem anomaly provided by the operation and maintenance expert role, the contingency plan tools are called to obtain contingency plan information to analyze whether the contingency plan can be executed. If the script can be executed automatically, the script of the contingency plan is executed automatically to resolve the anomaly. If not, the session is initiated and relevant personnel are notified.

[0048] Based on the above description, the following examples will further illustrate the data processing method of this application. Specific embodiments of the data processing method are described below.

[0049] When the large-scale model's agent detects that a large number of tasks in the real-time cluster are failing checkpoint execution and offline tasks are also failing in large numbers, the alarm triggers analysis logic. Because these situations occur simultaneously and on a large scale, the large model will analyze and determine that these situations have triggered a platform-level anomaly. Next, the agent can obtain audit information to rule out platform-level anomalies caused by updates to the big data platform. Simultaneously, it checks cluster monitoring to determine if the cluster machines or network are functioning normally, searches logs for critical error logs, and checks for sudden increases or decreases in monitoring data. Finally, based on a comprehensive analysis of the above data, it rules out big data platform changes, cluster anomalies, etc. If a large number of HDFSrouter connection anomalies are found in the critical anomaly logs, the agent can determine that this is a platform-level anomaly caused by an external HDFS router and can contact HDFS operations and maintenance for resolution.

[0050] In this embodiment, traffic fluctuations exceeding a first specified threshold compared to the same hourly rate, and the number of similar anomalies exceeding a second specified threshold, can be used to determine if a platform-level anomaly exists. In this embodiment, the large model performs polling to detect anomalies, first examining the aggregated data from the big data platform. If a platform-level anomaly is found, subsequent polling of single tasks can be paused to conserve resources.

[0051] For example, if a big data task encounters an anomaly due to dirty data, the AI ​​of the big data model will determine it as a single-task anomaly. After searching for contingency plans, it will find that the anomaly of this big data task can be ignored and parsed. So it will execute a script to modify the big data task configuration to allow parsing anomalies, and then restart the big data task.

[0052] For example, if multiple people in the cluster report a database connection failure, it is determined to be an external database anomaly. The contingency plan indicates that the automatic script cannot be executed automatically. The contingency plan also includes contact information for business line A and business line B. If the system determines that the database of business line A is abnormal, it will use a public account (such as an account on an internal chat tool) to initiate a multi-person conversation with the person in charge of the affected task and the platform manager through the system, thereby simultaneously notifying all parties to intervene and handle the situation quickly.

[0053] For details, please refer to Figure 3 , Figure 3 This is a flowchart of a big data platform provided in this application embodiment. The specific processing logic is as follows: (1) Obtain the monitoring data, log information, rule alarms (i.e., historical alarm rules), and manual operation records corresponding to the big data platform, and send the above content to the intelligent agent analysis expert (i.e., the analysis expert intelligent agent) for summary and classification. Specifically, the analysis expert intelligent agent connects to monitoring tools such as Prometheus and Grafana, log analysis tools such as ELK, and the platform operation audit system through tool plugins. It combines the current alarm information with data such as Flink / Spark task running metrics, business data indicators, full-link logs, historical alarm rules, and manual operation records to comprehensively judge whether a big data task has abnormal problems. It should be noted that the original alarm rules are still triggered, and the AI ​​system will cover the existing alarm rules.

[0054] Rometheus acquires Kafka traffic data, while Flink Metrics acquires basic monitoring metrics such as memory, CPU, and network traffic input / output, as well as user-defined metrics, such as click-through rates aggregated for a specific advertising scenario, forming data in a specified reporting format. Grafana primarily collects cluster platform monitoring data, including cluster resource status, CPU, memory, network card traffic, and load data for each machine, as well as basic platform information such as log collection efficiency and network connectivity. Log analysis tools like ELK primarily acquire log data, such as the execution logs of big data tasks, Spark, and Flink. ELK allows querying big data task execution logs within a specified time range using key exceptions and keywords. The platform operation audit system obtains data including user-modified tasks, platform version updates, plugins (dependencies for big data tasks), and changes to user-defined dependency files, which also affect the execution of big data tasks.

[0055] Specifically, the entire analysis tool aggregates all alarm anomalies (weight P0), anomaly logs (weight P1), release and manual operation records (weight P2), and platform-level monitoring data (such as overall traffic, overall cluster resource utilization, and alarms) within a certain time period (e.g., five minutes) (weight P1) to comprehensively determine whether a platform-level anomaly exists. If no platform-level anomaly is found, it checks for occasional single-task alarms, triggering a single-task anomaly determination. The determination logic is similar to that of platform-level anomalies, querying the logs and monitoring data of a single task (i.e., a single big data task). The intelligent agent analysis expert uses this data, combined with prior knowledge (e.g., a pre-defined rule document), to determine if there is a problem that needs to be addressed. If so, it determines the approximate type of anomaly (which may include network type, memory type, dirty data type, cluster type, and machine type, etc.) and also determines the anomaly's severity level, for example, setting several levels from 1 to 9 from low to high. After this stage, this data is transmitted to the next operational intelligent agent for processing. For example, an example of the content of a pre-defined rule document is as follows: 1) Type: XX Exception 2) Exception: org.apache.spark.SparkException: Job aborted due to stagefailure: ExecutorLostFailure Container killed by YARN for exceeding memorylimits. 36.4 GB of 36 GB physical memory used; 3) Note: Insufficient memory. For a single task, you can modify the memory configuration and restart. For a large number of tasks, there may be insufficient resources in a certain queue. You can switch resources. 4) Weight: 5 (from low to high 1~10) 5) Contingency plan: Modify configuration and restart; 6) Contact person: Platform duty officer, task manager.

[0056] In this context, P0, P1, and P2 are weighted prompt words, primarily used to label data when it's fed to the large model. For example, an alarm anomaly has a weight of P0, indicating it's a key reference indicator; an anomaly log has a weight of P1, indicating it's a general reference indicator. If the prompt words input to the large model are: "I have 3000 tasks, no alarms in the past hour, 200 anomaly logs, aggregate result: xxxx", the large model will correspondingly lower the weights, judging the platform to be without problems; conversely, if there are 200 alarms, it will judge the platform to have an anomaly.

[0057] For example, if polling monitoring reveals that the output traffic has dropped to zero, but there was output before, all data can be sent to the large model. The large model's agent can then determine that this zeroing is an anomaly (the input has not changed, but the output has dropped to 0). It can then search various data, such as finding that the backpressure monitoring is severe and there is a timeout in the logs. This may indicate that the downstream data storage service is malfunctioning, and a platform-level anomaly will be detected.

[0058] The timing for anomaly detection in big data tasks includes the following two moments. The first moment is when big data task alarms are triggered, such as traffic monitoring, write latency, read latency, memory GC counts, CKP counts and times, etc. The second moment is during polling, where traffic data, the number of alarms, and the number of abnormal logs are comprehensively evaluated every specified interval (e.g., 30 minutes) to determine whether platform anomaly detection needs to be triggered. This is triggered by inputting prompt words into the big data model. The prompt words can be: "View key anomalies of the big data task and alarm information in the past hour." Combining the input and output data volumes in the past hour with previous data, it is determined whether the big data task may have anomalies.

[0059] The timing for platform anomaly detection on the big data platform is triggered by round-robin training at a specified time interval (e.g., 30 minutes). The judgment is made by aggregating a report of data. For example, the report data may include: the cluster traffic trend in the past hour compared with the previous hour, the number and type of cluster alarms, the number and type of big data task alarms triggered, the top 10 key anomalies and their frequency, etc. Alternatively, the platform anomaly check can be triggered by alarms. The large model can be used to determine whether to trigger the platform anomaly check.

[0060] (2) After receiving the data transmitted from the previous analysis expert agent, the operation and maintenance agent will call the contingency plan description through the corresponding tools. A standard contingency plan description document includes the following: the type of problem to be pre-processed; typical cases; the highest level of processing in this contingency plan; examples that this contingency plan cannot handle; the on-duty and operation and maintenance personnel who can be contacted for this problem; a description of the basic principles and methods; the expected effect; and whether it can be automated. Each contingency plan description is set with a pre-defined script, which is used to handle abnormal problems and resolve them. The script is a handling plan for the current abnormal problem.

[0061] Once the operations and maintenance (O&M) agent obtains the contingency plan documentation through tools, it performs a comprehensive evaluation. The evaluation result can include: a first determination: execute the script in the contingency plan document; a second determination: cannot be handled; a third determination: O&M personnel intervention is required. If the determination is the first determination, the script in the contingency plan document is executed according to the pre-defined script, and the execution status of the contingency plan document is monitored continuously. If the determination is the second determination: cannot be handled, the current round of anomaly analysis ends and a report is submitted to the system administrator. If the determination is the third determination: O&M personnel intervention is required, the contact person O&M agent directly adds relevant personnel to a group chat and notifies them of the alert through their system accounts.

[0062] Specifically, the AI ​​model can recall the keyword "restart task," and this plan has the highest weight when searched through Elasticsearch: { “id”:1 "Plan Name": Modify configuration and restart "Principles and Methods": xxx Supported types: Memory modification (mem=xxx), Concurrent modification xxxx } The AI ​​model determines whether the script can be executed based on the recall content. If it can, it determines whether to notify the user. If it determines that the script should be executed, the workflow will directly connect to a downstream interface to notify the user. The recall ID 1 will be passed down. This is implemented in code. The AI ​​model only outputs whether to call the downstream function. If the function is called, it will input an ID=1, jobid=xxxx, param=mem=15g. The downstream method is implemented in code. Restart and other operations are implemented in code according to the plan. Each plan has an independently developed set of automated scripts to automatically execute the corresponding plan. Optionally, plans that do not support automated scripts can also be set.

[0063] Specifically, the complete logical chain of this application embodiment includes tools for acquiring data through the Model Context Protocol (MCP), overall analysis of the current cluster and single task status through AI analysis, design of contingency plans, and resolution of abnormal issues through AI. This application embodiment can perform a 5-minute cycle scan of the entire system according to the above logical chain, and use CrewAI and Dify to organize the data. CrewAI is responsible for defining the roles of AI, and other business processes are connected by the Workflow.

[0064] For example, please refer to the following: Figure 4 and Figure 5 , Figure 4 This is a schematic flowchart of another embodiment of the data processing method provided in this application. Figure 5 This is a schematic diagram of the pre-plan script provided in the embodiments of this application. The specific processing logic of the data processing method provided in the embodiments of this application is as follows: (1) Tools for acquiring data through Model Context Protocol (MCP). There are two ways for large AI models to acquire data: one is to aggregate all the data and submit it to the large AI model for analysis; the other is to have the AI ​​model's agent call various MCP interfaces to obtain information. Due to the intelligence of the agent, it is more efficient to have the agent autonomously decide which data is needed than to acquire all the data at once. Currently provided data interfaces include... Figure 6 As shown, Figure 6 This is a schematic diagram illustrating the data acquisition capability of the data processing method provided in the embodiments of this application. Figure 6The MCPTOOL interface showcases the data types it can retrieve and the details of the retrieved data. All MCP interfaces return results based on the input start and end times, thus supporting AI dynamic semantics. Simply tell the large model using prompts that you need to query the data for the most recent hour and compare it with the data for the previous hour to check volatility. The large model will then use the tools' functions to retrieve data from the current and previous periods. The large model's agent can perform self-analysis and iteration; for example, when comparing year-on-year and month-on-month data, the AI ​​large model will automatically input historical timestamps to obtain historical statistics.

[0065] (2) In this embodiment of the application, one or more intelligent agents are defined as operation and maintenance experts through the AI ​​big model to guide the AI ​​big model to perform an overall analysis of the current cluster and single task status. In the prompt words description of the AI ​​big model, it is necessary to specify: 1. Analyze single task anomalies and alarms in all periods; 2. Analyze the cluster status and alarms; 3. A large number of single task anomalies need to be judged by themselves to see if there is a commonality, and whether it is a platform-level error caused by a problem. Multiple examples need to be given in the prompt words, such as a large number of IP connection anomalies may be caused by target database failure, such as network instability; 4. It is necessary to determine whether it is a cluster-level anomaly, aggregate the anomalies, and standardize the output structure. Specifically, a complete thought chain description needs to be provided to the intelligent agent. A specific thought chain description includes: Step 1: Obtain alarm information, monitoring information, inspection anomalies, dynamic rules, cluster load, etc. within the period; Step 2: Obtain relevant logs of abnormal tasks or clusters, and some basic information of the tasks (priority, tags, etc.); Step 3: Classify and count the anomaly types of log and alarm issues; Step 4: If possible, obtain the possible anomaly types corresponding to various anomaly logs through tools; Step 5: Output in standard JSON format, including anomaly type, occurrence count, whether it is a cluster anomaly, severity level, log and monitoring data, judgment basis, etc. The detailed description of the thought chain also needs to provide enough descriptive examples. For example, the data volume should not differ too much from yesterday's year-on-year, otherwise it is a traffic anomaly; another example is that if the number of failovers does not occur consecutively within 3 periods, it is not counted as an anomaly. This example can be dynamically updated based on historical experience.

[0066] The operations and maintenance expert role includes three agents: the first agent is used to execute plan (thinking chain) prompts, the second agent is used to determine analysis prompts, including analysis steps and solutions, and the third agent is used to format and aggregate output model prompts.

[0067] (3) Contingency Plan Design. A contingency plan refers to the handling method for a specified type of exception, including automated execution scripts. It includes: supported exception types, contingency script ID, contingency plan description, the highest severity level for automated handling, and a list of relevant contacts. For example, please refer to... Figure 5 The contingency plan includes automated scripts. This diagram illustrates the automated solution for data parsing anomalies, modifying configuration parameters (e.g., modifying configuration parameters for running big data tasks, big data platform configuration parameters, and configuration parameters in the SQL of user tasks). The contingency plan addresses dirty data parsing anomalies. Executing the script follows the logic shown in the diagram, allowing configuration modifications and task restarts. The corresponding priority is P2, representing tasks that can lose data, while high-priority tasks cannot. An AI model can be used to determine whether the script can be executed. Each big data task has a priority. The priority level in the contingency plan determines whether the current big data task's configuration can be modified. For example, if the priority level in the contingency plan is P2, and the big data task's priority is also P2, since P2 is internally defined as unimportant and can lose data, the contingency plan script can be executed to modify the task configuration.

[0068] (4) In this embodiment of the application, one or more intelligent agents are defined as analysis expert roles through the AI ​​big model. Based on the information such as the type, frequency, and description of the problem anomaly provided by the operation and maintenance expert role, the contingency plan tools are called to obtain contingency plan information to analyze whether the contingency plan can be executed. If the script can be executed automatically, the script of the contingency plan is executed automatically to solve the abnormal problem. If not, the session is tried to be started and relevant personnel are notified. Specifically, a complete thought chain description needs to be provided to the intelligent agent. A specific thought chain description includes: step1: Obtain the solution description through tools based on the anomaly information; step2: Determine whether the solution is suitable for solving the problem. If so, output the script ID; step3: Output the people who need to be notified of the problem. If there are contacts in the solution, output these. If not, specify the platform duty officer xxxx; step4: Output in standard format (anomaly, whether to execute the script, script ID, notifier). The downstream calls the code program to start the session and execute the standard script according to the results output by the AI ​​big model.

[0069] In one specific embodiment, the keyword "prompt" for platform-level anomaly monitoring can be: Please determine the severity level based on the number of times the problem occurs, the type of problem, and the description of the type: "anomaly," "warning," or "notification." Then, the severity of this type of problem can be described using levels 1-10 in the recall information, allowing you to make your own judgment. For example, if the test Kafka cluster is disconnected, many tasks will trigger alarms indicating that Kafka is inaccessible. The AI ​​large model will collect data to trigger platform-level anomaly detection. After retrieving task and log information, it will be found that Kafka is inaccessible, with a severity level of 8. The tasks are all test tasks with a priority of P2. Based on this comprehensive judgment, a warning is given.

[0070] In summary, this application provides a data processing method. Based on an anomaly analysis guidance scheme input for an anomaly analysis task, a first prompt word for the anomaly analysis task is generated. The anomaly analysis guidance scheme instructs the analysis logic for anomaly objects in a target data platform based on anomaly object data. Then, based on the first prompt word, the anomaly object data of the target data platform is obtained. The anomaly analysis task is then executed based on the first prompt word and the anomaly object data to determine the target anomaly object currently experiencing an anomaly on the target data platform. Next, based on the anomaly handling guidance scheme for the target anomaly object, a second prompt word for an anomaly handling task is generated. The anomaly handling guidance scheme instructs the processing logic for the identified target anomaly object. Finally, the anomaly handling task is executed based on the second prompt word to process the anomaly of the target anomaly object. The method provided in this application embodiment can construct a first prompt word for an anomaly analysis task based on an anomaly analysis guidance scheme, and execute the anomaly analysis task using the first prompt word to identify the target anomaly object currently experiencing an anomaly on the target data platform; then, based on the anomaly handling guidance scheme for the target anomaly object, a second prompt word for anomaly handling task is generated, and the anomaly handling task is executed using the second prompt word to handle the anomaly of the target anomaly object; this provides an automated processing flow for the determination and handling of anomaly objects in the target data platform, meeting the high-efficiency operation and maintenance needs of big data platforms, realizing automated data operation and maintenance processing of big data platforms, and effectively improving the operation and maintenance efficiency of big data platforms.

[0071] This embodiment also provides a data processing device, which can be specifically integrated into a terminal device. For example, such as Figure 7 As shown, the data processing device may include: The first generation unit 201 is used to generate a first prompt word for the anomaly analysis task based on the anomaly analysis guidance scheme input for the anomaly analysis task, wherein the anomaly analysis guidance scheme is used to indicate: the analysis logic of anomaly objects in the target data platform based on anomaly object data; The first execution unit 202 is used to obtain abnormal object data of the target data platform based on the first prompt word, and to perform the abnormal analysis task based on the first prompt word and the abnormal object data to determine the target abnormal object that is currently abnormal in the target data platform; The second generation unit 203 is used to generate a second prompt word for the exception handling task based on the exception handling guidance scheme of the target exception object. The exception handling guidance scheme is used to indicate the processing logic for the identified target exception object. The second execution unit 204 is used to execute the exception handling task based on the second prompt word in order to handle the exception of the target exception object.

[0072] In some embodiments, the data processing apparatus includes a processing subunit for: Based on the first prompt word, the task abnormality data and / or cluster abnormality data within a specified period are obtained through the first interface corresponding to the first agent of the target large model.

[0073] In some embodiments, the data processing apparatus includes a processing subunit for: Based on the first prompt word, the task anomaly data, and / or the cluster anomaly data, the anomaly analysis task is executed to determine the anomaly task or anomaly platform currently experiencing anomalies in the target data platform, which is then identified as the target anomaly object.

[0074] In some embodiments, the data processing apparatus includes a processing subunit for: Based on the first prompt word, the task anomaly data, and / or the cluster anomaly data, the anomaly analysis task is executed to determine whether there is a platform anomaly. If a platform anomaly exists, the abnormal platform corresponding to the platform anomaly is identified as the target anomaly object, and the anomaly type and anomaly level of the abnormal platform are determined based on the first preset rule.

[0075] In some embodiments, the data processing apparatus includes a processing subunit for: If there is no platform anomaly, the anomaly analysis task will continue to be executed based on the first prompt word, the task anomaly data and / or the cluster anomaly data to determine whether there is a task anomaly. If so, the abnormal task corresponding to the task abnormality problem is determined as the target abnormal object, and the abnormality problem type and abnormality problem level of the abnormal task are determined based on the second preset rule.

[0076] In some embodiments, the data processing apparatus includes a processing subunit for: Based on the aforementioned anomaly level and the third preset rule, determine whether the anomaly handling scheme can be used to process the target anomaly object; If so, the anomaly of the target anomaly object is processed based on the anomaly handling scheme.

[0077] In some embodiments, the data processing apparatus includes a processing subunit for: If not, an exception message is generated based on the target exception object, and the exception message is sent to the target user account through the designated communication channel to notify the user of the target user account that the target exception object has an exception.

[0078] This application discloses a data processing apparatus. A first generation unit 201 generates a first prompt word for the anomaly analysis task based on an anomaly analysis guidance scheme input for the anomaly analysis task. The anomaly analysis guidance scheme instructs the analysis logic for anomaly objects in a target data platform based on anomaly object data. A first execution unit 202 obtains the anomaly object data of the target data platform based on the first prompt word and executes the anomaly analysis task based on the first prompt word and the anomaly object data to determine the target anomaly object currently experiencing an anomaly on the target data platform. A second generation unit 203 generates a second prompt word for the anomaly handling task based on the anomaly handling guidance scheme for the target anomaly object. The anomaly handling guidance scheme instructs the processing logic for the identified target anomaly object. A second execution unit 204 executes the anomaly handling task based on the second prompt word to process the anomaly of the target anomaly object. This application embodiment can construct a first prompt word for an anomaly analysis task based on an anomaly analysis guidance scheme, and use the first prompt word to execute the anomaly analysis task to identify the target anomaly object currently experiencing an anomaly on the target data platform; then, based on the anomaly handling guidance scheme for the target anomaly object, generate a second prompt word for anomaly handling task, and use the second prompt word to execute the anomaly handling task to handle the anomaly of the target anomaly object; this provides an automated processing flow for the determination and handling of anomaly objects in the target data platform, meeting the high-efficiency operation and maintenance needs of big data platforms, realizing automated data operation and maintenance processing of big data platforms, and effectively improving the operation and maintenance efficiency of big data platforms.

[0079] Accordingly, this application also provides an electronic device, which can be a terminal, such as a smartphone, tablet computer, laptop computer, touch screen, game console, personal computer (PC), personal digital assistant (PDA), or other terminal device. Alternatively, the electronic device can be a server.

[0080] like Figure 8 As shown, Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 300 includes a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, and a computer program stored in the memory 302 and executable on the processor. The processor 301 and the memory 302 are electrically connected. Those skilled in the art will understand that the electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0081] The processor 301 is the control center of the electronic device 300. It connects various parts of the electronic device 300 via various interfaces and lines. By running or loading software programs and / or units stored in the memory 302, and by calling data stored in the memory 302, it executes various functions and processes data of the electronic device 300, thereby providing overall monitoring of the electronic device 300. The processor 301 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0082] In this embodiment, the processor 301 in the electronic device 300 loads the instructions corresponding to the processes of one or more applications into the memory 302 according to the following steps, and the processor 301 runs the applications stored in the memory 302 to realize various functions, such as: Based on the anomaly analysis guidance scheme for the anomaly analysis task input, a first prompt word for the anomaly analysis task is generated, wherein the anomaly analysis guidance scheme is used to indicate: the analysis logic of anomaly objects in the target data platform based on anomaly object data; Based on the first prompt word, obtain the abnormal object data of the target data platform, and perform the abnormal analysis task based on the first prompt word and the abnormal object data to determine the target abnormal object that is currently experiencing an anomaly in the target data platform; A second prompt word for the exception handling task is generated based on the exception handling guidance scheme for the target exception object. The exception handling guidance scheme is used to indicate the processing logic for the identified target exception object. The exception handling task is executed based on the second prompt word to handle the exception of the target exception object.

[0083] The electronic device provided in this application embodiment can construct a first prompt word for an anomaly analysis task based on an anomaly analysis guidance scheme, and execute the anomaly analysis task using the first prompt word to determine the target anomaly object currently experiencing an anomaly on the target data platform; then, based on the anomaly handling guidance scheme for the target anomaly object, it generates a second prompt word for anomaly handling task, and executes the anomaly handling task using the second prompt word to handle the anomaly of the target anomaly object; this provides an automated processing flow for the determination and handling of anomaly objects in the target data platform, meeting the high-efficiency operation and maintenance needs of big data platforms, realizing automated data operation and maintenance processing of big data platforms, and effectively improving the operation and maintenance efficiency of big data platforms.

[0084] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0085] Optional, such as Figure 8 As shown, the electronic device 300 also includes: a touch display screen 303, a radio frequency circuit 304, an audio circuit 305, an input unit 306, and a power supply 307. The processor 301 is electrically connected to the touch display screen 303, the radio frequency circuit 304, the audio circuit 305, the input unit 306, and the power supply 307. Those skilled in the art will understand that... Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0086] The touch display screen 303 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. The touch display screen 303 may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Optionally, the display panel can be configured using a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar technologies. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), generate corresponding operation commands, and execute the corresponding program according to the operation commands. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, transmitting the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 301. It can also receive and execute commands from the processor 301. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits the information to the processor 301 to determine the type of touch event. Subsequently, the processor 301 provides corresponding visual output on the display panel based on the type of touch event. In this embodiment, the touch panel and the display panel can be integrated into the touch display screen 303 to achieve input and output functions. However, in some embodiments, the touch panel and the touch display screen 303 can be implemented as two independent components to achieve input and output functions. That is, the touch display screen 303 can also be used as part of the input unit 306 to achieve input functions.

[0087] The radio frequency circuit 304 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other electronic devices, and to transmit and receive signals with network devices or other electronic devices.

[0088] Audio circuitry 305 can be used to provide an audio interface between a user and an electronic device via a speaker and a microphone. Audio circuitry 305 converts received audio data into electrical signals, transmits them to the speaker, and the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 305, converted back into audio data, and then processed by processor 301 before being transmitted via radio frequency circuitry 304 to, for example, another electronic device, or output to memory 302 for further processing. Audio circuitry 305 may also include an earphone jack to facilitate communication between peripheral headphones and electronic devices.

[0089] The input unit 306 can be used to receive input numbers, characters, or user characteristic information (such as fingerprints, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.

[0090] Power supply 307 is used to supply power to various components of electronic device 300. Optionally, power supply 307 can be logically connected to processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 307 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0091] although Figure 8 As not shown in the diagram, the electronic device 300 may also include a camera, sensor, wireless fidelity module, Bluetooth module, etc., which will not be described in detail here.

[0092] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0093] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0094] Therefore, embodiments of this application provide a computer-readable storage medium storing multiple computer programs that can be loaded by a processor to execute any of the data processing methods provided in embodiments of this application. The computer program can execute the steps of the following data processing method: Based on the anomaly analysis guidance scheme for the anomaly analysis task input, a first prompt word for the anomaly analysis task is generated, wherein the anomaly analysis guidance scheme is used to indicate: the analysis logic of anomaly objects in the target data platform based on anomaly object data; Based on the first prompt word, obtain the abnormal object data of the target data platform, and perform the abnormal analysis task based on the first prompt word and the abnormal object data to determine the target abnormal object that is currently experiencing an anomaly in the target data platform; A second prompt word for the exception handling task is generated based on the exception handling guidance scheme for the target exception object. The exception handling guidance scheme is used to indicate the processing logic for the identified target exception object. The exception handling task is executed based on the second prompt word to handle the exception of the target exception object.

[0095] Because the computer program stored in this storage medium can construct a first prompt word for an anomaly analysis task based on an anomaly analysis guidance scheme, and execute the anomaly analysis task using the first prompt word to identify the target anomaly object currently experiencing an anomaly on the target data platform; then, based on the anomaly handling guidance scheme for the target anomaly object, a second prompt word for anomaly handling task is generated, and the anomaly handling task is executed using the second prompt word to handle the anomaly of the target anomaly object; this provides an automated processing flow for the determination and handling of anomaly objects in the target data platform, meeting the high-efficiency operation and maintenance needs of big data platforms, realizing automated data operation and maintenance processing of big data platforms, and effectively improving the operation and maintenance efficiency of big data platforms.

[0096] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0097] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0098] Since the computer program stored in the computer-readable storage medium can execute any of the data processing methods provided in the embodiments of this application, the beneficial effects that any of the data processing methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0099] According to one aspect of this application, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in the various optional implementations of the above embodiments.

[0100] In the above embodiments of the data processing apparatus, computer-readable storage medium, electronic device, and computer program product, the descriptions of each embodiment have different focuses. Parts not described in detail in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes and beneficial effects of the data processing apparatus, computer-readable storage medium, computer program product, electronic device, and their corresponding units described above can be referred to the description of the data processing methods in the above embodiments, and will not be repeated here.

[0101] The foregoing has provided a detailed description of a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A data processing method, characterized in that, include: Based on the anomaly analysis guidance scheme for the anomaly analysis task input, a first prompt word for the anomaly analysis task is generated, wherein the anomaly analysis guidance scheme is used to indicate: the analysis logic of anomaly objects in the target data platform based on anomaly object data; Based on the first prompt word, obtain the abnormal object data of the target data platform, and perform the abnormal analysis task based on the first prompt word and the abnormal object data to determine the target abnormal object that is currently experiencing an anomaly in the target data platform; A second prompt word for the exception handling task is generated based on the exception handling guidance scheme for the target exception object. The exception handling guidance scheme is used to indicate the processing logic for the identified target exception object. The exception handling task is executed based on the second prompt word to handle the exception of the target exception object.

2. The method according to claim 1, characterized in that, The step of obtaining abnormal object data of the target data platform based on the first prompt word includes: Based on the first prompt word, the task abnormality data and / or cluster abnormality data within a specified period are obtained through the first interface corresponding to the first agent of the target large model.

3. The method according to claim 2, characterized in that, The step of performing the anomaly analysis task based on the first prompt word and the anomaly object data to determine the target anomaly object currently experiencing an anomaly on the target data platform includes: Based on the first prompt word, the task anomaly data, and / or the cluster anomaly data, the anomaly analysis task is executed to determine the anomaly task or anomaly platform currently experiencing anomalies in the target data platform, which is then identified as the target anomaly object.

4. The method according to claim 3, characterized in that, The step of performing the anomaly analysis task based on the first prompt word, the task anomaly data, and / or the cluster anomaly data to determine the anomaly task or anomaly platform currently experiencing anomalies on the target data platform, as the target anomaly object, includes: Based on the first prompt word, the task anomaly data, and / or the cluster anomaly data, the anomaly analysis task is executed to determine whether there is a platform anomaly. If a platform anomaly exists, the abnormal platform corresponding to the platform anomaly is identified as the target anomaly object, and the anomaly type and anomaly level of the abnormal platform are determined based on the first preset rule.

5. The method according to claim 4, characterized in that, The method further includes: If there is no platform anomaly, the anomaly analysis task will continue to be executed based on the first prompt word, the task anomaly data and / or the cluster anomaly data to determine whether there is a task anomaly. If so, the abnormal task corresponding to the abnormal task problem is determined as the target abnormal object, and the abnormal task problem type and abnormal problem level are determined based on the second preset rule.

6. The method according to claim 1, characterized in that, The exception handling guidance scheme includes specifying the exception problem level, a third preset rule, and an exception problem handling scheme; The step of executing the exception handling task based on the second prompt word to handle the exception of the target exception object includes: Based on the aforementioned anomaly level and the third preset rule, determine whether the anomaly handling scheme can be used to process the target anomaly object; If so, the anomaly of the target anomaly object is processed based on the anomaly handling scheme.

7. The method according to claim 1, characterized in that, The exception handling guidance scheme also includes specifying communication channels and the target user account associated with the target exception object, and the method further includes: If not, an exception message is generated based on the target exception object, and the exception message is sent to the target user account through the designated communication channel to notify the user of the target user account that the target exception object has an exception.

8. A data processing apparatus, characterized in that, include: The first generation unit is used to generate a first prompt word for the anomaly analysis task based on the anomaly analysis guidance scheme input for the anomaly analysis task, wherein the anomaly analysis guidance scheme is used to indicate: the analysis logic of anomaly objects in the target data platform based on anomaly object data; The first execution unit is configured to obtain abnormal object data of the target data platform based on the first prompt word, and execute the abnormal analysis task based on the first prompt word and the abnormal object data to determine the target abnormal object that is currently experiencing an anomaly in the target data platform; The second generation unit is used to generate a second prompt word for the exception handling task based on the exception handling guidance scheme of the target exception object. The exception handling guidance scheme is used to indicate the processing logic for the identified target exception object. The second execution unit is used to execute the exception handling task based on the second prompt word in order to handle the exception of the target exception object.

9. An electronic device, characterized in that, The device includes a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to perform the steps of the data processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the data processing method as described in any one of claims 1 to 7.