Event processing method and device, computer equipment and readable storage medium

By introducing an intelligent agent system into the machine learning platform, environmental change events can be automatically identified and processed, solving the inefficiency problem of traditional platforms relying on manual monitoring and achieving efficient and intelligent event response and decision-making.

CN121882307APending Publication Date: 2026-04-17SPEEDBOT ROBOTICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SPEEDBOT ROBOTICS CO LTD
Filing Date
2025-12-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional machine learning platforms rely on manual monitoring and troubleshooting, resulting in inefficiency, delayed response, and difficulty in adapting to dynamically changing training environments.

Method used

By introducing an intelligent agent system, the system can perceive the operating environment of the machine learning platform, automatically identify environmental change events, and determine the event handling process based on preset mapping relationships and platform status, thereby achieving automated decision-making and event response.

Benefits of technology

This improved the platform's efficiency and intelligence in handling tasks, reduced manual intervention, and enabled timely and automated event response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121882307A_ABST
    Figure CN121882307A_ABST
Patent Text Reader

Abstract

The invention relates to an event processing method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: carrying out environment perception on a running environment of a machine learning platform, and determining an environment change event when the environment of the machine learning platform changes; analyzing the environment change event, and determining an event action matched with the environment change event; and determining an event processing flow of the machine learning platform to the environment change event according to the event action and the operation state of the machine learning platform. By adopting the method, manual intervention can be reduced, and the response timeliness of a platform for processing events is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an event processing method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Technology

[0002] The machine learning platform is an integrated machine learning training platform that provides a series of functions such as dataset management, model management, training management, packaging, and deployment. It supports fully automated operations from data preprocessing, model training, performance evaluation to model packaging and deployment, aiming to help developers efficiently build, iterate, and maintain models.

[0003] In the traditional approach, the training process needs to be monitored manually, faults need to be handled manually, and the deployment process needs to be configured one by one. This relies heavily on human intervention, which is not only inefficient and has a delayed response time, but also makes it difficult to adapt to the dynamically changing training environment. Summary of the Invention

[0004] Therefore, it is necessary to provide an event processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can reduce human intervention and improve the timeliness of platform event handling in response to the above-mentioned technical problems.

[0005] Firstly, this application provides an event handling method, including:

[0006] The machine learning platform's operating environment is perceived to identify environmental change events when the environment in which the machine learning platform operates changes.

[0007] Analyze the environmental change events to determine the event actions that match the environmental change events;

[0008] Based on the event action and the operating status of the machine learning platform, determine the event handling process of the machine learning platform for the environmental change event.

[0009] In one embodiment, the step of environmental awareness of the machine learning platform's operating environment and determining environmental change events when the environment in which the machine learning platform operates changes includes:

[0010] The system is designed to be aware of the operating environment of the machine learning platform and monitor user actions triggered by the machine learning platform.

[0011] If the operation matches the preset behavior, it is determined that the operating environment has changed;

[0012] The event corresponding to the preset behavior is determined as an environmental change event when the environment in which the machine learning platform is located changes.

[0013] In one embodiment, the step of environmental awareness of the machine learning platform's operating environment and determining environmental change events when the environment in which the machine learning platform operates changes includes:

[0014] The machine learning platform's operating environment is monitored to detect task-related information of training tasks within the machine learning platform.

[0015] If the task-related information matches the preset information, it is determined that the operating environment has changed;

[0016] The event corresponding to the preset information is identified as the environmental change event when the environment in which the machine learning platform is located changes.

[0017] In one embodiment, the step of environmental awareness of the machine learning platform's operating environment and determining environmental change events when the environment in which the machine learning platform operates changes includes:

[0018] The machine learning platform's operating environment is monitored, and model-related parameters of deployed models in the machine learning platform are periodically obtained.

[0019] If the model-related parameters match the preset parameters, it is determined that the operating environment has changed;

[0020] The events corresponding to the preset parameters are identified as environmental change events that occur when the environment in which the machine learning platform operates changes.

[0021] In one embodiment, analyzing the environmental change event and determining the event action matching the environmental change event includes:

[0022] The environmental change event is matched with the pre-set mapping relationship between preset events and preset event actions to determine the event action that matches the environmental change event.

[0023] In one embodiment, the event action includes a plurality of first actions; determining the event handling process of the machine learning platform for the environmental change event based on the event action and the operating status of the machine learning platform includes:

[0024] The initial processing flow of the machine learning platform for the environmental change event is determined according to the execution order among the plurality of first actions;

[0025] Identify candidate actions that match the operating state of the machine learning platform;

[0026] If a target action that matches the candidate action exists among the plurality of first actions, the event processing flow of the machine learning platform for the environmental change event is obtained by adding the candidate action before the target action in the initial processing flow; the candidate action is adjacent to the target action.

[0027] Secondly, this application provides an event processing apparatus, the apparatus comprising:

[0028] The processing module is used to perform environmental awareness of the machine learning platform's operating environment and determine environmental change events when the environment in which the machine learning platform operates changes.

[0029] The analysis module is used to analyze the environmental change events and determine the event actions that match the environmental change events;

[0030] The determination module is used to determine the event handling process of the machine learning platform for the environmental change event based on the event action and the running status of the machine learning platform.

[0031] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0032] The machine learning platform's operating environment is perceived to identify environmental change events when the environment in which the machine learning platform operates changes.

[0033] Analyze the environmental change events to determine the event actions that match the environmental change events;

[0034] Based on the event action and the operating status of the machine learning platform, determine the event handling process of the machine learning platform for the environmental change event.

[0035] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0036] The machine learning platform's operating environment is perceived to identify environmental change events when the environment in which the machine learning platform operates changes.

[0037] Analyze the environmental change events to determine the event actions that match the environmental change events;

[0038] Based on the event action and the operating status of the machine learning platform, determine the event handling process of the machine learning platform for the environmental change event.

[0039] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0040] The machine learning platform's operating environment is perceived to identify environmental change events when the environment in which the machine learning platform operates changes.

[0041] Analyze the environmental change events to determine the event actions that match the environmental change events;

[0042] Based on the event action and the operating status of the machine learning platform, determine the event handling process of the machine learning platform for the environmental change event.

[0043] The aforementioned event handling methods, devices, computer equipment, computer-readable storage media, and computer program products, by perceiving the operating environment of the machine learning platform, identify environmental change events that occur when the environment in which the machine learning platform operates changes. This achieves a shift from the traditional passive monitoring mode that relies on manual polling and judgment to a proactive intelligent monitoring model. This shift avoids the shortcomings of low monitoring efficiency, response delays, and difficulty in covering complex dynamic scenarios caused by the inherent limitations of human resources, thereby improving the overall efficiency and intelligence level of the platform's task processing. By analyzing environmental change events, matching event actions are determined, thus encoding event diagnosis and handling decisions that rely on engineer experience into automated rules. This avoids the inefficient process of manually analyzing and searching for event solutions one by one, achieving instant transformation from problem identification to solution determination, and ensuring the timeliness of generating handling decisions. By analyzing event actions and the operational status of the machine learning platform, the event handling process of the machine learning platform in response to environmental changes can be determined. In this way, by comprehensively analyzing event actions and the operational status of the machine learning platform, the event handling process adapted to the platform can be dynamically determined, which can reduce human intervention, improve the platform's response time to events, and thus improve the overall efficiency and intelligence level of the platform in handling tasks. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart illustrating an event handling method in one embodiment;

[0046] Figure 2This is a flowchart illustrating the process of environmental awareness of the machine learning platform's operating environment and determining environmental change events when the environment in which the machine learning platform operates changes, as shown in one embodiment.

[0047] Figure 3 This is a flowchart illustrating the process of environmental awareness of the machine learning platform's operating environment and determining environmental change events when the environment in which the machine learning platform operates changes, as described in another embodiment.

[0048] Figure 4 This is a flowchart illustrating the process of environmental awareness of the machine learning platform's operating environment and determining environmental change events when the environment in which the machine learning platform operates changes, as described in another embodiment.

[0049] Figure 5 This is a flowchart illustrating the event handling process of a machine learning platform in response to environmental change events, based on event actions and the platform's operating status, in one embodiment.

[0050] Figure 6 This is a flowchart illustrating the event handling method in another embodiment;

[0051] Figure 7 This is a structural block diagram of an event handling device in one embodiment;

[0052] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0054] It should be noted that the terms "comprising" and "having," and any variations thereof, as used in this application, are intended to cover non-exclusive inclusion. The term "multiple" as used in this application refers to two or more. The term "and / or" as used in this application refers to one of the solutions, or any combination of multiple solutions.

[0055] In recent years, large language models (LLMs) have made groundbreaking progress. Models such as the GPT series and QWEN, trained with massive amounts of parameters, have demonstrated powerful capabilities in natural language tasks such as text generation, question-answering, and knowledge reasoning, becoming a core driving force for artificial intelligence technology. However, LLMs have inherent limitations: their context window is limited by the length of the dialogue, making it difficult to handle extremely long texts or continuous data streams; they lack real-time environmental awareness and dynamic decision-making capabilities, and cannot autonomously invoke external tools or execute complex plans; single models are inefficient in multi-task parallel scenarios and have insufficient fault tolerance.

[0056] To overcome these shortcomings, agent technology emerged. By integrating core capabilities such as memory storage, tool invocation, and task planning, agents endow artificial intelligence (AI) systems with the ability to autonomously perceive their environment and dynamically adjust their strategies, solving the key problem that LLMs, as black-box models, cannot interact with the external world. Large language models, such as GPT-4 and LLaMA, are large-scale AI models with natural language understanding and generation capabilities. They are the core engines driving agents in task planning and decision-making.

[0057] In this context, an intelligent agent can refer to an autonomous software entity capable of perceiving its environment, planning, invoking tools (or its own capabilities), and executing actions to achieve a goal. An intelligent agent typically encapsulates complex logic for solving a specific type of problem and possesses the ability to autonomously understand, perceive, plan, remember, and use tools. In an exemplary embodiment, the intelligent agent can employ a "thinking + action" reasoning pattern when determining the task for the query content. Thinking refers to the intelligent agent analyzing the current situation and considering what to do next; action refers to the actions performed by the intelligent agent, typically by invoking tools.

[0058] Specifically, an Agent can encapsulate complex logic and capabilities for solving specific types of problems. Optionally, an Agent can be implemented using a single process or through the collaborative efforts of multiple processes. For example, an Agent can correspond to two processes: one responsible for decision-making and another for hardware control. The deployment method of Agents on physical machines is not unique. Optionally, the server can adopt a single-machine multi-Agent mode, a multi-machine multi-Agent mode, or a hybrid deployment mode. For example, in a single-machine multi-Agent mode, the code and logic of multiple Agents can be deployed on a single physical machine. These Agents are logically independent, possessing their own data, goals, and behaviors, and interacting through in-process message passing. The multi-machine multi-Agent mode is suitable for distributed systems. In this mode, each Agent or multiple Agents run on different physical machines or virtual machines, communicating and collaborating via a network. In a hybrid deployment mode, some Agents can run on the same machine, while others run on other machines, collectively forming a heterogeneous, distributed multi-Agent system.

[0059] Multi-Agent systems further utilize a distributed collaborative architecture to break down complex tasks into sub-goals and assign them to specialized agents for parallel processing, achieving a highly efficient collaborative mode of one master and multiple assistants, or multiple masters and multiple assistants. This architecture not only overcomes the performance bottleneck of a single model through task distribution but also enhances system robustness through communication and coordination mechanisms between agents. Even if some nodes fail, the overall system can still maintain operation, providing a novel solution for complex scenarios. In the Agent system, the environment refers to the external environment that the Agent can perceive. When an Agent changes a state of the environment, and the environment reacts accordingly, the Agent can obtain the changed state and determine its next action. In a game Agent, the environment is the current state of the game; in a system software Agent, the system or software is the state. The Agent is the "brain" placed within the environment, recognizing environmental changes, reacting accordingly to influence those changes, and then making the next decision based on those changes, repeating this process to complete the assigned task. Because the longer the context window of an LLM (Long Message Model) is used in practice, the worse its performance becomes. Since agent systems typically require multi-turn processing and need to save historical context, it becomes difficult for an agent system to complete a long-running task without processing historical dialogues. Therefore, the memory module in an agent system can be used for long-term memory compression and short-term memory construction. Long-term memory focuses on the global dialogue, extracting and compressing key information; short-term memory is used to accurately identify user questions in the current state using recent memory to make decisions.

[0060] Furthermore, based on the number and functions of agents in an agent system, systems can be categorized into several types. For example, a system with only one agent is a single-agent system, where one agent is responsible for all tasks. This type of system has relatively simple functionality. A system with one master agent and multiple worker agents is a master-slave agent system. The master agent is responsible for assigning tasks to worker agents and collecting their results; worker agents handle individual tasks. This type of system is highly scalable, with various capabilities being hot-swappable. This master-slave agent system is the mainstream, and its core is the interaction between agents. A system with multiple master agents is a distributed agent system. In this system, each agent has an equal status, and communication mechanisms are designed to allow interaction between agents. This type of system is more complex in design.

[0061] As shown above, there are various types of agent systems currently available. However, the basic paradigm of existing agent systems is that the user acts as the operational unit, and the agent acts as the execution unit; that is, the user issues commands, and the agent passively executes the commands. In some scenarios, this paradigm is not optimal. In other words, traditional agents require the user to issue specific requests, the agent plans based on those requests, and then the process loops as follows: execution, sensing environmental changes, formulating the next action based on those changes, until the user request is fulfilled or a response indicating task failure is output.

[0062] Therefore, when deploying agents on a machine learning platform, a system can be proposed where the agent is the primary operator and the user is secondary, defining the background and objectives for the agent. When the environment of the machine learning platform changes, the agent automatically recognizes the environmental change and determines subsequent operations without requiring user commands. This allows the agent to become the primary operator of the machine learning platform. In other words, the agent can automatically complete all subsequent operations after changes in the machine learning platform's operating environment, rather than waiting for user requests.

[0063] Specifically, machine learning training platforms typically employ a complete workflow. Users first upload data, select a model, trigger training, and then package and run the trained model on the target server. However, false positives and false negatives are fed back in, requiring users to retrain new models based on this erroneous data before deployment. This cyclical process continues multiple times until the desired results are achieved. This entire process requires manual user intervention and evaluation of the model's performance. However, with the integration of an Agent into the machine learning platform, this entire process can be managed by the Agent. Users simply upload data from the site, and the Agent automatically detects and executes everything else. Agent-managed platforms are highly fault-tolerant; regardless of where an error occurs, the Agent can automatically detect it and determine how to recover, or, if recovery is impossible, seek assistance from the user. The Agent also treats users as assistants, providing customized processes. Therefore, the Agent is not isolated from the user but rather sees the user as an assistant while remaining the primary agent in the decision-making process.

[0064] In view of this, such as Figure 1 As shown, this application provides an event handling method. Taking the application of this method to a first intelligent agent as an example, the first intelligent agent is integrated with a machine learning platform to perceive and autonomously make decisions about the platform's operating environment, including the following steps:

[0065] S102, perform environmental awareness on the operating environment of the machine learning platform, and determine the environmental change events when the environment in which the machine learning platform is located changes.

[0066] The machine learning platform is an integrated machine learning training platform that provides a series of training processes, including dataset management, model management, training management, packaging, and deployment. For example, it can easily train models and deploy them to target machines, and it can iteratively train and maintain the latest models online. The target machine refers to the final computing environment or hardware device that hosts and runs the deployed model. For example, the target machine can be a cloud server, a physical server, a private cloud virtual machine, or an edge server.

[0067] A machine learning platform is configured with a data layer, a computation layer, and a task layer. Therefore, the operating environment of a machine learning platform refers to the set of states of the data layer, computation layer, and task layer. An environment change event is an event generated when the state of any one of the data layer, computation layer, or task layer changes; that is, an environment change event corresponds to a change in the state of any one layer.

[0068] For example, environmental change events include events triggered by user-initiated actions. These events are initiated by the user through the platform, causing changes to the machine learning platform's preset configuration or input data. User-initiated actions include, but are not limited to: adjusting training parameters, uploading incremental datasets (such as adding labeled samples), manually terminating unfinished training tasks, and triggering model deployment commands.

[0069] For example, environmental change events include events triggered by changes in the state of the training process. These events are generated by the platform itself during task execution, including but not limited to: the training task has completed iterations according to preset parameters; hardware resource anomalies occur during model training; and data reading anomalies occur during model training.

[0070] In this case, with the generation of model weight files and evaluation reports, the training task can be considered to have completed iterations according to preset parameters. Hardware resource anomalies may include GPU memory overflow, CPU overload, etc. Data reading anomalies may include invalid dataset paths, incompatible data formats, etc.

[0071] For example, environmental change events include events automatically triggered when the model's state meets set conditions, including but not limited to: periodic inspections detecting that the model is outdated, and model test metrics falling below preset targets. Specifically, if the number of data reflows exceeds a first preset number, but the model's performance metrics are below the preset target, it indicates that the model is outdated, which will trigger an environmental change event. Similarly, if the number of data reflows is less than a second preset number, and the model's corresponding test accuracy is lower than the expected accuracy, it indicates that the model's test metrics are below the expected target, which will also trigger an environmental change event.

[0072] S104, Analyze environmental change events and determine event actions that match the environmental change events.

[0073] Among them, event actions refer to predefined standardized workflows for responding to environmental change events.

[0074] In one embodiment, analyzing environmental change events and determining event actions that match the environmental change events includes: matching the environmental change events with a pre-set mapping relationship between preset events and preset event actions to determine event actions that match the environmental change events.

[0075] For example, if the environment change event is a dataset upload completion event, the corresponding event action could be an action performed after the data is ready, such as initiating dataset validation and submitting the training task. If the environment change event is an insufficient GPU for training, the corresponding event action could be reallocating GPUs.

[0076] In an optional embodiment, the method further includes: if the environmental change event is preprocessed, determining an event action matching the environmental change event based on the preprocessed environmental change event. The preprocessing may include completing missing parameters, etc.

[0077] In an optional embodiment, the method further includes: when there are at least two event actions, outputting a prompt message to prompt the user to select one event action from multiple event actions; and determining the event action selected by the user as an event action that matches the environmental change event. Thus, by making the user the final decision-maker, the user's decision-making burden can be reduced while ensuring the controllability of key aspects.

[0078] For example, when optimizing the environmental change event representation model, the event action may include adjusting the learning rate and adding a regularization term, so that the user can select the optimal event action.

[0079] By adopting the method of the above embodiments, by introducing a mapping relationship between preset events and preset event actions, abstract environmental change events can be quickly and accurately mapped into specific, executable technical response solutions. This enables automated decision-making in a rule-driven manner, eliminating the need for complex real-time reasoning every time an event occurs, thereby improving the timeliness of event response.

[0080] S106, Based on the event actions and the operating status of the machine learning platform, determine the event handling process of the machine learning platform for environmental change events.

[0081] The operational status of a machine learning platform refers to a quantitative set of real-time resource availability, task execution progress, and service health statuses across its data, computation, and task layers. This set includes, but is not limited to: computational resource status, storage resource status, task queue status, data source status, and service health status. For example, computational resource status represents the number of GPUs or CPUs, computing power type, video memory or system memory utilization, and load prediction. Storage resource status represents the available capacity of dataset storage and model repositories. Task queue status represents the priority, resource consumption, estimated execution time, and dependencies of all running and pending tasks. Data source status represents dataset accessibility and validation results. Service health status represents the heartbeat status of tools such as the training engine, deployment services, and scheduler.

[0082] An event handling process refers to an automated work sequence that is dynamically planned or invoked in response to environmental changes, based on matching event actions and the real-time operating status of the platform, and has a clear execution order, operable steps, and resource constraints.

[0083] By employing the method described in the above embodiments, the operating environment of the machine learning platform is perceived, and environmental change events are identified when the environment in which the machine learning platform operates change. This achieves a transformation from the traditional passive monitoring mode that relies on manual polling and judgment to a proactive intelligent monitoring model. This transformation avoids the shortcomings of low monitoring efficiency, response delays, and difficulty in covering complex dynamic scenarios caused by the inherent limitations of human power, thereby improving the overall efficiency and intelligence level of the platform in handling tasks. By analyzing environmental change events, event actions matching the environmental change events are determined. Thus, event diagnosis and handling decisions that rely on engineer experience are encoded into automated rules, avoiding the inefficient process of manually analyzing and searching for event solutions one by one. This achieves an instant transformation from identifying problems to determining solutions, ensuring the timeliness of generating handling decisions. By determining the event handling process of the machine learning platform in response to environmental change events based on event actions and the operating status of the machine learning platform, a comprehensive analysis of event actions and the operating status of the machine learning platform can dynamically determine the event handling process adapted to the platform. This reduces human intervention, improves the platform's response time for handling events, and further improves the overall efficiency and intelligence level of the platform in handling tasks.

[0084] In one embodiment, such as Figure 2 As shown, a flowchart illustrates a process for environmental awareness of a machine learning platform's operating environment and for determining environmental change events when the platform's environment changes. Taking the application of this method to a first intelligent agent as an example, the process includes the following steps:

[0085] S202 performs environmental awareness of the machine learning platform's operating environment and monitors user actions triggered by the machine learning platform.

[0086] Among them, operational behavior refers to the actions performed by users on the machine learning platform through the platform's web-based graphical interface, console, command line, or application programming interface (API) calls.

[0087] S204, if the operating behavior matches the preset behavior, it is determined that the operating environment has changed.

[0088] Predefined behaviors refer to pre-defined user actions that can affect the state of the data layer, computation layer, or task layer of a machine learning platform. Examples of such user behaviors include adjusting training parameters, uploading datasets, manually terminating unfinished training tasks, and triggering model deployment commands.

[0089] For example, when a user adjusts training parameters, a complex related response process needs to be executed (such as initiating hyperparameter optimization and model retraining tasks), which changes the state of the platform's task layer and data layer. Therefore, a preset behavior for user adjustment of training parameters can be set.

[0090] S206, the event corresponding to the preset behavior is determined as the environmental change event when the environment in which the machine learning platform is located changes.

[0091] In this context, the event corresponding to a preset behavior refers to a specific environmental event triggered by a user action that causes a change in the state of any layer within the machine learning platform, whether it's the data layer, computation layer, or task layer. For example, if the preset behavior is a user adjusting training parameters, the corresponding event is the training parameter update event. If the preset behavior is a user uploading incremental data, the corresponding event is the dataset incremental update event.

[0092] For example, if a user enters the reason for training failure through the dialogue entry on the platform's web interface, the first intelligent agent will perform intent recognition on the user's input, match the user's input with multiple preset contents, and identify that the user's intent is to request root cause analysis of a failed training task. The corresponding environmental change event is the training failure diagnosis event.

[0093] By employing the method described in the above embodiments, user behavior can be automatically transformed into structured environmental change events through real-time matching of user behavior and preset behavior. This achieves the transformation from user intent to platform response, enabling the platform to intelligently process corresponding events in an event-driven manner, thereby improving the platform's intelligence in environmental perception, decision-making, and response.

[0094] In one embodiment, such as Figure 3 As shown, a flowchart illustrates a process for environmental awareness of a machine learning platform's operating environment and for determining environmental change events when the platform's environment changes. Taking the application of this method to a first intelligent agent as an example, the process includes the following steps:

[0095] S302 performs environmental awareness of the machine learning platform's operating environment and monitors task-related information of training tasks within the machine learning platform.

[0096] Task-related information refers to information related to the training task when the machine learning platform processes the training task. For example, task-related information may include training loss and validation loss, CPU utilization, CPU load, GPU memory usage, memory usage, progress, and data reading status. For instance, progress may indicate that the training task has completed iterations according to preset parameters (i.e., generating model weight files and evaluation reports), the training task is in progress, or the training task is paused. Data reading status may include data reading errors or normal data reading. Data reading errors are used to indicate dataset path failures, data format incompatibility, etc.

[0097] S304, if the task-related information matches the preset information, it is determined that the operating environment has changed.

[0098] Among them, preset information refers to predefined content used to indicate whether the acquired task-related information will affect the state of any layer in the machine learning platform, whether it is the data layer, the computation layer, or the task layer. In other words, the core function of preset information is to analyze the task-related information collected in real time and determine whether the situation represented by the task-related information will affect the state of any layer in the data layer, the computation layer, or the task layer.

[0099] For example, the preset information includes, but is not limited to: GPU memory usage is greater than the usage threshold, CPU load is greater than the load threshold; progress is represented by generating model weight files and evaluation reports; and data reading status is data reading error.

[0100] S306, the event corresponding to the preset information is identified as an environmental change event when the environment in which the machine learning platform is located changes.

[0101] The events corresponding to the preset information are related to the preset information itself. For example, if the preset information is "data read status is abnormal," the corresponding event is the "data read abnormal event." If the preset information is "GPU memory usage exceeds the usage threshold," the corresponding event is the "GPU memory overload warning event."

[0102] It's easy to understand that when the event corresponding to the preset information refers to a data read exception, in the data layer, the availability status of the specified data version or path on which the current task depends immediately becomes unavailable. In the task layer, the task's execution status changes from running to failed. In the computation layer, because the task is about to fail, the computing resources (CPU or GPU) allocated to it will enter an idle or await-reclaim state.

[0103] By using the method described in the above embodiments, task-related information can be automatically transformed into structured environmental change events through real-time matching of task-related information and preset information. This realizes the transformation from manual passive monitoring of task progress to proactive platform operation and maintenance, enabling the platform to intelligently handle corresponding events in an event-driven manner, thereby improving the platform's intelligence in environmental perception, decision-making, and response.

[0104] In one embodiment, such as Figure 4 As shown, a flowchart illustrates a process for environmental awareness of a machine learning platform's operating environment and for determining environmental change events when the platform's environment changes. Taking the application of this method to a first intelligent agent as an example, the process includes the following steps:

[0105] S402 performs environment awareness on the machine learning platform's operating environment and periodically obtains model-related parameters of the deployed models in the machine learning platform.

[0106] Here, "deployed model" refers to a machine learning model that has been trained, validated, and deployed to the target machine to provide real-time inference services. Model-related parameters refer to parameters associated with the deployed model, including the number of times the model receives feedback data and its performance metrics. Performance metrics include testing metrics (such as test accuracy and recall) and service performance metrics (such as latency and error rate).

[0107] S404 indicates that the operating environment has changed if the model-related parameters match the preset parameters.

[0108] In an optional embodiment, the model-related parameters include the number of times data is re-flowed and the model's performance metrics. The preset parameters include a first preset number of times and a preset target. If the number of times data is re-flowed is greater than the first preset number of times, and the model's performance metrics are less than the preset target, it indicates that the model is outdated, and the model-related parameters are determined to match the preset parameters.

[0109] In an optional embodiment, the model-related parameters include the number of times the data was fed back and the corresponding test accuracy of the model. The preset parameters include a second preset parameter and the expected accuracy. If the number of times the data was fed back is less than the second preset number, and the test accuracy of the model is less than the expected accuracy, then the model-related parameters are determined to match the preset parameters.

[0110] S406, determine the event corresponding to the preset parameters as the environmental change event corresponding to the change in the environment in which the machine learning platform is located.

[0111] By using the method described in the above embodiments, the model's state deviation can be automatically converted into an executable and semantically clear environmental change event by periodically matching the model's relevant parameters with preset parameters. This allows the platform to automatically execute the corresponding processing flow based on the environmental change event without user intervention, thereby improving the platform's intelligence in perceiving, making decisions, and responding to the environment.

[0112] In one embodiment, the machine learning platform's data layer is configured with a data management module, its computation layer with a resource monitoring module, and its task layer with a training scheduling module. The method further includes: when the data management module, resource monitoring module, and training scheduling module act as event producers, the event producers perform environmental awareness of the machine learning platform's operating environment to determine environmental change events when the environment in which the machine learning platform operates changes.

[0113] In an optional embodiment, environmental awareness is performed on the operating environment of the machine learning platform, and task-related information of training tasks in the machine learning platform is monitored, including: environmental awareness of the operating environment of the machine learning platform is performed through event producers, and user-triggered operation behaviors on the machine learning platform are monitored.

[0114] In an optional embodiment, environmental awareness is performed on the operating environment of the machine learning platform, and task-related information of training tasks in the machine learning platform is monitored, including: environmental awareness of the operating environment of the machine learning platform is performed through event producers, and task-related information of training tasks in the machine learning platform is monitored.

[0115] In an optional embodiment, the machine learning platform's operating environment is made aware of, and task-related information of training tasks in the machine learning platform is monitored, including: making the machine learning platform's operating environment aware of through event producers, and periodically obtaining model-related parameters of deployed models in the machine learning platform.

[0116] In an optional embodiment, the method further includes: obtaining standardized information written by the event producer from a message queue. The standardized information is encapsulated information about the generated environment change event when the event producer determines that the operating environment has changed. The standardized information can be used to characterize operational behavior, task-related information, and model-related parameters, etc. The method includes: parsing the standardized information to obtain any one of the following: information characterizing operational behavior, task-related information, and model-related parameters. Then, the first agent determines that the operating environment has changed if the operational behavior matches a preset behavior; or, if the task-related information matches preset information; or, if the model-related parameters match preset parameters.

[0117] In an optional embodiment, the method further includes: a first intelligent agent periodically inspecting the message queue to avoid perception omissions caused by message loss.

[0118] In one embodiment, the event action includes multiple first actions. For example... Figure 5 As shown, a flowchart illustrating the event handling process of a machine learning platform in response to environmental change events is provided, based on event actions and the platform's operational status. Taking the application of this method to a first intelligent agent as an example, the flowchart includes the following steps:

[0119] S502, determine the initial processing flow of the machine learning platform for environmental change events according to the execution order among multiple first actions.

[0120] Specifically, the initial processing flow of the machine learning platform for environmental change events is obtained by combining the first actions according to their execution order. Optionally, each first action also includes at least one second action. For example, taking a data-ready event as an example, the initial processing flow is as follows: data verification, training resource application, and training task submission. Data verification may include second actions such as sample quantity verification, annotation format verification, missing value detection, and outlier filtering.

[0121] S504, identify candidate actions that match the running state of the machine learning platform.

[0122] Specifically, based on the mapping relationship between the platform's operating status and preset actions, candidate actions that match the operating status of the machine learning platform are determined.

[0123] For example, if the running status indicates that the remaining GPU resources of the platform are less than a threshold, the corresponding candidate action is to pause low-priority tasks. In this way, by pausing low-priority tasks, the GPU resources occupied by low-priority tasks can be reclaimed, thereby increasing the total amount of available GPU resources on the platform.

[0124] S506, if there is a target action that matches the candidate action among multiple first actions, the event processing flow of the machine learning platform for environmental change events is obtained by adding the candidate action before the target action in the initial processing flow; the candidate action is adjacent to the target action.

[0125] For example, taking the data-ready event as an example, the initial processing flow is as follows: data verification, training resource request, and training task submission. Due to insufficient remaining GPU resources on the platform (i.e., the remaining GPU resources are less than the threshold), a candidate action of "pausing low-priority tasks" is added before the target action "training resource request". The final event processing flow is as follows: data verification, pausing low-priority tasks, training resource request, and training task submission.

[0126] By using the method described in the above embodiments, the initial processing flow can be dynamically adjusted based on the platform's operating state, thereby generating a customized flow that is most suitable for the current state. This helps to improve the platform's model training success rate in complex training scenarios.

[0127] In an optional embodiment, the method further includes: for each action in the event handling process, invoking the corresponding second agent to execute the relevant processing logic. Thus, by introducing multiple second agents, the processing pressure on the first agent can be reduced, which helps improve the platform's concurrent processing capabilities.

[0128] In an optional embodiment, the method further includes: during the process of handling environmental change events according to the event handling process, obtaining the action execution result corresponding to the current action in the event handling process; and adjusting the actions following the current action in the event handling process according to the action execution result.

[0129] For example, if the execution result of the current action is a data verification failure, such as 10 missing sample labels, then the action following the current action in the event handling process will be adjusted to notify the user of the data anomaly and wait for the user to confirm before re-verifying. This can avoid the waste of resources caused by mechanically executing the preset process, enabling the platform to have the ability to be context-aware and dynamically correct errors, thereby improving the intelligence of the platform in handling tasks.

[0130] In an optional embodiment, the method further includes: during the process of handling environmental change events according to the event handling process, obtaining user configuration information that matches the current action in the event handling process; and generating a processing strategy for the current action based on the user configuration information when the action execution result corresponding to the current action is a failure.

[0131] For example, if the user configuration information representation training fails and is retried once, when the action execution result representation training for the current action fails, a processing strategy for the current action is generated. The processing strategy can be to modify the training parameters, and then retrain according to the modified training parameters. No additional human decision-making is required, which can improve the autonomy and intelligence of the platform's training tasks.

[0132] In an optional embodiment, the method further includes: in response to the platform receiving an access request from an external tool, obtaining a tool capability description file; and registering the tool capability description file in the platform's tool registry. Thus, the first agent parses the tool capability description file registered in the tool registry by calling the Model Context Protocol (MCP) parser, and automatically generates a tool invocation template based on the parsed input parameter format, output result structure, and invocation method. This allows the first agent to invoke the corresponding tool to complete the corresponding task through the tool invocation template without needing to develop dedicated adaptation code for the tool.

[0133] The tool capability description file defines the tool's capability metadata specifications, including the tool name, function description, input parameter format (JSON Schema), output result structure, error code list (e.g., 1001 for parameter error, 1002 for execution timeout), and invocation method. Invocation methods include the Representational State Transfer Application Programming Interface (REST API) or the Command-Line Interface (CLI).

[0134] For example, the tool capability description file provides standardized functional semantic descriptions, input / output patterns, and invocation method declarations for various tools integrated into the platform (such as training, storing, and deploying models). This file decouples the tool's capability declaration from its internal processing logic, enabling the first intelligent agent to automatically discover and integrate tools without code by parsing the file.

[0135] In one embodiment, the method further includes: invoking a target tool that matches the environmental change event, and processing the environmental change event according to the event handling process; obtaining the tool execution status of the target tool; and, if the tool execution status is abnormal, triggering a corresponding preset fault tolerance strategy according to the type of abnormal status.

[0136] Optionally, obtaining the execution status of the target tool includes: obtaining the execution status of the target tool through a status query interface provided by the target tool; or receiving the execution status feedback from the target tool. The status query interface may include a task interface or a status interface, etc.

[0137] The preset fault tolerance strategy includes at least the following: if the abnormal state is characterized as a recoverable training failure error, then return to the target tool that matches the environmental change event and handle the environmental change event according to the event handling process; if the abnormal state is characterized as a resource shortage error, then trigger the resource re-application action.

[0138] In one embodiment, the method further includes: determining the execution result of the target task to which the environmental change event belongs; and storing the execution result in a database. The execution result may include the path to the trained model file, evaluation metrics, etc.

[0139] In one embodiment, the method further includes: analyzing the execution results to determine the environmental state impact of the execution results on the current operating environment of the machine learning platform; and, if the environmental state impact is taken as a new environmental change signal, initiating a new round of environmental perception process based on the environmental change signal. Thus, a closed loop of perception, thinking, execution, and re-perception can be formed.

[0140] For example, environmental conditions can affect the platform's model repository status by changing it from "pending update" to "ready" after model files are generated. The model repository refers to the core storage and metadata center that centralizes, versiones, and standardizes the management of all model assets on the platform. For instance, if the execution result indicates that the model's evaluation metrics are not up to standard, i.e., accuracy (ACC) < 90%, a new model optimization event will be triggered. The first agent will re-enter the thought process, planning a chain of actions for parameter tuning and retraining.

[0141] In summary, this application provides a full-chain design for environmental change perception, event understanding, event action decomposition, and dynamic adjustment, which can realize the dynamic adaptation and automated management of the agent to the machine learning platform, significantly reduce the cost of manual intervention, and improve the success rate of training tasks and the utilization rate of platform resources.

[0142] On the one hand, the core limitation of traditional agents is their strong reliance on user-initiated triggers. Users must manually input commands (such as checking training task status), click actions (such as starting training), or call APIs to initiate the processing flow. If the user is unaware of system anomalies (such as background GPU memory overflow during training) or forgets operational steps (such as failure to trigger verification after dataset upload), the agent remains dormant and unable to intervene. In contrast, the agent provided in this application uses environmental changes as its core trigger source and can autonomously start without any user requests. Through message queue listening and periodic status checks, once it detects dynamic changes in platform data, tasks, and resources (such as dataset upload completion, GPU load exceeding limits, or insufficient model accuracy), it automatically initiates a process of perception, consideration, and execution. For example, when a training task on the platform is detected to be interrupted due to a data format error, the agent provided in this application will immediately trigger the entire process of error location, format repair, and task restart without user notification or feedback, eliminating dependence on user intervention.

[0143] On the other hand, traditional agents, relying on user triggers, suffer from significant response delays. This involves multiple stages: user discovery of the problem, understanding of the problem, initiation of an agent request, and agent execution. This delay can lead to severe losses, especially in machine learning scenarios. For example, if a training task is interrupted due to GPU resource leakage, the user may only notice and trigger the agent two hours later, during which time GPU resources are continuously wasted, and intermediate training results may be lost. The agent provided in this application, however, possesses autonomous platform awareness, enabling a closed-loop response. Upon the generation of an environmental change event, the agent can directly complete perception and preprocessing, and initiate the execution process without waiting for user intervention. Taking incremental dataset uploads as an example, traditional agents require the user to upload the data and then request the agent to start training. In contrast, the agent provided in this application, after the dataset is uploaded to the platform, perceives the change and automatically executes data verification, format conversion, resource request, and training submission. This fully automates the process from data readiness to training initiation, significantly improving user efficiency when using the platform.

[0144] On the other hand, traditional agents require users to have a certain level of professional knowledge and operational understanding of the platform's functions. Users need to clearly understand how to complete a full training and deployment process. For example, novice users may not know how to deploy after training. The agent provided in this application, through autonomous environment awareness and automated processing, minimizes user operational costs, achieving a minimalist platform experience with zero professional barriers. Users only need to complete core operations (such as uploading datasets), and all subsequent processes (validation, resource coordination, anomaly handling, and result optimization) are autonomously completed by the agent using the corresponding tools, without the need to learn commands or understand technical details. For example, after a non-professional user uploads a dataset, they don't need to understand data format requirements or GPU resource application procedures. The agent automatically completes data format validation, determines the required GPU model and quantity, applies for resources, and starts training. If data is missing during the process, it will automatically push completion prompts, eliminating the need for manual troubleshooting. This significantly lowers the barrier to entry for machine learning platforms, allowing even non-professional users to efficiently and easily use the platform's functions.

[0145] In summary, such as Figure 6 The diagram illustrates an event handling method, using the application of this method to a first intelligent agent as an example. The method includes the following steps:

[0146] S602 performs environmental awareness of the machine learning platform's operating environment and identifies environmental change events when the environment in which the machine learning platform operates changes.

[0147] S604, Match the environmental change event with the pre-set mapping relationship between preset events and preset event actions to determine the event action that matches the environmental change event; the event action includes multiple first actions.

[0148] S606, determine the initial processing flow of the machine learning platform for environmental change events according to the execution order among multiple first actions.

[0149] S608, identify candidate actions that match the running state of the machine learning platform.

[0150] S610, if there is a target action that matches the candidate action among multiple first actions, the candidate action is added before the target action in the initial processing flow to obtain the event processing flow of the machine learning platform for environmental change events; the candidate action is adjacent to the target action.

[0151] The contents of S602 to S610 can be adapted to the description above.

[0152] In summary, this application proposes a new agent-led, user-collaborative paradigm, completely breaking away from the passive tool nature of traditional agents and constructing a new collaborative system with the agent as the system's central hub and the user as the key enabler. The core positioning of traditional agents is that of user assistants, and their working logic relies entirely on human commands, resulting in three fundamental limitations:

[0153] Trigger-dependent behavior means that tasks can only be initiated by human intervention (such as inputting commands, clicking, or calling APIs), and the agent cannot autonomously perceive system dynamics. For example, in machine learning platforms, traditional agents only intervene when the user manually sends a command to check for training anomalies. If the user does not notice the training interruption, the agent remains dormant and cannot intervene proactively.

[0154] Functional fragmentation means that it can only complete partial tasks for a single user request and lacks system-level management capabilities. For example, after a user triggers a data cleaning command, a traditional agent can only complete the cleaning operation and cannot independently connect the entire chain of resource application, training submission, and result optimization. The process can only be advanced by the user issuing multiple commands.

[0155] Decision-making gaps mean that traditional agents lack independent decision-making capabilities, and all key judgments rely on explicit human instructions. For example, when encountering insufficient training GPU resources, traditional agents must wait for user instructions (such as pausing low-priority tasks or postponing the current training task), and cannot make decisions autonomously based on the platform's resource status and task priorities, leading to process interruptions.

[0156] Therefore, traditional agents adopt a passive paradigm of human command and agent execution, which keeps agents as tools and prevents them from taking on the role of system management and proactive collaboration, thus greatly limiting the value of agents in complex scenarios.

[0157] In the method provided in this application, the Agent completely sheds its assistant role and upgrades to become the core manager and proactive executor of the platform, undertaking four core functions to achieve "managing everything, perceiving everything, and working proactively." Specifically, the Agent provided in this application has the following roles:

[0158] A full-dimensional perceiver, through message queue listening, timed status inspections, and multi-module data interaction (such as interfacing with the platform's data layer, task layer, and computation layer), captures real-time dynamics across all dimensions of the system, including environmental changes (dataset updates, hardware anomalies), task status (training progress, result metrics), and resource fluctuations (GPU / CPU usage, storage capacity). It obtains complete system information without human intervention. For example, it automatically detects implicit changes such as a 5% drop in model test set accuracy or the dataset exceeding the effective time window.

[0159] A global task planner is an agent that autonomously constructs the entire task flow based on the perceived system state, rather than relying on human instructions to break it down. For example, in a dataset upload scenario, the agent can autonomously plan the complete process of data verification, format conversion, resource matching, training submission, and anomaly contingency plans without requiring humans to issue instructions step by step. At the same time, it can dynamically adjust the plan according to the real-time status (such as automatically inserting low-priority task pauses when insufficient resources are detected).

[0160] A closed-loop executor is someone who actively invokes system tools and modules to drive the entire task process and achieve a closed loop of "execution-feedback-adjustment". For example, if a memory overflow occurs during training, the agent can autonomously perform closed-loop operations such as terminating the abnormal process, releasing invalid resources, reconfiguring the batch size, and restarting training without human intervention.

[0161] A dynamic decision provider is an agent that autonomously generates decision-making schemes based on preset rules and real-time system status, triggering collaboration only at high-risk or critical nodes requiring human judgment. For example, regarding model optimization, the agent can autonomously analyze training logs and generate two schemes: adjusting the learning rate to 0.001 or adding a regularization term. It also provides the expected effects of each scheme (such as a 3% increase in accuracy or an increase of 1 hour in training time) for the user to make the final decision, reducing the burden of human decision-making while ensuring the controllability of key processes.

[0162] As shown above, this application provides an Agent-centric, user-supported architecture. The user role shifts from command sender to core enabler, focusing on two high-value functions to achieve efficient collaboration with the Agent. The user has the following roles:

[0163] The data and goal enabler means that users only need to provide core inputs (such as labeled datasets and task goal requirements) and basic constraints (such as training cost limits and model accuracy thresholds), without needing to participate in the specific process. For example, a user uploads a medical image dataset and sets "tumor recognition accuracy ≥ 95%", and the entire process of subsequent data verification, model selection, training optimization, etc., is completed autonomously by the agent, without the user needing to intervene in the intermediate steps.

[0164] The key decision-maker, i.e., the user, only intervenes in decision-making in high-risk, highly subjective scenarios where the agent cannot make independent judgments. For example, when the agent generates two options for rerunning a training task that is interrupted (involving rerunning millions of data points) and rerunning it during off-peak hours the next day (saving 50% of resources), the user can make the final choice based on the urgency of the business (such as whether the model needs to be delivered the next day). This avoids decision-making bias caused by the agent's lack of business context and minimizes the operational burden on humans.

[0165] Therefore, the method provided in this application achieves a fundamental transformation of the Agent from a passive tool to the platform's central hub. On the one hand, through the Agent's autonomous perception, planning, execution, and decision-making, the operational efficiency of the system platform is significantly improved, and the cost of human intervention is reduced. On the other hand, by dividing the work between the Agent undertaking basic tasks and humans focusing on core decisions, the efficiency and stability of AI are leveraged while retaining the user's judgment in key scenarios. This constructs an intelligent collaborative system that is more adaptable to complex systems (such as machine learning platforms and industrial control systems), providing paradigm support for the platform's implementation in highly complex inspection scenarios.

[0166] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0167] Based on the same inventive concept, this application also provides an event processing apparatus for implementing the event processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more event processing apparatus embodiments provided below can be found in the limitations of the event processing method described above, and will not be repeated here.

[0168] In one exemplary embodiment, such as Figure 7 As shown, an event processing apparatus is provided, including: a processing module 702, an analysis module 704, and a determination module 706, wherein:

[0169] The processing module 702 is used to perform environmental awareness of the operating environment of the machine learning platform and determine the environmental change event when the environment in which the machine learning platform is located changes; the analysis module 704 is used to analyze the environmental change event and determine the event action that matches the environmental change event; the determination module 706 is used to determine the event processing flow of the machine learning platform for the environmental change event based on the event action and the operating status of the machine learning platform.

[0170] In one embodiment, the processing module 702 is further configured to: perform environmental awareness on the operating environment of the machine learning platform and monitor user-triggered operations on the machine learning platform; determine that the operating environment has changed if the operation matches a preset behavior; and determine the event corresponding to the preset behavior as an environmental change event when the environment of the machine learning platform changes.

[0171] In one embodiment, the processing module 702 is further configured to: perform environmental awareness of the operating environment of the machine learning platform, and monitor task-related information of training tasks in the machine learning platform; determine that the operating environment has changed if the task-related information matches preset information; and determine the event corresponding to the preset information as the environmental change event corresponding to the change in the environment of the machine learning platform.

[0172] In one embodiment, the processing module 702 is further configured to: perform environmental awareness of the operating environment of the machine learning platform, periodically acquire model-related parameters of the deployed models in the machine learning platform; determine that the operating environment has changed if the model-related parameters match preset parameters; and determine the event corresponding to the preset parameters as the environmental change event corresponding to the change in the environment of the machine learning platform.

[0173] In one embodiment, the analysis module 704 is further configured to: match the environmental change event with a pre-set mapping relationship between preset events and preset event actions, and determine the event action that matches the environmental change event.

[0174] In one embodiment, the event action includes a plurality of first actions; the determining module 706 is further configured to: determine the initial processing flow of the machine learning platform for the environmental change event according to the execution order among the plurality of first actions; determine candidate actions that match the running state of the machine learning platform; if there is a target action that matches the candidate action among the plurality of first actions, obtain the event processing flow of the machine learning platform for the environmental change event by adding the candidate action before the target action in the initial processing flow; the candidate action is adjacent to the target action.

[0175] Each module in the aforementioned event processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can invoke and execute the operations corresponding to each module.

[0176] In one exemplary embodiment, a computer device is provided, which may be a terminal. The terminal has a machine learning platform installed on it, and the internal structure diagram of the terminal may be as follows: Figure 8 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data during event processing. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements an event processing method.

[0177] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0178] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0179] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0180] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0181] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0182] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0183] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0184] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An event processing method, characterized by, The method includes: The machine learning platform's operating environment is perceived to identify environmental change events when the environment in which the machine learning platform operates changes. Analyze the environmental change events to determine the event actions that match the environmental change events; Based on the event action and the operating status of the machine learning platform, determine the event handling process of the machine learning platform for the environmental change event.

2. The method of claim 1, wherein, The process of achieving environmental awareness of the machine learning platform's operating environment and determining environmental change events when the environment in which the machine learning platform operates changes includes: The system is designed to be aware of the operating environment of the machine learning platform and monitor user actions triggered by the machine learning platform. If the operation matches the preset behavior, it is determined that the operating environment has changed; The event corresponding to the preset behavior is determined as an environmental change event when the environment in which the machine learning platform is located changes.

3. The method according to claim 1, characterized in that, The process of achieving environmental awareness of the machine learning platform's operating environment and determining environmental change events when the environment in which the machine learning platform operates changes includes: The machine learning platform's operating environment is monitored to detect task-related information of training tasks within the machine learning platform. If the task-related information matches the preset information, it is determined that the operating environment has changed; The event corresponding to the preset information is identified as the environmental change event when the environment in which the machine learning platform is located changes.

4. The method according to claim 1, characterized in that, The process of achieving environmental awareness of the machine learning platform's operating environment and determining environmental change events when the environment in which the machine learning platform operates changes includes: The machine learning platform's operating environment is monitored, and model-related parameters of deployed models in the machine learning platform are periodically obtained. If the model-related parameters match the preset parameters, it is determined that the operating environment has changed; The events corresponding to the preset parameters are identified as environmental change events that occur when the environment in which the machine learning platform operates changes.

5. The method of claim 1, wherein, The analysis of the environmental change event to determine the event action matching the environmental change event includes: The environmental change event is matched with the pre-set mapping relationship between preset events and preset event actions to determine the event action that matches the environmental change event.

6. The method of claim 1, wherein, The event action includes multiple first actions; determining the event handling process of the machine learning platform for the environmental change event based on the event action and the operating status of the machine learning platform includes: The initial processing flow of the machine learning platform for the environmental change event is determined according to the execution order among the plurality of first actions; Identify candidate actions that match the operating state of the machine learning platform; If a target action that matches the candidate action exists among the plurality of first actions, the event processing flow of the machine learning platform for the environmental change event is obtained by adding the candidate action before the target action in the initial processing flow; the candidate action is adjacent to the target action.

7. An event processing apparatus, characterized by comprising: The device includes: The processing module is used to perform environmental awareness of the machine learning platform's operating environment and determine environmental change events when the environment in which the machine learning platform operates changes. The analysis module is used to analyze the environmental change events and determine the event actions that match the environmental change events; The determination module is used to determine the event handling process of the machine learning platform for the environmental change event based on the event action and the running status of the machine learning platform. 8.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-7. When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.