Artificial intelligence model processing method and device
By generating anomaly tasks on the detection server and combining them with model instance data for verification and updates, the problem of anomaly detection and repair in artificial intelligence models is solved, thereby improving the stability and efficiency of the models.
Patent Information
- Application Number
- CN202510928102.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-11-04
AI Technical Summary
As the complexity of artificial intelligence models increases, the stability and availability of these models become key challenges affecting service continuity, and existing technologies struggle to effectively detect and repair anomalies.
By generating abnormal tasks after detecting anomalies in model instances of artificial intelligence models on the detection server, and combining the instance deployment data and anomaly handling parameters of the model instances, task execution verification and status updates are performed to achieve orderly processing and repair of abnormal tasks.
It enhances the flexibility and proactivity of anomaly detection, ensures the timeliness and stability of model repair, avoids the impact of a large number of abnormal tasks on the system in a short period of time, and improves the stability and efficiency of the system.
Smart Images

Figure CN120892231A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present document relates to the technical field of data processing, and particularly relates to a processing method and device of an artificial intelligence model. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, the application range of artificial intelligence models is wider and wider, covering many fields such as search, recommendation, content generation, intelligent customer service. In these fields, the stability and usability of artificial intelligence models are crucial to the continuity of services. The wide application of artificial intelligence models promotes the intelligent development and progress of various fields, but as the model complexity of artificial intelligence models gradually rises, the use requirements for artificial intelligence models are increasingly improved, which brings certain challenges to the application of artificial intelligence models by service providers in various fields. SUMMARY
[0003] One or more embodiments of the present specification provide a processing method of an artificial intelligence model, comprising: querying an abnormal task in an initialization state in a task set for model anomaly repair. The abnormal task is generated after an anomaly detection server detects model instances of each artificial intelligence model. According to instance deployment data and abnormal processing parameters of the model instance contained in the abnormal task, task execution verification is performed, and task state updating is performed based on the verification result. Based on the abnormal processing parameters contained in the abnormal task of the target task state in the task set, task execution is performed to perform abnormal repair processing of the model instance of the corresponding artificial intelligence model.
[0004] One or more embodiments of the present specification provide a processing device of an artificial intelligence model, comprising: a task query module configured to query an abnormal task in an initialization state in a task set for model anomaly repair. The abnormal task is generated after an anomaly detection server detects model instances of each artificial intelligence model. A verification module is configured to perform task execution verification according to instance deployment data and abnormal processing parameters of the model instance contained in the abnormal task, and to perform task state updating based on the verification result. An execution module is configured to perform task execution based on the abnormal processing parameters contained in the abnormal task of the target task state in the task set, to perform abnormal repair processing of the model instance of the corresponding artificial intelligence model.
[0005] One or more embodiments of the present specification provide a processing device of an artificial intelligence model, comprising: a processor; and a memory configured to store computer executable instructions that, when executed, cause the processor to: query an abnormal task of an initialization state in a task set for model anomaly repair. The abnormal task is generated after an anomaly detection server detects an anomaly of a model instance of each artificial intelligence model. According to instance deployment data and anomaly processing parameters of the model instance contained in the abnormal task, task execution verification is performed, and task state updating is performed based on the verification result. Based on the anomaly processing parameters contained in the abnormal task of the target task state in the task set, task execution is performed to perform anomaly repair processing of the model instance of the corresponding artificial intelligence model.
[0006] One or more embodiments of the present specification provide a computer readable storage medium for storing computer executable instructions that, when executed, implement the following steps: query an abnormal task of an initialization state in a task set for model anomaly repair. The abnormal task is generated after an anomaly detection server detects an anomaly of a model instance of each artificial intelligence model. According to instance deployment data and anomaly processing parameters of the model instance contained in the abnormal task, task execution verification is performed, and task state updating is performed based on the verification result. Based on the anomaly processing parameters contained in the abnormal task of the target task state in the task set, task execution is performed to perform anomaly repair processing of the model instance of the corresponding artificial intelligence model. BRIEF DESCRIPTION OF DRAWINGS
[0007] In order to more clearly illustrate the technical solutions in the one or more embodiments of the present specification or the prior art, brief introductions to the drawings needed in the embodiment or prior art description will be given below. Obviously, the drawings in the following description are only some embodiments described in the present specification, and those skilled in the art can also obtain other drawings according to these drawings without creative labor; Figure 1 A schematic diagram of a processing method implementation environment of an artificial intelligence model provided by one or more embodiments of the present specification; Figure 2 A processing flowchart of a processing method of an artificial intelligence model provided by one or more embodiments of the present specification; Figure 3 A schematic diagram of an anomaly processing process of an artificial intelligence model provided by one or more embodiments of the present specification; Figure 4 A processing flowchart of a processing method of an artificial intelligence model applied to an AI model scenario provided by one or more embodiments of the present specification; Figure 5A schematic diagram of an embodiment of a processing device of an artificial intelligence model provided for one or more embodiments of the present specification; Figure 6 A structural schematic diagram of a processing device of an artificial intelligence model provided for one or more embodiments of the present specification. DETAILED DESCRIPTION
[0008] In order to enable those skilled in the art to better understand the technical solutions in the one or more embodiments of the present specification, the technical solutions in the one or more embodiments of the present specification will be described clearly and completely below in conjunction with the drawings in the one or more embodiments of the present specification. Obviously, the described embodiments are only a part of the embodiments of the present specification, rather than all the embodiments. Based on the one or more embodiments of the present specification, all other embodiments obtained by those skilled in the art without creative work should belong to the protection scope of the present document.
[0009] The processing method of the artificial intelligence model provided by the one or more embodiments of the present specification can be applied to the implementation environment of the artificial intelligence model, and the implementation environment at least includes Figure 1 a server 101; in addition, the implementation environment can also include a detection server 102, model instances 103 of each artificial intelligence model; The server 101 is configured to perform task execution verification on the abnormal tasks in the initialization state, and perform task state updating based on the verification result, perform task execution on the abnormal tasks in the target task state, and perform abnormal repair processing on the model instances of the corresponding artificial intelligence model. The server 101 can be one or more servers, a server cluster composed of several servers, or a cloud server of a cloud computing platform. The detection server 102 is configured to perform abnormal detection on the model instances of each artificial intelligence model and generate abnormal tasks. The detection server 102 can be one or more servers, a server cluster composed of several servers, or a cloud server of a cloud computing platform. The model instances 103 of each artificial intelligence model can correspond to the detection server 102 one by one. The model instances 103 of each artificial intelligence model can be deployed on the detection server 102 or other servers. The detection server 102 can include a detection server 102-1, a detection server 102-2, and a detection server 102-n. The model instances 103 of each artificial intelligence model can include a model instance 103-1, a model instance 103-2, and a model instance 103-n. The artificial intelligence models to which the model instances 103-1, 103-2, and 103-n belong can be the same or different.
[0010] In the implementation environment, the server 101 queries an abnormal task in an initialization state in a task set, performs task execution verification according to instance deployment data and abnormal processing parameters of a model instance contained in the abnormal task, and performs task state updating based on a verification result, and performs task execution on the abnormal task in the target task state in the task set to perform abnormal repair processing on the model instance of the corresponding artificial intelligence model, so as to realize abnormal repair of the model instance of the artificial intelligence model through execution of the abnormal task. The abnormal task in the initialization state is generated after the detection server 102 performs abnormal detection on the model instance 103 of each artificial intelligence model.
[0011] One or more embodiments of the processing method of the artificial intelligence model provided in the specification are as follows: With reference to Figure 2 The processing method of the artificial intelligence model provided in the embodiment can be applied to a local server, and specifically includes steps S202 to S206.
[0012] Step S202: Query an abnormal task in an initialization state in a task set for model abnormal repair.
[0013] The task set in the embodiment refers to a set composed of one or more abnormal tasks, and the task set includes at least one abnormal task in a task state. The task set can be stored in any data structure, such as a task table or a task database. The task state includes an initialization state, a to-be-executed state, an executing state, an execution success state, an execution failure state, and / or a timeout state. Optionally, the abnormal task is generated after the detection server performs abnormal detection on the model instance of each artificial intelligence model.
[0014] The artificial intelligence model refers to an intelligent model with a large number of parameters constructed by an artificial neural network. The artificial intelligence model can be an AI (Artificial Intelligence) model or an AI large model. The artificial intelligence model can include a large language model. The architecture of the large language model can be a neural network architecture with a large number of parameters, a Transform architecture, or other architectures. For example, the large language model is a chatGPT (chat Generative Pre-trained Transformer) or various open-source large language models such as DeepSeek.
[0015] Optionally, each artificial intelligence model is deployed in an application service of an application platform and corresponds to the application service one-to-one, the application platform includes a platform corresponding to an application program, and the application platform can be any platform corresponding to an application program, for example, the application platform is a payment platform, a government affair platform, a ticketing platform, or a tourism platform; each application service deployed by the application platform can be any type of service, for example, each application service includes a medical service, an insurance service, or a resource management service; the model structure of each artificial intelligence model corresponding to each application service can be the same or different.
[0016] Each model instance of each artificial intelligence model can be one or more, each model instance of each artificial intelligence model can be a specific model of different versions trained based on the same model architecture but through different training data sets or hyperparameter settings, or can be a specific model trained based on the same model architecture and through the same training data set or hyperparameter setting.
[0017] In a specific implementation, an abnormal task in an initialization state is queried in the task set for model anomaly repair, specifically, an abnormal task in an initialization state is queried in the task set for model anomaly repair according to a task query period, so as to realize processing of abnormal tasks in sequence; the model anomaly repair includes anomaly repair of an artificial intelligence model, a model instance of the artificial intelligence model, and / or an application service corresponding to the artificial intelligence model.
[0018] As described above, the abnormal task in an initialization state can be generated after the detection server detects the model instance of each artificial intelligence model, so as to improve the flexibility and initiative of anomaly detection by detecting the model instance through the detection server corresponding to the model instance of each artificial intelligence model; in an optional implementation provided in this embodiment, the detection server detects the model instance in the process of anomaly detection, calculates an instance running index according to the instance running data of the model instance to obtain the instance running index, and detects a target anomaly processing strategy in the anomaly processing strategy based on the instance running index, specifically, the detection server can detect the model instance in the following manner: Calculate an instance running index according to the instance running data of each model instance of each artificial intelligence model to obtain the instance running index; If the instance running index triggers an instance anomaly condition in the anomaly processing strategy, detect a target anomaly processing strategy in the anomaly processing strategy based on the instance deployment data.
[0019] The instance running data of the model instance refers to running data generated by the model instance in a running process, and can be, for example, processing data of the model instance of each application service corresponding to an artificial intelligence model in an instance running process for request processing of a service request. The instance running data includes, for example, a number of successful request processing times, cache utilization data, and a number of request processing times. The instance running index refers to a processing index of the model instance in the instance running process for request processing of the service request. The instance running index includes, for example, a success rate, a cache utilization rate, and / or a throughput.
[0020] Specifically, the instance running index can be obtained by performing running index calculation according to the instance running data of the model instance of each artificial intelligence model. In a case where the instance running index triggers an instance exception condition in the exception handling strategy, the instance deployment data is matched with the exception handling strategy for matching processing to obtain a target exception handling strategy matched with the instance deployment data. In a case where the instance running index does not trigger the instance exception condition in the exception handling strategy, no processing is performed.
[0021] In the above, the instance running index is obtained by performing running index calculation according to the instance running data of the model instance of each artificial intelligence model. If the instance running index triggers an instance exception condition in the exception handling strategy, the instance deployment data is used to detect a target exception handling strategy in the exception handling strategy. Optionally, the instance running data is obtained by performing instance running data collection on the model instance based on a data collection strategy. The detection server is deployed in the cloud. The cloud corresponds to the model instance in a one-to-one manner. The data collection strategy and / or the exception handling strategy are input by calling an interface of the instance deployment data in the cloud. The data is obtained after the local server is called to query the interface. The local server includes a host server. The abnormality processing strategy refers to a strategy for processing or repairing an abnormality of a model instance of an artificial intelligence model. The abnormality processing strategy can be one or more. Each abnormality processing strategy can include an instance abnormality condition, a strategy name, instance deployment data, a fault code, a repair object field, a repair action field, an enable state, and / or a running mode. The running model can include a normal mode and / or an alarm mode. For example, the instance abnormality condition is that the success rate is less than 80% within 5 minutes and / or the cache utilization rate is greater than 80% within 10 minutes. The instance deployment data refers to environment data of software and / or hardware deployed by the model instance, i.e., the instance deployment data can be instance environment data. For example, the instance deployment data includes an IDC (Internet Data Center, i.e., a computer room), an artificial intelligence model to which the model instance belongs, and / or a deployment unit. For another example, the deployment unit includes a test deployment unit and / or a disaster recovery deployment unit. The repair object field refers to a field of an object to be repaired. The repair object field can uniquely represent the repair object. For example, the object to be repaired is a computer room, and the repair object field is a computer room field. The repair action field refers to a field of an action to be taken for repairing. The repair action field can include a repair action code. The repair action field can uniquely represent the repair action. For example, the repair action field is an expansion field.
[0022] In a specific implementation process, the detection server corresponding to the model instance of each artificial intelligence model can perform abnormality detection on the model instance. After the abnormality detection, an abnormality task is generated and sent to the local server. The local server can mark the abnormality task as an initialization state and write the abnormality task in the initialization state into a task set. Based on the detection of the target abnormality processing strategy in the abnormality processing strategy based on the instance deployment data, in an optional implementation of the embodiment, the abnormality task in the initialization state is generated according to the instance deployment data and the repair action field, the repair object field, and / or the execution module field. A specific abnormality task can be generated in the following manner: extracting the repair action field, the repair object field, and the execution module field from the target abnormality processing strategy; generating the abnormality task in the initialization state according to the instance deployment data and the repair action field, the repair object field, and the execution module field.
[0023] The execution module field refers to a field of a module for repairing an abnormality. The execution module field can uniquely represent the execution module. The execution module includes an executor. The execution module can be deployed on the local server, the cloud, and / or the operation and maintenance platform. The operation and maintenance platform can be an internal operation and maintenance platform or a third-party operation and maintenance platform.
[0024] Specifically, the repair action field, the repair object field and / or the execution module field can be extracted from the target exception handling strategy, and an exception task in an initialization state is generated according to the instance deployment data and the repair action field, the repair object field and / or the execution module field. Specifically, the exception task can be generated according to the instance deployment data, the repair action field, the repair object field, the execution module field and / or a trigger event. The trigger event includes a trigger source, and specifically can be an object that generates the exception task.
[0025] In order to comprehensively detect the artificial intelligence model, and improve the comprehensiveness of detecting the artificial intelligence model, the cloud can synchronize the instance running data to the local server after obtaining the instance running data based on the data collection strategy on the model instance. The local server can count the instance running data of the model instance of the artificial intelligence model of each application service, calculate the instance running index of the model instance according to the instance running data and / or calculate the model running index of the artificial intelligence model according to the instance running data, generate an exception task after the instance running index and / or the model running index triggers the exception repair condition in the exception handling strategy, mark the exception task as an initialization state and write it into a task set. In addition, the local server can also receive an exception service sent by the operation and maintenance platform or an external system and write it into the task set.
[0026] It should be noted that the above operation of querying the exception task in the initialization state in the task set for model exception repair can be replaced by querying the exception task in the initialization state in the task set, or by querying the exception task in the initialization state in the task set for exception repair of the artificial intelligence model. The above and other processing steps provided in the embodiment form a new implementation manner.
[0027] In step S204, task execution verification is performed according to the instance deployment data and the exception handling parameters of the model instance contained in the exception task, and task state updating is performed based on the verification result.
[0028] In the above operation of querying the exception task in the initialization state in the task set for model exception repair, in this step, the instance deployment data and the exception handling parameters of the model instance contained in the exception task are combined to perform task execution verification on the exception task, and the task state of the exception task is updated based on the verification result. In this way, the task execution verification avoids a large number of executions of the exception task in a short time, and improves the stability of the local server and the cloud.
[0029] The abnormality processing parameter in the embodiment refers to a processing parameter for performing abnormality repair or abnormality processing on a model instance of an artificial intelligence model. The abnormality processing parameter includes a repair action field. In addition, the abnormality processing parameter can further include a repair object field and / or an execution module field. The task execution verification includes task flow limiting verification. Specifically, the task flow limiting verification can be verification of whether to limit the flow of abnormal tasks.
[0030] In a specific implementation, in order to improve the orderliness of execution of abnormal tasks and avoid a large number of abnormal tasks being executed in a short time to cause an impact on the system and improve stability, in an optional implementation provided in the embodiment, the following operation is performed in the process of performing task execution verification according to the instance deployment data of the model instance contained in the abnormal task and the abnormality processing parameter: query the task execution limiting parameter matched by the abnormal task according to the abnormality processing parameter of the abnormal task containing the instance deployment data and the initialization state; perform task execution verification on the abnormal task in the initialization state based on the task execution limiting parameter.
[0031] The task execution limiting parameter refers to a limiting parameter for limiting execution of the abnormal task. The task execution limiting parameter includes a task flow limiting parameter, i.e., a flow limiting parameter for limiting the flow of the abnormal task. For example, the task execution limiting parameter includes a task limit concurrency number, a time window parameter, and / or a task limit time period.
[0032] Specifically, the task execution limiting parameter matched by the abnormal task can be queried according to the instance deployment data and the abnormality processing parameter of the abnormal task. If the query result is not empty, the abnormal task is verified based on the task execution limiting parameter. If the query result is empty, it is determined that the task execution verification is passed.
[0033] On this basis, in order to improve the comprehensiveness and refinement of the task execution verification, in a first optional implementation provided in the embodiment, in the process of performing task execution verification on the abnormal task in the initialization state based on the task execution limiting parameter, the abnormal task matched by a specific task state is queried according to the instance deployment data and / or the abnormality processing parameter of the abnormal task in the initialization state, and the abnormal task is verified based on the number of abnormal tasks and the task limit concurrency number, and the following operation can be specifically performed: query the abnormal task matched by the to-be-executed state and the executing state according to the instance deployment data and the abnormality processing parameter of the abnormal task in the initialization state; If the number of abnormal tasks is less than the task limit concurrency number, it is determined that the task execution verification is passed.
[0034] The specific task state includes a to-be-executed state and / or an executing state; the task limit concurrency refers to limiting the concurrency of abnormal tasks, for example, the task limit concurrency is maxConcurrency.
[0035] Specifically, in the process of querying the abnormal tasks in the to-be-executed state and the executing state that match the instance deployment data and the abnormal processing parameters of the abnormal task in the to-be-executed state and the initialized state according to the instance deployment data and / or the abnormal processing parameters, the abnormal tasks in the to-be-executed state and the executing state that match the instance deployment data and the abnormal processing parameters can be queried; in the process of determining that the task execution verification passes if the number of tasks is less than the task limit concurrency in the task execution limit parameter, specifically, if the task limit concurrency is single, it is determined that the task execution verification of the abnormal task passes if the number of tasks is less than the task limit concurrency, and it is determined that the task execution verification of the abnormal task fails if the number of tasks is greater than or equal to the task limit concurrency; if the task limit concurrency is multiple, the task limit concurrency whose ranking position is before a preset position in the multiple task limit concurrencies is determined, it is determined that the task execution verification of the abnormal task passes if the number of tasks is less than the determined task limit concurrency, and it is determined that the task execution verification of the abnormal task fails if the number of tasks is greater than or equal to the determined task limit concurrency; here, the sorting can be performed in the order from small to large.
[0036] In addition, in order to avoid resource waste caused by re-execution of the abnormal tasks that are successfully executed for the same processing dimension in a short time, improve resource utilization, and improve the efficiency of abnormal task processing, the abnormal tasks that do not need to be executed can be comprehensively identified; in the second optional implementation provided in the embodiment, in the process of performing task execution verification on the abnormal tasks in the initialized state based on the task execution limit parameter, the abnormal tasks in the executed state are obtained according to the time window parameter in the task execution limit parameter, and the task execution verification is performed according to the number of target abnormal tasks in the obtained abnormal tasks that match the instance deployment data and / or the abnormal processing parameters of the abnormal task in the initialized state, and the following operations can be performed: The abnormal tasks in the executed state are obtained according to the time window parameter in the task execution limit parameter, and the target abnormal tasks that match the instance deployment data and the abnormal processing parameters are selected from the obtained abnormal tasks in the executed state; If the number of target abnormal tasks is less than a preset number of tasks, it is determined that the task execution verification passes.
[0037] Specifically, the abnormal task in the corresponding time window in the execution success state can be acquired according to the time window parameter, the target abnormal task is filtered from the acquired abnormal task according to the instance deployment data and the abnormal processing parameter, if the number of the target abnormal task is less than the preset number of tasks, it is determined that the task execution verification is passed, if the number of the target abnormal task is greater than or equal to the preset number of tasks, it is determined that the task execution verification is not passed; the preset number of tasks can be any number, for example, 1; more specifically, if the time window parameter is multiple, the time window parameter whose ranking position is before the preset position can be determined from the multiple time window parameters, and the abnormal task in the execution success state in the corresponding time window is acquired according to the time window parameter, here the sorting can be performed in the order from small to large.
[0038] For example, the time window parameter is adjacent timeWindow seconds, that is, the past timeWindow seconds, the abnormal task in the execution success state in the past timeWindow seconds is queried, the target abnormal task is filtered from the queried abnormal task according to the instance deployment data and the abnormal processing parameter, if the number of the target abnormal task is less than the preset number of tasks, it is determined that the task execution verification is passed, if the number of the target abnormal task is greater than or equal to the preset number of tasks, it is determined that the task execution verification is not passed.
[0039] In addition, in order to avoid repeated execution of abnormal tasks for the same repair object and / or the same task type in the limited time period, and to improve the accuracy and efficiency of task execution, in the third optional implementation provided by the embodiment, in the process of performing task execution verification on the abnormal task in the initialization state based on the task execution limitation parameter, the abnormal task in the preset task state matched by the abnormal task in the task limitation time period is queried based on the repair object field and the task type contained in the abnormal task, and the task execution verification is performed according to the number of the queried abnormal task, here the task execution verification can be idempotent verification, and the following operations can be specifically performed: The abnormal task in the preset task state matched by the abnormal task in the task limitation time period is queried based on the repair object field and the task type contained in the abnormal task in the initialization state. If the number of the queried abnormal task is less than the task number threshold, it is determined that the task execution verification is passed.
[0040] Optionally, the preset task state includes an execution success state, an execution state and / or a to-be-executed state.
[0041] The task type refers to the repair type of the abnormal repair performed by the abnormal task, for example, the task type includes a model replacement type and / or an expansion type, and in addition, the task type can also include other types.
[0042] Specifically, the abnormal task matching the preset task state in the time period limited by the task execution limit parameter can be queried based on the repair object field and the task type contained in the abnormal task. If the number of the queried abnormal tasks is less than the task number threshold, it is determined that the task execution verification passes. If the number of the queried abnormal tasks is greater than or equal to the task number threshold, it is determined that the task execution verification fails. Alternatively, if the query result is empty, it is determined that the task execution verification passes. If the query result is not empty, it is determined that the task execution verification fails.
[0043] For example, the time period is adjacent to idempotentWindow seconds, that is, the past idempotentWindow seconds. Based on the repair object field and the task type contained in the abnormal task, it is queried whether there is an abnormal task in the past idempotentWindow seconds that matches the preset task state. If the query result is empty, it is determined that the task execution verification passes. If the query result is not empty, it is determined that the task execution verification fails.
[0044] It should be noted that the three implementation manners of the task execution verification of the abnormal task based on the task execution limit parameter provided above can be used independently or in combination. In the specific combination process, any one, two or three of the three implementation manners can be selected for serial execution, that is, the execution order is executed, and the latter is executed in the case where the former verification passes. The execution order is not limited, for example, the execution order is the first, the second, and then the third. Alternatively, any one, two or three of the three implementation manners can be selected for parallel execution, that is, in the case where any one, two or three of the three implementation manners passes the verification, it is determined that the task execution verification passes. Otherwise, it is determined that the task execution verification fails.
[0045] After the task execution verification, a verification result is obtained, and the task state is updated based on the verification result. In the process of updating the task state based on the verification result, the following can be specifically performed. If the task execution verification passes, the abnormal task is updated from the initialization state to the to-be-executed state. If the task execution verification fails, the reason for the failure of the verification is recorded for the abnormal task. The reason for the failure of the verification includes the failure of the concurrency limit, the failure of the time window limit, and / or the failure of the time period limit. The failure of the concurrency limit means that the task execution verification fails in the first implementation manner. The failure of the time window limit means that the task execution verification fails in the second implementation manner. The failure of the time period limit means that the task execution verification fails in the third implementation manner. The reason for the failure of the verification can be recorded in the flow limiting field, for example, the flow limiting field is extInfo.
[0046] In step S206, task execution is performed based on the abnormal task containing abnormal processing parameters of the target task state in the task set, to perform abnormal repair processing of the model instance of the corresponding artificial intelligence model.
[0047] In the above, based on the instance deployment data and abnormal processing parameters of the model instance contained in the abnormal task of the target task state, task execution verification is performed, and task state updating is performed based on the verification result. In this step, the abnormal task of the target task state is executed based on the abnormal processing parameters contained in the abnormal task of the target task state in the task set, to perform abnormal repair processing of the model instance of the corresponding artificial intelligence model, so as to realize the orderliness of the abnormal repair of the model instance; wherein the target task state includes a to-be-executed state and / or an executing state; specifically, after the task execution condition of the task set is triggered, the abnormal task of the target task state in the task set is executed based on the abnormal processing parameters contained in the abnormal task of the target task state, to perform abnormal repair processing of the model instance of the corresponding artificial intelligence model; here, the task execution condition can include expiration of the task execution period of the task set.
[0048] In specific implementation, the abnormal task in the to-be-executed state in the task set can be determined, and the abnormal task is updated from the to-be-executed state to the executing state, the execution module corresponding to the execution module field contained in the abnormal task in the executing state is called, and the abnormal repair processing of the model instance of the artificial intelligence model corresponding to the abnormal task is performed according to the repair action field contained in the abnormal task in the executing state; or the abnormal task in the executing state can be sent to an operation and maintenance platform, and the abnormal repair processing of the model instance of the artificial intelligence model corresponding to the abnormal task is performed according to the repair action field contained in the abnormal task in the executing state.
[0049] In actual application, the resources of the local server can be limited, so as to more efficiently perform abnormal repair and improve the stability of the artificial intelligence model. In an optional implementation provided by the embodiment, in the process of performing task execution based on the abnormal processing parameters contained in the abnormal task of the target task state in the task set, to perform abnormal repair processing of the model instance of the corresponding artificial intelligence model, the abnormal task of the target task state is uploaded to the detection server corresponding to the execution module field contained in the abnormal task of the target task state, to perform abnormal repair of the model instance according to the repair action field contained in the abnormal task of the target task state. Specifically, the following operations can be performed: The execution module field contained in the abnormal task of the target task state is read. The abnormal task of the target task state is uploaded to the detection server corresponding to the execution module field, to perform abnormal repair of the model instance according to the repair action field contained in the abnormal task of the target task state.
[0050] Specifically, in the process of reading the execution module field contained in the target task state of the abnormal task, the abnormal task in the task set in the to-be-executed state can be determined, and the abnormal task is updated from the to-be-executed state to the executing state, and the execution module field in the executing state of the abnormal task is read; in the process of performing abnormal repair of the model instance according to the repair action field contained in the target task state of the abnormal task, the repair process can be arranged according to the repair action field and the repair object field contained in the target task state of the abnormal task, and the abnormal repair process is obtained, and the abnormal repair of the model instance of the corresponding artificial intelligence model is performed according to the abnormal repair process.
[0051] After the above task execution based on the abnormal processing parameter contained in the target task state of the abnormal task in the task set, the task state is updated based on the task execution result, for example, the executing state of the abnormal task is updated to the execution success state, the execution failure state or the timeout state; In addition, the timestamp of the abnormal task can also be recorded, and the timestamp includes the timestamp of updating the abnormal task to the initialization state, the to-be-executed state, the executing state, the execution success state, the execution failure state and / or the timeout state.
[0052] The artificial intelligence model in the embodiment can be multiple, each artificial intelligence model has a corresponding application service, and each model instance of the artificial intelligence model can be one or more. Each model instance of each artificial intelligence model has a corresponding cloud, that is, the cloud adopts a distributed architecture, and the local server adopts a centralized architecture. The local server can generate a data collection strategy, an abnormal processing strategy and / or a task execution restriction parameter for each model instance of each artificial intelligence model; for example Figure 3 As shown, only the cloud corresponding to one model instance is shown, and the other model instances are similar. The cloud includes a data collection module (Collector), a decision unit (Biz-Diagnosis-Controller) and an execution engine (Biz-Failover-Operator). The data collection module is used to collect instance running data of the model instance. The index processor in the decision unit calculates the instance running index based on the instance running data, determines the target abnormal processing strategy based on the instance running index and the instance deployment data of the model instance in the abnormal processing strategy, extracts the repair action field, the repair object field and the execution module field from the target abnormal processing strategy, generates an abnormal task according to the instance deployment data and the repair action field, the repair object field and the execution module field, and sends the abnormal task to the local server through an interface call. The instance running data is obtained based on the running data collection of the data collection strategy, and the data collection strategy and the abnormal processing strategy are obtained by periodically calling the query interface of the strategy management center of the local server (main station) for data query. The local server comprises an exception detection system and a high availability center, the exception detection system comprises a monitor, logs and resource management, and the high availability center comprises an emergency module, a detection module, a policy management center, an audit center and a data dashboard; after receiving an exception task, the emergency module of the local server marks the exception task as an initialization state and writes the exception task into a task set, the task query module of the local server queries the exception task in the initialization state in the task set at a regular time, queries a task execution restriction parameter matched with the exception task, calls the task verification module to perform task execution verification on the exception task in the initialization state based on the task execution restriction parameter, updates the task state of the exception task according to the verification result, calls the task execution module to query the exception task in a to-be-executed state in the task set at a regular time, and performs task execution on the queried exception task to perform exception repair processing on the model instance of the corresponding artificial intelligence module; the local server can report the task execution result to the storage backup module of the decision unit of the cloud or store the task execution result in the audit center by the local server; the policy management center of the local server is used for generating a data collection strategy, an exception processing strategy and a task execution restriction parameter for the model instance of the artificial intelligence model, and the data dashboard is used for visually displaying the execution of the exception task, the running state of the artificial intelligence model and the like. After updating the task state of the exception task according to the verification result, the local server can also send the exception task in the to-be-executed state to the corresponding exception repair execution engine of the cloud, perform task execution on the exception task through the exception repair execution engine, and send the task execution result to the audit center of the local server.
[0053] It should be noted that the steps S204 to S206 can be replaced by performing task execution verification according to the instance deployment data and / or the exception processing parameter of the model instance contained in the exception task, and after the verification passes, performing task execution based on the exception processing parameter of the exception task in the target task state in the task set to perform exception repair processing on the model instance of the corresponding artificial intelligence model; or the step S204 can be replaced by performing task execution verification according to the instance deployment data and / or the exception processing parameter of the model instance contained in the exception task, and updating the task state based on the verification result; and the step S206 can be replaced by performing exception repair processing on the model instance of the corresponding artificial intelligence model based on the exception processing parameter of the exception task in the target task state in the task set.
[0054] It should be noted that each optional implementation and each feasible execution manner in steps S202 to S206 provided by the embodiment can be independently executed as needed, or can be combined with each other and referred to each other, and each specific execution step in each optional implementation or each feasible execution manner can also be independently executed or executed in combination as needed. The execution condition of "if" or "under what circumstances" involved in each step or operation can be directly deleted, and the operation after the execution condition can be executed subsequently, and the embodiment does not make a specific limitation on this.
[0055] In summary, the processing method of one or more artificial intelligence models provided by the embodiment, after querying the abnormal task in the initialization state in the task set for model anomaly repair, queries the task execution restriction parameter matched with the abnormal task according to the instance deployment data and the abnormal processing parameter contained in the abnormal task, performs task execution verification on the abnormal task based on the task execution restriction parameter, updates the task state of the abnormal task based on the verification result, and then changes the abnormal task in the to-be-executed state in the task set to the executing state after the task execution condition of the task set is triggered. The abnormal processing parameter contained in the abnormal task in the executing state is executed to perform anomaly repair processing on the model instance of the corresponding artificial intelligence model. In this way, the task execution verification avoids executing too many abnormal tasks in a short time to cause impact on the system, improves the orderliness and stability of anomaly repair, and improves the automation and initiative of anomaly detection by the detection server corresponding to the model instance of the artificial intelligence model to generate the abnormal task, and improves the timeliness and efficiency of anomaly repair.
[0056] The application of the processing method of an artificial intelligence model provided by the embodiment in an AI model scenario is taken as an example to further illustrate the processing method of an artificial intelligence model provided by the embodiment, which is shown in Figure 4 The processing method of an artificial intelligence model applied in an AI model scenario specifically includes the following steps.
[0057] Step S402: Querying an abnormal task in an initialization state in a task set for model anomaly repair.
[0058] Optionally, the abnormal task is generated after the detection server performs anomaly detection on the model instance of the AI model corresponding to each application service.
[0059] Step S404: Querying a task flow limiting parameter matched with the abnormal task according to the instance deployment data of the model instance and a repair action field contained in the abnormal task.
[0060] Step S406: Performing task flow limiting verification on the abnormal task based on the task flow limiting parameter, and updating the task state based on the verification result.
[0061] Step S408, determine the abnormal task in the to-be-executed state in the task set, and update the abnormal task in the to-be-executed state to the executing state.
[0062] Step S410, read the execution module field contained in the abnormal task in the executing state.
[0063] Step S412, upload the abnormal task in the executing state to the cloud corresponding to the execution module field, so as to perform abnormal repair on the model instance of the corresponding AI model according to the repair action field contained in the abnormal task in the executing state.
[0064] It should be noted that any one or any combination of steps S402 to S412 can be replaced by the corresponding technical means provided in steps S202 to S206 according to the needs of implementation and deployment, and steps S402 to S412 can also be combined into a new implementation manner according to the needs of implementation and deployment. Any one or any combination of steps S402 to S412 can also be combined into a new implementation manner according to the needs of actual deployment and one or more steps provided in steps S202 to S206, or combined into a new implementation manner with one or more optional embodiments provided in steps S202 to S206, which will not be described here.
[0065] The processing device of the artificial intelligence model provided in the specification implements, for example: In the above embodiments, a processing method of an artificial intelligence model is provided, and a processing device of an artificial intelligence model is also provided. The following will be described with reference to the accompanying drawings.
[0066] Reference Figure 5 It shows a schematic diagram of an embodiment of a processing device of an artificial intelligence model provided in the embodiment.
[0067] Since the device embodiment corresponds to the method embodiment, the description is relatively simple, and the related parts can be seen from the above-mentioned corresponding description of the method embodiment. The device embodiment described below is only illustrative.
[0068] The processing device of the artificial intelligence model provided in the embodiment comprises: The task query module 502 is configured to query the abnormal task in the initialization state in the task set for model abnormal repair; the abnormal task is generated after the detection server detects the model instance of each artificial intelligence model; The checking module 504 is configured to perform task execution checking according to the instance deployment data and the exception handling parameter of the model instance contained in the exception task, and perform task state updating based on the checking result. The execution module 506 is configured to perform task execution based on the exception handling parameter of the exception task in the target task state in the task set, so as to perform exception repair processing on the model instance of the corresponding artificial intelligence model.
[0069] The present specification provides an artificial intelligence model processing device, which implements the following, for example: Based on the same technical concept, one or more embodiments of the present specification also provide an artificial intelligence model processing device for executing the artificial intelligence model processing method provided in the present specification, Figure 6 A structural schematic diagram of an artificial intelligence model processing device provided by one or more embodiments of the present specification.
[0070] The artificial intelligence model processing device provided in the present embodiment comprises: As Figure 6 As shown, the artificial intelligence model processing device can have great differences due to different configurations or performances, and can comprise one or more processors 601 and memories 602. The memories 602 can store one or more application programs or data. The memories 602 can be temporary memories or persistent memories. The application programs stored in the memories 602 can comprise one or more modules (not shown in the figure), and each module can comprise a series of computer executable instructions in the artificial intelligence model processing device. Further, the processor 601 can be configured to communicate with the memory 602 and execute the series of computer executable instructions in the memory 602. The artificial intelligence model processing device can further comprise one or more power supplies 603, one or more wired or wireless network interfaces 604, one or more input / output interfaces 605, one or more keyboards 606, and the like.
[0071] In a specific embodiment, the artificial intelligence model processing device comprises a memory and one or more programs, wherein one or more programs are stored in the memory, and the one or more programs can comprise one or more modules, and each module can comprise a series of computer executable instructions in the artificial intelligence model processing device, and the one or more processors are configured to execute the one or more programs, which comprise the following computer executable instructions: query an abnormal task in an initialization state in a task set for model anomaly repair; the abnormal task is generated after an anomaly detection server detects an anomaly of a model instance of each artificial intelligence model; perform task execution verification according to instance deployment data and anomaly processing parameters of the model instance contained in the abnormal task, and perform task state updating based on a verification result; perform task execution based on anomaly processing parameters contained in an abnormal task in a target task state in the task set, to perform anomaly repair processing of a model instance of a corresponding artificial intelligence model.
[0072] The computer-readable storage medium provided in the specification implements, for example, the following: According to the same technical concept, one or more embodiments of the specification also provide a computer-readable storage medium corresponding to the processing method of the artificial intelligence model described above.
[0073] The computer-readable storage medium provided in the embodiment is used to store computer executable instructions, and the computer executable instructions realize the following steps when executed: query an abnormal task in an initialization state in a task set for model anomaly repair; the abnormal task is generated after an anomaly detection server detects an anomaly of a model instance of each artificial intelligence model; perform task execution verification according to instance deployment data and anomaly processing parameters of the model instance contained in the abnormal task, and perform task state updating based on a verification result; perform task execution based on anomaly processing parameters contained in an abnormal task in a target task state in the task set, to perform anomaly repair processing of a model instance of a corresponding artificial intelligence model.
[0074] It should be noted that the embodiments of the computer-readable storage medium in the specification and the embodiments of the processing method of the artificial intelligence model in the specification are based on the same inventive concept, so the specific implementation of the embodiments can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0075] The computer program product provided in the specification implements, for example, the following: According to the same technical concept, one or more embodiments of the specification also provide a computer program product corresponding to the processing method of the artificial intelligence model described above.
[0076] A computer program product includes computer programs / instructions, which, when executed by a processor, realize the following steps: query an initialization state of an abnormal task in a task set of model anomaly repair; the abnormal task is generated after the detection server detects the model instance of each artificial intelligence model; perform task execution verification according to the instance deployment data and the abnormal processing parameter of the model instance contained in the abnormal task, and perform task state updating based on the verification result; perform task execution based on the abnormal processing parameter contained in the abnormal task of the target task state in the task set, to perform abnormal repair processing on the model instance of the corresponding artificial intelligence model.
[0077] It should be noted that the embodiments of the computer program product in the present specification and the embodiments of the processing method of the artificial intelligence model in the present specification are based on the same inventive concept, and therefore the specific implementation of the embodiments can be referred to the foregoing implementation of the corresponding method, and the repeated parts will not be described.
[0078] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments, such as the device embodiment, the equipment embodiment and the computer readable storage medium embodiment, which are similar to the method embodiment, so the description is relatively simple. The related content in the device embodiment, the equipment embodiment and the computer readable storage medium embodiment can be referred to the part of the description of the method embodiment.
[0079] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in the embodiments and still achieve desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order in order to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous or possible.
[0080] In the 1930s, it was clear to distinguish whether an improvement in a technology was in hardware (e.g., improvement in circuit structure of diodes, transistors, switches, etc.) or in software (e.g., improvement in method flow). However, as technology has evolved, many improvements in method flow today can be considered as direct improvements in hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using hardware entity modules. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming it, rather than by asking a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating integrated circuit chips, this programming is now mostly implemented using "logic compiler" software, which is similar to software compilers used in program development, and the original code before compilation is written in a specific programming language, which is called a hardware description language (HDL), and there are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should be aware that only a little logical programming of the method flow in the above-mentioned hardware description languages and programming into an integrated circuit can easily obtain a hardware circuit that implements the logical method flow.
[0081] The controller can be implemented in any suitable way, for example, the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to implementing the controller in pure computer readable program code, it is also possible to implement the controller in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers, etc. to perform the same functions by logically programming the method steps. Such a controller can therefore be considered as a hardware component, and the means included therein for performing various functions can also be considered as structures within the hardware component. Alternatively, the means for performing various functions can even be considered as both a software module implementing the method and a structure within the hardware component.
[0082] The systems, apparatuses, modules or units illustrated by the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0083] For the sake of description, the above apparatuses are described in various units with functions respectively. Of course, the functions of the units can be implemented in one or more software and / or hardware in implementing the embodiments of the present specification.
[0084] Those skilled in the art will understand that one or more embodiments of the present specification can be provided as a method, a system or a computer program product. Therefore, one or more embodiments of the present specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROMs, optical storage devices, etc.) containing computer usable program code.
[0085] The specification is presented with reference to flow diagrams and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the specification. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing system or other programmable test processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable test processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks. Figure 1 The flow diagram and / or block diagram in the flow diagrams and / or block diagrams can represent one or more of any appropriate circuitry configured to perform the specified functions. In this regard, one or more flow diagrams and / or block diagrams in the flow diagrams and / or block diagrams can represent a device, such as a system, that is configured to perform specified functions, e.g., as described below in the discussion of the flow diagrams and / or block diagrams. Figure 1 The flow diagram and / or block diagram in the flow diagrams and / or block diagrams can represent one or more of any appropriate circuitry configured to perform the specified functions. In this regard, one or more flow diagrams and / or block diagrams in the flow diagrams and / or block diagrams can represent a device, such as a system, that is configured to perform specified functions, e.g., as described below in the discussion of the flow diagrams and / or block diagrams.
[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable test processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flow diagrams and / or block diagrams flow or flows and / or block or blocks specified in the flow diagrams and / or block diagrams. Figure 1 The flow diagram and / or block diagram in the flow diagrams and / or block diagrams can represent one or more of any appropriate circuitry configured to perform the specified functions. In this regard, one or more flow diagrams and / or block diagrams in the flow diagrams and / or block diagrams can represent a device, such as a system, that is configured to perform specified functions, e.g., as described below in the discussion of the flow diagrams and / or block diagrams. Figure 1 The flow diagram and / or block diagram in the flow diagrams and / or block diagrams can represent one or more of any appropriate circuitry configured to perform the specified functions. In this regard, one or more flow diagrams and / or block diagrams in the flow diagrams and / or block diagrams can represent a device, such as a system, that is configured to perform specified functions, e.g., as described below in the discussion of the flow diagrams and / or block diagrams.
[0087] These computer program instructions can also be loaded into a computer or other programmable test processing apparatus to cause a series of operational steps to be performed on the computer or other programmable test apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable test apparatus provide steps for implementing the flow diagrams and / or block diagrams flow or flows and / or block or blocks specified in the flow diagrams and / or block diagrams. Figure 1 The flow diagram and / or block diagram in the flow diagrams and / or block diagrams can represent one or more of any appropriate circuitry configured to perform the specified functions. In this regard, one or more flow diagrams and / or block diagrams in the flow diagrams and / or block diagrams can represent a device, such as a system, that is configured to perform specified functions, e.g., as described below in the discussion of the flow diagrams and / or block diagrams. Figure 1 The flow diagram and / or block diagram in the flow diagrams and / or block diagrams can represent one or more of any appropriate circuitry configured to perform the specified functions. In this regard, one or more flow diagrams and / or block diagrams in the flow diagrams and / or block diagrams can represent a device, such as a system, that is configured to perform specified functions, e.g., as described below in the discussion of the flow diagrams and / or block diagrams.
[0088] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0089] The memory can include non-persistent memory and / or volatile memory, e.g., random access memory (RAM) and / or non-volatile memory, e.g., read-only memory (ROM) or flash memory, among others. The memory is an example of computer-readable media.
[0090] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0091] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, such that processes, methods, articles or devices that include a series of features are not only limited to those features, but also include other features not explicitly listed, or inherent to such processes, methods, articles or devices. Without more limitations, the feature defined by the statement "comprising a" does not exclude the presence of other identical features in the process, method, article or device comprising the feature.
[0092] One or more embodiments of the specification can be described in the general context of computer-executable instructions being executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the specification can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0093] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment focuses on the difference from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
[0094] The above merely provides the example of the present document and is not intended to limit the present document. For those skilled in the art, the present document can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present document shall be included in the scope of claims of the present document.
Claims
1. A method for processing an artificial intelligence model, the method comprising: Query the set of tasks that are in the initialization state for model anomaly repair; The anomaly task is generated after the detection server performs anomaly detection on the model instances of each artificial intelligence model; Based on the instance deployment data and exception handling parameters of the model instance included in the abnormal task, the task execution is verified, and the task status is updated based on the verification results. Based on the exception handling parameters contained in the exception tasks of the target task state in the task set, the task is executed to perform exception repair processing on the model instance of the corresponding artificial intelligence model.
2. The processing method for the artificial intelligence model according to claim 1, wherein the anomaly detection includes: Based on the instance running data of each artificial intelligence model instance, the running metrics are calculated to obtain the instance running metrics; If the instance's runtime metrics trigger an instance exception condition in the exception handling strategy, the target exception handling strategy is detected based on the instance deployment data in the exception handling strategy.
3. The processing method for the artificial intelligence model according to claim 2, wherein the abnormal task in the initialization state is generated in the following manner: Extract the repair action field, repair object field, and execution module field from the target anomaly handling strategy; Based on the instance deployment data, the repair action field, the repair object field, and the execution module field, an abnormal task in the initialization state is generated.
4. The processing method for the artificial intelligence model according to claim 2, wherein the detection server is deployed in the cloud, and the cloud corresponds one-to-one with the model instance; The instance running data is obtained by collecting instance running data from the model instances of each artificial intelligence model based on the data collection strategy; the data collection strategy is obtained by calling the query interface of the local server in the cloud with the instance deployment data as the interface input.
5. The method for processing an artificial intelligence model according to claim 1, wherein the step of performing task execution verification based on instance deployment data and exception handling parameters of the model instance included in the abnormal task comprises: Based on the instance deployment data and the exception handling parameters contained in the exception task of the initialization state, query the task execution restriction parameters that match the exception task of the initialization state. Based on the task execution restriction parameters, the abnormal tasks in the initialization state are checked for task execution.
6. The processing method for the artificial intelligence model according to claim 5, wherein the step of performing task execution verification on the abnormal task in the initialization state based on the task execution constraint parameter includes: Based on the instance deployment data and the exception handling parameters contained in the exception task of the initialization state, query the matching exception tasks in the pending execution state and the execution state. If the number of abnormal tasks found in the query is less than the task concurrency limit, the task execution verification is considered successful.
7. The processing method for the artificial intelligence model according to claim 5, wherein the step of performing task execution verification on the abnormal task in the initialization state based on the task execution constraint parameter includes: According to the time window parameter in the task execution limit parameters, obtain the abnormal tasks with successful execution status, and filter out the target abnormal tasks that match the instance deployment data and abnormal handling parameters from the obtained abnormal tasks. If the number of target abnormal tasks is less than the preset number of tasks, the task execution verification is deemed successful.
8. The processing method for the artificial intelligence model according to claim 5, wherein the step of performing task execution verification on the abnormal task in the initialization state based on the task execution constraint parameter includes: Based on the repair object field and task type included in the abnormal task in the initialization state, query the abnormal task in the initialization state that matches the abnormal task in the preset task state within the task restriction time period. If the number of abnormal tasks found is less than the task number threshold, the task execution verification is considered successful. The preset task status includes execution successful status, execution in progress status, and / or pending execution status.
9. The processing method for the artificial intelligence model according to claim 1, wherein the step of executing the task based on the exception handling parameters contained in the exception task of the target task state in the task set to perform exception repair processing for the model instance of the corresponding artificial intelligence model includes: Read the execution module fields contained in the abnormal task of the target task status; The abnormal task in the target task status is uploaded to the detection server corresponding to the field of the execution module, so as to repair the abnormality of the model instance according to the repair action field contained in the abnormal task in the target task status.
10. The processing method of the artificial intelligence model according to claim 1, wherein each artificial intelligence model is deployed on the application service of the application platform and corresponds one-to-one with the application service.
11. A processing device for an artificial intelligence model, comprising: The task query module is configured to query abnormal tasks in the initialization status from the task set for model anomaly repair. The anomaly task is generated after the detection server performs anomaly detection on the model instances of each artificial intelligence model; The verification module is configured to perform task execution verification based on the instance deployment data and exception handling parameters of the model instance contained in the abnormal task, and update the task status based on the verification results. The execution module is configured to perform task execution based on the exception handling parameters contained in the exception task of the target task state in the task set, so as to perform exception repair processing of the model instance of the corresponding artificial intelligence model.
12. A processing device for an artificial intelligence model, comprising: processor; And, a memory configured to store computer-executable instructions, which, when executed, cause the processor to: The system queries the set of tasks for model anomaly repair to identify tasks in their initialization state; these anomaly tasks are generated by the detection server after it performs anomaly detection on the model instances of each artificial intelligence model. Based on the instance deployment data and exception handling parameters of the model instance included in the abnormal task, the task execution is verified, and the task status is updated based on the verification results. Based on the exception handling parameters contained in the exception tasks of the target task state in the task set, the task is executed to perform exception repair processing on the model instance of the corresponding artificial intelligence model.
13. A computer-readable storage medium for storing computer-executable instructions that, when executed, implement the steps of the method of claim 1.