Scene application operation maintenance platform based on artificial intelligence model
By sensing resource consumption, identifying abnormal tasks and optimizing resource scheduling, the problem of insufficient dynamic adaptability of resource scheduling and task deployment in the artificial intelligence model operation and maintenance platform is solved, and efficient balance of resource allocation and improved stability of task deployment are achieved.
Patent Information
- Application Number
- CN202510795997.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-16
AI Technical Summary
In the existing technology, the scenario application operation and maintenance platform of the artificial intelligence model lacks adaptability to dynamic changes in resource scheduling and task deployment, resulting in delayed abnormal state recognition, resource competition conflicts and performance bottlenecks, affecting the overall reasoning efficiency and stability.
The inference state perception module obtains resource consumption data, identifies abnormal tasks, and generates an abnormal inference list; the load balancing scheduling module screens resource-conflicting task groups and generates a conflicting resource allocation structure diagram; the resource intelligent allocation module migrates tasks to nodes with sufficient idle rates; and the deployment configuration adjustment module screens stable paths based on multi-dimensional indicators and generates a model configuration deployment mapping table.
It achieves accurate identification of resource fluctuations and improvement of task stability, improves the targeted resource allocation and balanced node utilization, ensures the latency and stability of task deployment, and enhances the accuracy of operation and maintenance response and scenario adaptability.
Smart Images

Figure CN120653398A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a scenario application operation and maintenance platform based on an artificial intelligence model. Background Art
[0002] The field of artificial intelligence technology encompasses multiple sub-directions, including computer vision, natural language processing, speech recognition, and machine learning. Its core focus is on enabling computers to simulate, extend, and expand human intelligent behavior by building models capable of autonomous learning and reasoning. Data acquisition, model training, inference execution, and continuous optimization are key components of the development of artificial intelligence. Artificial intelligence technology encompasses a complete process, from raw data collection and preprocessing to model building and training, model deployment and inference, and subsequent performance monitoring and maintenance. This technology also emphasizes the synergy between algorithms and computing resources, as well as cross-platform and cross-scenario deployment and management capabilities, to meet diverse application needs across diverse industries, including industry, healthcare, transportation, and finance.
[0003] Among them, the scenario application operation and maintenance platform refers to a management and control platform built for the configuration deployment, status monitoring, fault diagnosis, resource scheduling and other matters involved in the operation process of artificial intelligence applications in actual scenarios. It covers technical matters such as operation parameter setting, reasoning resource allocation, execution process tracking, operation anomaly identification and response mechanism after the artificial intelligence model is launched in a specific scenario. It is mainly achieved by building a unified configuration interface, establishing operation status detection rules based on log analysis, introducing multi-dimensional resource allocation strategies, and constructing a diagnostic mechanism with conditional trigger logic.
[0004] Existing technologies for monitoring the running status of inference tasks rely on preset rules and static thresholds, making them difficult to adapt to dynamic changes in task resource usage. This results in delayed recognition of abnormal conditions and the platform's inability to respond promptly to performance degradation during peak load periods. During resource scheduling, existing methods often allocate tasks based on fixed strategies, lacking a comprehensive analysis of task conflict structures. This makes it difficult to identify resource competition among multiple tasks within the same node, leading to the accumulation of performance bottlenecks. For example, in scenarios with limited GPU resources, scheduling multiple high-intensity tasks simultaneously to the same node can lead to GPU saturation and task response delays, impacting overall inference efficiency. Regarding task migration and deployment adjustments, existing processes ignore the comprehensive evaluation of multi-dimensional performance indicators of the target node and instead migrate tasks based solely on resource availability. This results in unstable execution of migrated tasks on the target node and causes performance fluctuations. At the operational feedback level, existing technologies lack an automatic mechanism to identify drift in task response intervals, making it difficult to promptly detect changes in execution efficiency. This impacts the ability to respond to performance anomalies in typical application scenarios such as urban visual recognition and industrial equipment identification, and hinders the sustainable optimization of intelligent inference. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings of the existing technology and propose a scenario application operation and maintenance platform based on artificial intelligence model.
[0006] To achieve the above objectives, the present invention adopts the following technical solutions: A scenario application operation and maintenance platform based on an artificial intelligence model includes: The inference state perception module obtains the CPU usage, memory call, GPU scheduling, and bandwidth usage of the inference process in the edge node device's running table. It makes judgments based on the time ratio of each resource consumption and the offset trend of the task response time, filters out tasks whose fluctuation values exceed the response cycle threshold, and generates an abnormal inference identification list. The load balancing scheduling module extracts the resource allocation status of the restricted tasks on the same node based on the abnormal reasoning identification list, filters the resource conflict task group, identifies the task conflict distribution by the ratio of the number of tasks to the number of allocation records, and obtains the conflict resource allocation structure diagram; The resource intelligent allocation module extracts the task operation level and resource offset rate according to the conflict resource allocation structure diagram, and migrates the tasks with excessive offset rates to nodes with sufficient idle rates according to the GPU idle rate mapping record, thereby obtaining the model reasoning migration path; The deployment configuration adjustment module calls the model inference migration path, screens task instances whose synchronization change interval is lower than the load fluctuation range based on the storage call frequency, memory margin and network latency of the target node, writes the deployment identifier into the migration configuration, and generates a model configuration deployment mapping table.
[0007] As a further solution of the present invention, the abnormal reasoning identification list includes task identification information, resource abnormality type, response offset amplitude, and fluctuation period label; the conflict resource allocation structure diagram includes resource conflict nodes, allocation conflict relationship, task conflict intensity, and node resource utilization; the model reasoning migration path includes migration node sequence, resource offset level, task priority label, and GPU idle rate matching result; the model configuration deployment mapping table includes target node parameters, deployment task index, configuration matching index, and load fluctuation tolerance range.
[0008] As a further solution of the present invention, the reasoning state perception module includes: The resource monitoring submodule obtains the CPU usage, memory call, GPU scheduling, and bandwidth usage of the inference process in the edge node device running table, identifies the corresponding sequence of resource indicators and time, and obtains the model resource usage trajectory; The performance offset identification submodule calls the model resource occupancy trajectory, calculates the occupancy ratio of each resource during the task execution cycle, extracts the model response time series, compares the difference ratio between the resource share and the response time change amplitude under the same index, and generates the performance response offset degree; The abnormal task screening submodule calls the performance response deviation, compares the task deviation with the response period threshold item by item, screens the model tasks whose deviation exceeds the threshold, sorts them by the deviation degree, and generates an abnormal reasoning identification list.
[0009] As a further solution of the present invention, the load balancing scheduling module includes: The conflict identification submodule extracts the resource allocation records of the restricted tasks on the same node based on the abnormal reasoning identification list, compares the task resource requirements with the node resource margin, filters out tasks whose resource requests exceed the node allocatable threshold, aggregates task numbers and resource conflict nodes, and generates a conflict task location table; The distribution comparison submodule calls the conflict task location table, counts the number of conflicting tasks and the total number of deployment records in the node, identifies the difference between task density and deployment pressure, calculates the deployment conflict distribution value of the current scenario node, hierarchically classifies the differentiated node operation loads according to the distribution value, analyzes the distribution of operation conflict hotspots in the scenario, and generates a node conflict risk layer; The structure construction submodule screens nodes in high-conflict areas based on the node conflict risk layer, extracts task execution trajectories and allocation path information, classifies the allocation frequency and sequence chain of differentiated tasks when migrating between nodes, analyzes the correlation between task paths and conflict density, and obtains a conflict resource allocation structure diagram.
[0010] As a further solution of the present invention, the resource intelligent allocation module includes: The offset identification submodule extracts the task operation level, GPU idle rate and resource offset rate according to the conflict resource allocation structure diagram, marks the task mapping with abnormal resource fluctuation according to the corresponding relationship between the task level and the offset rate, and generates a task offset mapping record; The node adaptation submodule calls the task offset mapping record, matches the task running level value with the GPU partition idle rate, compares the adaptation gap between the idle rate and the resource offset rate, calculates the node adaptation offset value, and screens the running nodes that meet the migration conditions in combination with the partition mapping table to establish a running node adaptation list; The path adjustment submodule calls the running node adaptation list, plans task migration according to the task running level, identifies the mapping sequence according to the node idle rate and level relationship, and marks the migration path to obtain the model reasoning migration path.
[0011] As a further solution of the present invention, the deployment configuration adjustment module includes: The path judgment submodule calls the model inference migration path, identifies the synchronous change range of node indicators between the current and previous weeks based on the storage call frequency, memory margin, and network latency of the target node, and makes a difference judgment with the load fluctuation amplitude, identifies the stable path and records the flag information, and generates a model inference stable path set; The task screening submodule calls the model to infer the stable path set, combines the task instance call frequency, resource allocation value and node load reconstruction index, calculates the task deployment interference degree, screens out task instances with interference degrees below the control threshold, associates them with the target path, and generates a low-interference deployment candidate set; The mapping identification submodule integrates the task deployment identifier and the stable path relationship information according to the low-interference deployment candidate set, updates the deployment status field and registers the migration configuration, maps the deployment mapping relationship between the model inference node and the task, and generates a model configuration deployment mapping table.
[0012] As a further solution of the present invention, the platform also includes a scene feedback integration module: Based on the model configuration deployment mapping table, the scenario feedback integration module extracts the scenario control terminal task response log and model interface call records, identifies the inference response duration, execution status, and call frequency in the dynamic task registry, and screens representative tasks with execution interval drift for typical applications such as urban visual recognition and industrial equipment health recognition to obtain an application scenario operation and maintenance list. The application scenario operation and maintenance list includes task execution frequency, response time change, execution status identification, and representative task records.
[0013] As a further solution of the present invention, the scene feedback integration module includes: The log parsing submodule collects the scenario control terminal task logs and model interface records based on the model configuration deployment mapping table, compares the timestamps according to the task numbers, identifies the task response intervals and interface call delays, and generates model call timing features; The response status evaluation submodule identifies the response duration and execution status fields based on the model call timing characteristics and the task number in the dynamic task registry, determines whether the response interval and duration offset exceed the set threshold, and filters the task fluctuation frequency in combination with the execution status to generate a model operation stable state set; The task offset identification submodule selects state samples of urban visual recognition and industrial equipment health identification task types according to the stable state set of the model operation, counts the number of continuous state fluctuations under the task number, and compares the call frequency to identify tasks with decreased stability and a frequency exceeding the benchmark value, and obtains the application scenario operation and maintenance list.
[0014] Compared with the prior art, the advantages and positive effects of the present invention are: By sensing fine-grained dynamic data on resource consumption in edge node inference processes and combining it with the drift trend of task response times, the present invention can accurately identify the impact of resource fluctuations on task stability, effectively screen out tasks with abnormal response period fluctuations, and enhance the sensitivity and real-time performance of inference status monitoring. By analyzing task deployment records on nodes, the distribution structure of resource conflicts is identified, improving the efficiency and visualization of identifying resource usage conflicts between tasks. Based on the conflict resource structure, the corresponding relationship between the model's operating level and GPU idle rate is extracted to optimize the migration path of tasks with high resource drift, improving the targeted resource allocation and balanced node utilization. During the migration deployment process, the coordination of multiple indicators such as the target node's storage call frequency, memory margin, and network round-trip time ensures that the latency and stability of task deployment remain within controllable ranges, reducing system jitter caused by load migration. By analyzing task response logs and interface call records, tasks with drifting execution intervals are identified, enabling performance backtracking and maintenance list generation for dynamic tasks on the application side, improving the accuracy of operation and maintenance responses and the flexibility of scenario adaptation, comprehensively enhancing the operational stability and resource scheduling intelligence of inference tasks in multiple scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a platform flow chart of the present invention; Figure 2 This is a flow chart of the reasoning state perception module in the present invention; Figure 3 This is a flow chart of the load balancing scheduling module in the present invention; Figure 4 This is a flow chart of the resource intelligent allocation module in the present invention; Figure 5 This is a flow chart of the deployment configuration adjustment module in the present invention; Figure 6 This is a flow chart of the scene feedback integration module in the present invention. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0017] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.
[0018] See also Figure 1 , the scenario application operation and maintenance platform based on artificial intelligence models includes: The inference state perception module obtains the CPU usage, memory call, GPU scheduling, and bandwidth usage of the inference process in the edge node device's running table. It makes judgments based on the time ratio of each resource consumption and the offset trend of the task response time, filters out tasks whose fluctuation values exceed the response cycle threshold, and generates an abnormal inference identification list. The load balancing scheduling module extracts the resource allocation status of restricted tasks on the same node based on the abnormal reasoning identification list, filters the resource conflict task groups, identifies the task conflict distribution through the ratio of the number of tasks to the number of allocation records, and obtains the conflict resource allocation structure diagram; The intelligent resource allocation module extracts the task operation level and resource offset rate from the inference scheduling area task control table, GPU partition mapping table, and model calculation level classification table based on the conflict resource allocation structure diagram. By comparing the task operation level with the idle rate in the GPU mapping record, it transfers tasks with offset rates exceeding the limit to nodes with sufficient idle rates, thereby obtaining the model inference migration path. The deployment configuration adjustment module calls the model inference migration path and selects task instances in the indicator group whose synchronous change interval is lower than the node load fluctuation range based on the storage call frequency, memory margin, and network round-trip time of the target node. It then writes the deployment identifier into the migration configuration and generates a model configuration deployment mapping table. The scenario feedback integration module extracts the scenario control terminal task response log and model interface call records based on the model configuration deployment mapping table, identifies the inference response duration, execution status and call frequency in the dynamic task registry, and screens representative tasks with execution interval drift for typical applications such as urban visual recognition and industrial equipment health recognition to obtain an application scenario operation and maintenance list.
[0019] The abnormal reasoning identification list includes task identification information, resource abnormality type, response offset amplitude, and fluctuation period label. The conflict resource allocation structure diagram includes resource conflict nodes, allocation conflict relationship, task conflict intensity, and node resource utilization. The model reasoning migration path includes migration node sequence, resource offset level, task priority label, and GPU idle rate matching result. The model configuration deployment mapping table includes target node parameters, deployment task index, configuration matching index, and load fluctuation tolerance range. The application scenario operation and maintenance list includes task execution frequency, response time change, execution status identification, and representative task records.
[0020] See also Figure 2 , the reasoning state perception module includes: The resource monitoring submodule obtains the CPU usage, memory call, GPU scheduling, and bandwidth usage of the inference process in the edge node device running table, identifies the corresponding sequence of resource indicators and time, and obtains the model resource usage trajectory; Usage data for various hardware and software resources is extracted from edge devices. This data may include CPU load (e.g., utilization percentage), memory usage, GPU load, and network bandwidth usage. Specifically, for an image recognition model running on an edge node, CPU load data is as follows: at the start of the task, CPU utilization is 20%, rising to 50% midway through inference, and returning to 10% at the end. Memory usage records show that during the task, memory usage increased from 200MB to 1GB, while GPU utilization also increased from 10% to 80%. This data is linked by timestamps, providing raw data support for subsequent analysis and evaluation. This monitoring process accurately tracks the dynamic changes and usage patterns of each resource, providing foundational data for identifying performance excursions. A time series of resource utilization at multiple points in time is generated, resulting in a model resource utilization trajectory that reflects changes in resource requirements during model operation.
[0021] The performance offset identification submodule calls the model resource occupancy trajectory, calculates the occupancy ratio of each resource during the task execution cycle, extracts the model response time series, compares the difference ratio between the resource share and the response time change under the same index, and generates the performance response offset degree; Extract response time series and resource usage from model execution data. For example, assume that within a task cycle, the response time is 100ms, 120ms, and 90ms, respectively, while the CPU usage during these time periods is 20%, 50%, and 30%, respectively. By calculating the resource usage ratio at each time point, the changing trend of resource usage can be determined. For example, the CPU usage of a task in the first phase of the cycle is 20%, increases to 50% in the second phase, and drops back to 30% in the third phase. Changes in the response time series are also extracted and compared, comparing the magnitude of changes in resource usage and response time at the same time points (indexes). Assume that while the CPU usage increases from 20% to 50%, the response time changes from 100ms to 120ms. During this change, the CPU usage increases by 30% and the response time increases by 20ms. By calculating the difference ratio, performance response deviation can be identified. For example, in this case, the ratio between the change in CPU usage and the change in response time can be calculated as 30% / 20ms. If the ratio exceeds the set threshold, it indicates that there is a performance deviation during this period, which is of great significance for subsequent task screening and optimization.
[0022] The abnormal task screening submodule calls the performance response deviation, compares the task deviation with the response cycle threshold item by item, screens the model tasks whose deviation exceeds the threshold, sorts them by the deviation degree, and generates an abnormal reasoning identification list; First, set a response cycle threshold. For example, a change in response time exceeding 20ms can be considered a significant performance offset. By comparing the performance response offset of each task with the response cycle threshold, screen out those tasks whose offset exceeds the threshold. Assuming that the offset of task 1 is 25ms and the offset of task 2 is 15ms, and the threshold is set to 20ms, then task 1 will be screened as an abnormal task, while task 2 will not be considered to have a significant offset. The screened abnormal tasks are sorted by the degree of offset, with tasks with larger offsets being prioritized. For example, if the offset of task 1 is 25ms and the offset of task 3 is 30ms, then task 3 will be ranked before task 1. Generate an abnormal reasoning identification list, which lists all tasks with large performance offsets, providing a basis for subsequent tuning and exception handling. Table 1: Example of resource usage data for inference tasks
[0023] As shown in Table 1, during the execution of the inference task, CPU usage increased from 20% to 50% over time, then dropped back to 30% at the end of the task. Memory usage increased from 200MB to 500MB, then dropped back to 400MB. GPU usage rose from 10% to 80%, reflecting significant changes in computing resource requirements during inference. Bandwidth usage fluctuated during the task, and response time also varied. This data allows us to further calculate the occupancy rate of each resource and the changing trend of response time, allowing for subsequent performance drift analysis.
[0024] See also Figure 3 , the load balancing scheduling module includes: The conflict identification submodule extracts the resource allocation records of restricted tasks on the same node based on the abnormal reasoning identification list, compares the task resource requirements with the node resource margin, filters out tasks whose resource requests exceed the node allocatable threshold, aggregates task numbers and resource conflict nodes, and generates a conflict task location table; Extract the original task execution records and the current node resource allocation log from the current scenario application operation and maintenance, and build a resource allocation table by combining the resource requirement fields and node resource capacity fields corresponding to each task in the scenario. The resource requirement fields specifically include the number of CPU cycles, memory consumption value and bandwidth occupancy rate, and the node resource capacity field comes from the node resource usage report periodically obtained by the current status acquisition module. For example, if the CPU requirement of task A is 3 cores and the memory requirement is 2GB, and the current remaining available resources of node N are CPU 2 cores and memory 1.5GB, then task A will be identified as a task with resource conflict. Then it is necessary to calculate whether the resource requirements of all tasks of the node exceed the current allocation range, and judge whether the three dimensions of CPU, memory and bandwidth exceed the current allocation range. There is a situation where the available capacity of the current node is exceeded. In each dimension, the warning threshold is set to 90% of the total capacity of the node. That is, if the resources required by a task exceed the remaining amount of the resource on the node × 0.9, it is marked as a conflict item. The threshold setting is based on the empirical statistics of the node's daily operation load limit. For example, a node can still run stably when 80% of the resources are occupied, while more than 90% frequently triggers performance degradation alarms. Therefore, 90% is the acceptable conflict threshold. After all conflicting task records are summarized, the conflicting node locations of tasks and resources are mapped one-to-one. For example, the resource conflict for task A on node N1 is that the CPU exceeds 1 core, and the memory for task B on node N2 exceeds 0.8GB. Finally, a conflict task location table is constructed, listing each conflicting task, the corresponding node, and the conflicting resource type.
[0025] The distribution comparison submodule calls the conflict task location table, counts the number of conflicting tasks and the total number of deployment records in the node, and identifies the difference between task density and deployment pressure using the formula: ; Calculate the deployment conflict distribution value of the current scenario node, classify the differentiated node operation loads according to the distribution value, analyze the distribution of operation conflict hotspots in the scenario, and generate a node conflict risk layer; in, Indicates the The number of deployment requests generated by conflicting tasks in the scenario, Indicates the The number of real-time dispatch paths that conflicting tasks can accept in the current cycle, Indicates the total number of conflicting tasks, Indicates the allocation conflict distribution value of the current scene node; The allocation conflict distribution value of the current scenario node is used to quantify the concentration of resource conflicts caused by the difference between task allocation density and actual allocation execution during the operation and maintenance of the artificial intelligence model. This value comprehensively measures the absolute deviation and interaction intensity between the number of tasks and the allocation record. Specifically, when the number of tasks is much higher than the number of comprehensible allocations or there are frequent allocations between tasks but the resources are insufficient to support their operation paths, the value will increase significantly, indicating that problems such as concentrated scheduling pressure, uneven resource allocation, and delayed allocation response are serious. The larger the allocation conflict distribution value, the higher the risk of node-bearing allocation instability. Conversely, it means that the current node allocation is balanced and the resource coordination status is good. This value can be used as the core judgment basis for identifying task scheduling anomalies, optimizing allocation paths, and adjusting node load balancing during operation and maintenance. Call the task numbers listed in the conflicting task location table and their node scheduling records in the operation and maintenance scenario. It is necessary to retrospectively search the original deployment of each task. The specific extraction content includes the total number of task scheduling request records and the number of deployment responses actually completed by the node. Build a mapping table between task numbers and deployment record numbers. This mapping relationship can be counted through daily records of the scheduling log system. The number of task requests is counted in units of scheduling instructions generated in the log, and the number of deployment records is counted in units of the number of events of successfully changing nodes in the task status record. After collecting the deployment information of each task group, the conflicting tasks under the same node are merged and the number of tasks of each node is counted. and the number of deployment records In order to unify the data dimensions and facilitate horizontal comparison, the two parameters need to be normalized. The normalization uses the minimum-maximum normalization formula ,in is the original value, 、 are the minimum and maximum values in the sample, respectively. The normalized results are used as parameters in subsequent calculations. Taking the actual scenario as an example, the deployment data of tasks 1 to 5 under node N1 are as follows: the number of tasks are 4, 6, 5, 7, and 8 respectively, and the number of deployment records are 10, 13, 12, 16, and 20. The normalized number of tasks is 、 、 、 、 , after normalization, the number of deployment records is 、 、 、 、 ; : After normalization Number of tasks in a task group; : After normalization Number of task group deployment records; : Number of task groups, here is 5; Substitute the data for calculation: Group 1: ; Group 2: ; Group 3: ; Group 4: ; Group 5: ; After accumulation: ; This value is the normalized dispatch conflict distribution value of node N1. By introducing the product root value of the number of dispatch records and the number of tasks, it reflects the combined influence of dispatch frequency and task density. The absolute difference is integrated to measure the dispatch balance. The numerical result of 0.5364 is much higher than the platform evaluation benchmark of 0.3, indicating that the current node scheduling has a dispatch structure load problem under task density. This result is used to generate a node conflict risk layer and identify the node as a medium-to-high conflict level node in the subsequent operation and maintenance platform graphical display.
[0026] The structure construction submodule screens nodes in high-conflict areas based on the node conflict risk layer, extracts task execution trajectories and deployment path information, classifies the deployment frequency and sequence chain of differentiated tasks when migrating between nodes, analyzes the correlation between task paths and conflict density, and obtains a conflict resource deployment structure diagram; Nodes with conflict distribution values exceeding the warning threshold are selected. The original migration paths, task allocation frequencies, and path temporal relationships of all conflicting tasks under this node are retrieved from the scheduling log. The allocation records of tasks from the start node to the target node are sorted by time to form a task execution trajectory list. The resource interaction nodes of each task during the allocation process in the trajectory are extracted and their frequencies of occurrence in the interaction path are recorded. For example, task A migrates from nodes N1→N3→N5 and is recorded as appearing twice at node N3. Task B migrates from N3→N4 and is recorded as appearing three times at N3, resulting in a conflict interaction frequency of 5 for N3. The allocation interaction frequencies of all nodes are classified and clustered. It is determined whether high-frequency nodes in the path structure are concentrated at the conflict node. If the frequency is greater than the average allocation frequency of all nodes plus the standard deviation, the node is marked as a node with high allocation conflict. A node allocation structure map is constructed. In the map, the weight of each edge represents the number of allocations between nodes, and the color of each node represents the conflict level. Finally, the task migration path and node conflict level mapping are integrated to obtain a conflict resource allocation structure map.
[0027] See also Figure 4 , the resource intelligent allocation module includes: The offset identification submodule extracts the task operation level, GPU idle rate, and resource offset rate based on the conflict resource allocation structure diagram. Based on the correspondence between task level and offset rate, it marks the task mapping with abnormal resource fluctuations and generates task offset mapping records. Identify the associated boundary between the GPU resource node and the task running status in the structure diagram. In the specific implementation, the model instance number and the scheduling task number mapped to each GPU node can be recorded through the connection matrix in the resource allocation structure diagram, and the running level of each scheduling task in the task control table of the inference scheduling area can be parsed. The level is divided according to the dimensions of computing intensity, data throughput, etc., and is set to a level range of 1 to 5. Level 5 represents a high computing density task. For example, calling the ResNet-152 model for image recognition tasks in a certain scenario is defined as level 5. Read the GPU partition mapping table and extract the current idle rate of each partition. The idle rate is obtained by the ratio of the current available video memory of the GPU to the total video memory. For example, the total video memory of a GPU node is 16GB, and 12GB is currently available, then the idle rate is 75%. Then, according to the model The resource offset rate of models at different levels is extracted from the computing level classification table. The offset rate is based on the average offset value of computing resource usage extracted from the original task operation record. For example, for the level 4 BERT model, its resource offset rate is 0.27, indicating that its resource demand deviates from the configuration benchmark by 27% during operation. A one-to-one correspondence is established among the extracted operation level, idle rate, and offset rate. Then, an offset analysis is performed to compare the resource offset rate with the task operation level. Tasks with higher task levels and abnormally increased offset rates are marked. For example, if the offset rate is higher than 25%, it is set as an abnormality. In practice, if the offset rate of a task with an operation level of 5 is as high as 32%, it is marked as a resource abnormality task. Finally, the task number, operation level, idle rate, offset rate and other items are aggregated to generate a task offset mapping record.
[0028] The node adaptation submodule calls the task offset mapping record, matches the task running level value with the GPU partition idle rate, and compares the adaptation gap between the idle rate and the resource offset rate using the formula: ; Calculate the node adaptation offset value, combine it with the partition mapping table to filter the running nodes that meet the migration conditions, and establish a running node adaptation list; in, Indicates the node adaptation offset value, is the task level value, Indicates the Node idle rate, Indicates the Node resource offset rate, Indicates the Node task running load, Indicates the total number of nodes currently participating in task allocation; The task running level is adapted and matched with the idle rate and resource offset rate of each GPU node. First, the task running level parameters are determined. Set the current task level to This level is an evaluation value of the model operation intensity based on the task control table of the inference scheduling area. It uses a level system of 1 to 5. The larger the value, the higher the model operation load demand. In the AI scenario, level 4 can correspond to scenarios such as BERT-base for complex natural language understanding. The current idle rate of each node is obtained from the GPU partition mapping table. The idle rate is obtained by dividing the idle memory of the GPU node by its total memory in real time. To ensure the uniformity of the values, all values are normalized to the range of 0-1. For example, the current idle memory of GPU1 is 9.6GB and the total memory is 12GB. The idle rate is calculated as , GPU2 idle rate is , GPU3 idle rate is ; Then extract the model to calculate the resource offset rate of each node in the hierarchical classification table , which is the offset of the resource benchmark after the current task is mapped to the node. The average fluctuation rate is calculated from the GPU load time series of the original task and normalized. For example, the offset rate of GPU1 is , GPU2 is , GPU3 is , task load Indicates the average computing cycle ratio of the task on the node GPU per unit time, which is defined as the proportion of the task in the original computing cycle of the node and normalized. For example, GPU1 is , GPU2 is , GPU3 is , according to the above data, substitute the adaptation offset formula for calculation; The first step is to calculate the deviation and (numerator) : ; Multiply by the task level : ; Step 2: Calculate the square of the denominator and sum it up : ; ; ; ; The third step is to calculate the final adaptation offset value: ; : Task running level (1-5), in this case it is 4; : No. The idle rate of each node is calculated by the ratio of GPU available video memory to total video memory, normalized; : No. The resource offset rate of each node is the normalized result of the average value of the original GPU load fluctuation; : No. The task load of each node is the ratio of tasks to GPU load per unit time. : The number of GPU nodes involved in scheduling matching, 3 in the example; By introducing the resource offset rate and running load into the normalization calculation, the adaptation relationship between the task level and the multi-node resource status is more carefully reflected as a quantifiable value, which is conducive to subsequent sorting and screening. The result shows that the current node adaptation offset is 3.616. If the preset offset upper limit is 5.0, it means that the node is within the offset control range and meets the scheduling migration conditions, and is then selected as an available running node and included in the running node adaptation list.
[0029] The path adjustment submodule calls the running node adaptation list, plans task migration according to the task running level, identifies the mapping sequence based on the relationship between node idle rate and level, and marks the migration path to obtain the model inference migration path; First, extract the node number and its idle rate value in the list, then extract the running level of the task to be migrated from the task offset mapping record, and compare them with the idle rates of each node in the list. For example, if the current migration task level is 5, you need to select a node with an idle rate higher than 70% for transfer. In the actual maintenance scenario, if the idle rate of node B is 72% and that of node D is 77%, both are migratable target nodes. Combined with the fitness value, priority is given to migrating to the node with the lowest fitness. During the migration process, it is necessary to record the original running node number, target node number, task level, and idle rate change value of the task, and obtain the model reasoning migration path in sequence. At the same time, identify the migration landing point number, jump order, and node remaining resource change interval of each step, and finally construct the reasoning task running migration path table for the scheduling control center to call in sequence.
[0030] See also Figure 5 , the deployment configuration adjustment module includes: The path judgment submodule uses the model to infer the migration path. Based on the target node's storage call frequency, memory margin, and network latency, it identifies the synchronous change range of node indicators between the current and previous weeks, and makes a difference judgment with the load fluctuation amplitude. It identifies the stable path and records the flag information to generate a set of stable model inference paths. The target node's hourly storage call count over the past seven days is used as raw input. The log fields are cleaned using the structured log collection module, and the storage operation type and access duration recorded in the fields are extracted. The average number of accesses per unit time is calculated to obtain the call frequency value for each node. For example, if node N1 records 55 storage calls in the first hour and 48 in the second hour, the daily average frequency is 51.5. The memory margin is then obtained by using a real-time monitoring tool to read the difference between the currently allocated memory and the used memory of each node, and converting it into a margin value in GB. For example, if node N1 has a total memory of 32GB and 26GB is currently used, the margin is 6GB. The ping return delay data collected from the layer is integrated into the average network round-trip time at a sampling frequency of 10 minutes. If the average round-trip time for node N1 is 45ms every 10 minutes, its T value is recorded as 45. The above three types of data are grouped according to nodes, and the change trends of the current cycle and the previous cycle are compared. The difference calculation method is used to obtain the change range of each node indicator. For example, the frequency in the current cycle is 52, and the frequency in the previous cycle is 50, with a difference of 2. The three types of difference data are set as the change vector V. The maximum and minimum difference between the load values recorded by the node in the original 10 cycles is used as the reference value of the node load fluctuation amplitude. If the load record interval of node N1 is between 0.42 and 0.63, the amplitude is 0.21. If the change value of any indicator in V is less than 0.21, the indicator can be regarded as a stable interval. The node whose three types of indicators all meet the requirements of being less than the fluctuation amplitude is marked as a stable node, and its path information is registered in the identification set to form a matching mapping between the model operation path and the node status, and obtain the model reasoning stable path set.
[0031] The task screening submodule calls the model to infer the stable path set, combining the task instance call frequency, resource allocation value and node load reconstruction index, and adopts the formula: ; Calculate the task deployment interference degree, filter out task instances with interference degrees below the control threshold, associate them with the target path, and generate a low-interference deployment candidate set; in, represents the task deployment interference degree, Represents the current calling frequency of the task, Represents the original path call frequency, Represents the resource allocation value, Represents the memory remaining. represents the load reconstruction degree, represents the migration impact; Task deployment interference refers to the comprehensive reflection value of multiple factors such as resource consumption changes, node scheduling structure reconstruction, and communication link fluctuations caused by migrating a task instance to the target node during the operation and maintenance of the artificial intelligence model. The higher the value, the more significant the impact of the task migration on the overall operation stability, causing problems such as node load imbalance and increased inference response latency. Conversely, the lower the value, the less interference the task deployment process has on the existing structure, and the task migration can be completed without reconstructing the node resource structure or increasing the communication load. Therefore, task deployment interference is a core indicator for quantitatively judging task migration risks and deployment adaptability. Combined with the key performance data of task instances, the task deployment interference degree is calculated and screened. First, the call frequency of task instances in the current cycle is obtained. By counting the task processing requests, for example, the call frequencies of Task 1, Task 2, and Task 3 are 95 times / second, 70 times / second, and 45 times / second, respectively. This value is directly extracted from the access frequency data of the model service in the log, and its unit is "times / second". Then, the call frequency of the original cycle path is obtained to represent the stability index of the corresponding path of the task before migration. For example, the call frequencies of the original paths of Task 1, Task 2, and Task 3 are 100 times / second, 65 times / second, and 60 times / second, respectively, and the unit is "times / second". This is obtained by calling the log data statistics of the previous cycle. Next, obtain the resource allocation value of the task instance. This value refers to the proportion of the task's total node resources occupied, which is mainly calculated from the CPU and memory usage recorded by the scheduling system. For example, the resource allocation values of Task 1, Task 2, and Task 3 are 0.75, 0.60, and 0.45, respectively, which have been normalized to the range of 0-1. At the same time, obtain the current memory margin of the node in GB, which needs to be normalized for unified calculation. For example, the remaining memory of nodes A, B, and C is 32GB, 24GB, and 16GB, respectively. The maximum node memory is set to 64GB, and the normalized values are 0.5, 0.375, and 0.25, respectively. Obtain the node load reconstruction degree. This value is the normalized CPU load fluctuation value. The collection method is to monitor the CPU load changes in the past five minutes and calculate the ratio of its fluctuation range to the set benchmark range. For example, the values of nodes A, B, and C are 0.6, 0.45, and 0.35, respectively. Finally, the migration impact is obtained, which is defined as the disturbance value of task migration on the balance of existing path call frequency. It is calculated by normalizing the path call offset value caused by task migration. For example, the values for tasks 1, 2, and 3 are 0.3, 0.25, and 0.2 respectively. Taking Task 2 as an example, , , , , , , substituting into the formula we get: The first step is to calculate the absolute value of the call frequency difference: ; The second step is to calculate the square root of the product of the resource allocation value and the memory margin: ; The third step is to add the denominators: ; Step 4: Overall calculation: ; The result of the task deployment interference is 7.82. If the deployment interference control threshold is set to 8.0, then task 2 is lower than the threshold and is a deployable task. It is added to the low-interference deployment candidate set, and finally the low-interference deployment candidate set is generated, where: Indicates the current call frequency of the task, in times / second. Indicates the original path call frequency, the unit is times / second. is the resource allocation value, dimensionless, calculated by normalizing the task resource ratio and the maximum node resource. It is the normalized value of memory remaining, in GB, and is normalized to 0-1 dimension. is the load reconstruction degree, dimensionless, obtained by normalizing the node load fluctuation, is the migration impact, dimensionless, obtained by normalizing the path disturbance value; By jointly modeling task load changes and node resource status, and integrating resource allocation and migration disturbance characteristics, a comprehensive evaluation of task deployment disturbance is achieved at a unified scale. Task 2 with an interference degree of 7.82 can be regarded as a low-interference task, which is suitable for model dynamic scheduling and deployment in scenario operation and maintenance.
[0032] The mapping identification submodule integrates the task deployment identifier and stable path relationship information based on the low-interference deployment candidate set, updates the deployment status field and registers the migration configuration, maps the deployment mapping relationship between the model inference node and the task, and generates a model configuration deployment mapping table; Integrate the task deployment identifier and stable path relationship information. First, obtain the deployment identifier of each task instance, for example, the deployment identifier of task 2 is T2, and that of task 3 is T3. Then obtain the corresponding stable path information, for example, task 2 corresponds to node B, and task 3 corresponds to node C. Integrate the task deployment identifier and path information to form a deployment mapping relationship, for example, T2→node B, T3→node C. Update the deployment status field, update the deployment status of the task instance to "deployed", and register the migration configuration. For example, add a record in the migration configuration table: Task 2 is deployed to node B, and Task 3 is deployed to node C. Map the deployment mapping relationship between the model inference node and the task, and generate a scenario operation model deployment mapping table, which includes information such as task identifier, deployment node, and deployment status.
[0033] See also Figure 6 , the scene feedback integration module includes: The log parsing submodule collects scenario control terminal task logs and model interface records based on the model configuration deployment mapping table, compares timestamps based on task numbers, identifies task response intervals and interface call delays, and generates model call timing features; The scene control terminal task logs and model interface records are collected. The collected data includes information such as task number, task execution time, task response time, and interface call timestamp. By comparing the task number and timestamp, the task response interval and interface call delay are identified, and the model call timing characteristics are calculated. The task response interval is calculated by subtracting the timestamp of each record in the task log, and the interface call delay is determined by subtracting the interface recording time from the actual task execution time. A time series data is then generated for each task, forming the task's timing characteristics. The characteristics reflect the task's response performance during execution and can be used to further analyze the task's stability and efficiency. For example, if the task response time for task number 001 is 50ms and the corresponding interface call time is 20ms, then the interface delay is 30ms. If this delay exceeds the set threshold, it will mark a potential performance issue, providing data support for subsequent optimization. Generating model call timing characteristics can provide basic data for subsequent task status evaluation and offset identification.
[0034] The response status evaluation submodule is based on the model call timing characteristics and the task number in the dynamic task registry. It identifies the response duration and execution status fields, determines whether the response interval and duration offset exceed the set threshold, and combines the execution status to filter the task fluctuation frequency and generate a model operation stable state set. Based on the task number in the dynamic task registry, the response duration and execution status fields of each task are identified. Further, it is determined whether the response interval and duration offset exceed the set threshold. By comparing the task response time in the timing characteristics with the predetermined maximum response duration (such as 100ms), if the actual response duration exceeds the threshold, it is marked as an abnormal task. The execution status field is also used to filter the execution status of the task. If a task experiences abnormal state fluctuations during execution, the frequency of such fluctuations is recorded. For example, if the response time of a task numbered 002 exceeds 100ms, and the execution status field shows that the task has entered the "failed" state multiple times, the task will be marked as abnormal and require further analysis and optimization. A model operation stable state set is generated, which can monitor the stability of the task in real time and evaluate whether the task is in a stable operation state based on the frequency of state fluctuations.
[0035] The task deviation identification submodule selects state samples of urban visual recognition and industrial equipment health recognition tasks based on the model's stable operation state set. It counts the number of consecutive state fluctuations under the task number and compares the call frequency. It identifies tasks with decreased stability and a frequency exceeding the baseline value, and obtains an application scenario operation and maintenance list. The analysis is based on the model's stable state set, which contains the execution status and fluctuations of tasks. During execution, representative task types are selected: urban visual recognition and industrial equipment health identification. These tasks are chosen based on their common and important nature in real-world applications. Each task is numbered and the number of consecutive state fluctuations is counted. Continuous state fluctuation refers to the number of times the task's state changes during execution. More fluctuations indicate unstable task execution. For example, if a task experiences five state fluctuations during execution, its consecutive state fluctuation count is 5. The task's call frequency is then compared, and tasks with higher frequency are monitored more closely. If a task's state fluctuation frequency exceeds a set baseline (for example, 3) and its call frequency is high, its stability is judged to have degraded and it is ultimately added to the application scenario's operation and maintenance checklist. The key to this process is the dual assessment of fluctuation count and call frequency. Only when both indicators reach a certain threshold is a task considered unstable and warrants attention.
[0036] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A scenario application operation and maintenance platform based on artificial intelligence models, characterized by: The platform includes: The inference state perception module obtains the CPU usage, memory call, GPU scheduling, and bandwidth usage of the inference process in the edge node device's running table. It makes judgments based on the time ratio of each resource consumption and the offset trend of the task response time, filters out tasks whose fluctuation values exceed the response cycle threshold, and generates an abnormal inference identification list. The load balancing scheduling module extracts the resource allocation status of the restricted tasks on the same node based on the abnormal reasoning identification list, filters the resource conflict task group, identifies the task conflict distribution by the ratio of the number of tasks to the number of allocation records, and obtains the conflict resource allocation structure diagram; The resource intelligent allocation module extracts the task operation level and resource offset rate according to the conflict resource allocation structure diagram, and migrates the tasks with excessive offset rates to nodes with sufficient idle rates according to the GPU idle rate mapping record, thereby obtaining the model reasoning migration path; The deployment configuration adjustment module calls the model inference migration path, screens task instances whose synchronization change interval is lower than the load fluctuation range based on the storage call frequency, memory margin and network latency of the target node, writes the deployment identifier into the migration configuration, and generates a model configuration deployment mapping table.
2. The scenario application operation and maintenance platform based on the artificial intelligence model according to claim 1 is characterized in that: The abnormal reasoning identification list includes task identification information, resource abnormality type, response offset amplitude, and fluctuation period label; the conflict resource allocation structure diagram includes resource conflict nodes, allocation conflict relationship, task conflict intensity, and node resource utilization; the model reasoning migration path includes migration node sequence, resource offset level, task priority label, and GPU idle rate matching result; the model configuration deployment mapping table includes target node parameters, deployment task index, configuration matching index, and load fluctuation tolerance range.
3. The scenario application operation and maintenance platform based on artificial intelligence model according to claim 1 is characterized in that: The reasoning state perception module includes: The resource monitoring submodule obtains the CPU usage, memory call, GPU scheduling, and bandwidth usage of the inference process in the edge node device running table, identifies the corresponding sequence of resource indicators and time, and obtains the model resource usage trajectory; The performance offset identification submodule calls the model resource occupancy trajectory, calculates the occupancy ratio of each resource during the task execution cycle, extracts the model response time series, compares the difference ratio between the resource share and the response time change amplitude under the same index, and generates the performance response offset degree; The abnormal task screening submodule calls the performance response deviation, compares the task deviation with the response period threshold item by item, screens the model tasks whose deviation exceeds the threshold, sorts them by the deviation degree, and generates an abnormal reasoning identification list.
4. The scenario application operation and maintenance platform based on the artificial intelligence model according to claim 3 is characterized in that: The load balancing scheduling module includes: The conflict identification submodule extracts the resource allocation records of the restricted tasks on the same node based on the abnormal reasoning identification list, compares the task resource requirements with the node resource margin, filters out tasks whose resource requests exceed the node allocatable threshold, aggregates task numbers and resource conflict nodes, and generates a conflict task location table; The distribution comparison submodule calls the conflict task location table, counts the number of conflicting tasks and the total number of deployment records in the node, identifies the difference between task density and deployment pressure, calculates the deployment conflict distribution value of the current scenario node, hierarchically classifies the differentiated node operation loads according to the distribution value, analyzes the distribution of operation conflict hotspots in the scenario, and generates a node conflict risk layer; The structure construction submodule screens nodes in high-conflict areas based on the node conflict risk layer, extracts task execution trajectories and allocation path information, classifies the allocation frequency and sequence chain of differentiated tasks when migrating between nodes, analyzes the correlation between task paths and conflict density, and obtains a conflict resource allocation structure diagram.
5. The scenario application operation and maintenance platform based on the artificial intelligence model according to claim 4 is characterized in that: The resource intelligent allocation module includes: The offset identification submodule extracts the task operation level, GPU idle rate and resource offset rate according to the conflict resource allocation structure diagram, marks the task mapping with abnormal resource fluctuation according to the corresponding relationship between the task level and the offset rate, and generates a task offset mapping record; The node adaptation submodule calls the task offset mapping record, matches the task running level value with the GPU partition idle rate, compares the adaptation gap between the idle rate and the resource offset rate, calculates the node adaptation offset value, and screens the running nodes that meet the migration conditions in combination with the partition mapping table to establish a running node adaptation list; The path adjustment submodule calls the running node adaptation list, plans task migration according to the task running level, identifies the mapping sequence according to the node idle rate and level relationship, and marks the migration path to obtain the model reasoning migration path.
6. The scenario application operation and maintenance platform based on the artificial intelligence model according to claim 5 is characterized in that: The deployment configuration adjustment module includes: The path judgment submodule calls the model inference migration path, identifies the synchronous change range of node indicators between the current and previous weeks based on the storage call frequency, memory margin, and network latency of the target node, and makes a difference judgment with the load fluctuation amplitude, identifies the stable path and records the flag information, and generates a model inference stable path set; The task screening submodule calls the model to infer the stable path set, combines the task instance call frequency, resource allocation value and node load reconstruction index, calculates the task deployment interference degree, screens out task instances with interference degrees below the control threshold, associates them with the target path, and generates a low-interference deployment candidate set; The mapping identification submodule integrates the task deployment identifier and the stable path relationship information according to the low-interference deployment candidate set, updates the deployment status field and registers the migration configuration, maps the deployment mapping relationship between the model inference node and the task, and generates a model configuration deployment mapping table.
7. The scenario application operation and maintenance platform based on artificial intelligence model according to claim 1 is characterized in that: The platform also includes a scene feedback integration module: Based on the model configuration deployment mapping table, the scenario feedback integration module extracts the scenario control terminal task response log and model interface call records, identifies the inference response duration, execution status, and call frequency in the dynamic task registry, and screens representative tasks with execution interval drift for typical applications such as urban visual recognition and industrial equipment health recognition to obtain an application scenario operation and maintenance list. The application scenario operation and maintenance list includes task execution frequency, response time change, execution status identification, and representative task records.
8. The scenario application operation and maintenance platform based on artificial intelligence model according to claim 7 is characterized in that: The scene feedback integration module includes: The log parsing submodule collects the scenario control terminal task logs and model interface records based on the model configuration deployment mapping table, compares the timestamps according to the task numbers, identifies the task response intervals and interface call delays, and generates model call timing features; The response status evaluation submodule identifies the response duration and execution status fields based on the model call timing characteristics and the task number in the dynamic task registry, determines whether the response interval and duration offset exceed the set threshold, and filters the task fluctuation frequency in combination with the execution status to generate a model operation stable state set; The task offset identification submodule selects state samples of urban visual recognition and industrial equipment health identification task types according to the stable state set of the model operation, counts the number of continuous state fluctuations under the task number, and compares the call frequency to identify tasks with decreased stability and a frequency exceeding the benchmark value, and obtains the application scenario operation and maintenance list.
Citation Information
Cited By
Entertainment equipment remote operation and maintenance system and method based on Internet of Things and SaaS platform
CN120935178A
Production process state monitoring scheduling optimization method based on real-time data acquisition
CN120952466A
Data center data management method based on artificial intelligence
CN121255476A
Server load balancing method and system based on edge computing
CN121301008A
Request processing method and device based on large model, electronic equipment and storage medium
CN121328705A