A multi-class input-oriented intelligent reasoning task scheduling method

CN122819508APending Publication Date: 2026-09-25SHAANXI ZHIANXUN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611295622.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

具体解决现有智能推理任务调度过程中模型适配结果与算力资源状态分阶段确定、二者之间的执行衔接性较弱,使得在算力状态变化或者适配模型未驻留时容易产生额外的模型加载、任务等待或重新调度,进而影响推理任务处理效率和执行连续性的问题

Benefits of technology

[0053]1、本发明基于对现有技术问题的进一步分析和研究,认识到模型适配结果与算力资源状态分阶段确定时,适配得到的推理模型可能因模型加载状态、实例占用状态或算力余量变化而无法直接执行,进而增加模型加载、任务等待或重新调度;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819508A_ABST
    Figure CN122819508A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computing resource scheduling, in particular to a kind of intelligent inference task scheduling method for multiple types of input, in the present application, task model adaptation range is formed by input task state, and candidate model power pairing is formed by combining model residence relationship, power load range, and target model power pairing is further determined;Inference execution configuration and inference processing timing are formed according to target model power pairing, input data processing is completed and result output routing is formed, so that model adaptation result and power resource state form continuous connection, reduce additional model loading, task waiting or rescheduling, improve inference task processing efficiency and execution continuity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computing resource scheduling technology, specifically to an intelligent reasoning task scheduling method oriented towards multiple types of input. Background Technology

[0002] With the application of intelligent inference technology in image analysis, video processing, and other data processing scenarios, inference systems typically need to determine the appropriate inference model based on the type and specifications of the input data and the task to be executed. They then utilize one or more computing resources, such as CPU, GPU, NPU, and FPGA, to complete the model inference. Since different inference models differ in input conditions, computational overhead, storage usage, and operating environment, and the resource usage of each computing device dynamically changes during task execution, model selection and computational resource scheduling significantly impact the processing efficiency and execution continuity of inference tasks.

[0003] Existing intelligent inference systems typically parse information such as the type, size, frame rate, or task instructions of the input data after receiving it, and select an inference model suitable for the current task from a pre-set set of models accordingly. Then, they allocate appropriate computing resources to the inference task based on the resource utilization of the computing device. When the selected inference model is already loaded onto the corresponding computing device, the existing model instance can be directly invoked to execute the inference; when the inference model has not yet been loaded, model loading and inference environment initialization must be completed before executing the inference task. Some systems also adjust the task distribution location based on the real-time load of the computing device and transmit the processing results according to a pre-configured output method after inference is completed.

[0004] However, in the above processing methods, model selection and computing resource allocation are typically performed at different processing stages based on task requirements and device operating status. When device load, available storage space, or model instance occupancy changes, a mismatch may occur between the previously selected model and subsequent available computing resources. This necessitates additional processing during scheduling, such as model loading, task waiting, or reselection of execution location. Especially when multiple inference tasks arrive consecutively or the computing resource status continuously changes, the model selection result and the actual execution resources are prone to repeated adjustments, increasing the processing overhead of task scheduling and model switching, and affecting the response efficiency and processing continuity of inference tasks. Summary of the Invention

[0005] The purpose of this invention is to provide an intelligent inference task scheduling method for multiple types of inputs, in order to solve the problems mentioned in the background art. Specifically, it addresses the problem that in the existing intelligent inference task scheduling process, the model adaptation result and the computing power resource status are determined in stages, and the execution connection between the two is weak. This makes it easy for additional model loading, task waiting, or rescheduling to occur when the computing power status changes or the adapted model is not resident, thereby affecting the processing efficiency and execution continuity of the inference task.

[0006] To achieve the above objectives, the present invention aims to provide an intelligent reasoning task scheduling method for multiple types of inputs, specifically including the following method steps:

[0007] S1. Obtain the task attributes and output destination of the input data to form the input task status; form the task model adaptation range from the model pool based on the input task status.

[0008] The process of forming the input task state specifically includes:

[0009] The task attributes for obtaining input data and the output destination are as follows: task attributes include data organization method, input size, input frame rate, inference task type, task priority, and allowed processing latency;

[0010] The task identifier, timestamp, image size, pixel encoding, and data length of the input data are validated.

[0011] The processing load is determined based on the input size, equivalent frame rate, and reference input, and the input task state consists of task identifier, data organization method, input size, input frame rate, processing load, inference task type, task priority, allowable processing latency, and output destination.

[0012] The process of forming the task model's adaptation range specifically includes:

[0013] Retrieve inference models and model configurations from the model pool;

[0014] Based on the inference task type, data organization method, input size, allowable processing latency, and minimum model accuracy level in the input task status, the inference model is adapted.

[0015] The expected processing time is determined based on the baseline processing time and processing load, and the set of inference models that meet the requirements of inference task type, data organization method, input size, expected processing time and model accuracy level is determined as the task model adaptation range.

[0016] Step S1 forms the input task status by considering task attributes, output destination, and processing load. Based on this, it selects inference models by combining task type, data organization method, input size, allowable processing latency, and model accuracy level, forming the task model adaptation range. This establishes a correspondence between the model adaptation results and the current input task requirements, providing a foundation for determining the executable model in conjunction with the computing power resource status.

[0017] S2. Obtain the model loading status, instance occupancy status, and computing power reserve of each computing power unit to form a model residency relationship; combine the processing load corresponding to the input task status and the model resource requirements within the task model adaptation range to form a computing power carrying range; pair the inference models within the task model adaptation range with computing power units that meet the requirements of computing power architecture, model version, and running status to form candidate model computing power pairings.

[0018] The formation process of model residency relationships specifically includes:

[0019] Obtain the model loading status, instance occupancy status, computing architecture, model version, remaining concurrency, sampling time, and status version for each computing unit;

[0020] Establish a callable residency relationship between the inference model and the computing power unit that has loaded the inference model and has remaining concurrency; establish an occupied residency relationship between the inference model and the computing power unit that has loaded the inference model and has no remaining concurrency; establish a deployable relationship between the inference model and the computing power unit that has not loaded the inference model but whose computing power architecture meets the model configuration.

[0021] The model residency relationship consists of callable residency relationships, occupied residency relationships, and deployable residency relationships.

[0022] The formation process of computing power carrying capacity specifically includes:

[0023] Model resource requirements include model computation requirements and model storage requirements;

[0024] The model calculation requirements are determined based on the processing load and the model calculation requirement coefficient, and the model storage requirements are determined based on the processing load, the model working storage coefficient, the model static storage requirements, and the model loading status.

[0025] Obtain the normalized available computing power, available storage power, and the number of new instances allowed for each computing power unit. Include combinations of inference models and computing power units whose model computing requirements do not exceed the normalized available computing power, whose model storage requirements do not exceed the available storage power, and which meet the instance creation requirements, within the computing power carrying capacity.

[0026] The process of forming candidate model computing power pairings specifically includes:

[0027] Read the valid computing power units in the inference model and status data within the task model adaptation range, enumerate the combinations in ascending order of model identifier and computing power unit identifier, and prevent combinations that do not meet the requirements of computing power architecture, model version or running status from forming candidate model computing power pairings.

[0028] The effective combination of model loading status, instance occupancy status, remaining concurrency, model computation requirements, model storage requirements, computing power margin, status version, and sampling time is written into the same pairing relationship to form candidate model computing power pairing.

[0029] Step S2 establishes model residency relationships based on model loading status, instance occupancy status, remaining concurrency, and computing power architecture. It also forms a computing power capacity range by combining processing load, model computation requirements, and model storage requirements. The adapted model is then paired with effective computing power units to form candidate model computing power pairings, so that model adaptation results, model residency status, and computing power capacity are uniformly associated, reducing the disconnect between model selection and computing power allocation.

[0030] S3. Based on the model residency relationship and computing power capacity, divide the candidate model computing power pairings into resided executable pairs and non-resided executable pairs; determine the target model computing power pairing from the resided executable pairs based on the computing power reserve; if there is no resided executable pairing, determine the target model computing power pairing from the non-resided executable pairs based on the computing power reserve.

[0031] The process of determining the target model's computing power pairing specifically includes:

[0032] Candidate model computing power pairings that have a callable residency relationship and are within the computing power capacity range are classified as resided executable pairings, while candidate model computing power pairings that have a deployable relationship and are within the computing power capacity range are classified as non-resided executable pairings.

[0033] For candidate model computing power pairing, normalized available computing power, available storage power, model computing requirements, and model storage requirements are obtained. The ratio of the normalized available computing power after deducting the model computing requirements to the normalized available computing power, and the ratio of the available storage power after deducting the model storage requirements to the available storage power, are taken as the smaller of the two as the margin ratio.

[0034] When the resident executable pairing is not empty, the target model computing power pairing is determined from the resident executable pairings in descending order of the remaining capacity ratio; when the resident executable pairing is empty, the target model computing power pairing is determined from the non-resident executable pairings in descending order of the remaining capacity ratio.

[0035] Step S3 distinguishes between resident executable pairs and non-resident executable pairs by model residency relationship and computing power capacity range, and determines the target model computing power pair from the resident executable pairs first based on computing power reserve. If there is no resident executable pair, the target model computing power pair is determined from the non-resident executable pairs. This takes into account both the model callability status and computing power reserve, and reduces the possibility of additional model loading, task waiting or rescheduling.

[0036] S4. Allocate computing resources according to the target model's computing power pairing; when the target inference model has been loaded, call the corresponding model instance; when it has not been loaded, load the target inference model into the target computing power unit and then call the corresponding model instance to form the inference execution configuration; form the inference processing sequence according to the input task status and the inference execution configuration.

[0037] The process of forming the inference execution configuration specifically includes:

[0038] Based on the target model's computing power pairing, atomic reservations are made for the computing and storage resources of the target computing power unit, and the computing power margin update is submitted after both computing and storage resources are successfully reserved.

[0039] When the target inference model has been loaded into the target computing unit, the model version, instance occupancy status and remaining concurrency are verified before the corresponding model instance is called; when the target inference model has not been loaded into the target computing unit, the model file and model configuration are verified, and the model weight loading, inference context establishment, model instance creation and instance calling are completed in sequence.

[0040] The inference execution configuration includes the target inference model, target computing power unit, model instance identifier, computing resource quota, storage resource quota, input data transmission method, inference batch, and number of execution streams.

[0041] The formation process of the reasoning processing sequence specifically includes:

[0042] Read the data organization method, input frame rate, and allowed processing latency from the input task status, and read the number of execution streams from the inference execution configuration;

[0043] In single-frame mode, the single-frame inference processing sequence is formed in the order of input verification, preprocessing, data transmission, model inference, postprocessing, and result encapsulation.

[0044] In continuous frame mode, a circular frame buffer is established, and a pipelined processing relationship is formed in the order of frame reading, preprocessing, data transmission, model inference and postprocessing. The length of the circular frame buffer is determined based on the input frame rate, allowable processing latency and buffer capacity, and the preprocessing, model inference and postprocessing are overlapped and advanced according to the number of execution streams.

[0045] Step S4 completes the reservation of computing and storage resources by matching the target model computing power, and directly calls the model instance or calls it after the model is loaded according to the model loading status to form the inference execution configuration; then, it combines the input task status to form the inference processing sequence, so that the model selection, computing power resource allocation and actual inference execution form a continuous correspondence, reducing the impact of changes in computing power status on task execution connection.

[0046] S5. Process the input data according to the reasoning processing sequence to obtain the reasoning result; form the result output route according to the correspondence between the reasoning result and the output destination, and distribute the reasoning result according to the result output route.

[0047] The specific process of generating the output route includes:

[0048] The model instances in the inference execution configuration are invoked according to the inference processing sequence to process the input data and obtain the inference results.

[0049] The result output route is formed based on the correspondence between the inference result type, task identifier, output destination type, output destination address and interface version. When multiple output destinations are configured, distribution paths are formed respectively and they all refer to the same inference result.

[0050] The output routing distribution inference results are used to exclude the corresponding distribution path when the address verification fails, and the local temporary storage area is used as the alternative output destination when all output destinations are invalid.

[0051] Step S5 processes the input data through the inference processing sequence and forms the result output route based on the correspondence between the inference result and the output destination. This allows the aforementioned model adaptation, computing power scheduling, and inference execution results to continue to the result distribution stage, thereby maintaining the continuity of the task processing link and reducing additional matching processing in the output stage.

[0052] Compared with the prior art, the beneficial effects of the present invention are:

[0053] 1. Based on further analysis and research of existing technical problems, this invention recognizes that when the model adaptation result and computing power resource status are determined in stages, the adapted inference model may not be able to be executed directly due to changes in model loading status, instance occupancy status or computing power reserve, thereby increasing model loading, task waiting or rescheduling.

[0054] To this end, the input task state is formed by the task attributes and the output destination, and the task model adaptation range is formed based on the input task state. Further, by combining the model residency relationship and the computing power capacity range, the inference models and computing power units within the task model adaptation range are paired with candidate model computing power, and the resident executable pairings and non-resident executable pairings are distinguished to determine the target model computing power pairing. Then, the inference execution configuration is formed based on the target model computing power pairing, and the inference processing sequence is formed by combining the input task state. Finally, the inference is completed according to the inference processing sequence and the result output route is formed.

[0055] This allows the model adaptation results to gradually converge into an executable combination of model and computing power unit as the model resides and computing power capacity increases, and further implement it in the actual inference execution and result distribution process. This helps to reduce additional model loading, task waiting or rescheduling, and improve the processing efficiency and execution continuity of intelligent inference tasks.

[0056] 2. After determining the target model computing power pairing, the present invention performs atomic reservation of computing resources and storage resources of the target computing power unit, and submits the computing power margin update after the computing resources and storage resources are successfully reserved at the same time. Then, it calls the corresponding model instance according to the model loading status of the target inference model or calls the corresponding model instance after the model loading is completed.

[0057] Therefore, it is possible to maintain a good connection between the computing resource status on which the target model computing power pairing is determined and the subsequent inference execution configuration, reducing the possibility that the target model computing power pairing cannot be executed according to the predetermined configuration due to changes in resource status, and further reducing the impact of task waiting or rescheduling on the continuity of inference execution. Attached Figure Description

[0058] Figure 1 This is a schematic diagram of the overall method steps of the present invention;

[0059] Figure 2 This is a schematic diagram of the core process of the overall method of the present invention;

[0060] Figure 3 This is a schematic diagram of the core process of inputting task status and task model adaptation range in step S1 of the present invention.

[0061] Figure 4 This is a schematic diagram of the core process of forming the model dwell relationship, computing power capacity range and candidate model computing power pairing in step S2 of the present invention;

[0062] Figure 5 This is a schematic diagram of the core process for determining the target model computing power pairing in step S3 of the present invention;

[0063] Figure 6This is a schematic diagram of the core process of inference execution configuration and inference processing timing formation in step S4 of the present invention;

[0064] Figure 7 This is a schematic diagram of the core process of reasoning result processing and result output routing in step S5 of the present invention. Detailed Implementation

[0065] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0066] Next, please refer to Figure 1-2 The purpose of this embodiment is to provide an intelligent reasoning task scheduling method for multiple types of input, including steps S1 to S5, wherein:

[0067] S1. Obtain the task attributes and output destination of the input data to form the input task status; based on the input task status, form the task model adaptation range from the model pool.

[0068] S2. Obtain the model loading status, instance occupancy status, and computing power reserve of each computing power unit to form a model residency relationship; combine the processing load corresponding to the input task status and the model resource requirements within the task model adaptation range to form a computing power carrying range; pair the inference models within the task model adaptation range with computing power units that meet the requirements of computing power architecture, model version, and running status to form candidate model computing power pairings.

[0069] S3. Based on the model residency relationship and computing power capacity, divide the candidate model computing power pairings into resided executable pairings and non-resided executable pairings; determine the target model computing power pairing from the resided executable pairings based on the computing power reserve; if there is no resided executable pairing, determine the target model computing power pairing from the non-resided executable pairings based on the computing power reserve.

[0070] S4. Allocate computing resources according to the target model's computing power pairing; when the target inference model is loaded, call the corresponding model instance; when it is not loaded, load the target inference model into the target computing power unit and then call the corresponding model instance to form the inference execution configuration; form the inference processing sequence according to the input task status and the inference execution configuration.

[0071] S5. Process the input data according to the reasoning processing sequence to obtain the reasoning result; form the result output route according to the correspondence between the reasoning result and the output destination, and distribute the reasoning result according to the result output route.

[0072] In this embodiment, multiple types of inputs are used to characterize input data from different input sources and with different data organization methods. Input sources include local file reading interfaces, network stream receiving interfaces, and camera acquisition interfaces. Input data includes single-frame images, image frames corresponding to video files, network video streams, and continuous frames from cameras. Data organization methods include single-frame methods and continuous-frame methods.

[0073] Specifically, the methods and steps are as follows:

[0074] Please see Figure 3 S1. Obtain the task attributes and output destination of the input data to form the input task status; form the task model adaptation range from the model pool based on the input task status.

[0075] Step S1 includes the process of forming the input task state and the process of forming the task model adaptation range.

[0076] The process of forming the input task state specifically includes:

[0077] Input data is obtained from the local file reading interface, network stream receiving interface, or camera acquisition interface; image files are read in the form of single-frame images and image header information, video files are read in the form of container header information and image frames arranged according to timestamps, and network video streams and continuous camera frames are received in the form of data blocks with acquisition timestamps and frame sequence numbers; when the task request arrives at the same time as the input data, the task request carries the task identifier, inference task type, task priority, allowed processing latency, and output destination; when the task request arrives before the input data, the task request is temporarily stored according to the task identifier, and step S1 is started when the first input data with the same task identifier arrives;

[0078] Task attributes include data organization method, input size, input frame rate, inference task type, task priority, and allowed processing latency; the data organization method can be single frame or continuous frame; the input size is in pixels, the input frame rate is in frames per second, and the allowed processing latency is in seconds; the task priority is represented by integer levels from 1 to 5; the output destination is represented by the destination type and the destination address, and the destination type includes local storage, display interface, visualization window, and external business interface;

[0079] Upon arrival of input data, the task identifier, timestamp, image size, pixel encoding, and data length are verified. If the data length is abnormal, it is reread once; if the image size is missing, it is reread from the image header or video container header. If the input frame rate is missing in continuous frame mode, the most recent valid input frame rate of the same input interface is read first. If no valid historical value exists, a frame interval is formed by adjacent valid timestamps. After forming three valid frame intervals, the arithmetic mean is taken, and its reciprocal is used to form a temporary input frame rate. The three frame intervals are used to reduce the impact of single timestamp jitter and meet the timing of short-latency task formation. Normal scheduling of the current input data is terminated if three valid frame intervals cannot be formed before the processing delay expires.

[0080] All timestamps are converted to microseconds in Coordinated Universal Time (UTC) and associated with the task identifier and frame sequence number; frames with timestamps in reverse order within the same task are excluded, and data blocks arriving earlier are retained when frame sequence numbers are duplicated; input width and input height The allowed range is 16 pixels to 8192 pixels, and the allowed range of input frame rate is 1 frame per second to 240 frames per second; when it is below the lower limit, it is corrected to the lower limit, and when it is above the upper limit, the normal scheduling of the current input data is terminated.

[0081] To ensure that different data organization methods and input sizes are included in the same computing power capacity assessment, a processing load is adopted. This represents the amount of inference data relative to the reference input; the width of the reference input. 640 pixels, reference input height 360 pixels, reference frame rate The frame rate is 25 frames per second; this reference input is consistent with the reference input used for model resource requirement coefficient calibration; the equivalent frame rate for continuous frame mode. Equal to the input frame rate, the equivalent frame rate in single-frame mode Equal to reference frame rate; processing load Determined according to the following formula:

[0082] ;

[0083] Handling load The reference input is a dimensionless positive number. When the reference input width, reference input height, or reference frame rate is equal to 0, the reference configuration is reread. If it is still equal to 0, the current task is terminated. When the processing load is not a finite value or is less than or equal to 0, the normal scheduling of the current input data is terminated. The input task status consists of task identifier, data organization method, input size, input frame rate, processing load, inference task type, task priority, allowed processing latency, and output destination. The input task status limits the task model adaptation range in step S1, provides the processing load in step S2, and provides the data organization method, input frame rate, and allowed processing latency in step S4. When the input data verification fails, the current task is terminated and an input verification failure record is written.

[0084] The process of forming the input task state completes the acquisition, verification, and unified representation of task attributes and output destination, resulting in the input task state, which provides input for the formation of the task model's adaptation range, computing power capacity range, and inference processing timing.

[0085] The process of forming the task model's adaptation range specifically includes:

[0086] The model pool reads inference models and model configurations from a version-verified model configuration library. The model configuration includes model identifier, model version, supported inference task types, supported data organization methods, allowed input size range, model accuracy level, benchmark processing time, model computation requirement coefficient, model static storage requirement, model working storage coefficient, and supported computing architecture. The model configuration carries a configuration version and integrity summary. If the configuration version is inconsistent or the summary verification fails, it is read again. If the verification still fails after the reread, the corresponding inference model is excluded.

[0087] When determining the task model fit range, the inference task type, data organization method, input size, allowable processing latency, and the minimum model accuracy level explicitly given in the task request are read from the input task status. First, inference models that do not support the inference task type or data organization method are excluded; then, inference models whose allowable input size range does not cover the current input size are excluded; finally, the baseline processing time is multiplied by the processing load. Once the estimated processing time is obtained, inference models whose estimated processing time exceeds the allowable processing latency are excluded; finally, inference models whose accuracy level is lower than the lowest model accuracy level are excluded.

[0088] The task model adaptation range is represented by a set of inference models that meet all conditions; the inference models are written into the task model adaptation range in ascending order of model identifier; when the task model adaptation range is empty, the inference task type, input size and model configuration version are re-verified. If it is still empty after verification, the current task is terminated and a model mismatch record is formed; when the model version or adaptation conditions change, the associated task model adaptation range becomes invalid, and the next input data is re-executed in step S1.

[0089] The process of forming the task model adaptation range completes the task adaptation, input adaptation, latency adaptation and accuracy adaptation judgment of the inference model, and obtains the task model adaptation range, which is used to limit the inference models participating in the candidate model computing power pairing in step S2.

[0090] Step S1 obtains the input task status and task model adaptation range, providing input for step S2 to form model residency relationship, computing power capacity range and candidate model computing power pairing.

[0091] Please see Figure 4 S2. Obtain the model loading status, instance occupancy status, and computing power reserve of each computing power unit to form a model residency relationship; combine the processing load corresponding to the input task status and the model resource requirements within the task model adaptation range to form a computing power carrying range; pair the inference models within the task model adaptation range with computing power units that meet the requirements of computing power architecture, model version, and running status to form candidate model computing power pairings.

[0092] Step S2 includes the process of forming the model residency relationship, the process of forming the computing power capacity range, and the process of forming the candidate model computing power pairing.

[0093] The formation process of model residency relationships specifically includes:

[0094] The computing power unit reports status data when the service starts, the model is loaded, the model is unloaded, the model instance is invoked, and the model instance is released. During operation, the status data is obtained at a fixed reading period of 20 milliseconds to 100 milliseconds. This embodiment uses 50 milliseconds. The status data includes the computing power unit identifier, computing power architecture, loaded model identifier, model version, total number of instances, number of instances occupied, remaining concurrency, sampling time, and status version. The total number of instances, number of instances occupied, and remaining concurrency together serve as specific data representations of the instance occupancy status.

[0095] The same Coordinated Universal Time (UTC) is used for both state sampling and scheduling. If the difference between the state sampling and scheduling times exceeds three reading cycles, the state data becomes invalid and is reread. If the state versions in two consecutive readings are in reverse order, the newer version is retained. If the number of instances occupied is greater than the total number of instances, the remaining concurrency is less than 0, the computing unit identifier is duplicated, or the loaded model version is inconsistent with the model version in the model pool, the corresponding state data is not used.

[0096] For each inference model within the task model adaptation range, a callable residency relationship is formed between the inference model and a computing power unit that has already loaded the inference model and has a remaining concurrency greater than 0; an occupied residency relationship is formed between the inference model and a computing power unit that has already loaded the inference model but has a remaining concurrency equal to 0; and a deployable relationship is formed between the inference model and a computing power unit that has not loaded the inference model but whose computing power architecture meets the model configuration. The model residency relationship consists of the above three types of relationships and is associated with the model identifier, computing power unit identifier, model loading status, instance occupied status, remaining concurrency, status version, and sampling time.

[0097] Callable residency relationships enable the corresponding candidate model computing power pairing to enter the judgment of already resided executable pairing; occupied residency relationships are re-judged to determine whether to convert to callable residency relationships after the model instance is released; deployable relationships enable the corresponding candidate model computing power pairing to only enter the judgment of non-resided executable pairing; when state data reading fails, the most recent valid result that has not exceeded the state expiration period is used, and the corresponding computing power unit is excluded when there is no valid historical result; model residency relationships are re-formed according to each reading cycle, and the absolute occupied state of the previous reading cycle is not directly carried over to the current reading cycle.

[0098] The process of forming the model residency relationship completes the consistency verification of model loading status, instance occupancy status, version and sampling time, and obtains the model residency relationship, which is used in step S3 to divide the resided executable pairings and the non-resided executable pairings.

[0099] The formation process of computing power carrying capacity specifically includes:

[0100] The computing power reserve is read from the operation monitoring interface of the computing power unit, including normalized available computing power, available storage power, and the number of allowed new instances. The normalized available computing power represents the reference inference computing power that can be provided to new tasks within a monitoring period, and the unit is normalized computing power unit. The available storage power represents the device storage space that can be allocated to model weights, inference context, and intermediate tensors, and the unit is gigabytes. The number of allowed new instances is a non-negative integer. The sampling time, status version, and computing power unit identifier of the computing power reserve should be consistent with the model residency relationship. If they are inconsistent, they should be read again. The normalized available computing power is given by the operation monitoring interface in normalized computing power unit. In this embodiment, the reference input used for processing the load is consistent with the reference input used for calibrating the model resource requirement coefficient. The model computing requirement and the normalized available computing power are both expressed in normalized computing power unit and used for comparison within the same unit. The model computing requirement coefficient is calibrated according to the model version and computing power architecture. The specific calibration process is described in the coefficient determination section below.

[0101] Model computational requirements and model storage requirements together serve as the specific data representation of model resource requirements; for the first [item] in the task model adaptation scope... A reasoning model, model computational requirements Processing load Demand coefficients calculated by the model Multiply to obtain; model calculates demand coefficients. This represents the normalized computational requirement corresponding to a unit of processing load, where the unit is normalized computing power unit, and the model computational requirement. Determined according to the following formula:

[0102] ;

[0103] The model's working storage requirements are determined by the processing load. With model working storage coefficients Multiply to obtain the model working storage coefficients. This represents the working storage requirement corresponding to a unit of processing load, in gigabytes; for the... The first inference model and the first Each computing unit, model storage requirements The storage requirement is determined based on the model loading status; when the inference model is already loaded into the computing unit, only the model's working storage requirement is considered; when the inference model is not loaded into the computing unit, the model's static storage requirement is also considered. And model working storage requirements; model storage requirements Determined according to the following formula:

[0104] ;in This is a numerical representation of the model loading state, set to 1 when loaded and 0 when not loaded. The numerical representation of the model loading state is the formula input for the current version of the model loading state, not a new object distinct from the model loading state. The normalized usable computational cost per computing unit is Available storage capacity is ; Not greater than and Not greater than When the corresponding inference model and computing unit combination is within the computing power carrying capacity, the corresponding combination is outside the computing power carrying capacity when any demand exceeds the corresponding margin, the allowed number of new instances does not meet the instance creation requirements of unloaded models, or the available resources are 0.

[0105] The computing power capacity is represented by a set of models and computing units that meet computing, storage, and instance creation requirements, and does not belong to a numerical range with a continuous start and end point; processing load. Model calculation of demand coefficients Static storage requirements for the model or model working storage coefficient If the model configuration is missing, reread the model configuration; if it is still missing, exclude the corresponding inference model. or If a value is missing, the remaining computing power is reread; if it is still missing, the corresponding computing power unit is excluded. If the model's computational requirements and storage requirements are less than 0 or not finite values, the corresponding combination is excluded. The remaining computing power is verified again before the target model computing power pairing is determined in step S3.

[0106] The process of forming the computing power capacity range involves comparing the processing load, model resource requirements, and computing power margin in the same unit to obtain the computing power capacity range, which is used in step S3 to determine whether the computing power pairing of candidate models is feasible.

[0107] The process of forming candidate model computing power pairings specifically includes:

[0108] Read the inference models and all valid computing power units with valid state data within the task model adaptation range, and enumerate combinations in ascending order of model identifier and computing power unit identifier; combinations that do not meet the requirements of computing power architecture, model version, or running status are not formed into candidate model computing power pairings; if the candidate record space can accommodate all combinations, enumerate them all at once; if it cannot, enumerate them in batches, and the number of each batch is determined by the integer part of the available candidate record space divided by the length of a single candidate model computing power pairing record; if the result is less than 1, stop the current scheduling; during the batch enumeration, step S3 continues to retain the current best executable pairing according to the same rules; where, a computing power unit with valid state data refers to a computing power unit whose state data has not been excluded due to exceeding the state expiration period, abnormal instance data, duplicate computing power unit identifier, or inconsistent loaded model version;

[0109] For each valid combination, the model loading status, instance occupancy status, remaining concurrency, model computation requirements, model storage requirements, computing power reserve, status version, and sampling time are written into the same pairing relationship to form a candidate model computing power pairing. The candidate model computing power pairing does not pre-exclude combinations that exceed the computing power carrying capacity, so that step S3 can distinguish between executable pairings and non-executable pairings based on the same candidate basis. When the inference model in the task model adaptation range cannot form a valid combination with all computing power units, the computing power architecture status and model version are refreshed once. If they are still empty after the refresh, the current task is terminated and a record of no deployable computing power is formed.

[0110] The process of forming candidate model computing power pairings involves the orderly enumeration and relationship solidification of inference models and computing power units within the task model adaptation range, resulting in candidate model computing power pairings. This provides a common processing object for step S3 to divide executable pairings and determine target model computing power pairings.

[0111] Step S2 obtains the model residency relationship, computing power capacity range, and candidate model computing power pairing, providing input for step S3 to determine the target model computing power pairing.

[0112] Please see Figure 5 S3. Based on the model residency relationship and computing power capacity, divide the candidate model computing power pairings into resided executable pairs and non-resided executable pairs; determine the target model computing power pairing from the resided executable pairs based on the computing power reserve; if there is no resided executable pairing, determine the target model computing power pairing from the non-resided executable pairs based on the computing power reserve.

[0113] Step S3 includes the process of determining the target model's computing power pairing.

[0114] The process of determining the target model's computing power pairing specifically includes:

[0115] Read the model loading status, instance occupancy status, remaining concurrency, model computation requirements, model storage requirements, and computing power reserve in the candidate model computing power pairings; pairings with callable residency relationships and within the computing power capacity range are assigned to resided executable pairings; pairings with deployable relationships and within the computing power capacity range are assigned to non-resided executable pairings; pairings with occupancy residency relationships, expired status versions, resource status conflicts, or located outside the computing power capacity range are not included in the executable pairings;

[0116] To compare computing power margins within the same type of executable pairings, for the first The first inference model and the first Computing capacity ratio of each computing unit ; The function represents taking the minimum value from the values ​​given in parentheses; normalization can be computed with a computational cost of... Available storage capacity is Balance ratio Take the smaller value between the calculated margin ratio and the storage margin ratio, and determine it according to the following formula:

[0117] ;of which the surplus ratio It is a dimensionless numerical value, ranging from 0 to 1; or When the result is 0, division is not performed; the remainder is proportional. The value is set to 0 and the corresponding pairing is excluded. , , or If the data is missing, the corresponding pairing is reread; if it is still missing, the pairing is excluded. If the result is less than 0, it indicates that the resource status has changed after the candidate model computing power pairing is formed, and the corresponding pairing is removed from the executable pairing. If the result is greater than 1, it is corrected to 1.

[0118] When an executable pair is already resident and not empty, it is executed according to the remaining proportion. The pairs are sorted in descending order, and the target model computing power pair is determined only among the pairs of resident executables. When the remaining ratios are the same, the order is determined by descending remaining concurrency, descending state sampling time, and ascending computing power unit identifier. When there are no resident executable pairs, the non-resident executable pairs are processed according to the same remaining ratio and remaining concurrency order. When the remaining ratios of non-resident executable pairs are the same, they are sorted first by the expected model loading time in ascending order, and then by the computing power unit identifier in ascending order. The first pair after sorting is determined as the target model computing power pair.

[0119] When all executable pairs are empty, the current input data is retained in the waiting cache; the retry interval is twice the state reading cycle, which is 100 milliseconds in this embodiment, so that at least one set of new state data is obtained between adjacent retries; the maximum number of retries is determined by the integer part of the remaining allowed processing delay divided by the retry interval, and is subject to the configuration upper limit; no retries are made when the remaining allowed processing delay is less than one retry interval; steps S2 and S3 are re-executed for each retry; if no executable pair is found after the maximum number of retries is reached, the current task is terminated and a scheduling timeout record is written; before entering step S4, the state version and sampling time of the target model computing power pair are checked again, and the next pair is selected if the check fails.

[0120] The process of determining the target model computing power pairing completes the priority of resident executable pairings, sorting of computing power reserves within the same type of pairing, and status version verification, to obtain the target model computing power pairing that simultaneously limits the target inference model and the target computing power unit. This pairing is used in step S4 to allocate computing power resources, call model instances, and form inference execution configuration.

[0121] Step S3 obtains the target model computing power pairing, providing input for step S4 to form the inference execution configuration and inference processing timing.

[0122] Please see Figure 6 S4. Allocate computing resources according to the target model computing power pairing; when the target inference model has been loaded, call the corresponding model instance; when it has not been loaded, load the target inference model into the target computing power unit and call the corresponding model instance to form the inference execution configuration; form the inference processing sequence according to the input task status and the inference execution configuration.

[0123] Step S4 includes the formation process of the inference execution configuration and the formation process of the inference processing timing.

[0124] The process of forming the inference execution configuration specifically includes:

[0125] Read the target inference model, target computing unit, model computing requirements, model storage requirements, model loading status, instance occupancy status, and status version from the target model computing power pairing; computing resources and storage resources are the specific resource types of computing power resources; lock the resource update record of the target computing unit and reread the status version; when the status version is consistent with the status version in the target model computing power pairing, perform atomic reservation for computing resources and storage resources. Atomic reservation means that the computing power reserve update is only submitted after both types of resources are successfully reserved simultaneously; if any resource reservation fails, cancel the completed reservation, exclude the pairing from the current candidate model computing power pairing, and return to step S3;

[0126] When the target inference model has been loaded into the target computing unit, the model version, instance occupancy status, and remaining concurrency are verified. If the model version is consistent and the remaining concurrency is greater than 0, the corresponding model instance is invoked. When the target inference model has not been loaded into the target computing unit, the model file and model configuration are read from the model pool, and the model file length, integrity summary, and runtime environment version are verified. After the verification is passed, the model weight loading, inference context establishment, model instance creation, and instance invocation are completed in sequence.

[0127] The estimated model loading time is read from valid historical loading records of the same model version and computing power architecture; if no valid historical records are found, the upper limit of loading time in the model configuration is used; the model loading timeout threshold is between 50 milliseconds and 2000 milliseconds, and in this embodiment, the threshold corresponding to the target inference model M3 is 120 milliseconds; if the model file reading fails, the integrity verification fails, the runtime environment version does not match, or the loading time exceeds the model loading timeout threshold, the model content that has been written but not initialized is unloaded, the reserved resources are released, the current target model computing power pair is excluded from the current task, and the process returns to step S3; if all executable pairings are excluded, the current task is terminated and a model loading failure record is generated.

[0128] The inference execution configuration includes the target inference model, target computing power unit, model instance identifier, computing resource quota, storage resource quota, input data transmission method, inference batch, and number of execution streams. The inference batch cannot exceed the remaining concurrency of the model instance. The number of execution streams is determined based on the three overlapping stages of preprocessing, model inference, and postprocessing, and does not exceed the allowable value of the target computing power unit. When one execution stream is allowed, they are executed sequentially; when two execution streams are allowed, adjacent stages are merged; when more than three execution streams are allowed, one execution stream is configured for each stage. The inference execution configuration is locked until the task ends.

[0129] The process of forming the inference execution configuration completes the atomic reservation of computing resources, model instance invocation, and model loading branch processing to obtain the inference execution configuration, which provides input for the formation of the inference processing sequence.

[0130] The formation process of the reasoning processing sequence specifically includes:

[0131] The system reads the data organization method, input frame rate, task priority, and allowed processing latency from the input task status, and reads the model instance identifier, input data transmission method, inference batch, and execution stream number from the inference execution configuration. In single-frame mode, the system forms a single-frame inference processing sequence in the order of input validation, preprocessing, data transmission, model inference, postprocessing, and result encapsulation. The system starts the next action after the previous action is completed and passes data validity validation. If any action fails, the system stops the subsequent actions of the current single frame and releases the single-frame temporary storage space.

[0132] In the continuous frame mode, a circular frame buffer is established, and a pipelined processing relationship is formed in the order of frame reading, preprocessing, data transmission, model inference and postprocessing; The function represents taking the minimum value from the values ​​within the parentheses. The function represents taking the maximum value from the values ​​within the parentheses. Indicates will Round up; Circular frame buffer length Based on input frame rate Allowed processing delay and cache balance Sure, This is the lower limit for the number of frames to be cached. The maximum number of frames to be buffered is determined by the following formula:

[0133] ;in Indicates the theoretical number of frames arriving within the allowed processing latency; rounding up preserves the position of complete frames; buffer remaining. This indicates the maximum number of frames in transit that have left the frame reading stage but have not yet completed processing. In this embodiment, preprocessing, model inference, and postprocessing overlap and advance, occupying a maximum of two additional frames in transit. Set to 2; when the number of execution flows or pipeline overlap changes, redetermine based on the actual maximum number of frames in transit. Lower limit of cached frames The minimum number of frames required to start the pipeline is determined based on the number of frames required; in this embodiment, it is set to 4. The maximum number of buffered frames is also considered. The value is determined by dividing the allocable buffer space by the single-frame buffer occupancy, and in this embodiment, it is set to 32. If the input frame rate or allowed processing latency is missing, return to step S1 for re-verification.

[0134] Frames are written to the circular frame buffer in ascending order of frame number. When the number of execution flows is greater than 1, preprocessing is performed on the current frame, model inference is performed on the previous frame, and postprocessing is performed on the frame before that. The asynchronous data transmission completion event serves as a prerequisite for starting the model inference action, and the model inference completion event serves as a prerequisite for starting the postprocessing action. When the task priority is 4 or 5 and the circular frame buffer is full, frame reading is paused until an empty position appears. When the task priority is 1 to 3 and the circular frame buffer is full, the earliest image frame that has not yet entered model inference is deleted, and the sequence number of the deleted frame is written into the missing frame range of the inference result. When the timestamp is reversed, the frame sequence number is duplicated, or the data transmission times out, the corresponding frame is excluded, and subsequent valid frames continue to be processed according to the original inference processing sequence.

[0135] The process of forming the inference processing timing sequence completes the action arrangement and start condition limitation of single-frame sequential processing or continuous frame pipelined processing to obtain the inference processing timing sequence, which is used to process the input data in step S5.

[0136] Step S4 obtains the inference execution configuration and inference processing timing, providing input for step S5 to process input data and form the result output route.

[0137] Please see Figure 7 S5. Process the input data according to the reasoning processing sequence to obtain the reasoning result; form the result output route according to the correspondence between the reasoning result and the output destination, and distribute the reasoning result according to the result output route.

[0138] Step S5 includes the process of forming the output route.

[0139] The specific process of generating the output route includes:

[0140] The model instances in the inference execution configuration are invoked according to the inference processing sequence; the inference results of the target recognition task include target location, target category, and confidence level; the inference results of the target tracking task include target location, target identifier, and trajectory points; the inference results of the image classification task include classification category and classification confidence level; each inference result is associated with the task identifier, frame number, acquisition timestamp, inference completion time, target inference model identifier, target computing unit identifier, and missing frame range;

[0141] When the output destination is local storage, the structured inference results are written to a file or database; when the output destination is a display interface or visualization window, the target location, trajectory points, or classification categories are overlaid onto the corresponding image frames before output; when the output destination is an external business interface, the task identifier, frame sequence number, collection timestamp, and structured inference results are encapsulated according to the interface version before being sent; the output destination address and interface version are verified before forming the distribution path, and no corresponding distribution path is formed if the versions do not match.

[0142] The output route consists of the correspondence between the inference result type, task identifier, output destination type, output destination address, and interface version; when multiple output destinations are configured, they form distribution paths and all refer to the same inference result; when address verification fails, the corresponding path is excluded; when all output destinations are invalid, the local temporary storage area is used as the alternative output destination.

[0143] When distribution fails, it will retry within the remaining time allowed for processing delay, with the number of retries ranging from 0 to 5, and 2 in this embodiment; the unavailability of a single output destination does not affect other output destinations; after the external business interface times out, the unsent results will be saved to the bounded retransmission cache; the capacity limit of the bounded retransmission cache is 100 to 10,000 results, and 1,000 in this embodiment; when the capacity limit is reached, the result with the earliest collection timestamp will be deleted and the deletion range will be recorded; after the inference execution is completed, the task is canceled, or abnormally terminated, the computing resource quota and storage resource quota in the inference execution configuration will be released, the instance occupancy of the corresponding model instance will be reduced by 1, and the computing power reserve will be updated.

[0144] The process of forming the output route involves matching the version of the inference result with the output destination, forming the path, and handling exception rollbacks to obtain the output route. The inference result is then distributed according to the output route.

[0145] Step S5 obtains the inference results and the result output route, completing the continuous processing of input data from task parsing, model adaptation, joint scheduling of model and computing unit, inference execution to distribution of inference results.

[0146] In this embodiment, regarding the determination of coefficients, the model calculates the required coefficients. and model working storage coefficient The model is calibrated according to its version and computing architecture. Each model version uses no fewer than 50 samples covering the allowed input size range and data organization method, and records the normalized computational and working storage increments. The increment of each sample is first divided by the corresponding processing load. Then, sort the obtained unit processing load increments in ascending order of their values. and Take the 95th percentile value of the corresponding sequence respectively; the two coefficients do not need to be normalized again; when the sample is insufficient, use the coefficients that have passed the verification of the previous model version; when there are no historical coefficients, do not write the corresponding inference model into the model pool; recalibrate when the model version, operating environment or computing power architecture changes.

[0147] In this embodiment, regarding threshold determination, the state reading period is set to 20 milliseconds to 100 milliseconds, and the state expiration period is set to three state reading periods; the task priority is set to 1 to 5, with 4 and 5 classified as high priority; the retry interval is set to twice the state reading period, and the maximum number of retries is jointly limited by the remaining allowed processing latency and the retry interval; the model loading timeout threshold is determined by adding one state reading period to the 99th percentile value of at least 30 valid loading times; and the lower limit for the number of cached frames is set. The maximum number of cached frames is determined based on the requirements for pipeline startup. The remaining cache space is determined based on the allocatable cache space. The maximum number of frames in transit is determined; when there are insufficient samples, the upper limit of loading time in the model configuration is used; when the threshold exceeds the limit, it is corrected to the allowable boundary, and if the condition is still not met after correction, the current processing is stopped.

[0148] To further illustrate the working process of this embodiment, a specific operating procedure is given below:

[0149] The target tracking task is performed using a continuous video stream of 1920 pixels by 1080 pixels at 25 frames per second; the task priority is 3, allowing for processing latency. The input time is 0.20 seconds; the output destination includes a high-definition display interface and a local results database; the reference input width is... 640 pixels, reference input height 360 pixels, reference frame rate The equivalent frame rate in continuous frame mode is 25 frames per second. It is 25 frames per second.

[0150] In step S1, Equal to 1920 Equal to 1080 Equals 25. Equals 640 Equal to 360 and Substituting 25 into the load handling formula, we get... Equals 9; Input task status includes continuous frame mode, input size 1920 pixels by 1080 pixels, input frame rate 25 frames per second, processing load 9, target tracking task, task priority 3, allowed processing delay 0.20 seconds, and two output destinations.

[0151] In this embodiment, the two inference models associated with the target tracking task in the model pool are respectively referred to as the first inference model. and the third reasoning model ; Model for calculating demand coefficient The static storage requirement for the model is 0.75. The model's working storage factor is 0.70 gigabytes. The size is 0.05 gigabytes, and the model accuracy level is 0.89; Model for calculating demand coefficient The static storage requirement for the model is 1.15. The model's working storage factor is 1.20 gigabytes. The size is 0.08 gigabytes, and the model accuracy level is 0.93; both inference models support continuous frame target tracking and current input size, and the task model's adaptability range is... and .

[0152] In step S2, the first computing unit Normalization available computational cost The available storage is 7.00. 2.00 gigabytes, already loaded And there exists a callable model instance; second computing unit Normalization available computational cost The available storage is 12.00. It is 4.50 gigabytes and has already been loaded. And there exists a callable model instance; the third computing unit Normalization available computational cost The available storage is 16.00. 6.00 gigabytes, not loaded and All three computing units support and The operating environment.

[0153] Model computation requirements It is 9 times 0.75, which equals 6.75 normalized computing power units; Model computation requirements It is 9 times 1.15, which equals 10.35 normalized computing power units; exist The above has already been loaded; model storage requirements. That's 9 times 0.05, which equals 0.45 gigabytes. exist and The above is not loaded; model storage requirements. Adding 0.45 to 0.70 equals 1.15 gigabytes; exist The above has already been loaded; model storage requirements. That's 9 times 0.08, which equals 0.72 gigabytes. exist and The above is not loaded; model storage requirements. Adding 0.72 to 1.20 equals 1.92 gigabytes.

[0154] Based on computing and storage requirements, and , and , and , and as well as and Located within the computing power capacity range; and The model's computational requirement is greater than 10.35. The normalized available computational cost is 7.00, which is outside the computing power capacity range; the computing power pairing of candidate models is determined by... and , and , and , and , and and and It consists of six architecture-compatible pairs.

[0155] In step S3, and Those with a callable residency relationship and located within the computing power capacity range are assigned to the already resided executable pair; and Those with a callable residency relationship and located within the computing power capacity range are assigned to the already resided executable pair; and , and as well as and Those with deployable relationships and located within the computing power capacity range are classified as non-resident executable pairs; and Located outside the computing power capacity range, it will not enter the executable pairing.

[0156] and Balance ratio Taking the smaller value between the calculated margin ratio of 0.0357 and the storage margin ratio of 0.7750, the result is 0.0357; and Balance ratio Taking the smaller value between the computational margin ratio of 0.1375 and the storage margin ratio of 0.8400, the result is 0.1375; therefore, in the resident executable pairing, and The target model's computing power was determined to be matched. and and and Although they have larger margin ratios of 0.5781 and 0.3531 respectively, they belong to non-resident executable pairings and do not participate in the determination of target model computing power pairing when the resident executable pairings are not empty.

[0157] In step S4, for and The atomic algorithm reserves 10.35 normalized computing units and 0.72 gigabytes of working storage space. Already loaded in Directly call the model instance M3-C2-01; Four execution flows are allowed, and there are currently three overlapping processing stages, so the number of execution flows is 3; the inference execution configuration also includes asynchronous data transfer mode and inference batch 1.

[0158] The theoretical number of frames allowed within the processing latency is 25 × 0.20 = 5 frames, and the maximum number of frames in transit is 2. Therefore Take 2; equal Function pair 32 and The function result takes the smaller value; The function result is 4 and The larger value in the value is 7, so a seven-frame circular frame buffer is used; the inference processing timing is such that the current frame performs preprocessing, the previous frame performs model inference, and the frame before that performs postprocessing.

[0159] In step S5, The system outputs the target location, target identifier, and trajectory points for each frame. The output routing distributes image frames with target bounding boxes and trajectories to the high-definition display interface, writing the task identifier, frame number, acquisition timestamp, target identifier, target location, and trajectory points to the local results database. Both distribution paths reference the same inference result and are not executed repeatedly. After the task is completed, 10.35 normalized computing units and 0.72 gigabytes of working storage space will be released, and the instance occupancy of model instance M3-C2-01 will be reduced by 1.

[0160] In the missing input branch, when the network video stream does not carry the input frame rate and there is no valid historical value, three valid frame intervals are formed based on the current frame timestamp; in this embodiment, the three frame intervals are all 0.040 seconds, the average frame interval is 0.040 seconds, and the temporary input frame rate is 25 frames per second; if three valid frame intervals cannot be formed before the allowable processing delay of 0.20 seconds expires, the current task is terminated.

[0161] In the empty result and in the no-solution branch, the task model adaptation range is empty. The model configuration version is re-verified. If it is still empty, a model incompatibility record is formed. If both executable pairing sets are empty, they are retried at 100-millisecond intervals. In this embodiment, the remaining allowable processing delay is 0.20 seconds. Therefore, it can be retried a maximum of 2 times. If it is still empty, a scheduling timeout record is formed.

[0162] In the model loading failure branch, the resident executable pair is empty and will... and When the target model's computing power is determined to be matched, if If the model file integrity digest verification fails, it will be released. The reserved computing resources will and Exclude from the current task and return to step S3 to re-determine from the remaining non-resident executable pairs; if all remaining pairs are unavailable, a model loading failure record is generated.

[0163] In the abnormal output branch, if the high-definition display interface address verification fails, the display distribution path is excluded, and the structured inference result is still written to the local result database; if both the local result database and the high-definition display interface are unavailable, the inference result is saved to the local temporary storage area; when the bounded retransmission cache reaches 1000 results, the result with the earliest collection timestamp is deleted, and the range of the deleted frame sequence number is written to the deletion record.

[0164] As can be seen from the above description, the intelligent reasoning task scheduling method for multiple types of input provided in this embodiment has the following technical effects:

[0165] First, the task model adaptation range is formed based on the input task status. Then, the candidate model computing power pairing is formed by combining the model residency relationship and computing power capacity range. The target model computing power pairing is determined by dividing the resident executable pairing and the non-resident executable pairing, so that the adaptation result of the inference model can be constrained by the model loading status, instance occupancy status, computing power reserve and model resource requirements at the same time.

[0166] Based on this, atomic reservations are made for the computing and storage resources of the target computing unit. Furthermore, the inference execution configuration and processing sequence are determined according to the model loading status of the target inference model. Finally, the inference results are distributed according to the output routing. This allows the task model adaptation range to gradually converge to the target model computing power pairing that meets the actual execution conditions. It also ensures a continuous correspondence between the target model computing power pairing and subsequent computing resource allocation, model instance invocation, and inference execution. This reduces task waiting or rescheduling caused by changes in computing resource status or the target inference model not being loaded, thereby improving the processing efficiency and execution continuity of intelligent inference tasks.

[0167] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for scheduling intelligent reasoning tasks oriented towards multiple types of input, characterized in that, The methods and steps include the following: S1. Obtain the task attributes and output destination of the input data to form the input task status; and form the task model adaptation range from the model pool based on the input task status. S2. Obtain the model loading status, instance occupancy status, and computing power reserve of each computing power unit to form a model residency relationship; combine the processing load corresponding to the input task status and the model resource requirements within the task model adaptation range to form a computing power carrying range; pair the inference models within the task model adaptation range with computing power units that meet the requirements of computing power architecture, model version, and running status to form candidate model computing power pairings. S3. Based on the model residency relationship and the computing power carrying capacity, the candidate model computing power pairings are divided into resided executable pairings and non-resided executable pairings; the target model computing power pairing is determined from the resided executable pairings based on the computing power reserve; if there is no resided executable pairing, the target model computing power pairing is determined from the non-resided executable pairings based on the computing power reserve. S4. Allocate computing resources according to the target model's computing power pairing; When the target inference model is loaded, the corresponding model instance is invoked; when it is not loaded, the target inference model is loaded into the target computing unit and the corresponding model instance is invoked to form an inference execution configuration; the inference processing sequence is formed based on the input task status and the inference execution configuration. S5. Process the input data according to the inference processing sequence to obtain the inference result; form a result output route according to the correspondence between the inference result and the output destination, and distribute the inference result according to the result output route.

2. The intelligent reasoning task scheduling method for multiple input types according to claim 1, characterized in that, The process of forming the input task state specifically includes: The task attributes of the input data and the output destination are obtained. The task attributes include the data organization method, input size, input frame rate, inference task type, task priority and allowed processing latency. The task identifier, timestamp, image size, pixel encoding, and data length of the input data are verified. The processing load is determined based on the input size, equivalent frame rate, and reference input, wherein the reference input includes reference input width, reference input height, and reference frame rate. The equivalent frame rate in continuous frame mode is equal to the input frame rate, and the equivalent frame rate in single frame mode is equal to the reference frame rate. The processing load is then determined according to the following formula: ;in, To handle the load, and These are the input width and input height, respectively. and These are the reference input width and reference input height, respectively. For equivalent frame rate, The reference frame rate is used as the reference frame rate; and the input task state is composed of task identifier, data organization method, input size, input frame rate, processing load, inference task type, task priority, allowed processing latency, and output destination.

3. The intelligent reasoning task scheduling method for multiple input types according to claim 1, characterized in that, The process of forming the task model adaptation range specifically includes: Obtain inference models and model configurations from the model pool; Based on the inference task type, data organization method, input size, allowable processing latency, and minimum model accuracy level in the input task status, an adaptation judgment is made on the inference model. The expected processing time is determined based on the baseline processing time and the processing load, and the set of inference models that meet the requirements of inference task type, data organization method, input size, expected processing time and model accuracy level is determined as the task model adaptation range.

4. The intelligent reasoning task scheduling method for multiple input types according to claim 1, characterized in that, The formation process of the model residency relationship specifically includes: Obtain the model loading status, instance occupancy status, computing architecture, model version, remaining concurrency, sampling time, and status version for each computing unit; Establish a callable residency relationship between the inference model and the computing power unit that has loaded the inference model and has remaining concurrency; establish an occupied residency relationship between the inference model and the computing power unit that has loaded the inference model and has no remaining concurrency; establish a deployable relationship between the inference model and the computing power unit that has not loaded the inference model but whose computing power architecture meets the model configuration. The model residency relationship is composed of the callable residency relationship, the occupied residency relationship, and the deployable relationship.

5. The intelligent reasoning task scheduling method for multiple types of input as described in claim 1, characterized in that, The process of forming the computing power carrying capacity specifically includes: The model resource requirements include model computation requirements and model storage requirements; The model computation requirements are determined based on the processing load and model computation requirement coefficient, and the model storage requirements are determined based on the processing load, model working storage coefficient, model static storage requirements, and model loading status. Obtain the normalized available computing power, available storage power, and the number of new instances allowed for each computing power unit. Include combinations of inference models and computing power units whose model computing requirements do not exceed the normalized available computing power, whose model storage requirements do not exceed the available storage power, and which meet the instance creation requirements, into the computing power carrying capacity range.

6. The intelligent reasoning task scheduling method for multiple types of input as described in claim 5, characterized in that, The process of forming the candidate model computing power pairing specifically includes: Read the valid computing power units in the inference model and state data within the task model adaptation range, enumerate the combinations in ascending order of model identifier and computing power unit identifier, and ensure that combinations that do not meet the requirements of computing power architecture, model version or running status do not form candidate model computing power pairings. The effective combination of model loading status, instance occupancy status, remaining concurrency, model computation requirements, model storage requirements, computing power reserve, status version, and sampling time is written into the same pairing relationship to form the candidate model computing power pairing.

7. The intelligent reasoning task scheduling method for multiple types of input as described in claim 4, characterized in that, The process of determining the target model computing power pairing specifically includes: Candidate model computing power pairs that have the callable residency relationship and are within the computing power capacity range are assigned to the already resided executable pairs; candidate model computing power pairs that have the deployable relationship and are within the computing power capacity range are assigned to the non-resided executable pairs. For candidate model computing power pairing, normalized available computing power, available storage power, model computing requirements, and model storage requirements are obtained. The ratio of the normalized available computing power after deducting the model computing requirements to the normalized available computing power, and the ratio of the available storage power after deducting the model storage requirements to the available storage power, are taken as the smaller of the two as the margin ratio. When the resident executable pairing is not empty, the target model computing power pairing is determined from the resident executable pairings in descending order according to the remaining margin ratio; when the resident executable pairing is empty, the target model computing power pairing is determined from the non-resident executable pairings in descending order according to the remaining margin ratio.

8. The intelligent reasoning task scheduling method for multiple types of input as described in claim 7, characterized in that, The process of forming the inference execution configuration specifically includes: Based on the target model computing power pairing, atomic reservation is performed on the computing resources and storage resources of the target computing power unit, and the computing power margin update is submitted after the computing resources and storage resources are successfully reserved at the same time. When the target inference model has been loaded into the target computing unit, the model version, instance occupancy status, and remaining concurrency are verified before the corresponding model instance is invoked; when the target inference model has not been loaded into the target computing unit, the model file and model configuration are verified, and the model weight loading, inference context establishment, model instance creation, and instance invocation are completed in sequence. The inference execution configuration includes the target inference model, target computing power unit, model instance identifier, computing resource quota, storage resource quota, input data transmission method, inference batch, and number of execution streams.

9. The intelligent reasoning task scheduling method for multiple types of input as described in claim 1, characterized in that, The formation process of the inference processing timing specifically includes: Read the data organization method, input frame rate, and allowed processing latency from the input task status, and read the number of execution streams from the inference execution configuration; In single-frame mode, the single-frame inference processing sequence is formed in the order of input verification, preprocessing, data transmission, model inference, postprocessing, and result encapsulation. In continuous frame mode, a circular frame buffer is established, and a pipelined processing relationship is formed in the order of frame reading, preprocessing, data transmission, model inference and postprocessing. The length of the circular frame buffer is determined based on the input frame rate, allowable processing latency and buffer capacity, and the preprocessing, model inference and postprocessing are overlapped and advanced according to the number of execution streams.

10. The intelligent reasoning task scheduling method for multiple types of input as described in claim 9, characterized in that, The process of forming the output route specifically includes: The input data is processed by calling the model instance in the inference execution configuration according to the inference processing sequence to obtain the inference result; The result output route is formed based on the correspondence between the inference result type, task identifier, output destination type, output destination address and interface version. When multiple output destinations are configured, distribution paths are formed respectively and they all refer to the same inference result. The inference results are distributed according to the output route. When the address verification fails, the corresponding distribution path is excluded. When all output destinations are invalid, the local temporary storage area is used as the alternative output destination.